Method for Rapid Identification of Total Amount of Air Ticket Itinerary Receipt
By preprocessing the aviation itinerary pictures and identifying the deep neural network model, the problem of poor results in the existing technology in identifying the total amount of the aviation itinerary is solved, and the rapid and accurate identification effect is achieved, reducing the workload of manual verification.
Patent Information
- Application Number
- CN202011396350.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-12-03
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2040-12-03
AI Technical Summary
It is difficult for the prior art to effectively identify the total amount information in the air travel order, especially in the case of background fine lines and complex printed characters, the traditional OCR method and OpenCV adaptive threshold method are not effective in scenarios where impurity pixels are scattered and have lower resolution.
By pre-processing the air travel single pictures, including cropping, rotating and removing fine lines on blue backgrounds, and then boxing and word splitting of the total amount numbers, a single character picture is identified using the deep neural network model and merged into the total amount recognition result.
It realizes the rapid identification of the total amount in the air travel list pictures of different scanning qualities. The average single itinerary is about 1s, and the recognition accuracy rate reaches more than 96%, reducing the manual verification workload of financial personnel.
Smart Images

Figure CN114663900B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a technology in the field of image processing, specifically a method for quickly identifying the total amount of an air ticket itinerary receipt. Background Art
[0002] Due to anti-counterfeiting and other reasons, the air ticket itinerary receipt itself has added many blue curved lines on the ticket background. These lines are dense, with varying shapes, and the thickness is similar to the characters of the valid information in the air ticket itinerary receipt. In addition, some tickets fade, change color, or there are significant differences in the imaging effects of different scanning / camera devices. Even the color of the printing ink of the valid information is similar to the color of the interference curves. The above situations cause great interference to the character recognition work of the air ticket itinerary receipt, so the traditional OCR method cannot be used to identify the total amount of the air ticket itinerary receipt.
[0003] Compared with the recognition of general tickets, there are both background fine lines and printed characters in the itinerary receipt, and the printing density or fading degree of the two is different in different batches. Although there is an adaptive threshold method adaptiveThreshold in the existing OpenCV technology, the two methods for calculating the neighborhood given by OpenCV are based on the Gaussian mean or average value around the pixel to be determined. The binary threshold obtained in the scenario where the impurity pixels are scattered will make the text blurred, and it will greatly affect the subsequent recognition effect in the scenario of low-resolution pictures. Summary of the Invention
[0004] Aiming at the above deficiencies of the existing binary threshold method, the present invention proposes a method for quickly identifying the total amount of an air ticket itinerary receipt, which automatically identifies the total amount of the air ticket itinerary receipt uploaded in the form of a picture in the enterprise reimbursement system for automatic verification with the filled information. This greatly reduces the workload of manual verification by financial personnel.
[0005] The present invention is realized through the following technical solutions:
[0006] The present invention relates to a method for quickly identifying the total amount of an air ticket itinerary receipt. By preprocessing the air ticket itinerary receipt picture, selecting the total amount digital box, and splitting single characters, each character picture in the total amount information is extracted from the itinerary receipt picture, and then a deep neural network model is used to identify the single character picture into character type data and merge them into the total amount recognition result.
[0007] The so-called picture preprocessing refers to: after unifying the picture material styles of different scanning batches through operations such as cropping, the general area where the total amount information is located is initially determined, specifically including:
[0008] Step 1: Read the picture, correct the picture through rotation operation, crop the non-itinerary part (such as the blank part), and compress the picture with too large dimensions;
[0009] The pictures with too large dimensions are: pictures with a horizontal resolution exceeding 1600 pixels.
[0010] Step 2: According to the position range of the key information to be recognized, roughly determine the horizontal and vertical coordinate intervals of the picture where the key information is located, and crop the original picture into the region of interest (RoI);
[0011] The specific operable range is: the vertical interval between the 4th and 6th horizontal lines of the air ticket itinerary receipt, and the horizontal interval of the right half of the air ticket itinerary receipt.
[0012] Step 3: Filter the blue fine lines of the background according to the color channel, that is, for the pixels with HSV values within the set range, set their HSV values to [255, 255, 255].
[0013] The box selection of the total amount number refers to: delimiting the text image representing the total amount with a rectangle of a specific color for use in the single-character segmentation process, specifically including:
[0014] Step i: Convert the color image to a grayscale image, and then perform a binarization operation on the grayscale image;
[0015] Step ii: Use the opening operation to remove impurity pixels, that is, discrete pixels that do not represent the total amount characters and have a pixel value of 255;
[0016] Step iii: After fusing the discrete pixels with a grayscale value of 255 representing the total amount through the closing operation, remove the protruding pixels at the upper and lower boundaries of the graph by judging the horizontal width;
[0017] Step iv: Screen out the positive image with the largest area from the image after removing the protruding pixels, and draw the positive circumscribed rectangle of the positive image to obtain the box selection area of the total amount information.
[0018] The opening operation refers to performing one erosion operation and then one dilation operation;
[0019] The closing operation refers to performing one dilation operation and then one erosion operation.
[0020] The erosion operation means that structure A is eroded by structure B, then
[0021] The dilation operation means that structure A is dilated by structure B, then
[0022] The single-character splitting mentioned above refers to: dividing the text area of the key information into multiple sub-areas, such that each sub-area contains and only contains one character, specifically including:
[0023] Step ①: After obtaining the boxed image, perform dynamic binarization processing and remove impurity pixels through opening operation;
[0024] For the dynamic binarization processing mentioned above, since the font of the same type of information on the same type of bill is determined, the parameter value min_black_rate that does not depend on the quality of the image to be recognized is determined, and the low-gray invalid pixels are removed on the premise of ensuring the least loss of effective pixels. Specifically:
[0025] i) Observe the thickness of the characters, and determine the minimum effective pixel rate min_black_rate through experiments, where: the min_black_rate of bold fonts is greater than that of non-bold fonts.
[0026] The observation mentioned above includes, but is not limited to, that the minimum effective pixel rate of bold characters is greater than that of thin characters.
[0027] The determination through experiments mentioned above means: by adjusting the parameter min_black_rate, observing the binarization effect of most sample images, achieving the effect of removing scattered invalid pixels to the greatest extent on the premise of ensuring that effective pixels are retained, so as to determine the min_black_rate suitable for the current font.
[0028] ii) Calculate the minimum effective pixel number target_black_count = w × h × min_black_rate, where: w and h are the length and width pixel values of the rectangular frame image respectively;
[0029] iii) Initialize the threshold to be determined target_v to 0 and set the current cumulative effective pixel curr_black_count to 0;
[0030] iv) curr_v starts to loop from 0 to 256:
[0031] a) Statistically calculate the number of pixels with a gray value equal to curr_v in the image through matrix operations and accumulate it to curr_black_count;
[0032] b) When curr_black_count is greater than or equal to target_black_count, end the loop and set target_v to curr_v, otherwise continue the loop;
[0033] v) Then, binarize the image based on the obtained target_v, that is, set the gray value greater than target_v as valid pixels (255), and set the gray value less than or equal to target_v as invalid pixels (0).
[0034] Step ②: Remove horizontal lines by traversing horizontally and judging the length of consecutive pixels.
[0035] Step ③: Traverse from left to right, perform threshold judgment based on the number of black pixels in each column of the image, and divide the rectangular image into multiple single-character regions, that is, an image that contains and only contains a single character after division. Specifically: For the currently traversed column, count the number of black pixels from top to bottom. When the number of black pixels is greater than the threshold, that is, the current column is not the column where the valid character is located, skip this column; when the number of black pixels is less than the threshold, that is, the current column belongs to a valid character, record the starting abscissa, and then continue to judge the next column until encountering the next column where the number of black pixels is greater than the threshold, that is, the valid character ends, and record the ending abscissa. Record the starting abscissa and ending abscissa of this step as a single-character region; add the current single-character region to the character list.
[0036] Step ④: Perform a width recheck on each single-character region in the character list, that is, determine whether a certain single-character region needs to be split or merged by setting a pixel width threshold. Specifically: When the width of the single-character region is too large to cause digital adhesion, return to Step ③ for re-segmentation; when the width of the single-character region is too small or the horizontal coordinate distance between two regions is too small to cause a single character to be split, merge the adjacent split parts into a single-character region.
[0037] The judgment criteria for being too large or too small are: whether it is greater than the preset maximum character width threshold or less than the preset minimum character width threshold.
[0038] For the recognition, after the single-character binary image returned by the single-character segmentation process is recognized into character type data through a trained neural network model, each digital image is separately recognized into a character result and then merged into a string result, that is, the total amount result. This training refers to supervised learning based on the MNIST dataset to obtain a machine learning model that can recognize a single character image.
[0039] Preferably, the present invention verifies the accuracy of the method by comparing the differences between the manually marked results and the recognition results.
[0040] The present invention relates to a system for implementing the above method, including: a picture preprocessing unit, a total amount detection and framing unit, a single-character splitting unit, and a single-character recognition unit, where: the picture preprocessing unit transmits the picture calibrated and with image interference removed to the total amount detection and framing unit, the total amount detection and framing unit transmits the picture with the total amount text framed by a specific color rectangle to the single-character splitting unit, the single-character splitting unit splits the image in the total amount text rectangle into multiple sub-pictures with one character as a unit and transmits them to the single-character recognition unit respectively, and the single-character recognition unit splices the character results recognized from each sub-picture into a result representing the total amount number and returns it.
[0041] Technical effects
[0042] As a whole, the present invention solves the technical problems that the existing image processing technology cannot process the business trip filling information and the verification and calculation of business trip vouchers, as well as the difficulties brought by the anti-counterfeiting patterns of the air ticket itinerary receipt itself and the limitations of the scanning quality to the detection and recognition of information such as the total amount of the air ticket itinerary receipt.
[0043] Compared with the prior art, the present invention uses a feature-oriented graphics processing method to quickly recognize the total amount content in air ticket itinerary receipt pictures with different scanning qualities. The average recognition time for a single itinerary receipt is about 1 s; through experiments, the overall recognition accuracy rate reaches more than 96%; because the present invention integrates a graphics detection method based on feature engineering and a recognition method based on neural network, it is universal when applied to recognizing other key information of air ticket itinerary receipts or other types of bills, and there is no need to provide a large amount of sample data for training. Description of the drawings
[0044] Figure 1 It is a flow chart of the present invention. Detailed implementation manners
[0045] As Figure 1 shown, this embodiment relates to a method for quickly recognizing the total amount of an air ticket itinerary receipt. By preprocessing the itinerary receipt picture, framing the total amount number, and splitting single characters, each character picture in the total amount information is extracted from the itinerary receipt picture, and then a deep neural network model is used to recognize a single character picture into character-type data and merge them into a total amount recognition result. Specifically, it includes:
[0046] Step 1. Picture preprocessing:
[0047] 1.1) Cut off the non-itinerary receipt part, such as the redundant white edges around, etc.
[0048] 1.2) When the horizontal resolution of the picture exceeds 1600 pixels, it is compressed proportionally to 1600 pixels, and the INTER_AREA mode is adopted for the compression method.
[0049] 1.3) By observing the experiment, determine that the HSV space of the blue fine lines is from [78, 43, 46] to [100, 255, 255]; for the pixels in the original image whose HSV values are within this HSV space, set their HSV values to [255, 255, 255], thereby roughly removing the interference of the blue background fine lines.
[0050] Step 2: Key information selection:
[0051] 2.1) Convert the color image into a grayscale image through color space conversion. The color space conversion uses the BGR2GRAY mode.
[0052] 2.2) Use morphological opening operation to remove impurities, and then obtain the coordinates of the 4th and 6th straight lines through the Hough line detection. Set the horizontal area range to 70%-95% of the abscissa of the image to determine the approximate area where the total amount text appears. The operation kernel of the opening operation is a rectangle of 25*2 pixels. The radius resolution (rho parameter) of the Hough line detection is set to 1.0; the angle resolution (theta parameter) is set to / 2; the straight line point threshold is set to 150, the line segment length threshold is set to 0.5 times the width of the original image, and the minimum threshold for the distance between two points of the line segment is set to 0.4 times the width of the original image.
[0053] Among them, the binarization threshold is set to a fixed value of 205 through experiments;
[0054] 2.3) Perform the opening operation again to connect the discrete numbers. The operation kernel is a rectangle of 70*9 pixels
[0055] 2.4) Perform the opening operation through the horizontal and vertical element kernels respectively to remove impurities and obtain a complete digital connected region block; the horizontal element kernel is a rectangle of 9*1 pixels, and the vertical element kernel is a rectangle of 1*7 pixels. The impurities refer to the thin horizontal and vertical straight lines that do not belong to the target graph and have not been completely eliminated in Step 2.
[0056] 2.5) Draw the minimum bounding rectangle of the connected region in BGR color of [0, 128, 0], where the width of the rectangle sides is set to 2 pixels.
[0057] Step 3: Character segmentation:
[0058] 3.1) Obtain the image within the selected rectangle according to the BGR color of [0, 128, 0].
[0059] 3.2) After binarization, remove impurities through the opening operation, and remove the horizontal lines by traversing horizontally and judging the length of continuous pixels. Specifically, the length threshold of continuous pixels is set to 35 pixels.
[0060] In this embodiment, the minimum effective pixel rate min_black_rate is set to 0.35 through dynamic binarization processing.
[0061] 3.3) Traverse from left to right, and divide the original rectangular image into multiple single-character regions according to whether the number of black pixels in each column of the image exceeds the threshold. Among them, the threshold is set to 1 pixel, and the specific steps include: for the currently traversed column, count the number of black pixels from top to bottom; when the number of black pixels is greater than the threshold, that is, the current column is not the column where the valid character is located, skip this column; when the number of black pixels is less than the threshold, that is, the current column belongs to a certain valid character, record the starting abscissa, and then continue to judge the next column. Until the next column with the number of black pixels greater than the threshold is encountered, that is, the valid character ends, record the ending abscissa. Record the starting abscissa and the ending abscissa as a single-character region; add the current single-character region to the character list.
[0062] 3.4) Perform width recheck on each single character in the character list:
[0063] 3.4.1) When the width of a single character exceeds the maximum character width limit, that is, there is digital adhesion, it is necessary to repeat the operation in step 3.3) for segmentation;
[0064] 3.4.2) When the width of a single character is less than the minimum character width limit, or the distance between two characters is less than the minimum character gap limit, that is, a single digit is split, it is necessary to merge the current single-character region and the next single-character region.
[0065] The setting of the pixel threshold is shown in the following table
[0066] Name Set value Description char_width 30 pixels Default character width min_char_width 4*char_width / 5 Minimum width limit of characters min_char_gap char_width / 3 Minimum gap between characters max_char_width 8*char_width / 5 Maximum width limit of characters
[0067] The opening operation refers to performing one erosion operation and then one dilation operation; the closing operation refers to performing one dilation operation and then one erosion operation.
[0068] Step 4, Total amount recognition based on neural network:
[0069] The neural network is a fully connected neural network constructed based on the keras framework and run through tensorflow, including: an input layer with an input structure of a (28, 28) matrix, a fully connected layer containing 128 nodes, and an output layer containing 10 nodes, where: the input layer receives the binary image returned by the character segmentation process as input, and the output layer outputs the probabilities of the recognition results for each digit from 0 to 9.
[0070] In the fully connected neural network of the total amount recognition process based on the neural network, the relu activation function is used, and the Dropout rate is set to 0.2.
[0071] The fully-connected neural network model in the total amount recognition process based on neural network is trained based on the MNIST dataset. The download address of the MNIST dataset is: http: / / yann.lecun.com / exdb / mnist /
[0072] Step 5: Use the number with the highest probability of each single character to be recognized in the current air ticket itinerary recognition task as the recognition result, and merge the recognition results into the final total amount result.
[0073] In this embodiment, 3,637 air ticket itinerary pictures with different scanning batches are manually labeled to verify the accuracy of the recognition results of the embodiment. The so-called labeling means that after manually checking the total amount number in the air ticket itinerary picture, the image file is named in the following format:
$default number
$total amount
[0074] Compared with the prior art, the present invention uses the key text information detection technology oriented to feature engineering to avoid the problem that a large amount of training data is required for text detection directly through deep learning technologies such as convolutional neural networks. For example, in this embodiment, only 300 sample data are observed during the development and training process, and finally a recognition accuracy of more than 96% is obtained.
[0075] The above specific implementation can be locally adjusted in different ways by those skilled in the art without departing from the principles and purposes of the present invention. The protection scope of the present invention is subject to the claims and is not limited by the above specific implementation. All implementation solutions within its scope are subject to the present invention.
Claims
1. A method for quickly identifying the total amount of an air ticket itinerary, characterized in that, by preprocessing the air ticket itinerary picture, selecting the digital frame of the total amount, and splitting single characters, each character picture in the total amount information is extracted from the itinerary picture, and then a deep neural network model is used to identify the single character picture into character type data and merge them into the total amount recognition result; The single character splitting mentioned above means: dividing the key information text area into multiple sub-areas so that each sub-area contains and only contains one character, specifically including: Step ①: After obtaining the selected image, perform dynamic binarization processing and remove impurity pixels through opening operation; Step ②: Traverse horizontally and judge the length of continuous pixels to remove horizontal lines; Step ③: Traverse from left to right, and perform threshold judgment according to the number of black pixels in each column of the image and divide the rectangular image into multiple single-character areas, that is, the image that contains and only contains a single character after division. Specifically: for the currently traversed column, count the number of black pixels from top to bottom: when the number of black pixels is greater than the threshold, that is, the current column is not the column where the valid character is located, skip this column; when the number of black pixels is less than the threshold, that is, the current column belongs to a certain valid character, record the starting abscissa, and then continue to judge the next column until the next column where the number of black pixels is greater than the threshold is encountered, that is, the valid character ends, record the ending abscissa; record the starting abscissa and ending abscissa of this step as a single-character area; add the current single-character area to the character list; Step ④: Perform width recheck on each single-character area in the character list, that is, determine whether a certain single-character area needs to be split or merged by setting a pixel width threshold. Specifically: when the width of the single-character area is too large to cause digital adhesion, return to Step ③ to re-segment; when the width of the single-character area is too small or the horizontal distance between two areas is too small to cause a single character to be split, merge the adjacent split parts into a single-character area; For the dynamic binarization processing, since the font of the same kind of information on the same kind of bill is determined, the parameter value min_black_rate that does not depend on the quality of the picture to be recognized is determined, and the invalid pixels with low gray level are removed on the premise of ensuring the least loss of effective pixels. Specifically: i) Observe the thickness of the characters and determine the minimum effective pixel rate min_black_rate through experiments, where: the min_black_rate of bold fonts is greater than that of non-bold fonts; ii) Calculate the minimum number of effective pixels target_black_count = w × h × min_black_rate, where: w and h are the length and width pixel values of the rectangular frame image respectively; iii) Initialize the threshold to be determined target_v to 0 and the current cumulative effective pixels curr_black_count to 0; iv) curr_v loops from 0 to 256: a) Statistically calculate the number of pixels with gray value equal to curr_v in the picture through matrix operation and accumulate it to curr_black_count; b) End the loop and set target_v to curr_v when curr_black_count is greater than or equal to target_black_count; otherwise, continue the loop; v) Then, binarize the image according to the obtained target_v, that is, set the gray value greater than target_v as valid pixels, i.e., 255, and set the gray value less than or equal to target_v as invalid pixels, i.e., 0.
2. The method for quickly identifying the total amount of air ticket itinerary receipts according to claim 1, characterized in that the image preprocessing refers to: after unifying the styles of image materials in different scanning batches through operations such as cropping, preliminarily determining the approximate area where the total amount information is located, specifically including: Step 1: Read the image, correct the image through rotation operation, crop the part that is not the itinerary receipt, and compress the image with too large size; Step 2: According to the position range of the key information to be recognized, roughly determine the horizontal and vertical coordinate intervals of the key information in the image, and crop the original image into the region of interest; The approximate area refers to: the vertical interval between the 4th and 6th horizontal lines of the air ticket itinerary receipt, and the horizontal interval of the right half of the air ticket itinerary receipt; Step 3: Filter the blue fine lines in the background according to the color channel, that is, for the pixels with HSV values within the set range, set their HSV values to [255, 255, 255].
3. The method for quickly identifying the total amount of air ticket itinerary receipts according to claim 1, characterized in that the total amount digital frame selection refers to: delineating the text image representing the total amount with a rectangular frame of the corresponding color for use in the single-character segmentation process, specifically including: Step i: Convert the color image to a grayscale image, and then perform a binarization operation on the grayscale image; Step ii: Use the opening operation to remove impurity pixels, that is, discrete pixels with a pixel value of 255 that do not represent the total amount characters; Step iii: After fusing the discrete pixels with a gray value of 255 representing the total amount through the closing operation, remove the protruding pixels at the upper and lower boundaries of the graph by judging the horizontal width; Step iv: Screen out the positive image with the largest area from the image after removing the protruding pixels, and draw the positive circumscribed rectangle of the positive image to obtain the frame selection area of the total amount information.
4. The method for quickly identifying the total amount of air ticket itinerary receipts according to claim 3, characterized in that the opening operation refers to performing one erosion operation and then one dilation operation; the closing operation refers to performing one dilation operation and then one erosion operation; The described corrosion operation means that structure A is corroded by structure B, then The dilation operation mentioned above means that structure A is dilated by structure B, then 5. The method for quickly identifying the total amount of air ticket itinerary receipts according to claim 1, characterized in that for the recognition, after the single-character binary images returned by the single-character segmentation process are recognized into character type data through a trained neural network model, each digital image is separately recognized into a character result and then merged into a string result, that is, the total amount result. This training refers to supervised learning based on the MNIST dataset to obtain a machine learning model capable of recognizing single-character images.
6. A system for implementing the method according to any one of claims 1-5, characterized in that it includes: An image preprocessing unit, a total amount detection and bounding unit, a single-character splitting unit, and a single-character recognition unit, where: the image preprocessing unit transmits the image that has been calibrated and has image interference removed to the total amount detection and bounding unit, the total amount detection and bounding unit transmits the image with the total amount text boxed with a corresponding color rectangle to the single-character splitting unit, the single-character splitting unit splits the image in the total amount text rectangle into multiple sub-images with one character as a unit and transmits them to the single-character recognition unit respectively, and the single-character recognition unit recognizes each sub-image into a character result, splices them into a result representing the total amount number, and returns it.