Voucher image processing method, device and equipment
By performing correction preprocessing and target area localization on the voucher image, combined with automated verification technology, the problems of template dependence and low recognition accuracy in existing voucher image processing methods are solved, achieving efficient and accurate voucher information extraction and verification.
Patent Information
- Application Number
- CN202511422784.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-30
- Publication Date
- 2025-11-18
- Estimated Expiration
- 2045-09-30
AI Technical Summary
Existing methods for processing voucher images have high template dependence, low efficiency in adapting to new voucher types, low accuracy in recognizing non-ideal images, low efficiency and high error rate in compliance verification, and are prone to missing duplicate vouchers during manual retrieval.
By acquiring and preprocessing the voucher image, locating the target region and extracting information, and combining it with automated verification, Gaussian filtering, median filtering, tilt correction, resolution normalization, text region segmentation and feature extraction techniques are used. The CNN+BiLSTM+CTC model is used for text recognition and hash value generation to achieve automated verification.
It significantly improves the accuracy of target information recognition under complex shooting conditions, enhances verification efficiency, avoids errors and omissions caused by manual operation, adapts to various certificate types, and shortens the adaptation time for new certificate types.
Smart Images

Figure CN120976950A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The embodiment of the present application relates to the technical field of image processing, in particular to a voucher image processing method, device and equipment. BACKGROUND
[0002] The voucher image refers to a visual image file converted from a paper or electronic voucher by means of shooting (mobile phone, scanner), screenshot or system export, and the core function is to carry verifiable business information, for example, payment voucher, etc. The information extraction in the existing voucher image processing method generally adopts a general OCR (optical character recognition) tool or an industry-specific template OCR, and there is high dependence on the template, and when a new voucher type is added, the template needs to be reconfigured for 1-2 weeks, the adaptation efficiency is low, the recognition accuracy is low when facing non-ideal shooting images such as inclination, blur and shadow, and the compliance verification in the existing voucher image processing method generally adopts a multi-system manual cross comparison, the efficiency is extremely low, the error rate is high, and manual retrieval of duplicate vouchers is easy to miss. SUMMARY
[0003] The technical problem to be solved by the embodiment of the present application is to provide a voucher image processing method, device and equipment, which can adapt to various voucher types, greatly shorten the adaptation time of newly added voucher types, and significantly improve the processing flexibility and efficiency.
[0004] To solve the above technical problems, the technical scheme of the embodiment of the present application is as follows: A voucher image processing method, comprising: acquiring a to-be-processed voucher image, wherein the to-be-processed voucher image comprises a target region containing target information; performing correction preprocessing on the to-be-processed voucher image to obtain a preprocessed voucher image; performing target region positioning processing on the preprocessed voucher image to obtain a plurality of target regions; extracting the target information in the plurality of target regions to obtain a plurality of target information; verifying the plurality of target information to obtain a verification result and outputting the verification result.
[0005] Optionally, the correction preprocessing on the to-be-processed voucher image to obtain a preprocessed voucher image comprises: performing noise removal processing on the to-be-processed voucher image to obtain a first intermediate voucher image; performing inclination correction processing on the first intermediate voucher image to obtain a second intermediate voucher image; performing resolution normalization processing on the second intermediate voucher image to obtain the preprocessed voucher image.
[0006] Optionally, the pre-processed credential image is subjected to target region positioning processing to obtain a plurality of target regions, comprising: The pre-processed credential image is subjected to text region segmentation to obtain a plurality of segmentation regions. The plurality of segmentation regions are subjected to positioning processing to obtain a plurality of target regions.
[0007] Optionally, the pre-processed credential image is subjected to text region segmentation to obtain a plurality of segmentation regions, comprising: The pre-processed credential image is input into a down-sampling processing layer of a preset network model for down-sampling processing to obtain a first target feature credential image. The target feature credential image is input into up-sampling processing of the preset network model to obtain a second target feature credential image. A segmentation probability map is output according to the second target feature credential image and an activation function of the preset network model. The pre-processed credential image is segmented according to a preset threshold and the segmentation probability map to obtain a plurality of segmentation regions.
[0008] Optionally, the plurality of segmentation regions are subjected to positioning processing to obtain a plurality of target regions, comprising: Coordinates of a bounding box of each segmentation region of the plurality of segmentation regions are obtained. A plurality of target regions are determined according to the coordinates of the bounding box of each segmentation region.
[0009] Optionally, target information in the plurality of target regions is extracted to obtain a plurality of target information, comprising: Text information in the plurality of target regions is recognized to obtain a text recognition result. The plurality of target information is obtained according to fuzzy matching of the text recognition result and fields in a target information table in a preset database.
[0010] Optionally, the plurality of target information is subjected to verification to obtain a verification result, comprising: Hash values and confidence levels of the plurality of target information are obtained. The plurality of target information is subjected to verification according to the hash values and confidence levels of the target information and a verification priority to obtain a verification result.
[0011] Embodiments of the present application also provide a credential image processing device, comprising: An acquisition module is configured to acquire a to-be-processed credential image, wherein the to-be-processed credential image comprises a target region of target information. The processing module is configured to perform correction preprocessing on the to-be-processed certificate image to obtain a preprocessed certificate image, perform target region positioning processing on the preprocessed certificate image to obtain a plurality of target regions, extract target information in the plurality of target regions to obtain a plurality of target information, and perform verification on the plurality of target information to obtain a verification result and output the verification result.
[0012] Embodiments of the present application also provide a computing device, comprising: one or more processors; a storage device configured to store one or more programs, when the one or more programs are executed by the one or more processors, the one or more processors implement the method as described above.
[0013] Embodiments of the present application also provide a computing device readable storage medium, the computing device readable storage medium stores a program, when the program is executed by a processor, the method as described above is implemented.
[0014] The above-mentioned scheme of the embodiments of the present application has at least the following beneficial effects: The above-mentioned scheme of the embodiments of the present application, by performing correction preprocessing on the certificate image, can effectively improve the image quality, combined with the optimized target region positioning and information extraction method, significantly improves the target information recognition accuracy under complex shooting conditions. After obtaining the target regions where the plurality of target information is located, the target information is extracted, and the plurality of target information is automatically verified, replacing the traditional manual operation, greatly improving the verification efficiency, and avoiding the errors and omissions caused by manual operation, ensuring the reliability of the verification result. And the processing method of the certificate image does not need to rely on a specific template, can adapt to various certificate types, greatly shortens the adaptation time of new certificate types, significantly improves the processing flexibility and efficiency. BRIEF DESCRIPTION OF DRAWINGS
[0015] Figure 1 is a flowchart of a certificate image processing method provided by an embodiment of the present application.
[0016] Figure 2 is a module schematic diagram of a certificate image processing device provided by an embodiment of the present application. DETAILED DESCRIPTION
[0017] Exemplary embodiments of the present application will be described in greater detail below with reference to the accompanying drawings. Although exemplary embodiments of the present application are shown in the drawings, it should be understood that the present application can be implemented in various forms and should not be limited by the embodiments described herein. On the contrary, these embodiments are provided so that the present application can be more thoroughly understood and the scope of the present application can be accurately conveyed to those skilled in the art.
[0018] AsFigure 1 As shown, the embodiment of the present application provides a voucher image processing method, comprising: Step 11, obtaining a to-be-processed voucher image, the to-be-processed voucher image comprising: a target region containing target information; here, the voucher image can be a payment voucher type image, comprising core business information, including customer name, amount, serial number and other target information, and the positions of these target information in the voucher image are the target region; Step 12, performing correction preprocessing on the to-be-processed voucher image to obtain a preprocessed voucher image; Step 13, performing target region positioning processing on the preprocessed voucher image to obtain a plurality of target regions; Step 14, extracting target information in the plurality of target regions to obtain a plurality of target information; Step 15, verifying the plurality of target information to obtain a verification result and outputting.
[0019] In this example, by performing correction preprocessing on the voucher type image, the image quality can be effectively improved, and by combining the optimized target region positioning and information extraction method, the target information recognition accuracy under complex shooting conditions can be significantly improved. After obtaining the target regions where the plurality of target information is located, the target information is extracted, and the plurality of target information is automatically verified, replacing the traditional manual operation, greatly improving the verification efficiency, avoiding errors and omissions caused by manual operation, and ensuring the reliability of the verification result. Moreover, the voucher type image processing method does not need to rely on a specific template, can adapt to various voucher types, greatly shortens the adaptation time of newly added voucher types, and significantly improves the processing flexibility and efficiency.
[0020] In an optional embodiment of the present application, in step 11, a to-be-processed voucher image is obtained, the to-be-processed voucher image comprising: a target region containing target information.
[0021] Specifically, the target region can be a region carrying core business information, and the core business information can include: a payer / recipient customer name ID, a transaction amount, a transaction serial number, a transaction time, and a bank / payment platform logo.
[0022] In this example, the target region is explicitly defined as a region carrying core business information, avoiding wasting computing power on irrelevant regions (such as blank spaces and decorative patterns) during subsequent image processing and information extraction, improving overall processing efficiency, and reducing recognition interference caused by irrelevant information, indirectly improving core field extraction accuracy.
[0023] The core business information is clearly defined, including the key items such as the payer / recipient customer name ID and the transaction amount, covers the core dimensions required for the payment voucher verification, avoids the problems that the manual verification or general OCR may miss the key information (such as serial number, bank identification) in the background technology, ensures that the subsequent rule verification (such as repeated verification, amount matching) can be carried out based on complete core information, and improves the verification accuracy.
[0024] The clear division of the target area and the core business information enables the subsequent image preprocessing to be optimized in the quality of the core area, the information extraction to be accurately positioned in the core field, and the basis to be provided for the hierarchical and accurate processing of each link.
[0025] In an optional embodiment of the present application, in step 12, the to-be-processed voucher image is subjected to correction preprocessing to obtain a preprocessed voucher image, which includes: In step 121, the to-be-processed voucher image is subjected to noise removal processing to obtain a first intermediate voucher image; specifically, it can include: In step 1211, according to The to-be-processed voucher image is subjected to smoothing processing; Wherein, x , y ) is the image coordinate of the to-be-processed voucher image, G x , y ) is a two-dimensional Gaussian function, which is used to describe the weight distribution of the Gaussian filter at the image coordinate x , y ), the image is subjected to smoothing processing through the function to reduce high-frequency noise, σ is the standard deviation of the two-dimensional Gaussian function, which determines the width of the Gaussian distribution; In step 1212, the voucher image subjected to the smoothing processing is subjected to median filtering to obtain the first intermediate voucher image; Specifically, the median value of all pixel gray values in a 3*3 window can be taken to replace the center pixel of the window to remove salt and pepper noise; In step 122, the first intermediate voucher image is subjected to inclination correction processing to obtain a second intermediate voucher image; specifically, it can include: In step 1221, according to The first intermediate voucher image is converted into a black-and-white binary image; Wherein, g x , y ) is the original gray value of the first intermediate voucher image, f x , y ) is the pixel value of the first intermediate voucher image after binarization, R To set the threshold value; Step 1222, edge detection is performed on the black and white binary image to extract the text edge; Specifically, if the pixel gradient value is greater than or equal to 150, it is determined as a strong edge and retained; If 50 < pixel gradient value < 150: only when connected with a strong edge, it is retained, otherwise it is rejected; If the pixel gradient value is less than or equal to 50, it is determined as a non-edge and rejected; Step 1223, polar coordinate transformation is used to detect the inclination angle of the straight line of the first intermediate credential image θ When θ > ± 0.5°, it is determined that the image is inclined; Wherein, the distance resolution is 1, and the angle resolution is π / 180; Step 1224, if the image is inclined, the image angle is adjusted according to To obtain the second intermediate credential image; Wherein, M is a rotation matrix, θ is the inclination angle of the straight line, dx , dy is the image center offset; Step 123, the second intermediate credential image is subjected to resolution normalization processing to obtain a preprocessed credential image; specifically, it can include: Step 1231, according to the original image resolution, different interpolation algorithms are used to unify the image to 300DPI (dots per inch), 2480x3508 pixels, 24-bit RGB format, to obtain a preprocessed credential image; Specifically, when the resolution is less than 300DPI, according to The bilinear interpolation is performed on the second intermediate credential image to obtain a preprocessed credential image; Wherein, H ( x , y ) is the pixel value at the coordinates ( x , y ) after bilinear interpolation, F ( x , y ) are the original pixel values of the four pixels around the interpolation point at the coordinates ( x i , y j ), i , j =0,1, x i , y j are the coordinates of the four pixels around the interpolation point, wi , w j The weights for bilinear interpolation are determined by the distance between the point to be interpolated and its four surrounding pixels. They are used to calculate the interpolated pixel value and avoid pixel block artifacts. When the resolution is greater than 300 DPI, according to Interpolate the second intermediate voucher image to obtain a preprocessed voucher image; in, Q ( x This is the interpolation function used when the resolution is greater than 300 DPI, used to calculate the interpolation weights to obtain the preprocessed voucher image. d The normalized distance is the result of normalizing the distance between the point to be interpolated and the relevant pixels.
[0026] In this example, Gaussian filtering and median filtering are used in series for processing. First,... σ The image is smoothed using a 5×5 convolution kernel Gaussian filter with a coefficient of 1.0 to suppress high-frequency noise, and then salt-and-pepper noise is removed using a 3×3 windowed mid-range filter. This results in an image signal-to-noise ratio ≥35dB and text edge preservation ≥98%. This avoids character breakage and misrecognition problems caused by noise in the background technology. For example, complex monetary numbers (such as 1234.56) will not be misjudged as 123.56 or 123456 due to noise, laying the foundation for accurate subsequent recognition.
[0027] Correcting angular deviations and eliminating text line breaks involves four steps: binarization, edge detection, angle detection, and rotation correction. First, the image is converted to a black-and-white binary image using an adaptive threshold. Then, edge detection (low threshold 50, high threshold 150) extracts the text edges. Next, Hough transform is used to detect the tilt angle. Finally, a rotation matrix is applied. M The correction angle is set to ≤±0.3°. This completely solves the text line breakage problem caused by tilted images in the background technology. For example, the tilted payer's name will not be split into two segments due to line breaks and cannot be completely extracted, ensuring the integrity of core fields such as customer name ID.
[0028] To unify image specifications and improve recognition consistency, a differentiated interpolation algorithm is used based on the original resolution: bilinear interpolation is used to avoid pixelation when the resolution is <300 DPI, and interpolation is used to preserve text details when the resolution is >300 DPI. The final output is a unified 300 DPI, 2480×3508 pixel, 24-bit RGB image with a pixel deviation of ≤2 pixels. This solves the problem of inconsistent image resolutions from different shooting devices (mobile phones, scanners) in the background technology, eliminating the need for subsequent recognition models to adapt to multiple specifications. This results in more stable cross-device and cross-scene recognition accuracy, directly supporting the goal of ≥95% cross-bank recognition accuracy.
[0029] In an optional embodiment of the present application, in step 13, the target region positioning processing is performed on the preprocessed credential image to obtain a plurality of target regions, including: In step 131, the preprocessed credential image is subjected to text region segmentation to obtain a plurality of segmented regions. In step 132, the plurality of segmented regions are subjected to positioning processing to obtain a plurality of target regions.
[0030] In this example, the target region is accurately locked by segmentation + positioning, solving the problem of no clear region focus in general OCR in the background technology and easy mis-extraction of irrelevant information.
[0031] In an optional embodiment of the present application, in step 131, the preprocessed credential image is subjected to text region segmentation to obtain a plurality of segmented regions, including: In step 1311, the preprocessed credential image is input into a down-sampling processing layer of a preset network model for down-sampling processing to obtain a first target feature credential image. Specifically, it can include: The preprocessed credential image (2480x3508 pixels) is cropped into an input block of 512x512x3 (RGB channel), input into a down-sampling layer (4 convolution blocks), each convolution block contains convolution + ReLU + pooling operation, and low-level features of the image are extracted to obtain a first target feature credential image (the size is gradually reduced and the number of channels is gradually increased); In step 1312, the target feature credential image is input into the up-sampling processing of the preset network model to obtain a second target feature credential image. Specifically, it can include: The first target feature credential image is input into an up-sampling layer (4 up-sampling blocks), each up-sampling block contains deconvolution + feature splicing (fused with the corresponding layer feature of the down-sampling) + convolution operation, the image size is restored, and a second target feature credential image (the size is consistent with the input block and the number of channels is 1) is obtained; In step 1313, a segmentation probability map is output according to the second target feature credential image and the activation function of the preset network model. Specifically, it can include: According to the activation function The second target feature credential image is processed to output the probability that each pixel belongs to a text region, i.e., a segmentation probability map; Wherein, S ( x , y ) is the feature value of the second target feature credential image at ( x , y ), P ( x , y ) is the probability that each pixel belongs to a text region, P (x , y ) ∈ [0, 1]; Model training uses a loss function for optimization. wherein, LOSS is a loss function, A is a predicted mask, B is a real mask, solving the sample imbalance problem. Step 1314, according to the preset threshold and the segmentation probability map, the preprocessed voucher image is segmented to obtain a plurality of segmentation regions; specifically, it can include: Set the preset threshold to 0.5, when P ( x , y ) ≥ 0.5, it is determined that the pixel belongs to the text region; otherwise, it is a background region, and finally a plurality of segmentation regions (such as customer name segmentation area, amount segmentation area, etc.) are obtained.
[0032] In this example, the text region segmentation process of step 131, through the down-sampling, up-sampling, probability calculation and threshold segmentation of the preset network model (U-Net), realizes the accurate extraction of the voucher text region, solves the problems of general OCR in the background technology, such as non-targeted segmentation, easy mixing with background interference, and inaccurate positioning of core region, and provides key support for high precision and high efficiency of subsequent information extraction.
[0033] Down-sampling processing, first cut the preprocessed voucher image of 2480x3508 pixels into an input block of 512x512x3, and then process it through 4 convolution blocks containing convolution+ReLU+pooling. In the process of gradually reducing the image size and increasing the number of channels, the low-level features of the image (such as edges and textures) are efficiently extracted. This step lays the foundation for subsequent feature analysis, avoids wasting computing power due to the large size of the original image, and at the same time, through the pooling operation, the key features are strengthened and the redundant information interference is reduced.
[0034] The up-sampling processing is through 4 up-sampling blocks containing deconvolution+feature splicing+convolution, to restore the first target feature voucher image obtained by down-sampling to the same size as the input block, and the number of channels is reduced to 1. The design of corresponding layer feature fusion of down-sampling can combine the detail features retained in the down-sampling stage with the global features in the up-sampling stage, avoid the loss of details caused by pure up-sampling, and ensure the accurate restoration of the edges and contours of the text region, providing high-quality feature maps for subsequent pixel-level segmentation.
[0035] In the segmentation probability map generation link, the feature value of the second target feature certificate image is converted into the probability of each pixel belonging to the text region through an activation function, and a loss function is used to optimize the model training. The loss function effectively solves the problem of unbalanced samples of the text region and the background region of the certificate image by calculating the intersection ratio of the predicted mask (A) and the real mask (B), avoids the model from being biased to misjudgment due to the high proportion of the background region, and improves the accuracy of the probability prediction of the text region.
[0036] Finally, threshold segmentation takes 0.5 as a preset threshold, and pixels with a probability greater than or equal to 0.5 are determined as text regions, and otherwise as background regions, and finally the exclusive segmentation regions of the customer name, the amount and the like are obtained. This quantitative determination method avoids the subjectivity of manual segmentation, and in combination with the foregoing feature processing, the coordinate error of the payee information column is less than or equal to 4 pixels, the amount / serial number column is less than or equal to 2 pixels, and the positioning accuracy rate reaches 99.1%.
[0037] The process does not need to rely on a special template of a traditional OCR, and does not need to be reconfigured when facing new types of certificates, and the adaptation efficiency is greatly improved; meanwhile, the accurate region segmentation can make the subsequent information extraction focus on the target region only, reduce irrelevant background interference, directly support the target of the core field extraction accuracy being greater than or equal to 96%, and lay a key foundation for the automation and high precision of the entire certificate verification process.
[0038] In an optional embodiment of the present application, in step 132, the plurality of segmentation regions are subjected to positioning processing to obtain a plurality of target regions, which includes: In step 1321, the coordinates of the bounding box of each segmentation region of the plurality of segmentation regions are obtained; specifically, it can include: For each segmentation region, the minimum circumscribed rectangle (bounding box) thereof is extracted by contour detection, and the coordinates of the upper left corner ( x min , y min ) and the lower right corner ( x max , y max ) of the bounding box are recorded, that is, the coordinates of the bounding box of each segmentation region are [ x min , y min ], x max , y max ]; In step 1322, a plurality of target regions are determined according to the coordinates of the bounding box of each segmentation region; specifically, it can include: The segmented regions are filtered according to business rules (such as the bounding box of the amount segmentation area must contain the ¥ symbol, and the customer name segmentation area must be located in the upper half of the image). The segmented regions after filtering are the target regions, and their bounding box coordinates are consistent with the coordinates obtained in step 1321.
[0039] In this example, the minimum bounding rectangle of each segmented region is extracted through contour detection, and the top left corner is recorded. x min , y min ) With the bottom right corner ( x max , y max The coordinate system transforms the originally abstract segmented area into a precise coordinate range. This avoids the ambiguity of approximate areas in traditional positioning, allowing subsequent information extraction to accurately focus on the range defined by the coordinates. For example, amount extraction is only performed within [( x min , y min ), ( x max , y max This process is performed within the []] area, reducing the erroneous extraction of surrounding irrelevant text (such as notes and descriptions). At the same time, the coordinate data also provides a standardized data format for subsequent cross-module interactions (such as matching fields in the information extraction layer), improving the efficiency of process connection.
[0040] By identifying the core target area and eliminating invalid interference, the segmented regions are filtered according to business rules (such as the amount area containing the ¥ symbol and the customer name area being located in the upper half of the image). This ensures that the final output target areas are all areas carrying core business information (payer ID, amount, etc.). This solves the problem in the background technology where general segmentation may include irrelevant text areas (such as advertisements and decorative text blocks) in the processing scope, avoiding the waste of computing power in non-core areas. At the same time, by filtering according to rules in advance, invalid identification is reduced during subsequent information extraction, indirectly improving the accuracy of core field extraction. This provides a preliminary guarantee for the target of core field extraction accuracy ≥96% and customer name variant matching accuracy ≥92%.
[0041] In an optional embodiment of the present invention, in step 14, target information is extracted from the plurality of target regions to obtain a plurality of target information, including: Step 141 involves recognizing the text information in the multiple target regions to obtain text recognition results; specifically, this may include: The CNN+BiLSTM+CTC model is used to recognize the text in the target region, and the text recognition results are obtained. CNN feature extraction: input the image of the target area (such as the amount area) into the ResNet50 model, keep the output of the conv5_x layer, and reduce the channel number to 256 through 1x1 convolution to obtain the visual feature sequence of the text X [ x 1, x 2,..., x T ] T is the length of the feature sequence, x t ∈ R 256 ; BiLSTM sequence modeling: input X into a 2-layer bidirectional LSTM (hidden layer dimension 256, dropout=0.3) to learn the context dependency of the feature sequence, and output the encoded sequence Y [ y 1, y 2,..., y T ] y t ∈ R C , C is the size of the character set, such as 6000 characters of numbers, letters and Chinese characters; CTC decoding: beam search is used to solve the problem of mismatch between the length of the feature sequence T and the length of the text sequence L , output the text sequence with the highest probability S [ s 1, s 2,..., s L ] (i.e. text recognition result, such as 12345.67 yuan, XX company); Step 142, according to the text recognition result and the field in the target information table in the preset database, a plurality of target information are obtained.
[0042] Specifically, the preset customer name library contains 3 tables (basic customer table: standard name + credit code; historical alias table: customer ID + alias; similar mapping table: customer ID + similar name + similarity), all of which are indexed for optimized query; Taking the customer name recognition result as an example, suppose the recognition result is S rec , and the customer name in the database is S db ; Exact match: if S rec = S dbThen confidence = 98%, direct match success; Fuzzy match: if S rec ≠ S db , match success; S rec and S db Segmentation, get word set C and D , according to , determine the similarity of word set C and D , when J ( C , D ) ≥ 0.8, determine match success; Wherein, J ( C , D ) is the similarity of word set C and D ; Edit distance sorting: according to LD ( S rec , S db ) = min {number of operations}, calculate S rec and S db The minimum number of edits; Where, LD ( S rec , S db ) is the minimum number of edits of S rec and S db , edit includes: insertion / deletion / replacement operation; When LD ( S rec , S db ) ≤ 3, take the top 5 candidates as the matching result; for auxiliary fields (such as serial number, time) regular matching is adopted.
[0043] In this example, the CNN+BiLSTM+CTC architecture is adopted, the visual features of the target area (such as the amount of money, the customer name area) are extracted through the conv5_x layer of ResNet50, then the context dependency relationship is learned through 2 layers of bidirectional LSTM, and finally the beam search decoding is used to solve the problem of mismatch between features and text sequence length. Compared with the general OCR of the background technology, this architecture can better cope with the font difference and character sticking of the voucher text, for example, the amount of money 12345.67 yuan will not be misjudged as 1234567 yuan due to font blur, and the text recognition accuracy of cross-bank and cross-platform is ≥95%, providing high-quality initial text data for subsequent matching.
[0044] Cover name variants and auxiliary fields to ensure information matching integrity Based on a customer name library containing 3 index tables, the customer name is processed through three layers of logic processing: exact matching ensures high confidence (98%) for completely consistent names; fuzzy matching with a similarity ≥0.8 and the top 5 candidates with an edit distance ≤3 solve the matching problem of variants such as XX Company Limited and XX Company Limited; and the auxiliary fields (serial number, time) are matched accurately using regular expressions. Compared with the single matching logic in the background technology, this method avoids matching failures caused by customer name variants, ensures that auxiliary fields are not missed, and the index optimization makes the time consumption of a single matching ≤100 ms, balancing accuracy and efficiency.
[0045] In an optional embodiment of the present application, in step 14, the target information in the plurality of target areas is extracted to obtain a plurality of target information, which includes: In step 143, the hash value generation and confidence evaluation of the plurality of target information are performed; specifically, it can include: In step 1431, according to Y =0.299 r +0.587 g +0.114 b , the target area RGB image is converted to grayscale; Wherein, Y is the grayscale value, r , g , b is the RGB channel value; In step 1432, the 16x16 block MD5 preprocessing is performed on the grayscale image; and a 64-bit hash value is generated by SHA-256; In step 1433, according to CF 1= PM 1× KT 1, the core field confidence is determined; Wherein, CF 1 is the core field confidence, PM 1 is the recognition model probability value,KT 1 is a field integrity coefficient (complete = 1.0, missing suffix = 0.8, missing prefix = 0.7, missing middle = 0.5); Step 1434, according to CF 2= PM 2x KT 2, determine the auxiliary field confidence; Wherein, CF 2 is the auxiliary field confidence, PM 2 is the regular matching score (complete match = 1.0, partial match = 0.5), KT 2 is the position matching score (target area = 1.0, non-target area = 0.5).
[0046] In this example, the target area RGB image is first converted to grayscale by the formula, then preprocessed by 16x16 block MD5 to offset the ±5% brightness jitter interference, and finally a 64-bit hash value is generated by SHA-256. This hash value can uniquely identify the target area features of the voucher, and after storing it in the Redis3 master 3 slave cluster, the query response is ≤50ms, supporting 1000+ queries per second, providing accurate basis for subsequent rule verification repeated interception (priority 2), directly supporting repeated voucher interception rate ≥99.98%, and completely solving the problem of easy omission of repeated vouchers in manual retrieval.
[0047] The information quality is screened and balanced, and the confidence is calculated for the core fields (such as customer name and amount); and the confidence is calculated for the auxiliary fields (such as serial number and time). Through the grading strategy of ≥95% automatic verification, 80%-94% pending confirmation, and <80% manual review, the low efficiency of full manual verification in the background technology is avoided, and the high error rate caused by direct verification of low-quality information is prevented, while the interface response is ≤50ms, ensuring efficient operation of the overall verification process.
[0048] In an optional embodiment of the present application, in step 15, the plurality of target information is verified to obtain a verification result, comprising: Step 151, obtaining the hash value and confidence of the plurality of target information; Specifically, the hash value is the 64-bit SHA-256 hash value generated in step 1432, and the confidence is the core field confidence CF 1 in step 1433 and the auxiliary field confidence CF 2 in step 1434; Step 152, according to the hash value and confidence of the target information and the verification priority, verifying the plurality of target information to obtain a verification result; specifically, it can include: The configurable verification process is realized based on a rule engine. The rules are stored in XML and MySQL. XML defines the verification logic, and MySQL stores the rule metadata. The parser converts the XML format into the DRL format, and then the verification is performed in the order of priority from 1 to 5. In the priority 1, the type verification is performed. The feature points of the bank or payment platform logo in the voucher are extracted, and compared with the preset feature library. If the matching degree is not less than 0.75, it is determined that the logo is valid. Meanwhile, the TF-IDF (term frequency-inverse document frequency) value of the fixed text (such as the bank return electronic payment voucher) in the voucher is determined. If the TF-IDF value is not less than 0.8, it is determined that the text is valid. If both the logo and the text are valid, the voucher type is qualified, otherwise, it is unqualified. In the priority 2, the repeated verification is performed. The hash value of the target information is queried in the cluster. If the hash value of the target information exists, it is determined that the information is repeated. Meanwhile, the joint index of the serial number and the customer ID is queried in the MySQL. If the combination exists, it is determined that the information is repeated. If any condition is met, the repeated voucher is intercepted. In the priority 3, the customer matching is performed. The customer ID in the order system is compared with the customer ID extracted from the target information. If the two are consistent, the customer matching is qualified, otherwise, it is unqualified. In the priority 4, the amount matching is performed. The amount in the target information is standardized, and converted into the decimal type after removing the non-numeric characters. Then, the amount is compared with the order amount. If the absolute error is not more than 0.01 yuan, the amount matching is qualified, otherwise, it is unqualified. In the priority 5, the abnormal interception is performed. If the voucher is repeated in the priority 2, and the customer ID is not matched in the priority 3, and the amount is not matched in the priority 4, the identification information, error code and error information of the voucher are written into the log library. Finally, according to the verification result and the confidence level, the processing is performed. If the priorities 1 to 4 are all qualified, and the confidence level is not less than 95%, the verification is passed automatically. If the priorities 1 to 4 are all qualified, but the confidence level is between 80% and 94%, the verification is marked as a to-be-confirmed state. If any of the priorities 1 to 4 is unqualified, or the confidence level is less than 80%, the verification is failed, and the specific failure reason is output, such as repeated voucher, customer ID mismatch, amount mismatch, etc.
[0049] In this example, the rules are stored in XML and MySQL (XML defines the logic, and MySQL stores the metadata). After being converted into the DRL format by the parser, the rules are executed. The rules can be edited and updated visually, and the effective time is less than or equal to 30 seconds, and the iteration period is less than or equal to 2 hours. In the background technology, the rules are hard-coded, and the iteration period is more than or equal to 72 hours. The efficiency is improved by more than 36 times. The system can quickly respond to the new verification requirements of the business (such as the addition of the payment platform voucher type verification). The system does not need to be reconstructed, and the adaptability is very strong.
[0050] Through SIFT feature matching (Logo matching degree ≥ 0.75) + TF-IDF text comparison (value ≥ 0.8), the voucher type recognition accuracy rate reaches 98.5%, solving the problem of traditional manual verification of easily confused voucher types (such as misjudging receipts as bank receipts); Redis hash value query + MySQL serial number-client ID joint index query, response ≤ 85ms, repeated interception rate ≥ 99.98%, completely eliminating the risk of missing repeated vouchers in manual retrieval in the background technology; Accurate comparison of customer ID (accuracy 100%), comparison of amount after standardization with an error of ≤ 0.01 yuan, avoiding character misreading and calculation errors during manual checking, and reducing error rate; For high-risk scenarios with repeated and unmatched customer / amount, automatically log, realize traceability of abnormal behavior, and fill the gap of no special monitoring of abnormal vouchers in traditional processes.
[0051] Combined with the confidence of core / auxiliary fields, the results are divided into automatic pass (≥ 95%), pending confirmation (80%-94%), and manual review (< 80%) three categories, which not only ensures that more than 95% of high-confidence vouchers do not need manual intervention, greatly improving efficiency (overall interface response ≤ 220ms), but also avoids misjudgment caused by excessive automation through grading. Compared with the background technology of full manual verification, the labor cost is reduced by more than 80%, and the verification accuracy is stable at more than 99%.
[0052] Example 1 A company receives a special value-added tax invoice image and needs to process the image, extract the invoice key information, verify it with order data, and determine whether the invoice is compliant. The image has slight tilt (visible to the naked eye that the picture is not horizontal) and a small amount of shadow, dust spots (pepper and salt noise); Example 1 provides a voucher image processing method, comprising: Step 21, obtaining a voucher image to be processed: After receiving the user uploaded invoice image, first determine the target area in the image that needs to extract information, including: purchaser information area (including purchaser name, taxpayer identification number, address phone, bank and account number), seller information area (containing the same content as the purchaser information area), invoice code area, invoice number area, goods or taxable labor name area, specification model area, unit area, quantity area, unit price area, amount area, tax rate area, tax amount area, total price including tax area (in both uppercase and lowercase), invoice date area, invoice person area, review area, seller (seal) area; Step 22, correction preprocessing: Smoothing processing: Gaussian filtering method is adopted, according to the weight distribution law of two-dimensional Gaussian function, the high-frequency gray fluctuation in the image caused by the shadow is smoothed, and the noise interference is reduced; wherein the standard deviation of the Gaussian function is set to 1.5, the filtering range is controlled through the parameter, and the overall gray transition of the image is more gentle; Median filtering: the smoothed image is subjected to 3*3 window median filtering, that is, the middle value of all pixel gray values in each 3*3 pixel window is taken to replace the gray value of the center pixel of the window, so as to remove the salt and pepper noise formed by dust in the picture during shooting, and obtain a first intermediate voucher image; Binaryzation processing: the gray threshold is set to 128, the pixels with a gray value greater than or equal to 128 in the first intermediate voucher image are converted to white (pixel value 255), and the pixels with a gray value less than 128 are converted to black (pixel value 0), so that the image becomes a black and white binary image, and the outline of the text and the invoice frame is highlighted; Edge detection: the gradient value of each pixel in the image is calculated, and the edge type is judged according to the gradient value; if the pixel gradient value is greater than or equal to 150, it is determined as a strong edge and directly retained; if the gradient value is between 50 and 150, it is retained only when the pixel is connected with a strong edge, otherwise it is rejected; if the gradient value is less than or equal to 50, it is determined as a non-edge and directly rejected, finally the edge features of the invoice frame and the text are extracted; Straight line inclination angle detection: polar coordinate transformation method is adopted to detect the straight line in the image, the distance resolution is set to 1 (that is, the distance is 1 unit in the polar coordinate), and the angle resolution is set to π / 180 (that is, the angle is 1 degree per step); after detection, it is found that the inclination angle of the straight line of the invoice image is 2.3 degrees, which is greater than ±0.5 degrees, and it is determined that the image is inclined; Angle adjustment: taking the center of the image as the rotation point (the image size is normalized to 2480*3508 pixels later, the center coordinates are 1240 pixels horizontally and 1754 pixels vertically), the image angle is adjusted according to the rotation matrix, the inclined image is rotated-2.3 degrees (reverse offset inclination angle), and the invoice is restored to the horizontal state, and a second intermediate voucher image is obtained; First, the original resolution of the second intermediate voucher image is detected, and it is found that it is 200DPI (lower than the target resolution of 300DPI), so the bilinear interpolation algorithm is used for resolution adjustment; this algorithm calculates the original pixel values of the four pixels around the interpolation point, determines the weight according to the distance between the interpolation point and the four pixels, and then calculates the pixel value after interpolation according to the weight, so as to avoid the block effect after adjustment; finally, the image is uniformly processed into a preprocessed voucher image with a resolution of 300DPI, a size of 2480*3508 pixels and a format of 24-bit RGB, at this time the image is clear and has no inclination, which meets the requirements of subsequent processing; Step 23, target area positioning processing: Down-sampling processing: the pre-processed credential image is cropped to an input block of 512x512 pixels containing 3 channels of RGB, and is input into a down-sampling processing layer of a preset network model; the down-sampling layer contains 4 convolution blocks, each of which sequentially performs convolution (using a 3x3 convolution kernel with a step size of 1), ReLU activation function processing (enhancing the non-linear expression ability of the model), and pooling (using a 2x2 pooling window to compress the image size) operations, gradually extracting low-level features (such as edges and textures) of the image, and finally obtaining a first target feature credential image (reduced in size to 32x32 pixels and increased in channel number to 256); Up-sampling processing: the first target feature credential image is input into an up-sampling processing layer of a preset network model, and the up-sampling layer contains 4 up-sampling blocks, each of which sequentially performs deconvolution (restoring the image size), feature concatenation (fusing the up-sampling layer features with the corresponding down-sampling layer features to supplement the detail information), and convolution operations, gradually restoring the image size, and finally obtaining a second target feature credential image (restored in size to 512x512 pixels and having a channel number of 1); Segmentation probability map generation: the second target feature credential image is processed using a Sigmoid activation function to output the probability of each pixel belonging to a text region (the probability value is between 0 and 1), forming a segmentation probability map; during model training, the model is optimized using a cross-entropy loss function to solve the sample imbalance problem (large difference in pixel number between text regions and background regions) and improve the accuracy of probability calculation; Region segmentation: set the probability threshold to 0.5, and determine the pixels with a probability value greater than or equal to 0.5 in the segmentation probability map as text regions and the pixels with a probability value less than 0.5 as background regions, and finally segment 12 segmentation regions from the pre-processed credential image, each corresponding to a column of the invoice; Boundary box coordinate acquisition: perform contour detection on each segmentation region to find the smallest rectangle (boundary box) that can completely enclose the region, and record the coordinates of the top-left corner (minimum horizontal value, minimum vertical value) and the bottom-right corner (maximum horizontal value, maximum vertical value) of each boundary box; for example, the boundary box coordinates of the buyer's name corresponding segmentation region are (horizontal 300 pixels, vertical 400 pixels) to (horizontal 800 pixels, vertical 460 pixels), and the boundary box coordinates of the total amount of tax (lowercase) corresponding segmentation region are (horizontal 1500 pixels, vertical 2300 pixels) to (horizontal 2000 pixels, vertical 2360 pixels); Target region determination: filter the 12 segmentation regions according to business rules, for example, the total amount of tax region must contain the ¥ symbol, and the invoice code region must be in the format of 10 digits, and exclude invalid interference regions (such as small blocks segmented from the edge blank area of the invoice) that do not meet the rules, and finally determine 10 target regions, the boundary box coordinates of which are consistent with those of the segmentation regions before filtering; Step 24, target information extraction: A CNN (Convolutional Neural Network) + BiLSTM (Bidirectional Long Short-Term Memory Network) + CTC (Connectionist Temporal Classification) model is used to recognize the text information of 10 target regions: CNN feature extraction: input the image of each target region (such as the total tax area image) into the ResNet50 model, retain the output features of the conv5_x layer in the model, and then reduce the feature channel number to 256 through a 1x1 size convolution kernel to obtain the visual feature sequence of the text (the sequence length is determined according to the size of the target region, for example, 80, and each feature element is a 256-dimensional vector); BiLSTM sequence modeling: input the visual feature sequence into a 2-layer bidirectional LSTM network (hidden layer dimension 256, dropout parameter 0.3 to prevent model overfitting), learn the context dependency of the feature sequence (such as the order of the text before and after, semantic association), and output the encoding sequence (the sequence length is consistent with the visual feature sequence, and the dimension of each encoding element is consistent with the size of the character set, which contains 6000 characters including numbers, letters, and Chinese characters); CTC decoding: use beam search algorithm to solve the problem of mismatch between visual feature sequence length and text sequence length, filter out the text sequence with the highest probability from the encoding sequence as the text recognition result; for example, the purchase party name recognition result is Shenzhen ABC Information Technology Co., Ltd., the taxpayer identification number is 91440300XXXXXXXXXX, the total tax (in lowercase) is ¥12960.00, and the invoice number is 00012345; Pre-set database preparation: the pre-set customer name database contains three tables, namely the basic customer table (storing customer standard name, credit code), the historical alias table (storing customer ID, customer's former alias), and the similar mapping table (storing customer ID, customer similar name, similarity), and indexes are established for the three tables to improve query speed; Customer name matching: Take the recognition result of the buyer name as an example, Shenzhen ABC Information Technology Co., Ltd. First, perform an exact match in the basic customer table, and find that the standard name of the customer in the database is Shenzhen ABC Information Technology Co., Ltd. The two are inconsistent, and enter the fuzzy matching stage. Tokenize the recognition result and the database name, and get the word set (Shenzhen, ABC, information technology, Co., Ltd.) after tokenizing the recognition result, and get the word set (Shenzhen, ABC, information technology, Co., Ltd.) after tokenizing the database name. Calculate the similarity of the two word sets (the number of intersection elements divided by the number of union elements), the result is 0.6 (the intersection has 3 elements, and the union has 5 elements), which is less than 0.8, and needs to be further judged by edit distance. Calculate the minimum number of edits (editing operations include inserting, deleting, and replacing characters) between the recognition result and the database name, the result is 1 (replace technology with technology), which is less than or equal to 3, and the database name is listed as a candidate result. Combined with the taxpayer identification number matching, it is found that the recognized 91440300XXXXXXXXXX is consistent with the credit code corresponding to ABC Information Technology Co., Ltd. in the database, and finally the matching is confirmed to be successful. Other information processing: Standardize the amount information format, convert ¥12960.00 to 12960.00 yuan; For auxiliary fields such as serial number and invoice date, use regular matching method (such as invoice date needs to meet the format XXXX-XX-XX) to confirm the information validity, and finally get multiple target information (including buyer, seller, amount, invoice number, etc. Key content); Gray scale conversion: According to the corresponding relationship between RGB channel value and gray scale value (gray scale value = 0.299 * red channel value + 0.587 * green channel value + 0.114 * blue channel value), convert the RGB image of the target area to a gray scale image; Hash value generation: Preprocess the gray scale image in 16x16 pixel blocks (segmented into multiple 16x16 pixel blocks to unify the processing size), and then generate a 64-bit hash value through the SHA-256 algorithm, which is used for subsequent repeated verification, for example, the generated hash value is 5f4dcc3b5aa765d61d8327deb882cf99a61c4195e894f2b; Core field confidence evaluation: The core fields include the buyer's name, taxpayer identification number, amount, and total price with tax. The confidence = recognition model probability value (model confidence probability for text recognition result, such as 0.97) * field integrity coefficient (field information integrity is 1.0, missing suffix content is 0.8, missing prefix content is 0.7, and missing middle content is 0.5); In this identification, the core field information is complete, and the recognition model probability value is 0.97, so the core field confidence = 0.97 * 1.0 = 0.97 (i.e. 97%); Auxiliary field confidence assessment: the auxiliary fields include invoice number, invoice date, and invoicer, and their confidence = regular matching score (1.0 for full match, 0.5 for partial match) x position matching score (1.0 for information located in the target area, 0.5 for located in a non-target area); this time, the auxiliary fields are all fully matched and located in the target area, so the auxiliary field confidence = 1.0 x 1.0 = 1.0 (i.e. 100%); Step 25, information verification, to obtain the verification result: Get the 64-bit SHA-256 hash value of the target information, as well as the core field confidence (97%) and the auxiliary field confidence (100%), as the basis for subsequent verification data; The system realizes a configurable verification process based on a rule engine, and the rules use dual storage methods of XML (to define verification logic) and MySQL (to store rule metadata), and the parser converts the XML format to DRL format for execution of verification: Priority 1 (type verification): extract the feature points of the seller (seal) area in the invoice image, and compare them with the pre-set invoice special seal feature library. If the matching degree is not less than 75%, it is determined that the seal is valid. At the same time, calculate the TF-IDF value (term frequency-inverse document frequency, reflecting the importance of text in the invoice) of the fixed text in the invoice. If the TF-IDF value is not less than 0.8, it is determined that the text is valid. This time, the seal matching degree is 82%, and the fixed text TF-IDF value is 85%, both of which meet the requirements, and the type of the voucher is qualified; Priority 2 (duplicate verification): query the historical voucher hash values stored in the system cluster. If the hash value of the current target information already exists, it is determined to be a duplicate. At the same time, query the joint index of the invoice number and the purchaser's taxpayer identification number in the MySQL database. If the combination already exists, it is also determined to be a duplicate. This time, it is found that neither the hash value nor the joint index exists, and it is determined that the voucher is not duplicated; Priority 3 (customer matching): retrieve the purchaser's taxpayer identification number corresponding to the order from the ERP system, and accurately compare it with the purchaser's taxpayer identification number (91440300XXXXXXXXXX) extracted from the target information. Both are consistent, and the customer matching is qualified; Priority 4 (amount matching): standardize the amount (12960.00 yuan) in the target information (remove non-numeric characters and ensure uniform format), and then compare it with the order amount (12960.00 yuan) in the ERP system. The absolute error is 0 yuan (not more than 0.01 yuan), and the amount matching is qualified; Priority 5 (abnormal interception): if the voucher is determined to be a duplicate in priority 2, and the customer ID does not match in priority 3, and the amount does not match in priority 4, the voucher identification information, error code, and error information need to be written into the log library. This time, there is no abnormal situation, and no interception is needed; Step 26, final verification result and output: After comprehensive verification, priority 1 to 4 are qualified, and the core field confidence is 97% or higher than 95%, and the certificate is automatically passed. The system outputs the extracted 10 target information (buyer, seller, amount, invoice number, etc.), verification pass identification, and generated 64-bit hash value for subsequent financial processes.
[0053] The present application clearly defines the target area of the processed certificate image as the core business information bearing area (including customer ID, amount, etc.). First, it avoids wasting computing power on irrelevant areas such as blank and decoration, directly improving overall efficiency. Second, it eliminates the problem of missing serial numbers, bank logos, and other key information during manual verification or general OCR, ensuring complete data support for subsequent verification. Third, it provides a basis for layered processing in each link, such as pre-processing to optimize the clarity of the amount area, making the processing more targeted.
[0054] Through three-level processing of denoising, tilt correction, and resolution normalization, the image quality is fundamentally improved: Gaussian filtering + median filtering in series for noise removal, with a signal-to-noise ratio ≥ 35 dB, avoiding misjudgment of amount digits due to noise (e.g. 1234.56 misjudged as 123.56); tilt correction controls the angle to ≤ ± 0.3° through a four-step process, reducing the text line break rate to 0%, ensuring complete extraction of fields such as customer name; resolution normalization uses bilinear / Lanczos interpolation according to differences, outputting a standard 300DPI image, solving the problem of different output specifications from devices such as mobile phones and scanners, supporting cross-bank recognition accuracy ≥ 95%.
[0055] Segmentation using the U-Net model, down-sampling to extract features, up-sampling to restore size, combined with DiceLoss optimization for sample imbalance, makes the payee information column coordinate error ≤ 4 pixels, and does not rely on special templates, and new certificate types do not need to be reconfigured; extract the bounding box coordinates and filter according to business rules (such as the amount area containing the ¥ symbol), convert abstract areas to precise coordinates, avoid misgrabbing unrelated text such as notes and advertisements during subsequent extraction, and provide standardized data for cross-module interaction, improving process connection efficiency.
[0056] Using CNN + BiLSTM + CTC architecture to deal with font differences and character sticking, cross-platform text recognition accuracy ≥ 95%; match customer names through precise + fuzzy + edit distance sorting, solving the problem of variants such as Limited Company and Limited Liability Company, and using regular expressions for accurate matching of auxiliary fields, with a variant matching accuracy of ≥ 92%; generate SHA-256 hash values to support repeated interception (response ≤ 50ms), and calculate confidence according to field type to provide a basis for hierarchical processing, avoiding the impact of low-quality information on verification results.
[0057] The rule engine uses XML+MySQL dual storage, and the iteration period is less than or equal to 2 hours (36 times higher than the efficiency of traditional hard coding), and quickly responds to business requirements; priority 1-4 covers types, repetition, customers, and amount verification, the repeated interception rate is greater than or equal to 99.98%, the customer matching accuracy is 100%, and the amount error is less than or equal to 0.01 yuan, which eliminates errors and omissions of manual checking; priority 5 is used for repeated and customer / amount mismatch scenarios to automatically log, realize abnormal traceability; combined with confidence grading (greater than or equal to 95% automatically passed, 80%-94% pending confirmation, and less than 80% reviewed), more than 95% of the vouchers do not need manual intervention, the labor cost is reduced by more than 80%, and the overall interface response is less than or equal to 220ms.
[0058] As shown in Figure 2 Embodiments of the present application also provide a voucher image processing apparatus 20, comprising: An acquisition module 21 is configured to acquire a to-be-processed voucher image, wherein the to-be-processed voucher image comprises a target region of target information. A processing module 22 is configured to perform correction preprocessing on the to-be-processed voucher image to obtain a preprocessed voucher image, perform target region positioning processing on the preprocessed voucher image to obtain a plurality of target regions, extract target information in the plurality of target regions to obtain a plurality of target information, and perform verification on the plurality of target information to obtain a verification result and output the verification result.
[0059] Optionally, the correction preprocessing on the to-be-processed voucher image to obtain the preprocessed voucher image comprises: Performing noise removal processing on the to-be-processed voucher image to obtain a first intermediate voucher image; Performing inclination correction processing on the first intermediate voucher image to obtain a second intermediate voucher image; Performing resolution normalization processing on the second intermediate voucher image to obtain the preprocessed voucher image.
[0060] Optionally, the target region positioning processing on the preprocessed voucher image to obtain a plurality of target regions comprises: Performing text region segmentation on the preprocessed voucher image to obtain a plurality of segmentation regions; Performing positioning processing on the plurality of segmentation regions to obtain the plurality of target regions.
[0061] Optionally, the text region segmentation on the preprocessed voucher image to obtain a plurality of segmentation regions comprises: Inputting the preprocessed voucher image into a down-sampling processing layer of a preset network model to perform down-sampling processing to obtain a first target feature voucher image; Inputting the target feature voucher image into an up-sampling processing of the preset network model to obtain a second target feature voucher image. According to the second target feature certificate image and the activation function output segmentation probability map of the preset network model; According to the preset threshold and the segmentation probability map, the preprocessed certificate image is segmented to obtain a plurality of segmentation regions.
[0062] Optionally, the plurality of segmentation regions are subjected to positioning processing to obtain a plurality of target regions, including: Coordinates of a bounding box of each segmentation region of the plurality of segmentation regions are obtained; According to the coordinates of the bounding box of each segmentation region, a plurality of target regions are determined.
[0063] Optionally, target information in the plurality of target regions is extracted to obtain a plurality of target information, including: Text information in the plurality of target regions is recognized to obtain a text recognition result; According to the text recognition result and a field in a target information table in a preset database, a plurality of target information are obtained.
[0064] Optionally, the plurality of target information are verified to obtain a verification result, including: Hash values and confidence levels of the plurality of target information are obtained; According to the hash values and the confidence levels of the target information and a verification priority, the plurality of target information are verified to obtain a verification result.
[0065] It should be noted that the device corresponds to the above method, and all implementation manners in the above method embodiments are applicable to this embodiment and can achieve the same technical effects.
[0066] Embodiments of the present application also provide a computing device, including: one or more processors; a storage device for storing one or more programs, when the one or more programs are executed by the one or more processors, so that the one or more processors implement the method as described above. All implementation manners in the above method embodiments are applicable to this embodiment and can achieve the same technical effects.
[0067] Embodiments of the present application also provide a computing device readable storage medium, storing instructions, when the instructions run on a computing device, so that the computing device executes the method as described above. All implementation manners in the above method embodiments are applicable to this embodiment and can achieve the same technical effects.
[0068] Those skilled in the art can clearly understand that the units and algorithm steps of each example described in combination with the embodiments disclosed in the present application can be realized by electronic hardware or a combination of software and electronic hardware of a computing device. Whether the functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.
[0069] Those skilled in the art can clearly understand that, for the convenience and brevity of the description, the specific working processes of the above-described system, device and unit can refer to the corresponding processes in the foregoing method embodiments, which will not be repeated here.
[0070] In the embodiments provided by the present application, it should be understood that the disclosed apparatus and method can be implemented in other ways. For example, the apparatus embodiments described above are merely schematic, for example, the division of the units is only a logical function division, and another division mode can be used in actual implementation, for example, a plurality of units or components can be combined or integrated into another system, or some features can be omitted or not executed. In addition, the coupling or direct coupling or communication connection between the units shown or discussed can be indirect coupling or communication connection through some interface, device or unit, and can be electrical, mechanical or other forms.
[0071] The units described as separate components can or can not be physically separate, and the components shown as units can or can not be physical units, i.e. can be located in one place or can be distributed on a plurality of network units. Part or all of the units can be selected according to actual needs to achieve the purpose of the embodiment.
[0072] In addition, each functional unit in each embodiment of the present application can be integrated into a processing unit, or each unit can exist physically independently, or two or more units can be integrated into one unit.
[0073] If the functions are implemented in the form of software function units and sold or used as independent products, they can be stored in a storage medium readable by a computing device. Based on this understanding, the technical solutions of the present application or the parts that essentially contribute to the prior art or the parts of the technical solutions can be embodied in the form of software products. The computing device software product is stored in a storage medium and includes a plurality of instructions for causing a computing device (which can be a personal computing device, a server, or a network device, etc.) to execute all or part of the steps of the method described in the various embodiments of the present application. The aforementioned storage medium includes a U disk, a mobile hard disk, a ROM, a RAM, a magnetic disk or an optical disk, and various program code storage media.
[0074] In addition, it should be noted that in the device and method of the present application, it is obvious that the components or steps can be decomposed and / or recombined. These decompositions and / or recombinations should be considered as equivalent solutions of the present application. Moreover, the steps of performing the above series of processes can naturally be executed in time sequence according to the order of description, but do not necessarily have to be executed in time sequence. Some steps can be executed in parallel or independently of each other. It can be understood by those skilled in the art that all or any steps or components of the method and device of the present application can be implemented in hardware, firmware, software or a combination thereof in any computing device (including a processor, a storage medium, etc.) or a network of computing devices, which can be implemented by those skilled in the art with basic programming skills after reading the description of the present application.
[0075] Therefore, the object of the present application can also be achieved by running a program or a set of programs on any computing device. The computing device can be a commonly known general-purpose device. Therefore, the object of the present application can also be achieved by merely providing a program product containing program code for implementing the method or device. That is, such a program product also constitutes the present application, and a storage medium storing such a program product also constitutes the present application. Obviously, the storage medium can be any commonly known storage medium or any storage medium developed in the future. It should be noted that in the device and method of the present application, it is obvious that the components or steps can be decomposed and / or recombined. These decompositions and / or recombinations should be considered as equivalent solutions of the present application. Moreover, the steps of performing the above series of processes can naturally be executed in time sequence according to the order of description, but do not necessarily have to be executed in time sequence. Some steps can be executed in parallel or independently of each other.
[0076] The above is the preferred embodiment of the present application. It should be noted that for those skilled in the art, without departing from the principles of the present application, a number of improvements and refinements can be made, which should also be considered within the scope of protection of the present application.
Claims
1. A method for processing voucher images, characterized in that, include: Acquire a voucher image to be processed, the voucher image to be processed including: a target region containing target information; The image of the voucher to be processed is subjected to correction preprocessing to obtain a preprocessed voucher image; The preprocessed voucher image is subjected to target region localization processing to obtain multiple target regions; Target information is extracted from the multiple target regions to obtain multiple target information. The multiple target information items are verified, the verification results are obtained, and then output.
2. The voucher image processing method according to claim 1, characterized in that, The image of the voucher to be processed is subjected to correction preprocessing to obtain a preprocessed voucher image, including: The image of the voucher to be processed is subjected to noise removal processing to obtain a first intermediate voucher image; The first intermediate voucher image is subjected to tilt correction processing to obtain the second intermediate voucher image; The resolution of the second intermediate voucher image is normalized to obtain a preprocessed voucher image.
3. The voucher image processing method according to claim 1, characterized in that, The preprocessed voucher image is subjected to target region localization processing to obtain multiple target regions, including: The preprocessed voucher image is segmented into text regions to obtain multiple segmented regions; The multiple segmented regions are then localized to obtain multiple target regions.
4. The voucher image processing method according to claim 3, characterized in that, The preprocessed voucher image is segmented into text regions to obtain multiple segmented regions, including: The preprocessed voucher image is input into the downsampling processing layer of the preset network model for downsampling processing to obtain the first target feature voucher image; The target feature certificate image is input into the preset network model for upsampling processing to obtain the second target feature certificate image; The segmentation probability map is output based on the second target feature certificate image and the activation function of the preset network model; The preprocessed voucher image is segmented according to a preset threshold and a segmentation probability map to obtain multiple segmented regions.
5. The voucher image processing method according to claim 3, characterized in that, The multiple segmented regions are localized to obtain multiple target regions, including: Obtain the coordinates of the bounding box of each of the multiple segmented regions; Multiple target regions are determined based on the coordinates of the bounding box of each segmented region.
6. The voucher image processing method according to claim 1, characterized in that, Target information is extracted from the multiple target regions to obtain multiple target information items, including: Text information in the multiple target regions is identified to obtain text recognition results; Based on the text recognition results, a fuzzy match is performed with the fields in the target information table of the preset database to obtain multiple target information.
7. The voucher image processing method according to claim 1, characterized in that, The multiple target information items are verified to obtain verification results, including: Obtain the hash value and confidence level of the multiple target information; Based on the hash value, confidence level, and verification priority of the target information, the multiple target information are verified to obtain the verification result.
8. A voucher image processing device, characterized in that, include: The acquisition module is used to acquire a voucher image to be processed, wherein the voucher image to be processed includes a target area of target information; The processing module is used to perform correction preprocessing on the voucher image to be processed to obtain a preprocessed voucher image; perform target region localization processing on the preprocessed voucher image to obtain multiple target regions; extract target information from the multiple target regions to obtain multiple target information; verify the multiple target information to obtain verification results, and output them.
9. A computing device, characterized in that, include: One or more processors; A storage device for storing one or more programs, which, when executed by one or more processors, cause the one or more processors to implement the method as described in any one of claims 1 to 7.
10. A computing device readable storage medium, characterized in that, The computing device readable storage medium stores a program that, when executed by a processor, implements the method as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Financial voucher image identification method and system
CN114782971A
Business voucher auditing method and device
CN117854095A
Bill identification method, apparatus and device, and storage medium
CN118552973A
Voucher information processing device, voucher information processing method, and program
JP2023120861A