A credential image processing method, apparatus and device

By performing correction preprocessing and target area localization on the voucher image, combined with automated verification, the problems of template dependence and low recognition accuracy in existing voucher image processing methods are solved, achieving efficient and accurate voucher information extraction and verification.

CN120976950BActive Publication Date: 2026-04-21BEIJING XINGHAN BONA PHARMACEUTICAL TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
BEIJING XINGHAN BONA PHARMACEUTICAL TECHNOLOGY CO LTD
Filing Date
2025-09-30
Publication Date
2026-04-21

AI Technical Summary

Technical Problem

Existing methods for processing voucher images have high template dependence, low efficiency in adapting to new voucher types, low accuracy in recognizing non-ideal images, low efficiency and high error rate in compliance verification, and are prone to missing duplicate vouchers during manual retrieval.

Method used

By performing corrective preprocessing on the voucher image, including noise removal, tilt correction, and resolution normalization, and combining it with a pre-set network model for target area localization and information extraction, automated verification is used to replace manual operation.

Benefits of technology

It significantly improves the accuracy of target information recognition under complex shooting conditions, enhances verification efficiency, avoids errors and omissions caused by manual operation, adapts to multiple certificate types, and shortens the adaptation time for new certificate types.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120976950B_ABST
    Figure CN120976950B_ABST
Patent Text Reader

Abstract

This invention provides a method, apparatus, and device for processing voucher images. The method includes: acquiring a voucher image to be processed, the voucher image including a target region containing target information; performing correction preprocessing on the voucher image to obtain a preprocessed voucher image; performing target region localization processing on the preprocessed voucher image to obtain multiple target regions; extracting target information from the multiple target regions to obtain multiple target information; verifying the multiple target information to obtain a verification result, and outputting it. This invention can adapt to various voucher types, significantly shortening the adaptation time for new voucher types, and significantly improving processing flexibility and efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present invention relate to the field of image processing technology, and in particular to a method, apparatus and device for processing voucher images. Background Technology

[0002] Voucher images refer to visual image files converted from paper or electronic vouchers through methods such as shooting (using mobile phones, scanners), screenshots, or system export. Their core function is to carry verifiable business information, such as payment vouchers. Existing voucher image processing methods typically extract information using general-purpose OCR (Optical Character Recognition) tools or industry-specific template-based OCR. This results in high template dependence; adding new voucher types requires 1-2 weeks of template reconfiguration, leading to low adaptation efficiency. Furthermore, the recognition accuracy is low for non-ideal images such as tilted, blurry, or shadowed images. Compliance verification in existing voucher image processing methods generally involves manual cross-comparison across multiple systems, which is extremely inefficient, has a high error rate, and is prone to missing duplicate vouchers during manual retrieval. Summary of the Invention

[0003] The technical problem to be solved by the embodiments of the present invention is to provide a method, apparatus and device for processing voucher images that can adapt to multiple voucher types, greatly shorten the adaptation time for new voucher types, and significantly improve processing flexibility and efficiency.

[0004] To solve the above-mentioned technical problems, the technical solution of the present invention is as follows:

[0005] A method for processing voucher images, comprising:

[0006] Acquire a voucher image to be processed, the voucher image to be processed including: a target region containing target information;

[0007] The image of the voucher to be processed is subjected to correction preprocessing to obtain a preprocessed voucher image;

[0008] The preprocessed voucher image is subjected to target region localization processing to obtain multiple target regions;

[0009] Target information is extracted from the multiple target regions to obtain multiple target information.

[0010] The multiple target information items are verified, the verification results are obtained, and then output.

[0011] Optionally, the image of the voucher to be processed is subjected to correction preprocessing to obtain a preprocessed voucher image, including:

[0012] The image of the voucher to be processed is subjected to noise removal processing to obtain a first intermediate voucher image;

[0013] The first intermediate voucher image is subjected to tilt correction processing to obtain the second intermediate voucher image;

[0014] The resolution of the second intermediate voucher image is normalized to obtain a preprocessed voucher image.

[0015] Optionally, the preprocessed voucher image is subjected to target region localization processing to obtain multiple target regions, including:

[0016] The preprocessed voucher image is segmented into text regions to obtain multiple segmented regions;

[0017] The multiple segmented regions are then localized to obtain multiple target regions.

[0018] Optionally, the preprocessed voucher image is segmented into text regions to obtain multiple segmented regions, including:

[0019] The preprocessed voucher image is input into the downsampling processing layer of the preset network model for downsampling processing to obtain the first target feature voucher image;

[0020] The target feature certificate image is input into the preset network model for upsampling processing to obtain the second target feature certificate image;

[0021] The segmentation probability map is output based on the second target feature certificate image and the activation function of the preset network model;

[0022] The preprocessed voucher image is segmented according to a preset threshold and a segmentation probability map to obtain multiple segmented regions.

[0023] Optionally, the multiple segmented regions are subjected to localization processing to obtain multiple target regions, including:

[0024] Obtain the coordinates of the bounding box of each of the multiple segmented regions;

[0025] Multiple target regions are determined based on the coordinates of the bounding box of each segmented region.

[0026] Optionally, target information is extracted from the multiple target regions to obtain multiple target information items, including:

[0027] Text information in the multiple target regions is identified to obtain text recognition results;

[0028] Based on the text recognition results, a fuzzy match is performed with the fields in the target information table of the preset database to obtain multiple target information.

[0029] Optionally, the multiple target information items are verified to obtain verification results, including:

[0030] Obtain the hash value and confidence level of the multiple target information;

[0031] Based on the hash value, confidence level, and verification priority of the target information, the multiple target information are verified to obtain the verification result.

[0032] Embodiments of the present invention also provide a voucher image processing apparatus, comprising:

[0033] The acquisition module is used to acquire a voucher image to be processed, wherein the voucher image to be processed includes a target area of ​​target information;

[0034] The processing module is used to perform correction preprocessing on the voucher image to be processed to obtain a preprocessed voucher image; perform target region localization processing on the preprocessed voucher image to obtain multiple target regions; extract target information from the multiple target regions to obtain multiple target information; verify the multiple target information to obtain verification results, and output them.

[0035] Embodiments of the present invention also provide a computing device, comprising:

[0036] One or more processors;

[0037] A storage device for storing one or more programs that, when executed by one or more processors, cause the one or more processors to perform the method as described above.

[0038] Embodiments of the present invention also provide a computing device readable storage medium storing a program that, when executed by a processor, implements the method described above.

[0039] The above-described solutions of the embodiments of the present invention have at least the following beneficial effects:

[0040] The above-described solution of this invention, through corrective preprocessing of voucher images, effectively improves image quality. Combined with optimized target area localization and information extraction methods, it significantly improves the accuracy of target information recognition under complex shooting conditions. After obtaining the target areas where multiple target information is located, the target information is extracted, and multiple target information are automatically verified, replacing traditional manual operations, greatly improving verification efficiency, and avoiding errors and omissions caused by manual operations, ensuring the reliability of verification results. Moreover, this method for processing voucher images does not rely on specific templates, can adapt to multiple voucher types, significantly shortens the adaptation time for new voucher types, and significantly improves processing flexibility and efficiency. Attached Figure Description

[0041] Figure 1This is a schematic flowchart of the voucher image processing method provided in an embodiment of the present invention.

[0042] Figure 2 This is a schematic diagram of the module of the voucher image processing device provided in an embodiment of the present invention. Detailed Implementation

[0043] Exemplary embodiments of the invention will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of the invention are shown in the drawings, it should be understood that the invention may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided so that this invention will be thorough and complete, and will fully convey the scope of the invention to those skilled in the art.

[0044] like Figure 1 As shown, an embodiment of the present invention provides a voucher image processing method, including:

[0045] Step 11: Obtain the image of the voucher to be processed. The image of the voucher to be processed includes: a target area containing target information; here, the voucher image can be a payment voucher image, including core business information, including target information such as customer name, amount, and transaction number. The location of this target information in the voucher image is the target area.

[0046] Step 12: Perform correction preprocessing on the image of the voucher to be processed to obtain a preprocessed voucher image;

[0047] Step 13: Perform target region localization processing on the preprocessed voucher image to obtain multiple target regions;

[0048] Step 14: Extract target information from the multiple target regions to obtain multiple target information;

[0049] Step 15: Verify the multiple target information, obtain the verification results, and output them.

[0050] In this example, by performing corrective preprocessing on voucher images, image quality can be effectively improved. Combined with optimized target region localization and information extraction methods, the accuracy of target information recognition under complex shooting conditions is significantly improved. After obtaining the target regions where multiple target information is located, the target information is extracted, and multiple target information are automatically verified, replacing traditional manual operations. This greatly improves verification efficiency and avoids errors and omissions caused by manual operations, ensuring the reliability of verification results. Moreover, this method for processing voucher images does not rely on specific templates and can adapt to various voucher types, significantly shortening the adaptation time for new voucher types and significantly improving processing flexibility and efficiency.

[0051] In an optional embodiment of the present invention, in step 11, a voucher image to be processed is obtained, the voucher image to be processed including a target area containing target information.

[0052] Specifically, the target area can be an area that carries core business information, which may include: payer / payee customer name ID, transaction amount, transaction serial number, transaction time, and bank / payment platform logo.

[0053] In this example, the target area is clearly defined as the area that carries core business information. This avoids wasting computing power on irrelevant areas (such as blank areas or decorative patterns) during subsequent image processing and information extraction, thereby improving overall processing efficiency, reducing recognition interference caused by irrelevant information, and indirectly improving the accuracy of core field extraction.

[0054] Clearly defining core business information includes key items such as payer / payee customer name ID and transaction amount, covering the core dimensions required for payment voucher verification. This avoids the problem of missing key information (such as transaction number and bank identification) that may occur in manual verification or general OCR in the background technology, ensuring that subsequent rule verification (such as duplicate verification and amount matching) can be carried out based on complete core information, thereby improving verification accuracy.

[0055] The clear division of target areas and core business information enables subsequent image preprocessing to optimize the quality of core areas in a targeted manner, and information extraction to accurately locate core fields, providing a basis for layered and precise processing at each stage.

[0056] In an optional embodiment of the present invention, step 12, performing correction preprocessing on the image of the voucher to be processed to obtain a preprocessed voucher image, includes:

[0057] Step 121 involves performing noise removal processing on the image of the voucher to be processed to obtain a first intermediate voucher image; specifically, this may include:

[0058] Step 1211, according to Smooth the image of the voucher to be processed;

[0059] in,( x , y () represents the image coordinates of the voucher image to be processed. G ( x , y ) is a two-dimensional Gaussian function used to describe Gaussian filtering in image coordinates ( x , y The weight distribution at position ) is used to smooth the image and reduce high-frequency noise. s The standard deviation of the two-dimensional Gaussian function determines the breadth of the Gaussian distribution;

[0060] Step 1212: Perform median filtering on the smoothed voucher image to obtain the first intermediate voucher image;

[0061] Specifically, the median grayscale value of all pixels within a 3×3 window can be used to replace the center pixel of the window and remove salt-and-pepper noise.

[0062] Step 122 involves performing tilt correction processing on the first intermediate voucher image to obtain a second intermediate voucher image; specifically, this may include:

[0063] Step 1221, according to The first intermediate voucher image is converted into a black and white binary image;

[0064] in, g ( x , y () represents the original grayscale value of the first intermediate voucher image. f ( x , y () represents the binarized pixel value of the first intermediate voucher image. R To set a threshold;

[0065] Step 1222: Perform edge detection on the black and white binary image to extract text edges;

[0066] Specifically, if the pixel gradient value is ≥150: it is determined to be a strong edge and is retained;

[0067] If 50 < pixel gradient value < 150: retain only if connected to a strong edge, otherwise discard;

[0068] If the pixel gradient value is ≤50: it is determined to be a non-edge and discarded;

[0069] Step 1223: Use polar coordinate transformation to detect the straight line tilt angle of the first intermediate voucher image. i ,when i Image tilt is determined when the angle is greater than ±0.5°;

[0070] The distance resolution is 1, and the angular resolution is π / 180.

[0071] Step 1224, if the image is tilted, according to Adjust the image angle to obtain the second intermediate voucher image;

[0072] in, M For rotation matrix, i The angle of inclination of the straight line. dx , day This is the image center offset;

[0073] Step 123 involves performing resolution normalization on the second intermediate voucher image to obtain a preprocessed voucher image; specifically, this may include:

[0074] Step 1231: Based on the original image resolution, different interpolation algorithms are used to unify the image to 300 DPI (dots per inch), 2480×3508 pixels, 24-bit RGB format, to obtain the preprocessed voucher image;

[0075] Specifically, when the resolution is <300 DPI, according to The second intermediate voucher image is subjected to bilinear interpolation to obtain a preprocessed voucher image;

[0076] in, H ( x , y After bilinear interpolation, at coordinates ( x , y The pixel value at ) F ( x , y ) represents the coordinates of the four pixels surrounding the point to be interpolated ( x i , y j The original pixel value at () i , j =0,1, x i , y j Let be the coordinates of the four pixels surrounding the point to be interpolated. w i , w j The weights for bilinear interpolation are determined by the distance between the point to be interpolated and its four surrounding pixels. They are used to calculate the interpolated pixel value and avoid pixel block artifacts.

[0077] When the resolution is greater than 300 DPI, according to Interpolate the second intermediate voucher image to obtain a preprocessed voucher image;

[0078] in, Q ( x This is the interpolation function used when the resolution is greater than 300 DPI, used to calculate the interpolation weights to obtain the preprocessed voucher image. d The normalized distance is the result of normalizing the distance between the point to be interpolated and the relevant pixels.

[0079] In this example, Gaussian filtering and median filtering are used in series for processing. First,... sThe image is smoothed using a 5×5 convolution kernel Gaussian filter with a coefficient of 1.0 to suppress high-frequency noise, and then salt-and-pepper noise is removed using a 3×3 windowed mid-range filter. This results in an image signal-to-noise ratio ≥35dB and text edge preservation ≥98%. This avoids character breakage and misrecognition problems caused by noise in the background technology. For example, complex monetary numbers (such as 1234.56) will not be misjudged as 123.56 or 123456 due to noise, laying the foundation for accurate subsequent recognition.

[0080] Correcting angular deviations and eliminating text line breaks involves four steps: binarization, edge detection, angle detection, and rotation correction. First, the image is converted to a black-and-white binary image using an adaptive threshold. Then, edge detection (low threshold 50, high threshold 150) extracts the text edges. Next, Hough transform is used to detect the tilt angle. Finally, a rotation matrix is ​​applied. M The correction angle is set to ≤±0.3°. This completely solves the text line breakage problem caused by tilted images in the background technology. For example, the tilted payer's name will not be split into two segments due to line breaks and cannot be completely extracted, ensuring the integrity of core fields such as customer name ID.

[0081] To unify image specifications and improve recognition consistency, a differentiated interpolation algorithm is used based on the original resolution: bilinear interpolation is used to avoid pixelation when the resolution is <300 DPI, and interpolation is used to preserve text details when the resolution is >300 DPI. The final output is a unified 300 DPI, 2480×3508 pixel, 24-bit RGB image with a pixel deviation of ≤2 pixels. This solves the problem of inconsistent image resolutions from different shooting devices (mobile phones, scanners) in the background technology, eliminating the need for subsequent recognition models to adapt to multiple specifications. This results in more stable cross-device and cross-scene recognition accuracy, directly supporting the goal of ≥95% cross-bank recognition accuracy.

[0082] In an optional embodiment of the present invention, step 13 involves performing target region localization processing on the preprocessed voucher image to obtain multiple target regions, including:

[0083] Step 131: Perform text region segmentation on the preprocessed voucher image to obtain multiple segmented regions;

[0084] Step 132: Perform localization processing on the multiple segmented regions to obtain multiple target regions.

[0085] In this example, the target area is accurately located by segmentation and positioning, which solves the problem in the background technology that general OCR has no clear area focus and is prone to extracting irrelevant information.

[0086] In an optional embodiment of the present invention, in step 131, the preprocessed voucher image is segmented into text regions to obtain multiple segmented regions, including:

[0087] Step 1311 involves inputting the preprocessed voucher image into the downsampling layer of a preset network model for downsampling processing to obtain the first target feature voucher image; specifically, this may include:

[0088] The preprocessed voucher image (2480×3508 pixels) is cropped into an input block of 512×512×3 (RGB channels), which is then input into a downsampling layer (4 convolutional blocks). Each convolutional block contains convolution + ReLU + pooling operations to extract low-level features of the image, resulting in the first target feature voucher image (the size is gradually reduced and the number of channels is gradually increased).

[0089] Step 1312 involves inputting the target feature voucher image into the upsampling process of the preset network model to obtain a second target feature voucher image; specifically, this may include:

[0090] The first target feature certificate image is input into the upsampling layer (4 upsampling blocks). Each upsampling block contains deconvolution + feature concatenation (fusion with the features of the corresponding downsampling layer) + convolution operation to restore the image size and obtain the second target feature certificate image (the size is the same as the input block, and the number of channels is 1).

[0091] Step 1313: Output a segmentation probability map based on the second target feature credential image and the activation function of the preset network model; specifically, this may include:

[0092] According to the activation function The second target feature certificate image is processed to output the probability that each pixel belongs to the text region, i.e., the segmentation probability map;

[0093] in, S ( x , y ) is the second target feature certificate image in ( x , y The eigenvalue at position ) P ( x , y The probability that each pixel belongs to the text region. P ( x , y )∈[0,1];

[0094] Model training uses a loss function Optimize;

[0095] in, LOSS For loss function, A To predict the mask, B This is a true mask to solve the imbalanced sample problem.

[0096] Step 1314: Segment the preprocessed voucher image according to a preset threshold and a segmentation probability map to obtain multiple segmented regions; specifically, this may include:

[0097] Set the preset threshold to 0.5, when P ( x , y When the value is greater than or equal to 0.5, the pixel is determined to belong to the text region; otherwise, it is the background region, resulting in multiple segmented regions (such as customer name segmented region, amount segmented region, etc.).

[0098] In this example, the text region segmentation process in step 131 achieves accurate extraction of the voucher text region through downsampling, upsampling, probability calculation, and threshold segmentation of the preset network model (U-Net). This fundamentally solves the problems of general OCR in background technology, such as lack of targeted segmentation, easy mixing of background interference, and inaccurate positioning of core regions, providing key support for high accuracy and efficiency of subsequent information extraction.

[0099] The downsampling process first crops the 2480×3508 pixel preprocessed voucher image into a 512×512×3 input block. Then, it is processed through four convolutional blocks containing convolution, ReLU, and pooling. By gradually reducing the image size and increasing the number of channels, low-level image features (such as edges and textures) are efficiently extracted. This step lays the foundation for subsequent feature analysis, avoiding wasted computational power due to excessively large original image sizes. Simultaneously, pooling operations enhance key features and reduce redundant information interference.

[0100] The upsampling process uses four upsampling blocks containing deconvolution, feature concatenation, and convolution to restore the first target feature certificate image obtained from downsampling to the same size as the input block, while reducing the number of channels to 1. The design of fusing features with the corresponding downsampling layer combines the detailed features retained in the downsampling stage with the global features in the upsampling stage, avoiding the loss of details caused by simple upsampling. This ensures accurate restoration of the edges and contours of the text region, providing high-quality feature maps for subsequent pixel-level segmentation.

[0101] In the segmentation probability map generation stage, an activation function is used to transform the feature values ​​of the second target feature voucher image into the probability that each pixel belongs to the text region. Simultaneously, a loss function is employed to optimize model training. This loss function effectively addresses the imbalance between text and background regions in the voucher image by calculating the intersection ratio of the predicted mask (A) and the true mask (B), preventing the model from being biased towards misjudgment due to an excessively high proportion of background regions and improving the accuracy of text region probability prediction.

[0102] Finally, threshold segmentation uses 0.5 as a preset threshold, classifying pixels with a probability ≥ 0.5 as text regions and those with a probability ≤ 0.5 as background regions, ultimately resulting in dedicated segmented regions for customer name, amount, etc. This quantitative judgment method avoids the subjectivity of manual segmentation and, combined with the aforementioned feature processing, ensures that the coordinate error of the payer information field is ≤ 4 pixels and the amount / serial number field is ≤ 2 pixels, achieving a positioning accuracy of 99.1%.

[0103] This process does not rely on traditional OCR templates and does not require reconfiguration when dealing with new voucher types, significantly improving adaptation efficiency. At the same time, precise region segmentation allows subsequent information extraction to focus only on the target area, reducing irrelevant background interference and directly supporting the goal of core field extraction accuracy ≥96%, laying a key foundation for the automation and high precision of the entire voucher verification process.

[0104] In an optional embodiment of the present invention, step 132 involves performing positioning processing on the plurality of segmented regions to obtain a plurality of target regions, including:

[0105] Step 1321: Obtain the coordinates of the bounding box of each of the multiple segmented regions; specifically, this may include:

[0106] For each segmented region, contour detection is used to extract its minimum bounding rectangle (bounding box), and the coordinates of the top-left corner of the bounding box are recorded. x min , y min ) and the coordinates of the lower right corner ( x max , y max ), that is, the bounding box coordinates of each segmented region are [( x min , y min ), ( x max , y max )];

[0107] Step 1322: Determine multiple target regions based on the coordinates of the bounding box of each segmented region; specifically, this may include:

[0108] The segmented regions are filtered according to business rules (such as the bounding box of the amount segmentation area must contain the ¥ symbol, and the customer name segmentation area must be located in the upper half of the image). The segmented regions after filtering are the target regions, and their bounding box coordinates are consistent with the coordinates obtained in step 1321.

[0109] In this example, the minimum bounding rectangle of each segmented region is extracted through contour detection, and the top left corner is recorded. xmin , y min ) With the bottom right corner ( x max , y max Coordinates transform the originally abstract segmented area into a precise coordinate range. This avoids the ambiguity of approximate areas in traditional positioning, allowing subsequent information extraction to accurately focus on the range defined by the coordinates. For example, amount extraction is only performed within [( x min , y min ), ( x max , y max This process is performed within the []] area, reducing the erroneous extraction of surrounding irrelevant text (such as notes and descriptions). At the same time, the coordinate data also provides a standardized data format for subsequent cross-module interactions (such as matching fields in the information extraction layer), improving the efficiency of process connection.

[0110] By identifying the core target area and eliminating invalid interference, the segmented regions are filtered according to business rules (such as the amount area containing the ¥ symbol and the customer name area being located in the upper half of the image). This ensures that the final output target areas are all areas carrying core business information (payer ID, amount, etc.). This solves the problem in the background technology where general segmentation may include irrelevant text areas (such as advertisements and decorative text blocks) in the processing scope, avoiding the waste of computing power in non-core areas. At the same time, by filtering according to rules in advance, invalid identification is reduced during subsequent information extraction, indirectly improving the accuracy of core field extraction. This provides a preliminary guarantee for the target of core field extraction accuracy ≥96% and customer name variant matching accuracy ≥92%.

[0111] In an optional embodiment of the present invention, in step 14, target information is extracted from the plurality of target regions to obtain a plurality of target information, including:

[0112] Step 141 involves recognizing the text information in the multiple target regions to obtain text recognition results; specifically, this may include:

[0113] The CNN+BiLSTM+CTC model is used to recognize the text in the target region, and the text recognition results are obtained.

[0114] CNN Feature Extraction: The image of the target region (e.g., the amount of money) is input into the ResNet50 model. The output of the conv5_x layer is retained, and the number of channels is reduced to 256 through a 1×1 convolution to obtain the visual feature sequence of the text. X =[ x 1, x 2, ..., xT ]( T The length of the feature sequence. x t ∈ R 256 );

[0115] BiLSTM sequence modeling: Input X into a 2-layer bidirectional LSTM (hidden layer dimension 256, dropout=0.3), learn the contextual dependencies of the feature sequence, and output the encoded sequence. Y =[ y 1, y 2, ..., y T ]( y t ∈ R C , C (This refers to the character set size, such as 6000 characters including numbers, letters, and Chinese characters).

[0116] CTC Decoding: Using Beam Search to Solve Feature Sequence Length Issues T With text sequence length L The mismatch problem outputs the text sequence with the highest probability. S =[ s 1, s 2, ... s L (i.e., text recognition results, such as 12345.67 yuan, XX company);

[0117] Step 142: Perform fuzzy matching between the text recognition results and the fields in the target information table of the preset database to obtain multiple target information.

[0118] Specifically, the default customer name database contains three tables (basic customer table: standard name + credit code; historical alias table: customer ID + alias; similarity mapping table: customer ID + similar name + similarity), all of which are indexed to optimize queries;

[0119] Taking customer name recognition results as an example, let the recognition result be... S rec The customer name in the database is S db ;

[0120] Exact match: If S rec = S db If the confidence level is 98%, then the match is successful.

[0121] Fuzzy matching: If S rec ≠S db ,right S rec and S db Perform word segmentation to obtain a word set. C and D ,according to Determine the vocabulary set C and D The similarity, when J ( C , D When the value is ≥0.8, the match is considered successful;

[0122] Among them, when J ( C , D (a collection of words) C and D Similarity;

[0123] Sort by edit distance: LD ( S rec , S db )= minutes {Number of operations}, calculation S rec and S db Minimum number of edits;

[0124] in, LD ( S rec , S db )for S rec and S db The minimum number of edits required, including insert / delete / replace operations;

[0125] when LD ( S rec , S db When the number of candidates is less than or equal to 3, the top 5 candidates are selected as the matching results; for auxiliary fields (such as serial number and time), regular expression matching is used.

[0126] In this example, a CNN+BiLSTM+CTC architecture is adopted. First, visual features of the target region (such as the amount and customer name area) are extracted through the conv5_x layer of ResNet50. Then, two layers of bidirectional LSTM are used to learn the contextual dependencies. Finally, beam search decoding is used to solve the problem of feature mismatch with text sequence length. Compared with general OCR techniques, this architecture can better handle problems such as font differences and character adhesion in voucher text. For example, the amount 12345.67 yuan will not be misjudged as 1234567 yuan due to font blurring. The text recognition accuracy is ≥95% across banks and platforms, providing high-quality initial text data for subsequent matching.

[0127] This approach covers name variations and auxiliary fields to ensure complete information matching. Based on a customer name database with three indexed tables, it processes customer names using a three-layer logic: exact matching, fuzzy matching, and edit distance sorting. Exact matching ensures high confidence (98%) for completely identical names; fuzzy matching with similarity ≥ 0.8 and the top 5 candidates with edit distance ≤ 3 resolves variation matching issues such as "XX Co., Ltd." and "XX Limited Liability Company"; auxiliary fields (serial number, time) are matched precisely using regular expressions. Compared to the single matching logic in the background technology, this approach avoids matching failures caused by customer name variations, ensures no auxiliary fields are missed, and optimizes the index so that a single match takes ≤ 100ms, balancing accuracy and efficiency.

[0128] In an optional embodiment of the present invention, in step 14, target information is extracted from the plurality of target regions to obtain a plurality of target information, including:

[0129] Step 143 involves generating hash values ​​and evaluating the confidence level of the multiple target information items; specifically, this may include:

[0130] Step 1431, according to Y =0.299 r +0.587 g +0.114 b Convert the RGB image of the target area to grayscale;

[0131] in, Y Grayscale value r , g , b These are RGB channel values;

[0132] Step 1432: Perform 16×16 block MD5 preprocessing on the grayscale image; generate a 64-bit hash value using SHA-256;

[0133] Step 1433, according to CF 1= PM 1× KT 1. Determine the confidence level of core fields;

[0134] in, CF 1 represents the confidence level of the core field. PM 1 represents the probability value of the recognition model. KT 1 represents the field integrity coefficient (complete = 1.0, missing suffix = 0.8, missing prefix = 0.7, missing middle = 0.5).

[0135] Step 1434, according to CF 2= PM 2× KT 2. Determine the confidence level of auxiliary fields;

[0136] in, CF 2 represents the confidence level of the auxiliary field. PM 2 represents the regular expression matching score (complete match = 1.0, partial match = 0.5). KT 2 represents the location matching score (target area = 1.0, non-target area = 0.5).

[0137] In this example, the RGB image of the target area is first converted to grayscale using a formula, then preprocessed with 16×16 blocks of MD5 to offset ±5% brightness jitter interference, and finally a 64-bit hash value is generated using SHA-256. This hash value can uniquely identify the characteristics of the target area of ​​the credential. After being stored in a Redis 3 master-3 slave cluster, the query response time is ≤50ms and it supports 1000+ queries / second. This provides accurate basis for the subsequent rule verification of duplicate interception (priority 2), directly supporting a duplicate credential interception rate of ≥99.98%, and completely solving the problem of easy omission of duplicate credentials in manual retrieval in the background technology.

[0138] Information quality is tiered and filtered to balance automation and accuracy. Confidence levels are calculated for core fields (such as customer name and amount) and auxiliary fields (such as transaction number and time). This tiered strategy, with ≥95% automated verification, 80%-94% pending confirmation, and <80% manual review, avoids the inefficiency of fully manual verification in the background technology while preventing high error rates caused by directly verifying low-quality information. Furthermore, the interface response time is ≤50ms, ensuring the efficient operation of the overall verification process.

[0139] In an optional embodiment of the present invention, step 15 involves verifying the plurality of target information to obtain a verification result, including:

[0140] Step 151: Obtain the hash value and confidence level of the multiple target information;

[0141] Specifically, the hash value comes from the 64-bit SHA-256 hash value generated in step 1432, and the confidence score comes from the core field confidence score in step 1433. CF Confidence of auxiliary fields in steps 1 and 1434 CF2;

[0142] Step 152: Verify the multiple target information items based on their hash value, confidence level, and verification priority to obtain verification results; specifically, this may include:

[0143] A configurable verification process is implemented based on a rule engine. The rules are stored in both XML and MySQL. XML defines the verification logic, while MySQL stores the rule metadata. The parser converts the XML format into DRL format and then performs the verification in order of priority from 1 to 5.

[0144] Priority 1 is type verification, which extracts feature points of the bank or payment platform logo in the voucher and compares them with a preset feature library. If the matching degree is not less than 0.75, the logo is deemed valid. At the same time, the TF-IDF (term frequency-inverse document frequency) value of the fixed text in the voucher (such as bank receipt electronic payment voucher) is determined. If the TF-IDF value is not less than 0.8, the text is deemed valid. If both the logo and the text are valid, the voucher type is qualified; otherwise, it is unqualified.

[0145] Priority 2 is for duplicate verification. It queries the hash value stored in the cluster. If the hash value of the target information already exists, it is determined to be duplicate. At the same time, it queries the combined index of serial number and customer ID in MySQL. If the combination already exists, it is also determined to be duplicate. If any condition is met, duplicate credentials will be blocked.

[0146] Priority 3 is customer matching, which involves accurately comparing the customer ID in the order system with the customer ID extracted from the target information. If they match, the customer is a qualified match; otherwise, it is not.

[0147] Priority 4 is amount matching. The amount in the target information is standardized, non-numeric characters are removed and converted to decimal type, and then compared with the order amount. The amount matching is qualified if the absolute error between the two does not exceed 0.01 yuan; otherwise, it is unqualified.

[0148] Priority 5 is for exception interception. If priority 2 determines that the voucher is duplicated, and priority 3 determines that the customer ID does not match and priority 4 determines that the amount does not match, the voucher's identification information, error code, and error message will be written to the log database.

[0149] Finally, processing is based on the verification results and confidence levels: if all priorities 1 to 4 are qualified and the confidence level is not lower than 95%, the verification will be automatically passed; if all priorities 1 to 4 are qualified but the confidence level is between 80% and 94%, they will be marked as pending confirmation; if any priority 1 to 4 fails verification or the confidence level is lower than 80%, the verification will be judged as failed, and the specific reason for failure will be output, such as duplicate voucher, customer ID mismatch, amount mismatch, etc.

[0150] In this example, a dual-storage rule system using XML and MySQL (XML defines the logic, and MySQL stores the metadata) is implemented. The rules are then converted to DRL format by a parser and executed, supporting visual editing and updates. The effective time is ≤30 seconds, and the iteration cycle is ≤2 hours. Compared to the background technology where hard-coded rules require ≥72 hours of iteration, this represents an efficiency improvement of over 36 times. It can quickly respond to new business verification needs (such as adding verification for new payment platform voucher types) without requiring system reconstruction, demonstrating extremely high adaptability.

[0151] By using SIFT feature matching (Logo matching degree ≥ 0.75) + TF-IDF text comparison (value ≥ 0.8), the accuracy rate of voucher type recognition reaches 98.5%, solving the problem of easily confusing voucher types in traditional manual verification (such as misjudging receipts as bank receipts);

[0152] Redis hash value query + MySQL serial number - customer ID composite index query, response ≤85ms, duplicate interception rate ≥99.98%, completely eliminating the risk of omission when manually retrieving duplicate credentials in the background technology;

[0153] Customer ID is accurately compared (100% accuracy), and the error after standardization of amount is ≤0.01 yuan, avoiding character misreading and calculation errors during manual verification and reducing the error rate;

[0154] Automatically log high-risk scenarios where there is no match between customers and amounts, enabling traceability of abnormal behavior and filling the gap in traditional processes where there is no dedicated monitoring of abnormal vouchers.

[0155] Based on the confidence levels of core and auxiliary fields, the results are divided into three categories: automatically approved (≥95%), pending confirmation (80%-94%), and manually reviewed (<80%). This ensures that over 95% of high-confidence credentials do not require manual intervention, significantly improving efficiency (overall interface response ≤220ms). Furthermore, the tiered approach avoids misjudgments caused by excessive automation. Compared to the background technology that relies entirely on manual verification, this reduces labor costs by over 80%, while maintaining a stable verification accuracy of over 99%.

[0156] Example 1

[0157] A company receives an image of a VAT invoice and needs to process it to extract key invoice information, verify it against order data, and determine the invoice's compliance. The image has slight tilt (visible horizontality) and a few shadows and dust spots (salt and pepper noise). Example 1 provides a method for processing voucher images, including:

[0158] Step 21, Obtain the image of the voucher to be processed:

[0159] After receiving the invoice image uploaded by the user, first identify the target areas in the image from which information needs to be extracted, specifically including: buyer information area (including buyer's name, taxpayer identification number, address and telephone number, bank and account number), seller information area (containing the same content as the buyer information area), invoice code area, invoice number area, goods or taxable services name area, specifications and model area, unit area, quantity area, unit price area, amount area, tax rate area, tax amount area, total price and tax area (in uppercase and lowercase), invoice date area, invoice issuer area, review area, and seller (seal) area;

[0160] Step 22, Correction Pretreatment:

[0161] Smoothing: Gaussian filtering is used to smooth high-frequency gray-level fluctuations caused by shadows in the image based on the weight distribution of the two-dimensional Gaussian function, thereby reducing noise interference. The standard deviation of the Gaussian function is set to 1.5, which controls the filtering range and makes the overall gray-level transition of the image smoother.

[0162] Median filtering: The smoothed image is subjected to median filtering in a 3×3 window, that is, the median value of all pixel gray values ​​in each 3×3 pixel window is taken and the gray value of the center pixel of the window is replaced, thereby removing the salt and pepper noise caused by dust in the image during shooting and obtaining the first intermediate certificate image.

[0163] Binarization: Set the grayscale threshold to 128, convert pixels with a grayscale value greater than or equal to 128 in the first intermediate voucher image to white (pixel value 255), and pixels with a grayscale value less than 128 to black (pixel value 0), turning the image into a black and white binary image, highlighting the outline of the text and the invoice border;

[0164] Edge detection: Calculate the gradient value of each pixel in the image and determine the edge type according to the gradient value. If the pixel gradient value is greater than or equal to 150, it is determined to be a strong edge and is directly retained. If the gradient value is between 50 and 150, it is retained only if the pixel is connected to a strong edge, otherwise it is discarded. If the gradient value is less than or equal to 50, it is determined to be a non-edge and is directly discarded. Finally, the edge features of the invoice border and text are extracted.

[0165] Line tilt angle detection: The polar coordinate transformation method is used to detect lines in the image. The distance resolution is set to 1 (i.e., each step distance in polar coordinates is 1 unit) and the angle resolution is π / 180 (i.e., each step angle is 1 degree). After detection, it was found that the tilt angle of the line in the invoice image is 2.3 degrees. This angle is greater than ±0.5 degrees, so it is determined that the image is tilted.

[0166] Angle adjustment: Using the image center as the rotation point (the image size is subsequently normalized to 2480×3508 pixels, and the center coordinates are 1240 pixels horizontally and 1754 pixels vertically), the image angle is adjusted according to the rotation matrix. The tilted image is rotated by -2.3 degrees (to counteract the tilt angle) to restore the invoice to a horizontal state, thus obtaining the second intermediate voucher image.

[0167] First, the original resolution of the second intermediate voucher image was checked and found to be 200 DPI (lower than the target resolution of 300 DPI). Therefore, a bilinear interpolation algorithm was used to adjust the resolution. This algorithm calculates the original pixel values ​​of the four pixels surrounding the point to be interpolated, determines the weights based on the distance between the point to be interpolated and these four pixels, and then calculates the interpolated pixel value based on the weights, thus avoiding pixel block effect after adjustment. Finally, the image was uniformly processed into a preprocessed voucher image of 300 DPI, 2480×3508 pixels, and 24-bit RGB format. At this point, the image is clear and without tilt, meeting the requirements for subsequent processing.

[0168] Step 23, Target Area Location Processing:

[0169] Downsampling: The preprocessed voucher image is cropped into a 512×512 pixel input block containing 3 RGB channels, and input into the downsampling layer of the preset network model. The downsampling layer contains 4 convolutional blocks. Each convolutional block is subjected to convolution (using a 3×3 convolutional kernel with a stride of 1), ReLU activation function processing (enhancing the non-linear expressive ability of the model), and pooling (using a 2×2 pooling window to compress the image size) operations in sequence to gradually extract low-level features of the image (such as edges and textures), and finally obtain the first target feature voucher image (the size is reduced to 32×32 pixels and the number of channels is increased to 256).

[0170] Upsampling Processing: The first target feature certificate image is input into the upsampling processing layer of the preset network model. The upsampling layer contains 4 upsampling blocks. Each upsampling block is subjected to deconvolution (to restore the image size), feature stitching (to fuse the features of the upsampling layer with the features of the corresponding downsampling layer to supplement detailed information), and convolution operations in sequence to gradually restore the image size and finally obtain the second target feature certificate image (the size is restored to 512×512 pixels and the number of channels is 1).

[0171] Segmentation probability map generation: The Sigmoid activation function is used to process the second target feature certificate image, and the probability of each pixel belonging to the text region (the probability value is between 0 and 1) is output to form a segmentation probability map; during the model training stage, the model is optimized by the cross-entropy loss function to solve the problem of sample imbalance (the number of pixels in the text region and the background region is very different) and improve the accuracy of probability calculation.

[0172] Region segmentation: The probability threshold is set to 0.5. Pixels with a probability value greater than or equal to 0.5 in the segmentation probability map are identified as text regions, and pixels with a probability value less than 0.5 are identified as background regions. Finally, 12 segmentation regions are segmented from the preprocessed voucher image, which correspond to the various columns of the invoice.

[0173] Boundary box coordinate acquisition: Perform contour detection on each segmented region to find the smallest rectangle (boundary box) that can completely enclose the region, and record the coordinates of the top left corner (horizontal minimum value, vertical minimum value) and the bottom right corner (horizontal maximum value, vertical maximum value) of each boundary box; for example, the boundary box coordinates of the segmented region corresponding to the buyer's name are (horizontal 300 pixels, vertical 400 pixels) to (horizontal 800 pixels, vertical 460 pixels), and the boundary box coordinates of the segmented region corresponding to the total price and tax (lowercase) are (horizontal 1500 pixels, vertical 2300 pixels) to (horizontal 2000 pixels, vertical 2360 pixels).

[0174] Target area determination: Based on business rules, the 12 segmented areas are filtered. For example, the total price and tax area must contain the ¥ symbol, and the invoice code area must be in 10-digit format. Invalid interference areas that do not meet the rules (such as small blocks segmented from the blank area at the edge of the invoice) are excluded. Finally, 10 target areas are determined, and their bounding box coordinates are consistent with the coordinates of the segmented areas before filtering.

[0175] Step 24, Target Information Extraction:

[0176] A CNN (Convolutional Neural Network) + BiLSTM (Bidirectional Long Short-Term Memory Network) + CTC (Connectivity Temporal Classification) model was used to recognize text information in 10 target regions.

[0177] CNN Feature Extraction: Input the image of each target region (such as the image of the total price and tax area) into the ResNet50 model, retain the output features of the conv5_x layer in the model, and then reduce the number of feature channels to 256 through a 1×1 convolutional kernel to obtain the visual feature sequence of the text (the sequence length is determined according to the size of the target region, for example, 80, and each feature element is a 256-dimensional vector).

[0178] BiLSTM sequence modeling: The visual feature sequence is input into a two-layer bidirectional LSTM network (the hidden layer dimension is 256, and the dropout parameter is 0.3 to prevent the model from overfitting). The network learns the contextual dependencies of the feature sequence (such as the order of words and semantic associations) and outputs an encoded sequence (the sequence length is the same as the visual feature sequence, and the dimension of each encoded element is the same as the size of the character set, which contains 6000 characters including numbers, letters, and Chinese characters).

[0179] CTC Decoding: The beam search algorithm is used to solve the problem of mismatch between the length of the visual feature sequence and the length of the text sequence. The text sequence with the highest probability is selected from the encoded sequence as the text recognition result. For example, the recognition result of the buyer's name is Shenzhen ABC Information Technology Co., Ltd., the taxpayer identification number is 91440300XXXXXXXXXX, the total price including tax (in lowercase) is ¥12960.00, and the invoice number is 00012345.

[0180] Pre-set database preparation: The pre-set customer name database contains 3 tables: the basic customer table (stores the standard customer name and credit code), the historical alias table (stores the customer ID and the customer's previous aliases), and the similarity mapping table (stores the customer ID, similar customer names, and similarity). Indexes are created for the 3 tables to improve query speed.

[0181] Customer Name Matching: Taking the buyer name recognition result Shenzhen ABC Information Technology Co., Ltd. as an example, a precise match is first performed in the basic customer table. It is found that the standard customer name in the database is Shenzhen ABC Information Technology Co., Ltd., which is inconsistent with the database name. The fuzzy matching stage is then entered. The recognition result and the database name are segmented into words. The segmented recognition result yields a word set (Shenzhen, ABC, Information Technology, Co., Ltd.), and the segmented database name yields a word set (Shenzhen, ABC, Information Technology, Co., Ltd.). The similarity between the two word sets (the number of elements in the intersection divided by the number of elements in the union) is calculated. The result is 0.6 (the intersection has 3 elements and the union has 5 elements), which is less than 0.8. Further judgment is needed through edit distance. The minimum number of edits (editing operations include inserting, deleting, and replacing characters) between the recognition result and the database name is calculated. The result is 1 (replacing "technology" with "science and technology"), which is less than or equal to 3. The database name is listed as a candidate result. Then, combined with the taxpayer identification number matching, it is found that the identified 91440300XXXXXXXXXX matches the credit code corresponding to ABC Information Technology Co., Ltd. in the database. Finally, the match is confirmed to be successful.

[0182] Other information processing: The amount information is standardized in format, converting ¥12960.00 to 12960.00 yuan; for auxiliary fields such as serial number and invoice date, regular expression matching is used (e.g., the invoice date must conform to the format XXXX-XX-XX) to confirm the validity of the information, and finally multiple target information (including key content such as buyer, seller, amount, and invoice number) are obtained.

[0183] Grayscale conversion: Based on the correspondence between RGB channel values ​​and grayscale values ​​(grayscale value = 0.299 × red channel value + 0.587 × green channel value + 0.114 × blue channel value), the RGB image of the target area is converted into a grayscale image;

[0184] Hash value generation: The grayscale image is preprocessed into 16×16 pixel blocks (divided into multiple 16×16 pixel blocks with uniform processing size), and then a 64-bit hash value is generated using the SHA-256 algorithm for subsequent duplicate verification. For example, the generated hash value is 5f4dcc3b5aa765d61d8327deb882cf99a61c4195e894f2b.

[0185] Core field confidence assessment: Core fields include buyer's name, taxpayer identification number, amount, and total price including tax. Their confidence level = recognition model probability value (the model's confidence probability of the text recognition result, e.g., 0.97) × field completeness coefficient (complete field information is 1.0, missing suffix content is 0.8, missing prefix content is 0.7, missing middle content is 0.5). In this recognition, the core field information was complete, and the recognition model probability value was 0.97. Therefore, the core field confidence level = 0.97 × 1.0 = 0.97 (i.e., 97%).

[0186] Auxiliary field confidence assessment: Auxiliary fields include invoice number, invoice date, and invoice issuer. Their confidence score = regular expression matching score (1.0 for exact match, 0.5 for partial match) × location matching score (1.0 for information within the target area, 0.5 for information outside the target area). In this case, all auxiliary fields are exact matches and located within the target area, therefore, the auxiliary field confidence score = 1.0 × 1.0 = 1.0 (i.e., 100%).

[0187] Step 25, information verification, obtain the verification result:

[0188] Obtain the 64-bit SHA-256 hash value of the target information, as well as the confidence level of the core field (97%) and the confidence level of the auxiliary field (100%), as the basic data for subsequent verification;

[0189] The system implements a configurable verification process based on a rule engine. Rules are stored using a dual-storage approach: XML (defining verification logic) and MySQL (storing rule metadata). The parser converts the XML format to DRL format before executing the verification.

[0190] Priority 1 (Type Verification): Extract feature points from the seller's (stamp) area in the invoice image and compare them with the preset invoice stamp feature library. If the matching degree is not less than 75%, the stamp is deemed valid. At the same time, calculate the TF-IDF value (term frequency-inverse document frequency, reflecting the importance of the text in the invoice) of the fixed text in the invoice. If the TF-IDF value is not less than 0.8, the text is deemed valid. In this case, the stamp matching degree is 82% and the fixed text TF-IDF value is 85%, both of which meet the requirements, and the voucher type is qualified.

[0191] Priority 2 (Duplicate Verification): Query the hash values ​​of historical vouchers stored in the system cluster. If the hash value of the current target information already exists, it is determined to be a duplicate. At the same time, query the combined index of the invoice number and the buyer's taxpayer identification number in the MySQL database. If this combination already exists, it is also determined to be a duplicate. In this query, it was found that neither the hash value nor the combined index exists, so the voucher is determined to be non-duplicate.

[0192] Priority 3 (Customer Matching): Retrieve the buyer's taxpayer identification number from the ERP system for the corresponding order and compare it precisely with the buyer's taxpayer identification number (91440300XXXXXXXXXX) extracted from the target information. If the two match, the customer matching is qualified.

[0193] Priority 4 (Amount Matching): Standardize the amount (12960.00 yuan) in the target information (remove non-numeric characters and ensure uniform format), and then compare it with the order amount (12960.00 yuan) in the ERP system. If the absolute error between the two is 0 yuan (not exceeding 0.01 yuan), the amount matching is qualified.

[0194] Priority 5 (Abnormal Interception): If Priority 2 determines that the voucher is duplicated, and Priority 3 indicates that the customer ID does not match and Priority 4 indicates that the amount does not match, the voucher identification information, error code, and error message must be written to the log database; if there is no abnormality in this case, no interception is required.

[0195] Step 26, Final Verification Results and Output:

[0196] Based on all verification items, if priority levels 1 to 4 are all qualified and the confidence level of the core fields is 97% or higher than 95%, the voucher is automatically approved. The system outputs the extracted 10 target information (buyer, seller, amount, invoice number, etc.), the verification approval mark, and the generated 64-bit hash value for use in subsequent financial processes.

[0197] This invention clearly defines the target area of ​​the voucher image to be processed as the core business information carrying area (including customer ID, amount, etc.). First, it avoids wasting computing power on blank, decorative, or other irrelevant areas in subsequent processing, directly improving overall efficiency. Second, it eliminates the problem of missing key information such as serial numbers and bank logos in manual verification or general OCR, ensuring that subsequent verification has complete data support. Third, it provides a basis for layered processing at each stage. For example, preprocessing can focus on optimizing the clarity of the amount area, making the processing more targeted.

[0198] Image quality is fundamentally improved through a three-stage processing approach: noise removal uses a Gaussian filter and a median filter in series to achieve a signal-to-noise ratio of ≥35dB, preventing misjudgments of monetary amounts due to noise (e.g., 1234.56 mistakenly interpreted as 123.56); tilt correction, through a four-step process, controls the angle to ≤±0.3°, reducing text line breakage to 0% and ensuring the complete extraction of fields such as customer names; resolution normalization uses bilinear / Lanczos interpolation based on differences, uniformly outputting a 300DPI standard image, resolving the issue of inconsistent output specifications from devices such as mobile phones and scanners, and supporting a cross-bank recognition accuracy of ≥95%.

[0199] The U-Net model is used for segmentation, downsampling to extract features, and upsampling to restore size. Combined with DiceLoss to optimize sample imbalance, the coordinate error of the payer information column is ≤4 pixels, and no special template is required. New voucher types do not need to be reconfigured. The bounding box coordinates are extracted and filtered according to business rules (such as the amount area containing the ¥ symbol), and the abstract area is converted into precise coordinates. This avoids subsequent extraction of irrelevant text such as notes and advertisements. At the same time, it provides standardized data for cross-module interaction and improves the efficiency of process connection.

[0200] Using a CNN+BiLSTM+CTC architecture, it addresses issues such as font differences and character adhesion, achieving a cross-platform text recognition accuracy of ≥95%. It employs precise + fuzzy + edit distance sorting to match customer names, resolving variations such as "Limited Company" and "Limited Liability Company." Auxiliary fields use regular expressions for precise matching, achieving a variation matching accuracy of ≥92%. It generates SHA-256 hash values ​​to support duplicate blocking (response ≤50ms), and calculates confidence levels by field type to provide a basis for tiered processing, preventing low-quality information from affecting verification results.

[0201] The rule engine uses XML+MySQL dual storage, with an iteration cycle of ≤2 hours (36 times more efficient than traditional hard coding), enabling rapid response to business needs; priority levels 1-4 cover type, duplicate, customer, and amount verification, with a duplicate interception rate of ≥99.98%, a customer matching accuracy of 100%, and an amount error of ≤0.01 yuan, eliminating errors and omissions from manual verification; priority level 5 automatically logs scenarios with duplicates and customer / amount mismatches, enabling anomaly tracing; combined with confidence level grading (≥95% automatic approval, 80%-94% pending confirmation, <80% requiring review), over 95% of vouchers do not require manual intervention, reducing labor costs by over 80%, and the overall interface response time is ≤220ms.

[0202] like Figure 2 As shown, embodiments of the present invention also provide a voucher image processing apparatus 20, comprising:

[0203] The acquisition module 21 is used to acquire a voucher image to be processed, wherein the voucher image to be processed includes a target area of ​​target information;

[0204] The processing module 22 is used to perform correction preprocessing on the voucher image to be processed to obtain a preprocessed voucher image; perform target region localization processing on the preprocessed voucher image to obtain multiple target regions; extract target information from the multiple target regions to obtain multiple target information; verify the multiple target information to obtain verification results, and output them.

[0205] Optionally, the image of the voucher to be processed is subjected to correction preprocessing to obtain a preprocessed voucher image, including:

[0206] The image of the voucher to be processed is subjected to noise removal processing to obtain a first intermediate voucher image;

[0207] The first intermediate voucher image is subjected to tilt correction processing to obtain the second intermediate voucher image;

[0208] The resolution of the second intermediate voucher image is normalized to obtain a preprocessed voucher image.

[0209] Optionally, the preprocessed voucher image is subjected to target region localization processing to obtain multiple target regions, including:

[0210] The preprocessed voucher image is segmented into text regions to obtain multiple segmented regions;

[0211] The multiple segmented regions are then localized to obtain multiple target regions.

[0212] Optionally, the preprocessed voucher image is segmented into text regions to obtain multiple segmented regions, including:

[0213] The preprocessed voucher image is input into the downsampling processing layer of the preset network model for downsampling processing to obtain the first target feature voucher image;

[0214] The target feature certificate image is input into the preset network model for upsampling processing to obtain the second target feature certificate image;

[0215] The segmentation probability map is output based on the second target feature certificate image and the activation function of the preset network model;

[0216] The preprocessed voucher image is segmented according to a preset threshold and a segmentation probability map to obtain multiple segmented regions.

[0217] Optionally, the multiple segmented regions are subjected to localization processing to obtain multiple target regions, including:

[0218] Obtain the coordinates of the bounding box of each of the multiple segmented regions;

[0219] Multiple target regions are determined based on the coordinates of the bounding box of each segmented region.

[0220] Optionally, target information is extracted from the multiple target regions to obtain multiple target information items, including:

[0221] Text information in the multiple target regions is identified to obtain text recognition results;

[0222] Based on the text recognition results, a fuzzy match is performed with the fields in the target information table of the preset database to obtain multiple target information.

[0223] Optionally, the multiple target information items are verified to obtain verification results, including:

[0224] Obtain the hash value and confidence level of the multiple target information;

[0225] Based on the hash value, confidence level, and verification priority of the target information, the multiple target information are verified to obtain the verification result.

[0226] It should be noted that this device is a device corresponding to the above method. All implementation methods in the above method embodiments are applicable to this embodiment and can achieve the same technical effect.

[0227] Embodiments of the present invention also provide a computing device, including: one or more processors; and a storage device for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement the method described above. All implementations in the above method embodiments are applicable to this embodiment and can achieve the same technical effect.

[0228] Embodiments of the present invention also provide a computing device readable storage medium storing instructions that, when executed on a computing device, cause the computing device to perform the method described above. All implementations in the above method embodiments are applicable to this embodiment and can achieve the same technical effect.

[0229] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed in this invention can be implemented in electronic hardware, or a combination of computing device software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.

[0230] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0231] In the embodiments provided by this invention, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative. For instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.

[0232] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0233] In addition, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.

[0234] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computing device-readable storage medium. Based on this understanding, the technical solution of this invention, essentially, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computing device software product is stored in a storage medium and includes several instructions to cause a computing device (which may be a personal computing device, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, ROM, RAM, magnetic disks, or optical disks.

[0235] Furthermore, it should be noted that in the apparatus and method of the present invention, it is obvious that the components or steps can be decomposed and / or recombined. These decompositions and / or recombinations should be considered equivalent solutions of the present invention. Moreover, the steps performing the above-described series of processes can naturally be executed in the order described, but are not necessarily required to be executed in chronological order; some steps can be executed in parallel or independently of each other. Those skilled in the art will understand that all or any step or component of the method and apparatus of the present invention can be implemented in any computing device (including processors, storage media, etc.) or network of computing devices, in hardware, firmware, software, or a combination thereof. This is something that those skilled in the art can achieve using basic programming skills after reading the description of the present invention.

[0236] Therefore, the object of the present invention can also be achieved by running a program or a set of programs on any computing device. The computing device can be a known general-purpose device. Therefore, the object of the present invention can also be achieved simply by providing a program product containing program code implementing the method or apparatus. That is, such a program product also constitutes the present invention, and the storage medium storing such a program product also constitutes the present invention. Obviously, the storage medium can be any known storage medium or any storage medium developed in the future. It should also be noted that in the apparatus and method of the present invention, it is obvious that the components or steps can be decomposed and / or recombined. These decompositions and / or recombinations should be considered equivalent to the present invention. Furthermore, the steps performing the above series of processes can naturally be performed in the order described, but are not necessarily required to be performed in chronological order. Some steps can be performed in parallel or independently of each other.

[0237] The above description represents the preferred embodiments of the present invention. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principles of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.

Claims

1. A method for processing voucher images, characterized in that, include: Acquire a voucher image to be processed, the voucher image to be processed includes: a target area containing target information; the target area is an area carrying core business information, the core business information includes: payer / payee customer name ID, transaction amount, transaction serial number, transaction time, and bank / payment platform logo; The image of the voucher to be processed is subjected to correction preprocessing to obtain a preprocessed voucher image; The preprocessed voucher image is subjected to target region localization processing to obtain multiple target regions; Target information is extracted from the multiple target regions to obtain multiple target information. The multiple target information items are verified to obtain the verification results, which are then output. Specifically, the preprocessed voucher image undergoes target region localization processing to obtain multiple target regions, including: The preprocessed voucher image is segmented into text regions to obtain multiple segmented regions; The multiple segmented regions are localized to obtain multiple target regions; The preprocessed voucher image is segmented into text regions to obtain multiple segmented regions, including: The preprocessed voucher image is input into the downsampling processing layer of the preset network model for downsampling processing to obtain the first target feature voucher image; the pixels of the preprocessed voucher image are cropped into input blocks and input into the downsampling layer containing 4 convolutional blocks, each convolutional block containing convolution + ReLU + pooling operations to extract low-level features of the image to obtain the first target feature voucher image; The target feature certificate image is input into the upsampling process of the preset network model to obtain the second target feature certificate image; the first target feature certificate image is input into an upsampling layer containing 4 upsampling blocks, each upsampling block containing deconvolution + feature concatenation + convolution operations to restore the image size and obtain the second target feature certificate image. The segmentation probability map is output based on the second target feature credential image and the activation function of the preset network model; based on the activation function... The second target feature certificate image is processed to output the probability that each pixel belongs to the text region, which is used as a segmentation probability map. in, S ( x , y ) is the second target feature certificate image in ( x , y The eigenvalue at position ) P ( x , y The probability that each pixel belongs to the text region. P ( x , y )∈[0,1]; Model training uses a loss function Optimize; in, LOSS For loss function, A To predict the mask, B This is a true mask to solve the imbalanced sample problem. The preprocessed voucher image is segmented according to a preset threshold and a segmentation probability map to obtain multiple segmented regions; the preset threshold is set to 0.5, when... P ( x , y If the value is greater than or equal to 0.5, the pixel is determined to belong to the text region; otherwise, it is the background region, resulting in multiple segmented regions. The multiple segmented regions are localized to obtain multiple target regions, including: Obtain the coordinates of the bounding box of each of the multiple segmented regions; for each segmented region, extract its minimum bounding rectangle using contour detection, and record the coordinates of the top left corner of the bounding box. x min , y min ) and the coordinates of the lower right corner ( x max , y max ), the bounding box coordinates of each segmented region are [( x min , y min ), ( x max , y max )]; Based on the coordinates of the bounding box of each segmented region, multiple target regions are determined; the segmented regions are filtered according to business rules, and the filtered segmented regions are used as target regions. Target information is extracted from the multiple target regions to obtain multiple target information items, including: The text information in the multiple target regions is identified to obtain text recognition results; the CNN+BiLSTM+CTC model is used to identify the text in the target regions to obtain text recognition results. CNN Feature Extraction: The image of the target region is input into a ResNet50 model, the output of the conv5_x layer is retained, and the number of channels is reduced to 256 through 1×1 convolution to obtain the visual feature sequence of the text. X =[ x 1, x 2, ..., x T ], T The length of the feature sequence; BiLSTM sequence modeling: Input X into a 2-layer bidirectional LSTM to learn the contextual dependencies of the feature sequence and output the encoded sequence. Y =[ y 1, y 2, ..., y T ]; CTC Decoding: Using Beam Search to Solve Feature Sequence Length Issues T With text sequence length L The mismatch problem outputs the text sequence with the highest probability. S =[ s 1, s 2, ... s L ]; Based on the text recognition result, a fuzzy match is performed with the fields in the target information table of the preset database to obtain multiple target information; let the recognition result be... S rec The customer name in the database is S db ; Exact match: If S rec = S db If the confidence level is 98%, then the match is successful. Fuzzy matching: If S rec ≠ S db ,right S rec and S db Perform word segmentation to obtain a word set. C and D ,according to Determine the vocabulary set C and D The similarity, when J ( C , D When the value is ≥0.8, the match is considered successful; Among them, when J ( C , D (a collection of words) C and D Similarity; Sort by edit distance: LD ( S rec , S db )= min {Number of operations}, calculation S rec and S db Minimum number of edits; in, LD ( S rec , S db )for S rec and S db The minimum number of edits required, including insert / delete / replace operations; when LD ( S rec , S db When ≤3, the top 5 candidates are taken as the matching result; regular expression matching is used for auxiliary fields; Target information is extracted from the multiple target regions to obtain multiple target information items, including: Hash values ​​are generated and confidence levels are evaluated for the multiple target information; based on E =0.299 r +0.587 g +0.114 b Convert the RGB image of the target area to grayscale; in, E Grayscale value r , g , b These are RGB channel values; Perform 16×16 block MD5 preprocessing on the grayscale image; generate a 64-bit hash value using SHA-256; according to CF 1= PM 1× KT 1. Determine the confidence level of core fields; in, CF 1 represents the confidence level of the core field. PM 1 represents the probability value of the recognition model. KT 1 represents the field integrity coefficient; according to CF 2= PM 2× KT 2. Determine the confidence level of auxiliary fields; in, CF 2 represents the confidence level of the auxiliary field. PM 2 represents the score for regular expression matching. KT 2 represents the position matching score; The verification of the multiple target information is performed to obtain the verification results, including: Obtain the hash value and confidence level of the multiple target information; Based on the hash value, confidence level, and verification priority of the target information, the multiple target information are verified to obtain the verification result. The rules adopt a dual storage method of XML and MySQL. XML defines the verification logic, and MySQL stores the rule metadata. After the parser converts the XML format into DRL format, it executes the verification in order of priority from 1 to 5. Priority 1 is type verification, which extracts the feature points of the bank or payment platform logo in the voucher and compares them with the preset feature library. If the matching degree is not less than 0.75, the logo is deemed valid. At the same time, the TF-IDF value of the fixed text in the voucher is determined. If the TF-IDF value is not less than 0.8, the text is deemed valid. If both the logo and the text are valid, the voucher type is qualified; otherwise, it is unqualified. Priority 2 is for duplicate verification. It queries the hash value stored in the cluster. If the hash value of the target information already exists, it is determined to be duplicate. At the same time, it queries the combined index of serial number and customer ID in MySQL. If the combination already exists, it is also determined to be duplicate. If any condition is met, duplicate credentials will be blocked. Priority 3 is customer matching, which involves accurately comparing the customer ID in the order system with the customer ID extracted from the target information. If they match, the customer is a qualified match; otherwise, it is not. Priority 4 is amount matching. The amount in the target information is standardized, non-numeric characters are removed and converted to decimal type, and then compared with the order amount. The amount matching is qualified if the absolute error between the two does not exceed 0.01 yuan; otherwise, it is unqualified. Priority 5 is for exception interception. If priority 2 determines that the voucher is duplicated, and priority 3 determines that the customer ID does not match and priority 4 determines that the amount does not match, the voucher's identification information, error code, and error message will be written to the log database. Finally, the verification results and confidence levels are processed accordingly: if all priorities 1 to 4 are qualified and the confidence level is not lower than 95%, the verification will be automatically passed; if all priorities 1 to 4 are qualified but the confidence level is between 80% and 94%, they will be marked as pending confirmation; if any priority 1 to 4 fails the verification or the confidence level is lower than 80%, the verification will be judged as a failure, and the specific reason for failure will be output.

2. The voucher image processing method according to claim 1, characterized in that, The image of the voucher to be processed is subjected to correction preprocessing to obtain a preprocessed voucher image, including: The image of the voucher to be processed is subjected to noise removal processing to obtain a first intermediate voucher image; The first intermediate voucher image is subjected to tilt correction processing to obtain the second intermediate voucher image; The resolution of the second intermediate voucher image is normalized to obtain a preprocessed voucher image.

3. A voucher image processing device, characterized in that, include: The acquisition module is used to acquire a voucher image to be processed. The voucher image to be processed includes a target area of ​​target information. The target area is an area that carries core business information, which includes: payer / payee customer name ID, transaction amount, transaction serial number, transaction time, and bank / payment platform logo. The processing module is used to perform correction preprocessing on the voucher image to be processed to obtain a preprocessed voucher image; perform target region localization processing on the preprocessed voucher image to obtain multiple target regions; extract target information from the multiple target regions to obtain multiple target information; verify the multiple target information to obtain verification results, and output them. Specifically, the preprocessed voucher image undergoes target region localization processing to obtain multiple target regions, including: The preprocessed voucher image is segmented into text regions to obtain multiple segmented regions; The multiple segmented regions are localized to obtain multiple target regions; The preprocessed voucher image is segmented into text regions to obtain multiple segmented regions, including: The preprocessed voucher image is input into the downsampling processing layer of the preset network model for downsampling processing to obtain the first target feature voucher image; the pixels of the preprocessed voucher image are cropped into input blocks and input into the downsampling layer containing 4 convolutional blocks, each convolutional block containing convolution + ReLU + pooling operations to extract low-level features of the image to obtain the first target feature voucher image; The target feature certificate image is input into the upsampling process of the preset network model to obtain the second target feature certificate image; the first target feature certificate image is input into an upsampling layer containing 4 upsampling blocks, each upsampling block containing deconvolution + feature concatenation + convolution operations to restore the image size and obtain the second target feature certificate image. The segmentation probability map is output based on the second target feature credential image and the activation function of the preset network model; based on the activation function... The second target feature certificate image is processed to output the probability that each pixel belongs to the text region, which is used as a segmentation probability map. in, S ( x , y ) is the second target feature certificate image in ( x , y The eigenvalue at position ) P ( x , y The probability that each pixel belongs to the text region. P ( x , y )∈[0,1]; Model training uses a loss function Optimize; in, LOSS For loss function, A To predict the mask, B This is a true mask to solve the imbalanced sample problem. The preprocessed voucher image is segmented according to a preset threshold and a segmentation probability map to obtain multiple segmented regions; the preset threshold is set to 0.5, when... P ( x , y If the value is greater than or equal to 0.5, the pixel is determined to belong to the text region; otherwise, it is the background region, resulting in multiple segmented regions. The multiple segmented regions are localized to obtain multiple target regions, including: Obtain the coordinates of the bounding box of each of the multiple segmented regions; for each segmented region, extract its minimum bounding rectangle using contour detection, and record the coordinates of the top left corner of the bounding box. x min , y min ) and the coordinates of the lower right corner ( x max , y max ), the bounding box coordinates of each segmented region are [( x min , y min ), ( x max , y max )]; Based on the coordinates of the bounding box of each segmented region, multiple target regions are determined; the segmented regions are filtered according to business rules, and the filtered segmented regions are used as target regions. Target information is extracted from the multiple target regions to obtain multiple target information items, including: The text information in the multiple target regions is identified to obtain text recognition results; the CNN+BiLSTM+CTC model is used to identify the text in the target regions to obtain text recognition results. CNN Feature Extraction: The image of the target region is input into a ResNet50 model, the output of the conv5_x layer is retained, and the number of channels is reduced to 256 through 1×1 convolution to obtain the visual feature sequence of the text. X =[ x 1, x 2, ..., x T ], T The length of the feature sequence; BiLSTM sequence modeling: Input X into a 2-layer bidirectional LSTM to learn the contextual dependencies of the feature sequence and output the encoded sequence. Y =[ y 1, y 2, ..., y T ]; CTC Decoding: Using Beam Search to Solve Feature Sequence Length Issues T With text sequence length L The mismatch problem outputs the text sequence with the highest probability. S =[ s 1, s 2, ... s L ]; Based on the text recognition result, a fuzzy match is performed with the fields in the target information table of the preset database to obtain multiple target information; let the recognition result be... S rec The customer name in the database is S db ; Exact match: If S rec = S db If the confidence level is 98%, then the match is successful. Fuzzy matching: If S rec ≠ S db ,right S rec and S db Perform word segmentation to obtain a word set. C and D ,according to Determine the vocabulary set C and D The similarity, when J ( C , D When the value is ≥0.8, the match is considered successful; Among them, when J ( C , D (a collection of words) C and D Similarity; Sort by edit distance: LD ( S rec , S db )= min {Number of operations}, calculation S rec and S db Minimum number of edits; in, LD ( S rec , S db )for S rec and S db The minimum number of edits required, including insert / delete / replace operations; when LD ( S rec , S db When ≤3, the top 5 candidates are taken as the matching result; regular expression matching is used for auxiliary fields; Target information is extracted from the multiple target regions to obtain multiple target information items, including: Hash values ​​are generated and confidence levels are evaluated for the multiple target information; based on E =0.299 r +0.587 g +0.114 b Convert the RGB image of the target area to grayscale; in, E Grayscale value r , g , b These are RGB channel values; Perform 16×16 block MD5 preprocessing on the grayscale image; generate a 64-bit hash value using SHA-256; according to CF 1= PM 1× KT 1. Determine the confidence level of core fields; in, CF 1 represents the confidence level of the core field. PM 1 represents the probability value of the recognition model. KT 1 represents the field integrity coefficient; according to CF 2= PM 2× KT 2. Determine the confidence level of auxiliary fields; in, CF 2 represents the confidence level of the auxiliary field. PM 2 represents the score for regular expression matching. KT 2 represents the position matching score; The verification of the multiple target information is performed to obtain the verification results, including: Obtain the hash value and confidence level of the multiple target information; Based on the hash value, confidence level, and verification priority of the target information, the multiple target information are verified to obtain the verification result. The rules adopt a dual storage method of XML and MySQL. XML defines the verification logic, and MySQL stores the rule metadata. After the parser converts the XML format into DRL format, it executes the verification in order of priority from 1 to 5. Priority 1 is type verification, which extracts the feature points of the bank or payment platform logo in the voucher and compares them with the preset feature library. If the matching degree is not less than 0.75, the logo is deemed valid. At the same time, the TF-IDF value of the fixed text in the voucher is determined. If the TF-IDF value is not less than 0.8, the text is deemed valid. If both the logo and the text are valid, the voucher type is qualified; otherwise, it is unqualified. Priority 2 is for duplicate verification. It queries the hash value stored in the cluster. If the hash value of the target information already exists, it is determined to be duplicate. At the same time, it queries the combined index of serial number and customer ID in MySQL. If the combination already exists, it is also determined to be duplicate. If any condition is met, duplicate credentials will be blocked. Priority 3 is customer matching, which involves accurately comparing the customer ID in the order system with the customer ID extracted from the target information. If they match, the customer is a qualified match; otherwise, it is not. Priority 4 is amount matching. The amount in the target information is standardized, non-numeric characters are removed and converted to decimal type, and then compared with the order amount. The amount matching is qualified if the absolute error between the two does not exceed 0.01 yuan; otherwise, it is unqualified. Priority 5 is for exception interception. If priority 2 determines that the voucher is duplicated, and priority 3 determines that the customer ID does not match and priority 4 determines that the amount does not match, the voucher's identification information, error code, and error message will be written to the log database. Finally, the verification results and confidence levels are processed accordingly: if all priorities 1 to 4 are qualified and the confidence level is not lower than 95%, the verification will be automatically passed; if all priorities 1 to 4 are qualified but the confidence level is between 80% and 94%, they will be marked as pending confirmation; if any priority 1 to 4 fails the verification or the confidence level is lower than 80%, the verification will be judged as a failure, and the specific reason for failure will be output.

4. A computing device, characterized in that, include: One or more processors; A storage device for storing one or more programs, which, when executed by one or more processors, cause the one or more processors to implement the method as described in any one of claims 1 to 2.

5. A computing device readable storage medium, characterized in that, The computing device readable storage medium stores a program that, when executed by a processor, implements the method as described in any one of claims 1 to 2.

Citation Information

Patent Citations

  • Business voucher auditing method and device

    CN117854095A

  • Bill identification method, apparatus and device, and storage medium

    CN118552973A