Financial document intelligent verification method and system

By locating the amount region and decimal point in financial document images, and combining geometric feature extraction and structural difference calculation, the problem of misjudgment of amount data caused by uneven brightness and artifacts in existing technologies has been solved, achieving high-precision intelligent verification and risk assessment.

CN121505641BActive Publication Date: 2026-04-17YOUFU TECHNOLOGY SERVICES (BEIJING) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
YOUFU TECHNOLOGY SERVICES (BEIJING) CO LTD
Filing Date
2025-11-27
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

Existing financial document recognition systems are prone to problems such as uneven brightness, artifact enhancement, and decimal point distortion during image capture and compression, resulting in inconsistencies between the monetary data and the original document information structure, leading to a high misjudgment rate. Traditional methods cannot identify non-malicious structural missing/blurred features, and the review burden is heavy.

Method used

By acquiring invoice image data, locating the amount region and decimal point, and combining geometric feature extraction and structural difference calculation, fine-grained and structured intelligent verification of the amount field is achieved. Multi-scale isolation degree calculation and confidence mapping are used to improve the accuracy of identification and the reliability of risk assessment.

Benefits of technology

In environments with weak strokes, blurred imaging, and local noise, the accuracy of monetary field recognition and the reliability of risk assessment are improved, achieving high robustness and engineering deployability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121505641B_ABST
    Figure CN121505641B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of intelligent verification, and particularly relates to a financial document intelligent verification method and system. The method comprises the following steps: obtaining bill image data; positioning an amount area according to the bill image data to obtain amount area data; recognizing amount characters according to the amount area data to obtain amount character data; positioning a decimal point according to the amount character data to obtain decimal point neighborhood image data; extracting features according to the decimal point neighborhood image data to obtain geometric feature data; calculating amount structure difference according to the geometric feature data to obtain structure difference data; and performing confidence mapping according to the structure difference data to obtain document risk label data. Through image structure difference and confidence mapping, the present application effectively identifies amount recognition abnormalities caused by non-malicious structure loss / fuzziness such as compression and shooting, and improves the accuracy and automation level of financial document image-text consistency verification.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent verification technology, and in particular to an intelligent verification method and system for financial documents. Background Technology

[0002] With the acceleration of corporate financial digitization, a large number of financial documents such as invoices, expense reports, and receipts are uploaded to the system via mobile devices as images, and then entered into the business process after being recognized by the OCR module. However, existing systems generally suffer from image layer noise such as uneven brightness, artifact enhancement, and decimal point distortion during the image capture and compression process, which makes the amount data in the OCR result structurally inconsistent with the original document information. Traditional document verification methods usually rely on the regularity of the amount field or template matching, which cannot identify non-malicious structural missing / blurred structures caused by the image link, resulting in a high false positive rate and a heavy review burden. Summary of the Invention

[0003] To address the aforementioned technical problems, this invention proposes an intelligent verification method and system for financial documents, thereby resolving at least one of the aforementioned technical issues.

[0004] This application provides an intelligent verification method for financial documents, including the following steps:

[0005] S1: Obtain the bill image data; locate the amount area based on the bill image data to obtain the amount area data;

[0006] S2: Recognize the amount characters based on the amount area data to obtain the amount character data; locate the decimal point in the amount character data to obtain the decimal point neighborhood image data;

[0007] S3: Extract features from the decimal point neighborhood image data to obtain geometric feature data;

[0008] S4: Calculate the monetary structure difference based on the geometric feature data to obtain the structure difference data; perform confidence mapping based on the structure difference data to obtain the document risk label data.

[0009] This invention achieves fine-grained, structured intelligent verification of the amount field in financial documents. Unlike traditional methods that rely on direct comparison of character recognition results, this invention employs a composite review of stroke isolation, directionality, and multi-scale geometric texture in the decimal point neighborhood. This effectively resists cross-media imaging disturbances such as image compression, HDR enhancement, and uneven lighting. Structural difference calculation compares the format structure of the image and numerical sides dimension by dimension, enabling the quantification of anomalies and their presentation in a residual mode. Confidence mapping achieves a fusion judgment of local structure and global amount area risk. This method improves the recognition accuracy and risk assessment reliability of the amount field in business scenarios such as weak strokes, blurred imaging, and local noise, and possesses good engineering deployability and versatility.

[0010] Preferably, the location of the amount area is specifically as follows:

[0011] Local brightness detection and artifact compression processing are performed on the ticket image data to obtain local brightness data and artifact compression data, respectively.

[0012] The ticket image data is filtered based on local brightness data and compressed artifact data to obtain image filtering data;

[0013] Based on the image-filtered data, structural line density analysis is performed to obtain the target structural line data;

[0014] Based on the target structural line data, monetary symbols are extracted from the image filtering data to obtain monetary symbol data;

[0015] Based on the numerical spacing analysis of the monetary symbol data, the monetary number column data is obtained;

[0016] The amount range data is constructed by using the amount symbol data and the amount number column data.

[0017] This invention employs dual-channel preprocessing—local brightness detection and compression artifact analysis—during the localization of the amount region. This effectively suppresses local interference caused by uneven lighting, reflections, and JPEG compression during document photography, allowing structural analysis to be built on a more stable data foundation. By performing structural line density analysis on the filtered images, the system can automatically identify the most layout-stable text lines on the document in non-template-based scenarios, thus avoiding the localization errors caused by traditional methods relying on fixed layouts. Combining amount symbol extraction and digit spacing analysis, the system can accurately capture the natural features of the amount field at the image structure level, making region construction not only dependent on character recognition results but also on character distribution patterns. This method maintains high robustness even in document environments with blurriness, tilt, partial occlusion, or compression noise, improving the accuracy and generalization ability of amount region localization.

[0018] Preferably, the structural line density analysis specifically includes;

[0019] Horizontal stroke enhancement is performed on the image-filtered data to obtain horizontal text enhancement data;

[0020] Vertical projection is performed on the horizontally enhanced text data to obtain the structural line density spectrum data;

[0021] Peak filtering is performed on the structure line density spectrum data to obtain structure line filtered data;

[0022] Based on the data filtered by the structural lines, the prior distribution of character height is determined to obtain the target structural line data.

[0023] This invention enhances the horizontal stroke structure of text lines on selected image data, highlighting the unique horizontal stroke structure while suppressing background shading, vertical lines, and blocky noise caused by compression artifacts. This provides a higher signal-to-noise ratio for structural line detection. The structural line density spectrum, formed by vertical projection of the enhancement results, transforms the spatial distribution of text lines into one-dimensional density peaks, allowing text line localization to rely on the image's own statistical characteristics rather than template structure. Peak filtering of the density spectrum effectively filters out false peaks caused by noise, wrinkles, or brightness abrupt changes, improving the stability of structural line recognition. Combining the prior distribution of character heights of candidate structural lines, the system can accurately identify the target structural line containing the amount under different document formats and resolutions, achieving robust text line localization capabilities across scenarios and enhancing the reliability and accuracy of amount region construction.

[0024] Preferably, the amount character recognition specifically includes:

[0025] Text orientation is corrected based on the amount area data to obtain amount text image data;

[0026] Character stroke enhancement is performed on the monetary text image data to obtain text image enhancement data;

[0027] Character connectivity processing is performed on the text image enhancement data to obtain preliminary character block data;

[0028] The initial character block data is sorted and integrated to obtain the character block data.

[0029] Local super-resolution enhancement is performed on the character block data to obtain character-enhanced data;

[0030] The enhanced character data is recognized using a preset dual-channel character recognition model to obtain monetary character data.

[0031] This invention introduces text direction correction during the recognition of monetary characters, effectively avoiding interference from tilt and rotation errors during document photography, ensuring that subsequent stroke analysis is based on consistent direction. Character stroke enhancement strengthens the visibility of digital strokes under complex backgrounds, compressed noise, or low-light blur conditions, improving the stability of character segmentation and extraction. Combining character connectivity processing and sorting integration steps, the system can automatically recover the digital sequence structure without templates, avoiding missegmentation caused by characters being too close together, broken strokes, or sticking together. Local super-resolution enhancement ensures that small characters (such as small, high-resolution numerals commonly found in monetary fields) retain sufficient detail in low-resolution or compressed images, providing stable and reliable input features for the dual-channel character recognition model. By employing a dual-channel recognition model that integrates stroke response and image texture, the accuracy of monetary character recognition in non-standard shooting environments can be improved.

[0032] Preferably, the decimal point positioning is specifically as follows:

[0033] The decimal point image is extracted from the amount character data to obtain the decimal point image data;

[0034] The decimal point position is extracted from the decimal point image data to obtain the decimal point position data;

[0035] Local stroke connectivity is filtered based on decimal point position data to obtain isolated position data;

[0036] Multi-scale isolation degree calculation is performed on isolated location data to obtain decimal point neighborhood image data.

[0037] This invention employs a layer-by-layer, structured detection approach during decimal point localization, effectively addressing challenges such as the small size, susceptibility to blurring, and susceptibility to compression artifacts in financial documents. By extracting the decimal point image based on monetary character data, the search space is limited within the character sequence, reducing the probability of false detections. The subsequent decimal point location extraction step utilizes local image saliency features to accurately locate potential point structures. Local stroke connectivity filtering effectively eliminates false candidate points that are connected to left and right digit strokes, blemishes, or noise, highlighting truly isolated points. Multi-scale isolation calculation is employed, using multi-dimensional features such as direction independence, local energy attenuation, and perturbation stability to measure the isolation consistency of candidate points at different scales, significantly improving the reliability of decimal point recognition under imaging conditions such as low light, blurring, compression artifacts, or excessive HDR enhancement.

[0038] Preferably, the multi-scale isolation degree calculation is specifically as follows:

[0039] Multi-scale local variance suppression convolution is performed on isolated location data to obtain stability feature data;

[0040] Local energy decay calculations are performed on the stability characteristic data to obtain local energy decay data;

[0041] The isolation probability data is obtained by calculating the isolation probability based on the local energy decay data and stability characteristic data;

[0042] Cross-scale similarity data is obtained by performing cross-scale similarity calculations based on isolated probability data;

[0043] The isolated location data corresponding to the maximum value in the cross-scale similarity data is determined as the decimal point neighborhood image data.

[0044] This invention utilizes multi-scale isolation degree calculation to cross-validate the stability of suspected decimal points at different spatial scales, effectively distinguishing between imaging noise, compression block effects, and local stroke adhesion. Multi-scale local variance suppression convolution highlights the local consistency characteristics of point structures, ensuring that true decimal points exhibit similar texture stability across multiple scale windows; while noise points and compression artifacts often fluctuate drastically at different scales. Based on this calculated local energy decay characteristic, the energy decay trend of candidate points during scale magnification can be characterized. True decimal points exhibit rapid and smooth decay, while pseudo-points maintain strong directionality or block structures. Through isolation probability modeling and cross-scale similarity calculation, consistency analysis of response trends at each scale is achieved, ensuring that the selected points not only satisfy local isolation but also cross-scale structural stability.

[0045] Preferably, S3 specifically comprises:

[0046] Local binarization is performed on the decimal point neighborhood image data to obtain local binary map data;

[0047] Isolated stroke connectivity analysis was performed on local binary map data to obtain isolated feature data.

[0048] Relative geometric feature data is obtained by extracting relative geometric features from the decimal point neighborhood image data based on the amount character data;

[0049] Texture microstructure features are extracted from the decimal point neighborhood image data to obtain texture microstructure feature data;

[0050] The isolated feature data, relative geometric feature data, and texture microstructure feature data are integrated to obtain geometric feature data.

[0051] This invention achieves a multi-dimensional and accurate characterization of the decimal point's morphological features by sequentially performing local binarization, isolated connectivity analysis, relative geometric feature extraction, and texture microstructure analysis within the decimal point's neighborhood. Local binarization preserves key stroke structures under varying lighting conditions and unclear, unconventional backgrounds while suppressing background noise, laying the foundation for stable geometric calculations. Isolated stroke connectivity analysis extracts structural attributes unique to genuine decimal points, such as single connected components, small areas, and high isolation, effectively distinguishing false points caused by digit adhesion, compression artifacts, or blemishes. Combining monetary character data with relative geometric feature extraction allows the system to capture the spatial relationship between the decimal point and the digits on either side, verifying its authenticity through dimensions such as positional consistency, alignment error, and horizontal spacing. Texture microstructure features characterize the impact of imaging quality on the decimal point's morphology using indicators such as high-frequency energy, directional entropy, and local roughness.

[0052] Preferably, the calculation of the difference in monetary structure is as follows:

[0053] Numerical side structure data is obtained by constructing the numerical side structure based on geometric feature data and preset amount format rules;

[0054] Image side structure data is obtained by constructing image side structure data based on geometric feature data.

[0055] The structural residual data are obtained by performing dimension-wise difference calculations based on the numerical and image-based structural data.

[0056] Multi-scale structural differencing is performed on the structural residual data to obtain structural differencing data.

[0057] This invention introduces dual-domain modeling of numerical and image-side structures during the structural differencing stage, enabling structural consistency verification of the amount field from two dimensions: format rules and imaging features. The numerical-side structure, constructed based on preset amount format rules, accurately reflects the theoretical constraints of amount text, such as decimal point position, number of digits, and spacing patterns. The image-side structure, generated based on real geometric features, represents the actual imaging form of the document under photographing, compression, or lighting variations. The structural residual data obtained through dimension-wise differencing quantifies the structural deviations on both sides into comparable numerical indicators, freeing anomalies from subjective judgment. Multi-scale structural differencing amplifies weak anomalies and reduces local noise, allowing for a comprehensive evaluation of the stability of structural deviations at different scales.

[0058] Preferably, the confidence mapping is as follows:

[0059] Prior probability data of the structure is obtained by fitting prior probability data based on structural difference data;

[0060] The risk level of the monetary region is calculated based on the prior probability data of the structure, and the risk level data of the monetary region is obtained.

[0061] Confidence calculations are performed based on prior structural probability data and risk level data for monetary regions to obtain document risk label data.

[0062] This invention achieves a hierarchical and quantifiable comprehensive judgment of the risk of monetary fields by introducing prior probability fitting, regional risk calculation, and global confidence fusion in the confidence mapping stage. After prior probability fitting, the structural difference data yields prior probabilities of different structural patterns (such as normal, imaging noise, or potential tampering), enabling the system to infer the possibility of anomalies from local geometric deviations. Based on the regional risk calculated from these prior probabilities, the overall imaging credibility of the monetary region is evaluated by combining character stability, stroke consistency, and regional texture variation, avoiding misjudgments caused by relying solely on local features. The system constructs a global confidence score by fusing structural prior probabilities and regional risk, achieving a clear distinction between normal, suspicious, and abnormal states.

[0063] Preferably, this application also provides a financial document intelligent verification system for executing the financial document intelligent verification method described above, the financial document intelligent verification system comprising:

[0064] The amount area positioning module is used to acquire ticket image data; and to locate the amount area based on the ticket image data to obtain the amount area data.

[0065] The decimal point positioning module is used to recognize monetary characters based on monetary area data to obtain monetary character data; and to locate the decimal point in the monetary character data to obtain decimal point neighborhood image data.

[0066] The geometric feature extraction module is used to extract features from the decimal point neighborhood image data to obtain geometric feature data.

[0067] The document risk label mapping module is used to perform monetary structure difference calculation based on geometric feature data to obtain structure difference data; and to perform confidence mapping based on the structure difference data to obtain document risk label data.

[0068] The beneficial effects of this invention are as follows: it achieves high-precision intelligent verification of the amount field in financial documents. In the amount region positioning stage, the system combines brightness distribution and compression artifact features to filter images, and accurately determines the area where the amount is located through structural line density analysis and sign / digit spacing resolution. In the character recognition and decimal point positioning stages, stroke enhancement, connected component analysis, and multi-scale isolation degree modeling effectively solve the industry problems of small decimal point size, easy blurring, and susceptibility to compression block interference, thereby ensuring the effective extraction of the core elements of the amount structure. In the feature extraction stage, through multi-source fusion of isolated structures, relative geometry, and texture microstructures, the system can comprehensively characterize the true shape of the decimal point and imaging perturbation features. Through amount structure difference calculation, the system compares the true structure on the image side with the amount format on the numerical side dimension by dimension, and combines confidence mapping to form quantifiable risk labels, achieving effective differentiation of imaging noise, format anomalies, and potential tampering. Attached Figure Description

[0069] Other features, objects, and advantages of this application will become more apparent from the following detailed description of the non-limiting embodiments, taken with reference to the accompanying drawings:

[0070] Figure 1 A flowchart illustrating the steps of an intelligent verification method for financial documents according to one embodiment is shown.

[0071] Figure 2 A flowchart illustrating the steps of a method for locating a monetary region according to an embodiment is shown.

[0072] Figure 3 A flowchart illustrating the steps of a method for recognizing monetary characters according to an embodiment is shown.

[0073] Figure 4 A flowchart illustrating the steps of a geometric feature extraction method according to an embodiment is shown.

[0074] Figure 5 A flowchart illustrating the steps of a method for calculating the difference in monetary structure according to an embodiment is shown. Detailed Implementation

[0075] The technical method of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. Based on the embodiments of this invention, all other embodiments obtained by those skilled in the art without inventive effort are within the scope of protection of this invention.

[0076] Furthermore, the accompanying drawings are merely illustrative of the invention and are not necessarily drawn to scale. The same reference numerals in the drawings denote the same or similar parts, and therefore repeated descriptions of them will be omitted. Some block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities. These functional entities can be implemented in software, in one or more hardware modules or integrated circuits, or in different network and / or processor methods and / or microcontroller methods.

[0077] It should be understood that although the terms "first," "second," etc., may be used herein to describe various units, these units should not be limited by these terms. These terms are used merely to distinguish one unit from another. For example, without departing from the scope of the exemplary embodiments, a first unit may be referred to as a second unit, and similarly, a second unit may be referred to as a first unit. The term "and / or" as used herein includes any and all combinations of one or more of the associated listed items.

[0078] Please see Figures 1 to 5 This application provides an intelligent verification method for financial documents, including the following steps:

[0079] S1: Obtain the bill image data; locate the amount area based on the bill image data to obtain the amount area data;

[0080] Specifically, the system acquires ticket image data, which can come from scanning devices or mobile devices. For the acquired images, the system performs preprocessing such as brightness normalization, rotation correction, and size standardization. The system uses a fixed-size sliding window to calculate the mean brightness and brightness fluctuations in local image regions. The system converts the image to the compression domain (e.g., JPEG DCT coefficients) and analyzes the energy distribution characteristics of high-frequency coefficients in the compressed image. If the high-frequency energy in a certain local area of ​​the image is significantly higher, or if the periodic distribution is abnormal, it indicates the presence of compression artifacts. The system traverses the image region using a sliding window, calculating the local variance and spectral flatness of the compression residual signal for each region. If the values ​​exceed a preset threshold, it is identified as a compression artifact region and marked as a non-priority analysis region. After the basic image quality is controlled, the system performs structural line density analysis on the image. For example, the system uses a horizontal stroke-enhanced convolution method to highlight the structural features in the text line direction, and then performs vertical integration statistics on the enhanced image to form a structural line density distribution map. The system identifies structural peaks with sufficient height and width to meet the minimum text line requirement in the structural line density distribution map. Simultaneously, it refers to a preset range of document character heights (system-preset parameters determined based on empirical data or industry / document-preset specifications) to determine the target structural line containing monetary text. After determining the target structural line, the system extracts the monetary symbol within the corresponding image region and analyzes the monetary number column. The system identifies monetary symbols such as "¥" and "¥" using template matching and uses these as reference points to extract consecutive numerical blocks from their right sides. Subsequently, the system filters out numerical sequences that meet the rules based on the interval variation between adjacent numbers, the consistency of stroke height, and possible decimal point positions, thus determining the row / column of the monetary number. Using the located monetary symbols and numerical columns as references, the system extracts the upper, lower, left, and right boundaries of the monetary region. The system determines the horizontal range of the monetary line based on the pixel positions of the leftmost and rightmost characters and adds appropriate buffers in the vertical direction to construct the final monetary region data.

[0081] S2: Recognize the amount characters based on the amount area data to obtain the amount character data; locate the decimal point in the amount character data to obtain the decimal point neighborhood image data;

[0082] Specifically, the system performs text direction correction based on the amount region data. For example, the system identifies the main direction of the amount line and calculates the text tilt angle accordingly. It also performs edge detection on the image (using the Canny operator) to extract obvious stroke edges. Subsequently, it uses Hough transform to detect straight line structures in the image. The system calculates the main direction angle distribution of the detected straight line set and selects the angle with the highest frequency as the main direction of the amount line, calculating the tilt angle accordingly. Based on the calculated text tilt angle, the system performs rotation correction on the amount region image to maintain the digit strokes in a standard horizontal arrangement. After the text direction is calibrated, the system performs character stroke enhancement processing on the corrected amount region data. For example, the system strengthens the main stroke structure of the digits through horizontal and vertical stroke enhancement convolution kernels and performs binarization processing combined with a preset threshold to extract the connected components of the characters. For regions that are too small or have abnormal shapes, the system identifies them as noise and deletes them, thus obtaining preliminary character candidate blocks. After obtaining the character candidate blocks, the system arranges them into a character sequence according to the horizontal coordinate order of the characters in the image. To address the issues of blurriness or insufficient resolution in some characters, the system performs local super-resolution enhancement on each character block, such as using a pre-defined lightweight image enhancement network or interpolation method. The system employs a dual-channel character recognition model. The model's input includes two channels: the first channel is the enhanced character image (the image region corresponding to the character sequence); the second channel is the stroke response map (obtained by calculating the horizontal and vertical grayscale change rates based on the image region corresponding to the character sequence), including gradient information in both directions. After fusing the two types of information, the model outputs a sequence of monetary character data. For example, for numbers in a monetary region, the model can recognize continuous monetary strings and output them as monetary character data.

[0083] The construction steps of the pre-defined lightweight image augmentation network are as follows: The input is the initially detected candidate image patch of characters. The system first performs normalization and size standardization processing on the input image (e.g., 32×32 pixels); then it is input into a lightweight network consisting of 3 convolutional modules. Each convolutional layer is followed by a ReLU activation function and pooling operation to extract the texture and edge features of the characters; the system uses mean squared error as the loss function and high-resolution images as supervision signals for training; during the training process, the system uses a paired sample set containing real character images and simulated blurred character images for learning; the output is the augmented character image.

[0084] The dual-channel character recognition model is constructed as follows: the first channel is the historical character image after the aforementioned enhancements, and the second channel is the historical stroke response map. The stroke response map is derived from the extraction of image gradient features. Specifically, the system performs Sobel horizontal and vertical gradient calculations on the historical character images to generate horizontal and vertical edge response maps. The model structure adopts a dual-branch convolutional network: each channel undergoes several layers of convolution processing, followed by feature fusion in the intermediate layer, and then outputs the probability distribution of character categories through a fully connected layer. The final output is a sequence of character recognition results, i.e., monetary character data.

[0085] After recognizing the complete sequence of monetary characters, the system locates the decimal point within the amount. For example, based on the position of the "." character in the character recognition result, the system crops the corresponding image block from the monetary area image as a candidate image region for the decimal point. The system then performs image filtering on this image block to extract existing candidate decimal point regions. For instance, by detecting local minima, the system identifies isolated point structures in the image and uses them as preliminary candidate decimal point locations. The system performs connectivity analysis on the small area surrounding each candidate location. For example, the system focuses on determining whether there is a connection between the point and the surrounding digit strokes. If there is a significant stroke connection in the left-right direction of the candidate point, such as the number of consecutive pixels exceeding a preset threshold, it is judged as a pseudo-decimal point and discarded; otherwise, the point is retained as an isolated point location as a valid candidate decimal point region. The system analyzes the isolation degree of candidate regions for effective decimal points at different scales. For example, the system evaluates the isolation characteristics of candidate regions for effective decimal points at multiple scale windows (such as small, medium, and large neighborhood ranges), including indicators such as variance suppression of the local image, energy distribution characteristics, and structural similarity at different scales. The system selects the scale window with the most significant isolation characteristics (preferably the scale window with the lowest variance, the most concentrated energy distribution, and the highest structural similarity), and extracts the image content corresponding to this window region as decimal point neighborhood image data.

[0086] S3: Extract features from the decimal point neighborhood image data to obtain geometric feature data;

[0087] Specifically, the system performs image processing and feature extraction operations based on the decimal point neighborhood image data to construct geometric feature data describing the decimal point structure. For example, the system performs threshold segmentation on the decimal point neighborhood image to generate a local binary image independent of imaging illumination. Based on the local binary image, the system analyzes the connectivity of isolated strokes in the image to obtain isolation features. For example, the system counts the number of connected regions / connected domains in the neighborhood (the system uses an 8-neighbor connected domain extraction algorithm to count the number of clusters of foreground pixels in the binary image, i.e., the number of different connected domains), the area of ​​each connected domain (by traversing the number of pixels contained in each connected domain), and the compactness of each connected domain (e.g., using roundness calculation, the area divided by the square of the perimeter), thereby determining whether there are artifacts generated during the compression process in the region, or abnormal adhesion structures between strokes.

[0088] The system combines the identified monetary character data to extract the relative geometric features of the decimal point in space, including the vertical offset of the decimal point relative to the center line of the digits (the system calculates the vertical difference between the centroid coordinates of the decimal point and the midline positions of the upper and lower boundaries of the adjacent digit regions), the horizontal distance between the decimal point and the digits on the left and right (the system calculates the nearest distance between the center of the decimal point and the boundary of the connected region of the adjacent digits on the left and right), and whether there is a very close distance between the decimal point and the adjacent digit strokes (the system performs a pixel-by-pixel nearest point search between the boundary of the decimal point and the boundary of the connected digit region to obtain the distance), in order to determine whether it may be stuck to the digit strokes. Beyond structural features, the system also extracts the texture microstructure of the decimal point's neighborhood to obtain texture-like features. These include local region orientation entropy (the system statistically analyzes the gradient direction distribution in the image, constructs an orientation histogram, and calculates Shannon entropy); energy values ​​of high-frequency residual signals (the system performs wavelet transform or high-pass filtering on the image and statistically calculates the accumulated energy value of high-frequency channels); and gray-level difference features within the neighborhood (the system constructs a gray-level co-occurrence matrix and extracts statistical features of gray-level differences between adjacent pixels, such as average and maximum differences). The system integrates / concatenates these isolated features, relative geometric features, and texture-like features, concatenating them in a fixed order and dimension to generate a decimal point geometric feature vector.

[0089] S4: Calculate the monetary structure difference based on the geometric feature data to obtain the structure difference data; perform confidence mapping based on the structure difference data to obtain the document risk label data.

[0090] Specifically, based on the extracted geometric feature data, the system analyzes the structural rationality of the amount region and generates risk label data for judging the reliability of the invoice. For example, the system constructs an ideal numerical structure model according to common amount format rules, including, for instance, at least one significant digit before the decimal point, usually two digits of precision after the decimal point, and the amount digits should have highly consistent character height and alignment characteristics. This model is encoded as a structural reference vector. The system constructs an image-side structural description vector based on the aforementioned geometric feature data. This image vector contains multiple key dimensions, such as the roundness and compactness of the decimal point or number region / connected region in the image, the alignment deviation between characters and the decimal point, the directional entropy of the image, and whether there are stroke adhesions between characters (the horizontal spacing of the character connected regions is less than a set threshold or there are connecting bridges). These indicators collectively represent the actual amount structure characteristics exhibited by the image.

[0091] After the structural model is constructed, the system compares the ideal structure (generated by the preset monetary structure template rules and transformed into a reference structure vector) with the image structure dimension by dimension, extracts the degree of difference between the two, and generates structural residual information / structural difference data.

[0092] The system performs multi-scale analysis on these residual information, extracting indicators including local maximum deviation, overall residual fluctuation level, sensitivity to changes in high-frequency features in the image, and isolation deviation of the decimal point region in the structure. These indicators are then combined into a measure / feature of the structural difference. Alternatively, the system prioritizes threshold judgment based on "alignment deviation of the decimal point region" and "consistency of character height": if the decimal point position deviates significantly from the central axis of the main characters (deviation exceeding 0.5 times the character height) or the standard deviation of the character height exceeds 25% of the average character height, the monetary structure is deemed abnormal, and a corresponding risk label is assigned. For example, if the "decimal alignment deviation" is greater than 0.5 times the character height, or the standard deviation of "character height consistency" is greater than 25% of the average height, it is marked as suspicious; if multiple unexpected sticky regions appear in the structure (such as character spacing < preset minimum threshold and the existence of connecting bridges), or the residual energy value of the decimal point neighborhood is significantly higher than the average value of its surrounding region (such as exceeding its outer neighborhood by 1.5 to 2 times), it indicates a structural abrupt change or texture compression artifact (within a 3030-pixel local area centered on the decimal point, the system calculates the high-frequency residual signal energy value of the decimal point neighborhood (such as a 9×9 pixel window) / the local energy square sum, and compares it with the average energy value of the remaining pixels in the local area excluding the 9×9 decimal point neighborhood. If the former exceeds the latter by 1.5 to 2 times, the system determines that there is a structural abrupt change or texture compression artifact at that location), it is judged as abnormal; if all differential indicators are within a reasonable range, or only have slight deviations (not exceeding the threshold), it is marked as normal. On the other hand, the system can also use rule-based trees or lightweight neural networks to map multiple structural difference features into probability values. For example, a lightweight classifier built using XGBoost or LightGBM can be used as training samples to assign a risk score to the current amount range and map it to a risk label. The system outputs the risk label data for each invoice, which is used for classification filtering, manual review marking, and review priority ranking in the automatic processing flow.

[0093] Preferably, the location of the amount area is specifically as follows:

[0094] S11: Perform local brightness detection and artifact compression processing based on the ticket image data to obtain local brightness data and artifact compression data respectively;

[0095] Specifically, the system performs local brightness detection and compression artifact processing on the acquired ticket image to evaluate the image quality and support subsequent recognition steps. For example, the system divides the entire image into several sliding window units of 3232 pixels each. For each window unit, the system calculates the average brightness and brightness fluctuation of its internal pixels, thereby generating a local brightness distribution map of the entire image. The system also performs frequency domain feature analysis on each window unit. By detecting the trend of image changes in the frequency domain, the system estimates the degree of compression artifacts present in the current region. For example, the system evaluates the magnitude of changes in frequency coefficients in the image and generates compression artifact data accordingly.

[0096] S12: Filter the ticket image data based on local brightness data and compressed artifact data to obtain image filtering data;

[0097] Specifically, the system uses the obtained local brightness data and compressed artifact data to perform quality screening on the original ticket image. For example, the system sets judgment thresholds for brightness stability and artifact energy to determine whether each region in the image meets the quality requirements. For regions where brightness fluctuations exceed the set thresholds or where the intensity of compressed artifacts is significantly higher than normal, the system marks them as interference regions and filters them out. The system then obtains image screening data.

[0098] S13: Perform structural line density analysis based on the image-filtered data to obtain the target structural line data;

[0099] Specifically, the system performs density analysis on structural lines in the ticket image based on image screening data to identify target structural lines containing monetary information. For example, the system uses horizontal edge enhancement, processing the image by applying a horizontal edge enhancement convolution kernel to strengthen the horizontal stroke features in the image. The system performs integral projection on the processed image along the vertical direction to generate a structural line density spectrum. The system detects local peak positions in this structural line density spectrum through a preset sliding window, filtering out candidate text lines with high structural density. The system sets that the peak height must exceed a preset multiple of the average density and requires the peak width to be within a reasonable range to exclude non-text areas or interfering lines. For the initially filtered candidate text lines, the system extracts the character connected components in their adjacent regions and calculates the average height and fluctuation of these characters. The system preferentially selects structural lines with character heights within a reasonable range (e.g., between [10, 30] pixels) and relatively uniform height distribution (mean / fluctuation less than 0.3) as valid text structural lines, marking the structural lines that meet the conditions as target structural lines.

[0100] S14: Extract monetary symbols from the image filtering data based on the target structural line data to obtain monetary symbol data;

[0101] Specifically, based on the obtained target structural line data, the system locates and extracts monetary symbols from the image filtering data. For example, within the target structural line region, the system uses template matching to search for image blocks similar to preset monetary symbol shapes (e.g., "¥" or "¥"). During the matching process, the system identifies possible monetary symbol regions by calculating the similarity between the image blocks and the template. The system uses a normalized image similarity calculation method and sets a corresponding matching threshold. When the matching degree of a candidate region reaches or exceeds the threshold, the system marks it as a monetary symbol. The monetary symbol is located on the far left of the monetary field, with a set of numeric characters immediately adjacent to its right. Therefore, the system not only relies on the shape matching results but also combines its relative position within the structural line for a comprehensive judgment, thereby extracting monetary symbol data.

[0102] S15: Analyze the number spacing based on the amount symbol data to obtain the amount number column data;

[0103] Specifically, the system uses the extracted monetary symbol data as a benchmark and identifies the corresponding numeric character column in the area to its right. For example, in the target structural line area to the right of the monetary symbol, the system extracts connected character blocks that meet preset area range and aspect ratio conditions as candidate numeric regions. The system sorts these character blocks according to their horizontal position in the image and calculates the horizontal spacing between adjacent characters. The system analyzes whether these spacings are within a reasonable range and judges the consistency of the candidate character heights. If the spacing between all adjacent characters is stable and the character height changes little, it indicates that these characters are arranged regularly, have a standardized layout format, and meet the structural characteristics of a typical monetary numeric column. Under the premise that all the above conditions are met, the system identifies the current area as monetary numeric column data.

[0104] S16: Construct the amount range based on the amount symbol data and the amount number column data to obtain the amount range data.

[0105] Specifically, based on the identified monetary symbol data and monetary number column data, the system constructs the monetary region in the document. For example, the system uses the left edge of the monetary symbol in the image as the left boundary of the monetary region, and the right edge of the rightmost digit in the numerical column as the right boundary, thus determining the horizontal range of the monetary region. Vertically, the system sets appropriate upper and lower buffer areas based on the average height of the identified characters, which are several times the character height, such as 1.2-1.5. Through this method, the system constructs the monetary region data.

[0106] Preferably, the structural line density analysis specifically includes;

[0107] Horizontal stroke enhancement is performed on the image-filtered data to obtain horizontal text enhancement data;

[0108] Specifically, the system performs stroke enhancement processing on the horizontal text structure in the ticket based on image screening data to obtain horizontal text enhancement data. If the image is in color format, the system converts it to a grayscale image. The system applies a horizontal edge enhancement operator to the grayscale image and extracts the gradient information in the horizontal direction of the image through convolution operation to form a horizontal edge response map. The system performs absolute value processing on the obtained edge response map and normalizes it to a uniform range of pixel values ​​to generate horizontal text enhancement data.

[0109] Vertical projection is performed on the horizontally enhanced text data to obtain the structural line density spectrum data;

[0110] Specifically, the system performs vertical integral projection based on the horizontal text enhancement data to generate structural line density spectrum data, which represents the horizontal stroke intensity distribution of each horizontal row in the image. For example, the system scans the image line by line along the vertical direction, calculates the sum of the intensity values ​​of all pixels in each line, and forms a numerical sequence representing the stroke density at different vertical positions, thus obtaining structural line density spectrum data.

[0111] Peak filtering is performed on the structure line density spectrum data to obtain structure line filtered data;

[0112] Specifically, after constructing the structure line density spectrum, the system filters local peaks in the spectrum to extract representative structure line locations, such as by setting the size of a sliding window. For each row position, the system determines whether its density value is higher than the average value in its neighborhood and compares it with a set threshold to identify potential structure line peaks. The system also sets a minimum span constraint, requiring that the continuous high-density region corresponding to the peak has a certain pixel height in the vertical direction, not less than the minimum height standard of a preset text line. The system sets a minimum peak spacing to ensure sufficient vertical separation between adjacent peaks. The system uses the row coordinates obtained through the above rules as the structure line filtering data.

[0113] Based on the data filtered by the structural lines, the prior distribution of character height is determined to obtain the target structural line data.

[0114] Specifically, after obtaining preliminary structural line screening data, the system combines the statistical characteristics of character height to determine the final target structural line data. For example, the system extracts an image band with a vertical height of approximately 10 to 15 pixels near each candidate structural line and performs adaptive binarization processing on this image band (e.g., using the Otsu method). The system performs connected component analysis in this binary image, extracts each character connected component, and statistically analyzes its vertical height information. For example, the system statistically analyzes the height values ​​of these character blocks and calculates the average height and the degree of height fluctuation. The system sets a priori rule range for character height, such as requiring that the character height be within the range of typical printed text and that the degree of fluctuation in the height distribution be small. When the character height distribution under a certain structural line meets the above rules, the system obtains the structural line screening data.

[0115] Preferably, the amount character recognition specifically includes:

[0116] S21: Correct the text orientation based on the amount area data to obtain the amount text image data;

[0117] Specifically, the system performs text orientation correction processing on the monetary area image to obtain monetary text image data with standardized layout. For example, the system applies an edge detection algorithm to the monetary area image to extract the edge contours of character strokes; based on these edge contours, the system performs line detection to extract the main horizontal text boundary lines in the image and calculates the average tilt angle of the overall text; when a significant tilt is detected in the image (e.g., the tilt angle exceeds a set threshold), the system performs inverse rotation correction on the image through affine transformation to restore the text to a horizontal arrangement. The rotation operation uses the image center as a reference point to ensure that the content is centered and the structure is not destroyed. During the rotation process, the system performs zero-value filling or mirror expansion processing on blank areas that exceed the original image range. The system outputs monetary image data with corrected text orientation.

[0118] S22: Perform character stroke enhancement on the monetary text image data to obtain text image enhancement data;

[0119] Specifically, the system performs character stroke enhancement processing on the corrected monetary text image data. For example, the system applies edge enhancement filters in the horizontal and vertical directions to extract the gradient response of the main strokes of the characters in both directions. The system integrates the gradient results in the two directions, superimposes their response values ​​to form a stroke response map, performs local contrast stretching on the stroke response map, and combines median filtering (such as a 3×3 filter window) to suppress background noise. In scenarios where it is necessary to enhance the character edge endpoints or stroke transition areas, the system performs Laplacian enhancement processing. The system then obtains the enhanced text image data.

[0120] S23: Perform character connectivity processing on the text image enhancement data to obtain preliminary character block data;

[0121] Specifically, the system performs character connectivity processing on the enhanced text image data to extract candidate character regions. For example, the system uses an 8-neighborhood connected component extraction algorithm to perform connectivity analysis on the foreground region in the binary image, thereby identifying spatially continuous pixel blocks, which are the initial candidate character regions. To eliminate non-character interference caused by noise, ink smudges, or printing defects, the system performs geometric screening on all extracted connected components, filtering out regions with excessively small areas or significantly abnormal aspect ratios. Connected components with areas smaller than a set threshold or with excessively flat or elongated shapes are identified as pseudo-character blocks and discarded. The system outputs a set of preliminary character block data, where each character block contains its position coordinates in the image, width and height information, and corresponding pixel mask data.

[0122] S24: Sort and integrate the initial character block data to obtain the character block data;

[0123] Specifically, the system sorts and integrates the extracted preliminary character block data to construct a character block sequence. For example, the system sorts each character block from left to right based on its horizontal position in the image. The system identifies character blocks that are too close to each other, such as adjacent character blocks with a spacing less than a certain proportion (e.g., 0.3) of their own width, and merges them into a single character unit. The system also judges large character blocks with abnormal areas. If the aspect ratio of a character block significantly deviates from the normal character shape (e.g., aspect ratio greater than 2), or its area is much larger than other character blocks, the system performs a segmentation operation, dividing it into two or more smaller character units. For example, the system calculates the pixel projection curve of the large character block in the horizontal direction, finds the valley points or minimum value regions in the curve, and prioritizes setting segmentation boundaries at these locations, thus naturally separating the connected characters. Through the above merging and segmentation strategies, the system outputs a set of character block data.

[0124] S25: Perform local super-resolution enhancement on the character block data to obtain character-enhanced data;

[0125] Specifically, the system performs local super-resolution enhancement processing on the sorted character block data. For example, based on the position of each character block in the original image, the system crops the corresponding region and uniformly scales it to a preset standard size, such as 32×32 or 48×48 pixels. After image scaling, the system uses a lightweight super-resolution reconstruction network to enhance the character image, such as using image super-resolution models like ESRGAN-lite or SRCNN. For scenarios with limited computing resources or high real-time requirements, the system can also use traditional image processing paths, such as a combination of bicubic interpolation and sharpness enhancement filters. The system outputs the processed character enhancement data.

[0126] S26: The character-enhanced data is recognized using a preset dual-channel character recognition model to obtain monetary character data.

[0127] Specifically, the system performs recognition operations on the enhanced character image data, outputting monetary character results using a pre-defined dual-channel character recognition model. This recognition model includes two input channels, each receiving different types of character information. The first is an image channel, which receives super-resolution processed character image blocks to provide the overall morphological features of the character. The second is a stroke channel, which receives the stroke response map of the corresponding character region (i.e., text image enhancement data), such as an edge map generated by a Sobel filter. The recognition model can employ a combination of a lightweight convolutional neural network structure and a fully connected classifier, such as a simplified LeNet variant, or a more efficient architecture like MobileNet combined with an attention mechanism, to adapt to the computational resource and recognition accuracy requirements of different scenarios. The model outputs a classification label for each character, covering numeric characters (0–9), decimal points, and common monetary symbols such as “¥” and “¥”, along with corresponding recognition confidence scores. For recognition results with low confidence, the system can mark the character as uncertain and include it in the candidate set for further processing. This can be achieved by implementing an N-Best candidate re-ranking strategy or marking it as an object for manual review, thus ensuring the accuracy and reliability of the recognition results. The system outputs monetary character data.

[0128] Preferably, the decimal point positioning is specifically as follows:

[0129] The decimal point image is extracted from the amount character data to obtain the decimal point image data;

[0130] Specifically, based on the identified monetary character data, the system extracts the decimal point region from the image to obtain decimal point image data for verification and precise positioning. For example, the system locates the index of the character identified as a decimal point (“.”) in the recognition sequence and extracts the position information of the corresponding image block, including its coordinate position and width and height data in the monetary region image. Using this character block as the center, the system sets the image cropping window size around it at a certain multiple (e.g., 1.5 to 2 times) of the character height, cropping the corresponding area from the monetary region image to generate preliminary decimal point image data. If the system does not explicitly identify the decimal point character in the character recognition result, it infers it based on the length of the monetary character and common monetary format rules. For example, if the recognition result contains 6 characters and usually has two decimal places, the system can automatically determine the logical insertion position of the fourth decimal place and calculate the estimated decimal point position accordingly, then extract the corresponding area from the image as preliminary decimal point image data.

[0131] The decimal point position is extracted from the decimal point image data to obtain the decimal point position data;

[0132] Specifically, the system processes the extracted decimal point image data to identify the specific location of the decimal point. For example, the system first performs a Gaussian filter on the image to eliminate local noise interference by setting appropriate smoothing parameters. The system uses a filtering method based on the Laplacian of Gaussian operator to extract dark response regions in the image, or directly searches for grayscale minima in the image to capture the decimal point location. The system sets a fixed-size local neighborhood window around each pixel, such as a 3×3 pixel region, and detects points with significantly negative response values ​​within it, using these as candidate decimal point locations. For each candidate decimal point location, the system calculates its local peak response index (calculating the average difference between its value and the surrounding pixel values) and filters out strong response points based on a set empirical threshold, retaining only the pixel coordinates that meet preset conditions. The system outputs the center coordinates of the pixels that meet the response requirements as the decimal point location data.

[0133] Local stroke connectivity is filtered based on decimal point position data to obtain isolated position data;

[0134] Specifically, the system performs local connectivity analysis on the acquired candidate decimal point positions to filter out truly independent decimal point positions. For example, the system constructs a fixed-size neighborhood window region, such as 11×11 pixels, centered on each candidate decimal point, and performs adaptive binarization processing (e.g., using the Otsu method) on this region to obtain a binary image. The system applies an 8-neighborhood connectivity analysis algorithm to this binary image to determine whether the current candidate point is connected to other foreground pixels in its neighborhood. The system counts the number of connected bridges between the point and its surrounding pixels. If a candidate point is connected only in one direction, or is completely isolated and not connected to any other pixels, the system classifies it as a structurally isolated point and retains it as a valid candidate decimal point position. Conversely, if a candidate point has a horizontal stroke connection with the characters on the left and right, the system classifies it as character fragments, adhering noise, or other misidentified regions and removes it to prevent it from being misidentified as a valid decimal point. The system obtains a set of decimal point position data as isolated position data.

[0135] Multi-scale isolation degree calculation is performed on isolated location data to obtain decimal point neighborhood image data.

[0136] Specifically, the system performs multi-scale isolation analysis on the selected isolated decimal point candidate locations to determine the neighborhood image data that best represents the characteristics of the real decimal point. For example, the system constructs multiple neighborhood window regions of different scales around each candidate location, with sizes including 5×5, 9×9, and 13×13 pixels. The system extracts the central image patch corresponding to each scale from the original image as the analysis object. For each scale window, the system performs the following three feature evaluations in sequence: The system calculates the variance of pixel gray values ​​within the window region. The real decimal point region has low gray value fluctuations, and the window region with a variance of less than a preset gray value threshold is determined as the matching region; The system constructs multiple concentric ring-shaped regions from the center outward within the window and compares the image energy distribution of the inner and outer regions. For example, the energy of the real decimal point is usually concentrated at the center and decays rapidly at the outer edge. The system calculates the local gradient direction distribution of the image within the region, forms a direction histogram, and evaluates its entropy value. If the direction distribution is relatively uniform, it indicates that the region lacks obvious stroke structure and is more likely to be a dot-like rather than a non-linear character, i.e., the entropy value is greater than the preset threshold. If multiple scales meet the conditions, the smallest matching block (e.g., 5×5) can be selected first to improve positioning accuracy; or the image block that meets the most conditions can be selected to obtain decimal neighborhood image data.

[0137] Preferably, the multi-scale isolation degree calculation is specifically as follows:

[0138] Multi-scale local variance suppression convolution is performed on isolated location data to obtain stability feature data;

[0139] Specifically, the system performs multi-scale perturbation analysis on the selected isolated locations and extracts stability feature data through local variance suppressed convolution operations. For example, the system extracts multiple neighboring image patches of different scales, such as 5×5, 9×9, and 13×13 pixels, centered on the coordinates of each isolated location. For each image patch, the system performs local variance suppressed convolution processing. This convolution process models the deviation between the current pixel and its neighborhood average by constructing a specific filter kernel. , For scale Variance-suppressed convolution kernels, This is the unit impulse function, which is 1 when (x,y)=(0,0) and 0 otherwise, used to retain the current pixel value. This represents the side length of the image patch at the current scale (e.g., n=5,9,13). After convolution, the system takes the absolute value of the output to obtain the scale response map. The system calculates the mean and variance of each scale response map. The mean represents the overall perturbation response intensity of the region, and the variance represents the fluctuation of its local distribution. The system combines these two values ​​as the stability feature data at that scale.

[0140] Local energy decay calculations are performed on the stability characteristic data to obtain local energy decay data;

[0141] Specifically, the system obtains multi-scale stability feature data and performs local energy decay calculations on the stability feature data at each scale. For example, the system divides the image into two regions, using the center pixel of each scale image block as a reference: a sub-block region located at the center (e.g., 3×3 pixels in size), and a ring-shaped edge region surrounding it. The system calculates the accumulated energy of the pixel responses in these two regions respectively, representing the distribution of image energy in the central and edge regions. , The sum of the squares of the stability feature data of all pixels within the central region is the energy. These are the pixel coordinates of a sub-region centered on the center pixel within the image patch. This represents the square of the stability feature data at coordinates (x, y) at scale s. The sum of the squares of the stability feature data of all pixels within the edge region is the energy. This refers to the pixel coordinates of the edge band surrounding the central region (e.g., the portion remaining after removing the center). By comparing the energy ratios of the two regions, the system evaluates whether the image patch exhibits a characteristic of energy concentration towards the center. If the energy in the central region is significantly higher than that at the edges, it indicates that the region has point-like concentration characteristics, consistent with the true representation of decimal points in an image; conversely, if the energy is widely distributed in the edge region, it is due to character stroke stretching, adhesion, or background interference resulting in a non-decimal point structure. The system uses this ratio as local energy attenuation data at the current scale.

[0142] The isolation probability data is obtained by calculating the isolation probability based on the local energy decay data and stability characteristic data;

[0143] Specifically, the system calculates the isolation probability of candidate decimal point positions based on stability feature data and local energy decay data at each scale. For example, the system integrates the feature parameters corresponding to each scale into a combined vector, which contains three key dimensions: the mean value in the scale response map, the variance in the scale response map, and the energy decay ratio between the center and edge regions. The system inputs the above feature vector into a preset probability fitting model, which can be a logistic regression model or a trained shallow neural network. The model has learned the statistical characteristics of typical decimal points in image structures through prior data and makes a probabilistic judgment on whether a candidate point is a true decimal point. The model outputs the probability value that the candidate position is a true decimal point at that scale, as the isolation probability data.

[0144] The model is trained by using a large number of labeled decimal point image regions as training samples during the training phase. Each sample contains known true decimal point locations (positive samples) and several non-decimal point interference locations (negative samples). Feature vectors (such as average response value, variance, and energy decay ratio) consistent with those extracted during the testing phase are extracted from these candidate points. These feature vectors, along with their corresponding labels (whether it is a true decimal point), are then input into the training model. If a logistic regression model is used, the system fits the parameters by maximizing the likelihood function, ensuring the output probability is closest to the true label of the sample. If a shallow neural network is used, it typically includes one or two fully connected network layers, using cross-entropy as the loss function, and combining backpropagation and gradient descent algorithms for parameter optimization. After training, the model can predict the probability that a candidate point is a true decimal point based on the input feature vectors.

[0145] Cross-scale similarity data is obtained by performing cross-scale similarity calculations based on isolated probability data;

[0146] Specifically, based on the multi-scale isolated probability data obtained in the previous step, the system evaluates the consistency of the determination of candidate decimal points at different scales, thereby obtaining cross-scale similarity data. For example, the system performs trend analysis on the sequence of isolated probability values ​​corresponding to each scale. If the sequence shows a stable decreasing trend as the scale increases, or reaches a peak at a medium scale and then slowly decays, it indicates that the candidate point exhibits consistent point-like characteristics at different scales and has strong structural stability. Conversely, if the probability value fluctuates greatly between different scales, such as showing a significant increase followed by a decrease, or a non-monotonic pattern of low followed by high, it indicates that the structural characteristics of the point are unstable at different scales and are greatly affected by surrounding noise or stroke interference. The system constructs a consistency index, namely... , As a consistency indicator, For scale indexing, This represents the total number of scales. For the first Isolated probability values ​​at each scale For the first The system uses the isolation probability value at each scale. A larger value indicates that the judgment results at different scales are more similar and the structural trend is more stable; conversely, a smaller value indicates that there are large differences in cross-scale judgments and the credibility decreases. The system uses this consistency index as cross-scale similarity data.

[0147] The isolated location data corresponding to the maximum value in the cross-scale similarity data is determined as the decimal point neighborhood image data.

[0148] Specifically, based on cross-scale similarity data corresponding to multiple isolated locations, the system determines image regions that conform to decimal point characteristics as decimal point neighborhood image data. For example, among all candidate isolated locations, the system compares their corresponding cross-scale similarity index values ​​and prioritizes the candidate location with the largest index value. Based on the selected location, the system finds the isolation probability value corresponding to that location at different scales and selects the scale with the highest probability value as the optimal receptive region scale. Using this location as the center, the system extracts image patches of the corresponding size at the selected scale (e.g., extracting images at a 9×9 scale), which are then used as decimal point neighborhood image data.

[0149] Preferably, S3 specifically comprises:

[0150] S31: Perform local binarization on the decimal point neighborhood image data to obtain local binary map data;

[0151] Specifically, the system performs local binarization processing on the decimal point neighborhood image data to generate image data that can be used for stroke and background separation. If the current image has not yet been converted to grayscale, the system converts it to a grayscale image. The system performs local binarization processing on the image. For example, within the current decimal point neighborhood image block, the system sets a sliding window (e.g., the window size is 1 / 3 of the image width). For each pixel, a local window region is extracted centered on that pixel, and the weighted average grayscale value of that region is calculated. The weights follow a two-dimensional Gaussian distribution (the central pixel has the largest weight, gradually decreasing towards the edges). This weighted average grayscale value is used as the binarization threshold for the current pixel. If the current pixel's grayscale value is below the threshold, it is set to foreground (stroke) 1; otherwise, it is set to background 0. Alternatively, the system dynamically adjusts the threshold calculation window size based on the width of the image block, setting it to one-third to one-half of the image width. Based on the above, the system obtains local binary image data.

[0152] S32: Perform isolated stroke connectivity analysis on local binary map data to obtain isolated feature data;

[0153] Specifically, based on local binary image data, the system performs connectivity analysis on isolated strokes in the foreground region of the image to extract isolated feature data. For example, the system uses 8-neighborhood connected component extraction to identify connected regions in the foreground pixels of the binary image, obtaining all spatially continuous pixel blocks. Among the extracted connected blocks, the system filters based on shape and position features, prioritizing candidate regions that conform to the shape of a decimal point. These include regions with small areas, compact outlines, a width-to-height ratio close to 1 (i.e., approximately circular), and whose center of gravity is located near the center of the image, such as regions within approximately 3 pixels above, below, to the left, and to the right of the center coordinates. These types of regions are more likely to correspond to real decimal points rather than character strokes or background noise. For connected blocks that meet the above conditions, the system extracts their geometric and morphological indicators such as area, roundness, compactness, and boundary complexity, and integrates these data into isolated feature data.

[0154] S33: Extract relative geometric features from the decimal point neighborhood image data based on the amount character data to obtain relative geometric feature data;

[0155] Specifically, based on the identified monetary character data, the system analyzes the relative spatial relationship of the current decimal point in the character sequence to extract relative geometric feature data. For example, the system determines the character index position corresponding to the current decimal point and, combined with its context information in the monetary character sequence, obtains the positions of its left and right adjacent character blocks. The system extracts the center coordinates of the left and right character blocks and calculates their geometric relationship with the centroid position of the decimal point. The system focuses on the following three features: first, the angle formed by the lines connecting the decimal point to the centers of the left and right characters; second, the horizontal spacing ratio between the decimal point and the left and right characters; and third, the relative offset of the decimal point in the vertical direction. Under normal business conditions, the decimal point should be located in the lower center between two digit characters, with a small vertical offset, a near-vertical angle, and a horizontal distance ratio that should not be too large (less than 0.4). The system uniformly constructs the above geometric quantities into relative geometric feature data.

[0156] S34: Extract texture microstructure features from the decimal point neighborhood image data to obtain texture microstructure feature data;

[0157] Specifically, the system performs texture microstructure analysis on the decimal point neighborhood image data to extract texture microstructure feature data. For example, the system performs gradient filtering in the horizontal and vertical directions within the image region (e.g., using the Sobel operator) to obtain the edge response of the image in different directions, thereby extracting the image gradient. Based on the image gradient, the system constructs a local gray-level co-occurrence matrix to statistically analyze the spatial distribution of gray values ​​in the image and extracts a series of texture feature indicators, including energy value, contrast, homogeneity, and information entropy. These indicators can effectively reflect whether the regional texture is concentrated, orderly, and whether there are abnormal interference patterns. As an alternative or enhancement method, the system can also directly construct an oriented gradient histogram (HOG) to partition and statistically analyze the edge directions in the image, dividing them into 4 to 8 angular channels to obtain directional features. Through the above processing, the system can determine whether the current decimal point neighborhood has the structural feature of "concentrated point texture," thereby eliminating unnatural interference caused by stains, image compression artifacts, or character adhesion. The system combines the extracted texture statistics and directional features into a texture microstructure feature vector.

[0158] S35: Integrate isolated feature data, relative geometric feature data, and texture microstructure feature data to obtain geometric feature data.

[0159] Specifically, the system fuses and combines the isolated feature data, relative geometric feature data, and texture microstructure feature data extracted in the previous steps to generate comprehensive geometric feature data for structural analysis. For example, the system standardizes various features, including adjusting each dimension to zero mean and unit variance. After standardization, the system concatenates and integrates these three sets of features to form a unified feature vector. The generated feature vector serves as the geometric feature data.

[0160] Specifically, isolated feature data, relative geometric feature data, and texture microstructure feature data are concatenated and quantized to obtain geometric feature data.

[0161] Preferably, the calculation of the difference in monetary structure is as follows:

[0162] S41: Construct the numerical side structure based on the geometric feature data and the preset amount format rules to obtain the numerical side structure data;

[0163] Specifically, based on the extracted geometric feature data and preset amount format rules, the system constructs standardized numerical structure data. For example, the system parses the composition of the amount string based on the position, aspect ratio, and character category information of each character extracted from the amount recognition results, forming a preliminary amount character sequence from left to right. The system sets preset amount writing specifications as matching rules, including structural combinations such as "integer digits plus a decimal point plus two decimal places" and "thousands separated by commas." The system matches each recognized character against these rules one by one, identifying its role in the amount structure, such as whether it is a number, a decimal point, or a separator, and constructs a standardized structural label sequence accordingly, for example, expressing structural features in the form of [number, number, number, decimal point, number, number]. For characters in the recognition results that are ambiguous, unrecognized, or have low confidence, the system uses placeholders to replace the actual characters and marks the corresponding confidence level, forming a structural vector with uncertainty annotations. The system outputs the numerical structure data.

[0164] S42: Construct the image-side structure based on the geometric feature data to obtain the image-side structure data;

[0165] Specifically, based on the generated geometric feature data, the system constructs a structural representation of the image. For example, the system extracts core structural information for each character block from the geometric feature data, including attributes such as the coordinate position of the character's centroid, stroke density, relative spacing, and isolation degree. The system constructs a character arrangement graph, connecting the character blocks in the image sequentially from left to right horizontally to form a chain-like graph structure representation. The system organizes the multi-dimensional structural indicators extracted from the graph into an image-side structural vector, and maintains consistency in dimensionality settings with the numerical-side structure.

[0166] S43: Perform dimension-wise difference calculations based on numerical and image-based structural data to obtain structural residual data;

[0167] Specifically, the system compares the numerical structural data and the image structural data dimension by dimension, calculating the degree of difference between them at the structural level to obtain structural residual data. For example, the system compares the attribute items of the same dimension in the two sets of structural vectors one by one, calculating residual values ​​according to different feature types. For coordinate information representing character positions, the system calculates the offset between the center position of the character in the image and its expected position in the numerical structure to indicate whether the character deviates from the standard layout. For character size indicators, such as height or width, the system compares the image detection results with the set values ​​in the numerical specifications to determine whether there are scaling, deformation, or layout abnormalities. The system also verifies the connectivity relationships between structures. If some characters are identified as separate blocks in the image, but should be connected structures according to the monetary structure (such as "1" and the decimal point being connected), it is considered a connectivity residual. For cases where the recognition result is inconsistent with the structural rule type, such as a position in the image being identified as the digit "1" when it should actually be a decimal point, this type of mismatch will be assigned a higher structural residual score. The system combines the residual information of all dimensions into a single structural residual dataset.

[0168] S44: Perform multi-scale structural differencing on the structural residual data to obtain structural differencing data.

[0169] Specifically, based on the generated structural residual data, the system performs multi-scale structural differencing to generate structural differencing data. For example, the system uses sliding windows of different sizes (e.g., window lengths of 3, 5, and 7) to perform convolutional traversal and analysis on the structural residual sequence, capturing the differential fluctuation characteristics within a local range. Within each sliding window, the system statistically analyzes indicators such as the variance of the structural residuals, the maximum jump amplitude, and the average difference within the current range, obtaining local differencing analysis results. The system fuses the local differencing analysis results obtained at multiple scales (weighted fusion of parts in the same region) to construct structural stability data. Based on the stability response map, the system identifies high-response regions (regions exceeding a preset threshold) and marks them as structural anomaly candidates, outputting structural differencing data.

[0170] Preferably, the confidence mapping is as follows:

[0171] Prior probability data of the structure is obtained by fitting prior probability data based on structural difference data;

[0172] Specifically, the system performs prior probability fitting based on structural difference data, outputting structural prior probability data. For example, the system pre-sets a set of structural difference confidence models trained from historical real document samples. This model can employ a Gaussian mixture model (GMM), a Bayesian estimation model, or other lightweight statistical learning methods. When the system receives a new structural difference vector, it uses it as input and compares it with the normal sample distribution in the model. The system analyzes the frequency of each structural residual signal in the real data dimension by dimension, assessing whether it falls within the normal range. For example, for residuals in dimensions such as location, size, and connectivity, the system calculates the probability of their occurrence in normal documents, i.e., the likelihood of the current residual value appearing in untampered samples. The system organizes the prior probabilities corresponding to each dimension of the residual into a structural prior probability vector.

[0173] The risk level of the monetary region is calculated based on the prior probability data of the structure, and the risk level data of the monetary region is obtained.

[0174] Specifically, based on the obtained prior structural probability data, the system quantifies the anomaly risk of the entire monetary field region and outputs risk level data for the monetary region. For example, the system aggregates the prior probabilities of all structural feature dimensions to assess the overall structural confidence level of the monetary region. During the aggregation process, the system quantifies the degree of anomaly in each dimension and identifies the concentration of low-probability residual terms. , This represents the risk level value / risk level data for the amount region. The total number of structural dimensions. For structural dimension indexing, For the first The system calculates the prior probabilities of each structural dimension. Optionally, the system performs weighted evaluations on key positions in the monetary structure. For example, if the prior probabilities of sensitive characters such as the decimal point or the high digit of the integer are significantly low, the system will trigger risk amplification, assigning higher risk contribution weights to the residual signals at these positions. The system outputs a normalized risk score for the monetary region, ranging from 0 to 1, where a higher value indicates that the monetary field is more likely to have structural anomalies or tampering.

[0175] Confidence calculations are performed based on prior structural probability data and risk level data for monetary regions to obtain document risk label data.

[0176] Specifically, the system combines prior structural probability data with regional risk data to perform a global confidence assessment, thereby generating document risk label data. For example, the system constructs a fusion discrimination model, using methods such as logistic regression or weighted functions to concatenate the previously obtained prior structural probability data and regional risk data into a joint vector as input for confidence calculation. For instance, the system inputs the joint vector into a trained logistic regression model. This model learns the sensitivity to anomalies in different dimensions based on historical samples. The output is a risk score in the [0,1] interval, representing the confidence probability that the document is judged as an "abnormal document." Multiple threshold levels are set (e.g., >0.7 for high risk, <0.3 for reliable, and others for medium risk) to output labels. These labels can serve as the basis for downstream business rule judgments, supporting the triggering of strategies such as automatic interception, risk warnings, or manual review. The document risk label data output by the system provides a reliable structural risk basis for the intelligent document review process.

[0177] Preferably, this application also provides a financial document intelligent verification system for executing the financial document intelligent verification method described above, the financial document intelligent verification system comprising:

[0178] The amount area positioning module is used to acquire ticket image data; and to locate the amount area based on the ticket image data to obtain the amount area data.

[0179] The decimal point positioning module is used to recognize monetary characters based on monetary area data to obtain monetary character data; and to locate the decimal point in the monetary character data to obtain decimal point neighborhood image data.

[0180] The geometric feature extraction module is used to extract features from the decimal point neighborhood image data to obtain geometric feature data.

[0181] The document risk label mapping module is used to perform monetary structure difference calculation based on geometric feature data to obtain structure difference data; and to perform confidence mapping based on the structure difference data to obtain document risk label data.

[0182] Therefore, the embodiments should be regarded as exemplary and non-limiting in all respects, and the scope of the invention is defined by the appended application documents rather than the foregoing description. Thus, it is intended that all variations falling within the meaning and scope of the equivalents of the application documents be incorporated into the invention.

[0183] The above description is merely a specific embodiment of the present invention, enabling those skilled in the art to understand or implement the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features of the invention herein.

Claims

1. A method for intelligent verification of financial documents, characterized in that, Includes the following steps: S1: Obtain the bill image data; locate the amount area based on the bill image data to obtain the amount area data; S2: Recognize the amount characters based on the amount area data to obtain the amount character data; locate the decimal point in the amount character data to obtain the decimal point neighborhood image data; S3: Extract features from the decimal point neighborhood image data to obtain geometric feature data; S4: Construct the numerical side structure based on the geometric feature data and the preset amount format rules to obtain the numerical side structure data; construct the image side structure based on the geometric feature data to obtain the image side structure data; Structural residual data is obtained by performing dimension-wise difference calculations based on numerical and image-based structural data; multi-scale structural difference processing is then performed on the structural residual data to obtain structural difference data. Confidence mapping is performed based on the structural difference data to obtain document risk label data.

2. The method of claim 1, wherein, The specific location of the amount area is as follows: Local brightness detection and artifact compression processing are performed on the ticket image data to obtain local brightness data and artifact compression data, respectively. The ticket image data is filtered based on local brightness data and compressed artifact data to obtain image filtering data; Based on the image-filtered data, structural line density analysis is performed to obtain the target structural line data; Based on the target structural line data, monetary symbols are extracted from the image filtering data to obtain monetary symbol data; Based on the numerical spacing analysis of the monetary symbol data, the monetary number column data is obtained; The amount range data is constructed by using the amount symbol data and the amount number column data.

3. The method of claim 2, wherein, The structural line density analysis specifically involves; Horizontal stroke enhancement is performed on the image-filtered data to obtain horizontal text enhancement data; Vertical projection is performed on the horizontally enhanced text data to obtain the structural line density spectrum data; Peak filtering is performed on the structure line density spectrum data to obtain structure line filtered data; Based on the data filtered by the structural lines, the prior distribution of character height is determined to obtain the target structural line data.

4. The method of claim 1, wherein, The specific method for recognizing monetary amounts is as follows: Text orientation is corrected based on the amount area data to obtain amount text image data; Character stroke enhancement is performed on the monetary text image data to obtain text image enhancement data; Character connectivity processing is performed on the text image enhancement data to obtain preliminary character block data; The initial character block data is sorted and integrated to obtain the character block data. Local super-resolution enhancement is performed on the character block data to obtain character-enhanced data; The enhanced character data is recognized using a preset dual-channel character recognition model to obtain monetary character data.

5. The method according to claim 1, characterized in that, The decimal point positioning is specifically as follows: The decimal point image is extracted from the amount character data to obtain the decimal point image data; The decimal point position is extracted from the decimal point image data to obtain the decimal point position data; Local stroke connectivity is filtered based on decimal point position data to obtain isolated position data; Multi-scale isolation degree calculation is performed on isolated location data to obtain decimal point neighborhood image data.

6. The method according to claim 5, characterized in that, The calculation of multi-scale isolation degree is as follows: Multi-scale local variance suppression convolution is performed on isolated location data to obtain stability feature data; Local energy decay calculations are performed on the stability characteristic data to obtain local energy decay data; The isolation probability data is obtained by calculating the isolation probability based on the local energy decay data and stability characteristic data; Cross-scale similarity data is obtained by performing cross-scale similarity calculations based on isolated probability data; The isolated location data corresponding to the maximum value in the cross-scale similarity data is determined as the decimal point neighborhood image data.

7. The method according to claim 1, characterized in that, S3 specifically refers to: Local binarization is performed on the decimal point neighborhood image data to obtain local binary map data; Isolated stroke connectivity analysis was performed on local binary map data to obtain isolated feature data. Relative geometric feature data is obtained by extracting relative geometric features from the decimal point neighborhood image data based on the amount character data; Texture microstructure features are extracted from the decimal point neighborhood image data to obtain texture microstructure feature data; The isolated feature data, relative geometric feature data, and texture microstructure feature data are integrated to obtain geometric feature data.

8. The method of claim 1, wherein, The confidence mapping is as follows: Prior probability data of the structure is obtained by fitting prior probability data based on structural difference data; The risk level of the monetary region is calculated based on the prior probability data of the structure, and the risk level data of the monetary region is obtained. Confidence calculations are performed based on prior structural probability data and risk level data for monetary regions to obtain document risk label data.

9. A financial document intelligent verification system, characterized in that, For executing the intelligent verification method for financial documents as described in claim 1, the intelligent verification system for financial documents includes: The amount area positioning module is used to acquire ticket image data; and to locate the amount area based on the ticket image data to obtain the amount area data. The decimal point positioning module is used to recognize monetary characters based on monetary area data to obtain monetary character data; and to locate the decimal point in the monetary character data to obtain decimal point neighborhood image data. The geometric feature extraction module is used to extract features from the decimal point neighborhood image data to obtain geometric feature data. The document risk label mapping module is used to perform monetary structure difference calculation based on geometric feature data to obtain structure difference data; and to perform confidence mapping based on the structure difference data to obtain document risk label data.

Citation Information

Patent Citations

  • Medical bill image structuring method and device and computer readable medium

    CN112926577A

  • Archive digitization method and system based on intelligent image enhancement and automatic classification

    CN119049066A