Machine learning-based intelligent identification and verification method for financial bills
By using real-time monitoring and machine learning algorithms to assess risk levels, analyze noise interference, and dynamically make decisions and verify processes, the stability and efficiency issues of financial bill identification and verification in existing technologies have been resolved, achieving stable operation and rational resource allocation under high-load conditions.
Patent Information
- Application Number
- CN202511126228.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-12
- Publication Date
- 2025-10-10
- Estimated Expiration
- 2045-08-12
AI Technical Summary
Existing financial document recognition and verification technologies are prone to collapse under high load conditions, cannot effectively handle noise interference, lack assessment of potential errors in the recognition process, and have unscientific verification processes, resulting in waste of resources and inefficiency.
By monitoring the number of bills in real time, using machine learning algorithms to generate feature interaction graphs to assess risk levels, analyze noise distribution patterns, calculate the degree of interference, dynamically make decisions to optimize the verification process, and calculate verification priorities based on the identification of complex factors.
It achieves stable operation under high load conditions, accurately handles noise interference, rationally allocates verification resources, improves the flexibility and adaptability of identification and verification, and reduces errors and resource waste.
Smart Images

Figure CN120612190B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of intelligent bill verification, and specifically to a method for intelligent identification and verification of financial bills based on machine learning. Background Art
[0002] In financial systems, bill recognition and verification are crucial for ensuring accurate financial data and standardized processes. Traditionally, this task is largely manual, requiring staff to verify bill information piece by piece. This not only consumes significant labor costs but also, when dealing with large numbers of bills, is prone to errors due to fatigue and negligence. With the advancement of information technology, automated bill recognition technology has been gradually applied to the financial sector, significantly improving processing efficiency.
[0003] Existing automation technology still has significant limitations. Regarding processing scale, when the number of bills arriving in a short period of time exceeds the system's preset processing capacity, the system often experiences operational lags, response delays, and even crashes. The lack of dynamic monitoring and intelligent control mechanisms for processing load makes it difficult to maintain stable operation under high loads.
[0004] Background noise in bill images is a major factor influencing recognition accuracy. Existing technologies often use unified filtering algorithms to address noise, failing to account for the diversity and dynamic nature of noise in different scenarios. For example, noise caused by blurred print, stains, and light reflections can vary significantly in distribution and intensity. This can often obscure key information on bills (such as the amount, seal pattern, and issue date) or lead to misidentification.
[0005] Existing technologies lack a systematic assessment of potential error risks during the identification process, making it impossible to predict the reliability of identification results in advance, resulting in a lack of targeted verification. Verification processes also often use a fixed sequence, without distinguishing between the difficulty and risk level of the bill. This results in complex and high-risk bills being delayed, while simple bills consume excessive verification resources, affecting the consistency and efficiency of the overall process. Furthermore, the attribute data contained in the identification path is not effectively mined, making it difficult to quantify the complexity of bill identification. This makes the division of verification priorities lack a scientific basis, further restricting the overall efficiency of bill processing. Summary of the Invention
[0006] The purpose of the present invention is to provide a method for intelligent identification and verification of financial bills based on machine learning to solve the problems raised in the above background technology.
[0007] To achieve the above objectives, the present invention provides a method for intelligently identifying and verifying financial bills based on machine learning, the method comprising:
[0008] Real-time acquisition of the original image data stream of the financial bill, monitoring the total quantity of input bills through the image acquisition device, judging whether the total quantity exceeds the preset system processing safety threshold;
[0009] When the total quantity exceeds the system processing safety threshold, the feature elements of the bill image are processed using a machine learning algorithm, a feature interaction graph is generated, and the potential error risk level in the bill recognition process is evaluated based on the feature interaction graph;
[0010] Applying signal decomposition technology to analyze the distribution pattern of bill image background noise, reconstructing a noise dynamic characteristic model, and calculating the interference degree index of background noise on the key area of the bill based on the noise dynamic characteristic model;
[0011] According to the potential error risk level and the interference degree index, it is decided whether to start the overall bill verification optimization process;
[0012] When it is decided to start the overall bill verification optimization process, for each independent bill image, the attribute data of its recognition path is parsed, and the recognition complexity factor of each bill is generated;
[0013] Based on the recognition complexity factor of each bill, the verification priority value of each bill is calculated, and the verification sequence is output according to the verification priority value.
[0014] Preferably, the real-time acquisition of the original image data stream of the financial bill, monitoring the total quantity of input bills through the image acquisition device, judging whether the total quantity exceeds the preset system processing safety threshold, specifically includes:
[0015] Continuously receiving bill image data packets from the image sensor, the bill image data packets containing resolution parameters, color depth parameters and image size parameters;
[0016] Classify the bill types in the bill image data packets using image recognition algorithms to determine the effective bill category list;
[0017] Compare the total number of bills in the effective bill category list with the system preset processing capacity upper limit value to generate an input load intensity index;
[0018] Based on the comparison result of the input load intensity index and the system processing safety threshold, it is judged whether there is a bill recognition safety hazard signal.
[0019] Preferably, the feature elements of the bill image are processed using a machine learning algorithm, a feature interaction graph is generated, and the potential error risk level in the bill recognition process is evaluated based on the feature interaction graph, specifically including:
[0020] extracting a pre-processed check image feature vector, the pre-processed check image feature vector being obtained by an edge enhancement algorithm and a feature dimension reduction technique;
[0021] inputting the pre-processed check image feature vector into a deep neural network model, training the model to generate a multi-dimensional feature interaction graph;
[0022] analyzing the correlation strength matrix between feature elements in the multi-dimensional feature interaction graph, and calculating a correlation anomaly probability;
[0023] quantifying a potential error risk level according to the correlation anomaly probability.
[0024] Preferably, the application signal decomposition technique is used to analyze the distribution pattern of the background noise of the check image, reconstruct a noise dynamic characteristic model, and calculate an interference degree index of the background noise on the key region of the check image based on the noise dynamic characteristic model, specifically including:
[0025] performing a frequency domain conversion operation on the denoised check image to extract noise spectrum distribution data;
[0026] processing the noise spectrum distribution data using a multi-scale decomposition algorithm to construct a spatio-temporal noise dynamic characteristic model;
[0027] analyzing the noise intensity variation curve in the spatio-temporal noise dynamic characteristic model to generate a key region interference evaluation value;
[0028] outputting the interference degree index based on the key region interference evaluation value.
[0029] Preferably, the potential error risk level and the interference degree index are used to determine whether to start an overall check verification optimization process, specifically including:
[0030] when the potential error risk level is high and the interference degree index is high, generating a start overall check verification optimization process instruction;
[0031] when the potential error risk level is medium or low, or the interference degree index is medium or low, generating a do not start overall check verification optimization process instruction;
[0032] outputting the start overall check verification optimization process instruction or the do not start overall check verification optimization process instruction as a decision result.
[0033] Preferably, the attribute data of the recognition path of each independent check image is analyzed to generate a recognition complexity factor for each check, specifically including:
[0034] calculating the feature similarity difference value of each check image with other check images;
[0035] Identify the conflict area distribution map in the feature similarity difference value and mark the coordinates of the low-definition area;
[0036] Integrate the number of regions and severity parameters in the conflict area distribution map to generate an identification path attribute dataset;
[0037] Based on the identification path attribute dataset, the identification complexity factor is calculated.
[0038] Preferably, the step of integrating the number of regions and severity parameters in the conflict region distribution map to generate an identification path attribute dataset specifically includes:
[0039] Count the total number of all marked areas in the conflict area distribution map;
[0040] Measure the average value of the blur index for each marked area;
[0041] The total quantity value and the average value of the fuzzy index are fused to construct the recognition path attribute dataset.
[0042] Preferably, the step of calculating the verification priority value of each bill based on the identification complexity factor of each bill and outputting the verification sequence according to the verification priority value specifically includes:
[0043] Assigning a type weight coefficient to each bill, wherein the type weight coefficient is determined based on the bill category and business importance;
[0044] Multiply the type weight coefficient by the recognition complexity factor to generate a verification priority score;
[0045] Compare the verification priority scores of all tickets and sort them from highest to lowest.
[0046] Preferably, the real-time acquisition of the original image data stream of the financial bills and monitoring the total number of input bills through the image acquisition device further includes:
[0047] Apply optical character recognition technology to preliminarily analyze text information fragments in the bill image data stream;
[0048] Store text information fragments in a temporary database, verify the integrity and consistency of text information fragments, and generate data quality reports.
[0049] Preferably, the input of the pre-processed bill image feature vector into the deep neural network model and the training model to generate a multi-dimensional feature interaction graph specifically includes:
[0050] A deep neural network model is constructed by combining convolutional neural network with attention mechanism. The deep neural network model is trained to learn the dependency between feature vectors and output a multi-dimensional feature interaction graph to the risk assessment module.
[0051] Compared with the prior art, the present application has the beneficial effects that:
[0052] The present method can timely grasp the current processing pressure of the system by acquiring the original image data stream of the financial bill in real time and monitoring the total quantity of input bills. When the total quantity exceeds the preset safety threshold, instead of simply refusing to process or forcibly continuing, a machine learning algorithm is introduced to process the feature elements of the bill image, and a feature interaction graph is generated to evaluate the potential error risk level. In this way, the system can still effectively perceive the risks in the identification process under high load conditions, avoiding a large number of errors caused by blind processing.
[0053] In terms of background noise processing, the present method breaks out of the framework of traditional fixed filtering methods and applies signal decomposition techniques to deeply analyze the distribution patterns of noise and further reconstruct a noise dynamic characteristic model. By calculating the interference degree index of background noise on the key regions of the bill through the model, the unique properties of noise on different bills can be accurately captured, making the noise processing more targeted and reducing the interference of noise on key information recognition, so that the identification result is closer to the actual content of the bill.
[0054] According to the potential error risk level and the interference degree index, it is decided whether to start the overall bill verification optimization process, so that the start of the verification process is more in line with actual needs. When the risk and interference are at a low level, the regular process can be maintained to avoid resource waste; when the risk or interference reaches a certain level, the optimization process is started to make the verification work more focused on bills with prominent problems.
[0055] For each independent bill image, the attribute data of the identification path is analyzed to generate an identification complexity factor, which provides a specific basis for the division of verification priority. The information complexity and identification difficulty of different bills differ, and based on the complexity factor, the verification priority value is calculated and sorted to output the verification sequence, which can make the verification resources more reasonably allocated, and the bills with high identification difficulty and potential problems are given priority, while simple bills are quickly passed through, reducing unnecessary waiting and repeated verification, and improving the smoothness of the overall verification link.
[0056] Overall, the present method forms a set of interconnected processing mechanisms from bill input to final verification sequence output, which can adapt to different quantity scales, different noise environments, and different complexity levels of bill processing scenarios. Through dynamic load perception, accurate risk assessment, targeted noise processing, intelligent process decision-making, and scientific priority division, the identification and verification process of financial bills is more flexible and adaptable, reducing various problems caused by technical rigidity, and making the entire processing flow more in line with the diversified needs of actual financial work. BRIEF DESCRIPTION OF DRAWINGS
[0057] Figure 1This is a working principle diagram of the financial bill intelligent identification and verification method based on machine learning according to the present invention;
[0058] Figure 2 Flowchart for feature interaction graph generation and error risk assessment;
[0059] Figure 3 Flowchart for background noise analysis and interference level assessment;
[0060] Figure 4 Flowchart generated for identifying complex factors. DETAILED DESCRIPTION
[0061] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0062] See also Figure 1 The present invention provides a method for intelligently identifying and verifying financial bills based on machine learning, the method comprising:
[0063] The system acquires raw image data streams of financial documents in real time. The image acquisition device monitors the total number of input documents and determines whether this total number exceeds a preset system processing safety threshold. If this total number exceeds the system processing safety threshold, a machine learning algorithm is used to process the characteristic elements of the document image, generating a feature interaction graph. Based on this feature interaction graph, the system assesses the potential error risk level during the document recognition process. Signal decomposition techniques are used to analyze the distribution pattern of background noise in the document image, reconstruct a noise dynamics model, and calculate, based on this noise dynamics model, an index of the degree of background noise interference with key document areas. Based on this potential error risk level and interference index, a decision is made whether to initiate the overall document verification optimization process. If the decision is made to initiate the overall document verification optimization process, attribute data of the recognition path for each individual document image is analyzed to generate a recognition complexity factor for each document. Based on the recognition complexity factor, a verification priority value is calculated for each document, and the verification sequence is output based on the verification priority value.
[0064] Example 1: See Figure 2, covering the processing flow and risk quantification mechanism of raw image data of financial bills. The image sensor continuously receives bill image data packets containing resolution parameters, color depth parameters and image size parameters. The resolution parameter reflects the image pixel density, the color depth parameter determines the number of color channel bits, and the image size parameter records the pixel width and height. These metadata are extracted into the processing system memory area as basic attributes. The image data packets are processed using a bill classification algorithm based on computer vision: first, a recognition rule set containing a template library of value-added tax invoices, itineraries, fiscal bills, etc. is constructed, and then the texture features and layout structure features of the bills to be classified are extracted. By calculating the similarity distance matrix between the features of the bill to be processed and the template library, the image units whose matching degree exceeds the preset standard are identified, a list of valid bill categories is generated, and the category label of each bill is marked.
[0065] The system reads the processing capacity cap within a pre-set memory area, which is determined by the number of processor cores, available memory, and historical average processing latency. It calculates the total number of tickets in the valid ticket category list and divides it by the capacity cap to generate a percentage-based indicator of the input load intensity. The processing safety threshold is set to a critical point in the load index (e.g., 75%). When the load index exceeds this critical point, a binary signal "1" is generated, indicating a ticket identification safety risk. Otherwise, a signal "0" is output.
[0066] When the security threat signal is activated, preprocessing is performed on the original bill image. The edge enhancement algorithm uses a Sobel operator convolution kernel to perform image convolution, detecting gradient changes at the bill's edges and enhancing outlines. Principal component analysis is then applied for feature dimensionality reduction: a covariance matrix is constructed for all bill image pixels, eigenvectors are calculated, and principal component vectors with a variance contribution exceeding 85% are selected. The preprocessed eigenvectors contain structured data such as bill type, amount area location, and seal area coordinates.
[0067] The input layer of the deep neural network model receives the preprocessed feature vector. The model architecture consists of four convolutional layers and two fully connected layers. Each convolutional layer is followed by a ReLU activation function and a max pooling operation. Through forward propagation, the model generates a multidimensional feature interaction graph in the hidden layer, which contains feature elements such as the amount recognition probability and seal matching degree. The interaction graph is stored as a three-dimensional tensor. The first dimension identifies the bill number, the second dimension stores the feature element type index, and the third dimension records the feature value confidence score.
[0068] When analyzing multidimensional feature interaction graphs, a feature correlation matrix is established: rows and columns correspond to different feature element types. The matrix element values are used to calculate the Pearson correlation coefficient of the co-occurrence frequency of feature elements in historical data for similar bills. The calculation process for correlation anomaly probability involves extracting the actual co-occurrence value of the current bill's feature elements, comparing this value with the theoretical co-occurrence value stored in the correlation matrix, and ultimately converting the deviation into an anomaly probability value ranging from 0 to 1. Risk level quantification uses a three-level classification: anomaly probability < 0.3 corresponds to low risk, 0.3 ≤ probability < 0.7 medium risk, and probability ≥ 0.7 high risk. The output includes a risk level label and the corresponding quantified probability value.
[0069] The bill classification algorithm maintains a dynamically updated bill template library. When a new bill type is added to the database, a semi-automatic verification process is initiated: the key area coordinates of three sample bills are manually marked, and the system automatically extracts the mean of the layout characteristics of the bills of this category and generates a new template. In the case of abnormal bill size parameters, the image processing system initiates an adaptive size normalization process: the difference ratio between the image size parameters and the standard value is detected. When the difference exceeds ±5%, the bilinear interpolation algorithm is called for pixel resampling. The parameter optimization in the edge enhancement process adopts an iterative parameter adjustment mechanism - the enhancement process of different parameters is repeatedly performed on the same batch of bills, and the parameter combination with the best bill edge continuity evaluation index is selected.
[0070] A data compression mechanism is implemented during the generation of the multidimensional feature interaction graph. Tensor data storage utilizes a quantized compression algorithm: 32-bit floating-point eigenvalues are mapped to 8-bit integer space, while retaining normalized scaling parameters for precision restoration. The update strategy for the feature association matrix sets a time decay factor: when new bill data enters, the weight of the historical correlation coefficient is reduced at an exponential decay rate of 0.95. The weight parameters of the deep neural network model are incrementally trained weekly, optimizing feature extraction capabilities by loading newly added bill datasets. The calculation process for risk anomaly probability includes anti-interference processing. When a single eigenvalue experiences a sudden anomaly, the feature credibility verification module is activated to verify whether there is a transmission delay in the sensor data acquisition timestamp.
[0071] Example 2: See Figure 3 , covering the bill image noise modeling technology and risk response decision-making mechanism. When receiving the denoised bill image data, a two-dimensional discrete Fourier transform operation is performed to complete the conversion from the time domain to the frequency domain. The spectrum distribution data is stored in the form of a complex matrix, the real part of which describes the amplitude information and the imaginary part records the phase information. The noise spectrum separation process implements the following steps: constructing an image mask matrix containing the text area, seal area, and background area; using the mask matrix to perform a point multiplication operation on the complex matrix; separating the spectral components of the background area as the noise analysis object; extracting the energy amplitude of each frequency point and normalizing it to a relative intensity value in the range of 0-100.
[0072] The multi-scale decomposition algorithm uses a third-order wavelet transform framework. The first-level decomposition divides the noise spectrum into low-frequency approximate components and high-frequency detail components. The second level further decomposes the low-frequency components into secondary high- and low-frequency combinations. The third level repeats this process to form a tree-structured spectrum diagram. During the reconstruction process, the energy entropy value, spectrum offset, and time-domain fluctuation characteristic parameters of each layer component are extracted: the variance statistics of the subband coefficients, the mean of the spectrum center distance, and the intensity variation coefficient within the continuous time window are calculated for each level. The method for constructing the spatiotemporal noise dynamic characteristic model includes: arranging the third-order decomposition parameters into a tensor structure in time sequence; the first dimension records the image acquisition time point; the second dimension stores the wavelet decomposition layer index; the third dimension stores the characteristic parameter value; and finally, applying tensor decomposition technology to reduce the dimension to a two-dimensional feature matrix.
[0073] During the noise intensity variation curve analysis phase, the model output includes two curves: time correlation and spectrum correlation. The time correlation curve plots the changing trend of the mean background noise intensity at a 30-second sampling interval; the spectrum correlation curve records the evolution of the energy contribution of the main noise frequency (50-100Hz). The key area interference assessment value is generated using a two-layer mapping strategy: the first layer calculates the frequency of abnormal fluctuations in the noise curve, and an outlier is identified when the intensity of a single point deviates from the historical mean by more than three standard deviations. The second layer measures the average energy contribution of the noise in the seal recognition area and the amount area during the abnormal period. The final output value is the weighted sum of the abnormal frequency and energy contribution.
[0074] The interference level indicator is quantified into a three-level system: a calculated value of ≤40 indicates a low interference level, 40-70 indicates a medium interference level, and a calculated value of ≥70 indicates a high interference level. The decision process is activated based on receiving the potential error risk level and interference level indicator from the aforementioned process. The system maintains a state transition matrix: rows correspond to high, medium, and low risk levels, columns indicate high, medium, and low interference levels, and matrix elements store decision instruction codes. When the risk level is high (code value 3) and the interference level is high (code value 3), instruction code D01 is triggered to initiate the overall bill verification optimization process. If the risk level is high but the interference level is not high, the activation frequency within the past 10 minutes is checked. If the cumulative activation exceeds three times, the optimization process is temporarily disabled.
[0075] When the risk level detection value is in the medium range (code value 2), the activation condition judgment mechanism adds a buffer rule: when the interference level reaches a high level and lasts for more than 5 minutes, the optimization process is activated; if the interference level is medium or below, the optimization process is skipped. When the risk level is low (code value 1), only basic monitoring is performed: the interference indicator change rate is recorded in real time, and a warning signal is issued when the change between adjacent sampling points exceeds 20%, but the core process is not activated. The final decision result is transmitted in a structured object format: it contains an 8-byte instruction code, a 32-byte timestamp, a 16-byte validity duration field, and a 256-byte additional parameter area. A data verification mechanism is implemented at the transport layer: a CRC32 cyclic redundancy check code field is set at the end of the additional parameter area. The receiving end only executes the instruction after the verification result matches.
[0076] The updating strategy for the spatiotemporal noise dynamic characteristic model uses a sliding window mechanism. For every 10 new bill images received, the system automatically removes the oldest 20% of feature data, retaining 80% of historical data for fusion modeling with the new data. Wavelet decomposition parameter optimization utilizes adaptive basis function selection: when persistent spikes are detected in the noise spectrum within a specific frequency band (e.g., 150-200Hz), the Daubechies wavelet basis is automatically switched to the Coiflet wavelet basis to improve decomposition accuracy. Outlier point detection in the time correlation curve implements noise isolation processing. When a single-point anomaly is detected, the system automatically captures 5 seconds of data before and after the period and performs a quadratic Fourier transform. If periodic spikes are detected, they are classified as inherent system noise rather than temporary interference.
[0077] The interference assessment calculation process incorporates regional weighting factors. Key areas of the bill are weighted based on their locational characteristics: the amount area is assigned a weighting factor of 0.6; the bill number area is assigned a weighting factor of 0.3; and the signature area is assigned a weighting factor of 0.1. The final assessment value is adjusted to the weighted average of the interference values for each area. The decision state transition matrix features dynamic reconstruction capabilities: the system conducts monthly statistical analysis of historical decision results and assigns a disable flag to decision paths with a state transition probability below 15%. The content stored in the additional parameter area is encrypted: the 128-bit parameter description text is AES encrypted using the bill batch number as the key for transmission, and the verification module performs simultaneous decryption upon process initiation.
[0078] Example 3: See Figure 4 The quantification process of the complexity of bill image recognition starts with the calculation module of feature similarity difference value. The system loads the pre-processed bill feature data set, which contains the convolution feature vector of the amount area of each bill, the seal matching confidence score and the key text area position encoding, and performs pairwise similarity comparison analysis. The feature vector is stored in the tensor buffer using 768-dimensional floating values, and the index rule is set to arrange the bills in the order of their storage time. Similarity difference value calculation engine: For the target bill image (where i is the unique identifier index of the bill) to query the database for all other bill image collections processed in the last 30 minutes ;extract The eigenvector of With each The eigenvector of ; Apply the standardized Euclidean distance algorithm to measure the vector difference, and the calculation formula is expressed as:
[0079]
[0080] In this formula, Indicates bill and The feature similarity difference value between them ranges from 0 (complete similarity) to 1 (maximum difference); and Corresponding bills and 768-dimensional feature vector of Calculate the Euclidean norm distance between two vectors; Logo vector The absolute difference of the value range of each dimension is calculated as the maximum characteristic element minus the minimum element. Logo vector The absolute difference of the value range of each dimension is calculated as the maximum characteristic element minus the minimum element. The engine processes all bill combinations in parallel and outputs a symmetric difference matrix D, where the matrix elements storage The main diagonal elements are automatically set to zero to avoid self-comparison. The matrix D is cached in a two-layer hash table structure, with the outer key being the ticket index i and the inner key being the difference value j.
[0081] The generation of the conflict area distribution map relies on difference matrix analysis and image space mapping. After reading the difference value, the system sets a dynamic difference threshold θ: the initial value is 0.7, and it is automatically lowered to 0.6 when the average difference value of the matrix D exceeds the historical mean for 5 consecutive sample points. Region Mapping Module: Extracting Difference Values All bill pair combinations of The image pixel grid is divided into a 64×64 unit block grid; the feature difference score of each block is calculated, and the difference value of the subset of feature vectors involved in the block (such as the 3rd-5th dimension vectors corresponding to the amount area) is calculated; the score is compared with the global threshold to mark the conflicting blocks. The conflict area distribution map is constructed as a binary mask matrix (k is the block index), the conflicting position is marked as 1, and the non-conflicting area is 0. The mask is stored in the image coordinate space, with the X axis recording the horizontal pixel coordinate (0-1023) and the Y axis recording the vertical pixel coordinate (0-767).
[0082] The identification of low-definition area coordinates integrates spatial frequency analysis and conflict mask data. The system applies Laplace gradient clarity detection to the conflict-marked blocks: the original bill grayscale image I is input, a 3×3 Laplace convolution kernel operation is performed on each conflict block, and the gradient change is calculated; the clarity threshold γ is set to 50 (adjusted according to the image resolution), and gradient values below γ are classified as low-definition areas. The coordinate marking process generates a point set ,in It is a two-dimensional coordinate tuple (x, y) that describes the center position of the block; it also records the fuzziness index of each coordinate , calculated as 1 minus the normalized gradient value ( The value range is 0-1, with 0 representing completely clear). Coordinate data is stored in a linked list structure, with a timestamp and ticket index as metadata headers.
[0083] The integration process of the number of regions and severity parameters includes statistics and average calculations. Conflict area distribution map analysis module: traversing the mask matrix , count the total number of blocks L marked as 1 (i.e. the total number of conflict areas). The severity parameter focuses on the ambiguity and performs an aggregation operation on the point set P: extract all The values form list B; the average value of the fuzziness index is calculated using the following cumulative averaging process. The identification path attribute dataset is defined as a structured data object, including the fields (storage value), (Store the fuzzy mean), and the severity vector S (serialized storage of each region After the data is generated, it is compressed into JSON format and output to the persistent storage area.
[0084] Identify complex factors and perform numerical comprehensive calculations based on attribute data sets. The system parses the attribute data set JSON object and reads value and value The factors are defined as linear combinations: weight coefficients are defined is 0.7 (region quantity weight), is 0.3 (fuzziness weight), and these coefficients are fixed values obtained through training of historical data sets. The calculation logic is as follows: Value to 0-1 range, divided by the maximum area reference value 100; weighted sum generates complexity factor value, formula uses a single description, no new formula here. The result value is written into the ticket metadata database as the complexity factor field, which is called by the verification sorting module.
[0085] The feature similarity difference calculation engine implements a real-time optimization mechanism: the difference matrix D maintains an incremental update strategy, and only calculates the difference value of the newly added combination when a new ticket is received, rather than rebuilding the full amount. The database memory management sets the recovery threshold: when the number of tickets exceeds 500, start the aging data cleaning, retain the latest 80% data to control the calculation load. The conflict block grid division uses a dynamic adjustment algorithm: detect the image content distribution, if the ticket layout structure contains multiple column text areas, automatically switch the grid size to 32x64 to improve accuracy.
[0086] The anti-noise processing of low-definition coordinate marking introduces a multiple measurement mechanism. Laplacian gradient detection performs three times of sampling: each time uses a different size of convolution kernel (3x3, 5x5, 7x7), and takes the median value as the final To reduce the influence of outliers. The fuzziness index calculation process implements a normalization preprocessing: adjust the clarity threshold γ according to the image brightness mean and contrast parameters, the specific rule is: when the brightness mean is lower than 100, increase γ to 60 to prevent false judgment in dark areas.
[0087] The attribute data set construction implements integrity check. Before fusing and fuzziness data, the system checks the validity of list B: if The value is empty or exceeds the 0-1 range, automatically call the original image to recalculate. The data format uses strong type constraint: Stored as a 16-bit unsigned integer, Using 32-bit floating-point numbers.
[0088] The recognition complexity factor calculation module integrates a security mechanism: in the application of weight coefficients and , embed coefficient check loop, verify coefficient and whether equal to 1.0 before each call to ensure calculation consistency. When processing high load scenarios, the system deploys parallel threads, starts an independent thread for each ticket to perform complexity factor calculation, and the shared memory area synchronizes data access through mutex lock. The ticket storage process includes a write-back link: after writing the complexity factor value, automatically trigger database index reconstruction to ensure subsequent query efficiency.
[0089] The coordinate data storage optimization adopts a spatial index technology: the point set P is encoded using an R-tree index structure, and the index key contains a ticket index and a timestamp to accelerate regional matching queries. The conflict region analysis is implemented under a GPU acceleration framework: mask matrix operations are migrated to CUDA core execution, and block processing is used to improve throughput. The persistence of the identified path attribute dataset uses a differential backup strategy: only when the dataset changes by more than 10% is the full amount written to disk, otherwise the incremental log is recorded.
[0090] The recovery mechanism for low-definition recognition copes with hardware failures: when gradient detection fails, the system falls back to a Fourier spectrum-based definition evaluation algorithm, extracting high-frequency component energy as an alternative value. The entire recognition complexity quantification process implements full-chain log tracking: when transmitting intermediate data between computing units, a UUID tracking code is attached to facilitate error tracing without affecting core logic execution speed.
[0091] Embodiment 4: Prioritize ticket verification and data quality verification. The system receives recognition complexity factor data from the previous process, combines it with the ticket type weight configuration table in the business rule library, and performs verification priority calculation. The business rule library maintains dynamically updated ticket category metadata, including processing weights for more than 20 common ticket types such as value-added tax special invoices, train tickets, and flight itineraries. The weight allocation principle is based on a comprehensive assessment of the importance of ticket amounts, anti-fraud requirements, and audit traceability needs. The latest version is revised and released by the financial risk control department every quarter. The system loads the current effective weight configuration table and associates it with the recognition complexity factor.
[0092] Taking five tickets processed in a batch as an example, the system records their basic attributes and intermediate values as shown in the following table:
[0093] Ticket Number Bill Type Identifying complex factors Type weight coefficient Verification priority score Text completeness score Data consistency flag INV20230715_001 Special value-added tax invoice 0.68 0.85 0.578 92 1 TRN20230715_002 train tickets 0.51 0.65 0.332 87 1 AIR20230715_003 Flight itinerary 0.73 0.72 0.526 95 1 REC20230715_004 Catering invoice 0.42 0.45 0.189 78 0 HOT20230715_005 Hotel accommodation invoice 0.59 0.68 0.401 83 1
[0094] The verification priority score calculation process uses a linear product rule: multiply the recognition complexity factor by the type weight coefficient directly, and the result is rounded to three decimal places. After the score is generated, the sorting engine rearranges the ticket processing order in descending order. The final verification sequence in the example is: value-added tax special invoice (0.578), flight itinerary (0.526), hotel accommodation invoice (0.401), train ticket (0.332), and catering invoice (0.189). The sequence is written to the task queue with a timestamp and operator ID, forming a traceable scheduling log.
[0095] Multi-stage verification is implemented in the optical character recognition technology processing link. When the original image data stream is parsed by the OCR engine, the system establishes a temporary text storage area and divides the recognition results by area: the text in the amount area is stored in the amount field buffer area, the invoice date is stored in the date field buffer area, and the bill code is stored in the code field buffer area. The text integrity score is calculated based on the area coverage: a list of text areas that must be included in each bill type is preset (for example, a value-added tax invoice must contain eight areas such as code, number, amount, and tax amount), and the actual percentage of valid areas recognized is counted and converted into a percentage score. The data consistency check implements cross-field logical verification: comparing the uppercase and lowercase values of the amount field to see if they match; verifying whether the check digit of the invoice code is correct; and verifying whether the date field complies with the chronological order. The consistency flag is set to a binary state, 1 if all checks pass, and 0 if any check fails.
[0096] The temporary database adopts a document-based storage structure, and each bill's text information fragment is stored as an independent document. The document model consists of two parts: fixed fields and dynamic fields. Fixed fields record the basic attributes of the bill (type, number, collection time); dynamic fields store recognized text key-value pairs (such as "amount": "480.00"). The database implements a version control mechanism, generating a new document version each time the text information is updated, and retaining historical modification records. The data quality report generation module periodically scans the database content: every hour, it calculates the average completeness score of each bill type; detects the proportion of documents missing key fields; and marks records with abnormal consistency indicators. The report output is a structured data object, which includes a quality score matrix and a list of abnormal document IDs.
[0097] The dynamic adjustment mechanism for type weight coefficients responds to business changes. When the finance department adds a new bill type, the system initiates a weight initialization process: 100 sample bills of that type are processed in a test environment to collect benchmark data such as recognition error rate and verification time. Risk control experts then assess the initial weights based on the benchmark data. After a two-week trial run, these weights are incorporated into the official weight configuration table. Trigger conditions for revising existing types of weights include: the recognition error rate of a certain type of bill increases by more than 5% for three consecutive months; the business department submits an application for a change in security level; or there are major adjustments to the audit rules. Each weight update requires triple confirmation: the system administrator verifies the new configuration format; the finance director reviews the business rationality; and the technical director checks system compatibility.
[0098] The text verification process implements progressively stricter controls. For bills with integrity scores below 80, the system initiates an enhanced recognition process: reprocessing the image using a high-precision OCR model; adding image preprocessing steps (including tilt correction and brightness balancing); and extending recognition time to three times that of the standard process. Failure of the consistency check triggers an exception handling protocol: locking the bill document from automatic verification; generating a manual review task and assigning it to the quality inspection queue; and marking problem areas on the bill image for review. The document version record in the temporary database supports rollback operations: if systematic deviations in the text recognition results are detected, the specified bill can be restored to any of the three previous versions.
[0099] The verification queue's scheduling strategy balances priority and system load. When the sorted bill sequence enters the distributor, the system monitors computing resource utilization in real time. When CPU utilization falls below 70%, the complete sequence is processed. When resources are limited, segmented loading is implemented, with only the top 20% of high-priority bills being sent to the execution unit, while the remaining tasks are temporarily stored in the buffer pool. The task executor is designed with a pluggable architecture, supporting the simultaneous execution of multiple verification algorithm modules (such as anti-counterfeiting detection, tax verification, and reimbursement compliance). Each module must declare its processing capacity level when pulling tasks from the central queue, with higher-level modules receiving higher-priority bills first.
[0100] Data quality report exception handling implements a tiered response. For individual bill text quality issues (completeness score 60-80), the system automatically adds the bill to the re-recognition queue. If more than 10% of a batch of bills exhibits consistency flag anomalies, a batch-level alarm is triggered and subsequent processing is suspended. If systemic recognition defects are discovered (such as date recognition errors caused by a specific OCR engine version), the problematic engine is automatically isolated and a backup recognition service is switched. Report analysis results are fed back to the preceding modules: the recognition complexity factor calculation module receives text quality data as adjustment parameters; and the image acquisition device optimizes the focus area based on information about commonly missing fields.
[0101] Example 5 details the technical implementation of deep neural network model construction and multi-dimensional feature interaction graph generation. The model adopts a hierarchical fusion architecture: the input layer is configured with a data normalization processing unit to receive preprocessed bill image feature vectors. These feature vectors are converted to a zero-mean unit variance numerical distribution through normalization, where the amount feature value range is mapped to the range of -1 to 1, and the bill type code is converted to a one-hot vector. The data normalization unit outputs a 768-dimensional standardized feature vector, which is projected into a 2048-dimensional high-dimensional space through a fully connected network as the first-layer output.
[0102] The convolutional neural network module consists of stacked core units: the first convolutional layer uses 128 7×7 convolutional kernels, sliding across the feature map with a stride of 2. The second layer uses 256 5×5 kernels, employing a dilated convolution pattern to expand the receptive field. The third layer deploys a combination of 512 3×3 kernels. The fourth layer uses 1024 3×3 kernels for depthwise separable convolution. Batch normalization is applied after each convolution layer, and a parameterized ReLU activation function is used to introduce nonlinear transformations. Max pooling layers are placed between each convolutional layer, with a fixed pooling window of 3×3 and a stride of 2 to compress the spatial dimensions.
[0103] The attention mechanism module is deployed in parallel with the convolutional layers, employing a dual-branch fusion architecture. The spatial attention branch receives the feature maps output by each convolution layer and generates channel-wise feature descriptors through a dual-path process of global average pooling and max pooling. The channel-wise attention branch utilizes 1×1 convolution to generate a spatial attention heatmap, identifying the coordinates of high-response regions within the feature map. The outputs of the two branches are fused at the connection layer through element-wise addition and multiplication. The matrix Hadamard product of the fused result and the original feature map is then performed to enhance the weights of key features. This process is repeated after each convolution layer, forming an iterative feature optimization path.
[0104] The model's backend processor generates the multidimensional feature interaction graph. The four-dimensional feature tensor (batch size × feature map height × width × number of channels) output by the final convolutional layer is fed into a conversion unit. This unit compresses the height and width dimensions through a three-dimensional convolution operation, outputting a three-dimensional feature graph structure. The first dimension retains the batch index identifying the bill serial number; the second dimension divides the feature element category index, including 12 categories of features such as amount recognition confidence and bill type matching; and the third dimension stores the 128-bit floating-point feature value vector. This three-dimensional feature array undergoes vector dimensionality reduction using a hierarchical clustering algorithm, ultimately generating a structured multidimensional feature interaction graph data object.
[0105] The model training mechanism implements a phased optimization approach. The initial training phase loads a historical bill dataset of millions of invoices, with a batch size of 256. The loss function uses a combination of focused cross-entropy and contrastive loss: cross-entropy targets bill type classification, while contrastive loss constrains the distribution of feature vector similarity. The optimizer uses stochastic gradient descent with hot restarts, with an initial learning rate of 0.1 and three-level step decay. Incremental training is performed every two weeks: newly added invoices with recognition confidence below 60% are screened to create a training subset. Only the parameters of the last two convolutional layers are unfrozen. Fine-tuning training is limited to 50 epochs to prevent overfitting.
[0106] The feature interaction graph output process implements a data validation protocol. Before being written into the risk assessment module, the system performs a feature value integrity check: checking for NaN outliers in the feature vector; verifying compliance with dimension matching rules; and confirming that the feature value distribution range meets the preset threshold. If more than 30% of the dimensions in a single bill's feature vector are missing, the data is automatically isolated and the recognition process is retriggered. Qualified feature interaction graphs are transmitted to the risk assessment module via a high-speed data bus using a reliable datagram format with a retransmission mechanism.
[0107] Model parameter updates implement automated version control. Trained model weights are stored in a distributed file system, and a 64-bit version hash is generated for each update. During the deployment phase, a grayscale release strategy is implemented: new model weights are applied only to 10% of input bills, and the old and new models are run in parallel for comparison. When the new model's recognition error rate does not exceed the baseline by more than 0.5 percentage points across 1,000 consecutive bills, the new version is fully switched to. Model files are stored encrypted, and key management is implemented through a hardware security module.
[0108] The network module design integrates dynamic resource allocation capabilities. A runtime monitoring system tracks GPU memory utilization and automatically activates feature map compression when free memory falls below 20%. This compresses 128-bit floating-point feature values to 32-bit half-precision format. When processor temperature reaches a critical threshold, the model automatically reduces the number of parallel convolution kernels to reduce computational load. An input buffer congestion control system is implemented: when concurrent input requests exceed processing capacity, request throttling is implemented based on the ticket's verification priority score.
[0109] The storage format for multidimensional feature interaction graphs uses a self-describing data structure. Data objects consist of a metadata header and a feature body. The header records the ticket number, model version, and feature dimension information; the feature body stores a sequential array of feature values. The file system uses columnar storage optimization: data indexed by the same feature category is stored contiguously to improve scanning efficiency. To support historical traceability, each feature graph is written to a persistent log after generation. Log entries include a timestamp and the execution node code.
[0110] The debugging mechanism includes a built-in feature visualization analysis tool. When a technical analyst activates debugging mode, the system captures the tensor data of the interactive graph and converts it graphically: spatial features are mapped into heat maps, and numerical features are converted into histograms. The resulting visualization is overlaid on the original bill image to create a diagnostic view, identifying the specific locations of high-risk feature elements within the bill. This feature is implemented via the Remote Desktop Protocol and does not affect the real-time performance of the main processing pipeline.
[0111] Feature vector exception handling sets up a multi-level recovery path. When a single bill feature anomaly is detected, the system attempts three correction options: re-execute the pre-processing process of the current image; switch to the model weights backed up three hours ago and recalculate; call a low-precision fast inference model to generate a backup result. During the correction process, the relevant bills are marked as "under inspection" in the task queue and will not be transferred to the subsequent links. The normal processing process for other bills remains unchanged. Feature calculation engine resource monitoring implements dynamic allocation. When the same type of anomaly occurs three times in a row, the hardware diagnosis process is automatically triggered to isolate the fault.
[0112] It should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that includes a list of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus.
[0113] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to these embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the appended claims and their equivalents.
Claims
1. A method for intelligent identification and verification of financial bills based on machine learning, characterized in that: The method comprises: Acquire the original image data stream of financial bills in real time, monitor the total number of input bills through image acquisition equipment, and determine whether the total number exceeds the preset system processing safety threshold; When the total amount exceeds the system processing safety threshold, a machine learning algorithm is used to process the characteristic elements of the bill image, generate a feature interaction graph, and evaluate the potential error risk level in the bill recognition process based on the feature interaction graph; Apply signal decomposition technology to analyze the distribution pattern of background noise in bill images, reconstruct the noise dynamic characteristic model, and calculate the interference degree index of background noise on key areas of bills based on the noise dynamic characteristic model; Based on the potential error risk level and the interference degree index, decide whether to initiate the overall bill verification optimization process; When the decision is made to start the overall bill verification optimization process, the attribute data of the recognition path of each independent bill image is analyzed to generate the recognition complexity factor of each bill; Based on the identification complexity factor of each bill, the verification priority value of each bill is calculated, and the verification sequence is output according to the verification priority value; The method of using a machine learning algorithm to process the characteristic elements of the bill image, generate a feature interaction graph, and evaluate the potential error risk level in the bill recognition process based on the feature interaction graph specifically includes: Extracting a pre-processed bill image feature vector, wherein the pre-processed bill image feature vector is obtained by an edge enhancement algorithm and a feature dimensionality reduction technique; Input the preprocessed bill image feature vector into the deep neural network model, and train the model to generate a multi-dimensional feature interaction graph; Analyze the correlation strength matrix between feature elements in the multidimensional feature interaction graph and calculate the probability of correlation anomaly; Quantify the potential error risk level based on the associated anomaly probability; The application of signal decomposition technology to analyze the distribution pattern of the background noise of the bill image, reconstruct the noise dynamic characteristic model, and calculate the interference degree index of the background noise on the key area of the bill based on the noise dynamic characteristic model, specifically including: Perform frequency domain conversion on the denoised bill image to extract noise spectrum distribution data; A multi-scale decomposition algorithm is used to process noise spectrum distribution data and construct a spatiotemporal noise dynamic characteristic model; Analyze the noise intensity variation curve in the spatiotemporal noise dynamic characteristic model to generate interference assessment values for key areas; Based on the interference assessment value of the key area, output the interference degree index; The method of analyzing the attribute data of the recognition path of each independent bill image and generating the recognition complexity factor of each bill specifically includes: Calculate the feature similarity difference value of each bill image with other bill images; Identify the conflict area distribution map in the feature similarity difference value and mark the coordinates of the low-definition area; Integrate the number of regions and severity parameters in the conflict area distribution map to generate an identification path attribute dataset; Based on the identification path attribute dataset, the identification complexity factor is calculated.
2. The method for intelligent identification and verification of financial bills based on machine learning according to claim 1, characterized in that: The real-time acquisition of the original image data stream of the financial bill, monitoring the total number of input bills through the image acquisition device, and determining whether the total number exceeds a preset system processing safety threshold specifically includes: Continuously receiving a bill image data packet from an image sensor, wherein the bill image data packet includes a resolution parameter, a color depth parameter, and an image size parameter; Using an image recognition algorithm to classify the bill types in the bill image data packet to determine a list of valid bill categories; Compare the total number of tickets in the valid ticket category list with the system's preset processing capacity limit to generate an input load intensity index; Based on the comparison results of the input load intensity index and the system processing safety threshold, it is determined whether there is a security risk signal for bill recognition.
3. The method for intelligent identification and verification of financial bills based on machine learning according to claim 1, characterized in that: The decision of whether to initiate the overall bill verification optimization process based on the potential error risk level and the interference degree index specifically includes: When the potential error risk level is high and the interference level indicator is high interference level, an instruction to start the overall bill verification optimization process is generated; When the potential error risk level is medium or low, or the interference level indicator is medium or low, an instruction to not start the overall bill verification optimization process is generated; An instruction to start the overall bill verification optimization process or an instruction not to start the overall bill verification optimization process is output as a decision result.
4. The method for intelligent identification and verification of financial bills based on machine learning according to claim 1, characterized in that: The integration of the number of regions and severity parameters in the conflict area distribution map to generate an identification path attribute dataset specifically includes: Count the total number of all marked areas in the conflict area distribution map; Measure the average value of the blur index for each marked area; The total quantity value and the average value of the fuzzy index are fused to construct the recognition path attribute dataset.
5. The method for intelligent identification and verification of financial bills based on machine learning according to claim 4 is characterized in that: The method of calculating the verification priority value of each bill based on the identification complexity factor of each bill and outputting the verification sequence according to the verification priority value specifically includes: Assigning a type weight coefficient to each bill, wherein the type weight coefficient is determined based on the bill category and business importance; Multiply the type weight coefficient by the recognition complexity factor to generate a verification priority score; Compare the verification priority scores of all tickets and sort them from highest to lowest.
6. The method for intelligent identification and verification of financial bills based on machine learning according to claim 1, characterized in that: The real-time acquisition of the original image data stream of the financial bill and monitoring the total number of input bills through the image acquisition device also include: Apply optical character recognition technology to preliminarily analyze text information fragments in the bill image data stream; Store text information fragments in a temporary database, verify the integrity and consistency of text information fragments, and generate data quality reports.
7. The method for intelligent identification and verification of financial bills based on machine learning according to claim 1, characterized in that: The input of the pre-processed bill image feature vector to the deep neural network model and the training model to generate a multi-dimensional feature interaction graph specifically include: A deep neural network model is constructed by combining convolutional neural network with attention mechanism. The deep neural network model is trained to learn the dependency between feature vectors and output a multi-dimensional feature interaction graph to the risk assessment module.
Citation Information
Patent Citations
Bill identification and verification method and device, computer equipment and storage medium
CN116704528A
Big data processing-based document information input method and system
CN119990719A