Bill amount identification method, apparatus and device, and storage medium

By preprocessing and segmentation curve adjustment of bill images, combined with character recognition model, the accuracy and universality of handwritten bill amount recognition are solved, and efficient bill amount extraction is achieved.

CN120340041APending Publication Date: 2025-07-18INDUSTRIAL AND COMMERCIAL BANK OF CHINA
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510499947.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-21
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The prior art is difficult to effectively identify the amount of bills in handwritten form, especially due to insufficient recognition accuracy and universality due to personalized writing styles and the diversity of different bill formats.

Method used

By preprocessing the initial bill image, the initial character segmentation curve is determined, and the segmentation curve is adjusted using positioning auxiliary data and grayscale difference data, the bill amount image is extracted, and then the character recognition model is used for amount recognition.

Benefits of technology

It improves the accuracy and universality of bill amount identification, and can effectively identify bill amounts without considering bill type.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120340041A_ABST
    Figure CN120340041A_ABST
Patent Text Reader

Abstract

The invention discloses a bill amount identification method and device, equipment and a storage medium, and can be applied to the field of financial science and technology. The method comprises the following steps: preprocessing an image area where a bill amount is located in an initial bill image to obtain a to-be-identified bill amount image; determining an initial character segmentation curve according to the to-be-recognized bill amount image; determining positioning auxiliary data and gray difference data according to the to-be-identified bill amount image and the initial character segmentation curve; adjusting the initial character segmentation curve according to the positioning auxiliary data and the gray difference data to obtain a target character segmentation curve, and extracting the bill amount in the bill amount image to be identified from the image background according to the target character segmentation curve to obtain a target bill amount image; and performing character recognition on the target bill amount image through a character recognition model to obtain bill amount data in the initial bill image. According to the technical scheme, the bill amount identification accuracy and universality are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the field of computer technology and can be applied to the field of fintech, and particularly to the field of image segmentation technology. Specifically, the present application relates to a method, apparatus, device and storage medium for recognizing the amount of a bill. Background Art

[0002] All kinds of bills are ubiquitous in commercial scenarios and are the main vouchers for transactions, reimbursements, etc. The informatization of bills mainly uses optical character recognition (OCR) and key information extraction and processing technologies, first recognizing the content of the bill taken in a natural scene and then extracting the key information of the bill.

[0003] However, since the amount of the bill may appear in a handwritten form, the diversity and individuality of handwritten fonts make the accurate recognition of the amount more difficult; everyone's writing style is different and may be affected by various factors such as writing speed and handwriting quality, which increases the difficulty of amount recognition; in addition, different types of bills have different formats, including checks, drafts, bank remittance forms, etc.; these bills may adopt different layouts and arrangements, resulting in different positions and structures of the amount, which leads to the need for different recognition methods for different types of bills, increasing the load on the system when recognizing the amount. Summary of the Invention

[0004] The present application provides a method, apparatus, device and storage medium for recognizing the amount of a bill to improve the accuracy and generality of bill amount recognition.

[0005] According to one aspect of the present application, a method for recognizing the amount of a bill is provided, and the method includes:

[0006] Preprocessing the image area where the bill amount is located in the initial bill image to obtain an image of the bill amount to be recognized; wherein, the image of the bill amount to be recognized is a black-and-white image;

[0007] Determining an initial character segmentation curve according to the image of the bill amount to be recognized; wherein, the initial character segmentation curve refers to a closed curve used to separate the bill amount from the background of the image of the bill amount to be recognized;

[0008] Determining the positioning auxiliary data and grayscale difference data of the image of the bill amount to be recognized according to the image of the bill amount to be recognized and the initial character segmentation curve;

[0009] Adjust the initial character segmentation curve according to the positioning auxiliary data and the grayscale difference data to obtain a target character segmentation curve, and extract the bill amount in the bill amount image to be recognized from the image background according to the target character segmentation curve to obtain a target bill amount image;

[0010] Perform character recognition on the target bill amount image through a character recognition model to obtain the bill amount data in the initial bill image.

[0011] According to another aspect of the present application, there is provided a bill amount recognition device, which includes:

[0012] An image preprocessing module, configured to preprocess the image area where the bill amount is located in the initial bill image to obtain a bill amount image to be recognized; wherein, the bill amount image to be recognized is a black and white image;

[0013] A segmentation curve determination module, configured to determine an initial character segmentation curve according to the bill amount image to be recognized; wherein, the initial character segmentation curve refers to a closed curve used to separate the bill amount from the background of the bill amount image to be recognized;

[0014] An image data determination module, configured to determine the positioning auxiliary data and the grayscale difference data of the bill amount image to be recognized according to the bill amount image to be recognized and the initial character segmentation curve;

[0015] A segmentation curve adjustment module, configured to adjust the initial character segmentation curve according to the positioning auxiliary data and the grayscale difference data to obtain a target character segmentation curve, and extract the bill amount in the bill amount image to be recognized from the image background according to the target character segmentation curve to obtain a target bill amount image;

[0016] A character recognition module, configured to perform character recognition on the target bill amount image through a character recognition model to obtain the bill amount data in the initial bill image.

[0017] According to another aspect of the present application, there is provided an electronic device, which includes:

[0018] One or more processors;

[0019] A memory, configured to store one or more programs;

[0020] When the one or more programs are executed by the one or more processors, the one or more processors implement any one of the bill amount recognition methods provided by the embodiments of the present application.

[0021] According to another aspect of the present application, there is provided a computer-readable storage medium having stored thereon a computer program, which when executed by a processor implements any one of the bill amount recognition methods provided by the embodiments of the present application.

[0022] According to another aspect of the present application, there is provided a computer program product including a computer program, which when executed by a processor implements any one of the bill amount recognition methods provided by the embodiments of the present application.

[0023] In the present application, the image area where the bill amount is located in the initial bill image is preprocessed to obtain a bill amount image to be recognized; wherein, the bill amount image to be recognized is a black-and-white image; according to the bill amount image to be recognized, an initial character segmentation curve is determined; wherein, the initial character segmentation curve refers to a closed curve used to separate the bill amount from the background of the bill amount image to be recognized; according to the bill amount image to be recognized and the initial character segmentation curve, positioning auxiliary data and gray-scale difference data of the bill amount image to be recognized are determined; according to the positioning auxiliary data and the gray-scale difference data, the initial character segmentation curve is adjusted to obtain a target character segmentation curve, and according to the target character segmentation curve, the bill amount in the bill amount image to be recognized is extracted from the image background to obtain a target bill amount image; through a character recognition model, character recognition is performed on the target bill amount image to obtain the bill amount data in the initial bill image. The above technical solution performs segmentation processing on the bill amount image by using a segmentation curve to separately extract the bill amount in the form of an image, and uses a character recognition model to perform character recognition on the extracted bill amount image to determine the required bill amount, and the bill amount can be recognized without considering the type of the bill, which can effectively improve the accuracy and versatility of bill amount recognition. Description of the Drawings

[0024] Figure 1 is a flowchart of a bill amount recognition method provided by Embodiment 1 of the present application;

[0025] Figure 2 is a flowchart of a bill amount recognition method provided by Embodiment 2 of the present application;

[0026] Figure 3 is a schematic structural diagram of a bill amount recognition device provided by Embodiment 3 of the present application;

[0027] Figure 4 is a schematic structural diagram of an electronic device implementing the bill amount recognition method of Embodiment 4 of the present application. Detailed Embodiments

[0028] To enable those skilled in the art to better understand the solution of this application, the following will clearly and completely describe the technical solution in the embodiments of this application in conjunction with the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of this application.

[0029] It should be noted that the terms "first", "second", etc. in the description and claims of this application and the above-mentioned accompanying drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that such data used can be interchanged under appropriate circumstances so that the embodiments of this application described here can be implemented in an order other than those illustrated or described here. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.

[0030] In addition, it should also be noted that in the technical solution of this application, the collection, storage, use, processing, transmission, provision, and disclosure of relevant data such as the initial bill image and the bill amount image to be recognized comply with the provisions of relevant laws and regulations and do not violate public order and good customs.

[0031] Embodiment 1

[0032] Figure 1 is a flowchart of a bill amount recognition method provided according to Embodiment 1 of this application. This embodiment is applicable to the situation of recognizing the bill amount of different types of bills, and can be executed by a bill amount recognition device. The bill amount recognition device can be implemented in the form of hardware and / or software, and the bill amount recognition device can be configured in a computer device, such as a server. As Figure 1 shown, the method includes:

[0033] S110. Preprocess the image area where the bill amount is located in the initial bill image to obtain the bill amount image to be recognized; wherein, the bill amount image to be recognized is a black-and-white image.

[0034] In this embodiment, the initial bill image refers to the bill image that needs to be recognized for the bill amount currently. The initial bill image can be bill images of different types. The bill amount image to be recognized refers to the image after preliminary processing, which only contains the part of the bill amount. A black-and-white image refers to an image whose color mode has only black and white.

[0035] Optionally, perform image denoising on the initial bill image, and crop the image area where the bill amount is located from the denoised initial bill image to obtain an initial bill amount image containing the bill amount; convert the initial bill amount image into a black-and-white image format to obtain the bill amount image to be recognized.

[0036] In this embodiment, image denoising refers to removing unnecessary noise in the image while trying to retain the details and important features of the image; among them, noise is usually random interference introduced by the shooting device, the transmission process, or other environmental factors, and it may be manifested as random pixel changes in the image. The initial bill amount image is cropped from the initial bill image and only contains the image part of the bill amount; by performing regional positioning on the initial image, the area containing the amount information is extracted to obtain this image. The black-and-white image format refers to an image format with only two colors, black and white.

[0037] Furthermore, perform grayscale processing on the initial bill image to obtain a candidate bill amount image in grayscale image format; perform binarization processing on the candidate bill amount image to obtain the bill amount image to be recognized in black-and-white image format.

[0038] In this embodiment, grayscale processing is the process of converting a color image into a grayscale image; a grayscale image is composed of pixels with different grayscale levels (from black to white), and is usually used for image analysis and processing because it reduces the complexity of color information and makes subsequent processing steps simpler and more efficient. The grayscale image format is an image format where each pixel has a grayscale value indicating the degree from black to white; grayscale images are usually used for subsequent image processing operations such as edge detection and segmentation. The candidate bill amount image is an image after grayscale processing and contains the image area initially judged to be possibly the bill amount area. Binarization processing is the process of converting a grayscale image into a black-and-white image; by setting a threshold, the part of the grayscale value in the image greater than the threshold becomes white, and the part less than the threshold becomes black; this processing helps to remove unnecessary details and makes important structures (such as the bill amount) more prominent, facilitating subsequent processing.

[0039] In an optional implementation manner, the bill amount image to be recognized can be obtained by performing image denoising, grayscale processing, binarization processing, and cropping on the initial bill image.

[0040] S120. Determine an initial character segmentation curve according to the bill amount image to be recognized; where the initial character segmentation curve refers to a closed curve used to separate the bill amount from the background of the bill amount image to be recognized.

[0041] In this embodiment, the initial character segmentation curve refers to a preliminarily determined closed curve used to separate the bill amount part from the background; this closed curve can be a rectangular or circular bounding box.

[0042] Optionally, an improved level set algorithm can be used to initialize the initial character segmentation curve according to the image of the bill amount to be recognized.

[0043] In this embodiment, the improved level set algorithm optimizes and enhances the algorithm based on the level set method to improve its performance, accuracy, stability, and applicability; the level set method itself is a powerful mathematical tool for tracking and describing the evolution of dynamic graphical interfaces or surfaces, and has extensive applications in fields such as computer vision, image processing, and fluid mechanics.

[0044] S130. Determine the positioning auxiliary data and gray-scale difference data of the image of the bill amount to be recognized according to the image of the bill amount to be recognized and the initial character segmentation curve.

[0045] In this embodiment, the positioning auxiliary data refers to the data that uses the global information of the image, such as the gradient of the image, to help locate the edge of the bill amount image. The gray-scale difference data refers to the difference data between different gray levels in the image, which is used to push the segmentation curve through the gray-scale difference between the inside and outside regions of the image to ensure that the segmentation is more in line with the local region characteristics of the image.

[0046] S140. Adjust the initial character segmentation curve according to the positioning auxiliary data and the gray-scale difference data to obtain the target character segmentation curve, and extract the bill amount in the image of the bill amount to be recognized from the image background according to the target character segmentation curve to obtain the target bill amount image.

[0047] In this embodiment, the target character segmentation curve refers to the character segmentation curve after adjustment. Compared with the initial character segmentation curve, it is optimized to make the segmentation result more accurate; this curve is used to accurately extract the bill amount part and remove irrelevant background information. The target bill amount image refers to the accurate region extracted from the image of the bill amount to be recognized, which only contains the bill amount part; it should be noted that there is at least one target bill amount image.

[0048] Optionally, adjust the initial character segmentation curve according to the positioning auxiliary data and the gray-scale difference data to obtain a candidate character segmentation curve; if the candidate character segmentation curve meets the image segmentation condition, then determine the candidate character segmentation curve as the target character segmentation curve; if the candidate character segmentation curve does not meet the image segmentation condition, then determine the positioning auxiliary data and the gray-scale difference data of the image of the bill amount to be recognized according to the candidate character segmentation curve and the image of the bill amount to be recognized until the candidate character segmentation curve meets the image segmentation condition or reaches the maximum number of iterations.

[0049] In this embodiment, the candidate character segmentation curve is a new segmentation curve obtained by adjusting the initial character segmentation curve. Guided by the positioning auxiliary data and the grayscale difference data, the candidate character segmentation curve is an optimized candidate result aimed at more precisely segmenting the character region in the image. If the candidate character segmentation curve meets the segmentation conditions, it will become the target character segmentation curve; otherwise, further adjustment is required. The image segmentation conditions are set based on a large number of practices and are manually set according to the actual situation or empirical values. Exemplarily, this condition can be that the total energy change rate in the improved level set algorithm calculation process is less than the change rate threshold. The maximum number of iterations refers to the maximum number of times allowed during the process of adjusting the candidate character segmentation curve. If the target character segmentation curve that meets the image segmentation conditions is still not found within these iterations, the process will stop. The setting of the maximum number of iterations can prevent the algorithm from falling into an infinite loop and limit the consumption of computing resources.

[0050] S150. Use the character recognition model to perform character recognition on the target bill amount image to obtain the bill amount data in the initial bill image.

[0051] In this embodiment, the character recognition model refers to a machine learning model, which is usually used to convert the characters in the image into readable text information. It can be trained based on deep learning technologies (such as convolutional neural networks) and can recognize various handwritten or printed characters. The bill amount data refers to the digital or amount information recognized from the target bill amount image, usually a specific value used to represent the amount of the bill.

[0052] Optionally, input the target bill amount image into the character recognition model for forward propagation to obtain the model bill amount image. Use the character recognition model to compare the target bill amount image and the model bill amount image to obtain the loss value of the model bill amount image. According to the loss value, adjust the character recognition model so as to perform character recognition on the target bill amount image according to the adjusted character recognition model to obtain the bill amount data in the initial bill image.

[0053] In this embodiment, forward propagation is a computational process in a neural network, which refers to the process of input data passing through the calculations of each layer of the neural network and finally outputting the final result. In character recognition, the target bill amount image undergoes forward propagation through the character recognition model, and finally the model outputs the prediction result for this image. The model bill amount image is the prediction result of the character recognition model for the target bill amount image during the forward propagation process. This image contains the recognition output generated by the model based on the input image (target bill amount image), usually manifested as the recognition result of the model for the bill amount area, which may be an image or a structured output marked with the amount data. The loss value (or the value of the loss function) is a numerical value that measures the gap between the model's prediction result and the true result. In character recognition, the loss value represents the error between the prediction result generated by the model (model bill amount image) and the actual content (true bill amount data) of the target bill amount image. Common loss functions include the cross-entropy loss function (for classification problems), etc.

[0054] Specifically, the character recognition model includes an encoder and a decoder. Among them, the encoder includes an input layer, a batch normalization layer, at least one encoding convolutional layer, and at least one max pooling layer. The decoder includes an upsampling layer, a decoding convolutional layer, a dropout layer, and an output layer. The output end of the input layer is connected to the input end of the first encoding convolutional layer among at least one encoding convolutional layer. The output end of each encoding convolutional layer is connected to the input end of the next encoding convolutional layer of this encoding convolutional layer. The output end of at least one encoding convolutional layer is connected to the input end of the batch normalization layer. The output end of at least one encoding convolutional layer is connected to the input end of the decoding convolutional layer in a skip connection manner. The output end of the batch normalization layer is connected to the input end of the first max pooling layer among at least one max pooling layer. The output end of the last max pooling layer among at least one max pooling layer is connected to the input end of the upsampling layer. The output end of the upsampling layer is connected to the input end of the decoding convolutional layer. The output end of the decoding convolutional layer is connected to the input end of the dropout layer. The output end of the dropout layer is connected to the input end of the output layer.

[0055] It can be understood that the encoder of the character recognition model is a convolutional network, which contains repeated convolutional layer operations. After each convolutional layer, there is a batch normalization layer (Batch Normalization, BN), a rectified linear unit (ReLU), and a max pooling layer. The role of the batch normalization layer is to accelerate network convergence and improve gradient dispersion. The BN layer is generally placed after the convolutional layer and before the rectified linear function because the output distribution of ReLU changes as the network learns, and normalization cannot eliminate the variance shift. The output of the convolutional layer is generally a symmetric non-sparse distribution, similar to a Gaussian distribution, and normalizing it can obtain a more stable distribution, accelerating network convergence. The role of the pooling layer includes reducing network parameters and making the convolutional neural network invariant to translation, rotation, scaling, etc. Max pooling means taking the maximum value of the feature points in the neighborhood. The max pooling layer can reduce the shift of the estimated mean caused by the parameter error of the convolutional layer, thereby retaining more texture information. The decoder of the character recognition model is also a convolutional network, which contains repeated convolutional layers. However, different from the encoder, the convolutional layers of the decoder are followed by an upsampling layer, a feature map fusion operation, and a dropout layer, aiming to increase the target detail information by fusing low-level feature information and finally obtaining a segmentation mask of the original image size. The role of the upsampling layer is to gradually enlarge the reduced feature map to the original image size. When the input image passes through the compression network part of the character recognition model, serious detail and information loss will occur. Directly performing an upsampling operation on such features will not well restore the detail information in the image, and detail information is necessary for the segmentation of complex bank system bill images. Therefore, the features before the upsampling operation are particularly important. Therefore, the convolutional layers of the encoder and the convolutional layers of the decoder are connected by means of skip connections. By fusing the features of the same resolution in the lower layer and the feature maps of the same resolution after upsampling, this method can effectively make the feature maps contain richer edge information.

[0056] In the embodiment of the present application, preprocessing is performed on the image area where the bill amount is located in the initial bill image to obtain the bill amount image to be recognized; wherein, the bill amount image to be recognized is a black-and-white image; according to the bill amount image to be recognized, an initial character segmentation curve is determined; wherein, the initial character segmentation curve refers to a closed curve used to separate the bill amount from the background of the bill amount image to be recognized; according to the bill amount image to be recognized and the initial character segmentation curve, positioning auxiliary data and gray level difference data of the bill amount image to be recognized are determined; according to the positioning auxiliary data and the gray level difference data, the initial character segmentation curve is adjusted to obtain a target character segmentation curve, and according to the target character segmentation curve, the bill amount in the bill amount image to be recognized is extracted from the image background to obtain a target bill amount image; through a character recognition model, character recognition is performed on the target bill amount image to obtain the bill amount data in the initial bill image. In the above technical solution, by using a segmentation curve to perform segmentation processing on the bill amount image, the bill amount in the form of an image is separately extracted, and a character recognition model is used to perform character recognition on the extracted bill amount image to determine the required bill amount, and the bill amount can be recognized without considering the type of the bill, which can effectively improve the accuracy and versatility of bill amount recognition.

[0057] Embodiment 2

[0058] Figure 2 FIG. is a flowchart of a bill amount recognition method provided according to Embodiment 2 of the present application. On the basis of the technical solutions of the above embodiments, "according to the bill amount image to be recognized and the initial character segmentation curve, determine the positioning auxiliary data and the gray level difference data of the bill amount image to be recognized" is refined to "perform region division on the bill amount image to be recognized according to the initial character segmentation curve to obtain a first image region located inside the initial character segmentation curve and a second image region located outside the initial character segmentation curve in the bill amount image to be recognized; according to the initial character segmentation curve and the bill amount image to be recognized, determine the first gray level mean value of the first image region and the second gray level mean value of the second image region; according to the bill amount image to be recognized, the first gray level mean value and the second gray level mean value, determine the positioning auxiliary data and the gray level difference data of the bill amount image to be recognized". It should be noted that for the parts not detailed in the embodiments of the present application, reference may be made to the relevant descriptions of other embodiments. As Figure 2 shown, the method includes:

[0059] S210. Perform preprocessing on the image area where the bill amount is located in the initial bill image to obtain the bill amount image to be recognized; wherein, the bill amount image to be recognized is a black-and-white image.

[0060] S220. Determine an initial character segmentation curve based on the image of the amount of the bill to be recognized. Herein, the initial character segmentation curve refers to a closed curve used to separate the amount of the bill from the background of the image of the amount of the bill to be recognized.

[0061] S230. Perform region division on the image of the amount of the bill to be recognized according to the initial character segmentation curve, and obtain a first image region located inside the initial character segmentation curve and a second image region located outside the initial character segmentation curve in the image of the amount of the bill to be recognized.

[0062] In this embodiment, the first image region refers to the image part inside the initial character segmentation curve; this region usually contains numbers or characters related to the amount of the bill; in the algorithm, the first image region is the key region to be recognized. The second image region refers to the image part outside the initial character segmentation curve; this region is usually not a directly relevant recognition region, but may contain background information or other irrelevant image regions.

[0063] S240. Determine a first grayscale mean value of the first image region and a second grayscale mean value of the second image region according to the initial character segmentation curve and the image of the amount of the bill to be recognized.

[0064] In this embodiment, the grayscale mean value refers to the average value of the pixel grayscale values in a certain part of the image. The grayscale value represents the brightness or intensity of each pixel in the image; in this scenario, the grayscale mean values of the first image region and the second image region are used as references for image analysis to help the system identify the brightness differences between different regions.

[0065] S250. Determine positioning auxiliary data and grayscale difference data of the image of the amount of the bill to be recognized according to the image of the amount of the bill to be recognized, the first grayscale mean value, and the second grayscale mean value.

[0066] Optionally, the first grayscale mean value includes a first global grayscale mean value and a first local grayscale mean value; the second grayscale mean value includes a second global grayscale mean value and a second local grayscale mean value.

[0067] Furthermore, calculate the squared differences between all pixel grayscale values of the image of the amount of the bill to be recognized and the first global grayscale mean value and the first local grayscale value respectively to obtain a first global grayscale value error and a first local grayscale value error; calculate the squared differences between all pixel grayscale values of the image of the amount of the bill to be recognized and the second global grayscale mean value and the second local grayscale mean value respectively to obtain a second global grayscale value error and a second local grayscale value error; sum the first global grayscale value error and the second global grayscale value error to obtain the positioning auxiliary data; sum the first local grayscale value error and the second local grayscale value error to obtain the grayscale difference data.

[0068] Optionally, the positioning assistance data is equivalent to the global energy term in the improved level set algorithm and can be determined by the following formula:

[0069] E G (φ)=λ1∫|I i (x)-c1| 2 H ε (φ(x))dx+λ2∫|I i (x)-c2| 2 (1-H ε (φ(x)))dx;

[0070] Where, λ1∫|I i (x)-c1| 2 H ε (φ(x))dx refers to the first global grayscale value error. λ2∫|I i (x)-c2| 2 (1-H ε (φ(x)))dx refers to the second global grayscale value error. λ1 and λ2 are preset weight parameters used to control the balance of the energy terms in the inner and outer regions of the curve. I i (x) refers to the grayscale values of all pixels in the image of the amount of the bill to be recognized. c1 refers to the first global grayscale mean value. H ε (φ(x)) is the regularized Heaviside function used to smoothly divide the inner and outer regions of the curve. c2 refers to the second global grayscale mean value.

[0071] Optionally, the grayscale difference data is equivalent to the local energy term in the improved level set algorithm and can be determined by the following formula:

[0072] E L (φ)=λ3∫|I i (x)-d1| 2 H ε (φ(x))dx+λ4∫|I i (x)-d2| 2 (1-H ε (φ(x)))dx;

[0073] Where, λ3∫|I i (x)-d1| 2 H ε (φ(x))dx refers to the first local grayscale value error. λ4∫|I i (x)-d2| 2 (1-H εThe integral of (φ(x))dx represents the second local grayscale value error. λ3 and λ4 are preset weight parameters used to control the balance of the energy terms inside and outside the curve. d1 represents the first local grayscale mean value. d2 represents the second local grayscale mean value. φ(x) is the level set function, which expresses the dynamically evolving segmentation curve as the zero level set of the function, representing the signed distance from the pixel point x to the segmentation curve. Generally, it is positive inside the curve and negative outside the curve.

[0074] Further, H ε (φ(x)) can be determined by the following formula:

[0075]

[0076] Among them, ε is the regularization parameter used to control the width of the function transition band. represents the arctangent function, which is used to map to the interval (-π / 2, π / 2) to smoothly switch the regions inside and outside the curve and avoid instability in numerical calculations.

[0077] S260. According to the positioning auxiliary data and the grayscale difference data, adjust the initial character segmentation curve to obtain the target character segmentation curve, and based on the target character segmentation curve, extract the bill amount in the bill amount image to be recognized from the image background to obtain the target bill amount image.

[0078] Optionally, perform weighted fusion on the positioning auxiliary data and the grayscale difference data to obtain the target adjustment data of the initial character segmentation curve; aiming at minimizing the target adjustment data, adjust the initial character segmentation curve until the change rate of the target adjustment data or the target adjustment data meets the curve generation condition to obtain the target character segmentation curve.

[0079] In this embodiment, the target adjustment data is used to quantify the matching degree between the current character segmentation curve and the target boundary; it is equivalent to the total energy function in the improved level set algorithm during the calculation process.

[0080] Exemplarily, the target adjustment data can be determined by the following formula:

[0081] E(φ) = αE L (φ) + βE G (φ);

[0082] Among them, α and β are preset weight parameters.

[0083] In an alternative embodiment, after weighted fusion of the positioning assistance data and the grayscale difference data, a regularization term can be introduced to avoid repeated initialization of the level set function and maintain the smoothness of the contour curve during evolution; the positioning assistance data, the grayscale difference data, and the regularization term are weighted and fused to obtain the target adjustment data for the initial character segmentation curve.

[0084] In this embodiment, the regularization term refers to a term used to regularize a certain function or model; the goal of regularization is to introduce additional constraints or penalty terms so that the model can avoid overfitting during optimization and ensure the smoothness or appropriate complexity of the solution.

[0085] S270. Through the character recognition model, perform character recognition on the target bill amount image to obtain the bill amount data in the initial bill image.

[0086] Specifically, through the character recognition model, assign a separate label to each character in the target bill amount image to form a multi-channel label image, and correspond each channel to a category; through the character recognition model, perform character recognition on the label image to obtain the bill amount data in the initial bill image.

[0087] Exemplarily, use the character recognition model to analyze the target bill amount image, identify each character in the image, and assign a separate label to each character; this label can be a number, a symbol, or other characters to be recognized; since each character has a label, you will obtain a multi-channel label image; each channel represents a character category (such as numbers 0-9, decimal points, etc.); this multi-channel image will contain the category information of each character in the image; when generating the label image, each pixel value of the label image will represent the character category at that position (such as numbers 0, 1, 2, or decimal points); once this label image is obtained, the label image can be further recognized through the character recognition model (such as a deep learning-based character recognition model) to extract the amount data in the bill image.

[0088] In the embodiment of the present application, the image area where the bill amount is located in the initial bill image is preprocessed to obtain the bill amount image to be recognized; wherein, the bill amount image to be recognized is a black-and-white image; according to the bill amount image to be recognized, an initial character segmentation curve is determined; wherein, the initial character segmentation curve refers to a closed curve used to separate the bill amount from the background of the bill amount image to be recognized; the bill amount image to be recognized is divided into regions according to the initial character segmentation curve, and a first image region located inside the initial character segmentation curve and a second image region located outside the initial character segmentation curve in the bill amount image to be recognized are obtained; according to the initial character segmentation curve and the bill amount image to be recognized, a first gray-scale average value of the first image region and a second gray-scale average value of the second image region are determined; according to the bill amount image to be recognized, the first gray-scale average value and the second gray-scale average value, positioning auxiliary data and gray-scale difference data of the bill amount image to be recognized are determined; according to the positioning auxiliary data and the gray-scale difference data, the initial character segmentation curve is adjusted to obtain a target character segmentation curve, and according to the target character segmentation curve, the bill amount in the bill amount image to be recognized is extracted from the image background to obtain a target bill amount image; through a character recognition model, character recognition is performed on the target bill amount image to obtain the bill amount data in the initial bill image. In the above technical solution, by using a segmentation curve to perform segmentation processing on the bill amount image, the bill amount in the form of an image is separately extracted, and a character recognition model is used to perform character recognition on the extracted bill amount image to determine the required bill amount. The bill amount can be recognized without considering the type of the bill, which can effectively improve the accuracy and versatility of bill amount recognition.

[0089] Embodiment III

[0090] Figure 3 FIG. 7 is a schematic structural diagram of a bill amount recognition device provided according to Embodiment III of the present application, which is applicable to the case of recognizing the bill amount of different types of bills. The bill amount recognition device can be implemented in the form of hardware and / or software, and the bill amount recognition device can be configured in a computer device, such as a server. As Figure 3 shown, the device includes:

[0091] An image preprocessing module 310, configured to preprocess the image area where the bill amount is located in the initial bill image to obtain a bill amount image to be recognized; wherein, the bill amount image to be recognized is a black-and-white image;

[0092] A segmentation curve determination module 320, configured to determine an initial character segmentation curve according to the bill amount image to be recognized; wherein, the initial character segmentation curve refers to a closed curve used to separate the bill amount from the background of the bill amount image to be recognized;

[0093] An image data determination module 330, configured to determine positioning auxiliary data and gray-scale difference data of the to-be-recognized bill amount image according to the to-be-recognized bill amount image and the initial character segmentation curve;

[0094] A segmentation curve adjustment module 340, configured to adjust the initial character segmentation curve according to the positioning auxiliary data and the gray-scale difference data to obtain a target character segmentation curve, and extract the bill amount in the to-be-recognized bill amount image from the image background according to the target character segmentation curve to obtain a target bill amount image;

[0095] A character recognition module 350, configured to perform character recognition on the target bill amount image through a character recognition model to obtain the bill amount data in the initial bill image.

[0096] In an embodiment of the present application, preprocessing is performed on the image area where the bill amount is located in the initial bill image to obtain a to-be-recognized bill amount image; wherein, the to-be-recognized bill amount image is a black-and-white image; according to the to-be-recognized bill amount image, an initial character segmentation curve is determined; wherein, the initial character segmentation curve refers to a closed curve used to separate the bill amount from the background of the to-be-recognized bill amount image; according to the to-be-recognized bill amount image and the initial character segmentation curve, positioning auxiliary data and gray-scale difference data of the to-be-recognized bill amount image are determined; according to the positioning auxiliary data and the gray-scale difference data, the initial character segmentation curve is adjusted to obtain a target character segmentation curve, and according to the target character segmentation curve, the bill amount in the to-be-recognized bill amount image is extracted from the image background to obtain a target bill amount image; through a character recognition model, character recognition is performed on the target bill amount image to obtain the bill amount data in the initial bill image. In the above technical solution, by using a segmentation curve to perform segmentation processing on the bill amount image, the bill amount in the form of an image is separately extracted, and a character recognition model is used to perform character recognition on the extracted bill amount image to determine the required bill amount, and the bill amount can be recognized without considering the type of the bill, which can effectively improve the accuracy and generality of bill amount recognition.

[0097] Optionally, the image data determination module 330 includes:

[0098] A region division unit, configured to divide the to-be-recognized bill amount image according to the initial character segmentation curve to obtain a first image region located inside the initial character segmentation curve and a second image region located outside the initial character segmentation curve in the to-be-recognized bill amount image;

[0099] A gray-scale mean determination unit, configured to determine a first gray-scale mean of the first image region and a second gray-scale mean of the second image region according to the initial character segmentation curve and the to-be-recognized bill amount image;

[0100] An image data determination unit, configured to determine positioning auxiliary data and gray - level difference data of a bill amount image to be recognized according to a first image region, a second image region, a first gray - level mean value, and a second gray - level mean value.

[0101] Optionally, the first gray - level mean value includes a first global gray - level mean value and a first local gray - level mean value; the second gray - level mean value includes a second global gray - level mean value and a second local gray - level mean value; correspondingly, the image data determination unit is specifically configured to:

[0102] Calculate the squared differences between all pixel gray - level values of the first image region and the first global gray - level mean value and the first local gray - level value respectively, to obtain a first global gray - level value error and a first local gray - level value error;

[0103] Calculate the squared differences between all pixel gray - level values of the second image region and the second global gray - level mean value and the second local gray - level mean value respectively, to obtain a second global gray - level value error and a second local gray - level value error;

[0104] Sum the first global gray - level value error and the second global gray - level value error to obtain the positioning auxiliary data;

[0105] Sum the first local gray - level value error and the second local gray - level value error to obtain the gray - level difference data.

[0106] Optionally, the segmentation curve adjustment module 340 is specifically configured to:

[0107] Perform weighted fusion on the positioning auxiliary data and the gray - level difference data to obtain target adjustment data of an initial character segmentation curve;

[0108] Adjust the initial character segmentation curve with the goal of minimizing the target adjustment data until the change rate of the target adjustment data or the target adjustment data meets the curve generation condition, to obtain a target character segmentation curve.

[0109] Optionally, the character recognition module 350 is specifically configured to:

[0110] Input the target bill amount image into a character recognition model for forward propagation to obtain a model bill amount image;

[0111] Compare the target bill amount image and the model bill amount image through the character recognition model to obtain a loss value of the model bill amount image;

[0112] Adjust the character recognition model according to the loss value, so as to perform character recognition on the target bill amount image according to the adjusted character recognition model to obtain the bill amount data in the initial bill image.

[0113] Optionally, the character recognition model includes an encoder and a decoder; wherein, the encoder includes an input layer, a batch normalization layer, at least one encoding convolutional layer, and at least one max pooling layer; the decoder includes an upsampling layer, a decoding convolutional layer, a dropout layer, and an output layer;

[0114] The output end of the input layer is connected to the input end of the first encoding convolutional layer in at least one encoding convolutional layer; the output end of each encoding convolutional layer is connected to the input end of the next encoding convolutional layer of this encoding convolutional layer; the output end of at least one encoding convolutional layer is connected to the input end of the batch normalization layer; the output end of at least one encoding convolutional layer is connected to the input end of the decoding convolutional layer in a skip connection manner; the output end of the batch normalization layer is connected to the input end of the first max pooling layer in at least one max pooling layer; the output end of the last max pooling layer in at least one max pooling layer is connected to the input end of the upsampling layer; the output end of the upsampling layer is connected to the input end of the decoding convolutional layer; the output end of the decoding convolutional layer is connected to the input end of the dropout layer; the output end of the dropout layer is connected to the input end of the output layer.

[0115] The bill amount recognition device provided by the embodiments of the present application can execute the bill amount recognition method provided by any embodiment of the present application, and has the corresponding functional modules and beneficial effects for executing each bill amount recognition method.

[0116] According to the embodiments of the present application, the present application also provides an electronic device, a readable storage medium, and a computer program product.

[0117] Embodiment 4

[0118] Figure 4 It is a schematic structural diagram of an electronic device 410 for implementing the bill amount recognition method of the embodiments of the present application. The electronic device is intended to represent various forms of digital computers, such as, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, personal digital processors, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present application described herein and / or required by the present application.

[0119] As Figure 4As shown, the electronic device 410 includes at least one processor 411 and a memory communicatively connected to the at least one processor 411, such as a read-only memory (ROM) 412, a random access memory (RAM) 413, etc. The memory stores a computer program executable by the at least one processor. The processor 411 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 412 or the computer program loaded from the storage unit 418 into the random access memory (RAM) 413. In the RAM 413, various programs and data required for the operation of the electronic device 410 can also be stored. The processor 411, the ROM 412, and the RAM 413 are connected to each other via a bus 414. An input / output (I / O) interface 415 is also connected to the bus 414.

[0120] Multiple components in the electronic device 410 are connected to the I / O interface 415, including: an input unit 416, such as a keyboard, a mouse, etc.; an output unit 417, such as various types of displays, speakers, etc.; a storage unit 418, such as a magnetic disk, an optical disc, etc.; and a communication unit 419, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 419 allows the electronic device 410 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0121] The processor 411 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the processor 411 include but are not limited to a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The processor 411 executes the various methods and processes described above, such as the bill amount recognition method.

[0122] In some embodiments, the bill amount recognition method can be implemented as a computer program tangibly embodied in a computer-readable storage medium, such as the storage unit 418. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 410 via the ROM 412 and / or the communication unit 419. When the computer program is loaded into the RAM 413 and executed by the processor 411, one or more steps of the bill amount recognition method described above can be executed. Alternatively, in other embodiments, the processor 411 can be configured for the bill amount recognition method by any other appropriate means (e.g., by means of firmware).

[0123] The various embodiments of the systems and techniques described above in this specification can be implemented in digital electronic circuitry, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on a chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be a special-purpose or general-purpose programmable processor that receives data and instructions from, and transmits data and instructions to, a storage system, at least one input device, and at least one output device.

[0124] The computer programs for implementing the methods of this application can be written in any combination of one or more programming languages. These computer programs can be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus, such that the computer programs, when executed by the processor, cause the functions / operations specified in the flowchart and / or block diagram to be implemented. The computer programs can be executed entirely on the machine, partly on the machine, as a stand-alone software package partly on the machine and partly on a remote machine or entirely on the remote machine or server.

[0125] In the context of this application, a computer-readable storage medium can be a tangible medium that can contain, or store a computer program for use by or in connection with an instruction execution system, apparatus, or device. The computer-readable storage medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. Alternatively, the computer-readable storage medium can be a machine-readable signal medium. A more specific example of the machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0126] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) through which the user can provide input to the electronic device. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0127] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), blockchain network, and the Internet.

[0128] The computing system can include a client and a server. The client and the server are generally far from each other and typically interact through a communication network. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system and solves the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services.

[0129] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in this application can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of this application can be achieved, and no limitation is imposed herein.

[0130] The above specific embodiments do not constitute a limitation on the protection scope of this application. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this application shall be included within the protection scope of this application.

Claims

1. A method for recognizing the amount of a bill, characterized in that, Including: Preprocess the image area where the bill amount is located in the initial bill image to obtain the bill amount image to be recognized; wherein, the bill amount image to be recognized is a black and white image; Determine an initial character segmentation curve according to the bill amount image to be recognized; wherein, the initial character segmentation curve refers to a closed curve used to separate the bill amount from the background of the bill amount image to be recognized; Determine the positioning auxiliary data and gray-scale difference data of the bill amount image to be recognized according to the bill amount image to be recognized and the initial character segmentation curve; Adjust the initial character segmentation curve according to the positioning auxiliary data and the gray-scale difference data to obtain a target character segmentation curve, and extract the bill amount in the bill amount image to be recognized from the image background according to the target character segmentation curve to obtain a target bill amount image; Perform character recognition on the target bill amount image through a character recognition model to obtain the bill amount data in the initial bill image.

2. The method according to claim 1, characterized in that Determine the positioning auxiliary data and gray-scale difference data of the bill amount image to be recognized according to the bill amount image to be recognized and the initial character segmentation curve, including: Divide the bill amount image to be recognized into regions according to the initial character segmentation curve to obtain a first image region located inside the initial character segmentation curve and a second image region located outside the initial character segmentation curve in the bill amount image to be recognized; Determine a first gray-scale mean value of the first image region and a second gray-scale mean value of the second image region according to the initial character segmentation curve and the bill amount image to be recognized; Determine the positioning auxiliary data and gray-scale difference data of the bill amount image to be recognized according to the bill amount image to be recognized, the first gray-scale mean value and the second gray-scale mean value.

3. According to the method described in claim 2, the mean value lies in that the first gray-scale mean value includes a first global gray-scale mean value and a first local gray-scale mean value; the second gray-scale mean value includes a second global gray-scale mean value and a second local gray-scale mean value; correspondingly, determining the positioning auxiliary data and gray-scale difference data of the bill amount image to be recognized according to the bill amount image to be recognized, the first gray-scale mean value and the second gray-scale mean value includes: Calculate the square differences between all pixel gray-scale values of the bill amount image to be recognized and the first global gray-scale mean value and the first local gray-scale value respectively to obtain a first global gray-scale value error and a first local gray-scale value error; Calculate the square differences between all pixel gray-scale values of the bill amount image to be recognized and the second global gray-scale mean value and the second local gray-scale mean value respectively to obtain a second global gray-scale value error and a second local gray-scale value error; Sum the first global gray-scale value error and the second global gray-scale value error to obtain positioning auxiliary data; Sum the first local gray-scale value error and the second local gray-scale value error to obtain gray-scale difference data.

4. The method according to claim 1, wherein Adjust the initial character segmentation curve according to the positioning auxiliary data and the gray-scale difference data to obtain a target character segmentation curve, including: Perform weighted fusion on the positioning auxiliary data and the grayscale difference data to obtain the target adjustment data of the initial character segmentation curve; With the goal of minimizing the target adjustment data, adjust the initial character segmentation curve until the change rate of the target adjustment data or the target adjustment data meets the curve generation condition to obtain the target character segmentation curve.

5. The method according to claim 1, wherein The character recognition of the target bill amount image through the character recognition model to obtain the bill amount data in the initial bill image includes: Input the target bill amount image into the character recognition model for forward propagation to obtain the model bill amount image; Compare the target bill amount image and the model bill amount image through the character recognition model to obtain the loss value of the model bill amount image; Adjust the character recognition model according to the loss value to perform character recognition on the target bill amount image according to the adjusted character recognition model to obtain the bill amount data in the initial bill image.

6. The method according to claim 5, wherein The character recognition model includes an encoder and a decoder; wherein, the encoder includes an input layer, a batch normalization layer, at least one encoding convolutional layer, and at least one max pooling layer; the decoder includes an upsampling layer, a decoding convolutional layer, a dropout layer, and an output layer; The output end of the input layer is connected to the input end of the first encoding convolutional layer in the at least one encoding convolutional layer; the output end of each encoding convolutional layer is connected to the input end of the next encoding convolutional layer of this encoding convolutional layer; the output end of the at least one encoding convolutional layer is connected to the input end of the batch normalization layer; the output end of the at least one encoding convolutional layer is connected to the input end of the decoding convolutional layer in a skip connection manner; the output end of the batch normalization layer is connected to the input end of the first max pooling layer in the at least one max pooling layer; the output end of the last max pooling layer in the at least one max pooling layer is connected to the input end of the upsampling layer; the output end of the upsampling layer is connected to the input end of the decoding convolutional layer; the output end of the decoding convolutional layer is connected to the input end of the dropout layer; the output end of the dropout layer is connected to the input end of the output layer.

7. A bill amount recognition device, characterized in that, Include: An image preprocessing module for preprocessing the image area where the bill amount is located in the initial bill image to obtain the bill amount image to be recognized; wherein, the bill amount image to be recognized is a black and white image; A segmentation curve determination module for determining an initial character segmentation curve according to the bill amount image to be recognized; wherein, the initial character segmentation curve refers to a closed curve used to separate the bill amount from the background of the bill amount image to be recognized; An image data determination module for determining the positioning auxiliary data and the grayscale difference data of the bill amount image to be recognized according to the bill amount image to be recognized and the initial character segmentation curve; A splitting curve adjustment module, configured to adjust the initial character splitting curve according to the positioning auxiliary data and the grayscale difference data to obtain a target character splitting curve, and extract the bill amount in the bill amount image to be recognized from the image background according to the target character splitting curve to obtain a target bill amount image; A character recognition module, configured to perform character recognition on the target bill amount image through a character recognition model to obtain the bill amount data in the initial bill image.

8. An electronic device, characterized in that, Comprising: One or more processors; A memory, configured to store one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the bill amount recognition method according to any one of claims 1-6.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by a processor, it implements the bill amount recognition method according to any one of claims 1-6.

10. A computer program product, comprising a computer program, where the computer program implements the bill amount recognition method according to any one of claims 1-6 when executed by a processor.

Citation Information

Cited By

  • Financial document intelligent verification method and system

    CN121505641A