Business voucher inspection method, device, equipment, medium and program product

By using the text recognition model to be identified and tested in the image in the business voucher inspection, the problems of long inspection time and error-prone in the prior art are solved, and fast and accurate automated inspection of business vouchers is achieved.

CN119992583APending Publication Date: 2025-05-13INDUSTRIAL AND COMMERCIAL BANK OF CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510208055.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-25
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

In the prior art, business credentials are checked for a long time and are prone to errors, resulting in slow business processing and poor customer experience.

Method used

The target cutting area is determined by the voucher type of the target voucher image, the image is cut, and the text recognition model is used to identify the text to be tested, and abnormality test is performed to realize automated inspection.

Benefits of technology

It improves the inspection efficiency of business certificates, reduces the time and error rate of manual inspection, and improves customer experience and business processing efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119992583A_ABST
    Figure CN119992583A_ABST
Patent Text Reader

Abstract

The invention provides a business voucher inspection method, device and equipment, a medium and a program product, which can be applied to the technical field of artificial intelligence and the field of financial science and technology. The method comprises the steps that at least one target cutting area is determined based on the voucher type of a target voucher image, the voucher type is determined based on attribute information of the target voucher image, and the target cutting area is the area where a to-be-detected text in the target voucher image is located in the target voucher image; based on the at least one target cutting area, cutting the target voucher image to obtain at least one cutting block image; respectively inputting the at least one cut block image into a text recognition model to obtain a to-be-detected text corresponding to the at least one cut block image; and performing anomaly detection on the to-be-detected text corresponding to the at least one cut block image to obtain a detection result of the target voucher image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the fields of artificial intelligence technology and financial technology, and specifically to a business credential verification method, device, equipment, medium and program product. Background Art

[0002] During the business processing process, the enterprise needs to check the business vouchers provided by the customer. Only after the business vouchers have been checked and found to be correct can the subsequent business processing be carried out. In the relevant technology, the staff usually check the business vouchers. However, due to the large number of business voucher elements, each business voucher may take up a lot of staff's time, especially during the peak business period, it is easy to have errors, omissions, etc., which will cause problems in the subsequent business processing.

[0003] In the process of realizing the concept of the present disclosure, the inventors found that there are at least the following problems in the related art: the related art usually adopts a manual method to check the business vouchers. Due to the large number of business voucher elements, it is easy for the inspection time to be too long and the inspection to be wrong, which leads to problems such as slow business processing and poor customer experience. Summary of the invention

[0004] In view of the above problems, the present disclosure provides a business credential verification method, apparatus, device, medium and program product.

[0005] According to one aspect of the present disclosure, a business voucher verification method is provided, comprising: determining at least one target cutting area based on a voucher type of a target voucher image, wherein the voucher type is determined based on attribute information of the target voucher image, and the target cutting area is an area in the target voucher image where text to be verified in the target voucher image is located; cutting the target voucher image based on the at least one target cutting area to obtain at least one cutting block image; inputting the at least one cutting block image into a text recognition model respectively to obtain text to be verified corresponding to each of the at least one cutting block images; and performing abnormality inspection on the text to be verified corresponding to each of the at least one cutting block images respectively to obtain an inspection result of the target voucher image.

[0006] Another aspect of the present disclosure provides a business voucher verification device, including: an area determination module, used to determine at least one target cutting area based on the voucher type of a target voucher image, wherein the voucher type is determined based on the attribute information of the target voucher image, and the target cutting area is the area in the target voucher image where the text to be verified in the target voucher image is located; a cutting module, used to cut the target voucher image based on the at least one target cutting area to obtain at least one cutting block image; a text recognition module, used to input the at least one cutting block image into a text recognition model respectively to obtain the text to be verified corresponding to each of the at least one cutting block images; and a verification module, used to perform abnormality inspection on the text to be verified corresponding to each of the at least one cutting block images respectively to obtain the inspection result of the target voucher image.

[0007] Another aspect of the present disclosure provides an electronic device, comprising: one or more processors; and a memory for storing one or more computer programs, wherein the one or more processors execute the one or more computer programs to implement the steps of the above method.

[0008] Another aspect of the present disclosure further provides a computer-readable storage medium having a computer program or instruction stored thereon, which implements the steps of the above method when the computer program or instruction is executed by a processor.

[0009] Another aspect of the present disclosure further provides a computer program product, including a computer program or instructions, which implement the steps of the above method when executed by a processor.

[0010] According to the business voucher verification method disclosed in the present invention, the target cutting area where the text to be verified in the target voucher image is located is determined by the voucher type of the target voucher image, and the target voucher image is cut based on the target cutting area to obtain at least one cutting block image. Each cutting block image is recognized by a text recognition model to obtain at least one text to be verified, and at least one text to be verified is respectively tested for abnormalities, so as to obtain an automated inspection result for the target voucher image more quickly. Since the area where the text to be verified in the target voucher image is located is determined by the voucher type of the target voucher image, and the cutting block image cut according to the area is recognized by the text recognition model, it is possible to realize subsequent automated and more accurate abnormality verification of the text to be verified, at least partially solving the technical problems of long inspection time and easy error in the related art, realizing automated inspection of business vouchers, and improving the overall business processing efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] The above contents and other purposes, features and advantages of the present disclosure will become more apparent through the following description of the embodiments of the present disclosure with reference to the accompanying drawings, in which:

[0012] Figure 1 The application scenario diagram of the business credential verification method, apparatus, device, medium and program product according to the embodiment of the present disclosure is schematically shown;

[0013] Figure 2 A flowchart of a business credential verification method according to an embodiment of the present disclosure is schematically shown;

[0014] Figure 3 A schematic diagram schematically shows a plurality of cut block images according to an embodiment of the present disclosure;

[0015] Figure 4 The structural diagram of the text recognition model according to the embodiment of the present disclosure is schematically shown;

[0016] Figure 5 The schematic diagram shows the structure of the second convolutional layer of the text recognition model according to the embodiment of the present disclosure;

[0017] Figure 6 The flowchart of determining the extended sub-area according to the embodiment of the present disclosure is schematically shown;

[0018] Figure 7 A flowchart of a business credential verification method according to another embodiment of the present disclosure is schematically shown;

[0019] Figure 8 A schematic diagram of a structure of a business credential verification device according to an embodiment of the present disclosure is shown; and

[0020] Fig. 9 A block diagram of an electronic device suitable for implementing a business credential verification method according to an embodiment of the present disclosure is schematically shown. DETAILED DESCRIPTION

[0021] Hereinafter, embodiments of the present disclosure will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are exemplary only and are not intended to limit the scope of the present disclosure. In the following detailed description, for ease of explanation, many specific details are set forth to provide a comprehensive understanding of the embodiments of the present disclosure. However, it is apparent that one or more embodiments may also be implemented without these specific details. In addition, in the following description, descriptions of known structures and technologies are omitted to avoid unnecessary confusion of the concepts of the present disclosure.

[0022] The terms used herein are only for describing specific embodiments and are not intended to limit the present disclosure. The terms "comprise", "include", etc. used herein indicate the existence of the features, steps, operations and / or components, but do not exclude the existence or addition of one or more other features, steps, operations or components.

[0023] All terms (including technical and scientific terms) used herein have the meanings commonly understood by those skilled in the art unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification and should not be interpreted in an idealized or overly rigid manner.

[0024] When using expressions such as "at least one of A, B, and C, etc.", they should generally be interpreted according to the meaning of the expression commonly understood by those skilled in the art (for example, "a system having at least one of A, B, and C" should include but is not limited to a system having A alone, B alone, C alone, A and B, A and C, B and C, and / or A, B, C, etc.).

[0025] The business credential verification method and device disclosed herein can be used in the fields of artificial intelligence technology and financial technology, and can also be used in any field other than the fields of artificial intelligence technology and financial technology, such as the field of computer technology, etc. There is no limitation on the application field of the business credential verification method and device disclosed herein.

[0026] In the technical solution of the present disclosure, the user information (including but not limited to user personal information, user image information, user device information, such as location information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved are all information and data authorized by the user or fully authorized by all parties, and the collection, storage, use, processing, transmission, provision, disclosure and application of the relevant data comply with relevant laws, regulations and standards, take necessary confidentiality measures, do not violate public order and good morals, and provide corresponding operation entrances for users to choose to authorize or refuse.

[0027] In the scenario of using personal information for automated decision-making, the methods, devices, and systems provided by the embodiments of the present disclosure provide users with corresponding operation portals for users to choose to agree or reject the automated decision-making results; if the user chooses to reject, the expert decision-making process will be entered. The expression "automated decision-making" here refers to the activity of automatically analyzing and evaluating a person's behavioral habits, interests and hobbies, or economic, health, credit status, etc. through computer programs, and making decisions. The expression "expert decision-making" here refers to the activity of making decisions by people who specialize in a certain field, have specialized experience, knowledge and skills, and have reached a certain level of professionalism.

[0028] During the research, it was found that many enterprises, such as banks, often need to handle centralized payment business in the daily work of their outlets. The business needs to accept business vouchers with customer reserved seals, such as power of attorney, transfer checks, cash checks, etc. When accepting business, it is first necessary to check whether the business voucher provided by the customer is filled in correctly, such as whether the payment password has 16 digits, whether the entrustment date of the business power of attorney is the acceptance date, etc. Often, a voucher needs to check multiple elements, which takes up a lot of staff time, resulting in long waiting time for customers and poor experience.

[0029] An embodiment of the present disclosure provides a business credential verification method, comprising: determining at least one target cutting area based on the credential type of a target credential image, wherein the credential type is determined based on the attribute information of the target credential image, and the target cutting area is the area in the target credential image where the text to be verified in the target credential image is located; based on the at least one target cutting area, cutting the target credential image to obtain at least one cutting block image; inputting the at least one cutting block image into a text recognition model respectively to obtain the text to be verified corresponding to each of the at least one cutting block images; performing abnormality inspection on the text to be verified corresponding to each of the at least one cutting block images respectively to obtain the inspection result of the target credential image.

[0030] Figure 1 The application scenario diagram of the business credential verification method, apparatus, device, medium and program product according to the embodiments of the present disclosure is schematically shown.

[0031] like Figure 1 As shown, the application scenario 100 according to this embodiment may include a first terminal device 101, a second terminal device 102, a third terminal device 103, a network 104, and a server 105. The network 104 is used to provide a medium for a communication link between the first terminal device 101, the second terminal device 102, the third terminal device 103, and the server 105. The network 104 may include various connection types, such as wired, wireless communication links, or optical fiber cables, etc.

[0032] The user can use the first terminal device 101, the second terminal device 102, and the third terminal device 103 to interact with the server 105 through the network 104 to receive or send messages, etc. Various communication client applications can be installed on the first terminal device 101, the second terminal device 102, and the third terminal device 103, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social platform software, etc. (only for example).

[0033] The first terminal device 101, the second terminal device 102, and the third terminal device 103 may be various electronic devices having display screens and supporting web browsing, including but not limited to smart phones, tablet computers, laptop computers, desktop computers, and the like.

[0034] The server 105 may be a server that provides various services, such as a background management server (only as an example) that provides support for websites browsed by users using the first terminal device 101, the second terminal device 102, and the third terminal device 103. The background management server may analyze and process the received data such as user requests, and feed back the processing results (such as web pages, information, or data obtained or generated according to user requests) to the terminal device.

[0035] It should be noted that the business credential verification method provided in the embodiment of the present disclosure can generally be executed by the server 105. Accordingly, the business credential verification device provided in the embodiment of the present disclosure can generally be set in the server 105. The business credential verification method provided in the embodiment of the present disclosure can also be executed by a server or server cluster that is different from the server 105 and can communicate with the first terminal device 101, the second terminal device 102, the third terminal device 103 and / or the server 105. Accordingly, the business credential verification device provided in the embodiment of the present disclosure can also be set in a server or server cluster that is different from the server 105 and can communicate with the first terminal device 101, the second terminal device 102, the third terminal device 103 and / or the server 105.

[0036] It should be understood that Figure 1 The number of terminal devices, networks and servers in the embodiment is only for illustration. Any number of terminal devices, networks and servers may be provided according to implementation requirements.

[0037] The following will be based on Figure 1 The scene described by Figure 2~Figure 7 The business credential verification method of the disclosed embodiment is described in detail.

[0038] Figure 2 The flowchart of the business credential verification method according to the embodiment of the present disclosure is schematically shown.

[0039] like Figure 2 As shown, the method includes operations S210 to S240.

[0040] In operation S210, at least one target cutting area is determined based on the credential type of the target credential image, wherein the credential type is determined based on attribute information of the target credential image, and the target cutting area is an area in the target credential image where the text to be verified in the target credential image is located.

[0041] In operation S220, the target voucher image is cut based on at least one target cutting area to obtain at least one cut block image.

[0042] In operation S230, at least one cut-out block image is respectively input into a text recognition model to obtain a text to be checked corresponding to each of the at least one cut-out block images.

[0043] In operation S240, the texts to be checked corresponding to the at least one cut-out block image are respectively checked for abnormalities to obtain a check result of the target voucher image.

[0044] According to an embodiment of the present disclosure, the target voucher image may be an image obtained by scanning and pre-processing a paper version of the business voucher to be processed.

[0045] According to the embodiments of the present disclosure, there is no limitation on the method for determining the credential type. The credential type may be determined by matching the attribute information of the target credential image with the preset attribute information. The preset attribute information may be associated with the preset credential type. Thus, after obtaining the preset attribute information that matches the attribute information, the preset credential type associated with the preset attribute information is determined to be the credential type of the target credential image.

[0046] According to an embodiment of the present disclosure, the attribute information is not limited and may include size information. In some embodiments, the attribute information may also be identification information.

[0047] According to an embodiment of the present disclosure, the target cutting area may be the area where the entire text of the text to be verified is located in the target voucher image.

[0048] According to an embodiment of the present disclosure, a target cutting area is discovered in the target credential image according to the credential type of the target credential image.

[0049] According to an embodiment of the present disclosure, the target voucher image may be cut along at least one target cutting area to obtain at least one cut block image.

[0050] According to the embodiments of the present disclosure, there is no limitation on the implementation method of the text recognition model, and the text recognition model can be implemented by a multi-layer neural network, for example, a recurrent neural network combined with a convolutional neural network.

[0051] According to an embodiment of the present disclosure, the text recognition model can recognize the text to be checked included in the cut block image.

[0052] According to the embodiments of the present disclosure, the manner of performing abnormality inspection on the text to be inspected is not limited, and the text to be inspected may be inspected using preset inspection rules corresponding to the text to be inspected.

[0053] According to an embodiment of the present disclosure, at least one text to be checked can be checked in parallel, thereby improving the checking speed.

[0054] According to an embodiment of the present disclosure, by summarizing the verification sub-results of at least one text to be verified, the verification result of the target credential image can be obtained.

[0055] According to the business voucher verification method disclosed in the present invention, the target cutting area where the text to be verified in the target voucher image is located is determined by the voucher type of the target voucher image, and the target voucher image is cut based on the target cutting area to obtain at least one cutting block image. Each cutting block image is recognized by a text recognition model to obtain at least one text to be verified, and at least one text to be verified is respectively checked for abnormalities, so as to obtain an automated inspection result for the target voucher image more quickly. Since the area where the text to be verified in the target voucher image is located is determined by the voucher type of the target voucher image, and the cutting block image cut according to the area is recognized by a text recognition model, so that the text to be verified can be automatically and more accurately checked for abnormalities in the future, thereby at least partially solving the technical problems of long inspection time and easy error in inspection in the related art, realizing automated inspection of business vouchers, and improving the overall business processing efficiency.

[0056] Figure 3 A schematic diagram schematically shows a plurality of cut-out block images according to an embodiment of the present disclosure.

[0057] like Figure 3 As shown, the target voucher image can be cut to obtain multiple cut block images, each of which is an independent image.

[0058] According to the embodiments of the present disclosure, Figure 3 The target credential images shown are for illustration only, and the content and format of different credential types may be different in specific scenarios.

[0059] According to an embodiment of the present disclosure, determining at least one target cutting area based on the credential type of the target credential image may include the following operations.

[0060] Based on the credential type of the target credential image, at least one preset cutting area corresponding to the credential type is determined, wherein at least part of the text to be checked is located in the preset cutting area, and the preset cutting area includes a plurality of pixels; for each preset cutting area, based on the grayscale values ​​of the plurality of pixels, a starting sub-area is determined from the preset cutting area; the area is expanded with the starting sub-area as the origin to obtain M expanded sub-areas, wherein M is an integer greater than or equal to 0; based on the M expanded sub-areas and the starting sub-area, a target cutting area is obtained.

[0061] According to an embodiment of the present disclosure, the preset cutting area may be an area where at least part of the text to be checked is located in the target voucher image. In some embodiments, the preset cutting area may be an area where a box or a table where the text to be checked is located in the target voucher image, such as Figure 3 The table where 123456 is located.

[0062] According to an embodiment of the present disclosure, in some embodiments, a preset cutting area may also be used as a target cutting area.

[0063] According to an embodiment of the present disclosure, the starting sub-region may be a starting point in the process of exploring the target cutting region.

[0064] According to an embodiment of the present disclosure, in each preset cutting area, a starting sub-area where the text to be checked may exist can be determined by the grayscale values ​​of multiple pixels, and the average grayscale value of the starting sub-area can be used as a reference, and the area can be expanded with the position of the starting sub-area as the origin, thereby obtaining M expanded sub-areas.

[0065] According to an embodiment of the present disclosure, the target cutting area can be obtained by splicing the extended sub-area and the initial sub-area.

[0066] According to an embodiment of the present disclosure, a preset cutting area where at least part of the text to be checked is located is determined by the voucher type of the target voucher image, and a starting sub-area is determined in the preset cutting area, and the area is expanded with the starting sub-area as the origin, thereby obtaining a target cutting area where all the text to be checked is located, thereby avoiding the situation where some business vouchers are handwritten and the customer fills in the text outside the box or table and cannot be recognized if the box or table where the text to be checked is located is directly intercepted.

[0067] According to an embodiment of the present disclosure, region expansion is performed with the starting sub-region as the origin to obtain M expanded sub-regions, which may include the following operations.

[0068] Determine an inscribed circle sub-region of a starting sub-region; based on a preset radius and an expansion threshold, determine at least one candidate sub-region tangent to the edge of the inscribed circle sub-region, wherein different candidate sub-regions have different orientations relative to the inscribed circle sub-region, and the coordinates of the pixel points located at the edge of the candidate sub-region are less than or equal to the expansion threshold; for each candidate sub-region, based on the grayscale value of at least one pixel point included in the candidate sub-region, determine the average grayscale value and the average gradient value of the candidate sub-region; based on the average grayscale value and the average gradient value of the candidate sub-region, perform scalability judgment on the candidate sub-region to obtain a judgment result; when the judgment result indicates that the candidate sub-region is an extended sub-region, use the extended sub-region as a new starting sub-region; when the new starting sub-region meets the outward expansion condition, repeat the above operation.

[0069] According to the embodiments of the present disclosure, the expansion threshold is used to characterize the maximum expandable boundary value. In some embodiments, it can be determined according to the coordinates of the nth pixel point that can be expanded in the four directions of the preset cutting area. Therefore, the expansion threshold can include the thresholds in the four directions of the upper, lower, left and right.

[0070] According to an embodiment of the present disclosure, in some embodiments, n may be 10.

[0071] According to the embodiments of the present disclosure, by setting the expansion threshold, it is possible to avoid misidentifying the text to be checked in other cutting blocks when exploring the target cutting area.

[0072] According to an embodiment of the present disclosure, the shape of the candidate sub-region may be a circle, a semicircle, or a closed arc tangent to the edge of the inscribed circle of the starting sub-region. The radius of the circle, semicircle, or closed arc is a preset radius. The edge of the semicircle or closed arc is tangent to the edge of the inscribed circle, but the entire circle may exceed the expansion threshold, so the area formed by the partial arc tangent to the inscribed circle and the straight line where the expansion threshold is located can be used as the candidate sub-region.

[0073] According to an embodiment of the present disclosure, during expansion, the candidate sub-regions of the inscribed circle sub-region in multiple directions may be expanded simultaneously, such as: simultaneously determining the candidate sub-regions located at the top, bottom, left and right.

[0074] According to an embodiment of the present disclosure, when the shape of the candidate sub-region is circular, its region definition may be as shown in the following formula (1).

[0075] ; (1)

[0076] in, Represents a candidate sub-region, whose shape is a point is the center of the circle, A circle with radius Represents the coordinates of the pixel points located in the candidate sub-region in the target voucher image.

[0077] According to an embodiment of the present disclosure, the grayscale values ​​of all pixels in the candidate sub-region in the target voucher image can be determined, and the average grayscale value and the average gradient value of all pixels can be calculated.

[0078] According to an embodiment of the present disclosure, the calculation formulas for the average grayscale value and the average gradient value are shown in the following formulas (2) to (3).

[0079] (2)

[0080] ; (3)

[0081] in, represents the average gray value of the candidate sub-region k in the t-th iteration, Characterizes the number of pixels included in the candidate sub-region k, Characterize the grayscale value and grayscale value of all pixels in the candidate sub-region k. Represents the gray value of the i-th pixel. represents the average gradient value of the candidate sub-region k in the t-th iteration, Characterizes the grayscale mean of candidate sub-region k.

[0082] According to the embodiments of the present disclosure, it is possible to determine whether the candidate sub-region may contain partial text of the text to be checked, through the average grayscale value and the average gradient value of the candidate sub-region. If so, the candidate sub-region is considered to be an extended sub-region. If there is no partial text, the candidate sub-region is considered not to be an extended sub-region, and the candidate sub-region is no longer used as a new starting sub-region.

[0083] According to the embodiments of the present disclosure, when the candidate sub-region is determined to be an extended sub-region through the judgment result, the extended sub-region can be used as a new starting sub-region. It is also judged whether it meets the outward expansion condition. If the outward expansion condition is met, it is considered that the new starting sub-region can be used as the origin for outward expansion. Therefore, the relevant steps for the starting sub-region can be repeatedly performed on the new extended sub-region, such as: determining at least one candidate sub-region tangent to the new starting sub-region, and performing subsequent operations such as average grayscale value and average gradient value calculation, scalability judgment, etc.

[0084] According to an embodiment of the present disclosure, when it is determined that the new starting sub-region does not satisfy the outward expansion condition, outward expansion is not performed with the new starting sub-region as the origin.

[0085] According to an embodiment of the present disclosure, when it is determined that all new starting sub-regions do not meet the outward expansion condition, the expanded sub-regions obtained in all iterative processes can be counted, and the expanded sub-regions and the starting sub-regions can be spliced ​​to obtain the target cutting region.

[0086] According to the embodiments of the present disclosure, in some embodiments, the circumscribed rectangle of each extended sub-region can be determined, that is, the extended sub-region is the inscribed circle sub-region or partially inscribed circle sub-region of the circumscribed rectangle, so that the circumscribed rectangle of each extended sub-region and the starting sub-region are spliced ​​to obtain the target cutting area.

[0087] According to an embodiment of the present disclosure, the outward expansion condition characterizes whether the new starting sub-region can continue to expand outward, and the outward expansion condition may include at least one of the following: the absolute value of the difference between the average grayscale value of the new starting sub-region and the average grayscale value of the starting sub-region is greater than or equal to the target value, and the new starting sub-region has a candidate sub-region and its candidate sub-region is not an explored candidate sub-region.

[0088] According to an embodiment of the present disclosure, the absolute value of the difference between the average grayscale value of the new starting sub-region and the average grayscale value of the starting sub-region in the outward expansion condition is greater than or equal to the target value, which can be shown as the following formula (4).

[0089] ; (4)

[0090] in, is the average gray value of the new starting sub-region, is the average gray value of the starting sub-region, is the target value.

[0091] According to the embodiments of the present disclosure, the target value is not limited and can be obtained through experiments.

[0092] According to the embodiments of the present disclosure, by constructing a circumscribed sub-region that is tangent to the inscribed sub-region of the starting sub-region, and determining the candidate sub-region through the circumscribed sub-region, a tight expansion with the starting sub-region as the origin can be achieved, and different candidate sub-regions have different orientations relative to the inscribed sub-region, so that the range can be expanded in multiple directions at the same time. The average grayscale value and average gradient value of the candidate sub-region are used to determine whether there may be text in the candidate sub-region. If there is text, the candidate sub-region is considered to be an extended sub-region, and the extended sub-region is used as the new starting sub-region, thereby achieving continuous iterative determination of the extended sub-region until all new starting sub-regions do not meet the outward expansion conditions.

[0093] According to an embodiment of the present disclosure, based on the average grayscale value and the average gradient value of the candidate sub-region, the scalability of the candidate sub-region is judged to obtain the judgment result, which may include the following operations.

[0094] Subtract the average grayscale value from the result of the first gradient operation to obtain a first grayscale threshold, wherein the first gradient operation result is obtained by multiplying the average gradient value by the first preset parameter value; add the average grayscale value and the second gradient operation result to obtain a second grayscale threshold, wherein the second gradient operation result is obtained by multiplying the average gradient value by the second preset parameter value; based on the grayscale value range formed by the first grayscale threshold and the second grayscale threshold and the average grayscale value of the starting sub-region, perform scalability judgment on the candidate sub-region to obtain a judgment result.

[0095] According to an embodiment of the present disclosure, when the average grayscale value of the initial subregion does not belong to the grayscale value range, it is considered that there is no text in the candidate subregion, that is, the judgment result indicates that the candidate subregion is not an extended subregion.

[0096] According to an embodiment of the present disclosure, when the average grayscale value of the initial subregion belongs to the grayscale value range, it is considered that there may be text in the candidate subregion, that is, the judgment result indicates that the candidate subregion is an extended subregion.

[0097] According to the embodiments of the present disclosure, the specific numerical values ​​of the first preset parameter value and the second preset parameter value are not limited and can be determined based on experiments.

[0098] According to an embodiment of the present disclosure, the scalability judgment is performed on the candidate sub-regions, and the implementation process of obtaining the judgment result can refer to the following formula (5).

[0099] ; (5)

[0100] in, is the average gray value of the starting sub-region in the tth iteration, is the first preset parameter, is the second preset parameter.

[0101] According to an embodiment of the present disclosure, at t=1, for That is, the average gray value of the starting sub-region. When t is greater than or equal to 2, for That is, the average gray value of the new starting sub-region in each iteration.

[0102] According to an embodiment of the present disclosure, by fusing the average grayscale value and the average gradient value of the candidate sub-region to determine the grayscale value range, the similarity between the candidate sub-region and the starting sub-region can be judged, and the grayscale value range can be reasonably controlled by the first preset parameter value and the second preset parameter value, thereby making the judgment more accurate.

[0103] According to an embodiment of the present disclosure, determining a starting sub-region from a preset cutting region based on the grayscale values ​​of each of a plurality of pixel points may include the following operations.

[0104] Divide the preset cutting area into a grid to obtain multiple grid sub-areas; for each grid sub-area, determine the average grayscale value of the grid sub-area based on the grayscale value of at least one pixel point included in the grid sub-area; and take the grid sub-area with the largest average grayscale value among the multiple grid sub-areas as the starting sub-area.

[0105] According to the embodiments of the present disclosure, the grid division method is not limited. The preset cutting area can be divided by a matrix-shaped grid, and each grid has the same size. In some embodiments, a 3×3 matrix-shaped grid can be used to traverse the entire preset cutting area.

[0106] According to an embodiment of the present disclosure, the average grayscale value of each grid sub-region can be determined simultaneously, and the grid sub-region with the largest average grayscale value can be determined as the starting sub-region.

[0107] According to the embodiments of the present disclosure, since the grid sub-region with the largest average grayscale value is most likely to contain text, the grid sub-region is used as the starting sub-region and expanded outward, so as to achieve more accurate recognition of the region where all the texts to be checked are located.

[0108] According to an embodiment of the present disclosure, performing abnormality checks on the texts to be checked that correspond to at least one cut-out block image respectively to obtain a check result of the target voucher image may include the following operations.

[0109] Based on the preset inspection rules corresponding to the at least one preset cutting area, the preset inspection rules corresponding to the at least one cutting block image are determined; for each cutting block image, based on the preset inspection rules corresponding to the cutting block image, the text to be inspected corresponding to the cutting block image is inspected for abnormality to obtain an inspection sub-result; based on the inspection sub-result of the at least one text to be inspected, the inspection result of the target voucher image is obtained.

[0110] According to an embodiment of the present disclosure, the correspondence between the preset inspection rules and the preset cutting areas can be set in advance and placed in a storage space, so that target credential images of different credential types can quickly determine the preset inspection rules for inspecting the text to be inspected included therein.

[0111] According to an embodiment of the present disclosure, the inspection sub-result includes the cause of the abnormality in the text to be inspected and the identifier of the text to be inspected, so that the inspection result of the target credential image can be obtained by splicing or embedding the inspection sub-results of all the texts to be inspected in the target credential image into a preset result template, so that the customer can find the root cause of the error through the inspection result and correct it as soon as possible, thereby improving the overall speed of business processing.

[0112] According to an embodiment of the present disclosure, in some embodiments, for some credential types, preset verification rules may be as shown in the following Table 1.

[0113] Table 1

[0114]

[0115] According to an embodiment of the present disclosure, the text recognition model includes a feature extraction module, a circulation module and a transcription module; the feature extraction module includes a first convolutional layer and a second convolutional layer, the first convolutional layer and the second convolutional layer are connected in series, and the second convolutional layer includes multiple parallel paths, and the paths are composed of convolutional sublayers or pooling layers; at least one cut block image is input into the text recognition model respectively to obtain a text to be checked corresponding to each of the at least one cut block images, which may include the following operations.

[0116] For each cut block image, the feature extraction module performs multi-layer feature extraction on the cut block image to obtain a target feature map; the target feature map is input into the circulation module for feature prediction to obtain predicted label data; the predicted label data is input into the transcription module to obtain the text to be verified corresponding to the cut block image, wherein the transcription module is used to map the preset label data into the initial text and perform normalization processing on the initial text.

[0117] According to an embodiment of the present disclosure, the feature extraction module can realize multi-layer feature extraction of the cut block image. Both the first convolution layer and the second convolution layer can be multiple.

[0118] According to the embodiments of the present disclosure, there is no limitation on the implementation of the cyclic module, and it can be implemented by a recurrent neural network or an improved network thereof, for example, a bidirectional gated recurrent neural network (BiGRU).

[0119] According to the embodiments of the present disclosure, there is no limitation on the implementation of the transcription module, and a Softmax function or the like may be used.

[0120] According to an embodiment of the present disclosure, the normalization process may include: deduplication, semantic correction, formatting, wrong word replacement and other processes.

[0121] Figure 4 The structural diagram of the text recognition model according to the embodiment of the present disclosure is schematically shown.

[0122] like Figure 4 As shown in the figure, the text recognition model includes a feature extraction module, a loop module and a transcription module. The feature extraction module is used to perform multi-layer feature extraction on the cut block image, so as to better extract the image feature information. The loop module can perform label prediction on the target feature map and generate the corresponding predicted label distribution. The transcription module accepts the predicted label distribution and maps it into specific characters through the objective function. The output sequence is normalized by merging, deduplication and other normalization processes to recognize the text.

[0123] According to an embodiment of the present disclosure, the feature extraction module may include multiple first convolutional layers, one second convolutional layer and an average pooling layer. The multiple first convolutional layers respectively include: first convolutional layer 1, first convolutional layer 2 and first convolutional layer 3. The multiple first convolutional layers are in a series relationship, and the first convolutional layer and the second convolutional layer are also in a series relationship.

[0124] According to the embodiments of the present disclosure, in some embodiments, the text to be checked in the cut block image is a path length of five. Through the transcription module, no matter whether the initial text obtained by mapping is path-path-length-length-degree-five-five or path-path-length-degree-degree-five-five, the correct result can be obtained, that is, the path length is five.

[0125] Figure 5 FIG1 shows a 1×1 convolution schematically showing the structure of the second convolution layer of the text recognition model according to an embodiment of the present disclosure.

[0126] like Figure 5 As shown, the second convolution layer can include 4 parallel paths. The first path uses a 1×1 convolution sublayer, the second path and the third path respectively use a combination of a 3×3 convolution sublayer and a 1×1 convolution sublayer, and the fourth path uses a combination of a 3×3 maximum pooling sublayer and a 1×1 convolution sublayer. The 1×1 convolution sublayer can change the number of input channels. By using a separate 1×1 convolution sublayer in each path or combining it with other convolution sublayers, feature information of different spatial regions can be extracted, and the maximum pooling layer can retain the main features and reduce parameter dimensionality.

[0127] According to the embodiments of the present disclosure, the role of the 1×1 convolution sublayer is equivalent to performing a full connection operation on a certain position in the cut block image, adding an n×n×m image, and using an m'×1×1 convolution kernel, you can get a n×n×m' feature map, that is, adding a nonlinear function such as RELU, etc., to achieve the effect of reducing and increasing channel communication, and the pooling sublayer changes the height and width of the feature map. That is, the reasonable use of the 1×1 convolution sublayer in the second convolution layer achieves the effect of reducing the computational cost, and at the same time can handle the filters required by the convolution sublayer and the pooling sublayer by itself, through 1×1, 3×3 convolution kernel convolution, 3×3 pooling to extract the required features, it is necessary to ensure that its dimension remains unchanged, and use the same in max-pooling, s=1. Finally, the feature maps obtained by each convolution sublayer are superimposed or feature fused to obtain the output feature map of the second convolution sublayer. The first branch path channel is 64, the second is 128, the third is 32, and the fourth is 32, which together form a 256 network structure. The network structure of the second convolution layer can reduce the waste of computing resources to a certain extent. The layer size is reduced by using a 1×1 convolution kernel. Under a reasonable design of its structure, it can ensure that the performance of its neural network is not affected and the computing cost can be reduced.

[0128] According to an embodiment of the present disclosure, Figure 4 In the figure, 1×1cov, 3×3cov, and 3×3Max-pooling can represent the 1×1 convolution sublayer, the 3×3 convolution sublayer, and the 3×3 maximum pooling sublayer respectively. Filters represent the number of channels.

[0129] According to an embodiment of the present disclosure, the specific structure of the second convolutional layer is only schematic. In some embodiments, convolution sublayers or pooling sublayers may be added or reduced in the path, or the number of convolution kernels of the convolution sublayer or the number of kernels of the pooling sublayer may be changed.

[0130] According to the embodiment of the present disclosure, the image text features obtained through the feature extraction layer are passed to the circulation layer. The circulation layer obtains the feature vector and generates the corresponding predicted label distribution. The transcription module accepts the predicted label distribution and maps it into specific characters through the objective function. The output sequence is merged, deduplicated and other related operations to recognize the text.

[0131] According to the embodiments of the present disclosure, by setting serial and parallel convolution layers in the feature extraction module, it is possible to achieve a more comprehensive extraction of the cut block image features and improve the recognition accuracy of the handwritten text to be inspected. A transcription module is also set to map and normalize the predicted label data output by the loop module, such as deduplication, semantic correction, formatting, etc., so as to increase the probability of outputting the correct text to be inspected and improve the accuracy of subsequent abnormality detection.

[0132] According to the embodiments of the present disclosure, in text recognition of cut block images, since the length of the predicted label is not fixed, a Chinese character may be a multi-column feature of the predicted sequence. If the text sequence is too long, it is not suitable to use the ordinary text loss optimization function, and it is difficult to keep the length of the real text sequence consistent with the predicted sequence. Therefore, when training the text recognition model, the CTC (Connectionist Temporal Classification Loss Function) loss function can be used, and end-to-end training can be performed, ignoring the intermediate intervention process. It is an overall connection training mode. CTC introduces blank characters, and based on the many-to-one mapping rule, the gap between the real text label and the model predicted label can be obtained, which can solve the problem of processing the difference between the actual label length and the predicted label length.

[0133] According to an embodiment of the present disclosure, the working principle of the CTC loss function is to obtain the final label by passing the label probability through β, and the β formula is shown in the following formula (6):

[0134] ; (6)

[0135] Among them, L is the character set that needs to be recognized, Indicates that the length of the output sequence does not exceed T.

[0136] According to an embodiment of the present disclosure, in some embodiments, the cycle time step T=20, if the character set to be recognized is "path length is five", 3 paths can deduce the true label.

[0137] β(π1)=(path-path-length-length-55)=the path length is five.

[0138] β(π2)=(path-length-five-five)=the path length is five.

[0139] β(π3 )=(---path--length---five)=the path length is five.

[0140] Among them, "-" represents a space character. The CTC loss function will not have a "-" symbol interval, and repeated characters will be merged into the same word. The formula for calculating the CTC loss function is shown in the following formula (7).

[0141] ; (7)

[0142] in, It represents the probability of output l given input x, where l represents the label. After β transformation, the set of all paths Π with label l is generated. Any path Π needs to meet the requirements of the following formula (8).

[0143] ; (8)

[0144] in, Characterizes a given input x output path The probability of Indicates the path The label at time t, Represents the output sequence output at time t probability.

[0145] According to an embodiment of the present disclosure, by adjusting the gradient , is the network parameter in the loop module, which can make under conditions Get the maximum value.

[0146] According to an embodiment of the present disclosure, after the text recognition model is trained, it is tested, and its recognition effect on the text in the image is shown in the following Table 2.

[0147] Table 2

[0148]

[0149] According to an embodiment of the present disclosure, the target credential image is obtained in the following manner.

[0150] In response to receiving an initial credential image, grayscale processing is performed on the initial credential image to obtain a grayscale credential image, wherein the initial credential image is obtained by scanning the image of the business credential to be inspected; target denoising processing is performed on the grayscale credential image to obtain a denoised credential image, wherein the target denoising processing includes Gaussian noise removal processing and salt and pepper noise removal processing; tilt correction is performed on the denoised credential image to obtain a target credential image.

[0151] According to an embodiment of the present disclosure, a scanner or an image scanning device stored in a local machine may be used to perform image scanning on the business voucher to be inspected, thereby obtaining an initial voucher image.

[0152] According to the embodiments of the present disclosure, there is no limitation on the grayscale processing method, and a maximum value method, an average value method, or a weighted average value method may be used.

[0153] According to an embodiment of the present disclosure, the grayscale value can be expressed by the following formula (9).

[0154] ; (9)

[0155] Among them, R(i, j), G(i, j), B(i, j) represent the red component, green component and blue component of pixel (i, j) respectively. Represents the grayscale value of pixel (i, j).

[0156] According to the embodiments of the present disclosure, it is found through testing that the best effect is achieved after graying using the weighted average method. The weighted average method is to perform weighted average on the three components with different weights according to the needs of different businesses, and finally obtain a value, which is used as the grayscale value of the grayscale image. The calculation formula is shown in formula (10).

[0157] ; (10)

[0158] in, Respectively represent the weighting of the three color components R(i, j), G(i, j), and B(i, j). The present disclosure has found through experiments that The grayscale effect is best when , so in some embodiments, the above weight value can be set to the above value.

[0159] According to the embodiments of the present disclosure, the grayscale conversion of the initial voucher image is a process of converting a color initial voucher image into a grayscale image. Grayscale conversion can reduce the image memory usage without affecting the image description information, improve the visual effect and highlight the required target area, which is beneficial to the subsequent voucher inspection, the recognition of the text to be inspected in the voucher and other functions.

[0160] According to the embodiments of the present disclosure, the presence of noise in the grayscale voucher image will greatly affect the recognition results of the subsequent text recognition model. For the grayscale voucher image, the noises with relatively large impact in the image are mainly Gaussian noise and salt and pepper noise, so the above two noises can be removed.

[0161] According to the embodiments of the present disclosure, there is no limitation on the method of removing the above two noises. For example, a Gaussian filter and a median filter may be used to remove Gaussian noise and salt and pepper noise respectively.

[0162] According to the embodiment of the present disclosure, the filtering of the Gaussian filter is a linear smoothing filter. In the application, the Gaussian convolution kernel needs to be calculated first, and the calculation method is shown in formula (11):

[0163] ; (11)

[0164] Among them, x and y represent the coordinates of the center point of the convolution kernel, n represents the size of the convolution kernel, represents the Gaussian function variance, Represents the value of the convolution kernel.

[0165] According to an embodiment of the present disclosure, after the convolution kernel is calculated, it needs to be applied to each pixel in the image. For each pixel in the input image, the convolution kernel is applied to perform weighted average calculation on the pixels around it in a cyclic traversal manner. Each weight in the convolution kernel is multiplied by the grayscale value of its corresponding pixel, and then all the products are added to obtain a weighted average grayscale value. Finally, the value of the central pixel is replaced by this weighted average grayscale value to obtain the output value of the pixel. After processing each pixel, the denoising effect of the image can be achieved.

[0166] According to the embodiments of the present disclosure, the median filter is a nonlinear filtering technology that can effectively remove salt and pepper noise or other non-Gaussian noise in an image. In the denoising process of the median filter, it is first necessary to determine the size of the convolution kernel, that is, the size of the current field. Generally, the size of the convolution kernel is also set to an odd number, and then the other pixels in the convolution kernel are sorted from small to large according to their pixel values, so that the sorted middle value can be obtained and used as the output value of the pixel. Finally, the output value is assigned to the pixel corresponding to the original image to obtain a denoised image.

[0167] According to the embodiment of the present disclosure, assuming that the current pixel point is (x, y), the area formed with it as the center is , the pixel value of the pixel point in the field is represented by g(s, t), then the pixel value calculation formula obtained by calculating with the median filter is shown in the following formula (12).

[0168] ; (12)

[0169] According to the embodiments of the present disclosure, there is no limitation on the tilt correction method, and methods such as a projection method, a Hough transform method, and a nearest neighbor clustering method may be used.

[0170] According to the embodiments of the present disclosure, after testing, the Hough transform method has a relatively better effect on the tilt correction of the denoised voucher image. Therefore, in some embodiments, the Hough transform method can be used for tilt correction.

[0171] According to an embodiment of the present disclosure, specifically, the Hough transform method can adopt the following steps: construct a parameter space, map each pixel point in the grayscale voucher image from a Cartesian coordinate system to a polar coordinate system ; Build an accumulator array to record each The number of times the coordinate point is passed by the straight line; traverse all pixel points, for all The coordinate points are counted using the constructed accumulator array; in the accumulator array, the largest value, i.e. the statistical peak, is found, and the coordinates of the point are recorded. The straight line where this coordinate point is located is the straight line that needs to be detected, and the angle value is the tilt angle. The grayscale voucher image is then rotated according to this value to achieve the correction effect.

[0172] According to an embodiment of the present disclosure, tilt correction can correct the direction of the grayscale voucher image so that the image remains aligned, thereby facilitating determination of a preset cutting area during subsequent image segmentation, thereby obtaining a target cutting area.

[0173] Figure 6 The flowchart of determining an extended sub-area according to an embodiment of the present disclosure is schematically shown.

[0174] like Figure 6 As shown, determining the extended sub-area may include operations S601 to S607.

[0175] In operation S601, an inscribed circle sub-region of a starting sub-region is determined.

[0176] In operation S602 , based on a preset radius and an expansion threshold, a candidate sub-region tangent to the edge of the inscribed circle sub-region is determined.

[0177] According to the embodiments of the present disclosure, in some embodiments, the shape of the candidate sub-region may be a circle, a closed arc, etc. In other embodiments, the shape of the candidate sub-region may also be a rectangle, etc., which is not limited thereto. When the candidate sub-region is a rectangle, the preset radius may be half the width of the rectangle.

[0178] According to an embodiment of the present disclosure, when there are multiple candidate sub-regions, operations S603 to S607 may be performed on each candidate sub-region, and all the finally determined extended sub-regions may be summarized to obtain M extended sub-regions.

[0179] In operation S603, it is determined whether the candidate sub-region is an extended sub-region. If it is determined that the candidate sub-region is an extended sub-region, operation S604 is performed. If it is determined that the candidate sub-region is not an extended sub-region, operation S607 is performed.

[0180] According to an embodiment of the present disclosure, based on the grayscale value of at least one pixel included in the candidate subregion, the average grayscale value and the average gradient value of the candidate subregion are determined. Based on the average grayscale value and the average gradient value of the candidate subregion, the candidate subregion is subjected to scalability judgment to obtain a judgment result, and whether the candidate subregion is an extended subregion can be determined by the judgment result.

[0181] In operation S604, the expanded sub-region is used as a new starting sub-region.

[0182] In operation S605, it is determined whether the new starting sub-region meets the outward expansion condition. If it is determined that the new starting sub-region meets the outward expansion condition, operations S601 to S604 are repeatedly performed on the new starting sub-region. If it is determined that the new starting sub-region does not meet the outward expansion condition, operation S606 is performed.

[0183] According to an embodiment of the present disclosure, when the new starting sub-region is circular, it is no longer necessary to repeatedly determine the inscribed circle sub-region of the starting sub-region. The new starting sub-region can be considered as the inscribed circle sub-region, and the repeatedly executed operations are S602 to S604.

[0184] In operation S606, the expansion with the new starting sub-region as the origin is stopped.

[0185] In operation S607 , the candidate sub-region is marked, and no expansion is performed with the candidate sub-region as the origin.

[0186] Figure 7 A flowchart of a business credential verification method according to another embodiment of the present disclosure is schematically shown.

[0187] like Figure 7 As shown, the method includes operations S701 to S709.

[0188] In operation S701, in response to receiving a request to scan the business credential to be verified, an image scanning device is controlled to scan the business credential to be verified to obtain an initial credential image.

[0189] In operation S702, target denoising is performed on the grayscale voucher image to obtain a denoised voucher image.

[0190] In operation S703, tilt correction is performed on the denoised voucher image to obtain a target voucher image.

[0191] In operation S704 , at least one target cutting area is determined based on the voucher type of the target voucher image.

[0192] In operation S705 , the target voucher image is cut based on at least one target cutting area to obtain at least one cut block image.

[0193] In operation S706, at least one cut-out block image is respectively input into a text recognition model to obtain a text to be checked corresponding to each of the at least one cut-out block images.

[0194] In operation S707, it is determined whether the target text to be checked matches the rule. The target text to be checked is randomly selected from at least one text to be checked. If it is determined that the target text to be checked matches the rule, operation S708 is performed. If it is determined that the target text to be checked matches the rule, operation S709 is performed.

[0195] In operation S708, a new target text to be checked is selected from the remaining texts to be checked, and operations S707 to 709 are performed on the new target text to be checked.

[0196] In operation S709 , it is determined that the inspection result of the target credential image is abnormal.

[0197] According to an embodiment of the present disclosure, when it is recognized that any target text to be inspected does not match the rule, the inspection result of the target voucher image is output as an abnormal result.

[0198] Based on the above business credential verification method, the present disclosure also provides a business credential verification device. Figure 8 The device is described in detail.

[0199] Figure 8 The structure block diagram of the business credential verification device according to the embodiment of the present disclosure is schematically shown.

[0200] like Figure 8 As shown, the business voucher verification device 800 of this embodiment includes a region determination module 810 , a cutting module 820 , a text recognition module 830 and a verification module 840 .

[0201] The area determination module 810 is used to determine at least one target cutting area based on the credential type of the target credential image, wherein the credential type is determined based on the attribute information of the target credential image, and the target cutting area is the area where the text to be verified in the target credential image is located in the target credential image.

[0202] The cutting module 820 is used to cut the target voucher image based on at least one target cutting area to obtain at least one cutting block image.

[0203] The text recognition module 830 is used to input at least one cut-out block image into a text recognition model to obtain a text to be checked corresponding to each of the at least one cut-out block images.

[0204] The inspection module 840 is used to perform abnormality inspection on the text to be inspected corresponding to each of the at least one cut block image, and obtain the inspection result of the target voucher image.

[0205] According to an embodiment of the present disclosure, the region determination module 810 includes: a first region determination submodule, a second region determination submodule, an expansion submodule and a third region determination submodule.

[0206] The first area determination submodule is used to determine at least one preset cutting area corresponding to the credential type based on the credential type of the target credential image, wherein at least part of the text to be checked is located in the preset cutting area, and the preset cutting area includes multiple pixel points.

[0207] The second region determination submodule is used to determine, for each preset cutting region, a starting subregion from the preset cutting region based on the grayscale values ​​of the plurality of pixels.

[0208] The expansion submodule is used to expand the area with the starting sub-area as the origin to obtain M expanded sub-areas, where M is an integer greater than or equal to 0.

[0209] The third region determination submodule is used to obtain a target cutting region based on the M extended subregions and the initial subregion.

[0210] According to an embodiment of the present disclosure, the expansion submodule includes: a first area determination unit, a second area determination unit, a calculation unit, a judgment unit, a third area determination unit and an iterative execution unit.

[0211] The first region determining unit is used to determine an inscribed circle sub-region of a starting sub-region.

[0212] The second area determination unit is used to determine at least one candidate sub-area tangent to the edge of the inscribed circle sub-area based on a preset radius and an expansion threshold, wherein different candidate sub-areas have different orientations relative to the inscribed circle sub-area, and the coordinates of the pixel points located at the edge of the candidate sub-area are less than or equal to the expansion threshold.

[0213] The calculation unit is used to determine, for each candidate sub-region, an average grayscale value and an average gradient value of the candidate sub-region based on the grayscale value of at least one pixel point included in the candidate sub-region.

[0214] The judgment unit is used to judge the scalability of the candidate sub-region based on the average grayscale value and the average gradient value of the candidate sub-region to obtain a judgment result.

[0215] The third region determining unit is configured to use the extended subregion as a new starting subregion when the determination result indicates that the candidate subregion is an extended subregion.

[0216] The iterative execution unit is used to repeatedly execute the above operation when the new starting sub-region meets the outward expansion condition.

[0217] According to an embodiment of the present disclosure, a judgment unit includes: a first operator unit, a second operator unit and a judgment unit.

[0218] The first operator unit is used to perform a subtraction operation on the average gray value and the first gradient operation result to obtain a first gray threshold, wherein the first gradient operation result is obtained by multiplying the average gradient value by a first preset parameter value.

[0219] The second operator unit is used to perform an addition operation on the average gray value and the second gradient operation result to obtain a second gray threshold, wherein the second gradient operation result is obtained by multiplying the average gradient value by the second preset parameter value.

[0220] The judgment subunit is used to perform scalability judgment on the candidate subregion based on the grayscale value range formed by the first grayscale threshold and the second grayscale threshold and the average grayscale value of the starting subregion to obtain a judgment result.

[0221] According to an embodiment of the present disclosure, the second region determination submodule includes: a division unit, a mean value determination unit and a sub-region determination unit.

[0222] The division unit is used to divide the preset cutting area into grids to obtain multiple grid sub-areas.

[0223] The mean value determination unit is used to determine the average gray value of each grid sub-region based on the gray value of at least one pixel point included in the grid sub-region.

[0224] The sub-region determining unit is used to take the grid sub-region with the largest average grayscale value among the multiple grid sub-regions as the starting sub-region.

[0225] According to an embodiment of the present disclosure, the verification module 840 includes: a rule determination submodule, a verification submodule, and a result determination submodule.

[0226] The rule determination submodule is used to determine a preset inspection rule corresponding to each of the at least one cutting block images based on a preset inspection rule corresponding to each of the at least one preset cutting areas.

[0227] The inspection submodule is used for performing an abnormality inspection on the to-be-inspected text corresponding to each cut-block image based on a preset inspection rule corresponding to the cut-block image to obtain an inspection subresult.

[0228] The result determination submodule is used to obtain the inspection result of the target voucher image based on the inspection sub-result of at least one text to be inspected.

[0229] According to an embodiment of the present disclosure, the text recognition model includes a feature extraction module, a loop module and a transcription module. The feature extraction module includes a first convolutional layer and a second convolutional layer, the first convolutional layer and the second convolutional layer are connected in series, and the second convolutional layer includes multiple parallel paths, and the paths are composed of convolutional sublayers or pooling sublayers. The text recognition module 830 includes: a feature extraction submodule, a prediction submodule and a transcription submodule.

[0230] The feature extraction submodule is used for performing multi-layer feature extraction on each cut block image to obtain a target feature map.

[0231] The prediction submodule is used to input the target feature map into the loop module for feature prediction to obtain predicted label data.

[0232] The transcription submodule is used to input the predicted label data into the transcription module to obtain the text to be checked corresponding to the cut block image, wherein the transcription module is used to map the preset label data into the initial text and normalize the initial text.

[0233] According to an embodiment of the present disclosure, the business credential verification device 800 further includes a grayscale module, a denoising module and a correction module.

[0234] The grayscale module is used for grayscale processing the initial credential image in response to receiving the initial credential image to obtain a grayscale credential image, wherein the initial credential image is obtained by image scanning the business credential to be inspected.

[0235] The denoising module is used to perform target denoising processing on the grayscale voucher image to obtain a denoised voucher image, wherein the target denoising processing includes Gaussian noise removal processing and salt and pepper noise removal processing.

[0236] The correction module is used to perform tilt correction on the denoised voucher image to obtain a target voucher image.

[0237] According to an embodiment of the present disclosure, any multiple modules of the region determination module 810, the cutting module 820, the text recognition module 830 and the inspection module 840 can be combined in one module for implementation, or any one of the modules can be split into multiple modules. Alternatively, at least part of the functions of one or more of these modules can be combined with at least part of the functions of other modules and implemented in one module. According to an embodiment of the present disclosure, at least one of the region determination module 810, the cutting module 820, the text recognition module 830 and the inspection module 840 can be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on a chip, a system on a substrate, a system on a package, an application specific integrated circuit (ASIC), or can be implemented by hardware or firmware such as any other reasonable way of integrating or packaging the circuit, or implemented in any one of the three implementation methods of software, hardware and firmware or in an appropriate combination of any of them. Alternatively, at least one of the region determination module 810, the cutting module 820, the text recognition module 830, and the inspection module 840 may be at least partially implemented as a computer program module, and when the computer program module is executed, a corresponding function may be performed.

[0238] Fig. 9 A block diagram of an electronic device suitable for implementing a business credential verification method according to an embodiment of the present disclosure is schematically shown.

[0239] like Fig. 9 As shown, the electronic device 900 according to an embodiment of the present disclosure includes a processor 901, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 902 or a program loaded from a storage part 908 to a random access memory (RAM) 903. The processor 901 may include, for example, a general-purpose microprocessor (e.g., a CPU), an instruction set processor and / or a related chipset and / or a dedicated microprocessor (e.g., an application-specific integrated circuit (ASIC)), etc. The processor 901 may also include an onboard memory for caching purposes. The processor 901 may include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of the present disclosure.

[0240] In RAM 903, various programs and data required for the operation of electronic device 900 are stored. Processor 901, ROM 902 and RAM 903 are connected to each other via bus 904. Processor 901 performs various operations of the method flow according to the embodiment of the present disclosure by executing the program in ROM 902 and / or RAM 903. It should be noted that the program can also be stored in one or more memories other than ROM 902 and RAM 903. Processor 901 can also perform various operations of the method flow according to the embodiment of the present disclosure by executing the program stored in the one or more memories.

[0241] According to an embodiment of the present disclosure, the electronic device 900 may further include an input / output (I / O) interface 905, which is also connected to the bus 904. The electronic device 900 may further include one or more of the following components connected to the input / output (I / O) interface 905: an input portion 906 including a keyboard, a mouse, etc.; an output portion 907 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage portion 908 including a hard disk, etc.; and a communication portion 909 including a network interface card such as a LAN card, a modem, etc. The communication portion 909 performs communication processing via a network such as the Internet. A drive 910 is also connected to the input / output (I / O) interface 905 as needed. A removable medium 911, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 910 as needed, so that a computer program read therefrom is installed into the storage portion 908 as needed.

[0242] The present disclosure also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments; or may exist independently without being assembled into the device / apparatus / system. The above computer-readable storage medium carries one or more programs, and when the above one or more programs are executed, the method according to the embodiment of the present disclosure is implemented.

[0243] According to an embodiment of the present disclosure, a computer-readable storage medium may be a non-volatile computer-readable storage medium, for example, may include but is not limited to: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In the present disclosure, a computer-readable storage medium may be any tangible medium containing or storing a program, which may be used by or in combination with an instruction execution system, an apparatus or a device. For example, according to an embodiment of the present disclosure, a computer-readable storage medium may include the ROM 902 and / or RAM 903 described above and / or one or more memories other than ROM 902 and RAM 903.

[0244] The embodiment of the present disclosure also includes a computer program product, which includes a computer program, and the computer program contains program code for executing the method shown in the flowchart. When the computer program product is run in a computer system, the program code is used to enable the computer system to implement the business credential verification method provided by the embodiment of the present disclosure.

[0245] The above functions defined in the system / device of the embodiment of the present disclosure are performed when the computer program is executed by the processor 901. According to the embodiment of the present disclosure, the system, device, module, unit, etc. described above can be implemented by a computer program module.

[0246] In one embodiment, the computer program may be based on a tangible storage medium such as an optical storage device, a magnetic storage device, etc. In another embodiment, the computer program may also be transmitted and distributed in the form of a signal on a network medium, and downloaded and installed through the communication part 909, and / or installed from a removable medium 911. The program code contained in the computer program may be transmitted using any appropriate network medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.

[0247] In such an embodiment, the computer program can be downloaded and installed from the network through the communication part 909, and / or installed from the removable medium 911. When the computer program is executed by the processor 901, the above functions defined in the system of the embodiment of the present disclosure are performed. According to the embodiment of the present disclosure, the system, device, means, module, unit, etc. described above can be implemented by a computer program module.

[0248] According to an embodiment of the present disclosure, the program code for executing the computer program provided by the embodiment of the present disclosure can be written in any combination of one or more programming languages. Specifically, these computing programs can be implemented using high-level process and / or object-oriented programming languages, and / or assembly / machine languages. Programming languages ​​include, but are not limited to, Java, C++, python, "C" language or similar programming languages. The program code can be executed entirely on the user computing device, partially on the user device, partially on the remote computing device, or entirely on the remote computing device or server. In the case of a remote computing device, the remote computing device can be connected to the user computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computing device (for example, using an Internet service provider to connect through the Internet).

[0249] The flow charts and block diagrams in the accompanying drawings illustrate the possible architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flow chart or block diagram can represent a module, a program segment, or a part of a code, and the above-mentioned module, program segment, or a part of a code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order from the order marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram or flow chart, and the combination of the boxes in the block diagram or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0250] It will be appreciated by those skilled in the art that the features described in the various embodiments of the present disclosure may be combined and / or combined in a variety of ways, even if such combinations or combinations are not explicitly described in the present disclosure. In particular, without departing from the spirit and teachings of the present disclosure, the features described in the various embodiments of the present disclosure may be combined and / or combined in a variety of ways. All of these combinations and / or combinations fall within the scope of the present disclosure.

[0251] The embodiments of the present disclosure are described above. However, these embodiments are only for illustrative purposes and are not intended to limit the scope of the present disclosure. Although the embodiments are described above, this does not mean that the measures in the various embodiments cannot be used in combination to advantage. Without departing from the scope of the present disclosure, those skilled in the art may make a variety of substitutions and modifications, which should all fall within the scope of the present disclosure.

Claims

1. A business voucher verification method, characterized in that: The method comprises: Determine at least one target cutting area based on the credential type of the target credential image, wherein the credential type is determined based on the attribute information of the target credential image, and the target cutting area is the area in the target credential image where the text to be verified in the target credential image is located; Based on at least one target cutting area, cutting the target voucher image to obtain at least one cutting block image; Inputting at least one of the cut block images into a text recognition model to obtain a text to be checked corresponding to each of the at least one of the cut block images; Anomaly checks are performed on the texts to be checked corresponding to at least one of the cut-out block images to obtain a check result of the target voucher image.

2. The method according to claim 1, characterized in that The step of determining at least one target cutting area based on the credential type of the target credential image includes: Based on the credential type of the target credential image, determining at least one preset cutting area corresponding to the credential type, wherein at least part of the text to be verified is located in the preset cutting area, and the preset cutting area includes a plurality of pixel points; For each of the preset cutting areas, Based on the grayscale values ​​of the plurality of pixels, determining a starting sub-region from the preset cutting region; Expand the region with the starting subregion as the origin to obtain M expanded subregions, where M is an integer greater than or equal to 0; The target cutting region is obtained based on the M extended sub-regions and the initial sub-region.

3. The method according to claim 2, characterized in that The region expansion is performed with the starting sub-region as the origin to obtain M expanded sub-regions, including: Determine the inscribed circle sub-region of the starting sub-region; Based on a preset radius and an expansion threshold, determining at least one candidate subregion tangent to the edge of the inscribed circle subregion, wherein different candidate subregions have different orientations relative to the inscribed circle subregion, and coordinates of pixels located at the edge of the candidate subregion are less than or equal to the expansion threshold; For each of the candidate sub-regions, based on the grayscale value of at least one pixel included in the candidate sub-region, determining an average grayscale value and an average gradient value of the candidate sub-region; Based on the average grayscale value and the average gradient value of the candidate sub-region, the scalability of the candidate sub-region is judged to obtain a judgment result; When the judgment result indicates that the candidate sub-region is the extended sub-region, taking the extended sub-region as a new starting sub-region; When the new starting sub-region meets the outward expansion condition, the above operation is repeated.

4. The method according to claim 3, characterized in that The step of performing scalability judgment on the candidate sub-region based on the average grayscale value and the average gradient value of the candidate sub-region to obtain a judgment result includes: Subtracting the average grayscale value from a first gradient operation result to obtain a first grayscale threshold, wherein the first gradient operation result is obtained by multiplying the average gradient value by a first preset parameter value; Performing an addition operation on the average grayscale value and a second gradient operation result to obtain a second grayscale threshold, wherein the second gradient operation result is obtained by multiplying the average gradient value by a second preset parameter value; Based on the grayscale value range formed by the first grayscale threshold and the second grayscale threshold and the average grayscale value of the starting sub-region, the scalability of the candidate sub-region is judged to obtain the judgment result.

5. The method according to claim 2, characterized in that: The determining a starting sub-region from the preset cutting region based on the respective grayscale values ​​of the plurality of pixel points comprises: Divide the preset cutting area into grids to obtain a plurality of grid sub-areas; For each of the grid sub-regions, based on the grayscale value of at least one pixel point included in the grid sub-region, determining an average grayscale value of the grid sub-region; The grid sub-region with the largest average grayscale value among the multiple grid sub-regions is used as the starting sub-region.

6. The method according to claim 2, characterized in that The step of performing abnormality inspection on the text to be inspected corresponding to at least one of the cut block images to obtain the inspection result of the target voucher image includes: Determining a preset inspection rule corresponding to each of at least one of the cutting block images based on a preset inspection rule corresponding to each of the at least one preset cutting area; For each of the cut block images, based on a preset inspection rule corresponding to the cut block image, an abnormality inspection is performed on the to-be-inspected text corresponding to the cut block image to obtain an inspection sub-result; Based on the verification sub-result of at least one of the texts to be verified, a verification result of the target voucher image is obtained.

7. The method according to claim 1, characterized in that The text recognition model includes a feature extraction module, a circulation module and a transcription module; the feature extraction module includes a first convolutional layer and a second convolutional layer, the first convolutional layer and the second convolutional layer are connected in series, and the second convolutional layer includes a plurality of parallel paths, and the paths are composed of convolutional sublayers or pooling sublayers; The step of inputting at least one of the cut block images into a text recognition model to obtain a text to be checked corresponding to each of the at least one of the cut block images comprises: For each of the cut block images, The feature extraction module performs multi-layer feature extraction on the cut block image to obtain a target feature map; Inputting the target feature map into the loop module for feature prediction to obtain predicted label data; The predicted label data is input into the transcription module to obtain the text to be checked corresponding to the cut block image, wherein the transcription module is used to map the preset label data into the initial text and perform normalization processing on the initial text.

8. The method according to claim 1, characterized in that The target credential image is obtained by: In response to receiving an initial credential image, grayscale processing is performed on the initial credential image to obtain a grayscale credential image, wherein the initial credential image is obtained by image scanning of the business credential to be inspected; Performing target denoising processing on the grayscale voucher image to obtain a denoised voucher image, wherein the target denoising processing includes Gaussian noise removal processing and salt and pepper noise removal processing; The denoised voucher image is tilt-corrected to obtain the target voucher image.

9. A business voucher verification device, characterized in that: The device comprises: an area determination module, configured to determine at least one target cutting area based on a credential type of a target credential image, wherein the credential type is determined based on attribute information of the target credential image, and the target cutting area is an area in the target credential image where the text to be verified in the target credential image is located; A cutting module, configured to cut the target voucher image based on at least one target cutting area to obtain at least one cutting block image; A text recognition module, used for inputting at least one of the cut block images into a text recognition model to obtain a text to be checked corresponding to each of the at least one of the cut block images; The inspection module is used to perform abnormality inspection on the text to be inspected corresponding to at least one of the cut block images to obtain the inspection result of the target voucher image.

10. An electronic device comprising: one or more processors; a memory for storing one or more computer programs, It is characterized in that the one or more processors execute the one or more computer programs to implement the steps of the method according to any one of claims 1 to 8.

11. A computer-readable storage medium having a computer program or instruction stored thereon, characterized in that: When the computer program or instruction is executed by a processor, the steps of the method according to any one of claims 1 to 8 are implemented.

12. A computer program product comprising a computer program or instructions, characterized in that When the computer program or instruction is executed by a processor, the steps of the method according to any one of claims 1 to 8 are implemented.