Identification model generation method and device, service calling method, equipment, medium
By generating a combination of identification models associated with the target bill type in the bill recognition system, the problem that the prior art cannot identify new types of bills online is solved, and the efficiency of bill automation processing and continuous optimization of the model is achieved.
Patent Information
- Application Number
- CN202310246424.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-10
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2043-03-10
AI Technical Summary
Existing bill recognition technology cannot identify new types of bills online, resulting in production disruptions and delays in document automation.
By determining whether the model library contains the target recognition model combination associated with the ticket image to be identified, if not, the basic recognition model combination is obtained, the sample image and label information are received, the basic recognition model combination is tuned based on this information, and the target recognition model combination is generated associated with the target ticket type.
The online identification of new types of tickets is realized, which avoids production interruptions, improves the efficiency of automated credential processing, and continuously optimizes the recognition rate and robustness of the model.
Smart Images

Figure CN116453250B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of image recognition technology, and in particular to a recognition model generation method and device, a service calling method, equipment, medium and program product. Background Art
[0002] In the business processing of financial institutions, various types of bills are involved, such as various transfer debit vouchers, transfer credit vouchers, credit vouchers, collection vouchers, non-interbank deposit transaction confirmations, real-time payment vouchers, etc., and with the development of the business, more types of bills will continue to be added.
[0003] In the process of realizing the concept of the present disclosure, the inventors found that there are at least the following problems in the related technologies: the existing bill recognition technologies are mostly based on pre-training models by annotating a large number of samples for different bill types, and the models are put into production after they are tested and qualified. In this way, if a new bill type is encountered during the online recognition process, the recognition cannot be completed online, and a new model needs to be retrained offline for the new bill type. This will cause production interruptions and delay the process of automatic voucher processing. Summary of the invention
[0004] In view of the above problems, the present disclosure provides a recognition model generation method and apparatus, a service calling method, device, medium and program product.
[0005] In one aspect of the present disclosure, a method for generating a recognition model is provided, comprising:
[0006] In response to a recognition request for a bill image to be recognized, determining whether a plurality of recognition model combinations pre-stored in a model library include a target recognition model combination associated with a target bill type of the bill image to be recognized;
[0007] When it is determined that the target recognition model combination is not included in the multiple recognition model combinations pre-stored in the model library, a basic recognition model combination is obtained from the model library;
[0008] Receiving a plurality of sample images and obtaining annotation information of the plurality of sample images, wherein the sample images and the bill image to be identified belong to the same target bill type;
[0009] The parameters of the basic recognition model combination are tuned based on the multiple sample images and the annotation information of the multiple sample images to generate a target recognition model combination associated with the target bill type.
[0010] According to an embodiment of the present disclosure, the above method further includes:
[0011] When it is determined that the target recognition model combination corresponding to the target bill type is included in the multiple recognition model combinations pre-stored in the model library, the target recognition model combination is obtained from the model library;
[0012] Input the bill image to be identified into the target recognition model combination, and output the recognition result of the bill image to be identified;
[0013] Correct the recognition result to obtain the bill information of the bill image to be recognized;
[0014] Sending the bill information of the bill image to be identified to the business processing system;
[0015] The parameters of the target recognition model combination are tuned based on the bill information of the bill image to be recognized.
[0016] According to an embodiment of the present disclosure, determining whether a plurality of recognition model combinations pre-stored in a model library include a target recognition model combination associated with a target bill type of a bill image to be recognized includes:
[0017] Input the bill image to be identified into the layout recognition model, and output the target layout information corresponding to the bill image to be identified;
[0018] Matching the target layout information with a plurality of standard bill templates pre-stored in the template library to generate a layout matching result, wherein the plurality of standard bill templates correspond one-to-one to a plurality of recognition model combinations pre-stored in the model library;
[0019] According to the layout matching result, it is determined whether the multiple recognition model combinations include a target recognition model combination associated with the target bill type.
[0020] According to an embodiment of the present disclosure, wherein:
[0021] The target layout information includes first position information and first quantity information of multiple target element areas included in the bill image to be identified, and second position information and second quantity information of multiple standard element areas marked in the bill standard template;
[0022] The target layout information is matched with multiple standard bill templates pre-stored in the template library to generate layout matching results including:
[0023] The first position information of the plurality of target element regions is matched with the second position information of the plurality of standard element regions, and the first quantity information of the target element regions is matched with the second quantity information of the plurality of standard element regions to generate a layout matching result.
[0024] According to an embodiment of the present disclosure, wherein:
[0025] The layout matching result includes a first matching result, and the first matching result is used to indicate whether the multiple standard bill templates include a target standard bill template that matches the target layout information;
[0026] According to the layout matching result, determining whether the multiple recognition model combinations include a target recognition model combination associated with the target bill type includes:
[0027] If the matching result indicates that the multiple standard bill templates do not include the target bill standard template that matches the target layout information, determining that the multiple recognition model combinations do not include the target recognition model combination associated with the target bill type; and
[0028] When the matching result indicates that the multiple standard bill templates include a target bill standard template matching the target layout information, it is determined that the multiple recognition model combinations include a target recognition model combination associated with the target bill type.
[0029] According to an embodiment of the present disclosure, wherein:
[0030] The layout matching result includes a second matching result, and the second matching result is used to characterize the matching degree between the target layout information and each bill standard template;
[0031] The basic recognition model combination obtained from the model library includes:
[0032] Determining a preferred bill standard template from a plurality of bill standard templates according to the second matching result;
[0033] A recognition model combination corresponding to the preferred bill standard template is obtained from multiple recognition model combinations as a basic recognition model combination.
[0034] According to an embodiment of the present disclosure, wherein:
[0035] The annotation information of the sample image includes: the region range annotation coordinate values, the element annotation name, and the element annotation value of each of the multiple element regions included in the sample image;
[0036] Obtaining annotation information for multiple sample images includes:
[0037] Reading a pre-annotated image corresponding to a pre-selected sample image from a plurality of sample images from a pre-annotation system, wherein each element region in the pre-annotated image is marked with a range frame line;
[0038] Based on the pre-annotated image, generate the region range annotation coordinate values of the element region included in the sample image;
[0039] Combined with the general recognition model, the feature annotation name and feature annotation value of the feature area contained in the sample image are obtained.
[0040] According to an embodiment of the present disclosure, generating the region range annotation coordinate value of the element region included in the sample image based on the pre-annotated image includes:
[0041] Input the pre-annotated image into the coordinate recognition model, and output the frame line coordinate value corresponding to the range frame line in the pre-annotated image;
[0042] The frame line coordinate values are determined as the area range annotation coordinate values of the feature area included in the sample image.
[0043] According to an embodiment of the present disclosure, in combination with a general recognition model, obtaining the element annotation name and element annotation value of the element area contained in the sample image includes:
[0044] Inputting the sample image into the universal recognition model, and outputting a first character recognition result for the printed characters in the sample image;
[0045] Reading a second character recognition result for the handwritten characters in the sample image from the pre-labeling system;
[0046] The first character recognition result and the second character recognition result are combined as the element annotation name and the element annotation value of the element area included in the sample image.
[0047] According to an embodiment of the present disclosure, wherein:
[0048] The basic recognition model combination includes a basic semantic segmentation model and a basic text recognition model, the target recognition model combination includes a target semantic segmentation model and a target text recognition model, and the annotation information of the sample image includes the region range annotation coordinate values, the element annotation name, and the element annotation value of each of the multiple element regions contained in the sample image;
[0049] Tuning the parameters of the basic recognition model combination based on multiple sample images and the annotation information of multiple sample images includes:
[0050] The sample image is input into the basic semantic segmentation model, and the predicted coordinate values of the respective area ranges of multiple element areas contained in the sample image are output;
[0051] Input the sample image and the predicted coordinate values of the area range of each of the multiple element areas contained in the sample image into the character recognition model, and output the predicted element name and predicted element value of each of the multiple element areas contained in the sample image;
[0052] Calculate the loss value of the network loss function according to the element annotation names and element annotation values of each of the multiple element regions, and the element prediction names and element prediction values of each of the multiple element regions;
[0053] When the loss value of the network loss function is less than a preset threshold, a target semantic segmentation model and a target text recognition model associated with the target bill type are generated.
[0054] Another aspect of the present disclosure provides a service calling method, comprising:
[0055] In response to a recognition request for a bill image to be recognized, calling a discrimination service to execute: determining whether a plurality of recognition model combinations pre-stored in a model library include a target recognition model combination associated with a target bill type of the bill image to be recognized;
[0056] When it is determined that the target recognition model combination corresponding to the target bill type is included in the multiple recognition model combinations pre-stored in the model library, the bill recognition service is called to execute: calling the target recognition model combination from the model library, and using the target recognition model combination to output a recognition result of the bill image to be recognized, the recognition result is used to characterize the target element names and target element values of multiple target element areas contained in the bill image to be recognized;
[0057] When it is determined that the target recognition model combination is not included in the multiple recognition model combinations pre-stored in the model library, the bill annotation service is called to perform an operation of annotating the bill image to be identified, wherein the annotation information obtained by annotating the bill image to be identified includes the target element names and target element values of the multiple target element areas contained in the bill image to be identified; and when it is determined that the target recognition model combination is not included in the multiple recognition model combinations pre-stored in the model library, the model tuning service is called to execute: calling the basic recognition model combination from the model library, and tuning the parameters of the basic recognition model combination, generating a target recognition model combination associated with the target bill type, and adding the target recognition model combination to the model library.
[0058] Another aspect of the present disclosure provides a recognition model generation device, comprising:
[0059] A determination module, configured to determine, in response to a recognition request for a bill image to be recognized, whether a plurality of recognition model combinations pre-stored in a model library includes a target recognition model combination associated with a target bill type of the bill image to be recognized;
[0060] A first acquisition module is used to acquire a basic recognition model combination from the model library when it is determined that the target recognition model combination is not included in the multiple recognition model combinations pre-stored in the model library;
[0061] A second acquisition module is used to receive a plurality of sample images and acquire annotation information of the plurality of sample images, wherein the sample images and the bill image to be identified belong to the same target bill type;
[0062] The tuning module is used to tune the parameters of the basic recognition model combination based on multiple sample images and annotation information of the multiple sample images to generate a target recognition model combination associated with the target bill type.
[0063] A service calling device, comprising:
[0064] The first calling module is used to call the discrimination service to execute, in response to a recognition request for a bill image to be recognized: determining whether a plurality of recognition model combinations pre-stored in the model library include a target recognition model combination associated with a target bill type of the bill image to be recognized;
[0065] The second calling module is used to call the bill recognition service to execute, when it is determined that the target recognition model combination corresponding to the target bill type is included in the multiple recognition model combinations pre-stored in the model library: calling the target recognition model combination from the model library, and using the target recognition model combination to output a recognition result of the bill image to be recognized, wherein the recognition result is used to characterize the target element names and target element values of the multiple target element areas included in the bill image to be recognized;
[0066] The third calling module is used to call the bill annotation service to perform the operation of annotating the bill image to be identified when it is determined that the target recognition model combination is not included in the multiple recognition model combinations pre-stored in the model library, and call the model tuning service to execute: calling the basic recognition model combination from the model library, and tuning the parameters of the basic recognition model combination, generating a target recognition model combination associated with the target bill type, and adding the target recognition model combination to the model library.
[0067] Another aspect of the present disclosure provides an electronic device, comprising: one or more processors; and a memory for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors execute the above-mentioned recognition model generation method or service calling method.
[0068] Another aspect of the present disclosure further provides a computer-readable storage medium having executable instructions stored thereon, which, when executed by a processor, causes the processor to execute the above-mentioned recognition model generation method or service calling method.
[0069] Another aspect of the present disclosure further provides a computer program product, including a computer program, which implements the above-mentioned recognition model generation method or service calling method when executed by a processor.
[0070] According to the embodiment of the present disclosure, the above-mentioned recognition model generation method provides an online model tuning method for bill information recognition. For bill information recognition, first determine whether the model library contains a target recognition model combination associated with the target bill type of the bill image to be recognized. If not, obtain the annotation information of the bill image to be recognized, and temporarily obtain the bill information by annotating the information of this type of bill image, as a bottom-line measure to ensure that information production is not interrupted. At the same time, the basic recognition model combination is optimized online in real time using the annotated image to obtain a new recognition model. In this way, when the bill of the same type is subsequently transmitted, the established new model can be directly called to perform information recognition to obtain bill information. It can be seen that by combining the annotated information collection and online real-time tuning, the annotated data is used as business information data for subsequent system processing, and the annotated data is used as a test set to continuously tune the model effect and continuously optimize the variable parameters in the neural network algorithm. The problem of collecting new bill information is solved, and the information of various paper documents and vouchers can be quickly converted into accurate structured information, further promoting the intelligent and efficient processing of financial institutions' business. This method is compatible with the real-time collection of new and old bills, improves the efficiency of automated processing of voucher collection, and ensures the effectiveness of model training and application. With the continuous learning of large amounts of real-time data in production, the robustness and recognition rate of the model will continue to improve. BRIEF DESCRIPTION OF THE DRAWINGS
[0071] The above contents and other objects, features and advantages of the present disclosure will become more apparent through the following description of the embodiments of the present disclosure with reference to the accompanying drawings, in which:
[0072] Figure 1 A schematic diagram showing an application scenario of a recognition model generation method, apparatus, device, medium, and program product according to an embodiment of the present disclosure;
[0073] Figure 2 A flowchart of a method for generating a recognition model according to an embodiment of the present disclosure is schematically shown;
[0074] Figure 3 The schematic diagram shows a principle diagram of a recognition model generation method according to an embodiment of the present disclosure;
[0075] Figure 4 A schematic diagram of a method for model tuning according to an embodiment of the present disclosure is shown;
[0076] Figure 5 A flowchart of a service calling method according to an embodiment of the present disclosure is schematically shown;
[0077] Figure 6 The structure block diagram of the recognition model generating device according to the embodiment of the present disclosure is schematically shown;
[0078] Figure 7 The structure block diagram of the service calling device according to the embodiment of the present disclosure is schematically shown;
[0079] Figure 8 A block diagram of an electronic device suitable for implementing a recognition model generation method or a service calling method according to an embodiment of the present disclosure is schematically shown. DETAILED DESCRIPTION
[0080] Hereinafter, embodiments of the present disclosure will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are exemplary only and are not intended to limit the scope of the present disclosure. In the following detailed description, for ease of explanation, many specific details are set forth to provide a comprehensive understanding of the embodiments of the present disclosure. However, it is apparent that one or more embodiments may also be implemented without these specific details. In addition, in the following description, descriptions of known structures and technologies are omitted to avoid unnecessary confusion of the concepts of the present disclosure.
[0081] The terms used herein are only for describing specific embodiments and are not intended to limit the present disclosure. The terms "comprise", "include", etc. used herein indicate the existence of the features, steps, operations and / or components, but do not exclude the existence or addition of one or more other features, steps, operations or components.
[0082] All terms (including technical and scientific terms) used herein have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification, and should not be interpreted in an idealized or overly rigid manner.
[0083] When using expressions such as "at least one of A, B, and C, etc.", they should generally be interpreted according to the meaning of the expression commonly understood by those skilled in the art (for example, "a system having at least one of A, B, and C" should include but is not limited to a system having A alone, B alone, C alone, A and B, A and C, B and C, and / or A, B, C, etc.).
[0084] In the embodiments of the present disclosure, the collection, updating, analysis, processing, use, transmission, provision, disclosure, storage, etc. of the data involved (for example, including but not limited to user personal information) are in compliance with the provisions of relevant laws and regulations, are used for legitimate purposes, and do not violate public order and good morals. In particular, necessary measures are taken for user personal information to prevent illegal access to user personal information data and maintain the security of user personal information, network security, and national security.
[0085] In the embodiments of the present disclosure, the user's authorization or consent is obtained before obtaining or collecting the user's personal information.
[0086] A voucher is a written proof that can prove the occurrence of economic business matters, clarify economic responsibilities and has legal effect. In various enterprises, institutions, organizations and departments, the processing and circulation of various paper vouchers are involved in the process of handling various affairs or businesses.
[0087] In recent years, with the rapid development of information technology, all industries have increased their investment and construction of IT systems, and promoted the efficient processing of business information through computer electronic, automated, and intelligent processes. All industries are in the process of digital transformation. Paper vouchers still have a wide range of usage scenarios and effectiveness in terms of law, system, and management, and financial institutions are no exception. For example, when accepting transfer and settlement services, in addition to online transfers and electronic remittances, when processing counter services, they will still accept a large number of paper documents such as savings certificates, checks, bills, payments, and deposit slips. The information on the face of the ticket needs to be scanned and entered, and the information is transmitted to the bank's core system to complete payment and settlement.
[0088] In the business processing of financial institutions, various types of bills are involved, such as various transfer debit vouchers, transfer credit vouchers, credit vouchers, collection vouchers, non-interbank deposit transaction confirmations, real-time payment vouchers, etc., and with the development of the business, more types of bills will continue to be added.
[0089] The relevant technical center can collect bill information by introducing artificial intelligence methods, for example, training a recognition model for each voucher type, and then completing the acquisition of voucher information through model recognition and manual collection.
[0090] For example, after the business is transmitted from the counter scanner to the backend system through the network, the recognition model is first used to recognize the elements on the bill image to obtain the recognition results of the element names and element values on this image of this business, and then the recognition result information is distributed to the operator of the operation center for manual information collection. At the same time, the recognition result is compared with the first collection data. If the comparison is consistent, it is used as the final result; if the comparison is inconsistent, it enters manual secondary collection and is used as the final result.
[0091] To obtain a bill recognition model, manual labeling and deep learning model training are required. First, we need to match the needs, determine the useful element names, element values and corresponding positions on the bill, and then determine the data labeling plan; then we need to obtain a large number of image data of such bills in production and start labeling production operations; the labeled data is handed over to the artificial intelligence team to use the labeled sample data to start the model training, model deployment and model online process, and finally the bank's back-end system completes the application of the model.
[0092] In the process of model training, the relevant technology firstly labels and trains a large number of samples to generate a bill recognition model. The whole process involves: extracting data in production, specifying labeling rules and performing labeling production operations, model training and other links. Overall, it takes a long time for a new bill to be automatically processed.
[0093] In addition, in related technologies, training is done before application. If problems occur during the application process, it is necessary to collect sample data again, perform labeling and model training, and correct algorithm parameters to complete the model tuning process before going online. The process of repairing the model and then applying it is also relatively time-consuming and resource-intensive.
[0094] It can be seen that there are at least the following problems in the relevant technology: the existing bill recognition technology is mostly to train the model by annotating a large number of samples for different bill types in advance, and then put it into production after the model is tested and qualified. In this way, if a new bill type is encountered during the online recognition process, the recognition cannot be completed online, and a new model needs to be retrained offline for the new bill type. This will cause production interruptions and delay the process of voucher automation processing.
[0095] In view of this, an embodiment of the present disclosure provides a recognition model generation method, comprising:
[0096] In response to a recognition request for a bill image to be recognized, determine whether a target recognition model combination associated with a target bill type of the bill image to be recognized is included in a plurality of recognition model combinations pre-stored in a model library; if it is determined that the target recognition model combination is not included in a plurality of recognition model combinations pre-stored in the model library, obtain a basic recognition model combination from the model library; receive a plurality of sample images, and obtain annotation information of the plurality of sample images, wherein the sample images and the bill image to be recognized belong to the same target bill type; and optimize parameters of the basic recognition model combination based on the plurality of sample images and the annotation information of the plurality of sample images to generate a target recognition model combination associated with the target bill type.
[0097] Figure 1 The application scenario diagram of the recognition model generation method, apparatus, device, medium and program product according to the embodiments of the present disclosure is schematically shown.
[0098] like Figure 1 As shown, the application scenario 100 according to this embodiment may include a first terminal device 101, a second terminal device 102, a third terminal device 103, a network 104, and a server 105. The network 104 is used to provide a medium for a communication link between the first terminal device 101, the second terminal device 102, the third terminal device 103, and the server 105. The network 104 may include various connection types, such as wired, wireless communication links, or optical fiber cables, etc.
[0099] The user may use at least one of the first terminal device 101, the second terminal device 102, and the third terminal device 103 to interact with the server 105 through the network 104 to receive or send messages, etc. Various communication client applications may be installed on the first terminal device 101, the second terminal device 102, and the third terminal device 103, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social platform software, etc. (only as examples).
[0100] The first terminal device 101, the second terminal device 102, and the third terminal device 103 may be various electronic devices having display screens and supporting web browsing, including but not limited to smart phones, tablet computers, laptop computers, desktop computers, and the like.
[0101] The server 105 may be a server that provides various services, such as a background management server (only as an example) that provides support for websites browsed by users using the first terminal device 101, the second terminal device 102, and the third terminal device 103. The background management server may analyze and process the received data such as user requests, and feed back the processing results (such as web pages, information, or data obtained or generated according to user requests) to the terminal device.
[0102] In the application scenario of the embodiment of the present disclosure, the user can use at least one of the first terminal device 101, the second terminal device 102, and the third terminal device 103 to interact with the server 105, and send a request to the server for obtaining a bill recognition result. In response to the recognition request for the bill image to be recognized, the server 105 can execute the method of the embodiment of the present disclosure, first determining whether the multiple recognition model combinations pre-stored in the model library include a target recognition model combination associated with the target bill type of the bill image to be recognized; when it is determined that the multiple recognition model combinations pre-stored in the model library include the target recognition model combination corresponding to the target bill type, obtaining the target recognition model combination from the model library; inputting the bill image to be recognized into the target recognition model combination, outputting the recognition result of the bill image to be recognized, and returning it to the user through at least one of the first terminal device 101, the second terminal device 102, and the third terminal device 103. When it is determined that the target recognition model combination is not included in the multiple recognition model combinations pre-stored in the model library, a basic recognition model combination is obtained from the model library; multiple sample images are received, and the annotation information of the multiple sample images is obtained, and the parameters of the basic recognition model combination are tuned based on the multiple sample images and the annotation information of the multiple sample images to generate a target recognition model combination associated with the target bill type. Subsequently, when bills of the same type are input, the bills can be recognized by the newly created target recognition model combination, and the recognition result of the bill image can be obtained, which is returned to the user through at least one of the first terminal device 101, the second terminal device 102, and the third terminal device 103.
[0103] It should be noted that the recognition model generation method provided in the embodiment of the present disclosure can generally be executed by the server 105. Accordingly, the recognition model generation device provided in the embodiment of the present disclosure can generally be set in the server 105. The recognition model generation method provided in the embodiment of the present disclosure can also be executed by a server or server cluster that is different from the server 105 and can communicate with the first terminal device 101, the second terminal device 102, the third terminal device 103 and / or the server 105. Correspondingly, the recognition model generation device provided in the embodiment of the present disclosure can also be set in a server or server cluster that is different from the server 105 and can communicate with the first terminal device 101, the second terminal device 102, the third terminal device 103 and / or the server 105.
[0104] It should be understood that Figure 1 The number of terminal devices, networks and servers in the embodiment is only for illustration. Any number of terminal devices, networks and servers may be provided according to implementation requirements.
[0105] The following will be based on Figure 1 The scene described by Figure 2 to Figure 8The recognition model generation method and device, and service calling method of the disclosed embodiment are described in detail.
[0106] Figure 2 A flowchart of a method for generating a recognition model according to an embodiment of the present disclosure is schematically shown; Figure 3 The schematic diagram of the recognition model generation method according to the embodiment of the present disclosure is shown. Figure 2 , Figure 3 The recognition model generation method of the embodiment of the present disclosure is described.
[0107] like Figure 2 As shown, the recognition model generation method of this embodiment includes operations S201 to S204.
[0108] In operation S201, in response to a recognition request for a bill image to be recognized, determining whether a plurality of recognition model combinations pre-stored in a model library includes a target recognition model combination associated with a target bill type of the bill image to be recognized;
[0109] In operation S202, when it is determined that the target recognition model combination is not included in the plurality of recognition model combinations pre-stored in the model library, a basic recognition model combination is obtained from the model library;
[0110] In operation S203, a plurality of sample images are received, and annotation information of the plurality of sample images is obtained, wherein the sample images and the to-be-recognized bill image belong to the same target bill type;
[0111] In operation S204, the parameters of the basic recognition model combination are optimized based on the multiple sample images and the annotation information of the multiple sample images to generate a target recognition model combination associated with the target bill type.
[0112] According to the embodiments of the present disclosure, the above method can be applied in the business processing of financial institutions. When accepting transfer and settlement services, financial institutions will still accept a large number of paper bills such as savings certificates, checks, bills, payments, and deposit slips when handling counter services, in addition to online transfers and electronic remittances. The information on the bills needs to be scanned and entered, and the information is transmitted to the financial institution's business system to complete payment and settlement. Various types of bills will be involved in the process, such as various transfer debit vouchers, transfer credit vouchers, credit vouchers, entrusted collection vouchers, non-interbank deposit transaction confirmations, real-time payment vouchers, etc., and more types of bills will continue to be added as the business develops.
[0113] like Figure 3 As shown, when accepting business, the teller of a financial institution's branch uploads the relevant paper bills to the financial institution's back-end system through a scanner, either individually or in batches, and initiates a recognition request for the image of the bill to be recognized to the back-end, thus completing the task initiation.
[0114] After the background receives the scanned image of the paper bill and the business description information, in operation S201, the bill type and model are first identified, that is, it is determined whether the multiple recognition model combinations pre-stored in the model library contain a target recognition model combination associated with the target bill type of the bill image to be identified. The multiple recognition model combinations pre-stored in the model library can be pre-trained model combinations for each bill type, and one bill type corresponds to one recognition model combination.
[0115] like Figure 3 As shown, each recognition model combination may include a semantic segmentation model and an ICR text recognition model. The semantic segmentation model is used to identify and extract the regions of each element of the bill image. The ICR text recognition model is used to recognize the text in the region of each element, identify the element name and element value, and obtain the recognition result. For example, a payment voucher includes multiple element areas: element area 1 is the payment bank name element area, element area 2 is the payee name element area, element area 3 is...; each element area includes an element name and an element value, such as the element name of element area 1 is "payment bank name" and the element value is "xx bank".
[0116] like Figure 3 As shown, if the bill is judged to be an original bill, the multiple recognition model combinations pre-stored in the model library include a target recognition model combination corresponding to the target bill type. The target recognition model combination (semantic segmentation model and ICR text recognition model) is directly obtained from the model library, and the bill image to be recognized is input into the target recognition model combination, and the recognition result of the bill image to be recognized is output.
[0117] like Figure 3 As shown, if the bill is judged to be a new bill, the multiple recognition model combinations pre-stored in the model library do not include the target recognition model combination corresponding to the target bill type, and the basic recognition model combination is obtained from the model library through operation S202, and through operation S203, multiple sample images and their annotation information are obtained, and through operation S204, the parameters of the basic recognition model combination are tuned based on the multiple sample images and the annotation information of the multiple sample images to generate a target recognition model combination associated with the target bill type.
[0118] The basic recognition model combination may be any one of a plurality of pre-trained recognition model combinations stored in the model library. The annotation information of the sample image may be obtained by manual or automatic annotation, and the annotation information is obtained by annotating the sample image of the same target bill type as the bill image to be identified. The annotation information of the sample image may, for example, be area range annotation (such as the range coordinates of the annotated element area), element name annotation, and element value annotation for multiple element areas contained in the sample image.
[0119] According to an embodiment of the present disclosure, the multiple sample images used for model tuning can be bill images of new bill types to be identified that are received online in real time. Figure 3 As shown in the figure, the bill information can be temporarily obtained by annotating these sample images as a bottom-line measure to ensure that information production is not interrupted. At the same time, the basic recognition model combination is tuned online in real time using the annotated sample images to obtain a new recognition model. In this way, when the same type of bill is subsequently transmitted, the newly established model can be directly called to perform information recognition to obtain the bill information.
[0120] According to the embodiments of the present disclosure, Figure 3 As shown in the figure, for the original bills, the recognition results obtained by the model are collected as bill information (element names and corresponding element values of paper bills), and for the new bills, the bill information is temporarily obtained by labeling these sample images. In this way, the bill information collected in real time for the original bills and the new bills is sent to the business processing system for business information processing, archiving, settlement and other related purposes.
[0121] According to the embodiments of the present disclosure, after the bill information is collected, the accuracy of the bill information can be manually reviewed and proofread. Specifically, the system will prompt the collection results of the previous steps on the input interface, and the operator can complete the information verification of the content of each element on the voucher image according to the recognition prompt result. If the review is correct, it will be directly passed as the final result; if the review finds that the content is incorrect, the operator will directly perform a second collection, and the content of the second collection will be the final result.
[0122] According to the embodiment of the present disclosure, the above-mentioned recognition model generation method provides an online model tuning method for bill information recognition. For bill information recognition, first determine whether the model library contains a target recognition model combination associated with the target bill type of the bill image to be recognized. If not, obtain the annotation information of the bill image to be recognized, and temporarily obtain the bill information by annotating the information of this type of bill image, as a bottom-line measure to ensure that information production is not interrupted. At the same time, the basic recognition model combination is optimized online in real time using the annotated image to obtain a new recognition model. In this way, when the bill of the same type is subsequently transmitted, the established new model can be directly called to perform information recognition to obtain bill information. It can be seen that by combining the annotated information collection and online real-time tuning, the annotated data is used as business information data for subsequent system processing, and the annotated data is used as a test set to continuously tune the model effect and continuously optimize the variable parameters in the neural network algorithm. The problem of collecting new bill information is solved, and the information of various paper documents and vouchers can be quickly converted into accurate structured information, further promoting the intelligent and efficient processing of financial institutions' business. This method is compatible with the real-time collection of new and old bills, improves the efficiency of automated processing of voucher collection, and ensures the effectiveness of model training and application. With the continuous learning of large amounts of real-time data in production, the robustness and recognition rate of the model will continue to improve.
[0123] According to the embodiments of the present disclosure, the basic recognition model combination can be any one of a plurality of trained recognition model combinations pre-stored in the model library. For the plurality of recognition models pre-stored in the model library, although their respective network parameters are different, they are all used for the bill recognition function, and the only difference is that the applicable bill types are different. Therefore, based on the preliminarily trained model, the tuning process can be performed, and the model can be tuned only through a small amount of sample data and a small number of iterations. Compared with the traditional deep learning process that requires data acquisition, training, deployment and online deployment, the cycle from model training to application is shortened to a great extent, thereby ensuring the timeliness of bill information collection.
[0124] According to the embodiments of the present disclosure, Figure 3 As shown, if the bill is judged to be an original bill, the multiple recognition model combinations pre-stored in the model library include a target recognition model combination corresponding to the target bill type. The target recognition model combination (semantic segmentation model and ICR text recognition model) is directly obtained from the model library, and the bill image to be recognized is input into the target recognition model combination, and the recognition result of the bill image to be recognized is output.
[0125] Further, the output recognition result is corrected to obtain the bill information of the bill image to be recognized, and the bill information of the bill image to be recognized is sent to the business processing system. At the same time, the parameters of the target recognition model combination are tuned based on the bill information of the bill image to be recognized.
[0126] According to the embodiments of the present disclosure, faced with the task of collecting information on various new and old bills, the above-mentioned method of the embodiments of the present disclosure achieves compatible processing of new and old bills through different processing methods for new and old bills, and can continuously adapt to various types of bills online to ensure uninterrupted production and improve business processing efficiency.
[0127] In addition, since the recognition results of the recognition model may not be accurate enough, the recognition results are further corrected. At the same time, the recognition model is reversely tuned using the corrected bill recognition information. In this way, as bills are continuously added to production, the parameters of the original model can be continuously updated and optimized through reverse tuning, so that the recognition accuracy of the model is also continuously improved.
[0128] According to an embodiment of the present disclosure, further, determining whether the multiple recognition model combinations pre-stored in the model library contain a target recognition model combination associated with the target bill type of the bill image to be recognized includes the following operations:
[0129] Operation 11, input the bill image to be identified into the layout recognition model, and output the target layout information corresponding to the bill image to be identified. The layout recognition model is used to perform image segmentation on the bill image and identify the element distribution information in the bill, including the position information and quantity information of each element area. Specifically, the target layout information includes the first position information and first quantity information of multiple target element areas contained in the bill image to be identified. The first position information can be, for example, the center coordinate position of the target element area, and the first quantity information is the total number of multiple target element areas contained in the bill.
[0130] Operation 12, matching the target layout information with a plurality of standard bill templates pre-stored in the template library to generate a layout matching result, wherein the plurality of standard bill templates correspond one-to-one to a plurality of recognition model combinations pre-stored in the model library.
[0131] According to an embodiment of the present disclosure, after the recognition model for any new bill type is trained and optimized through online tuning of the model, the model is stored in the model library. At the same time, a bill standard template for the bill type that the model can recognize is generated and stored in the template library. The bill standard template is marked with the second position information and second quantity information of multiple standard element areas. The second position information can be, for example, the center coordinate position of the standard element area, and the second quantity information is the total number of multiple standard element areas contained in the bill.
[0132] According to an embodiment of the present disclosure, matching the target layout information with a plurality of standard bill templates pre-stored in the template library to generate a layout matching result includes: matching the first position information of the plurality of target element regions with the second position information of the plurality of standard element regions, and matching the first quantity information of the target element regions with the second quantity information of the plurality of standard element regions to generate a layout matching result. For example, it may be determined whether the coordinate difference between the first position information and the second position information is less than a preset error threshold, and at the same time, it is determined whether the first quantity information and the second quantity information are the same. If the above conditions are met at the same time, it is considered that the target layout information matches the standard bill template, otherwise, it is not matched.
[0133] Operation 13: Determine, based on the layout matching result, whether the multiple recognition model combinations include a target recognition model combination associated with the target bill type.
[0134] The layout matching result includes a first matching result, and the first matching result is used to indicate whether the multiple standard bill templates include a target standard bill template that matches the target layout information, that is, whether the target layout information matches the standard bill template.
[0135] When the matching result indicates that multiple bill standard templates do not include the target bill standard template that matches the target layout information, it is determined that the multiple recognition model combinations do not include the target recognition model combination associated with the target bill type; conversely, when the matching result indicates that multiple bill standard templates include the target bill standard template that matches the target layout information, it is determined that the multiple recognition model combinations include the target recognition model combination associated with the target bill type.
[0136] According to an embodiment of the present disclosure, by utilizing a layout recognition model to output target layout information corresponding to the bill image to be identified, and matching the target layout information with multiple bill standard templates pre-stored in a template library, it is possible to accurately determine whether the model library contains a target recognition model combination associated with the target bill type based on the matching results.
[0137] According to the embodiments of the present disclosure, further, the layout matching result includes a second matching result, and the second matching result is used to characterize the degree of matching between the target layout information and each bill standard template. In the process of matching the first position information of multiple target element areas with the second position information of multiple standard element areas, and matching the first quantity information of the target element areas with the second quantity information of multiple standard element areas, the second matching result can be output according to the degree of matching. For example, the smaller the coordinate difference between the first position information and the second position information, and the smaller the difference between the first quantity information and the second quantity information, the higher the matching degree is considered, and vice versa.
[0138] Based on this, the basic recognition model combination obtained from the model library can be: according to the second matching result, determine the preferred bill standard template from multiple bill standard templates; obtain the recognition model combination corresponding to the preferred bill standard template from multiple recognition model combinations as the basic recognition model combination.
[0139] Through the above method, the layout structure applicable to the basic recognition model combination is most similar to the layout structure of the bill image to be identified. Therefore, based on the model, the model tuning can be completed with only a small amount of sample data and a few iterations, which greatly shortens the model tuning cycle and ensures the timeliness of bill information collection.
[0140] According to an embodiment of the present disclosure, the annotation information of the sample image includes: the area range annotation coordinate values of each of the multiple element areas contained in the sample image (such as the coordinates of multiple boundary points of the range frame line), the element annotation name (key), and the element annotation value (value).
[0141] For new types of bills, they are first transferred to the information annotation stage to annotate and collect image information. Manual annotation can be used. Manually mark the specific area of each element, and annotate the element name and element value. For example, the operator observes the location where the bill name is printed on the bill, frames the specific area of the bill name, and marks the text content corresponding to the element name and element value.
[0142] For new types of bills, image information can be annotated and collected in a semi-manual and semi-automatic manner to improve the efficiency of annotation. Specifically, obtaining annotation information of multiple sample images may include the following operations:
[0143] Operation 21, read the pre-annotated image corresponding to the pre-selected sample image from the multiple sample images from the pre-annotation system, and each element area in the pre-annotated image is marked with a range frame line. The pre-selected sample image can be any one or more images from the multiple sample images, and each element area on the ticket surface can be framed by manual annotation.
[0144] Operation 22, based on the pre-annotated image, generates the region range annotation coordinate values of the element region included in the sample image. Specifically, the pre-annotated image is input into the coordinate recognition model, and the frame line coordinate values corresponding to the range frame line in the pre-annotated image are output; the frame line coordinate values are determined as the region range annotation coordinate values of the element region included in the sample image.
[0145] According to the embodiments of the present disclosure, for the same type of bill, the various fields on the page are the same in format and layout. Therefore, for the annotation of element areas, the embodiments of the present disclosure obtain the regional annotation information of a small number of bill samples by manually annotating a small number of bills. Since the same type of bill has the same page layout, the regional annotation information of the small number of bills can be used as the regional annotation information of all bills of this type. Subsequently, the frame line coordinate values corresponding to the range frame lines in the pre-annotated image can be directly assigned to the unannotated bill image, so as to realize the automatic annotation of the area positions of the remaining sample images.
[0146] Furthermore, the annotation information can be further corrected manually. The operator can fine-tune the position of the pre-annotated frame line by dragging and dropping to ensure the correctness of the area annotation. Specifically, the system will prompt the collection results on the input interface. The operator can observe the area range mark of the area segmentation on the image according to the recognition prompt results, and fine-tune the segmented area according to the actual situation of the ticket.
[0147] Operation 23, combining with the general recognition model, obtains the element annotation name and element annotation value of the element area contained in the sample image.
[0148] According to the embodiments of the present disclosure, for new bill types, the element annotation names and element annotation values of the element areas contained in the sample images can be annotated manually or semi-manually and semi-automatically to improve the annotation efficiency.
[0149] By using manual marking, the operator can observe the text in each element area of the ticket and record the text content.
[0150] A semi-manual and semi-automatic approach may be adopted, which may be combined with a general recognition model to obtain the element annotation name and element annotation value of the element area included in the sample image. Specifically: the sample image may be input into the general recognition model, and a first text recognition result for the printed text in the sample image may be output; a second text recognition result (which may be a manual annotation) for the handwritten text in the sample image may be read from a pre-annotation system; and the first text recognition result and the second text recognition result may be combined as the element annotation name and element annotation value of the element area included in the sample image.
[0151] According to the embodiments of the present disclosure, the OCR universal recognition model can be used to recognize the printed text on the ticket, including the voucher title, element name, and element value; and manual assistance can be used to recognize handwritten text, so that the marking efficiency can be improved.
[0152] Figure 4 The schematic diagram shows a method principle diagram for model tuning according to an embodiment of the present disclosure.
[0153] According to the embodiments of the present disclosure, Figure 4 As shown, the basic recognition model combination includes a basic semantic segmentation model and a basic text recognition model. The basic semantic segmentation model and the basic text recognition model can be used to process the applicable type of bill recognition requests online in real time, and can participate in model tuning based on annotation information. The same model can be deployed in two service systems, where the model in one service system processes recognition requests for online real-time image information recognition, and the model in the other service system participates in model tuning.
[0154] Based on this, tuning the parameters of the basic recognition model combination based on multiple sample images and the annotation information of multiple sample images includes the following operations:
[0155] Operation 31, input the sample image into the basic semantic segmentation model, and output the predicted coordinate values of the area range of each of the multiple element areas contained in the sample image. The semantic segmentation model is used to identify and extract the areas of each element on the voucher image. The R-CNN network can be used. Compared with the traditional CNN structure with image classification as the main purpose, R-CNN can handle more complex tasks such as target detection and image segmentation. R-CNN performs semantic segmentation based on the detection results by converting region-based predictions into pixel predictions. R-CNN first uses selective search to extract a large number of target proposals, and then calculates the CNN features of each proposal. Finally, a class-specific linear support vector machine is used to classify each area. Finally, the input voucher image is detected and output as the key element slice to be identified.
[0156] Operation 32, input the sample image and the predicted coordinate values of the area range of each of the multiple element areas contained in the sample image into the text recognition model, and output the predicted element names and predicted element values of each of the multiple element areas contained in the sample image. The text recognition model can adopt models such as Auto-ML, BERT, Fine-turn, etc., which can accurately identify the handwriting on the bill. Specifically, Encode-Decode (encoding layer-decoding layer) can be combined with Auto-ML to construct a neural network based on NAS with adaptive receptive field. The network structure of each layer is a "suitable receptive field" substructure searched by NAS, and the information around the word is used to assist in the recognition of the word. This flexible solution also supports dynamic adjustment of the amount of calculation, which can ensure the flexible deployment of the model in various scenarios such as the cloud and local end. Finally, the output of structured text information is generated through the decoding layer.
[0157] Operation 33, calculating the loss value of the network loss function according to the element annotation names and element annotation values of each of the multiple element regions, and the element prediction names and element prediction values of each of the multiple element regions.
[0158] Operation 34 , when the loss value of the network loss function is less than a preset threshold, a target semantic segmentation model and a target text recognition model associated with the target bill type are generated.
[0159] The model tuning strategy uses the backpropagation (BP) strategy. Backpropagation is a basic method for model training and iterative optimization of various machine learning and deep learning algorithms based on neural networks. It can be used in combination with optimization methods (such as gradient descent). This method calculates the gradient of the loss function for all weights in the network. The calculated gradient is fed back to the optimization method to update the weights to minimize the loss function (backward propagation of error).
[0160] The learning process of the BP algorithm consists of a forward propagation process and a back propagation process. In the forward propagation process, the input information passes through the input layer and the hidden layer, and is processed layer by layer and transmitted to the output layer. If the expected output value is not obtained in the output layer, the sum of the squares of the error between the output and the expected value is taken as the objective function, and the back propagation is turned on. The partial derivatives of the objective function to the weights of each neuron are obtained layer by layer, and the gradient of the objective function to the weight vector is formed as the basis for modifying the weights, so that the output results of the network are constantly close to the expected performance.
[0161] The labeled bill sample images continuously input in production are used as the input of the basic semantic segmentation model and the basic text recognition model. The BP neural network at the bottom of the two model algorithms will continuously iterate the model through incremental training. When the loss value of the network loss function is less than the preset threshold, the target semantic segmentation model and target text recognition model associated with the target bill type are generated.
[0162] Another aspect of the present disclosure provides a service calling method. Figure 5 The flowchart of the service calling method according to the embodiment of the present disclosure is schematically shown.
[0163] like Figure 5 As shown, the service calling method of the embodiment of the present disclosure includes operations S501 to S503.
[0164] In operation S501, in response to a recognition request for a bill image to be recognized, a discrimination service is called to perform: determining whether a plurality of recognition model combinations pre-stored in a model library include a target recognition model combination associated with a target bill type of the bill image to be recognized.
[0165] In operation S502, when it is determined that the target recognition model combination corresponding to the target bill type is included in the multiple recognition model combinations pre-stored in the model library, the bill recognition service is called to execute: the target recognition model combination is called from the model library, and the target recognition model combination is used to output the recognition result of the bill image to be recognized, and the recognition result is used to characterize the target element names and target element values of the multiple target element areas contained in the bill image to be recognized.
[0166] In operation S503, when it is determined that the target recognition model combination is not included in the multiple recognition model combinations pre-stored in the model library, the bill annotation service is called to perform the operation of annotating the bill image to be identified, and the model tuning service is called to perform: calling the basic recognition model combination from the model library, and tuning the parameters of the basic recognition model combination, generating the target recognition model combination associated with the target bill type, and adding the target recognition model combination to the model library. The annotation information obtained by annotating the bill image to be identified includes the target element names and target element values of the multiple target element areas contained in the bill image to be identified.
[0167] According to an embodiment of the present disclosure, the specific implementation of the above operations S501 to S503 can refer to the description in the aforementioned embodiment of the recognition model generation method, which will not be repeated here.
[0168] According to an embodiment of the present disclosure, the basic recognition model combination includes a basic semantic segmentation model and a basic text recognition model. The basic semantic segmentation model and the basic text recognition model can be used to process applicable types of bill recognition requests online in real time on the one hand, and can participate in model tuning based on annotation information on the other hand. The same model can be deployed in two service systems, one of which is deployed in the bill recognition service system for online recognition of bill image element information; the other is deployed in the model tuning service system for participating in online model tuning.
[0169] After calling the discrimination service to determine whether the model library contains the target recognition model combination associated with the target bill type of the bill image to be recognized, a service identification flag is generated. If the model library contains the target recognition model combination, fag=1; if the model library does not contain the target recognition model combination, flag=0.
[0170] When the model library contains the target recognition model combination, the operation of calling the bill recognition service is triggered according to the service identifier flag=1, and the target recognition model combination in the bill recognition service system is used to online recognize the bill image element information.
[0171] When the model library contains the target recognition model combination, the service identifier flag = 0 triggers the call of the bill annotation service and the model tuning service. By obtaining the annotation information of the bill image to be identified and multiple sample images from the bill annotation service system (the sample image and the bill image to be identified belong to the same bill type), the bill information is temporarily obtained by annotating the information of this type of bill image as a bottom-line measure to ensure that information production is not interrupted. At the same time, in the model tuning service system, the basic recognition model combination is tuned online in real time using the annotated images to obtain a new recognition model.
[0172] After the model tuning and updating is completed, the service status can be switched. The service system that originally executed the model tuning is switched to the bill recognition service (because the model update is completed and needs to be put into production in time); conversely, the service system that originally executed the bill recognition service is switched to the model tuning service to perform the tuning operation for the new round of bill recognition models. In this way, the uninterrupted execution of online production and online tuning is guaranteed.
[0173] According to the embodiment of the present disclosure, by calling the discrimination service, it can be determined whether the model library contains a target recognition model combination associated with the target bill type of the bill image to be recognized, and then the corresponding service call operation is performed according to different situations. In this way, when there is a target recognition model combination, the bill recognition service is called to complete the online information recognition in time; when there is no target recognition model combination, the bill annotation service is called to obtain the annotation information of the bill image to be recognized, and the bill information is temporarily obtained by annotating the information of this type of bill image as a bottom-line measure to ensure that the information production is not interrupted. At the same time, the model tuning service is called, and the basic recognition model combination is tuned online in real time using the annotated image to obtain a new recognition model. In this way, after the subsequent bills of the same type are passed in, the newly established model can be directly called to perform information recognition to obtain the bill information. It can be seen that by combining the online model recognition, annotation information collection and online real-time tuning, combined with different processing methods for new and old bills, compatible processing of new and old bills is achieved, and various bill types can be continuously adapted online to ensure uninterrupted production and improve business processing efficiency. The problem of collecting new bill information has been solved, and the information of various paper documents and vouchers can be quickly converted into accurate structured information, further promoting the intelligent and efficient processing of financial institutions' business. This method is compatible with the real-time collection of new and old bills, improves the efficiency of automatic processing of voucher collection, and ensures the effectiveness of model training and application. With the continuous learning of a large amount of real-time data in production, the robustness and recognition rate of the model will continue to improve.
[0174] Based on the above recognition model generation method, the present disclosure also provides a recognition model generation device. Figure 6The device is described in detail.
[0175] Figure 6 The structural block diagram of the recognition model generating device according to an embodiment of the present disclosure is schematically shown.
[0176] like Figure 6 As shown, the recognition model generating device 600 of this embodiment includes a determination module 601 , a first acquisition module 602 , a second acquisition module 603 , and a first tuning module 603 .
[0177] The determination module 601 is used to determine, in response to a recognition request for a bill image to be recognized, whether a plurality of recognition model combinations pre-stored in the model library include a target recognition model combination associated with a target bill type of the bill image to be recognized.
[0178] The first acquisition module 602 is used to acquire a basic recognition model combination from the model library when it is determined that the target recognition model combination is not included in the multiple recognition model combinations pre-stored in the model library.
[0179] The second acquisition module 603 is used to receive a plurality of sample images and acquire annotation information of the plurality of sample images, wherein the sample images and the bill image to be identified belong to the same target bill type.
[0180] The first tuning module 604 is used to tune the parameters of the basic recognition model combination based on the multiple sample images and the annotation information of the multiple sample images to generate a target recognition model combination associated with the target bill type.
[0181] According to the embodiment of the present disclosure, the determination module 601 first determines whether the model library contains a target recognition model combination associated with the target bill type of the bill image to be recognized. If not, the second acquisition module 603 obtains the annotation information of the bill image to be recognized, and temporarily obtains the bill information by annotating the information of this type of bill image, as a bottom-line measure to ensure that information production is not interrupted. At the same time, the first tuning module 604 uses the annotated image to perform online real-time tuning on the basic recognition model combination to obtain a new recognition model. In this way, when the bills of the same type are subsequently transmitted, the established new model can be directly called to perform information recognition to obtain bill information. It can be seen that by combining the annotated information collection and online real-time tuning, the annotated data is used as business information data for subsequent system processing, and the annotated data is used as a test set to continuously optimize the model effect and continuously optimize the variable parameters in the neural network algorithm. The problem of collecting new bill information is solved, and the information of various paper documents and vouchers can be quickly converted into accurate structured information, further promoting the intelligent and efficient processing of financial institutions' business. This method is compatible with the real-time collection of new and old bills, improves the efficiency of automated processing of voucher collection, and ensures the effectiveness of model training and application. With the continuous learning of large amounts of real-time data in production, the robustness and recognition rate of the model will continue to improve.
[0182] According to an embodiment of the present disclosure, the above-mentioned device also includes a third acquisition module, an identification module, a correction module, a sending module, and a second tuning module.
[0183] The third acquisition module is used to acquire the target recognition model combination from the model library when it is determined that the target recognition model combination corresponding to the target bill type is included in the multiple recognition model combinations pre-stored in the model library.
[0184] The recognition module is used to input the bill image to be recognized into the target recognition model combination and output the recognition result of the bill image to be recognized.
[0185] A correction module, used to correct the recognition result to obtain the bill information of the bill image to be recognized;
[0186] A sending module, used for sending the bill information of the bill image to be identified to the business processing system;
[0187] The second tuning module is used to tune the parameters of the target recognition model combination based on the bill information of the bill image to be recognized.
[0188] According to an embodiment of the present disclosure, the determination module includes a recognition unit, a matching unit, and a first determination unit.
[0189] Among them, the recognition unit is used to input the bill image to be recognized into the layout recognition model, and output the target layout information corresponding to the bill image to be recognized.
[0190] The matching unit is used to match the target layout information with multiple standard bill templates pre-stored in the template library to generate a layout matching result, wherein the multiple standard bill templates correspond one-to-one to the multiple recognition model combinations pre-stored in the model library.
[0191] The first determination unit is used to determine whether the multiple recognition model combinations include a target recognition model combination associated with the target bill type according to the layout matching result.
[0192] According to an embodiment of the present disclosure, the target layout information includes first position information and first quantity information of multiple target element areas contained in the bill image to be identified, and second position information and second quantity information of multiple standard element areas marked in the bill standard template.
[0193] The matching unit includes a matching subunit, which is used to match the first position information of multiple target element areas with the second position information of multiple standard element areas, and to match the first quantity information of the target element areas with the second quantity information of multiple standard element areas to generate a layout matching result.
[0194] According to an embodiment of the present disclosure, the layout matching result includes a first matching result, and the first matching result is used to indicate whether the plurality of standard bill templates include a target standard bill template that matches the target layout information.
[0195] The first determining unit includes a first determining subunit and a second determining subunit.
[0196] The first determination subunit is used to determine that multiple recognition model combinations do not include a target recognition model combination associated with a target bill type when the matching result indicates that multiple bill standard templates do not include a target bill standard template matching the target layout information.
[0197] The second determining subunit is used to determine that the multiple recognition model combinations include a target recognition model combination associated with the target bill type when the matching result indicates that the multiple bill standard templates include a target bill standard template matching the target layout information.
[0198] According to an embodiment of the present disclosure, the layout matching result includes a second matching result, and the second matching result is used to characterize the matching degree between the target layout information and each bill standard template.
[0199] The first acquisition module includes a second determination unit and a first acquisition unit.
[0200] The second determining unit is used to determine a preferred bill standard template from multiple bill standard templates according to the second matching result.
[0201] The first acquisition unit is used to acquire a recognition model combination corresponding to the preferred bill standard template from multiple recognition model combinations as a basic recognition model combination.
[0202] According to an embodiment of the present disclosure, the annotation information of the sample image includes: area range annotation coordinate values, element annotation names, and element annotation values of each of the multiple element areas included in the sample image.
[0203] The second acquisition module includes a reading unit, a first generating unit, and a second acquisition unit.
[0204] The reading unit is used to read a pre-annotated image corresponding to a pre-selected sample image from a plurality of sample images from the pre-annotation system, and each element area in the pre-annotated image is marked with a range frame line.
[0205] The first generating unit is used to generate region range annotation coordinate values of the element region included in the sample image based on the pre-annotated image.
[0206] The second acquisition unit is used to acquire the element annotation name and the element annotation value of the element area contained in the sample image in combination with the general recognition model.
[0207] According to an embodiment of the present disclosure, the first generating unit includes a first identifying subunit and a third determining subunit.
[0208] The first recognition subunit is used to input the pre-annotated image into the coordinate recognition model, and output the frame line coordinate value corresponding to the range frame line in the pre-annotated image.
[0209] The third determining subunit is used to determine the frame line coordinate values as the area range annotation coordinate values of the element area included in the sample image.
[0210] According to an embodiment of the present disclosure, the second acquisition unit includes a second identification subunit, a reading subunit, and a combining subunit.
[0211] The second recognition subunit is used to input the sample image into the universal recognition model and output a first text recognition result for the printed text in the sample image.
[0212] The reading subunit is used to read the second character recognition result for the handwritten characters in the sample image from the pre-marking system.
[0213] The combining subunit is used to combine the first character recognition result and the second character recognition result as the element annotation name and element annotation value of the element area included in the sample image.
[0214] According to an embodiment of the present disclosure, the basic recognition model combination includes a basic semantic segmentation model and a basic text recognition model, the target recognition model combination includes a target semantic segmentation model and a target text recognition model, and the annotation information of the sample image includes the area range annotation coordinate values, element annotation names, and element annotation values of each of the multiple element areas contained in the sample image.
[0215] The tuning module includes a semantic segmentation unit, a text recognition unit, a calculation unit, and a second generation unit.
[0216] The semantic segmentation unit is used to input the sample image into the basic semantic segmentation model, and output the predicted coordinate values of the area range of each of the multiple element areas contained in the sample image.
[0217] The character recognition unit is used to input the sample image and the predicted coordinate values of the area range of each of the multiple element areas contained in the sample image into the character recognition model, and output the predicted element name and predicted element value of each of the multiple element areas contained in the sample image.
[0218] The calculation unit is used to calculate the loss value of the network loss function according to the element annotation names and element annotation values of each of the multiple element regions, and the element prediction names and element prediction values of each of the multiple element regions.
[0219] The second generating unit is used to generate a target semantic segmentation model and a target text recognition model associated with the target bill type when the loss value of the network loss function is less than a preset threshold.
[0220] Based on the above service calling method, the present disclosure also provides a service calling device. Figure 7 The device is described in detail.
[0221] Figure 7 The structural block diagram of the service invoking device 700 according to the embodiment of the present disclosure is schematically shown.
[0222] like Figure 7 As shown, the service calling device 700 of this embodiment includes a first calling module 701 , a second calling module 702 , and a third calling module 703 .
[0223] Among them, the first calling module 701 is used to respond to the recognition request for the bill image to be recognized, call the discrimination service to execute: determine whether the multiple recognition model combinations pre-stored in the model library contain a target recognition model combination associated with the target bill type of the bill image to be recognized.
[0224] The second calling module 702 is used to call the bill recognition service to execute, when it is determined that the target recognition model combination corresponding to the target bill type is included in the multiple recognition model combinations pre-stored in the model library: calling the target recognition model combination from the model library, and using the target recognition model combination to output the recognition result of the bill image to be recognized, and the recognition result is used to characterize the target element names and target element values of the multiple target element areas contained in the bill image to be recognized.
[0225] The third calling module 703 is used to call the bill marking service and the model tuning service when it is determined that the target recognition model combination is not included in the multiple recognition model combinations pre-stored in the model library, wherein the bill marking service is called to perform the operation of marking the bill image to be identified, and the model tuning service is called to execute: calling the basic recognition model combination from the model library, and tuning the parameters of the basic recognition model combination, generating a target recognition model combination associated with the target bill type, and adding the target recognition model combination to the model library.
[0226] According to the embodiment of the present disclosure, the first calling module 701 calls the discrimination service to determine whether the model library contains a target recognition model combination associated with the target bill type of the bill image to be recognized, and then performs the corresponding service calling operation according to different situations. In this way, when there is a target recognition model combination, the bill recognition service is called by the second calling module 702 to complete the online information recognition in time; when there is no target recognition model combination, the bill annotation service is called by the third calling module 703 to obtain the annotation information of the bill image to be recognized, and the bill information is temporarily obtained by annotating the information of this type of bill image as a bottom-line measure to ensure that the information production is not interrupted. At the same time, the model tuning service is called by the third calling module 703, and the basic recognition model combination is tuned online in real time using the annotated image to obtain a new recognition model. In this way, after the bill of the same type is subsequently transferred, the newly established model can be directly called to perform information recognition to obtain the bill information. It can be seen that by combining the online model recognition, annotation information collection and online real-time tuning, combined with different processing methods for new and old bills, compatible processing of new and old bills is achieved, and various bill types can be continuously adapted online to ensure uninterrupted production and improve business processing efficiency.
[0227] According to an embodiment of the present disclosure, any multiple modules in the determination module 601, the first acquisition module 602, the second acquisition module 603, the first tuning module 604, or the first call module 701, the second call module 702, and the third call module 703 can be combined in one module for implementation, or any one of the modules can be split into multiple modules. Alternatively, at least part of the functions of one or more of these modules can be combined with at least part of the functions of other modules and implemented in one module. According to an embodiment of the present disclosure, at least one of the determination module 601, the first acquisition module 602, the second acquisition module 603, the first tuning module 604, or the first call module 701, the second call module 702, and the third call module 703 can be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on a chip, a system on a substrate, a system on a package, an application specific integrated circuit (ASIC), or can be implemented by hardware or firmware such as any other reasonable way of integrating or packaging the circuit, or implemented in any one of the three implementation methods of software, hardware, and firmware or in an appropriate combination of any of them. Alternatively, the determination module 601, the first acquisition module 602, the second acquisition module 603, the first tuning module 604, or at least one of the first calling module 701, the second calling module 702, and the third calling module 703 can be at least partially implemented as a computer program module, which can perform corresponding functions when the computer program module is executed.
[0228] Figure 8 A block diagram of an electronic device suitable for implementing a recognition model generation method or a service calling method according to an embodiment of the present disclosure is schematically shown.
[0229] like Figure 8 As shown, the electronic device 800 according to an embodiment of the present disclosure includes a processor 801, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 802 or a program loaded from a storage part 808 to a random access memory (RAM) 803. The processor 801 may include, for example, a general-purpose microprocessor (e.g., a CPU), an instruction set processor and / or a related chipset and / or a special-purpose microprocessor (e.g., an application-specific integrated circuit (ASIC)), etc. The processor 801 may also include an onboard memory for caching purposes. The processor 801 may include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of the present disclosure.
[0230] In RAM 803, various programs and data required for the operation of electronic device 800 are stored. Processor 801, ROM 802 and RAM 803 are connected to each other via bus 804. Processor 801 performs various operations of the method flow according to the embodiment of the present disclosure by executing the program in ROM 802 and / or RAM 803. It should be noted that the program can also be stored in one or more memories other than ROM 802 and RAM 803. Processor 801 can also perform various operations of the method flow according to the embodiment of the present disclosure by executing the program stored in the one or more memories.
[0231] According to an embodiment of the present disclosure, the electronic device 800 may further include an input / output (I / O) interface 805, which is also connected to the bus 804. The electronic device 800 may further include one or more of the following components connected to the I / O interface 805: an input portion 806 including a keyboard, a mouse, etc.; an output portion 807 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage portion 808 including a hard disk, etc.; and a communication portion 809 including a network interface card such as a LAN card, a modem, etc. The communication portion 809 performs communication processing via a network such as the Internet. A drive 810 is also connected to the I / O interface 805 as needed. A removable medium 811, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 810 as needed, so that a computer program read therefrom is installed into the storage portion 808 as needed.
[0232] The present disclosure also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments; or may exist independently without being assembled into the device / apparatus / system. The above computer-readable storage medium carries one or more programs, and when the above one or more programs are executed, the method according to the embodiment of the present disclosure is implemented.
[0233] According to an embodiment of the present disclosure, the computer-readable storage medium may be a non-volatile computer-readable storage medium, for example, may include but is not limited to: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In the present disclosure, a computer-readable storage medium may be any tangible medium containing or storing a program, which may be used by or in combination with an instruction execution system, an apparatus or a device. For example, according to an embodiment of the present disclosure, a computer-readable storage medium may include the ROM 802 and / or RAM 803 described above and / or one or more memories other than ROM 802 and RAM 803.
[0234] The embodiment of the present disclosure also includes a computer program product, which includes a computer program, and the computer program contains program code for executing the method shown in the flowchart. When the computer program product is run in a computer system, the program code is used to enable the computer system to implement the recognition model generation method or service calling method provided by the embodiment of the present disclosure.
[0235] The above functions defined in the system / device of the embodiment of the present disclosure are executed when the computer program is executed by the processor 801. According to the embodiment of the present disclosure, the system, device, module, unit, etc. described above can be implemented by a computer program module.
[0236] In one embodiment, the computer program may rely on tangible storage media such as optical storage devices, magnetic storage devices, etc. In another embodiment, the computer program may also be transmitted and distributed in the form of signals on a network medium, and downloaded and installed through the communication part 809, and / or installed from a removable medium 811. The program code contained in the computer program may be transmitted using any appropriate network medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.
[0237] In such an embodiment, the computer program can be downloaded and installed from the network through the communication part 809, and / or installed from the removable medium 811. When the computer program is executed by the processor 801, the above functions defined in the system of the embodiment of the present disclosure are performed. According to the embodiment of the present disclosure, the system, device, means, module, unit, etc. described above can be implemented by a computer program module.
[0238] According to an embodiment of the present disclosure, the program code for executing the computer program provided by the embodiment of the present disclosure can be written in any combination of one or more programming languages. Specifically, these computing programs can be implemented using high-level process and / or object-oriented programming languages, and / or assembly / machine languages. Programming languages include, but are not limited to, Java, C++, python, "C" language or similar programming languages. The program code can be executed entirely on the user computing device, partially on the user device, partially on the remote computing device, or entirely on the remote computing device or server. In the case of a remote computing device, the remote computing device can be connected to the user computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computing device (for example, using an Internet service provider to connect through the Internet).
[0239] The flow charts and block diagrams in the accompanying drawings illustrate the possible architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flow chart or block diagram can represent a module, a program segment, or a part of a code, and the above-mentioned module, program segment, or a part of a code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order from the order marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram or flow chart, and the combination of the boxes in the block diagram or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0240] It will be appreciated by those skilled in the art that the features described in the various embodiments and / or claims of the present disclosure may be combined and / or combined in a variety of ways, even if such combinations and / or combinations are not explicitly described in the present disclosure. In particular, the features described in the various embodiments and / or claims of the present disclosure may be combined and / or combined in a variety of ways without departing from the spirit and teachings of the present disclosure. All of these combinations and / or combinations fall within the scope of the present disclosure.
[0241] The embodiments of the present disclosure are described above. However, these embodiments are only for illustrative purposes and are not intended to limit the scope of the present disclosure. Although the embodiments are described above separately, this does not mean that the measures in the various embodiments cannot be used in combination to advantage. The scope of the present disclosure is defined by the attached claims and their equivalents. Without departing from the scope of the present disclosure, those skilled in the art may make a variety of substitutions and modifications, which should all fall within the scope of the present disclosure.
Claims
1. A recognition model generation method, include: In response to a recognition request for a bill image to be recognized, determining whether a plurality of recognition model combinations pre-stored in a model library include a target recognition model combination associated with a target bill type of the bill image to be recognized; When it is determined that the target recognition model combination is not included in the plurality of recognition model combinations pre-stored in the model library, a basic recognition model combination is obtained from the model library; Receiving a plurality of sample images and acquiring annotation information of the plurality of sample images, wherein the sample images and the to-be-identified bill image belong to the same target bill type; The parameters of the basic recognition model combination are tuned based on a plurality of sample images and the annotation information of the plurality of sample images to generate a target recognition model combination associated with the target bill type.
2. The method according to claim 1, further comprising: include: When it is determined that the target recognition model combination corresponding to the target bill type is included in the multiple recognition model combinations pre-stored in the model library, the target recognition model combination is obtained from the model library; Inputting the to-be-recognized bill image into the target recognition model combination, and outputting a recognition result of the to-be-recognized bill image; Correcting the recognition result to obtain the bill information of the bill image to be recognized; Sending the bill information of the bill image to be identified to the business processing system; as well as The parameters of the target recognition model combination are optimized based on the bill information of the bill image to be recognized.
3. The method according to claim 1, in, Determining whether a plurality of recognition model combinations pre-stored in the model library includes a target recognition model combination associated with the target bill type of the bill image to be recognized includes: Input the to-be-recognized bill image into a layout recognition model, and output target layout information corresponding to the to-be-recognized bill image; Matching the target layout information with a plurality of standard bill templates pre-stored in a template library to generate a layout matching result, wherein the plurality of standard bill templates correspond one-to-one to the plurality of recognition model combinations pre-stored in a model library; According to the layout matching result, it is determined whether the multiple recognition model combinations include a target recognition model combination associated with the target bill type.
4. The method according to claim 3, in: The target layout information includes first position information and first quantity information of multiple target element areas included in the to-be-recognized bill image, and second position information and second quantity information of multiple standard element areas marked in the bill standard template; The target layout information is matched with a plurality of standard bill templates pre-stored in the template library to generate a layout matching result including: The first position information of the multiple target element areas is matched with the second position information of the multiple standard element areas, and the first quantity information of the target element areas is matched with the second quantity information of the multiple standard element areas to generate a layout matching result.
5. The method according to claim 3, in: The layout matching result includes a first matching result, and the first matching result is used to indicate whether the plurality of standard bill templates include a target standard bill template that matches the target layout information; Determining, according to the layout matching result, whether the multiple recognition model combinations include a target recognition model combination associated with the target bill type comprises: In a case where the matching result indicates that the multiple standard bill templates do not include a target bill standard template matching the target layout information, determining that the multiple recognition model combinations do not include a target recognition model combination associated with the target bill type; as well as When the matching result indicates that the multiple standard bill templates include a target standard bill template matching the target layout information, it is determined that the multiple recognition model combinations include a target recognition model combination associated with the target bill type.
6. The method according to claim 3, in: The layout matching result includes a second matching result, and the second matching result is used to represent the matching degree between the target layout information and each of the bill standard templates; The basic recognition model combination obtained from the model library includes: Determining a preferred bill standard template from the plurality of bill standard templates according to the second matching result; A recognition model combination corresponding to the preferred bill standard template is obtained from the multiple recognition model combinations as the basic recognition model combination.
7. The method according to any one of claims 1 to 6, in: The annotation information of the sample image includes: the region range annotation coordinate values, the element annotation names, and the element annotation values of each of the multiple element regions included in the sample image; Acquiring the annotation information of the plurality of sample images includes: Reading a pre-annotated image corresponding to a pre-selected sample image from a plurality of sample images from a pre-annotation system, wherein each element region in the pre-annotated image is marked with a range frame line; Based on the pre-annotated image, generating area range annotated coordinate values of the element area included in the sample image; In combination with the general recognition model, the element annotation name and the element annotation value of the element area included in the sample image are obtained.
8. The method according to claim 7, in, Generating the region range annotation coordinate values of the element region included in the sample image based on the pre-annotated image includes: Inputting the pre-annotated image into a coordinate recognition model, and outputting frame line coordinate values corresponding to the range frame lines in the pre-annotated image; The frame line coordinate values are determined as area range marking coordinate values of the element area included in the sample image.
9. The method according to claim 7, in, In combination with the general recognition model, obtaining the element annotation name and element annotation value of the element area contained in the sample image includes: Inputting the sample image into a universal recognition model, and outputting a first character recognition result for the printed characters in the sample image; Reading a second text recognition result for the handwritten text in the sample image from the pre-labeling system; The first character recognition result and the second character recognition result are combined as the element annotation name and the element annotation value of the element area included in the sample image.
10. The method according to any one of claims 1 to 6, in: The basic recognition model combination includes a basic semantic segmentation model and a basic text recognition model, the target recognition model combination includes a target semantic segmentation model and a target text recognition model, and the annotation information of the sample image includes the region range annotation coordinate values, the element annotation names, and the element annotation values of each of the multiple element regions included in the sample image; Optimizing the parameters of the basic recognition model combination based on the multiple sample images and the annotation information of the multiple sample images includes: Inputting the sample image into a basic semantic segmentation model, and outputting predicted coordinate values of the area ranges of each of the multiple element areas included in the sample image; Inputting the sample image and the predicted coordinate values of the area ranges of the multiple element areas contained in the sample image into the character recognition model, and outputting the predicted element names and predicted element values of the multiple element areas contained in the sample image; Calculating a loss value of a network loss function according to the element annotation names and element annotation values of each of the plurality of element regions, and the element prediction names and element prediction values of each of the plurality of element regions; When the loss value of the network loss function is less than a preset threshold, a target semantic segmentation model and a target text recognition model associated with the target bill type are generated.
11. A service calling method, include: In response to a recognition request for a bill image to be recognized, calling a discrimination service to execute: determining whether a plurality of recognition model combinations pre-stored in a model library include a target recognition model combination associated with a target bill type of the bill image to be recognized; In the case where it is determined that the target recognition model combination corresponding to the target bill type is included in the multiple recognition model combinations pre-stored in the model library, the bill recognition service is called to execute: calling the target recognition model combination from the model library, and using the target recognition model combination to output a recognition result of the bill image to be recognized, wherein the recognition result is used to characterize the target element names and target element values of the multiple target element areas included in the bill image to be recognized; When it is determined that the target recognition model combination is not included in the multiple recognition model combinations pre-stored in the model library, the bill labeling service and the model tuning service are called, wherein the bill labeling service is called to perform the operation of labeling the bill image to be recognized, and the model tuning service is called to execute: calling the basic recognition model combination from the model library, and tuning the parameters of the basic recognition model combination, generating a target recognition model combination associated with the target bill type, and adding the target recognition model combination to the model library.
12. A recognition model generation device, include: A determination module, configured to determine, in response to a recognition request for a bill image to be recognized, whether a plurality of recognition model combinations pre-stored in a model library includes a target recognition model combination associated with a target bill type of the bill image to be recognized; A first acquisition module is used to acquire a basic recognition model combination from the model library when it is determined that the target recognition model combination is not included in the multiple recognition model combinations pre-stored in the model library; A second acquisition module is used to receive a plurality of sample images and acquire annotation information of the plurality of sample images, wherein the sample images and the to-be-identified bill image belong to the same target bill type; The first tuning module is used to tune the parameters of the basic recognition model combination based on multiple sample images and the annotation information of the multiple sample images to generate a target recognition model combination associated with the target bill type.
13. A service calling device, include: The first calling module is used to call the discrimination service to execute, in response to a recognition request for a bill image to be recognized: determining whether a plurality of recognition model combinations pre-stored in a model library include a target recognition model combination associated with a target bill type of the bill image to be recognized; The second calling module is used to call the bill recognition service to execute, when it is determined that the target recognition model combination corresponding to the target bill type is included in the multiple recognition model combinations pre-stored in the model library: calling the target recognition model combination from the model library, and using the target recognition model combination to output a recognition result of the bill image to be recognized, wherein the recognition result is used to characterize the target element names and target element values of the multiple target element areas included in the bill image to be recognized; The third calling module is used to call the bill marking service and the model tuning service when it is determined that the target recognition model combination is not included in the multiple recognition model combinations pre-stored in the model library, wherein the bill marking service is called to perform the operation of marking the bill image to be identified, and the model tuning service is called to execute: calling the basic recognition model combination from the model library, and tuning the parameters of the basic recognition model combination, generating a target recognition model combination associated with the target bill type, and adding the target recognition model combination to the model library.
14. An electronic device, include: one or more processors; a storage device for storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors are enabled to execute the method according to any one of claims 1 to 11.
15. A computer-readable storage medium having executable instructions stored thereon, which, when executed by a processor, causes the processor to execute the method according to any one of claims 1 to 11.
16. A computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the method according to any one of claims 1 to 11 is implemented.
Citation Information
Patent Citations
Certificate classification method and device, electronic equipment and storage medium
CN112801086A
Method and apparatus for extracting information about a negotiable instrument, electronic device and storage medium
US20220148324A1