Card text recognition method and device, electronic equipment and storage medium

By inputting the card certificate image into the text detection model, obtaining the detection results of text lines and text blocks, and performing character recognition, the problem of poor generalization ability of card certificate recognition in the existing technology is solved, and effective identification of different card surface typesetting and newly issued version card certificates is achieved.

CN120088807APending Publication Date: 2025-06-03BEIJING HISIGN TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510103014.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-22
Publication Date
2025-06-03

AI Technical Summary

Technical Problem

The existing technology is difficult to meet various card identification needs, especially the new version of card certificates generated by changes in card surface layout, and the generalization ability is poor.

Method used

By obtaining the image to be recognized by the card certificate, input it to the pre-trained text detection model, obtain the text line and text block detection results, and determine the text line to be recognized based on these results, and finally character recognition is performed on the text line to realize the text recognition of the card certificate.

Benefits of technology

It realizes text recognition of different card typesettings and newly issued versions of card certificates, improves the generalization ability and scope of application of recognition, and provides a more general card certificate text recognition solution.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120088807A_ABST
    Figure CN120088807A_ABST
Patent Text Reader

Abstract

The invention provides a card text recognition method and device, electronic equipment and a storage medium, and belongs to the technical field of artificial intelligence, and the method comprises the steps: obtaining a to-be-recognized image of a card; inputting the to-be-recognized image into a pre-trained text detection model to obtain a first text line detection result and a first text block detection result output by the text detection model; determining a to-be-identified text line of the card based on the first text line detection result and the first text block detection result; and performing character recognition on the to-be-recognized text line to obtain a text recognition result of the card. According to the method, the card image is input into the text detection model to detect the text line and the text block, and the text line needing to be recognized is determined from the text line and the text block according to the requirements of items, services and the like, so that the situation that specific text content cannot be directly positioned due to different card typesetting, and finally text recognition fails can be avoided; the invention provides a card text recognition scheme which is wider in application range and higher in universality.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular, to a method, device, electronic device and storage medium for identifying card and certificate text. Background Art

[0002] With the continuous enrichment of the card surface customization services for cards and certificates such as bank cards, transportation cards, medical insurance cards, and integrated circuit cards, the font, position, layout, etc. of the card numbers on the card surface often vary, increasing the difficulty of identifying bank cards, transportation cards and other cards and certificates.

[0003] Currently, the commonly used method for identifying card and certificate text is mainly to train a targeted recognition model for the card surface content such as the card numbers, Chinese and English characters to be recognized, and then apply the trained recognition model to the card and certificate recognition inference. Taking the identification of a bank card number as an example, the bank card image is input into the trained recognition model, and the recognition model will directly locate and recognize the bank card number, and then obtain the bank card number recognition result.

[0004] However, in the case of the continuous enrichment of the card surface customization services and the continuous release of new versions of cards and certificates, the method of training a targeted recognition model to recognize specific card surface content can no longer meet the various card and certificate recognition requirements, nor can it effectively recognize the cards and certificates with changed card surface layout and newly released versions, and has poor generalization ability. Summary of the Invention

[0005] The present invention provides a method, device, electronic device and storage medium for identifying card and certificate text, so as to solve the defects in the prior art that it cannot meet the various card and certificate recognition requirements, nor can it effectively recognize the cards and certificates with changed card surface layout and newly released versions, and has poor generalization ability, and realizes a card and certificate recognition solution with strong generalization ability and applicable to various different card surface layouts.

[0006] The present invention provides a method for identifying card and certificate text, including: Obtaining a to-be-recognized image of a card and certificate; Inputting the to-be-recognized image into a pre-trained text detection model to obtain a first text line detection result and a first text block detection result output by the text detection model; Based on the first text line detection result and the first text block detection result, determining the to-be-recognized text line of the card and certificate; Performing character recognition on the to-be-recognized text line to obtain the text recognition result of the card and certificate.

[0007] According to the method for identifying card and certificate text provided by the present invention, the determining the to-be-recognized text line of the card and certificate based on the first text line detection result and the first text block detection result includes: Normalize the detection result of the first text block to obtain a text block image to be recognized with the same size as the image to be recognized. Input the text block image to be recognized into the text detection model to obtain the second text line detection result and the second text block detection result output by the text detection model. Based on the first text line detection result and the second text line detection result, determine the text line to be recognized.

[0008] According to a card text recognition method provided by the present invention, the text detection model is trained based on a plurality of training samples; each training sample includes a recognition image sample and its corresponding text line label and text block label; each training sample is obtained based on the following method: Obtain the recognition image sample. Annotate the text lines in the recognition image sample to obtain the text line label corresponding to the recognition image sample. When the number of the text line labels is more than two, based on the text lines with a line spacing less than a preset pixel value, determine the small-spacing text line labels from the text line labels, and merge the small-spacing text line labels into the text block label corresponding to the recognition image sample.

[0009] According to a card text recognition method provided by the present invention, the character recognition of the text line to be recognized to obtain the text recognition result of the card includes: Input the text line to be recognized into a pre-trained character classification model to obtain the character type with the highest confidence output by the character classification model. Based on the character type, determine at least one character recognition model to be called. When there is one character recognition model, input the text line to be recognized into the character recognition model to obtain the text recognition result output by the character recognition model. When there are more than two character recognition models, input the text line to be recognized into more than two character recognition models respectively to obtain the text recognition results output by more than two character recognition models respectively.

[0010] According to a card text recognition method provided by the present invention, the types of the character recognition models include Chinese character recognition models, English character recognition models, and digital character recognition models.

[0011] According to a card text recognition method provided by the present invention, the obtaining of the image to be recognized of the card includes: Obtain the image to be corrected of the card. Input the image to be corrected into a pre-trained text direction recognition model to obtain the text angle with the highest confidence output by the text direction recognition model; Based on the text angle, perform rotation correction on the image to be corrected to obtain the image to be recognized.

[0012] According to a card text recognition method provided by the present invention, the obtaining of the image to be corrected of the card includes: Obtain the original captured image of the card; Input the original captured image into a pre-trained target detection model to obtain the rectangular detection frame output by the target detection model; Use the rectangular detection frame to crop the original captured image to obtain a rough card image; Input the rough card image into a pre-trained edge detection model to obtain the vertex coordinate values output by the edge detection model; Based on the vertex coordinate values, perform perspective transformation and cropping on the rough card image to obtain the image to be corrected.

[0013] The present invention also provides a card recognition device, including: An image acquisition module for acquiring an image to be recognized of a card; An image detection module for inputting the image to be recognized into a pre-trained text detection model to obtain a first text line detection result and a first text block detection result output by the text detection model; A text line detection module for determining the text line to be recognized of the card based on the first text line detection result and the first text block detection result; A text line recognition module for performing character recognition on the text line to be recognized to obtain the text recognition result of the card.

[0014] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and running on the processor. When the processor executes the computer program, the card text recognition method as described in any one of the above is implemented.

[0015] The present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the card text recognition method as described in any one of the above is implemented.

[0016] The ID card text recognition method, device, electronic device and storage medium provided by the present invention input the image of the ID card into a text detection model to obtain the text line detection result and text block detection result in the ID card image, and screen out the text lines that need to be recognized and detected from the detected text line detection results and text block detection results according to the actual project, business and other requirements, so as to recognize the text lines corresponding to the content on the ID card surface that needs to be detected. It can avoid the situation where it is impossible to directly locate the specific text content on different ID cards due to different layouts of the ID card surface content, resulting in text recognition failure. It can perform text recognition on ID cards of different types, different versions, especially newly issued versions with any text layout, and provides an ID card text recognition solution with a wider application range and stronger versatility. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0018] Figure 1 is one of the flow diagrams of the ID card text recognition method provided by the present invention.

[0019] Figure 2 is the second flow diagram of the ID card text recognition method provided by the present invention.

[0020] Figure 3 is the structural diagram of the ID card text recognition device provided by the present invention.

[0021] Figure 4 is the structural diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0022] To make the objectives, technical solutions and advantages of the present invention clearer, the following will clearly and completely describe the technical solutions in the present invention with reference to the accompanying drawings in the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art without creative efforts based on the embodiments in the present invention belong to the scope of protection of the present invention.

[0023] It should be noted that in the description of the present invention, the terms "comprising", "including" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements but also includes other elements not expressly listed, or elements inherent to such process, method, article or device. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.

[0024] The terms "first", "second", etc. in the present invention are used to distinguish similar objects and are not used to describe a specific order or sequence. It should be understood that such used data can be interchanged under appropriate circumstances so that the embodiments of the present invention can be implemented in an order other than those illustrated or described herein, and the objects distinguished by "first", "second", etc. are generally of the same category, and the number of objects is not limited. For example, the first object can be one or multiple. In addition, "and / or" means at least one of the connected objects, and the character " / " generally indicates an "or" relationship between the associated objects before and after.

[0025] The following Figures 1-4 describes the card and certificate text recognition method, device, electronic device and storage medium provided by the present invention.

[0026] Figure 1 is one of the flow diagrams of the card and certificate text recognition method provided by the present invention. As Figure 1 shown, the card and certificate text recognition method includes but is not limited to steps 101 to 104.

[0027] It should be noted that the execution subject of the card and certificate text recognition method provided by the present invention can be a server, a computer device, such as a mobile phone, a tablet computer, a notebook computer, a handheld computer, a vehicle-mounted electronic device, a wearable device, an Ultra-Mobile Personal Computer (UMPC), a netbook or a Personal Digital Assistant (PDA), etc.

[0028] Step 101: Obtain a to-be-recognized image of the card and certificate.

[0029] Among them, the card and certificate includes but is not limited to any one of bank cards, transportation cards, bus cards, identity cards, medical insurance cards, etc. with different layouts, versions and / or different card surface contents.

[0030] Specifically, the to-be-recognized image of the card certificate can be retrieved from the card certificate image database after certain image preprocessing, or the original captured image of the card certificate obtained by real-time shooting can be subjected to real-time image preprocessing such as rotation correction, perspective transformation, and cropping to obtain the to-be-recognized image of the card certificate. In some cases, if the original captured image obtained by real-time shooting meets the preset image requirements (such as the proportion of the background part other than the card certificate in the original captured image is less than 10%, etc.), the original captured image obtained by shooting can be directly used as the to-be-recognized image.

[0031] Step 102: Input the to-be-recognized image into a pre-trained text detection model to obtain the first text line detection result and the first text block detection result output by the text detection model.

[0032] Among them, the text detection model is a pre-trained detection model used to recognize text lines and text blocks in the to-be-recognized image.

[0033] Specifically, input the to-be-recognized image of the card certificate into a pre-trained text detection model, and the text detection model processes the text lines and text blocks on the card surface of the to-be-recognized image to obtain the first text line detection result and the first text block detection result output by the text detection model.

[0034] Generally speaking, the first text line detection result includes several text line mask maps, corresponding to the text lines on the to-be-recognized image of the card certificate at the same positions as the text line mask maps; the first text block detection result includes several text block mask maps, corresponding to the text blocks composed of more than two text lines at the same positions as the text block mask maps on the to-be-recognized image of the card certificate.

[0035] For the text line mask maps in the first text line detection result and the text block mask maps in the first text block detection result output by the text detection model, their positions can correspond to any positions on the card surface of the card certificate where text content is typeset. By detecting text lines and text blocks, rather than directly locating the positions where specific text content is typeset on the card certificate image, even if the positions where the same specific text content is typeset on different types and versions of card certificates are different, all text-related content on the card certificate can be obtained through the text detection model.

[0036] Taking a bank card as an example, regardless of whether the text content such as the bank card number, expiration date, and bank card name of the bank card is on the front or back of the bank card, or whether it is on the upper side, center, or lower side of the card surface, etc., the text line mask maps and text block mask maps corresponding to the text content such as the bank card number, expiration date, and bank card name can be detected through the text detection model.

[0037] As an alternative embodiment, the text detection model uses efficientNet as the backbone network and adopts a Feature Pyramid Network (FPN) for feature fusion. It combines low-level detailed information and high-level semantic information in a top-down manner, and finally sets up a binary classification convolutional layer, where the text line is one category and the text block is another category, so that the output of the text detection model is divided into two channels. The first text line detection result is used as the output of one channel, and the first text block detection result is used as the output of the other channel.

[0038] Step 103: Based on the first text line detection result and the first text block detection result, determine the text lines to be recognized on the card certificate.

[0039] Specifically, after obtaining the first text line detection result and the first text block detection result output by the text detection model, it is necessary to further determine the text lines to be recognized on the card certificate according to the first text line detection result and the first text block detection result. At this time, it is further determined whether the text block mask image is included in the first text block detection result, and whether it is necessary to recognize the card certificate content corresponding to the first text block detection result according to the requirements of the current project, business, etc.

[0040] In the case where the text block mask image is not included in the first text block detection result or it is not necessary to recognize the card certificate content corresponding to the first text block detection result, the first text block detection result can be ignored and discarded, and the vertex coordinates of the text lines detected in the text line mask image are directly calculated according to the several text line mask images included in the first text line detection result, and then several text lines to be recognized are determined according to the vertex coordinates of the text lines.

[0041] In the case where it is necessary to recognize the card certificate content corresponding to the first text block detection result, the several text block mask images included in the first text block detection result are further processed to more precisely segment the text lines in the text block mask image, and then the text lines in the text block mask image and the text lines in the first text block detection result are used as the text lines to be recognized.

[0042] Step 104: Perform character recognition on the text lines to be recognized to obtain the text recognition result of the card certificate.

[0043] Specifically, after simultaneously detecting the text lines and text blocks of the image to be recognized of the card certificate, and determining the text lines to be recognized of the card certificate required by the current project and business according to the obtained first text line detection result and the first text block detection result, character recognition processing is performed on the text lines to be recognized to obtain the text recognition result of the card certificate.

[0044] Taking the card certificate as a bank card as an example, the text recognition results obtained include at least one of the content on the card surface such as Chinese characters, English characters, and / or digital characters, including but not limited to bank card numbers, expiration dates, bank card names, etc.

[0045] The card certificate text recognition method provided by the present invention inputs the image of the card certificate into the text detection model to obtain the text line detection result and text block detection result in the card certificate image, and screens out the text lines to be recognized and detected from the detected text line detection result and text block detection result according to the requirements of actual projects, businesses, etc., so as to recognize the text lines corresponding to the content on the card surface to be detected. It can avoid the situation where the specific text content on different card certificates cannot be directly located due to different layouts of the content on the card surface of the card certificate, resulting in text recognition failure. It can perform text recognition on different types and different versions of card certificates, especially newly issued versions, with any text layout, providing a card certificate text recognition solution with a wider application range and stronger versatility.

[0046] Based on the above embodiments, as an optional embodiment, determining the text line to be recognized of the card certificate based on the first text line detection result and the first text block detection result includes: Performing normalization processing on the first text block detection result to obtain a to-be-recognized text block image with the same size as the to-be-recognized image; Inputting the to-be-recognized text block image into the text detection model to obtain the second text line detection result and the second text block detection result output by the text detection model; Determining the to-be-recognized text line based on the first text line detection result and the second text line detection result.

[0047] Specifically, if, according to the requirements of the current project, business, etc., after the to-be-recognized image is first input into the pre-trained text detection model to obtain the first text line detection result and the first text block detection result, and it is determined that the content of the card certificate corresponding to the first text block detection result needs to be recognized, then according to several text block mask graphs included in the first text block detection result, calculate the vertex coordinates of the text blocks detected in the text block mask graphs, and then determine the to-be-recognized text blocks that need to be detected again and include more than two text lines according to the vertex coordinates of the text blocks. Perform normalization processing on the image size of the to-be-recognized text blocks so that the image size of the to-be-recognized text blocks is the same as the size of the to-be-recognized image, and obtain several to-be-detected to-be-recognized text block images.

[0048] Input the image of the text block to be recognized into the text detection model, and obtain the second text line detection result and the second text block detection result output by the text detection model. At this time, according to a number of text line mask maps included in the first text line detection result and the second text line detection result, calculate the vertex coordinates of the detected text lines in the text line mask maps, and then determine a number of text lines to be recognized according to the vertex coordinates of the text lines.

[0049] Generally speaking, for recognizing the content on the card surface of an identification card, after being processed by the text detection model twice, according to the text line mask maps included in the first text line detection result and the second text line detection result output by the text detection model, the text lines to be recognized that meet different project and business requirements can be basically detected.

[0050] In some exceptional cases, for example, when the second text block detection result still includes text block mask maps of interest to the project or business, for the second text block detection result, the step of normalizing the text block detection result can be repeatedly executed to obtain a text block image to be recognized with the same size as the image to be recognized, input the text block image to be recognized into the text detection module, and determine the text lines to be recognized that are of interest to the project or business according to the text line mask maps included in the text line detection results output by the text detection module multiple times.

[0051] The identification card text recognition method provided by the present invention can detect the text lines in the text block by re-inputting the text block detected for the first time into the text detection model, and can more precisely segment the text lines required for different identification cards with different layouts in different identification card text recognition tasks, avoiding the situation where the specific text content on different identification cards cannot be directly located due to different layouts of the card surface content, and ultimately resulting in text recognition failure, and has a wider application range and stronger versatility.

[0052] Based on the above embodiments, as an alternative embodiment, the text detection model is trained based on a plurality of training samples; each training sample includes an identification image sample and its corresponding text line label and text block label; each training sample is obtained in the following manner: Obtain the identification image sample; Annotate the text lines in the identification image sample to obtain the text line label corresponding to the identification image sample; When the number of the text line labels is more than two, based on the text lines with a line spacing less than a preset pixel value, determine small-spacing text line labels from the text line labels, and merge the small-spacing text line labels into the text block label corresponding to the identification image sample.

[0053] Specifically, in the training stage of the text detection model, it is first necessary to obtain training samples for training the text detection model.

[0054] When obtaining each training sample, first collect recognition image samples that meet the requirements of model training, and adjust the size of the recognition image samples to meet the input requirements of the text detection model. Further, use annotation tools such as PPOCRLabel to annotate several text lines of the recognition image samples to obtain text line labels corresponding to the recognition image samples.

[0055] In the case where the number of text line labels of the recognition image sample is more than two, if there are two text line labels with a line spacing less than the preset pixel value, the text line labels with a line spacing less than the preset pixel value are used as small-spacing text line labels, and then the small-spacing text line labels are merged and annotated as text block labels corresponding to the recognition image sample.

[0056] Among them, the preset pixel value can be determined according to the actual situation such as the type of card certificate and the recognition task, and there is no limit to this. Taking a bank card as an example, the line spacing of text lines such as card numbers and bank names often exceeds the preset pixel value, and text line labels are obtained after annotation, while the line spacing of text lines such as the precautions at the bottom of the back of the bank card often is less than the preset pixel value, and small-spacing text line labels are obtained after annotating each text line of precautions, and the small-spacing text line labels are merged to obtain the corresponding text block labels.

[0057] Combining the recognition image sample and its corresponding text line labels and text block labels obtained by annotation together, a training sample is obtained. Repeating the steps of annotating text line labels and merging small-spacing text line labels, multiple training samples for training the text detection model are obtained.

[0058] Finally, use multiple training samples to train the text prediction model, and use loss functions such as Dice loss function (DiceLoss), cross-entropy loss, and L1 loss to perform convergence training on the text prediction model until the training iteration times meet the maximum iteration times or meet iteration termination conditions such as preset prediction accuracy, and a pre-trained text prediction model is obtained.

[0059] It can be understood that if the text block detection results (the first text block detection result and the second text block detection result) output by the text detection model pre-trained based on the foregoing multiple training samples, and the line spacing between each text line in the text block mask map is less than the preset pixel value.

[0060] The card and certificate text recognition method provided by the present invention, by considering that the sizes and line spacings of texts on the same card and certificate are different, adopts a combined design of text line and text block detection during text detection. For small texts with a line spacing less than a preset pixel value, they are marked as text blocks for supervised training, and for large texts with a line spacing greater than or equal to the preset pixel value, they are marked as text lines for supervised training. A text detection model capable of finely segmenting text lines with different line spacings can be trained to obtain the text lines required for different card and certificate text recognition tasks for various types of card and certificate layouts, thereby improving the recognition effect of card and certificate texts.

[0061] Based on the above embodiments, as an optional embodiment, the character recognition of the text line to be recognized to obtain the text recognition result of the card and certificate includes: Input the text line to be recognized into a pre-trained character classification model to obtain the character type with the highest confidence output by the character classification model; Based on the character type, determine at least one character recognition model to be called; When there is one character recognition model, input the text line to be recognized into the character recognition model to obtain the text recognition result output by the character recognition model; When there are two or more character recognition models, input the text line to be recognized into two or more of the character recognition models respectively to obtain the text recognition results output by the two or more character recognition models respectively.

[0062] Wherein, each character type is determined based on at least one of character types obtained from Chinese characters, numerical characters, English characters and their combinations.

[0063] For example, each character type is at least one of "Chinese characters", "numerical characters", "English characters", "Chinese characters + numerical characters", "Chinese characters + English characters", "numerical characters + English characters", "Chinese characters + numerical characters + English characters".

[0064] Optionally, the types of character recognition models include but are not limited to Chinese character recognition models, English character recognition models, numerical character recognition models, Greek letter character recognition models, and so on.

[0065] Specifically, after obtaining the text lines to be recognized of the card certificate using the text detection model, character recognition needs to be performed on each text line to be recognized, that is, each text line to be recognized is input into a pre-trained character classification model to obtain the character type with the highest confidence output by the character classification model. According to the character type output by the character classification model, the characters included in the text line to be recognized can be determined, that is, it can be determined whether the text line to be recognized includes Chinese characters, numerical characters, and / or English characters, and then at least one character recognition model for recognizing the text line to be recognized is determined according to the characters included in the text line to be recognized.

[0066] When there is one determined character recognition model, the text line to be recognized is input into the character recognition model to obtain the text recognition result output by the character recognition model.

[0067] For example, if it is determined according to the character type that the character recognition model is a Chinese character recognition model, the text line to be recognized is input into the Chinese character recognition model to obtain the Chinese recognition result output by the Chinese character recognition model as the text recognition result of the text line to be recognized.

[0068] When there are two or more determined character recognition models, the text line to be recognized is respectively input into the two or more character recognition models to obtain the text recognition results respectively output by the two or more character recognition models. For example, if it is determined according to the character type that the character recognition models are a Chinese character recognition model and an English character recognition model, the text line to be recognized is respectively input into the Chinese character recognition model and the English character recognition model to obtain the Chinese recognition result and the English recognition result respectively output by the Chinese character recognition model and the English character recognition model, which are jointly used as the text recognition result of the text line to be recognized.

[0069] Optionally, the character classification model is pre-trained based on the following method: Text line images with different lengths and widths are collected, normalized to a preset size (such as 480*320 or 64*320, etc.) to obtain text line samples, and multi-classification is performed on the text line samples according to the character types of the texts to obtain the character type labels corresponding to the text line samples.

[0070] For example, character type tags can be divided into three categories: The first type of character type tag is the "tag containing Chinese characters". For text line samples that only include Chinese characters, and those that include both Chinese characters and other characters (English characters and / or numerical characters), they are all set as the "tag containing Chinese characters"; The second type of character type tag is the "tag containing English characters and not containing Chinese characters". For text line samples that only include English characters, and those that include both English characters and numerical characters, they are all set as the "tag containing English characters and not containing Chinese characters"; The third type of character type tag is the "tag only containing numerical characters". For text line samples that only include numerical characters, they are all set as the "tag only containing numerical characters".

[0071] Further, a character classification model is built based on the Lcnet network, and the channel coefficient is set to 0.25. The number of categories is set according to the type of character type tag (such as 3). Multiple text line samples are input into the character classification model, and the cross-entropy loss between the model output value and the character type tag corresponding to the text line sample is calculated to perform regression iterative training, and a pre-trained character classification model is obtained.

[0072] The card and certificate text recognition method provided by the present invention first performs multi-classification on the character type of the text line to be recognized, and then calls the corresponding character recognition model according to the classification result of the text line to be recognized for character recognition to obtain the text recognition result, avoiding calling an irrelevant character recognition model to process the text line to be recognized. While reducing the waste of computing resources, it improves the speed of card and certificate text recognition and can be applied to the recognition of card and certificate texts with various layouts and various demand services.

[0073] Based on the above embodiments, as an optional embodiment, the types of the character recognition models include a Chinese character recognition model, an English character recognition model, and a numerical character recognition model.

[0074] For example, the Chinese character recognition model is built based on the Structured Visual Reasoning Task (SVRT) and the Long Short-Term Memory (LSTM). Common Chinese characters and rare Chinese characters in the printing scenario are collected as the training set to train the Chinese character recognition model, and a Chinese character recognition model that can recognize a Chinese character dictionary composed of 27,000 common Chinese characters and rare Chinese characters is obtained.

[0075] For another example, the English character recognition model is constructed based on Relational and Episodic Visual Reasoning and Temporal Task (REPSVRT) and the LSTM network. English letters in scenarios such as printed scenes and relief scenes are collected as the training set to train the English character recognition model, and an English character recognition model that can recognize the English dictionary library composed of all letters A-Z, a-z, and English symbols is obtained.

[0076] For still another example, the digital character recognition model is constructed based on REPSVRT and the LSTM network. Digits in scenarios such as printed scenes and relief scenes are collected as the training set to train the digital character recognition model, and a digital character recognition model that can recognize the digital dictionary library composed of all digits 0-9 and digital symbols is obtained.

[0077] Optionally, according to the specific usage scenarios of different services and project requirements, different character recognition models can be selected individually or in combination to recognize the text line to be recognized.

[0078] For example, for the card and certificate text recognition task that only needs to recognize the bank card number, only the digital character recognition model needs to be called to recognize the text line to be recognized.

[0079] By classifying the character recognition model into a Chinese character recognition model, an English character recognition model, and a digital character recognition model, the interference recognition between the recognition of numbers, Chinese, and English during character recognition can be reduced.

[0080] Based on the above embodiments, as an optional embodiment, the obtaining of the image to be recognized of the card and certificate includes: Obtaining the image to be corrected of the card and certificate; Inputting the image to be corrected into a pre-trained text direction recognition model to obtain the text angle with the highest confidence output by the text direction recognition model; Based on the text angle, performing rotation correction on the image to be corrected to obtain the image to be recognized.

[0081] Among them, the text direction recognition model is trained based on the image samples to be corrected and their corresponding text angle labels.

[0082] For example, the text angle label is any one of angle labels such as 0° label, 90° label, 180° label, and 270° label.

[0083] With the popularization and enrichment of card customization services, the text layout on cards appears in different forms such as horizontal and vertical arrangements. Even when fixing the card in a horizontal card form (i.e., the width of the card is greater than the height of the card) or a vertical card form (i.e., the height of the card is greater than the width of the card) during card image processing, it is still impossible to ensure that the text direction on the card is positive, and the card image cannot be input into the text detection image for text line and text block detection. Therefore, further rotation correction of the card image is required.

[0084] Specifically, obtain the image to be corrected of the card in the horizontal card form or the vertical card form, input the image to be corrected into the trained text direction recognition model, and let the text direction recognition model classify the text direction of the image to be corrected to obtain the text angle with the highest confidence output by the text direction recognition model. Further, according to this text angle with the highest confidence, perform rotation correction on the image to be corrected to obtain the image to be recognized with the positive text direction of the card, so that the image to be recognized can be input into the text detection image for text line and text block detection.

[0085] Optionally, the text direction recognition model is pre-trained based on the following method: collect card images in different forms of horizontal card form and vertical card form, and normalize the card images to a preset size (such as 128*128 size, etc.) as the image sample to be corrected; ignore the non-text background and its pattern direction in the image sample to be corrected, and label the text angle label of the image sample to be corrected according to the text direction in the image sample to be corrected to obtain the training set for training the text direction recognition model; build the text direction recognition model based on the LCNet network and set the channel coefficient to 0.25; use the training set to train the text direction recognition model, calculate the cross-entropy loss between the model output value and the text angle label corresponding to the image sample to be corrected for regression iterative training, and obtain the pre-trained text direction recognition model.

[0086] The card text recognition method provided by the present invention, by inputting the card image into the text direction recognition model, having the text direction recognition model classify the text direction of the image to be corrected, and performing rotation correction on the image to be corrected according to the text angle output by the model to obtain the image to be recognized with the positive text direction, helps to improve the accuracy of subsequent text line and text block detection and character recognition, and is applicable to card text recognition with various layouts.

[0087] Based on the above embodiments, as an optional embodiment, the obtaining of the image to be corrected of the card includes: Obtain the original captured image of the card; Input the original captured image into the pre-trained target detection model to obtain the rectangular detection frame output by the target detection model; Crop the original captured image using the rectangular detection frame to obtain a rough card image; Input the rough card image into a pre-trained edge detection model to obtain the vertex coordinate values output by the edge detection model; Based on the vertex coordinate values, perform perspective transformation and cropping on the rough card image to obtain an image to be corrected.

[0088] Among them, the object detection model is trained based on captured image samples and their corresponding detection frame labels; the edge detection model is trained based on rough card image samples and their corresponding vertex coordinate labels.

[0089] Since the distances and angles at which the user terminal captures the card vary greatly, the card appears in the original captured image with situations such as being too small in size and having an irregular shape (such as a trapezoidal shape with an angle), which cannot be directly corrected, detected, or recognized. Therefore, it is necessary to perform certain preprocessing on the original captured image of the card.

[0090] Specifically, obtain the original captured image of the card uploaded or captured by the user terminal, input the original captured image into a pre-trained object detection model to obtain the rectangular detection frame output by the object detection model, and use the rectangular detection frame to crop the original captured image, deleting the scene background part of the original captured image that is not the card image, ensuring that the scene background part meets the preset requirements (such as less than 20% of the whole image, etc.), to obtain a rough card image.

[0091] Since object detection can only obtain a rectangular detection frame, the rough card image obtained after cropping can only remove part of the scene background part and cannot handle the situation of irregular shapes. Therefore, it is necessary to further process the rough card image.

[0092] Further input the rough card image into a pre-trained edge detection model, use the edge detection model to detect the edge of the card to obtain the vertex coordinate values output by the edge detection model, and perform perspective transformation on the rough card image using the vertex coordinate values to transform the card part in the image into a rectangle to obtain an intermediate image, and further crop the intermediate image to delete the scene background part of the non-card image to obtain an image to be corrected that has not been rotated and corrected according to the text direction.

[0093] In addition, if the object detection model does not output a rectangular detection frame, it means that there is no card in the original captured image. At this time, return the detection result of not detecting the card. Optionally, the object detection model is deployed on the front end (such as the H5 end).

[0094] Optionally, the object detection model is trained using the PaddleDetection framework, with the lightweight yunnet detection network as the backbone feature extraction module, and is constructed using a multi-head mechanism that includes an object position regressor and a class classifier. After the object detection model is trained, the onnx2ncnn tool is used to convert the trained model format into the float16 ncnn format, and the weights are compressed and quantized to control the volume of the object detection model within 1M, so as to meet the basic requirements of small algorithm volume and high accuracy for the deployment of the object detection model on the front end.

[0095] Optionally, the edge detection model is constructed based on the Gated Shape Convolutional Neural Network (Gated-SCNN) with geometric structure information, and canny edge detection is added in different sizes to extract shape features, improve noise resistance, and increase the edge feature information of the certificate.

[0096] Optionally, by collecting card image data and performing preprocessing, a captured image sample and / or a rough card image sample are obtained. Further, the PPOCRLabel tool is used for edge annotation to obtain the detection box label corresponding to the captured image sample and the vertex coordinate label corresponding to the rough card image sample.

[0097] The card text recognition method provided by the present invention performs object detection on the original captured image of the card, obtains a rectangular detection box to crop the original captured image to obtain a rough card image, and then performs edge detection on the rough card image to obtain vertex coordinate values for perspective transformation and cropping of the rough card image. Even if the forms, positions, sizes, lengths, and widths of the cards in the original captured image are all different, a card image suitable for subsequent correction, detection, and recognition steps can be obtained through object detection and edge detection, which can meet the requirements of various card recognition scenarios, further broaden the application scope of the card text recognition method, and enhance the versatility of the card text recognition method.

[0098] To better illustrate the card text recognition method provided by the present invention, an embodiment is provided below to illustrate the whole process of the card recognition method.

[0099] Figure 2 is the second flow diagram of the card text recognition method provided by the present invention, as Figure 2As shown, after obtaining the original captured image of the card or certificate uploaded or captured by the user, the original captured image is input into a pre-trained object detection model to obtain the rectangular detection frame output by the object detection model, and the original captured image is cropped using the rectangular detection frame to obtain a rough card or certificate image. The rough card or certificate image is input into a pre-trained edge detection model to obtain the vertex coordinate values output by the edge detection model, so as to perform perspective transformation on the rough card or certificate image using the vertex coordinate values to obtain a to-be-corrected image of the card or certificate with standardized width and height, thereby standardizing the situation where the card or certificate image is deformed due to factors such as the shooting angle and distance of the card or certificate.

[0100] Further, the to-be-corrected image is input into a text direction recognition model, and the text direction recognition model classifies the text direction of the to-be-corrected image to obtain the text angle with the highest confidence output by the text direction recognition model, so as to perform rotation correction based on the text direction on the to-be-corrected image using the text angle to obtain a to-be-recognized image after text direction correction.

[0101] The to-be-recognized image is input into a text detection model, and the text detection model detects the text lines and text blocks of the card or certificate in the to-be-recognized image to obtain the first text line detection result and the first text block detection result output by the text detection model. If, based on project and business requirements, the first text block detection result is not of interest, then the first text block detection result is ignored, and only the text line mask map included in the first text line detection result is used to determine the to-be-recognized text line, so as to perform character recognition on the to-be-recognized text line.

[0102] If the first text block detection result is of interest, it is necessary to further extract the text lines in the text block mask included in the first text block detection result. Therefore, the first text block detection result is normalized in terms of image size to obtain a to-be-recognized text block image with the same size as the to-be-recognized image. The to-be-recognized text block image is then input into the text detection model to obtain the second text line detection result and the second text block detection result output by the text detection model. The to-be-recognized text line is determined according to the text line mask maps in the first text line detection result and the second text line detection result output by the text detection model.

[0103] The to-be-recognized text line is input into a character classification model, and the character classification model classifies the characters included in the to-be-recognized text line to obtain the character type with the highest confidence, and at least one character recognition model among the Chinese character recognition model, English character recognition model, and / or digital character recognition model to be called is determined according to the character type. The to-be-recognized text line is input into the character recognition model, and finally the text recognition result output by the character recognition model is obtained, completing the general and accurate recognition of the card or certificate text.

[0104] This embodiment can perform text detection on cards and certificates with arbitrary angles and layouts, select text lines of interest according to actual business, and select an identification model based on the character types of the text lines, which can improve both the recognition speed and accuracy, and provides a card and certificate text recognition solution with a wider application range and stronger versatility.

[0105] Figure 3 It is a schematic structural diagram of the card and certificate text recognition device provided by the present invention. As Figure 3 shown, the card and certificate text recognition device includes, but is not limited to, an image acquisition module 301, an image detection module 302, a text line detection module 303, and a text line recognition module 304.

[0106] The image acquisition module 301 is used to acquire the image to be recognized of the card and certificate.

[0107] The image detection module 302 is used to input the image to be recognized into a pre-trained text detection model to obtain a first text line detection result and a first text block detection result output by the text detection model.

[0108] The text line detection module 303 is used to determine the text lines to be recognized of the card and certificate based on the first text line detection result and the first text block detection result.

[0109] The text line recognition module 304 is used to perform character recognition on the text lines to be recognized to obtain the text recognition result of the card and certificate.

[0110] It should be noted that the card and certificate text recognition device provided by the present invention can execute the card and certificate text recognition method described in any of the above embodiments during specific operation, and this embodiment will not be elaborated here.

[0111] The card and certificate text recognition device provided by the present invention inputs the image of the card and certificate into a text detection model to obtain the text line detection result and text block detection result in the card and certificate image, and screens out the text lines that need to be recognized and detected from the detected text line detection result and text block detection result according to the requirements of actual projects, businesses, etc., so as to recognize the text lines corresponding to the content on the card surface to be detected, which can avoid the situation where it is impossible to directly locate the specific text content on different cards and certificates due to different layouts of the content on the card surface of the card and certificate, resulting in text recognition failure, and can perform text recognition on different types and versions of cards and certificates with arbitrary text layouts, especially newly issued versions, and provides a card and certificate text recognition solution with a wider application range and stronger versatility.

[0112] Figure 4 It is a schematic structural diagram of the electronic device provided by the present invention. As Figure 4As shown in the figure, the electronic device may include: a processor 410, a communications interface 420, a memory 430, and a communication bus 440. Among them, the processor 410, the communications interface 420, and the memory 430 communicate with each other through the communication bus 440. The processor 410 may call the logical instructions in the memory 430 to execute the card and certificate text recognition method provided in any of the above embodiments. The card and certificate text recognition method includes but is not limited to the following steps: obtaining a to-be-recognized image of the card and certificate; inputting the to-be-recognized image into a pre-trained text detection model to obtain a first text line detection result and a first text block detection result output by the text detection model; determining the to-be-recognized text line of the card and certificate based on the first text line detection result and the first text block detection result; performing character recognition on the to-be-recognized text line to obtain the text recognition result of the card and certificate.

[0113] In addition, when the logical instructions in the above-mentioned memory 430 are implemented in the form of software functional units and sold or used as independent products, they may be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, may be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs that can store program codes.

[0114] On the other hand, the present invention also provides a computer program product. The computer program product includes a computer program. The computer program can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the card and certificate text recognition method provided in any of the above embodiments. The card and certificate text recognition method includes but is not limited to the following steps: obtaining a to-be-recognized image of the card and certificate; inputting the to-be-recognized image into a pre-trained text detection model to obtain a first text line detection result and a first text block detection result output by the text detection model; determining the to-be-recognized text line of the card and certificate based on the first text line detection result and the first text block detection result; performing character recognition on the to-be-recognized text line to obtain the text recognition result of the card and certificate.

[0115] In another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the card and certificate text recognition method provided in any of the above embodiments. The card and certificate text recognition method includes but is not limited to the following steps: obtaining a to-be-recognized image of a card and certificate; inputting the to-be-recognized image into a pre-trained text detection model to obtain a first text line detection result and a first text block detection result output by the text detection model; determining the to-be-recognized text line of the card and certificate based on the first text line detection result and the first text block detection result; performing character recognition on the to-be-recognized text line to obtain a text recognition result of the card and certificate.

[0116] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without creative labor.

[0117] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the above technical solution, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0118] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments or perform equivalent replacements for some of the technical features. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of each embodiment of the present invention.

Claims

1. A card text recognition method, characterized in that: include: Obtain the image of the card to be identified; Inputting the image to be recognized into a pre-trained text detection model to obtain a first text line detection result and a first text block detection result output by the text detection model; Determining a text line to be recognized of the card based on the first text line detection result and the first text block detection result; Character recognition is performed on the text line to be recognized to obtain a text recognition result of the card.

2. The card text recognition method according to claim 1, characterized in that: The determining the to-be-recognized text line of the card based on the first text line detection result and the first text block detection result includes: Normalizing the first text block detection result to obtain a text block image to be recognized having the same size as the image to be recognized; Inputting the to-be-recognized text block image into the text detection model to obtain a second text line detection result and a second text block detection result output by the text detection model; The text line to be recognized is determined based on the first text line detection result and the second text line detection result.

3. The card text recognition method according to claim 1, characterized in that: The text detection model is obtained by training based on multiple training samples; each training sample includes a recognition image sample and its corresponding text line label and text block label; each training sample is obtained based on the following method: Acquire the recognition image sample; Labeling the text lines in the recognition image samples to obtain text line labels corresponding to the recognition image samples; When the number of the text line labels is more than two, based on the text lines whose line spacing is less than a preset pixel value, small-pitch text line labels are determined from the text line labels, and the small-pitch text line labels are merged into the text block label corresponding to the recognized image sample.

4. The card text recognition method according to claim 1, characterized in that: The performing character recognition on the text line to be recognized to obtain the text recognition result of the card includes: Inputting the text line to be recognized into a pre-trained character classification model to obtain the character type with the highest confidence output by the character classification model; Based on the character type, determining at least one character recognition model to call; In the case where there is only one character recognition model, inputting the text line to be recognized into the character recognition model to obtain the text recognition result output by the character recognition model; When there are more than two character recognition models, the text lines to be recognized are respectively input into the more than two character recognition models to obtain the text recognition results respectively output by the more than two character recognition models.

5. The card text recognition method according to claim 4, characterized in that: The types of the character recognition models include Chinese character recognition models, English character recognition models and digital character recognition models.

6. The card text recognition method according to claim 1, characterized in that: The step of obtaining the image to be identified of the card includes: Acquire the image to be corrected of the card; Inputting the image to be corrected into a pre-trained text direction recognition model to obtain the text angle with the highest confidence output by the text direction recognition model; Based on the text angle, rotation correction is performed on the image to be corrected to obtain the image to be recognized.

7. The card text recognition method according to claim 6, characterized in that: The step of obtaining the card image to be corrected includes: Obtaining an original photographed image of the card; Inputting the original captured image into a pre-trained target detection model to obtain a rectangular detection frame output by the target detection model; Using the rectangular detection frame to crop the original captured image to obtain a rough card image; Inputting the rough card image into a pre-trained edge detection model to obtain vertex coordinate values ​​output by the edge detection model; Based on the vertex coordinate values, perspective transformation and cropping are performed on the rough card image to obtain an image to be corrected.

8. A card identification device, characterized in that: include: An image acquisition module, used to acquire the image to be identified of the card; An image detection module, used for inputting the image to be recognized into a pre-trained text detection model to obtain a first text line detection result and a first text block detection result output by the text detection model; A text line detection module, used for determining a text line to be identified of the card based on the first text line detection result and the first text block detection result; The text line recognition module is used to perform character recognition on the text line to be recognized to obtain the text recognition result of the card.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the computer program, the card text recognition method according to any one of claims 1 to 7 is implemented.

10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the card text recognition method according to any one of claims 1 to 7 is implemented.