A card and certificate recognition method, device, terminal and storage medium

Through the position training model and the detection training model, the problem of low recognition reliability caused by the infill or missing card certificates is solved, and the accurate identification of card certificate characteristic information is achieved.

CN110472602BActive Publication Date: 2025-07-29TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN201910770808.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2019-08-20
Publication Date
2025-07-29
Estimated Expiration
2040-06-06

AI Technical Summary

Technical Problem

In the prior art, card certificate identification is susceptible to the problem that card certificate does not fill the camera area or is missing, resulting in low recognition reliability.

Method used

The position training model and detection training model are used to identify the feature information of the card certificate by obtaining the area and feature information of the card certificate in the image.

Benefits of technology

Even if the image contains non-card certificate areas or the card certificate is incomplete, the characteristic information of the card certificate can be accurately identified, improving the reliability of card certificate identification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN110472602B_ABST
    Figure CN110472602B_ABST
Patent Text Reader

Abstract

The present application discloses a card identification method, apparatus, terminal and storage medium. The method includes: obtaining a target image; using a position training model to obtain an image area where the card is located in the target image; wherein, the position training model is trained using at least two image samples with regional position labels; obtaining at least one target area in the image area where the card is located, and the target area is an area containing feature information; identifying each of the target areas to obtain an identification result of the card, and the identification result includes at least one piece of feature information. It can be seen that in the present application, even if the image contains areas other than the card or the card in the image is incomplete, the position training model trained using image samples with regional position labels can be used to obtain the area where the card is located in the image for identifying the feature information of the card, thereby improving the reliability of card identification.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image recognition technology, and in particular, to a card and certificate recognition method, device, terminal, and storage medium. Background Art

[0002] With the development of technology, there are more and more scenarios for card and certificate recognition. For example, when entering identity information, an ID card can be scanned to identify identity information such as name and ID number; another example is that when entering financial information, a credit card or debit card can be scanned to identify information such as the opening bank and account number, and so on.

[0003] Currently, when recognizing a card and certificate, it is usually necessary for the user to align the camera with the card and certificate according to the prompt information on the shooting interface and try to fill the shooting area with the card and certificate as much as possible in order to identify the card and certificate information.

[0004] It can be seen that in the above recognition scheme, there may be a situation where the user cannot fill the shooting area with the card and certificate or there is a lack of part of the card and certificate in the shooting area. At this time, the card and certificate cannot be recognized, resulting in relatively low recognition reliability of the card and certificate. Summary of the Invention

[0005] In view of this, this application provides a card and certificate recognition method, device, terminal, and storage medium to improve the recognition reliability of the card and certificate.

[0006] To achieve the above object, on the one hand, this application provides a card and certificate recognition method, including:

[0007] Obtain a target image;

[0008] Use a position training model to obtain the image area where the card and certificate is located in the target image; wherein, the position training model is trained using at least two image samples with regional position labels;

[0009] Obtain at least one target area in the image area where the card and certificate is located, and the target area is an area containing feature information;

[0010] Recognize each target area to obtain the recognition result of the card and certificate, and the recognition result includes at least one piece of feature information.

[0011] In a possible implementation manner, the step of using a position training model to obtain the image area where the card and certificate is located in the target image includes:

[0012] Input the target image into the position training model to obtain a position vector output by the position training model, and the position vector includes the regional vertex coordinate information of the card and certificate in the target image;

[0013] Intercept the image area where the target card is located in the target image based on the regional vertex coordinate information.

[0014] Optionally, the regional position label of the image sample includes a regional vertex coordinate label, and the position training model outputs the position vector based on the regional vertex coordinate label;

[0015] Wherein, the regional vertex coordinate label is the coordinate value after normalizing the regional vertex coordinate information in the image sample.

[0016] Optionally, it further includes: optimizing the model parameters of the position training model by using a preset loss function.

[0017] In a possible implementation manner, obtaining at least one target area in the image area where the card is located includes:

[0018] Input the image area where the card is located into a detection training model to obtain at least one image mask output by the detection training model, and the pixel value of the pixel point on the image mask indicates the probability value that the pixel point belongs to a feature information pixel point; wherein, the detection training model is trained by using at least two image samples with true value labels;

[0019] Based on the pixel values of the pixel points on each image mask, obtain the target area in each image mask respectively.

[0020] Optionally, the obtaining the target area in each image mask based on the pixel values of the pixel points on each image mask includes:

[0021] Take the connected domain composed of adjacent pixel points with at least approximately the same pixel value in each image mask as the target area in the image mask.

[0022] Optionally, the regional position label of the image sample of the position training model includes a position attitude label, and the position training model outputs a classification vector based on the position attitude label, and the classification vector includes the position attitude information of the card in the target image;

[0023] Wherein, the detection training model is used to output at least one image mask corresponding to the position attitude information.

[0024] In a possible implementation manner, the recognizing each target area to obtain the recognition result of the card includes:

[0025] Use a preset recognition model to perform feature recognition on each target area to obtain the recognition result of the card, and the recognition result includes the feature information in each target area.

[0026] On the other hand, the present application also provides a card and certificate recognition device, including:

[0027] An image acquisition unit, configured to acquire a target image;

[0028] A model operation unit, configured to use a position training model to obtain an image area where the card and certificate is located in the target image; wherein, the position training model is obtained by training with at least two image samples with regional position labels;

[0029] A region acquisition unit, configured to acquire at least one target region in the image area where the card and certificate is located, and the target region is a region containing feature information;

[0030] A region recognition unit, configured to recognize each target region to obtain a recognition result of the card and certificate, and the recognition result includes at least one piece of feature information.

[0031] On the other hand, the present application also provides a terminal, including:

[0032] A processor and a memory;

[0033] Wherein, the processor is configured to execute a program stored in the memory;

[0034] The memory is configured to store a program, and the program is at least used for:

[0035] Acquire a target image;

[0036] Use a position training model to obtain an image area where the card and certificate is located in the target image; wherein, the position training model is obtained by training with at least two image samples with regional position labels;

[0037] Acquire at least one target region in the image area where the card and certificate is located, and the target region is a region containing feature information;

[0038] Recognize each target region to obtain a recognition result of the card and certificate, and the recognition result includes at least one piece of feature information.

[0039] On the other hand, the present application also provides a storage medium, in which computer-executable instructions are stored. When the computer-executable instructions are loaded and executed by a processor, the card and certificate recognition method described in any one of the above is implemented.

[0040] As can be seen from the above solution, for a card identification method, device, terminal, and storage medium provided by this application, after obtaining a target image, a position training model is used to obtain the image area where the card is located in the target image. Then, after obtaining a target area containing feature information in the image area where the card is located, by identifying the target area, an identification result containing at least one feature information of the card is obtained. It can be seen that in this application, even if the image contains areas other than the card or the card in the image is incomplete, the position training model trained with image samples with regional position labels can be used to obtain the area where the card is located in the image for identifying the feature information of the card. Therefore, in this application, regardless of whether the card in the image is full or complete, the identification of the feature information of the card can be achieved, thereby improving the reliability of card identification. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] To more clearly illustrate the technical solutions of the embodiments of this application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of this application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0042] Figure 1 FIG. shows a schematic composition framework diagram of a card identification system according to an embodiment of this application;

[0043] Figures 2 - 4 FIGS. respectively show other architecture schematic diagrams of the card identification system according to the embodiments of this application;

[0044] Figure 5 FIG. shows a schematic diagram of the hardware composition structure of a terminal for implementing card identification according to an embodiment of this application;

[0045] Figure 6 FIG. shows a schematic flowchart of a card identification method according to an embodiment of this application;

[0046] Figures 7 - 11 FIGS. respectively show schematic diagrams related to obtaining the image area where the card is located according to the embodiments of this application;

[0047] Figure 12 FIG. shows a schematic diagram related to obtaining the target area where the feature information is located according to the embodiments of this application;

[0048] Figures 13 - 16 FIGS. respectively show schematic diagrams of the process of a card identification method for implementing identity card identification according to the embodiments of this application;

[0049] Figure 17 FIG. shows a schematic composition structure diagram of an embodiment of a card identification device according to an embodiment of this application. Specific implementation manners

[0050] The solution of the present application can be applicable to the identification of cards and certificates with a fixed template, and is mainly applied in fields such as security inspection, finance, transportation, and smart government affairs. For example, the identification of cards and certificates such as identity cards, driving licenses, and business licenses is performed.

[0051] The inventors of the present application have found through research that: currently, when identifying cards and certificates, it is often necessary for staff to fill the entire imaging area of the camera with the card and certificate and ensure that the card and certificate are complete in the imaging area. Only at this time can effective identification be carried out. For example, when scanning information such as the name, date of birth, and ID number on an identity card through a mobile phone camera, the user needs to align the camera with the identity card and, under the guidance of the mobile phone interface, fill the mobile phone imaging area with the identity card, and the identity card needs to be complete in the imaging area for the mobile phone to identify the information on the identity card; another example is that when entering the account information on a bank card through the bank system, the staff needs to place the bank card in the fixed area under the fixed camera of the bank system so as to ensure that the bank card fills the acquisition area of the camera and is complete. Only at this time can the bank system effectively identify the account information on the bank card. It can be seen that currently, when identifying cards and certificates, there may be a situation where card and certificate identification cannot be carried out due to the card and certificate not being able to fill the imaging area or the card and certificate being missing, which affects the reliability of card and certificate identification.

[0052] Therefore, the inventors of the present application have further studied and found that even if the card and certificate cannot fill the imaging area or is missing in the imaging area, the card and certificate itself has unique characteristics, such as presenting a rectangular, square or other shaped area in the imaging area, and various feature information such as fields or images in the card and certificate are in fixed positions of the card and certificate. Therefore, in order to improve the reliability of card and certificate identification, a position training model can be trained using image samples with regional position tags, and such a position training model is used to process the acquired images that may contain non-card and certificate areas or have missing card and certificates. For example, first obtain the image area where the card and certificate is located in the image, remove the non-card and certificate areas, and then further identify the target area containing feature information in the image area where the card and certificate is located, and then identify the feature information in the target area, so as to obtain the identification result of the card and certificate, thereby avoiding the situation where card and certificate identification cannot be carried out due to the image containing non-card and certificate areas or missing card and certificates, and thus improving the reliability of card and certificate identification.

[0053] For the convenience of understanding, the system applicable to the solution of the present application is introduced first in this article. Refer to Figure 1 , which shows a schematic diagram of a composition architecture of a card and certificate identification system of the present application.

[0054] As shown by Figure 1As can be seen, the system may include: a server 10 and a terminal 20, and the server 10 and the terminal 20 are communicatively connected via a network.

[0055] Among them, the server 10 can be a front-end server or a back-end server, etc., and the terminal 20 can be a client such as a mobile phone, a pad, a computer, a handheld scanner, etc. At this time, the user can collect an image through the terminal 20 and send the image to the server 10, and the server 10 performs image processing to achieve card identification.

[0056] It should be noted that in another implementation, the card identification system in this application may not include the server 10, only the terminal 20, and the function of the server 10 for image processing is integrated into the terminal 20. After the terminal 20 collects an image, it processes the image to achieve card identification; or, in another implementation, the card identification system in this application may not include the terminal 20, only the server 10, and the function of the terminal 20 for image collection is integrated into the server 10. After the server 10 collects an image, it processes the image to achieve card identification.

[0057] For example, after the bank counter terminal collects an image of a bank card, it transmits the image to the bank back-end server, and the bank back-end server performs image processing to identify the characteristic information on the bank card, such as Figure 2 ; after the mobile phone terminal collects an image of an ID card through the camera, it transmits the image to the central processing unit of the mobile phone, and the central processing unit of the mobile phone processes the image to obtain the characteristic information on the ID card, such as Figure 3 as shown in; after the government affairs server collects an image of a business license through the camera, it transmits the image to the processor of the government affairs server, and the processor of the government affairs server processes the image to obtain the characteristic information on the business license, such as Figure 4 shown.

[0058] Among them, a device capable of image collection, such as a camera or a video recorder, etc., can be configured on the terminal 20 or the server 10 to collect an image of the card to be identified. The collected image may include the entire area of the card, or may include a partial area of the card, or may include other non-card areas.

[0059] Among them, in order to implement the corresponding image processing function on the terminal or the server, a program for implementing the corresponding function needs to be stored in the memory of the terminal or the server. To facilitate understanding of the hardware composition of the terminal or the server, the terminal will be introduced as an example below. As Figure 5 shown, it is a schematic diagram of a composition structure of the terminal of this application. The terminal 20 in this embodiment may include: a processor 201, a memory 202, a communication interface 203, an input unit 204, a display 205, and a communication bus 206.

[0060] Among them, the processor 201, the memory 202, the communication interface 203, the input unit 204, and the display 205 all complete mutual communication through the communication bus 206.

[0061] In this embodiment, the processor 201 can be a central processing unit (CPU), an application specific integrated circuit, a digital signal processor, a field programmable gate array, or other programmable logic devices, etc.

[0062] The processor 201 can call the program stored in the memory 202. Specifically, the processor 201 can execute the operations performed by the terminal in the following embodiments of the card identification method.

[0063] The memory 202 is used to store one or more programs. The program can include program codes, and the program codes include computer operation instructions. In the embodiments of the present application, the memory stores at least programs for implementing the following functions:

[0064] Obtain a target image;

[0065] Using the position training model, obtain the image area where the card is located in the target image; wherein, the position training model is trained using at least two image samples with regional position labels;

[0066] Obtain at least one target area in the image area where the card is located, and the target area is an area containing feature information;

[0067] Identify each of the target areas to obtain the identification result of the card, and the identification result includes at least one piece of feature information.

[0068] In a possible implementation manner, the memory 202 can include a program storage area and a data storage area. Among them, the program storage area can store an operating system and application programs required for at least one function (such as image display, etc.); the data storage area can store data created during the use of the computer, such as image data, feature information, etc.

[0069] In addition, the memory 202 can include a high-speed random access memory, and can also include a non-volatile memory, such as at least one disk storage device or other volatile solid-state storage devices.

[0070] The communication interface 203 can be the interface of a communication module, such as the interface of a GSM module.

[0071] Of course, Figure 5The structure of the terminal shown does not limit the terminal in the embodiments of the present application. In practical applications, the terminal may include more or fewer components than those Figure 5 shown, or combine some components. It can be understood that the hardware composition of the server can refer to Figure 5 the hardware composition of the terminal in

[0072] Combining the above commonalities, referring to Figure 6 , which shows a schematic flowchart of an embodiment of a card and certificate recognition method of the present application. The method in this embodiment may include:

[0073] S601: Obtain a target image.

[0074] Specifically, in this embodiment, a card and certificate can be imaged by an image acquisition device such as a mobile phone camera to obtain a target image.

[0075] Among them, the target image may include an image area where the card and certificate is located, and may also include an image area that is not a card and certificate, as Figure 7 shown; and in the card and certificate image area in the target image, the card and certificate may be complete or incomplete, as Figure 8 shown.

[0076] S602: Use a position training model to obtain the image area where the card and certificate is located in the target image.

[0077] Among them, the position training model is trained using at least two image samples with regional position labels. Correspondingly, in this embodiment, the target image is used as the input of the position training model, and the pre-trained position training model is run, so as to be able to obtain the image area where the card and certificate is located in the target image, as Figure 9 shown, ignoring the image area in the target image that has nothing to do with the card and certificate, and intercepting the image area where the card and certificate is located from the target image.

[0078] It should be noted that the position training model can be based on a convolutional neural network model. Specifically, the convolutional neural network may include but is not limited to neural networks such as VGG (Visual Geometry Group Network), Googlenet, resnet, densenet, mobilenet or shufflenet, etc. The position training model takes the convolutional neural network as the backbone network and trains at least two image samples with regional position labels to form a model that can output a position vector.

[0079] Specifically, in S602, when obtaining the image area where the card and certificate is located, it can be implemented in the following way:

[0080] First, input the target image into the position training model to obtain the position vector output by the position training model. The position vector includes the region vertex coordinate information of the card in the target image, such as the vertex coordinate information of the four corners when the card is complete, such as Figure 10 shown in; or when one corner of the card is missing, the vertex coordinate information of the missing corner and other corners predicted by the position training model, or the vertex coordinate information of the four corners appearing in the target area, etc., such as Figure 11 shown in.

[0081] It should be noted that the region vertex coordinate information in the position vector forms a vector in the form of horizontal and vertical coordinates. Taking the case where the card has 4 vertices in the target image as an example, the position vector contains 8 vector components, which are the horizontal and vertical coordinates of the 4 vertices, such as Figure 10 the vertex coordinates in: (x1, y1), (x2, y2), (x3, y3), (x4, y4), forming the position vector (x1, y1, x2, y2, x3, y3, x4, y4).

[0082] After that, based on the region vertex coordinate information in the position vector, intercept the image region where the target card is located in the target image.

[0083] Taking the Figure 10 image in as an example, in this embodiment, the region vertex coordinate information output by the position training model is used to intercept the image in the target image. For example, connect the adjacent points of the 4 coordinate points corresponding to the region vertex coordinate information to form a regular closed image region, and this region is the image region where the card is located;

[0084] Or, taking the Figure 11 image in as an example, in this embodiment, the region vertex coordinate information output by the position training model is used to intercept the image in the target image. For example, for the 4 coordinate points corresponding to the region vertex coordinate information (one or more of which are not within the range of the target image or one or more are at the edge of the target image), form a regular closed image region, and intercept the overlapping part between the image region formed by the 4 coordinate points and the target image in the target image, which is the image region where the card is located.

[0085] It should be noted that the image samples used to train the position training model have known region position labels, and the region position labels include region vertex coordinate labels. Correspondingly, after training, the position training model can accurately identify the region vertex coordinate information of the image region where the card is located in the target image based on the training of these region vertex coordinate labels, and form a position vector for output.

[0086] During the training process of the position training model, first, the regional vertex coordinate information of the image sample is normalized to form the regional vertex coordinate label, which is then used for model training. For example, the regional vertex coordinate information in the image sample is represented by the position of pixel points. In an image sample of 200*400, the vertex coordinates of the upper left corner of the card certificate are (100, 300). Before using the image sample for model training, the vertex coordinates of the upper left corner of the card certificate are normalized first, that is, the ratios of 100 / 200 and 300 / 400 are used as the horizontal and vertical coordinate labels of the upper left corner in the regional position label.

[0087] Correspondingly, after the position training model is trained, the regional vertex coordinate information in the position vector output by inputting the target image into the position training model is also the coordinate value normalized relative to the target image. If you want to represent the regional vertex coordinate information in the position vector in pixel points, you can multiply the horizontal and vertical coordinate values of the regional vertex coordinate information in the position vector by the total number of corresponding horizontal and vertical pixel points in the target image.

[0088] In one implementation, during the training process of the position training model, in this embodiment, in order to improve the output accuracy of the position training model in subsequent use, the loss function can be used to optimize the model parameters in the position training model. For example, the regression loss function is used to optimize the position training model during training, so that the finally trained position training model can output a position vector closer to the real data, making the regional vertex coordinate information in the position vector more accurate.

[0089] S603: Obtain at least one target region in the image region where the card certificate is located.

[0090] Among them, the target region is the region containing feature information. That is to say, in this embodiment, further region screening is performed in the intercepted image region where the card certificate is located, and the target regions containing feature information in the image region where the card certificate is located are screened out, and other regions not containing feature information are ignored.

[0091] It should be noted that in the image region where the card certificate is located, there may be one or more feature information of the card certificate. For example, an identity card contains at least information such as name, date of birth, and identity card number. Therefore, in this embodiment, by performing region screening on the image region where the card certificate is located, one or more target regions can be screened out. And the types of feature information contained in different target regions may be different, and the layout positions of the corresponding different feature information on the card certificate may be different. For example, the name and identity card number on the identity card are in the first row and the last row of the card certificate respectively. Therefore, the target regions screened out in this embodiment may be in different specific positions in the image region where the card certificate is located.

[0092] In one implementation, when obtaining the target area in this embodiment, S603 can be specifically implemented in the following manner:

[0093] First, input the image area where the card is located into the detection training model to obtain at least one image mask output by the detection training model. Among them, the pixel value of the pixel point on the image mask indicates the probability value that the pixel point belongs to the feature information pixel point.

[0094] Specifically, the detection training model is trained using at least two image samples with ground truth labels. For example, there are corresponding annotation ground truths in the image samples, and each field in the image sample has a ground truth, and the so-called ground truth is the exact value obtained through the annotation information. For example, for each field in the image sample, the ground truth is marked: the position of the field (x1, y1, x2, y2, x3, y3, x4, y4) and the type of the field (such as name, gender, etc.). Thus, when training the detection training model, each field can be separated according to the type of the field, and each field in the output result of the model training corresponds to a mask, so the model will output at least one image mask. When training the detection training model, the pixel values of the pixel points in the mask in the training output result can be set according to the position of the field. For example, during training, on the mask corresponding to the name field, the pixel points within the name box are set to 1, and the pixel points at other positions except the name box are set to 0.

[0095] Thus, in this embodiment, at least one image mask can be output by the detection training model first for the image area where the card is located, and each image mask is for one type of feature information. Then, the pixel points in each image mask are subjected to multiple binary classification processes, and the pixel values of the pixel points on the image mask output by the detection training model are set according to the training results of the ground truth labels in the image samples.

[0096] For example, based on the training results of the ground truth labels in the image samples, the detection training model determines whether each pixel point in each image mask is name feature information, whether it is the feature information of the date of birth, whether it is the feature information of the ID number, whether it is the feature information of the head portrait, and so on. Correspondingly, finally, the detection training model outputs an image mask with at least one pixel point having a pixel value. The pixel value of each pixel point on each image mask indicates the probability that the pixel point belongs to a pixel point of a certain specific feature information. For example: on the mask of the name field, the pixel values of some pixel points are 1, indicating that these pixel points correspond to the area where the name box is located, and the pixel values of other pixel points are 0, indicating that these pixel points correspond to the area outside the name box; for another example, on the mask of the ID number field, the pixel values of some pixel points are 1, indicating that these pixel points correspond to the area where the ID number box is located, and the pixel values of other pixel points are 0, indicating that these pixel points correspond to the area outside the ID number box.

[0097] It should be noted that the detection training model can be a model trained based on a fully convolutional network, such as the fully convolutional network FCN (Fully Convolutional Networks for Semantic Segmentation), etc. The detection training model uses the fully convolutional network as the backbone network and trains on image samples with at least two ground truth labels to form a model that can output an image mask with pixel points set with pixel values.

[0098] After that, based on the pixel values of the pixel points on each image mask, the image regions in each image mask are obtained respectively. For example, in this embodiment, the pixel points that are adjacent and have the same or approximately the same pixel value in each image mask are connected, and then the connected domain composed of adjacent pixel points with at least approximately the same pixel value is used as the target region in the image mask. Thus, the target regions containing feature information in each image mask are obtained, as Figure 12 shown.

[0099] In a specific implementation, the regional position label of the image sample of the position training model may further include a position and pose label, such as the front and back poses of a card or certificate. Correspondingly, after being trained based on the position and pose label, the position training model can also output a classification vector for the input target image. The classification vector includes the position and pose information of the card or certificate in the target image, such as the pose information that the front side of the card or certificate faces up or the back side of the card or certificate faces up. The position and pose information in the classification vector determines the type of the image mask output by the subsequent detection training model. For example, when the front side of the card or certificate faces up, the image mask output by the detection training model is a mask for the card or certificate with the front side facing up, and the feature information in each image mask is for the feature information on the front side of the card or certificate; when the back side of the card or certificate faces up, the image mask output by the detection training model is a mask for the card or certificate with the back side facing up, and at this time, the feature information in each image mask is for the feature information on the back side of the card or certificate. That is to say, the detection training model processes the image area where the card or certificate is located in the input target image based on the position and pose information of the card or certificate in the target image, and obtains at least one image mask containing the feature information on the front or back side of the card or certificate, that is: each image mask contains specific feature information, and whether the feature information belongs to the front or back side of the card or certificate is determined by the position and pose information of the card or certificate in the target image.

[0100] In addition, during the training process of the position training model, in this embodiment, in order to improve the output accuracy of the position training model during subsequent use, the model parameters in the position training model can be optimized by using a loss function. For example, the classification loss function is used to optimize the position training model during training, so that the finally trained position training model can output a classification vector closer to the real data, and the position and pose information in the classification vector is more accurate.

[0101] It should be noted that the position vector and the classification vector in this embodiment can form a vector, that is to say, the position training model can only output one vector, and this vector contains the components of the position vector and the components of the classification vector.

[0102] Optionally, during the training process of the detection training model, in this embodiment, in order to improve the output accuracy of the detection training model during subsequent use, the model parameters in the detection training model can be optimized by using a loss function. For example, the classification loss function is used to optimize the detection training model during training, so that the finally trained detection training model can output a sample true value closer to the real data, and the pixel values of the mask pixel points are more accurate.

[0103] S604: Identify each target area to obtain the recognition result of the card or certificate.

[0104] Among them, each target area contains its unique feature information. Therefore, the recognition result of the card certificate includes at least one feature information on the card certificate, and these feature information correspond to different feature types, such as the type of name, the type of date of birth, the type of ID number, and so on.

[0105] In one implementation, when S604 recognizes the target area, it can be specifically implemented in the following way:

[0106] Using a preset recognition model, perform feature recognition on each target area to obtain the recognition result of the card certificate, and the recognition result of the card certificate contains the feature information in each target area, such as name, date of birth, and ID number.

[0107] Among them, the recognition model can be a model based on a preset recognition algorithm, which can be but not limited to tesseract or the end-to-end text recognition algorithm CRNN (Convolutional Recurrent Neural Network), etc.

[0108] It should be noted that the feature information in the recognition result can include a string composed of any one or more of Chinese characters, English characters, and numerical characters, can also include other special symbols, such as * or #, etc., and can also include images, such as special marking information such as a portrait.

[0109] From the above solution, it can be seen that for a card certificate recognition method provided in this embodiment, after obtaining the target image, use the position training model to obtain the image area where the card certificate is located in the target image, and then after obtaining the target area containing the feature information in the image area where the card certificate is located, by recognizing the target area, the recognition result containing at least one feature information of the card certificate is obtained. It can be seen that in this embodiment, even if the image contains a non-card certificate area or the card certificate in the image is incomplete, the position training model trained with the image sample with area position labels can be used to obtain the area where the card certificate is located in the image for the recognition of the feature information of the card certificate. Therefore, in this application, regardless of whether the card certificate in the image is full or complete, the recognition of the feature information of the card certificate can be realized, thereby improving the reliability of card certificate recognition.

[0110] For the convenience of understanding, the following combines Figure 13 The following logic architecture diagram of the terminal for recognizing the card certificate to introduce the example of this solution in actual application:

[0111] First of all, the process of card certificate recognition in this solution mainly includes three steps: card certificate detection, field detection, and field recognition. For the above three steps, three neural network-based models are pre-trained in this embodiment to complete the corresponding step tasks:

[0112] 1) Card and certificate detection:

[0113] As Figure 14 shown in, in this embodiment, taking the ID card as an example, the image containing the ID card collected is input into the network model (such as the position training model in the previous text). Correspondingly, the result output by the network model is a 9-dimensional vector. Among them, the first 8 dimensions represent the relative position information of the card, that is, the relative positions of the 4 vertices of the card in the image, represented by the normalized coordinate values x1, y1, x2, y2, x3, y3, x4, y4.

[0114] In addition, there is another dimension in the vector to represent the front and back postures of the card (for example, 0 is the front side and 1 is the back side).

[0115] Among them, the backbone network of the network model in the card and certificate detection stage is a convolutional neural network (CNN, Convolutional Neural Networks), including but not limited to VGG, Googlenet, resnet, densenet, mobilenet, shufflenet, etc.

[0116] It should be noted that when training the network model in the card and certificate detection stage, the vertex position coordinates of the card in the image used as the training sample can be first normalized (normalized) to between 0 and 1 by dividing the absolute pixel point coordinates by the total number of pixel points of the image size. And, use a regression loss function such as smooth-L1 to optimize the model parameters regarding the output position vector in the model, and use a classification loss function such as BCE Loss to optimize the model parameters regarding the output classification vector in the model, so as to improve the accuracy of the model output result.

[0117] 2) Field detection:

[0118] As Figure 15 shown in, in this embodiment, taking the ID card as an example, the image where the ID card is located (not containing or only containing a little non-ID card content) is input into the network model (such as the detection training model in the previous text). The network model outputs multiple masks with the same size as the input image, and these masks respectively correspond to multiple fields of the ID card, such as: name, gender, date of birth, address, ID number, portrait, issuing authority, validity period.

[0119] Among them, the backbone network of the network model in the field detection stage is a fully convolutional network such as FCN. During the training process of the model, in this embodiment, corresponding masks are generated based on the known positions of each field in the training sample, and each pixel point on the mask has a pixel value.

[0120] In addition, in this embodiment, a classification loss function can be used to optimize the network model during the training in the field detection stage. The classification loss function can include but is not limited to BCE loss, Dice loss, etc.

[0121] After that, in this embodiment, after obtaining the predicted mask as described above, based on the magnitude of the pixel values of the pixel points in the mask, the detection frame (such as the target area in the foregoing) where the field is located on each mask is obtained. Specifically, in this embodiment, for each mask, through the connected component algorithm, the connected components on each mask are obtained, and then the minimum bounding rectangle is taken for each connected component to obtain the detection frames of each field, as Figure 16 shown in the detection frames of each field in the ID card picture. It can be seen that in this embodiment, each field and its valid information can be directly obtained through the network model in the field detection stage, and there is no need to filter out irrelevant fields through post-processing again after recognition.

[0122] 3) Field recognition:

[0123] Among them, in this embodiment, based on the field position information represented by the detection frame obtained by field detection, any recognition model can be used to recognize the field information in the detection frame, such as including but not limited to tesseract or CRNN, etc.

[0124] It should be noted that the applicable network models, loss functions, etc. in this embodiment are not limited to the models and loss functions used in this article, and all modules or functions with the same functions can be used as alternative solutions.

[0125] In summary, the solution in this embodiment can effectively solve the problem of identifying fixed-template cards and certificates photographed at any direction and any angle, and the process is simple, and the performance and accuracy are also more competitive compared with other methods.

[0126] On the other hand, the present application also provides a card and certificate recognition device, as Figure 17 shown in, which shows a schematic diagram of the composition of an embodiment of a card and certificate recognition device of the present application. The device in this embodiment can be applied to a terminal, and the device can include:

[0127] An image acquisition unit 1701, configured to acquire a target image;

[0128] A model operation unit 1702, configured to use a position training model to acquire an image area where the card and certificate are located in the target image; wherein, the position training model is obtained by training with at least two image samples with region position labels;

[0129] An area acquisition unit 1703, configured to acquire at least one target area in the image area where the card is located, where the target area is an area containing feature information;

[0130] An area recognition unit 1704, configured to recognize each of the target areas to obtain an identification result of the card, where the identification result includes at least one piece of feature information.

[0131] Optionally, the model operation unit 1702 is specifically configured to:

[0132] Input the target image into a position training model to obtain a position vector output by the position training model, where the position vector includes regional vertex coordinate information of the card in the target image;

[0133] Based on the regional vertex coordinate information, intercept the image area where the target card is located in the target image.

[0134] Wherein, the regional position label of the image sample includes a regional vertex coordinate label, and the position training model outputs the position vector based on the regional vertex coordinate label; correspondingly, the regional vertex coordinate label is a coordinate value obtained by normalizing the regional vertex coordinate information in the image sample.

[0135] Optionally, the area acquisition unit 1703 is specifically configured to:

[0136] Input the image area where the card is located into a detection training model to obtain at least one image mask output by the detection training model, where the pixel value of a pixel point on the image mask indicates the probability value that the pixel point belongs to a feature information pixel point; wherein, the detection training model is trained using at least two image samples with ground truth labels;

[0137] Based on the pixel values of the pixel points on each image mask, respectively obtain the target areas in each image mask. For example, a connected domain formed by adjacent pixel points with at least approximately the same pixel value in each image mask is used as the target area in the image mask.

[0138] Wherein, the regional position label of the image sample of the position training model includes a position and pose label, and the position training model outputs a classification vector based on the position and pose label, where the classification vector includes the position and pose information of the card in the target image; correspondingly, the detection training model is used to output at least one image mask corresponding to the position and pose information.

[0139] Optionally, the area recognition unit 1704 is specifically configured to:

[0140] Using a preset recognition model, perform feature recognition on each of the target regions to obtain the recognition result of the card certificate, where the recognition result includes the feature information in each of the target regions.

[0141] On the other hand, an embodiment of the present application further provides a storage medium, in which computer-executable instructions are stored. When the computer-executable instructions are loaded and executed by a processor, the card certificate recognition method executed by the terminal in any of the above embodiments is implemented.

[0142] It should be noted that the various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. The same or similar parts among the various embodiments can be referred to each other. For device-type embodiments, since they are basically similar to method embodiments, they are described relatively simply. The relevant parts can be referred to the partial description of the method embodiments.

[0143] Finally, it should also be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article or device including the element.

[0144] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

[0145] The above are only the preferred embodiments of the present invention. It should be pointed out that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.

Claims

1. A card and certificate recognition method, characterized in that, Including: Obtain a target image; The target image is an image of a certificate with a fixed template taken at any direction and any angle; Input the target image into a position training model to obtain a position vector output by the position training model, where the position vector includes regional vertex coordinate information of the certificate in the target image; the regional vertex coordinate information of the image region where the certificate is located includes vertex coordinate information of the missing angle and other angles predicted by the position training model when part of the angle of the certificate is missing; Based on the regional vertex coordinate information, intercept the image region where the certificate is located in the target image; wherein, the position training model is trained using at least two image samples with regional position labels; the regional position labels of the image samples include regional vertex coordinate labels and position attitude labels, and the regional vertex coordinate labels are coordinate values after normalizing the regional vertex coordinate information in the image samples; use the position training model to output a classification vector based on the position attitude labels, and the classification vector includes position attitude information of the certificate in the target image; Input the image region where the certificate is located into a detection training model to obtain at least one image mask output by the detection training model corresponding to the position attitude information, where one image mask corresponds to one type of feature information, and the pixel value of a pixel point on the image mask indicates the probability value that the pixel point belongs to a feature information pixel point; wherein, the detection training model is trained using at least two image samples with ground truth labels, and the ground truth labeled for each field in the image samples includes: the coordinate position of the field and the type of the field; Based on the pixel values of the pixel points on each image mask, obtain the target region in each image mask respectively; the target region is a region containing feature information; Identify each target region to obtain the recognition result of the certificate, and the recognition result includes at least one type of feature information.

2. The method according to claim 1, characterized in that The obtaining the target region in each image mask respectively based on the pixel values of the pixel points on each image mask includes: Take the connected domain composed of adjacent pixel points with at least approximately the same pixel values in each image mask as the target region in the image mask.

3. The method according to claim 1, wherein The identifying each target region to obtain the recognition result of the certificate includes: Use a preset recognition model to perform feature recognition on each target region to obtain the recognition result of the certificate, and the recognition result includes the feature information in each target region.

4. A card and certificate recognition device, characterized in that, Including: An image acquisition unit for obtaining a target image; The target image is an image of a certificate with a fixed template taken at any direction and any angle; A model running unit for inputting the target image into a position training model to obtain a position vector output by the position training model, where the position vector includes regional vertex coordinate information of the card in the target image; the regional vertex coordinate information of the image region where the card is located includes vertex coordinate information of the missing corner and other corners predicted by the position training model when the card lacks some corners; based on the regional vertex coordinate information, intercept the image region where the card is located in the target image; wherein, the position training model is trained by using at least two image samples with regional position labels, and the regional position labels include regional vertex coordinate labels and position attitude labels, and the regional vertex coordinate labels are coordinate values after normalizing the regional vertex coordinate information in the image samples; using the position training model to output a classification vector based on the position attitude labels, and the classification vector includes position attitude information of the card in the target image; A region obtaining unit for inputting the image region where the card is located into a detection training model to obtain at least one image mask output by the detection training model corresponding to the position attitude information, and one image mask corresponds to one kind of feature information, and the pixel value of a pixel point on the image mask indicates the probability value that the pixel point belongs to a feature information pixel point; wherein, the detection training model is trained by using at least two image samples with ground truth labels, and the ground truth labeled for each field in the image samples includes: the coordinate position of the field and the type of the field; Based on the pixel values of the pixel points on each image mask, respectively obtain the target regions in each image mask; the target region is a region containing feature information; A region recognition unit for recognizing each target region to obtain an identification result of the card, and the identification result includes at least one kind of feature information.

5. The card and certificate recognition device according to claim 4, characterized in that, The region obtaining unit is used for: Taking the connected domain composed of adjacent pixel points with at least approximately the same pixel values in each image mask as the target region in the image mask.

6. The card and certificate recognition device according to claim 4, characterized in that, The region recognition unit is specifically used for: Using a preset recognition model to perform feature recognition on each target region to obtain an identification result of the card, and the identification result includes the feature information in each target region.

7. A terminal, characterized in that, Including: A processor and a memory; Wherein, the processor is used to execute the program stored in the memory; The memory is used to store the program, and the program is at least used for: Obtaining a target image; the target image is an image of a card with a fixed template taken at any direction and any angle; Inputting the target image into a position training model to obtain a position vector output by the position training model, where the position vector includes regional vertex coordinate information of the card in the target image; the regional vertex coordinate information of the image region where the card is located includes vertex coordinate information of the missing corner and other corners predicted by the position training model when the card lacks some corners; Intercept the image region where the card is located in the target image based on the regional vertex coordinate information; wherein, the position training model is trained using at least two image samples with regional position labels; the regional position labels include regional vertex coordinate labels and position attitude labels, and the regional vertex coordinate labels are the coordinate values after normalizing the regional vertex coordinate information in the image samples; use the position training model to output a classification vector based on the position attitude label, and the classification vector includes the position attitude information of the card in the target image; Input the image region where the card is located into the detection training model, and obtain at least one image mask output by the detection training model corresponding to the position attitude information. One image mask corresponds to one type of feature information, and the pixel value of the pixel point on the image mask indicates the probability value that the pixel point belongs to the feature information pixel point; wherein, the detection training model is trained using at least two image samples with ground-truth labels, and the ground-truth of each field annotation in the image sample includes: the coordinate position of the field and the type of the field; Based on the pixel values of the pixel points on each image mask, obtain the target region in each image mask respectively; the target region is the region containing the feature information; Identify each target region to obtain the recognition result of the card, and the recognition result includes at least one piece of feature information.

8. A storage medium, characterized in that, The computer-executable instructions are stored in the storage medium, and when the computer-executable instructions are loaded and executed by the processor, the card recognition method according to any one of claims 1 to 3 above is implemented.

Citation Information

Patent Citations

  • A method and device for identifying a bank card number.

    CN109447059A

  • Identity card image detection method, device and equipment

    CN110059680A