Card authentication method and apparatus

By acquiring card images from multiple angles on the client side and using a card authentication model to verify authenticity, the problem of forged card image injection attacks is solved, thus improving the security of the eKYC system.

WO2026060779A1PCT designated stage Publication Date: 2026-03-26ANT BLOCKCHAIN TECHNOLOGY (SHANGHAI) CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-10-31
Publication Date
2026-03-26

AI Technical Summary

Technical Problem

Existing technologies are insufficient to effectively identify forged card images generated through image generation techniques, leading to injection attacks on eKYC systems and impacting customer information and property security.

Method used

The client instructs the user to capture card images in multiple ways, and uses a trained card identification model to determine whether the image was captured in multiple ways. The server then performs authenticity verification and uses differential images and spectrograms to enhance the accuracy of the verification.

Benefits of technology

It enables effective identification of counterfeit card images, improves the security of the eKYC system, and protects customer information and property security.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024128715_26032026_PF_FP_ABST
    Figure CN2024128715_26032026_PF_FP_ABST
Patent Text Reader

Abstract

Provided in the embodiments of the present description are a card authentication method and apparatus. The method comprises: a server receiving from a client a plurality of first card image frames for a card to be authenticated, wherein the client is used for instructing a user to collect images of said card in a plurality of ways that enable the client to collect different images of said card; and using a trained card authentication model to determine whether the plurality of first card image frames are collected in the plurality of ways, so as to determine an authenticity identification result for said card. In this way, the authenticity of cards is identified.
Need to check novelty before this filing date? Find Prior Art

Description

Card identification method and device

[0001] The present application claims priority to the Chinese patent application No. 2024113120594, filed on September 19, 2024, and entitled "Card identification method and device", the content of which is incorporated herein by reference in its entirety. TECHNICAL FIELD

[0002] The present specification relates to the technical field of image processing, and in particular to a card identification method and device. BACKGROUND

[0003] Currently, in order to provide users with better and safer services, institutions in the fields of financial services, banks, insurance, telecommunications, and e-commerce generally use online eKYC (electronic know your customer) systems to enable the corresponding institutions to understand the identity and background of customers with whom they establish business relationships, in order to prevent risks such as money laundering, fraud, and illegal financing.

[0004] In the eKYC system, the card of the customer needs to be verified and identified to verify and identify the identity of the customer. Currently, there are some malicious parties that use image generation technology to forge card images (such as card images generated by AIGC (Artificial Intelligence Generated Content, content generated by artificial intelligence technology)), and then use such forged card images to attack the eKYC system through injection attacks. In order to improve the security of their eKYC systems and protect customer information and property safety, relevant institutions need to identify and defend against injection attacks using such forged card images. The identification and defense against injection attacks using such forged card images to some extent depends on the identification of such forged card images.

[0005] Therefore, how to provide a card identification method that can identify the above-mentioned forged card images becomes a problem to be solved.

[0006] SUMMARY

[0007] One or more embodiments of the present specification provide a card identification method and device to identify the authenticity of a card.

[0008] According to a first aspect, a card identification method is provided, applied to a server, the method comprising:

[0009] receiving, from a client, a plurality of first card images of a card to be identified, wherein the client is configured to instruct a user to collect images of the card to be identified in a plurality of ways, and the plurality of ways enable the client to collect different images of the card to be identified;

[0010] The trained card identification model is used to determine whether the plurality of first card images are collected in the plurality of manners, so as to determine the authenticity identification result of the card to be identified.

[0011] According to a second aspect, a card identification method applied to a client is provided, including:

[0012] In the process of instructing a user to collect images of a card to be identified in a plurality of manners, a plurality of first card images of the card to be identified are acquired, wherein the plurality of manners enable the client to collect different images of the card to be identified.

[0013] The plurality of first card images are sent to a server, so that after the server receives the plurality of first card images, a trained card identification model is used to determine whether the plurality of first card images are collected in the plurality of manners, so as to determine the authenticity identification result of the card to be identified.

[0014] According to a third aspect, a card identification device deployed on a server is provided, including:

[0015] A receiving module is configured to receive, from a client, a plurality of first card images of a card to be identified, wherein the client is used to instruct a user to collect images of a card to be identified in a plurality of manners, and the plurality of manners enable the client to collect different images of the card to be identified.

[0016] A first determining module is configured to use a trained card identification model to determine whether the plurality of first card images are collected in the plurality of manners, so as to determine the authenticity identification result of the card to be identified.

[0017] According to a fourth aspect, a card identification device deployed on a client is provided, including:

[0018] An acquiring module is configured to acquire, in the process of instructing a user to collect images of a card to be identified in a plurality of manners, a plurality of first card images of the card to be identified, wherein the plurality of manners enable the client to collect different images of the card to be identified.

[0019] A sending module is configured to send the plurality of first card images to a server, so that after the server receives the plurality of first card images, a trained card identification model is used to determine whether the plurality of first card images are collected in the plurality of manners, so as to determine the authenticity identification result of the card to be identified.

[0020] According to a fifth aspect, a computer readable storage medium is provided, having stored thereon a computer program which, when executed in a computer, causes the computer to perform the method according to the first aspect or the second aspect.

[0021] According to a sixth aspect, a computing device is provided, comprising a memory and a processor, wherein the memory has stored thereon executable code that, when executed by the processor, implements the method according to the first aspect or the second aspect.

[0022] According to the card identification method and device provided by the embodiments of the present specification, the server receives a plurality of first card images of a card to be identified from a client, wherein the client is used to instruct a user to collect images of the card to be identified in a plurality of ways, and the plurality of ways enable the client to collect different images of the card to be identified; and a trained card identification model is used to determine whether the plurality of first card images are collected in a plurality of ways, so as to determine a true or false identification result of the card to be identified. In the above process, considering that there is generally no change between the plurality of images injected by the injection attack, when the client collects images of the card to be identified, the user is instructed to collect images of the card to be identified in a plurality of ways, so that the plurality of card images obtained for the real card are different images; subsequently, after the server obtains the plurality of first card images of the card to be identified, the trained card identification model is used to determine whether the plurality of first card images are collected in a plurality of ways, so as to determine whether the images collected by the client are real images collected, and then determine the true or false identification result of the card to be identified, thereby realizing identification of the true or false of the card. BRIEF DESCRIPTION OF DRAWINGS

[0023] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings needed to be used in the embodiment description. Obviously, the drawings in the following description are only some embodiments of the present application, and for those skilled in the art, other drawings can also be obtained from these drawings without creative labor.

[0024] Fig. 1 is a schematic diagram of an implementation framework of one embodiment disclosed by the present specification;

[0025] Fig. 2 is a schematic diagram of one flow of a training process of a card identification model provided by the embodiment;

[0026] Fig. 3 is a schematic diagram of one flow of a card identification method provided by the embodiment;

[0027] Fig. 4 is another schematic diagram of a flow of a card identification method provided by the embodiment;

[0028] Fig. 5 is another schematic diagram of a flow of a card identification method provided by the embodiment;

[0029] FIG. 6 is another flow diagram of a card identification method according to an embodiment;

[0030] FIG. 7 is a schematic block diagram of a card identification device deployed on a server according to an embodiment;

[0031] FIG. 8 is a schematic block diagram of a card identification device deployed on a client according to an embodiment. DETAILED DESCRIPTION

[0032] The technical solutions of the embodiments of the present specification will be described in detail below with reference to the accompanying drawings.

[0033] The present specification discloses a card identification method and device. First, the application scenario and technical concept of the method will be introduced as follows.

[0034] As mentioned above, in the eKYC system, the card (for example, including but not limited to bank card, identity card, etc.) of the customer needs to be verified and identified to verify and identify the identity of the customer. Currently, there are some malicious parties that use image generation technology to forge card images (for example, card images generated by AIGC (Artificial Intelligence Generated Content, content generated by artificial intelligence technology)), and then use such forged card images to attack the eKYC system through injection attacks. In order to improve the security of the eKYC system and protect the customer information and property safety, related institutions need to identify and defend against injection attacks using such forged card images. Among them, the identification and defense against injection attacks using such forged card images to a certain extent depends on the identification of such forged card images.

[0035] In the analysis of forged card images used in the scenario of attacking the eKYC system through injection attacks, it is found that the multiple frames of forged card images used in one injection attack generally do not change, for example, their corresponding collection angles are often indistinguishable, that is, the collection angles corresponding to the multiple frames of forged card images injected through injection attacks are generally unchanged (for example, the multiple frames of forged card images used in one injection attack may be the same image).

[0036] In view of this, the inventors propose a card identification method to identify the authenticity of the card, and thus to identify and defend against injection attacks using forged card images. For example, FIG. 1 shows an implementation scenario according to an embodiment of the present specification. An exemplary structure of a card identification system is shown in the implementation scenario. As shown in FIG. 1, the card identification system includes a client and a server.

[0037] In some possible examples, the card identification system can be a related system of any authentication system (e.g., an eKYC system) that needs to utilize card verification to verify the identity of a client, so as to first verify the authenticity of the card in the card image uploaded by the client through the related system, and then trigger the authentication system to perform a process of utilizing the authentic card to verify the identity of the client in the case where the card in the card image uploaded by the client is verified to be an authentic card. In yet some possible examples, the card identification system is the aforementioned authentication system, so as to first verify the authenticity of the card in the card image uploaded by the client, and then trigger a process of utilizing the authentic card to verify the identity of the client in the case where the card in the card image uploaded by the client is verified to be an authentic card.

[0038] On the client side, the user is instructed to collect images of the card to be identified in multiple manners, and multiple images X are obtained in the instruction process; then the client sends the multiple images X to the server. The multiple manners enable the client to collect different images of the card to be identified. For example, the multiple collection manners can include, but are not limited to, a collection manner of collecting images of the card at different collection angles, and / or a collection manner of collecting images of the card in a flash light setting and a non-flash light setting. The collection manner of collecting images of the card at different collection angles can include a front collection manner and other non-front collection manners. The front collection manner and the other non-front collection manners can be collection manners in the non-flash light setting or the flash light setting.

[0039] Hereinafter, when referring to the collection of images of the card in the front collection manner and the other non-front collection manners, without mentioning whether the flash light or the non-flash light is set, it can be understood as the collection of images of the card in the front collection manner and the other non-front collection manners in the non-flash light setting.

[0040] For example, the front collection manner can refer to a collection manner in which the card is not tilted, for example, a collection manner in which the plane on which the card is located is parallel to the plane on which the display page of the client is located, or a collection manner in which the plane on which the card is located is parallel to the plane on which the image collection device called by the client is located. The parallel can refer to that the included angle between the plane on which the card is located and the plane on which the display page of the client is located is less than a preset included angle threshold.

[0041] The other non-frontal collection manner can include, but is not limited to: a collection manner of tilting the card certificate upward by X degrees, for example, an angle between a plane where the card certificate is located and a plane where a display page of the client is located in a vertical direction is X degrees; a collection manner of tilting the card certificate downward by Y degrees, for example, an angle between the plane where the card certificate is located and the plane where the display page of the client is located in the vertical direction is -Y degrees; a collection manner of tilting the card certificate to the left by Z degrees, for example, an angle between the plane where the card certificate is located and the plane where the display page of the client is located in a horizontal direction is Z degrees; and a collection manner of tilting the card certificate to the right by Q degrees, for example, an angle between the plane where the card certificate is located and the plane where the display page of the client is located in the horizontal direction is -Q degrees. In some examples, considering that the user may need to manually assist in achieving the above (tilting) angle, the above (tilting) angle can be allowed to have an error, as long as the error is within an allowed range. The preset angle threshold is smaller in value than the above X degrees, Y degrees, Z degrees, and Q degrees.

[0042] As shown in FIG. 1, the process in which the client instructs the user to collect images of the card certificate in multiple ways can include: first instructing the user to collect an image of the card certificate in a frontal collection manner (without a flash), to obtain a first frame image X1; after obtaining the first frame image X1, the client can then randomly determine another collection manner 1, to instruct the user to collect an image of the card certificate in the other collection manner 1, to obtain a second frame image X2.

[0043] The other collection manner 1 can be any one of the above-mentioned other non-frontal collection manners, or a collection manner with a flash (i.e., the client can call the flash of the device where it is located to enable the user to collect an image of the card certificate under the condition of being provided with a flash, at this time, the card certificate can not be tilted).

[0044] In some examples, as shown in FIG. 1, after obtaining the first frame image X1 and the second frame image X2, the client obtains multiple frame images X, and can directly upload (send) the multiple frame images X including the first frame image X1 and the second frame image X2 to the server. In yet some examples, after obtaining the second frame image X2, the client can continue to randomly determine another collection manner 2 (which can be different from the above-mentioned other collection manner 1), to instruct the user to collect an image of the card certificate in the other collection manner 2, to obtain a third frame image X3, and so on, to obtain multiple frame images X.

[0045] If the user is a normal user, i.e., a user who needs to identify the real card certificate, he will cooperate with the instructions of the client to collect the images of the real card certificate he needs to identify in multiple ways through the image collection device called by the client. Correspondingly, the multiple images X obtained by the client are different images collected by the client for the user's real card certificate in the aforementioned multiple ways. If the user is an abnormal user (such as he maliciously steals other users, for example, other users' accounts), i.e., a user who needs to attack the corresponding authentication system through injection attack (for example, needs to pretend to be himself as the user who steals other users), although the client will instruct the user to collect the images of the card certificate he needs to identify in multiple ways through the image collection device called by the client, but such users will inject the card certificate images they have forged in advance (for example, card certificate images forged by image generation technology, which include the card certificate of the other user stolen in advance) through injection attack, i.e., the card certificate images they have forged in advance will be collected by the image collection device called as images, and will be obtained by the client. Correspondingly, the multiple images X obtained by the client are the forged card certificate images injected by the injection attack.

[0046] It can be understood that there are differences between the multiple images collected in multiple ways on the client side, which include but are not limited to different imaging angles caused by different collection angles and / or different brightness rendering effects in the images caused by different settings of the flash when collecting (for example, the images collected when the flash is set have bright spots, while the images collected without the flash do not have bright spots). And as mentioned earlier, the multiple forged card certificate images injected by the injection attack generally have the same corresponding collection angle, i.e., there is no angle difference between the multiple forged card certificate images, and it is also impossible to forge different brightness rendering effects when the flash is set and when the flash is not set (for example, the images collected when the flash is set have bright spots, while the images collected without the flash do not have bright spots).

[0047] In view of the above, on the server side, a trained card identification model is pre-stored, which is used to determine whether the input multi-frame images are collected in the multiple manners indicated by the client. Specifically, as shown in FIG. 1, after the server obtains the multi-frame images X uploaded by the client, in order to verify whether the multi-frame images X uploaded by the client are images collected by the user for the card to be identified or are fake card images injected by the injection attack, in an implementation manner, the server can directly input the multi-frame images X received from the client into the card identification model to determine whether the multi-frame images X are collected in the multiple manners indicated by the client by using the card identification model, so as to determine the authenticity identification result of the card to be identified. If it is determined that the multi-frame images X are collected in the multiple manners indicated by the client, it is determined that the card to be identified is a real card, and if it is determined that the multi-frame images X are not collected in the multiple manners indicated by the client, it is determined that the card to be identified is a fake card.

[0048] In another implementation manner, after the server obtains the multi-frame images X uploaded by the client, the server can generate the difference image C and the spectrum image P corresponding to the multi-frame images X based on the multi-frame images X received from the client. The difference image C and the spectrum image P corresponding to the multi-frame images X can better highlight the differences between the multi-frame images X from the gradient dimension and the frequency domain dimension, and help to better improve the accuracy of the card authenticity identification result. Then, the server inputs the multi-frame images X and the difference image C and the spectrum image P corresponding to the multi-frame images X into the card identification model to determine whether the multi-frame images X are collected in the multiple manners indicated by the client by using the card identification model, so as to determine the authenticity identification result of the card to be identified.

[0049] The trained card identification model can be a network model trained based on specified training data. The training data for training the card identification model can at least include a plurality of sample image groups and corresponding label data. Each sample image group at least includes multi-frame sample card images collected in the multiple manners, and the corresponding label data of the sample image group includes data indicating that the sample card corresponding to the sample image group is a real card. The multiple manners can include but are not limited to collecting card images at different collection angles and / or collecting card images in the condition of setting a flash and without a flash.

[0050] In some examples, in the case that a single sample image group includes multi-frame sample card images, the single sample image group can further include a sample difference image and / or a sample spectrum image generated based on the multi-frame sample card images.

[0051] In some possible examples, the training data can further include a sample image group containing a fake card image and label data of the sample image group, where each sample image group at least includes a plurality of frame fake card images of a fake card (there can be no difference between the plurality of frame fake card images, or there can be some fake differences), and the label data of the sample image group includes data indicating that the corresponding card is a fake card. Each sample image group can further include a sample difference image and / or a sample spectrum image generated based on the frame fake card image.

[0052] In the above process, considering that there is generally no change between the plurality of frame images injected by the injection attack, when the client collects images of the card to be identified, that is, instructs the user to collect images of the card to be identified in multiple ways, so that the plurality of frame images obtained for the real card are different images. Subsequently, after the server obtains the plurality of frame images X, the trained card identification model is used to determine whether the plurality of frame images X are collected in multiple ways, to determine whether the images collected by the client are real images collected, and to determine the authenticity identification result of the card corresponding to the plurality of frame images X, to achieve identification of the authenticity of the card.

[0053] The card identification method provided in the specification will be described in detail below in conjunction with specific embodiments.

[0054] In the card identification process, the server needs to use the card identification model to verify the authenticity of the card in the plurality of frame images uploaded by the client. The training process of the card identification model is also provided in the embodiments of the specification. The training process of the card identification model will be introduced first.

[0055] As shown in FIG. 2, a schematic diagram of a training process of a card identification model is shown, which can be executed by a first electronic device. The first electronic device can be implemented by any device, apparatus, platform, device cluster, etc. having computing and processing capabilities. The training process of the card identification model can include the following steps S210-S230.

[0056] In step S210, a card identification model to be trained is obtained. The card identification model to be trained can be any network model capable of classifying image sequences in related technologies, such as a network model based on a transformer, a long short-term memory neural network model, and other network models for image classification processing. Specifically, it can be a neural network model based on ViT (Visual Transformer). The first electronic device can obtain the card identification model to be trained at a preset storage location.

[0057] After the first electronic device obtains the card certificate identification model to be trained, in step S220, a first sample image group and first label data thereof are obtained. The first label data at least includes data indicating whether the card certificate corresponding to the first sample image group is a true certificate.

[0058] The first sample image group and the first label data thereof can be any sample image group and label data thereof obtained from the aforementioned training data. In one case, the first sample image group can at least include sample card certificate images obtained by collecting real sample card certificates in multiple ways, i.e., the card certificate corresponding to the first sample image group is a real card certificate, and accordingly, the corresponding first label data includes data indicating that the card certificate corresponding to the first sample image group is a true certificate. The card certificate corresponding to the first sample image group can be a bank card or an identity certificate.

[0059] Exemplarily, the multiple frames of sample card certificate images in the first sample image group can include images obtained by collecting real card certificates under different collection angles, i.e., the corresponding collection angles between the multiple frames of images in the first sample image group are different.

[0060] Exemplarily, the multiple frames of sample card certificate images in the first sample image group can include images obtained by collecting sample card certificates under the condition of setting a flash and the condition of not setting a flash, and the multiple frames of sample card certificate images include an image obtained by collecting a real card certificate under the condition of not setting a flash and an image obtained by collecting a real card certificate under the condition of setting a flash.

[0061] Exemplarily, the multiple frames of sample card certificate images in the first sample image group can include images obtained by collecting real card certificates under different collection angles and images obtained by collecting real card certificates under the condition of setting a flash and the condition of not setting a flash. Specifically, the first sample image group includes an image obtained by collecting a real card certificate under the condition of collection angle 1 and not setting a flash, an image obtained by collecting a real card certificate under the condition of collection angle 2 (different from collection angle 1) and not setting a flash, and an image obtained by collecting a real card certificate under the condition of collection angle 3 (the same as collection angle 1 or 2, or different) and setting a flash.

[0062] In another case, the first sample image group can at least include multiple frames of sample card certificate images containing fake card certificates, i.e., the card certificate corresponding to the first sample image group is a fake card certificate, and accordingly, the corresponding first label data at least includes data indicating that the card certificate corresponding to the first sample image group is a fake certificate (i.e., a false certificate). The fake card certificate can be a fake bank card or a fake identity certificate.

[0063] In some examples, considering that the difference image between images can better highlight the differences between images collected in multiple ways, the first sample image group can further include a first sample difference image generated based on the multiple sample card image frames included therein. In some possible examples, the image collected under the flash setting can better highlight the anti-counterfeiting mark when the multiple sample card image frames include images collected under the flash setting and the no flash setting, and the real card corresponding to the multiple sample card image frames has an anti-counterfeiting mark. Accordingly, the difference image generated based on such multiple sample card image frames can better represent that the card corresponding to the multiple sample card image frames is a real card.

[0064] In yet some examples, considering that the frequency spectrum generated based on an image can reflect the features of the image from the frequency domain dimension, it can better help to achieve identification of the authenticity of the card. Accordingly, the first sample image group can further include a first sample frequency spectrum generated based on a specified sample card image in the multiple sample card image frames included therein. The specified sample card image can be one or more of the multiple sample card image frames.

[0065] For example, when the multiple sample card image frames in the first sample image group include an image of a real card collected under a flash setting, the first sample frequency spectrum can be generated based on the image of the real card collected under the flash setting. The resulting frequency spectrum can reflect the features of the image from the frequency domain dimension, and can better assist in identifying the detailed information of the card material (e.g., highlighting that the card material causes the image to have a light spot feature under the action of the flash). When the multiple sample card image frames in the first sample image group include images of a real card collected under different collection angles, the first sample frequency spectrum can be generated based on the image of the real card collected under the aforementioned non-front collection manner (e.g., the collection manner of collecting the real card by tilting X degrees upward, tilting Y degrees downward, etc.).

[0066] In the embodiments of the present specification, the multiple sample card image frames included in the first sample image group, the first sample difference image, and / or the first sample frequency spectrum can be referred to as sample images in the first sample image group.

[0067] In some examples, the first label data corresponding to the first sample image group can further include data indicating a plurality of manners corresponding to the plurality of sample images in the first sample image group. For example, when the plurality of sample card image in the first sample image group includes images of the real card captured at different angles, the data indicating the plurality of manners corresponding to the plurality of sample images in the first sample image group can include data indicating different angles, or can further include specific tilt directions (e.g., upward tilt, downward tilt, left tilt, or right tilt).

[0068] For another example, when the plurality of sample card image in the first sample image group includes images of the real card captured at different angles, and images of the real card captured with and without flash, the data indicating the plurality of manners corresponding to the plurality of sample images in the first sample image group can include data indicating different angles and different flash settings.

[0069] For another example, when the plurality of sample card image in the first sample image group includes images of the real card captured with and without flash, the data indicating the plurality of manners corresponding to the plurality of sample images in the first sample image group can include data indicating different flash settings.

[0070] In some examples, when the card corresponding to the first sample image group is a fake card, the first label data corresponding to the first sample image group can further include data indicating a plurality of manners corresponding to the plurality of sample images in the first sample image group. The data can be null.

[0071] Then, in step S230, the first sample image group and the first label data are used to train the card identification model to be trained.

[0072] In some examples, step S230 can include steps 01-02. In step 01, the first electronic device can use the card identification model to be trained to process the first sample image group to determine whether the first sample image group is captured in the plurality of manners to determine a prediction result corresponding to the first sample image group. The prediction result includes data indicating whether the card corresponding to the first sample image group is a real card. In step 02, the first label data of the first sample image group and the prediction result are used to adjust the model parameters of the card identification model to be trained.

[0073] Exemplarily, the first electronic device can input the first sample image group into the card identification model to be trained, process sample images in the first sample image group by the card identification model to be trained, determine whether the first sample image group is collected in multiple manners, determine a prediction result corresponding to the first sample image group, and the prediction result at least includes data indicating whether the card corresponding to the first sample image group is a real card. Then, a preset loss function is used to construct a prediction loss based on a difference between the data indicating whether the card corresponding to the first sample image group is a real card in the prediction result and the data indicating whether the card corresponding to the first sample image group is a real card in the first label data. Then, the model parameters of the card identification model to be trained are adjusted to minimize the prediction loss, that is, to minimize the difference between the prediction result and the first label data.

[0074] In some examples, the data indicating whether the card corresponding to the first sample image group is a real card in the first label data can be a label value to indicate whether the card corresponding to the first sample image group is a real card. If the label value is 1, it indicates that the card corresponding to the first sample image group is a real card. If the label value is 0, it indicates that the card corresponding to the first sample image group is a fake card. The data indicating whether the card corresponding to the first sample image group is a real card in the prediction data can be a probability value indicating whether the card corresponding to the first sample image group is a real card predicted by the model. The closer the probability value is to 1, the greater the possibility that the card corresponding to the first sample image group is a real card predicted by the model.

[0075] Subsequently, the first electronic device can use a preset loss function to construct a prediction loss using a difference between the prediction data (probability value) and the first label data (label value), adjust the model parameters of the card identification model to be trained to minimize the prediction loss, that is, to minimize the difference between the prediction data and the first label data, so that the card identification model has the ability to identify the authenticity of the card corresponding to the input image sequence.

[0076] In yet some examples, in addition to the data indicating whether the card corresponding to the first sample image group is a real card in the first label data, the prediction data can include data indicating whether the card corresponding to the first sample image group is a real card predicted by the card identification model, and data indicating multiple manners of multiple sample images in the first sample image group predicted by the card identification model, in the case that the data indicating multiple manners of multiple sample images in the first sample image group is included in the first label data.

[0077] In the above example, the first electronic device adopts a preset loss function, utilizes a difference between the prediction data and the first label data (i.e., utilizes a first difference between data indicating whether the card certificate corresponding to the first sample image group is a true certificate in the first label data and data indicating whether the card certificate corresponding to the first sample image group is a true certificate predicted by the card certificate identification model in the prediction data, and a second difference between data indicating a plurality of manners corresponding to a plurality of sample images in the first sample image group in the first label data and data indicating the plurality of manners corresponding to the plurality of sample images in the first sample image group predicted by the card certificate identification model in the prediction data), and constructs the prediction loss.

[0078] Then, the model parameters of the card certificate identification model to be trained are adjusted to minimize the prediction loss, i.e., to minimize the difference between the prediction data and the first label data (i.e., to minimize the first difference and the second difference), so that the card certificate identification model has the ability to identify the authenticity of the card certificate corresponding to the input image sequence and the ability to identify the plurality of manners corresponding to the plurality of images in the input image sequence (i.e., to predict whether the plurality of images in the input image sequence are collected at different angles, or are set at different flash settings, or are collected at different angles and are set at different flash settings).

[0079] It can be understood that the above steps S210-S230 are a model iteration training process for the card certificate identification model. In order to train a better card certificate identification model, the above process can be executed multiple times. That is, after step S230, the model parameters of the card certificate identification model are adjusted, and the step S220 is returned to be executed.

[0080] The stop condition of the above model iteration training process can include that the number of iteration training reaches a preset number threshold, or the iteration training duration reaches a preset duration, or the prediction loss is less than a set loss threshold, and the like.

[0081] In the above model iteration training process, the sample image group containing the plurality of sample certificates with real card needles and the label data thereof, and the sample image group containing the plurality of sample certificates with fake card certificates and the label data thereof are used to jointly train the card certificate identification model, so as to improve the accuracy and stability of the card certificate identification result of the trained card certificate identification model.

[0082] In some possible examples, the card identification model to be trained includes a blocking layer, a mapping layer, an attention mechanism-based extraction layer, and a classification layer. In an example, the blocking layer in the card identification model to be trained can block the input images based on a preset size w0*h0, which can be set according to actual needs. The mapping layer can be set as a linear mapping layer or other types of mapping layers (for example, a nonlinear mapping layer), to map each input image group into a feature vector (that is, a token) of a preset length. The attention mechanism-based extraction layer can be a Transformer structure extraction layer, and the classification layer can be implemented by a fully connected layer, for example, a multilayer perceptron or other classifiers.

[0083] Correspondingly, the step 01 can include the following steps 011-015.

[0084] In step 011, the blocking layer is used to block each sample image in the first sample image group to obtain a sample image block corresponding to each sample image. After the first electronic device obtains the first sample image group and the first label data, the first electronic device inputs the sample image group in the first sample image group into the blocking layer to block each sample image in the first sample image group to obtain a sample image block corresponding to each sample image. When the blocking layer blocks each sample image in the first sample image group, each sample image is blocked according to a preset size without overlapping.

[0085] It can be understood that each sample image in the first sample image group is sorted according to a specified sorting order, for example, the first sample image group includes a sample card image 1 collected in a front-facing manner, one or more sample card images 2 collected in other non-front-facing manner or under the condition of setting a flash, and a first sample difference image and a first sample spectrum image generated based on the sample card image 1 and the sample card image 2. Correspondingly, the sorting order between each sample image can be that the sample card image 1 is located at the first position, each sample card image 2 is located after the sample card image 1, the first sample difference image is located after the last sample card image 2, and the first sample spectrum image is located after the first sample difference image. Each sample image in the first sample image group can also be sorted according to other sorting orders, for example, the first sample spectrum image is located before the first sample difference image.

[0086] Then in step 012, the sample image blocks corresponding to each sample image are combined based on the inter-frame relationship between each sample image to obtain a plurality of sample image block groups, each of which includes one or more sample image blocks, and the plurality of sample image blocks in a single sample image block group are located at the same position in their respective sample images.

[0087] In this step, the card certificate identification model to be trained considers not only the spatial dimension information between two-dimensional image blocks in a single frame image in the image group (i.e., image sequence), but also the time dimension information between multiple images in the image group during the card certificate identification process, that is, a three-dimensional image block group is formed by combining image blocks at corresponding positions in multiple images in the image group to reflect the time dimension information between the multiple images; then, the authenticity of the card certificate corresponding to the multiple images in the image group is identified by combining the spatial dimension information between the two-dimensional image blocks and the time dimension information between the multiple images reflected by the three-dimensional image block group.

[0088] Specifically, the first electronic device combines sample image blocks corresponding to each sample image by using the inter-frame relationship between the sample images to obtain multiple groups of sample image block groups, wherein each group of sample image block groups includes one or more sample image blocks, and the multiple sample image blocks in a single sample image block group have the same position in their respective sample images.

[0089] For example, a first sample image group includes sample image A, sample image B, sample image C, and sample image D. A sample image block group includes multiple sample image blocks, such as sample image block A1 and sample image block B1. In this case, the sample image block A1 in the sample image block group has a position in the sample image A, which is the top-left pixel position (1, 1) and the area indicated by the bottom-right pixel position (w0, h0). Correspondingly, the sample image block B1 in the sample image block group has a position in the sample image B, which is the top-left pixel position (1, 1) and the area indicated by the bottom-right pixel position (w0, h0). For another example, a sample image block group includes multiple sample image blocks, such as sample image block A2, sample image block B2, sample image block C2, and sample image block D2. In this case, the sample image block A2 in the sample image block group has a position in the sample image A, which is the top-left pixel position (1, 1+h0) and the area indicated by the bottom-right pixel position (w0, 2h0). Correspondingly, the sample image block B1 in the sample image block group has a position in the sample image B, which is the top-left pixel position (1, 1+h0) and the area indicated by the bottom-right pixel position (w0, 2h0). The sample image block C2 in the sample image block group has a position in the sample image C, which is the top-left pixel position (1, 1+h0) and the area indicated by the bottom-right pixel position (w0, 2h0). The sample image block D2 in the sample image block group has a position in the sample image D, which is the top-left pixel position (1, 1+h0) and the area indicated by the bottom-right pixel position (w0, 2h0).

[0090] It can be understood that in the case that the first sample image group includes N sample images, the number of sample image blocks included in a sample image block group can be an integer from 1 to N. Wherein, N is an integer greater than or equal to 2.

[0091] After the first electronic device obtains the plurality of sample image block groups, in step 013, the first electronic device processes each sample image block group by using a mapping layer to obtain a sample mapping sequence. In this step, the first electronic device inputs each sample image block group into the mapping layer to process each sample image block group by using the mapping layer to obtain a sample mapping corresponding to each sample image block group (i.e., the aforementioned feature vector of a preset length, i.e., token).

[0092] After the first electronic device obtains the sample mapping corresponding to each sample image block group, in order to preserve the spatiotemporal position information of the sample image blocks in each sample image block group and the image sequence composed of the sample images in the first sample image group, the first electronic device can use a preset position encoding algorithm to perform position encoding on the sample mapping corresponding to each sample image block group to obtain the position encoding of the sample mapping corresponding to each sample image block group; then, the first electronic device sorts the sample mapping corresponding to each sample image block group with position encoding according to the sorting order indicated by the position encoding to obtain a sample mapping sequence.

[0093] The preset position encoding algorithm can include but is not limited to a sine position encoding algorithm, a cosine position encoding algorithm, or a position encoding algorithm based on a trainable embedding vector, in which case the model parameters that need to be trained are carried.

[0094] After that, in step 014, the first electronic device processes the sample mapping sequence by using an extraction layer to obtain a sample image representation. In this step, the first electronic device inputs the sample mapping sequence into the extraction layer based on the attention mechanism to perform self-attention calculation on the sample mapping sequence based on the attention mechanism by using the extraction layer to capture the spatiotemporal dependency between the sample mappings in the sample mapping sequence to obtain an image representation that can better represent the feature information of the sample images in the first sample image group, which is referred to as a sample image representation hereinafter. The sample image representation can include more image feature information that is beneficial for card identification.

[0095] Then, in step 015, the first electronic device processes the sample image representation by using a classification layer to determine prediction data. After the first electronic device obtains the sample image representation, the first electronic device inputs the sample image representation into the classification layer to process the sample image representation by using the classification layer to obtain prediction data. Then, based on the prediction data and the aforementioned first label data, the first electronic device adjusts the model parameters of the card identification model to be trained.

[0096] The first electronic device obtains the trained card certificate identification model after the card certificate identification model training is completed, i.e., the model iteration training process of the card certificate identification model reaches the stopping condition, and stores the trained card certificate identification model in a designated storage space. The trained card certificate identification model can be used for card certificate identification.

[0097] FIG. 3 shows a flowchart of a card certificate identification method in an embodiment of the present specification. The method is applied to a client and a server in a card certificate identification system. The server and the client can run on different electronic devices. Any electronic device can be implemented by any device, equipment, platform, device cluster, etc. with computing and processing capabilities.

[0098] The card certificate identification system can be an embedded system of any authentication system (e.g., an eKYC system) that needs to verify the identity of a client using a card certificate, or the card certificate identification system can be the aforementioned authentication system.

[0099] In the card certificate identification process, as shown in FIG. 3, the method includes the following steps S310-S340:

[0100] In step S310, the client obtains a plurality of first card certificate images of a card certificate to be identified in the process of instructing the user to collect images of the card certificate to be identified in multiple ways. The multiple ways enable the client to collect different images of the card certificate to be identified.

[0101] In some exemplary scenarios, the client can call the image collection device of the device it is located in when it determines that the user has a demand to upload card certificate images, aiming to collect images through the image collection device of the device it is located in. The client calls the image collection device of the device it is located in, i.e., starts the image collection device of the device it is located in, to collect images. Subsequently, the client can output first instruction information to instruct the user to collect images of the card certificate to be identified in a corresponding manner (e.g., the aforementioned front-side collection manner) to obtain card certificate images. Then, the client can randomly determine the collection manner of the next frame and instruct the user to collect images of the card certificate to be identified based on the randomly determined collection manner of the next frame. For example, the client continues to output second instruction information to continue to instruct the user to collect images of the card certificate to be identified in the collection manner of the next frame (e.g., the aforementioned other non-front-side collection manner), or the client can directly call and start the flash to instruct the user to collect images of the card certificate to be identified under the condition of setting the flash to obtain card certificate images.

[0102] In a possible example, the client obtains two frames of the card image, and takes the two frames of the card image as the plurality of frames of the first card image. In another possible example, after the client obtains the two frames of the card image, the client can continue to randomly determine the acquisition manner of the next frame, and instruct the user to acquire the image of the card to be identified based on the randomly determined acquisition manner of the next frame; for example, outputting third instruction information to continue to instruct the user to acquire the image of the card to be identified in a new manner (for example, the other non-front acquisition manner described above or the acquisition manner of acquiring the image under the condition of setting the flash), and so on, to obtain a specified number of frames of the card image, and determine the plurality of frames of the first card image.

[0103] In some possible cases, the card image obtained as described above can be an image actually acquired by the image acquisition device as described above, and the image actually acquired by the image acquisition device as described above contains a real card. In some other possible cases, a malicious party can use a pre-forged card image to replace an image actually acquired by the image acquisition device as described above by means of an injection attack, and accordingly, the card image obtained as described above can be a pre-forged card image injected by means of the injection attack, to pretend that the card image acquired by the image acquisition device as instructed by the client contains a forged card (i.e., a fake card).

[0104] In some examples, the plurality of manners include an acquisition manner of acquiring images at different acquisition angles, and / or an acquisition manner of acquiring images under the condition of setting the flash and the condition of not setting the flash. In some possible examples, in the case where the client exists in the H5 browser mode, there can be a case where the client cannot call the image acquisition device (for example, a camera) of the device in which the client exists, and accordingly, the client can instruct the user to acquire the image of the card to be identified at different acquisition angles. In some other possible examples, in the case where the client exists in the native SDK (Software Development Kit) mode, the client can instruct the user to acquire the image of the card to be identified at different acquisition angles, and / or acquire the image of the card to be identified under the condition of setting the flash and the condition of not setting the flash.

[0105] In some possible examples, in order to ensure the accuracy of the identification result obtained by the server, it is necessary to ensure that the image quality of the plurality of frames of the first card image obtained by the client is high (for example, including a clear and complete card to be identified), and specifically, the plurality of frames of the first card image all satisfy a preset image quality condition, and the preset image quality condition can include at least one of the following conditions: the image clarity is not lower than a preset clarity threshold, the image brightness value is not lower than a preset brightness threshold, the image contains a specified corner point of the card to be identified, and the card to be identified is located in a specified acquisition area during image acquisition.

[0106] In some specific examples, after the client instructs the user to collect the image of the card to be identified in a manner a to obtain a card image a (wherein the card image a can be an image collected by the client through the called image collection device, or an image injected in an injection attack manner containing a fake card), the client can determine whether the card image a meets the preset image quality condition, and if it is determined that the card image a does not meet the preset image quality condition, the card image a can be discarded, and the user is instructed to collect the image of the card to be identified in a manner b (which can be the same as or different from the aforementioned manner a) to obtain a new card image b (wherein the card image b can be an image collected by the client through the called image collection device, or an image injected in an injection attack manner containing a fake card), and the determination of whether the card image b meets the preset image quality condition is continued; if it is determined that the card image b meets the preset image quality condition, the card image b is taken as the first card image, and then the user is instructed to collect the image of the card to be identified in another manner c (different from the manner b) to obtain another card image c (wherein the card image b can be an image collected by the client through the called image collection device, or an image injected in an injection attack manner containing a fake card), and the determination of whether the card image c meets the preset image quality condition is continued, and so on, until a specified number of multiple first card images are obtained.

[0107] After the client obtains multiple first card images, in step S320, the client sends the multiple first card images to the server. Correspondingly, in step S330, the server receives the multiple first card images from the client.

[0108] After the server receives the multiple first card images from the client, in order to ensure the security of the corresponding authentication system, the server needs to determine whether the card to be identified corresponding to the multiple first card images is a true card after receiving the multiple first card images from the client, and correspondingly, in step S340, the server determines whether the multiple first card images are collected in multiple manners by using the trained card identification model to determine the true or false identification result of the card to be identified.

[0109] In some examples, after receiving the plurality of first card images from the client, the server assembles the plurality of first card images into an image sequence, and inputs the image sequence into the card identification model to obtain the authenticity identification result. In this step, after obtaining the plurality of first card images, the server directly assembles the plurality of first card images into an image sequence, wherein the plurality of first card images can be sorted according to a specified sorting order (for example, according to the order in which the client instructs the user to collect images) to form an image sequence; then the image sequence (i.e. the sorted plurality of first card images) is input into the card identification model to make the card identification model process the image sequence to determine whether the plurality of first card images are collected in the aforementioned multiple ways, so as to determine the authenticity identification result for the card to be identified.

[0110] In yet other examples, considering that the difference image between images can better highlight the differences between images collected in multiple ways, in order to better improve the accuracy of the identification result, after receiving the plurality of first card images from the client, the server at step S340 can include steps 11-12 as follows:

[0111] At step 11, the server determines the difference image between the plurality of first card images based on the plurality of first card images. In this step, when the plurality of first card images is two, the difference image between the two first card images can be determined based on the two first card images. The difference image can reflect the gradient information between the two first card images, and can better highlight the differences between the two first card images.

[0112] When the plurality of first card images is three or more, a difference image can be generated based on each two images in the plurality of first card images, to obtain a plurality of difference images between the plurality of first card images. For example, the plurality of first card images can be sorted according to the order in which the client instructs the user to collect images, and accordingly, the difference image between each adjacent two images in the plurality of first card images can be calculated in sequence according to the order to obtain the difference image between the plurality of first card images.

[0113] Afterwards, at step 12, the server groups the plurality of first card images and the difference images into an image sequence, and inputs the image sequence into the card identification model to obtain the authenticity identification result. In this step, the server groups the plurality of first card images and the difference images into an image sequence. For example, in the image sequence, the plurality of first card images are arranged in front, and the difference images are arranged in the back. The image sequence is input into the card identification model to process the image sequence by the card identification model to obtain the authenticity identification result. In some possible examples, in the image sequence, the plurality of first card images are sorted according to the order of image acquisition indicated by the client, the difference images between the plurality of first card images are calculated in the order, and the difference images are sorted according to the front and back order of the respective first card images.

[0114] In yet some examples, considering that the spectrum map generated based on the image can reflect the features of the image from the frequency domain dimension, and can better help to achieve the identification of the authenticity of the card, correspondingly, in order to better improve the accuracy of the identification result, after receiving the plurality of first card images from the client, the server can include the following steps 21-22 at step S340:

[0115] At step 21, the server determines the spectrum map corresponding to the plurality of first card images based on a specified image in the plurality of first card images. The specified image can be any one or more images in the plurality of first card images.

[0116] In some possible examples, in the case that the plurality of first card images include images acquired under the condition of setting the flash, the specified image can be the image acquired under the condition of setting the flash in the plurality of first card images, so that the spectrum map obtained can reflect the features of the image from the frequency domain dimension, and can better assist in identifying the detailed information of the card material (for example, highlighting that the card material causes the image to appear light spots under the action of the flash), so as to better help to achieve the identification of the authenticity of the card. In the case that the plurality of first card images include images acquired at different acquisition angles, the specified image can be an image acquired by other non-front acquisition mode in the plurality of first card images, so that the spectrum map obtained can reflect the features of the image from the frequency domain dimension, and better help to achieve the identification of the authenticity of the card.

[0117] Afterwards, in step 22, the server will form an image sequence with the multiple first card images and the spectrum image, and input the image sequence into the card identification model to obtain the authenticity identification result. In this step, the server forms an image sequence with the multiple first card images and the spectrum image. For example, the multiple first card images are arranged in the front of the image sequence, and the spectrum image is arranged in the back. The image sequence is input into the card identification model, so that the card identification model processes the image sequence to obtain the authenticity identification result. In some possible examples, the multiple first card images in the image sequence are arranged in the order of image acquisition according to the indication of the client.

[0118] In yet some examples, in order to improve the accuracy of the identification result, after the server receives the multiple first card images from the client, in step S340, the server can include the following steps 31-33:

[0119] In step 31, the server determines the difference images between the multiple first card images based on the multiple first card images. The specific implementation process of step 31 can refer to the specific implementation process of the aforementioned step 11, and will not be described here.

[0120] In step 32, the server determines the spectrum image corresponding to the multiple first card images based on the specified image in the multiple first card images. The specific implementation process of step 32 can refer to the specific implementation process of the aforementioned step 21, and will not be described here.

[0121] In step 33, the server forms an image sequence with the multiple first card images, the difference images, and the spectrum image, and inputs the image sequence into the card identification model to obtain the authenticity identification result. In this step, the server forms an image sequence with the multiple first card images, the difference images, and the spectrum image. For example, the multiple first card images are arranged in the front of the image sequence, the difference images are arranged after the multiple first card images, and the spectrum image is arranged at the end. Alternatively, the multiple first card images are arranged in the front of the image sequence, the spectrum image is arranged after the multiple first card images, and the spectrum image is arranged at the end of the difference images. Then, the server inputs the image sequence into the card identification model, so that the card identification model processes the image sequence to obtain the authenticity identification result. The sorting order between the multiple first card images and the sorting order between the difference images when the difference images are multiple can refer to the sorting order between the multiple first card images and the sorting order between the difference images when the difference images are multiple, which will not be described here.

[0122] In the above examples, the multiple first card images, the difference images, and the spectrum image are combined as the card identification model, so that the card identification model can combine the multi-dimensional features such as the spatial domain, the time domain, the gradient, and the frequency domain in the multiple first card images to determine the authenticity of the card to be identified corresponding to the multiple first card images, so that a more accurate identification result can be obtained.

[0123] In some possible examples, the card identification model includes a blocking layer, a mapping layer, an attention mechanism-based extraction layer, and a classification layer. Correspondingly, step S340 can include steps 41-45 as follows:

[0124] At step 41, the blocking layer is used to block each first card image in the plurality of first card images to obtain image blocks corresponding to each first card image. The implementation principle of step 41 is similar to that of step 011, and the implementation process can be referred to the implementation process of step 011, which will not be repeated here.

[0125] At step 42, the image blocks corresponding to each first card image are combined based on the inter-frame relationship between the first card images to obtain a plurality of image block groups, wherein each image block group includes one or more image blocks, and the plurality of image blocks in a single image block group have the same position in their respective first card images. The implementation principle of step 42 is similar to that of step 012, and the implementation process can be referred to the implementation process of step 012, which will not be repeated here.

[0126] At step 43, the mapping layer is used to process each image block group to obtain a mapping sequence. The implementation principle of step 43 is similar to that of step 013, and the implementation process can be referred to the implementation process of step 013, which will not be repeated here.

[0127] At step 44, the extraction layer is used to process the mapping sequence to obtain an image representation. The implementation principle of step 44 is similar to that of step 014, and the implementation process can be referred to the implementation process of step 014, which will not be repeated here.

[0128] At step 45, the classification layer is used to process the image representation to determine the authenticity identification result of the card to be identified. The implementation principle of step 45 is similar to that of step 015, and the implementation process can be referred to the implementation process of step 015, which will not be repeated here.

[0129] In an example, when the authenticity identification result indicates that the card to be identified is a true card, it is determined that the card to be identified is a true card. Subsequently, the server can continue to trigger the user identity authentication process using the card to be identified to provide subsequent business services for the user. In another example, when the authenticity identification result indicates that the card to be identified is a fake card (i.e., a false card), it is determined that the card to be identified is a false card, that is, the plurality of first card images are images injected by injection attack. Subsequently, the server can lock the client and perform subsequent specified security processes to protect the security of the account of the real user.

[0130] In some possible examples, in the process of training the card certificate identification model, the label data thereof includes data indicating whether the corresponding sample card certificate image group is a true card certificate, and data indicating a plurality of manners corresponding to a plurality of sample card certificate images in the corresponding sample card certificate image group. In this case, the card certificate identification model can determine not only the authenticity identification result of the card certificate to be identified, but also the plurality of manners corresponding to the plurality of first card certificate images (for example, whether the plurality of first card certificate images are collected at different angles, or whether the flash settings are different, or whether the collection angles and the flash settings are different).

[0131] In this example, in order to better improve the accuracy of the authenticity identification result of the card certificate to be identified, the client can also send the plurality of manners indicated by the user when collecting the images of the card certificate to be identified to the server. Correspondingly, the server can determine the final authenticity identification result of the card certificate to be identified based on the authenticity identification result of the card certificate to be identified determined by the card certificate identification model, and the plurality of manners corresponding to the plurality of first card certificate images determined by the card certificate identification model and the plurality of manners indicated by the user when collecting the images of the card certificate to be identified sent by the client.

[0132] For example, when the plurality of manners corresponding to the plurality of first card certificate images determined by the card certificate identification model and the plurality of manners indicated by the user when collecting the images of the card certificate to be identified sent by the client are the same, it is considered that the authenticity identification result of the card certificate to be identified determined by the card certificate identification model is accurate.

[0133] If the plurality of manners corresponding to the plurality of first card certificate images determined by the card certificate identification model and the plurality of manners indicated by the user when collecting the images of the card certificate to be identified sent by the client are different, if the authenticity identification result of the card certificate to be identified determined by the card certificate identification model indicates that the card certificate to be identified is a fake card certificate, it is determined that the result is accurate. If the authenticity identification result of the card certificate to be identified determined by the card certificate identification model indicates that the card certificate to be identified is a true card certificate, it can be determined that the result is doubtful. In order to better protect the security of the user account, the server can feed back information that the user needs to re-upload the images of the card certificate to the client, so that the client instructs the user to collect the images of the card certificate in a plurality of manners.

[0134] In this embodiment, considering that there is generally no change between the multiple injected images injected by the injection attack mode, when the client collects images of the card to be identified, the user is instructed to collect images of the card to be identified in multiple ways, so that the multiple images of the card obtained for the real card are different images. Subsequently, after the server obtains the multiple images of the card to be identified, the trained card identification model is used to determine whether the multiple images are collected in multiple ways, to determine whether the images collected by the client are real images collected, and then determine the authenticity identification result of the card to be identified, to identify the authenticity of the card.

[0135] FIG. 4 shows a flowchart of a card identification method in another embodiment of the present specification. The method is applied to a client and a server in a card identification system, and the server and the client can run on different electronic devices. Any electronic device can be realized by any device, equipment, platform, equipment cluster, etc. with computing and processing capabilities. In the card identification process, as shown in FIG. 4, the method comprises the following steps S410-S450:

[0136] In step S410, the client acquires multiple first card images of the card to be identified in the process of instructing the user to collect images of the card to be identified in multiple ways, wherein the multiple ways enable the client to collect different images of the card to be identified.

[0137] In step S420, the client sends the multiple first card images to the server. The implementation principle of steps S410-S420 is similar to that of the aforementioned steps S310-S320, and the implementation process can be referred to the implementation process of the aforementioned steps S310-S320, which will not be repeated here.

[0138] In step S430, the server receives the multiple first card images from the client. The implementation principle of step S430 is similar to that of the aforementioned step S330, and the implementation process can be referred to the implementation process of the aforementioned step S330, which will not be repeated here.

[0139] In step S440, the server detects whether the image contents of the multiple first card images are consistent by using a preset content recognition algorithm.

[0140] In some possible examples, in order to better confuse the server in identifying the authenticity of the card, the malicious party may inject different fake card images when the client instructs the user to collect the images of the card in multiple ways, so that there are differences between the multiple frames of fake card images injected by the injection attack. Considering that the real card images collected by the client in multiple ways are collected from the same card, the image contents (i.e., the contents in the card) of the corresponding multiple frames of real collected card images are consistent. In view of the above, after receiving the multiple frames of first card images, the server can perform consistency checking on the image contents of the multiple frames of first card images. Specifically, the preset content recognition algorithm is used to detect whether the image contents of the multiple frames of first card images are consistent, to obtain a detection result; and then a subsequent process is performed based on the detection result.

[0141] For example, the foregoing preset content recognition algorithm can include, but is not limited to, an OCR (Optical Character Recognition) algorithm, a neural network-based content recognition algorithm, and any algorithm that can recognize image text and / or non-text content in related technologies. In some other possible examples, the server can also jointly use a preset content recognition algorithm (such as an OCR algorithm) and an angle checking and classification algorithm for detecting the tilt angle of a target in an image to jointly detect whether the image contents of the multiple frames of first card images are consistent.

[0142] The server can recognize the image contents of the first card images based on the preset content recognition algorithm, and then determine the similarity values between the image contents of the first card images based on the preset similarity algorithm and the image contents of the first card images. Then, based on the similarity values between the image contents of the first card images and a certain threshold, a detection result indicating whether the image contents of the multiple frames of first card images are consistent is determined. If the similarity values between the image contents of the multiple frames of first card images exceed a certain threshold, it is determined that the detection result indicates that the image contents of the multiple frames of first card images are consistent. If the similarity values between the image contents of the multiple frames of first card images do not exceed a certain threshold, it is determined that the detection result indicates that the image contents of the multiple frames of first card images are inconsistent.

[0143] In step S450, if the detection result indicates that the image contents of the multiple frames of first card images are consistent, the server determines whether the multiple frames of first card images are collected in the foregoing multiple ways by using the card identification model, to determine the authenticity identification result of the card to be identified. The implementation principle of step S450 is similar to that of the foregoing step S340, and the implementation process can be referred to the implementation process of the foregoing step S340, which will not be described herein.

[0144] In some examples, if the detection result indicates that the image contents of the multiple first card image frames are inconsistent, the method can further include step S460, and the server determines that the card to be identified is a fake card. If the detection result indicates that the image contents of the multiple first card image frames are inconsistent, the server can consider that the multiple first card image frames are different fake card image frames injected by a malicious party in order to confuse the server, and then determine that the card to be identified is a fake card. In this way, the authenticity of the card can be identified, and the accuracy of the identification result of the authenticity of the card can be improved.

[0145] Corresponding to the method embodiments, the specification embodiments provide a card identification method applied to a server in a card identification system. The card identification system further includes a client. The server and the client can run in different electronic devices. Any electronic device can be implemented by any device, apparatus, platform, device cluster, etc. that has computing and processing capabilities. As shown in FIG. 5, in a card identification process, the method includes the following steps S510-S520:

[0146] In step S510, multiple first card image frames of a card to be identified are received from a client. The client is used to instruct a user to collect images of the card to be identified in multiple ways. The multiple ways enable the client to collect different images of the card to be identified. The implementation principle of step S510 is similar to that of the foregoing step S330. For details, refer to the implementation process of the foregoing step S330, which will not be described here.

[0147] In step S520, whether the multiple first card image frames are collected in the multiple ways is determined by using a trained card identification model, to determine a true-false identification result of the card to be identified. The implementation principle of step S520 is similar to that of the foregoing step S340. For details, refer to the implementation process of the foregoing step S340, which will not be described here.

[0148] In some possible implementation manners, the multiple ways include a collection way of collecting images at different collection angles and / or a collection way of collecting images in a flash light setting and a no flash light setting.

[0149] In some possible implementation manners, in step S520, the following steps are included: based on the multiple first card image frames, difference images between the multiple first card image frames are determined; the multiple first card image frames and the difference images are combined to form an image sequence, and the image sequence is input into the card identification model to obtain the true-false identification result.

[0150] In some possible implementation manners, at step S520, the method further includes: determining a frequency spectrum corresponding to the plurality of first card images based on a specified image in the plurality of first card images; grouping the plurality of first card images and the frequency spectrum into an image sequence, and inputting the image sequence into the card identification model to obtain the authenticity identification result.

[0151] In some possible implementation manners, at step S520, the method further includes: grouping the plurality of first card images into an image sequence, and inputting the image sequence into the card identification model to obtain the authenticity identification result.

[0152] In some possible implementation manners, at step S520, the method further includes: detecting, by using a preset content recognition algorithm, whether image contents of the plurality of first card images are consistent; and if the detection result indicates that the image contents of the plurality of first card images are consistent, determining, by using the card identification model, whether the plurality of first card images are collected in the plurality of manners, to determine the authenticity identification result of the card to be identified.

[0153] In some possible implementation manners, at step S520, the method further includes: if the detection result indicates that the image contents of the plurality of card images are inconsistent, determining that the card to be identified is a fake card.

[0154] In some possible implementation manners, the card identification model is a network model based on a transformer.

[0155] In some possible implementation manners, the card identification model includes a blocking layer, a mapping layer, an extraction layer based on an attention mechanism, and a classification layer; and at step S520, the method further includes: grouping, by using the blocking layer, each first card image in the plurality of first card images to obtain an image block corresponding to each first card image.

[0156] combining, by using an inter-frame relationship between the first card images, the image blocks corresponding to the first card images to obtain a plurality of image block groups, wherein each image block group includes one or more image blocks, and the plurality of image blocks in a single image block group are located at the same position in the respective first card images;

[0157] processing, by using the mapping layer, each image block group to obtain a mapping sequence;

[0158] processing, by using the extraction layer, the mapping sequence to obtain an image representation;

[0159] processing, by using the classification layer, the image representation to determine the authenticity identification result of the card to be identified.

[0160] In some possible implementation manners, the training data for training the card identification model comprises a plurality of sample image groups and corresponding label data, and each sample image group at least comprises a sample card image obtained by capturing a sample card in the plurality of manners.

[0161] In some optional implementation manners, each sample image group further comprises a first sample differential image and / or a first sample spectrum image generated based on the plurality of first sample card images.

[0162] Corresponding to the method embodiments, the card identification method provided by the embodiments of the present disclosure is applied to a client in a card identification system, and the card identification system further comprises a server. The server and the client can run in different electronic devices, and any electronic device can be implemented by any device, apparatus, platform or device cluster having computing and processing capabilities. As shown in FIG. 6, in the card identification process, the method comprises the following steps S610-S620:

[0163] In step S610, in the process of instructing a user to capture images of a card to be identified in a plurality of manners, a plurality of first card images of the card to be identified are obtained, wherein the plurality of manners enable the client to capture different images of the card to be identified. The implementation principle of step S610 is similar to that of the aforementioned step S310, and the implementation process can refer to the implementation process of the aforementioned step S310, which will not be described here.

[0164] In step S620, the plurality of first card images are sent to the server, so that after the server receives the plurality of first card images, the trained card identification model is used to determine whether the plurality of first card images are captured in the plurality of manners, to determine the authenticity identification result of the card to be identified. The implementation principle of step S620 is similar to that of the aforementioned step S320, and the implementation process can refer to the implementation process of the aforementioned step S320, which will not be described here.

[0165] The above describes specific embodiments of the present disclosure, and other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in a different order than the order in which they are recited, and still achieve desirable results. In addition, the processes depicted in the figures do not necessarily have to be performed in the specific order shown, or sequentially, to achieve desirable results. In certain implementations, multitasking and parallel processing can be advantageous, or can be desirable.

[0166] Corresponding to the above method embodiments, the specification embodiments provide a card identification device 700, which is deployed on a server, and a schematic block diagram thereof is shown in FIG. 7, including: a receiving module 710, configured to receive, from a client, a plurality of first card images of a card to be identified, wherein the client is used to instruct a user to collect images of the card to be identified in a plurality of manners, and the plurality of manners enable the client to collect different images of the card to be identified; and a first determining module 720, configured to determine, by using a trained card identification model, whether the plurality of first card images are collected in the plurality of manners, to determine a true or false identification result of the card to be identified.

[0167] In some optional embodiments, the plurality of manners include: a collection manner of collecting images at different collection angles, and / or a collection manner of collecting images in a flash light setting and a no flash light setting.

[0168] In some optional embodiments, the first determining module 720 is specifically configured to determine, based on the plurality of first card images, difference images between the plurality of first card images.

[0169] The plurality of first card images and the difference images are combined into an image sequence, and the image sequence is input into the card identification model to obtain the true or false identification result.

[0170] In some optional embodiments, the first determining module 720 is specifically configured to determine, based on a specified image in the plurality of first card images, a frequency spectrum corresponding to the plurality of first card images.

[0171] The plurality of first card images and the frequency spectrum are combined into an image sequence, and the image sequence is input into the card identification model to obtain the true or false identification result.

[0172] In some optional embodiments, the first determining module 720 is specifically configured to combine the plurality of first card images into an image sequence, and input the image sequence into the card identification model to obtain the true or false identification result.

[0173] In some optional embodiments, the first determining module 720 is specifically configured to detect, by using a preset content recognition algorithm, whether image contents of the plurality of first card images are consistent.

[0174] If the detection result indicates that the image contents of the plurality of first card images are consistent, the card identification model is used to determine whether the plurality of first card images are collected in the plurality of manners, to determine the true or false identification result of the card to be identified.

[0175] In some optional embodiments, the first determining module 720 is further configured to determine that the card to be identified is a fake card if the detection result indicates that the image contents of the multiple frames of card image are inconsistent.

[0176] In some optional embodiments, the card identification model is a transformer-based network model.

[0177] In some optional embodiments, the card identification model comprises a block division layer, a mapping layer, an attention mechanism-based extraction layer, and a classification layer.

[0178] The first determining module 720 is specifically configured to divide each first card image in the multiple frames of first card image by using the block division layer to obtain image blocks corresponding to each first card image.

[0179] Each image block corresponding to each first card image is combined by using the inter-frame relationship between each first card image to obtain multiple groups of image block groups, wherein each group of image block groups comprises one or more image blocks, and the multiple image blocks in a single group of image block groups have the same position in their respective first card images.

[0180] The mapping layer is used to process each group of image block groups to obtain a mapping sequence.

[0181] The extraction layer is used to process the mapping sequence to obtain an image representation.

[0182] The classification layer is used to process the image representation to determine the authenticity identification result for the card to be identified.

[0183] In some optional embodiments, the training data for training the card identification model comprises a plurality of sample image groups and corresponding label data, and each sample image group at least comprises sample card images obtained by collecting sample cards in the multiple manners.

[0184] In some optional embodiments, each sample image group further comprises a first sample differential image and / or a first sample spectrum image generated based on the multiple frames of first sample card image.

[0185] Corresponding to the above method embodiments, the specification embodiments provide a card identification device 800, which is deployed at a client, and a schematic block diagram of the device is shown in FIG. 8, and the device includes: an indication module 810, configured to acquire a plurality of first card images of a card to be identified in a process of instructing a user to collect images of the card to be identified in a plurality of manners, wherein the plurality of manners enable the client to collect different images of the card to be identified; and a sending module 820, configured to send the plurality of first card images to a server, so that, after the server receives the plurality of first card images, the server determines whether the plurality of first card images are collected in the plurality of manners by using a trained card identification model, to determine a true or false identification result for the card to be identified.

[0186] The device embodiments correspond to the method embodiments, and specific descriptions can be referred to the descriptions of the method embodiments, which will not be repeated here. The device embodiments are based on the corresponding method embodiments, and have the same technical effects as the corresponding method embodiments. Specific descriptions can be referred to the corresponding method embodiments.

[0187] The specification embodiments further provide a computer-readable storage medium, which stores a computer program, and when the computer program is executed in a computer, the computer program causes the computer to execute the card identification method provided in the specification.

[0188] The specification embodiments further provide a computing device, which includes a memory and a processor, the memory stores executable code, and when the processor executes the executable code, the card identification method provided in the specification is implemented.

[0189] Each of the embodiments in the specification is described in a progressive manner, and the same or similar parts of each of the embodiments can be referred to each other. Each of the embodiments mainly describes differences from other embodiments. In particular, the storage medium and the computing device embodiments are described more simply, and the relevant parts can be referred to the descriptions of the method embodiments.

[0190] Those skilled in the art should realize that, in one or more of the examples described above, the functions described in the embodiments of the present application can be implemented by using hardware, software, firmware or any combination thereof. When implemented by using software, the functions can be stored in a computer readable medium or transmitted as one or more instructions or codes on a computer readable medium.

[0191] The above detailed description of the specific implementation has further detailed the purposes, technical solutions and beneficial effects of the embodiments of the present application. It should be understood that the above description is only a specific implementation of the embodiments of the present application and is not used to limit the protection scope of the present application. Any modification, equivalent replacement, improvement, etc. made on the basis of the technical solutions of the present application shall be included in the protection scope of the present application.

Claims

1. A card identification method applied to a server, the method comprising: receiving, from a client, a plurality of first card images of a card to be identified, wherein the client is configured to instruct a user to capture images of the card to be identified in a plurality of manners, and the plurality of manners enable the client to capture different images of the card to be identified; determining, by using a trained card identification model, whether the plurality of first card images are captured in the plurality of manners, to determine a true or false identification result of the card to be identified.

2. The method of claim 1, wherein, The plurality of manners include a capturing manner of capturing images at different capturing angles, and / or a capturing manner of capturing images in a flash light setting and a non-flash light setting.

3. The method of claim 1, wherein, The determining of the true or false identification result of the card to be identified comprises: determining, based on the plurality of first card images, difference images between the plurality of first card images; combining the plurality of first card images and the difference images into an image sequence, and inputting the image sequence into the card identification model to obtain the true or false identification result.

4. The method of claim 1, wherein, The determining of the true or false identification result of the card to be identified comprises: determining, based on a specified image in the plurality of first card images, a frequency spectrum corresponding to the plurality of first card images; combining the plurality of first card images and the frequency spectrum into an image sequence, and inputting the image sequence into the card identification model to obtain the true or false identification result.

5. The method of claim 1, wherein, The determining of the true or false identification result of the card to be identified comprises: combining the plurality of first card images into an image sequence, and inputting the image sequence into the card identification model to obtain the true or false identification result.

6. The method of claim 1, wherein, The determining of the true or false identification result of the card to be identified comprises: detecting, by using a preset content recognition algorithm, whether image contents of the plurality of first card images are consistent; if the detection result indicates that the image contents of the plurality of first card images are consistent, determining, by using the card identification model, whether the plurality of first card images are captured in the plurality of manners, to determine the true or false identification result of the card to be identified.

7. The method of claim 6, further comprising: if the detection result indicates that the image contents of the plurality of card images are inconsistent, determining that the card to be identified is a fake card.

8. The method of any one of claims 1-7, wherein, The card identification model is a transformer-based network model.

9. The method of any one of claims 1-7, wherein, The card identification model comprises a block division layer, a mapping layer, an attention mechanism-based extraction layer, and a classification layer. The determining, by using the trained card identification model, whether the plurality of first card images are captured in the plurality of manners, to determine the true or false identification result of the card to be identified, comprises: dividing each first card image in the plurality of first card images into image blocks by using the block division layer; combining the image blocks corresponding to each first card image by using inter-frame relationships between the first card images, to obtain a plurality of image block groups, wherein each image block group comprises one or more image blocks, and the plurality of image blocks in a single image block group are located at the same position in their respective first card images; processing each image block group by using the mapping layer, to obtain a mapping sequence; The mapping sequence is processed by using the extraction layer to obtain an image representation; The image representation is processed by using the classification layer to determine the authenticity identification result of the card to be identified.

10. The method of any one of claims 1-7, wherein, The training data of the card identification model includes a plurality of sample image groups and corresponding label data. Each sample image group at least includes a sample card image collected in the plurality of manners.

11. A card identification method applied to a client, the method comprising: In the process of instructing a user to collect images of a card to be identified in a plurality of manners, a plurality of first card images of the card to be identified are acquired, wherein the plurality of manners enable the client to collect different images of the card to be identified; The plurality of first card images are sent to a server, so that after the server receives the plurality of first card images, a trained card identification model is used to determine whether the plurality of first card images are collected in the plurality of manners, to determine the authenticity identification result of the card to be identified.

12. A card identification device deployed on a server, comprising: A receiving module configured to receive a plurality of first card images of a card to be identified from a client, wherein the client is used to instruct a user to collect images of a card to be identified in a plurality of manners, and the plurality of manners enable the client to collect different images of the card to be identified; A first determining module configured to use a trained card identification model to determine whether the plurality of first card images are collected in the plurality of manners, to determine the authenticity identification result of the card to be identified.

13. A card identification device deployed on a client, comprising: An acquiring module configured to acquire a plurality of first card images of a card to be identified in the process of instructing a user to collect images of the card to be identified in a plurality of manners, wherein the plurality of manners enable the client to collect different images of the card to be identified; A sending module configured to send the plurality of first card images to a server, so that after the server receives the plurality of first card images, a trained card identification model is used to determine whether the plurality of first card images are collected in the plurality of manners, to determine the authenticity identification result of the card to be identified.

14. A computing device comprising a memory and a processor, wherein, The memory stores executable code, and the processor executes the executable code to implement the method of any one of claims 1-10 or 11.

Citation Information

Patent Citations

  • Certificate authenticity identification method and device

    CN111324874A

  • Certificate authentic identification method and device, electronic equipment and storage medium

    CN111898538A

  • Method and device for identifying authenticity of signature image

    CN112836636A

  • Video-Processing Method, Electronic Device, and Computer-Readable Storage Medium

    US20210168441A1