A method of identifying a credential and a computing device
Patent Information
- Application Number
- CN202610894918.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-18
- Publication Date
- 2026-09-04
AI Technical Summary
[0006]为解决客户端采集的证件图像质量不佳,以及图像增强处理与防伪检测之间的矛盾问题,本说明书实施例提供一种基于客户端超分合成与双轨验证的eKYC证件识别方法及计算设备
[0017] This embodiment actively fuses effective information from multiple frames of images using a synthesis model, resulting in images that are clearer and have higher resolution than any original single frame, fundamentally breaking through the image quality ceiling of single-frame acquisition. This embodiment places the most computationally intensive super-resolution synthesis task on the client side; the client only needs to upload the optimized synthesized image and the target image. This significantly reduces network bandwidth consumption and user waiting time, alleviating the computational pressure and cost on the server side. This embodiment employs a dual-track upload strategy, uploading a high-quality "super-resolution synthesized image" and an "original real image" that retains the original image features to the server separately. On the server side, OCR recognition is performed based on the synthesized image, and anti-counterfeiting detection is performed based on the target image. These two processes perform their respective functions in parallel, fundamentally resolving the contradiction between recognition and anti-counterfeiting, and achieving a significant improvement in recognition rate and security.
Smart Images

Figure CN122695641A_ABST
Abstract
Description
Technical Field
[0001] The embodiments in this specification belong to the field of identity verification technology, and in particular relate to a method and computing device for identifying identification documents. Background Technology
[0002] With the acceleration of global digitalization, remote electronic identity authentication (electronic Know Your Customer, eKYC) has become the infrastructure of the digital economy and is widely used in scenarios such as finance and the Internet.
[0003] In the traditional eKYC process, users need to use a mobile device to photograph their ID card on the client side and upload the image to the server. The server then performs Optical Character Recognition (OCR), information verification, and anti-counterfeiting detection on the image. The main problem with this eKYC process is that the image quality captured by the client directly determines the success rate and security of the entire eKYC process. Furthermore, due to uncontrollable factors such as the user's shooting environment (e.g., uneven lighting, camera shake) and device hardware (e.g., differences in camera quality), single-frame ID card images captured by the client often suffer from problems such as blurriness, glare, and low resolution, which are difficult to avoid.
[0004] To improve the quality of image acquisition on the client side, some researchers have proposed a scheme to select the "optimal frame," which involves the client acquiring multiple frames of images and selecting the "optimal frame" with the highest image quality for document recognition. However, the "optimal frame" is relative. Under poor shooting conditions, the so-called "optimal frame" may still be of poor quality and cannot meet the requirements of high-precision OCR, thus not fundamentally solving the image quality bottleneck problem.
[0005] In addition, researchers have proposed a server-side image enhancement or super-resolution synthesis scheme. This scheme uploads multiple frames of images captured by the client to the server, where image enhancement or super-resolution synthesis is performed based on these frames to obtain a higher-quality document image for document recognition. While this scheme can guarantee image quality, image enhancement or super-resolution synthesis essentially involves lossy modification or regeneration of image pixel values to make the image clearer. During this process, to smooth the image and remove interference, subtle and random original physical features in the image (such as sensor noise PRNU, moiré patterns) are treated as noise and erased. These original physical features are important criteria for determining the authenticity of documents in anti-counterfeiting detection. Image enhancement or super-resolution synthesis also introduces processing traces into the image (such as color banding, ringing effects, frequency domain anomalies, local texture anomalies, etc.). These processing traces are highly similar to those left by malicious attackers generating or tampering with images. During anti-counterfeiting detection, it is difficult to determine whether the processing traces in the image are caused by image enhancement or super-resolution synthesis or left by malicious attackers when forging the image. This greatly increases the difficulty of anti-counterfeiting detection, making it difficult to obtain accurate results. Furthermore, all the complex calculations in this solution (such as super-resolution synthesis and image enhancement) are completed on the server side. The client needs to upload all the original video images it has acquired to the server, which greatly increases network bandwidth consumption and user waiting time. It also places high demands on the server's computing power, resulting in high deployment and operation costs. Summary of the Invention
[0006] To address the issue of poor image quality of documents captured by the client and the conflict between image enhancement processing and anti-counterfeiting detection, this specification provides an eKYC document recognition method and computing device based on client-side super-resolution synthesis and dual-track verification.
[0007] The first aspect of this specification provides a method for identifying identification documents, executed by a client, including:
[0008] Capture multiple frames of images containing identification documents;
[0009] A composite image is generated based on the multi-frame images using a synthesis model, and the resolution of the composite image is higher than that of the multi-frame images.
[0010] Select several target images from the multiple frames of images;
[0011] The composite image and the target image are sent to the server for document recognition, and the composite image is used for character recognition.
[0012] The second aspect of this specification provides a method for identifying identification documents, executed by a server, including:
[0013] The system receives a composite image and a target image from the client; the composite image is synthesized based on multiple frames of images containing the document, and the resolution of the composite image is higher than that of the multiple frames; the target image is selected from the multiple frames.
[0014] Perform optical character recognition (OCR) on the synthesized image; perform anti-counterfeiting detection on the target image;
[0015] The document recognition result is obtained based on the OCR recognition result and the anti-counterfeiting detection result.
[0016] A third aspect of this specification provides a computing device including a memory and a processor, wherein the memory stores executable code, and the processor, when executing the executable code, implements the method described in the first aspect.
[0017] This embodiment actively fuses effective information from multiple frames of images using a synthesis model, resulting in images that are clearer and have higher resolution than any original single frame, fundamentally breaking through the image quality ceiling of single-frame acquisition. This embodiment places the most computationally intensive super-resolution synthesis task on the client side; the client only needs to upload the optimized synthesized image and the target image. This significantly reduces network bandwidth consumption and user waiting time, alleviating the computational pressure and cost on the server side. This embodiment employs a dual-track upload strategy, uploading a high-quality "super-resolution synthesized image" and an "original real image" that retains the original image features to the server separately. On the server side, OCR recognition is performed based on the synthesized image, and anti-counterfeiting detection is performed based on the target image. These two processes perform their respective functions in parallel, fundamentally resolving the contradiction between recognition and anti-counterfeiting, and achieving a significant improvement in recognition rate and security. Attached Figure Description
[0018] To more clearly illustrate the technical solutions of the embodiments in this specification, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this specification. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0019] Figure 1 This is a multi-terminal interaction diagram of a document recognition method in one embodiment of this specification;
[0020] Figure 2 This is an architecture diagram of a document recognition method in one embodiment of this specification;
[0021] Figure 3 This is an architecture diagram of a video super-resolution model in one embodiment of this specification;
[0022] Figure 4This is a schematic diagram of a local feature fusion module in one embodiment of this specification;
[0023] Figure 5 This is a schematic diagram of a cyclic bidirectional propagation module in one embodiment of this specification. Detailed Implementation
[0024] To enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this specification, and not all embodiments. Based on the embodiments in this specification, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of this specification.
[0025] In this manual, remote electronic Know Your Customer (eKYC) refers to the process of verifying a customer's true identity without face-to-face interaction through remote information collection and identity verification using digital technology. This process typically includes collecting image information of the user's ID document, recognizing the document content, and verifying the document's authenticity. It aims to break down the physical limitations of traditional offline branches, achieving convenience and automation in business transactions while ensuring the accuracy, security, and compliance of identity verification. It is widely used in scenarios requiring strict identity verification, such as financial account opening, digital government services, and e-commerce.
[0026] In this manual, Optical Character Recognition (OCR) refers to the technology that uses optical and computer technologies to convert text in an image into editable text that a computer can understand. In the eKYC process, OCR technology can be used to identify specific content in documents (such as customer name, document number, identity information, etc.), thereby cross-checking the document content information with an authoritative identity database to achieve remote identity verification.
[0027] In this specification, Video Super-Resolution (VSR) is a technique that generates high-resolution images using information from multiple video frames. Its core advantage lies in its ability to fuse complementary detail information from multiple video frames, thereby reconstructing an image with clearer texture and richer high-frequency details than a single frame. This technology aims to overcome resolution bottlenecks caused by hardware limitations of imaging devices or bandwidth compression, thereby improving image clarity and visual quality.
[0028] Figure 1 and Figure 2The flowchart and architecture of the eKYC document recognition method based on client-side super-resolution synthesis and dual-track verification provided in the embodiments of this specification are shown respectively. It is understood that this method can be executed by any device, equipment, platform, or device cluster with computing and processing capabilities.
[0029] like Figure 1 and Figure 2 As shown, in step S101, multiple frames of images containing the document are captured.
[0030] In one implementation, step S101 specifically includes:
[0031] S1011, Take a picture of the ID card and obtain a video stream containing the ID card;
[0032] In the eKYC process, the user points the camera at their ID card, and the client receives a video stream containing the ID card in real time through the camera, capturing each frame of the video stream in real time.
[0033] S1012, obtain the location of the document in each frame of the image, and track the location of the document in real time during the video capture process;
[0034] The document location includes the coordinates of the four corners of the document in each frame of the image, the border positions of the four sides of the document, and the coordinates of the center point of the document.
[0035] The target detection algorithm is used to detect the target in each frame of the video stream to obtain the location of the document in the image. Then, the image features corresponding to the document location are extracted. The extracted image features are then input into the tracking algorithm network. The tracking algorithm associates the document locations in different video frames, thereby realizing real-time tracking of the document location and motion trajectory.
[0036] In one implementation, the document positions of different video frames are associated, specifically including: expanding a certain size around the document position in the current frame to obtain a candidate box; capturing the document position of the next frame within the candidate box based on image features; and recording the differences and associations of document positions in previous and subsequent frames to obtain the document motion trajectory; after obtaining the document motion trajectory, the document position of the next frame can be predicted based on the existing motion trajectory; and the position of the candidate box can be adjusted based on the prediction result to improve the tracking success rate.
[0037] In one implementation, the object detection algorithm may be R-CNN, YOLO, RetinaNet, etc.
[0038] In one implementation, the tracking algorithm may be SORT.
[0039] In other embodiments, joint tracking and detection networks such as JDE, RetinaTrack, and CenterTrack can be used to simultaneously perform target detection and target tracking of the document.
[0040] Real-time tracking of document position during video capture can improve image quality. Specifically, the camera can adjust its focusing strategy based on the document position tracking results, making the image of the document clearer; the client can guide the user to place the document in the prescribed direction, position, and angle based on the document position tracking results, thereby obtaining images of the document placed in the prescribed direction, position, and angle, improving the comprehensiveness of the image material.
[0041] S1013, Filter images in the video stream based on image quality;
[0042] Each captured video frame undergoes image quality inspection. Based on the inspection results, severely blurry, overly dark, or overexposed frames are removed to obtain high-quality frames. For example, image quality inspection can be performed using no-reference image quality assessment algorithms such as BRISQUE, NIQE, and BIQI. The decision to retain a frame is based on the image quality score output by the algorithm.
[0043] S1014, identify the document angle in each remaining frame image, and retain multiple frame images under different document angles based on the identification results;
[0044] After the image quality inspection is completed, for each remaining frame of the image, the angle of the document (including pitch angle, plane angle, etc.) is identified with the document as the center, and multiple frames of video images under different document angles are retained based on the identification results.
[0045] For example, based on the recognition results, multiple frames of video images from different document angles are retained. Specifically, this includes: classifying the video images according to the range of document angles (e.g., the first category is where both the pitch angle and the plane angle are close to 0°, and the second category is where the pitch angle is close to -5° and the plane angle is close to 0°); then, setting an image acquisition index for each category of document angle range (e.g., acquiring 10 frames for the first category and 5 frames for the second category); when all images from each category of document angle range in the retained images reach the acquisition index, the acquisition is considered complete, thus obtaining multiple high-quality multi-frame video images from different angles.
[0046] By acquiring video images from multiple angles, information loss due to local blurring or reflections in the video images can be avoided, thereby improving the quality of the super-resolution composite image and the original optimal frame. Specifically, during video acquisition, uneven lighting and camera shake are unavoidable when the user is holding the shooting device, making it difficult to ensure that the image quality of every location on the document is the same. Even images with high average quality will have blurry or reflective areas, and these blurry or reflective areas often change position as the document's angle changes. Based on this, the embodiments in this specification acquire images of the document from different angles and perform super-resolution composites on them, so that the blurry areas of the images from different angles complement each other, thereby improving the clarity of the super-resolution composite image.
[0047] S1015, Align the document positions of all frames in a multi-frame image.
[0048] By reconstructing the images, the document positions in all frames across multiple images are aligned to the same coordinate system. Specifically, one frame is selected from the retained multiple images as a reference frame, and the coordinate information of the four corners of the document in the reference frame is recorded. Each frame other than the reference frame is then reconstructed so that the coordinates of the four corners of the document in each frame are the same as those in the reference frame, thus achieving document position alignment.
[0049] In one implementation, aligning the document positions across all frames in the multi-frame image is accomplished using a geometric correction method based on perspective transformation. This method specifically includes: first, using a trained keypoint detection model (e.g., Hourglass Network or an improved YOLO Pose branch) to detect the coordinates P1, P2, P3, and P4 of the four vertices of the document in each frame; then, selecting the frame with the highest quality score as the reference frame and recording its four vertex coordinates as the target coordinates. Then, for each other frame, based on its four vertex coordinates... Target coordinates of the reference frame The perspective transformation matrix H of 3x3 is solved according to formula (1).
[0050] (1)
[0051] Where w' is the perspective depth factor.
[0052] Next, using bilinear interpolation or grid sampling operators, the current frame image is warped to the coordinate system of the reference frame based on matrix H.
[0053] If the four points detected are collinear or the area of the quadrilateral formed is abnormal (such as too small or too large), the frame is determined to be misaligned. For frames that fail to align, other methods (such as manual alignment) can be used to re-align them, or frames that fail to align can be discarded directly and will not participate in subsequent super-resolution fusion.
[0054] The collection, storage, use, processing, transmission, provision, and disclosure of video frames and video sequences mentioned in this specification are performed under authorization and comply with relevant laws and regulations.
[0055] The above describes the acquisition and preprocessing of multi-frame video images on the client side. The following describes the multi-frame super-resolution synthesis process.
[0056] Next, in step S102, a composite image is generated based on the multi-frame images using a synthesis model, wherein the resolution of the composite image is higher than that of the multi-frame images.
[0057] The synthesis model can be a VSR model, used to generate high-resolution super-resolution synthesized images based on multiple frames of images.
[0058] Figure 3 This specification illustrates an algorithmic framework for a lightweight VSR model provided in an embodiment. The VSR model is deployed on a client-side, and the core computational tasks for multi-frame super-resolution synthesis are executed on the mobile client. Figure 3 As shown, the VSR model includes a feature extraction module, a Local Feature Matching (LFM) module, a cyclic bidirectional propagation module, and a reconstruction module, which are connected sequentially. The feature extraction module is used to extract features from the input multi-frame video images to obtain the initial features g of each frame. i The LFM module is used to process the initial features g of each frame of image. i The initial feature g of its fixed number of neighboring frames i+1 g i-1 Perform feature fusion to obtain the fused features. The cyclic bidirectional propagation module is used to propagate the fused features bidirectionally to obtain reconstructed features; the reconstruction module is used to reconstruct the image based on the reconstructed features to obtain a super-resolution composite image.
[0059] In one implementation, the feature extraction module is built on a Transformer network structure. The Transformer's self-attention mechanism can capture the correlation between global pixels, enabling the model to better understand the overall layout of the image.
[0060] In one implementation, such as Figure 4As shown, the local feature fusion module includes an alignment submodule (AlignModule) and a fusion submodule. The alignment submodule is used to integrate the target frame g. i Its neighboring frame g i+1 g i-1 The features are aligned, and the fusion submodule fuses the aligned features using the concat function and residual blocks to obtain the fused features. .
[0061] Before feature propagation, local feature fusion is performed first, allowing the features of the current frame to be fused with information from its neighboring frames. The fused features are then passed to the next stage of the cyclic bidirectional propagation module, which can strengthen cross-frame feature fusion in feature propagation and improve the feature fusion effect.
[0062] In one implementation, the fusion submodule of the local feature fusion module is constructed based on a Transformer network structure, and the local feature fusion is performed based on the Transformer's cross-attention mechanism. The Transformer's cross-attention mechanism has a global perspective, and can simultaneously capture the correlation between global pixels of the source modality during fusion, thereby improving the feature fusion effect by utilizing global pixel relationships.
[0063] In one embodiment of this implementation, the alignment submodule aligns the features of neighboring frames based on the optical flow estimation method, reducing the search space so that the Transformer only needs to process the residual displacement, further reducing the computational complexity.
[0064] In one embodiment of this implementation, in the local feature fusion module, for the features of the current frame at time step t... It is mapped to a query vector through a linear projection layer. .Right now ,in This is a learnable weight matrix.
[0065] Select the feature F of each of the N frames before and after the current frame (e.g., N=2, i.e., t-2, t-1, t+1, t+2). t-i and F t+j Each is mapped to a key vector through a linear projection layer. Sum value vector ,Right now , .
[0066] calculate and The dot product similarity, after Softmax normalization, is then compared with... Weighted summation yields the fused features. As shown in equation (2).
[0067] (2)
[0068] In one embodiment, the local feature fusion module further includes a dynamic frame selection submodule, which is used to perform text line quality detection on the video frames to be fused before local feature fusion, obtain the text line quality score of each frame image, and assign a fusion weight score to each frame image according to the text line quality score of each frame image during feature fusion, so that the weight of video frames with low text line quality scores (such as blurred or severely occluded frames) is reduced in subsequent fusion, or even directly discarded.
[0069] The text line quality detection specifically includes: obtaining the position of each text line in the image using a text detection algorithm (such as the Dbnet text line algorithm), where the position of the text line can be represented by a text line bounding box; and obtaining the label (e.g., "Name", "ID_Number", "Issuing Authority", "Validity", etc.) and quality score S for each text line based on its position. line The quality score S for each line of text line The weighted summation is performed to obtain the text line quality score Q of the image frame, where the weight of each text line is determined according to the importance of the text line label.
[0070] The label for each text line is obtained based on its position. Specifically, this includes classifying the detected text lines into different semantic labels based on the text line position and prior knowledge of the document layout. The prior knowledge of the document layout includes the text line position and the size of the text line bounding box corresponding to each type of text line label. For example, the name "Name" is located in the first line at the top right corner of the document, 3cm from the top boundary, and the bounding box includes 2 to 3 Chinese characters; the ID number "ID_Number" is located in the last line at the bottom of the document, 2cm from the bottom boundary, and the bounding box includes 18 numeric characters.
[0071] The weight of each text line is determined based on the importance of its text line label. Specifically, according to eKYC business requirements, among various text line labels, the ID_Number is the unique identifier of the document and has the highest importance, therefore it is assigned the highest weight. The name is considered secondary key information and is given higher weight. Information such as the issuing authority's authority and the validity period is of lower importance and is therefore assigned a lower weight. .
[0072] Under the above weight settings, text line quality detection allows frames with clearly visible key fields such as ID numbers and names to occupy higher fusion weights in LFM fusion, while giving lower fusion weights to frames with blurred key fields but clear backgrounds. This allows the model to focus on frames with clearly visible key fields when synthesizing images, improving the machine readability of the super-resolution synthesized image.
[0073] The quality score S for each line of text is obtained based on its position. line Specifically, this involves: for each detected text line, extracting its image patch based on its position, and inputting the image patch into a lightweight quality assessment subnetwork to obtain a quality score for a single text line. The quality score of a single text line is jointly determined by the sharpness score and integrity score of the text line image patch. The sharpness score is based on the Laplacian variance or FFT high-frequency energy proportion, measuring the sharpness of the text edges; the integrity score is determined based on whether the text line is occluded, truncated, or exceeds the boundary. Finally, the sharpness score and integrity score are weighted, summed, and normalized to obtain the quality score S of a single text line. line Next, the quality score S for each line of text is calculated. line We perform a weighted summation to obtain the overall text line quality score Q of the i-th frame image. i As shown in equation (3)
[0074] (3)
[0075] Where M is the number of valid text lines detected in the i-th frame, j is the text line number, and label j Let j be the label of the j-th line of text.
[0076] Next, according to Q i Determine the fusion weights for the i-th frame during LFM fusion. As shown in formula (4).
[0077] (4)
[0078] in Temperature coefficient is used to adjust the smoothness of the weights; Window is the set of neighboring frames that participate in the fusion.
[0079] If the text line quality score Qi of a certain frame image is lower than the minimum quality threshold If the value is 0.4, then the frame is determined to be an "invalid frame," and its fusion weights are directly applied during LFM fusion. Set to 0, or discard the frame directly during the feature alignment stage to prevent blurry, heavily reflective, or occluded frames from introducing noise and contaminating the super-resolution results.
[0080] In LFM, not all frames contribute equally; some frames may be blurry or useless, and forced fusion can introduce noise. Therefore, before performing LFM, a dynamic frame selection submodule assigns a fusion weight score to each frame, reducing the weight of frames with low text quality scores (such as blurry or heavily occluded frames) in subsequent fusion, thus avoiding the introduction of noise during LFM. Furthermore, since the super-resolution composite image is ultimately used for OCR recognition, the text quality in the document is used as the standard for setting the fusion weights, giving higher weights to frames with clear, unoccluded text, thereby improving the text clarity in the super-resolution composite image and the final OCR recognition performance.
[0081] like Figure 5 As shown, the cyclic bidirectional propagation includes backward propagation and forward propagation, where forward propagation propagates from frame 1 to frame n, and backward propagation propagates from frame n back to frame 1. The forward propagation includes: transferring the current frame features... Forward propagation features of the previous frame Perform feature fusion to obtain the forward propagation features of the current frame. And pass it to the (i+1)th frame; the initial forward propagation features at the start of forward propagation. Empty; the backpropagation includes: taking the current frame features Backpropagation features of the next frame Perform feature fusion to obtain the backpropagation features of the current frame. And pass it to the (i-1)th frame; backpropagation is based on the result of forward propagation, that is, the initial backpropagation features at the start of backpropagation. Forward propagation characteristics at the end of forward propagation After the bidirectional propagation cycle is completed, the backpropagation features of each frame are... The features are used as input to the reconstruction module. Since backpropagation is performed based on forward propagation, the initial backpropagation features are... The current frame already carries information from all past frames, and the fusion process during backpropagation adds temporal information from future frames to the current frame. Therefore, the backpropagation features of each frame... Both methods aggregate the bidirectional context of the entire sequence. Cyclic bidirectional propagation, through forward and backward passes, enables each frame in the video to aggregate global spatiotemporal information from the entire video sequence. This allows each frame to be fused with complementary details from its past and future frames, significantly improving the quality of reconstructed features and thus enhancing the detail recovery capability during image reconstruction, avoiding the problem of long-distance information loss in unidirectional propagation. In one implementation, the cyclic bidirectional propagation module is built based on a Transformer network structure, and feature fusion during cyclic bidirectional propagation is based on the Transformer's cross-attention mechanism. The Transformer's cross-attention mechanism has a global perspective, simultaneously capturing the correlation between global pixels of the source modality during fusion, and utilizing global pixel relationships to improve feature fusion results.
[0082] In one implementation, during the cyclic bidirectional propagation process, the current frame features are... Forward propagation features of the previous frame Feature fusion is performed, specifically including: aligning the current frame features based on optical flow. Forward propagation features of the previous frame Alignment; aligning the forward propagation features. With current frame features By performing element-wise addition, the forward propagation features of the current frame are obtained. .
[0083] In one implementation, the VSR model is trained using the following method:
[0084] Acquire multi-frame image samples containing the document, as well as high-resolution images corresponding to the samples and the text content of the document in the images;
[0085] The multi-frame image samples are input into the video super-resolution model to be trained to obtain the first synthetic image;
[0086] The main loss is obtained based on the difference between the first synthesized image and the high-resolution image;
[0087] The first synthesized image is input into a fixed OCR recognition model, and a first auxiliary loss is obtained based on the difference between the recognition result and the text content of the document.
[0088] Finally, the parameters of the VSR model are adjusted based on the main loss and the first auxiliary loss, and training is iterated until the main loss and the first auxiliary loss reach the preset standard.
[0089] The super-resolution composite image in the embodiments of this specification is used for document OCR recognition. Therefore, OCR auxiliary loss is added during the training phase of the lightweight VSR model. That is, at the end of the training phase, the super-resolution composite image generated by the VSR model is fed into a fixed OCR model, and its recognition loss is added to the total loss, thereby guiding the model to generate a super-resolution composite image that is more conducive to OCR recognition.
[0090] In some embodiments, the fixed OCR recognition model can be a mature, lightweight OCR recognition network (e.g., CRNN, SVTR, or MobileNet-V3 + CTC Head). The OCR model has been pre-trained on a large-scale general-purpose document dataset, and its parameters are frozen during VSR model training, not participating in gradient updates, and are only used for feature extraction and loss calculation. To ensure gradient differentiability, the logits output before the last softmax layer of the OCR model or the visual feature maps of intermediate layers are used for loss calculation, avoiding the use of non-differentiable Argmax operations.
[0091] In some embodiments, the total loss function of the VSR model Main loss and first auxiliary loss Weighted composition, i.e. ,in and The balance coefficient is usually set to... In the early stages of training, the main loss is primarily used to recover the image structure; in the later stages of training, the loss is gradually increased. This increases the weight of the first auxiliary loss, enhancing text clarity. The main loss... Super-resolution composite graphs are calculated using L1 loss or Charbonnier loss. With high-resolution images The pixel-level differences between them ensure the underlying visual quality of the image.
[0092] In one specific embodiment, the first auxiliary loss is obtained based on the recognition probability. Specifically, the first synthesized image is input into a fixed OCR recognition model to obtain the predicted probability distribution of the text sequence. ;calculate Labels with text content on documents The cross-entropy loss or CTC loss between them is used as the first auxiliary loss. This method can directly optimize the final accuracy of document text recognition.
[0093] In another implementation, the VSR model is trained using the following method:
[0094] Acquire multi-frame image samples containing the identification document and high-resolution images corresponding to the samples;
[0095] The multi-frame image samples are input into the video super-resolution model to be trained to obtain the first synthetic image;
[0096] The main loss is obtained based on the difference between the first synthesized image and the high-resolution image;
[0097] The first composite image and the high-resolution image are respectively input into a fixed OCR recognition model to obtain the OCR model output of the first composite image and the OCR model output of the high-resolution image.
[0098] The first auxiliary loss is obtained based on the difference between the OCR model output of the first synthetic image and the OCR model output of the high-resolution image;
[0099] Finally, the parameters of the VSR model are adjusted based on the main loss and the first auxiliary loss, and training is iterated until the main loss and the first auxiliary loss reach the preset standard.
[0100] In one specific embodiment, the first auxiliary loss is obtained based on feature perception. Specifically, after inputting the first synthesized image and its corresponding high-resolution image into a fixed OCR recognition model, feature maps of the intermediate convolutional layers of the fixed OCR model are extracted. The L2 distance between the intermediate feature maps extracted from the first synthesized image and the high-resolution image is calculated and used as the first auxiliary loss. This method enables the super-resolution image to maintain consistency with the high-resolution ground truth in the text-sensitive regions (such as stroke edges and character structures) that the OCR model focuses on, allowing the synthesis model to focus on improving the generation quality of the text-sensitive regions of the synthesized image, thereby improving the machine readability of the synthesized image.
[0101] The above describes the model structure and training method of a lightweight VSR model provided in the embodiments of this specification. It can be understood that, in addition to the lightweight VSR model in the above embodiments, the VSR model in step S102 can also be a lightweight version of advanced VSR models in the industry such as BasicVSR and Real-ESRGAN, as long as it can meet the requirement of near real-time operation on mobile devices.
[0102] The above describes a multi-frame super-resolution synthesis process provided by the embodiments of this specification. This embodiment actively fuses effective information (such as subpixel displacement and multi-angle texture) from multiple video frames using the VSR algorithm, allowing different video frames to complement each other's details. This fills in the blurriness and details of each original frame, synthesizing a super-resolution composite image that is theoretically clearer, has higher resolution, and less noise than any original single frame, fundamentally breaking through the image quality ceiling of single-frame acquisition.
[0103] Next, in step S103, several target images are selected from the multi-frame images.
[0104] The target image is the highest quality "original real image" selected directly from real video frames captured from the original video stream without any synthetic processing. This image retains the most original physical world information and has not been subject to any artificial algorithmic processing, making it suitable for document anti-counterfeiting detection.
[0105] In one embodiment, based on the image quality detection result obtained in step S1013, the frame with the highest image quality score among the multiple frames is selected as the target image.
[0106] In another embodiment, the target image comprises image frames acquired under at least two different lighting conditions.
[0107] In a more specific embodiment, the target image comprises two frames: one is the image frame with the highest image quality score under normal lighting, and the other is the image frame captured after the client actively controls the flash to be turned on. Before and after the flash is turned on, the reflective effect of the screen or printed paper is significantly different from that of a genuine document. Therefore, comparing the best frame under normal lighting with the best frame after the flash is turned on can greatly enhance the detection capability of screen re-photographs and some printed documents, providing a stronger anti-counterfeiting solution.
[0108] Next, in step S104, the composite image and the target image are sent to the server.
[0109] The client packages the composite image and the target image into two separate data sets and uploads them to the server via different data tracks. The super-resolution composite image is specifically used for OCR recognition, while the original real image is specifically used for document anti-counterfeiting detection.
[0110] By uploading the composite image and the target image to the server through two different data tracks, parallel transmission of the composite image and target image data can be achieved, thereby greatly reducing data transmission time and user waiting time.
[0111] Next, in step S105, optical character recognition (OCR) is performed on the super-resolution composite image on the server, and anti-counterfeiting detection is performed on the original real image.
[0112] After receiving the composite image and target image uploaded by the client, the server processes the two sets of data in parallel: high-precision OCR recognition is performed on the super-resolution composite image, and anti-spoofing detection is performed on the original real image. Parallel processing can improve data processing efficiency, shorten data processing time, and reduce user waiting time.
[0113] When performing high-precision OCR recognition on super-resolution synthesized images, the accuracy and recall of OCR recognition are significantly improved because the super-resolution synthesized images have extremely high image quality and the VSR model used for super-resolution synthesis is specifically designed and trained for OCR recognition scenarios.
[0114] During anti-counterfeiting detection, a series of anti-counterfeiting detection algorithms are executed on the original authentic image, including but not limited to: screen capture detection, print / copy detection, PS counterfeiting detection, and physical anti-counterfeiting feature verification.
[0115] Next, in step S106, the document recognition result is obtained based on the OCR recognition result and the anti-counterfeiting detection result.
[0116] After parallel processing is completed, the server-side integrated decision-making module receives the OCR recognition result and the anti-counterfeiting detection result, and obtains the final document recognition result based on the OCR recognition result and the anti-counterfeiting detection result.
[0117] In one implementation, the eKYC process only determines final authentication success when both OCR recognition and anti-counterfeiting detection are successful. Failure at any stage will result in process interruption or transfer to manual review.
[0118] In some more specific embodiments, the server-side integrated decision-making module outputs basic document information based on the OCR recognition result and outputs document authenticity information based on the anti-counterfeiting detection result, and outputs the document authenticity information and basic document information together as evidence recognition result.
[0119] In one implementation, after obtaining the document recognition result, the method further includes: S107, the server sends the document recognition result to the client, and the client feeds back the recognition result to the user.
[0120] In summary, the embodiments in this specification fundamentally decouple and resolve the long-standing industry-wide core contradiction between "image recognition clarity" and "anti-counterfeiting detection fidelity" by generating a "super-resolution composite image" specifically for OCR recognition and an "original real image" specifically for anti-counterfeiting detection on the client side, and uploading them separately to the server for parallel processing. The embodiments in this specification, through the design of a lightweight VSR model and the combination of an intelligent dynamic frame selection strategy, enable the completion of complex multi-frame video super-resolution synthesis calculations on resource-constrained mobile clients. This not only reduces the demand for network bandwidth and server resources but also optimizes the real-time performance of user interaction. Furthermore, by introducing a dynamic frame selection strategy based on text line quality into the VSR model and incorporating OCR-assisted loss during VSR model training, the super-resolution synthesis process focuses more on improving the clarity of key text information, guiding the VSR model to generate super-resolution composite images with stronger "machine readability," thereby improving the accuracy and recall rate of OCR recognition.
[0121] According to another embodiment, a computing device is also provided, including a memory and a processor, wherein executable code is stored in the memory, and when the processor executes the executable code, it implements the steps of the method as described in any of the above embodiments.
[0122] According to another embodiment, a computer-readable storage medium is also provided, having stored thereon a computer program that, when executed in a computer, causes the computer to perform the steps of the method as described in any of the above embodiments.
[0123] According to yet another embodiment, a computer program product is also provided, including a computer program / instructions that, when executed by a processor, implement the steps of the method as described in any of the above embodiments.
[0124] While one or more embodiments of this specification provide the operational steps of the methods described in the embodiments or flowcharts, more or fewer operational steps may be included based on conventional or non-inventive means. The order of steps listed in the embodiments is merely one possible order of execution among many steps and does not represent the only possible order. In actual device or end product execution, the methods shown in the embodiments or drawings may be executed sequentially or in parallel (e.g., in a parallel processor or multi-threaded processing environment, or even a distributed data processing environment). The terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, product, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, product, or apparatus. Without further limitations, the presence of other identical or equivalent elements in the process, method, product, or apparatus that includes said elements is not excluded. For example, the use of terms such as "first," "second," etc., is to denote names and does not indicate any particular order.
[0125] For ease of description, the above devices are described in terms of function, divided into various modules. Of course, when implementing one or more of these specifications, the functions of each module can be implemented in one or more software and / or hardware components, or a module that performs the same function can be implemented by a combination of multiple sub-modules or sub-units. The device embodiments described above are merely illustrative. For example, the division of units is only a logical functional division; in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces, indirect coupling or communication connection between devices or units, and may be electrical, mechanical, or other forms.
[0126] This specification is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this specification. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0127] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0128] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0129] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.
[0130] Memory may include non-persistent storage in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.
[0131] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can store information by any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic disk storage, graphene storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.
[0132] Those skilled in the art will understand that one or more embodiments of this specification can be provided as a method, system, or computer program product. Therefore, one or more embodiments of this specification may take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, one or more embodiments of this specification may take the form of a computer program product implemented on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0133] One or more embodiments of this specification can be described in the general context of computer-executable instructions, such as program modules, that are executed by a computer. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform a particular task or implement a particular abstract data type. One or more embodiments of this specification can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.
[0134] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, system embodiments are basically similar to method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions in the method embodiments. In the description of this specification, the terms "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., refer to specific features, structures, materials, or characteristics described in connection with that embodiment or example, which are included in at least one embodiment or example of this specification. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described can be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification and the features of different embodiments or examples.
[0135] The above description is merely an embodiment of one or more embodiments of this specification and is not intended to limit the scope of these embodiments. Various modifications and variations can be made to these embodiments by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this specification should be included within the scope of the claims.
Claims
1. A method for identifying identification documents, executed by a client, comprising: Capture multiple frames of images containing identification documents; A composite image is generated based on the multi-frame images using a synthesis model, and the resolution of the composite image is higher than that of the multi-frame images. Select several target images from the multiple frames of images; The composite image and the target image are sent to the server for document recognition, and the composite image is used for character recognition.
2. The method according to claim 1, wherein capturing multiple frames of images containing the identification document specifically includes: Photograph the ID document to obtain a video stream containing the ID document; The document's location is captured in each frame of the video stream, and its location is tracked in real time during the shooting process. Images in the video stream are filtered based on image quality; Identify the document angle in the remaining images, and retain multiple frames of images under different document angles based on the identification results; Align the positions of documents in multiple frames of images.
3. The method according to claim 1, wherein, The process of generating a composite image based on the multi-frame images using a synthesis model specifically includes: Feature extraction is performed on the multiple frames of images; Each frame image is fused with features from its neighboring frame images to obtain the fused features; The fused features are then propagated bidirectionally in a loop to obtain the reconstructed features. The image is reconstructed based on the reconstructed features to obtain a synthetic image.
4. The method according to claim 3, wherein, The feature fusion of each frame image with its neighboring frame images specifically includes: During feature fusion, the fusion weight of each frame image is determined based on the text line quality score of each frame image; the text line quality score is obtained through text line quality detection.
5. The method according to claim 4, wherein, The text line quality detection specifically includes: Get the position of each line of text in the image; The label and quality score for each line of text are obtained based on its position. The quality scores of each text line are weighted and summed to obtain the text line quality score of the image frame; the weight of each text line is determined according to the text line label.
6. The method according to claim 1, wherein, The synthetic model was trained using the following method: Acquire multi-frame image samples containing the document, as well as high-resolution images and document text content corresponding to the samples; The multi-frame image samples are input into the synthesis model to be trained to obtain the first synthesized image; The main loss is obtained based on the difference between the first synthesized image and the high-resolution image; The first synthesized image is input into a fixed OCR recognition model, and a first auxiliary loss is obtained based on the difference between the recognition result and the text content of the document. The model parameters of the synthetic model are adjusted based on the main loss and the first auxiliary loss.
7. The method according to claim 1, wherein, The target image includes: image frames acquired under at least two different lighting conditions.
8. A method for identifying identification documents, executed by a server, comprising: Receive the composite image and the target image from the client; The composite image is synthesized based on multiple frames of images containing the document, and the resolution of the composite image is higher than that of the multiple frames; the target image is selected from the multiple frames; the multiple frames are directly obtained by the client. Perform optical character recognition (OCR) on the synthesized image; Perform anti-counterfeiting detection on the target image; The document recognition result is obtained based on the OCR recognition result and the anti-counterfeiting detection result.
9. The method according to claim 8, wherein, The anti-counterfeiting detection includes: screen capture detection, print / copy detection, PS counterfeiting detection, and physical anti-counterfeiting feature verification.
10. A computing device comprising a memory and a processor, wherein the memory stores executable code, and the processor, when executing the executable code, implements the method of any one of claims 1-9.