Semantic Segmentation-based Document Location and Classification Method
Through the document positioning and classification method based on semantic segmentation, the problem that the existing technology cannot accurately describe the edge of the document is solved, and the accurate positioning and classification of the edge of the document is realized, and the direction and damage of the document can be judged.
Patent Information
- Application Number
- CN202211707800.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-28
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2042-12-28
AI Technical Summary
The prior art cannot accurately describe the edge of the document, especially the perspective or damaged document, and cannot determine whether the document has any problems such as damage, incompleteness or flip.
The document positioning and classification method based on semantic segmentation is used to predict the document and portrait area through the semantic segmentation model, and the document area is represented by masks to finely represent the document area, and the document orientation is judged based on the relative position of the portrait area.
It realizes the precise positioning and classification of the edges of the certificate, and can determine whether the certificate is damaged or incomplete, and accurately determine the orientation of the certificate, such as the problem of flipping or rotating by 180°.
Smart Images

Figure CN116704166B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision technology, and specifically to a method for document positioning and classification based on semantic segmentation. Background Art
[0002] Object detection aims to locate predefined object instances in an image, representing the object position with a bounding box, which is a standard rectangular box and can be uniquely determined by two coordinate points (the upper left corner coordinate and the lower right corner coordinate) or the center point coordinate and the width and height values. Common object detection algorithms based on deep learning include the two-stage R-CNN series and single-stage algorithms designed for real-time inference, such as SSD, RetinaNet, and YOLO series. Identity document photos are widely used in the fields of online finance, Internet finance, and e-commerce. For the document images taken by users, they often contain a large amount of background, while the document only occupies a partial area. Therefore, it is extremely important to accurately locate the document area, especially to pay attention to whether the uploaded document contains problems such as rotation, tilt, flip, damage, mutilation, and occlusion.
[0003] 1. A prior patent (Publication No.: CN112017245A) discloses a document positioning method. Obtain a to-be-detected image, input the to-be-detected image into a target detection model, perform target detection on the to-be-detected image through the target detection model to obtain a detection result; the detection result includes the type information of the document in the to-be-detected image, the position information of the document in the to-be-detected image, the position information of the vertices of the document in the to-be-detected image, and the direction information of the document in the to-be-detected image. By increasing the number of prediction structures to change the structure of the existing target detection model, two detection data, namely the overall direction of the document and the position of the document vertices, are newly added in the target detection of the document, achieving improved document detection effect while faster document detection; 2. A prior patent (Publication No.: CN114240952A) discloses a document positioning method, device, electronic device, and readable storage medium. The method first obtains a sample image, where the sample image can be obtained by transforming the original image of the target document, then annotates all vertex coordinates of the target document in the sample image, constructs a loss function based on the sample image, and iterates the model parameters based on this loss function until convergence to obtain a model positioning module. The model positioning model is trained with a sample image obtained by transforming the original image of the target document and annotates all vertices of the target document. The transformation includes rotation, shearing, distortion, perspective, affine, etc., enabling the trained document positioning model to accurately locate non-regularly shaped target documents in the image, providing a basis for subsequent detection, recognition, extraction, etc. of information in the document and expanding the applicability of document positioning.
[0004] Based on the object detection algorithm, Patent 1 locates and classifies the certificates simultaneously and obtains the certificate orientation information. In view of the problem that Patent 2 cannot effectively detect non-regular rectangular certificates, it proposes to label the vertices of the certificates and predict the four vertices of the certificates based on the YOLO framework.
[0005] The above methods have obvious disadvantages, that is, neither the standard rectangular box nor the quadrilateral box with four points can accurately describe the edge of the certificate (such as a certificate with perspective or damage), and it is impossible to judge whether there are damage or incomplete quality problems of the certificate based on the detection results. Although the method of detecting four vertices can judge the inclination angle of the certificate to a certain extent, it cannot judge the orientation of the certificate, such as whether it is flipped or rotated 180°. Summary of the Invention
[0006] (1) Technical problems to be solved
[0007] In view of the deficiencies of the prior art, the present invention provides a method for certificate location and classification based on semantic segmentation, which solves the problem that neither the standard rectangular box nor the quadrilateral box with four points can accurately describe the edge of the certificate.
[0008] (2) Technical solutions
[0009] To achieve the above objectives, the present invention is realized through the following technical solutions: A method for certificate location and classification based on semantic segmentation, the method includes:
[0010] Based on the feature extraction network, feature fusion and prediction branch network of the model, wherein the prediction branch network includes a semantic segmentation branch and a certificate type classification branch.
[0011] Preferably, the method includes model design, which is the selection of the main framework of the model. Any lightweight real-time semantic segmentation model can be used, and it needs to be improved to adapt to the location of the certificate and portrait areas and the classification of certificates.
[0012] Preferably, the method also includes a processing process, which is to process the input image to be detected. In the preprocessing stage, the image is converted into a floating-point tensor, the size is scaled to a specified size, and the pixel values are normalized between 0 and 1.
[0013] Preferably, the model design specifically includes the following steps:
[0014] S1. Since the portrait area is completely included in the certificate area, for the feature map output by the segmentation branch, the general semantic segmentation model performs multi-classification at a certain pixel point, that is, a pixel can only belong to a certain class, while the pixels in the overlapping area of the certificate and the portrait can belong to both the certificate and the portrait at the same time;
[0015] S2. Decouple the classification and segmentation of certificates. The segmentation branch is only responsible for outputting the feature maps of certificates and portraits, and the certificate classification branch is specifically responsible for classifying the types of certificates.
[0016] Preferably, the processing flow specifically includes the following steps:
[0017] 1). The feature extraction network uses a CNN-based model to extract image features;
[0018] 2). In the feature fusion stage, the multi-scale feature maps obtained by the feature extraction network are fused;
[0019] 3). Design the prediction branch network, including the semantic segmentation branch and the certificate type classification branch, sharing the feature extraction network, and each branch is responsible for different tasks;
[0020] 4). Regarding the characteristics of certificates, they are different from other object objects processed by general semantic segmentation models, such as the sky and grasslands.
[0021] (III) Beneficial effects
[0022] The present invention provides a method for certificate positioning and classification based on semantic segmentation. It has the following beneficial effects:
[0023] The present invention provides a method for certificate positioning and classification based on semantic segmentation. Based on a semantic segmentation model rather than an object detection model, it predicts the certificate and portrait areas, uses a mask to represent, can more precisely represent the certificate area, can accurately crop the certificate based on the mask for subsequent processing, and can judge the orientation of the certificate, such as flipping and rotation, according to the relative position of the portrait area in the certificate.
[0024] The present invention provides a method for certificate positioning and classification based on semantic segmentation. The present invention can decouple certificate segmentation and certificate classification, which is convenient for separate optimization of the segmentation branch and the classification branch, improving the performance of each branch. In the model training stage, it can effectively avoid problems caused by data imbalance. Suppose there are 5000 ID cards and 10 driver's licenses in the training dataset. If a general semantic segmentation model is used, almost only the information of ID cards can be learned and certificate classification cannot be achieved. After decoupling, the segmentation branch only needs to focus on the certificate and does not need to care whether it is an ID card or a driver's license. In addition, it is convenient for model expansion. Suppose a new type of certificate is added. When optimizing the model, the feature extraction network and the segmentation branch can be frozen, and only new type of certificates need to be collected to optimize the classification branch, thus avoiding the annotation of the certificate areas of new type of data and saving a large amount of manpower and material resources.
[0025] The present invention provides a method for document positioning and classification based on semantic segmentation. The present invention can replace the multi-classification of semantic segmentation with binary classification, which is convenient for the special optimization of the document feature map. A specific MASK region loss function is designed based on the characteristics of the document region, so that the obtained MASK is more refined, and the document MASK obtained by the model can be used to judge whether there are quality problems such as damage, mutilation, and occlusion of the document. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] Figure 1 It is a schematic diagram of the network structure of the present invention;
[0027] Figure 2 It is an example diagram of the processing of the feature map by the general semantic segmentation model and the improved segmentation model of the present invention;
[0028] Figure 3 It is a schematic diagram of the visualized DEMO output by the segmentation branch of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0029] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0030] Embodiment:
[0031] As Figures 1-3 shown, the embodiment of the present invention provides a method for document positioning and classification based on semantic segmentation, and the method includes:
[0032] Based on the feature extraction network, feature fusion, and prediction branch network of the model, where the prediction branch network includes a semantic segmentation branch and a document type classification branch.
[0033] The method includes model design, which is the selection of the main framework of the model. Any lightweight real-time semantic segmentation model can be used, and it needs to be improved to adapt to the positioning of the document and portrait areas and the classification of documents.
[0034] The model design specifically includes the following steps:
[0035] S1. Since the portrait area is completely contained in the document area, for the feature map output by the segmentation branch, the general semantic segmentation model performs multi-classification at a certain pixel point, that is, a pixel can only belong to a certain class, while the pixels in the overlapping area of the document and the portrait can belong to both the document and the portrait at the same time;
[0036] Therefore, when splitting, two layers of feature maps are output. One layer is responsible for document segmentation, and the other layer is responsible for portrait segmentation. Each layer adopts a binary classification method. In the prediction stage, a threshold is set to determine whether a certain point in this layer belongs to the object category it is responsible for.
[0037] S2. Decouple the classification and segmentation of documents. The segmentation branch is only responsible for outputting the feature maps of documents and portraits, and the document classification branch is specifically responsible for classifying document categories.
[0038] For example, there are three document types: ID card, driver's license, and residence permit. The output feature map dimension of the general semantic segmentation method is (5, H, W), where 5 represents the object categories (ID card, driver's license, residence permit, portrait, background), and H and W represent the height and width of the feature map respectively. After multi-classification, a certain pixel point can only be one of the 5 categories; after decoupling the classification and segmentation of documents, (1) the output feature map dimension of the segmentation network is (2, H, W), where 2 represents the object categories (document, portrait). Since binary classification is performed on each layer and it is specifically responsible for segmenting the object, it is determined whether it is the background through a threshold without having to separately consider the background as a category. (2) The classification branch network is responsible for classifying the three types of documents (ID card, driver's license, and residence permit).
[0039] The method also includes a processing flow. For the input image to be detected, in the preprocessing stage, the image is converted into a floating-point tensor, scaled to a specified size, and the pixel values are normalized between 0 and 1.
[0040] The specific processing flow includes the following steps:
[0041] 1). The feature extraction network uses a CNN-based model to extract image features;
[0042] This CNN model can adopt an image classification model pre-trained on the ImageNet dataset, such as ResNet, VGG
[16] , DenseNet, MobileNet, EfficientNet. Remove its fully connected layer, and only the output feature maps of each stage need to be retained.
[0043] 2). In the feature fusion stage, the multi-scale feature maps obtained by the feature extraction network are fused;
[0044] Assume the input image dimension is (3, H, W), where 3 represents the three RGB channels of the image, and H and W are the height and width of the image. The feature maps of five scales extracted by the feature extraction network are respectively (C1, H / 2, W / 2), (C2, H / 4, W / 4), (C3, H / 8, W / 8), (C4, H / 16, W / 16), (C5, H / 32, W / 32), where C1 - C5 are the dimension sizes of the feature maps, and the dimensions vary with different feature extraction networks. Any of the above fusion methods of lightweight real-time semantic segmentation models can be used to obtain the final feature map, whose dimension is (C, H / 8, W / 8), where C is the dimension size of the fused feature map.
[0045] 3). Design of the prediction branch network, including the semantic segmentation branch and the document type classification branch, sharing the feature extraction network, with each branch responsible for different tasks;
[0046] (1) Among them, the semantic segmentation branch takes the feature map fused in the previous stage as input and outputs a feature map with the dimension of (2, H / 8, W / 8), that is, two feature maps of size (H / 8, W / 8). It is specified that the first layer is responsible for document classification, and the second layer is responsible for portrait classification, and binary classification is performed on each layer. (2) The document type classification network contains two layers of CNN and one layer of fully connected. Taking the feature map (C3, H / 8, W / 8) obtained by the feature extraction network as input, the output is a vector with the size of num_classes, representing the number of document types; each value in the vector represents the confidence of its corresponding type, which is a probability value between 0 and 1, and the sum of the entire vector is 1.
[0047] 4). Regarding the characteristics of documents, different from other object objects processed by general semantic segmentation models, such as the sky and grasslands;
[0048] A document is a continuous closed area. Therefore, in order to obtain a better-quality and finer document MASK, an additional loss function is designed to optimize the number of regions of the MASK. where a maski is the number of maski output by the model, and a i is the actual number of documents.
[0049] Train the above model on the training dataset to obtain the model weight file (implemented using pytorch, and the obtained weight file is in PTH format).
[0050] Secondly, convert the model weight file in PTH format to the Open Neural Network Exchange (ONNX) format.
[0051] Inference is implemented using ONNX, based on the onnxruntime library. It includes preprocessing of image normalization and size scaling, recording the scaling ratios RatioH and RatioW, which represent the height and width scaling factors respectively. Two outputs are obtained from the inference:
[0052] (1) One is the feature map of size (2, H / 8, W / 8) of the segmentation branch, where H and W are the height and width of the preprocessed image. First, perform sigmoid processing to convert it into probability values between 0 and 1. Assume that the first layer is responsible for the segmentation of the certificate, and the second layer is responsible for the segmentation of the portrait. Separate thresholds T1 and T2 (such as 0.7 and 0.9) are set for the two layers. In the first layer, the areas with probability values greater than T1 are regarded as the certificate regions, and in the second layer, the areas with probability values greater than T2 are regarded as the portrait regions. The certificate MASK and portrait MASK of size (H / 8, W / 8) are obtained respectively, and are enlarged to the size of the original input image according to RatioH and RatioW, which is used as the final MASK output;
[0053] (2) The other output is the classification branch, a vector with a size of N, representing the number of certificate types. Perform softmax operation on the vector to make it into probability values, take out the maximum probability value and its corresponding index, and obtain the type of the certificate according to the index in the pre-defined certificate category set.
[0054] As Figure 2 shown, assume there are two types of certificates, the second-generation ID card and the driver's license. (a) is the feature map output by the general semantic segmentation model, using multi-classification, and only one class can be predicted at a certain point. (b) is the output of the improved segmentation branch, which only focuses on the certificate and the portrait, and performs binary classification for each layer. A certain point can belong to both the certificate and the portrait categories at the same time.
[0055] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principle and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A method for document location and classification based on semantic segmentation, characterized in that, the method includes: The feature extraction network uses a CNN-based model to extract image features, and in the feature fusion stage, the multi-scale feature maps obtained by the feature extraction network are fused; Assume that the input image dimension is (3, H, W), where 3 represents the RGB three channels of the image, and H and W are the height and width of the image. The feature maps of 5 scales extracted by the feature extraction network are respectively (C1, H / 2, W / 2), (C2, H / 4, W / 4), (C3, H / 8, W / 8), (C4, H / 16, W / 16), (C5, H / 32, W / 32), where C1 - C5 are the dimension sizes of the feature maps. The dimensions of different feature extraction networks are different. After fusion, the final feature map is obtained, and its dimension is (C, H / 8, W / 8), where C is the dimension size of the fused feature map; The prediction branch network includes a semantic segmentation branch and a document type classification branch, sharing the feature extraction network, and each branch is responsible for different tasks; (1) Among them, the semantic segmentation branch takes the feature map fused in the previous stage as input and outputs a feature map with a dimension of (2, H / 8, W / 8), that is, two feature maps with a size of (H / 8, W / 8). It is specified that the first layer is responsible for the classification of documents, and the second layer is responsible for the classification of portraits. Binary classification is performed on each layer, and separate thresholds T1 and T2 are set for the two layers. In the first layer, the area with a probability value greater than T1 is regarded as the document area, and in the second layer, the area with a probability value greater than T2 is regarded as the portrait area, respectively obtaining a document MASK and a portrait MASK with a size of (H / 8, W / 8). (2) The document type classification branch includes two layers of CNN and one layer of fully connected. Taking the feature map (C3, H / 8, W / 8) obtained by the feature extraction network as input, the output is a vector with a size of num_classes, representing the number of document types; each value in the vector represents the confidence of its corresponding type, which is a probability value between 0 and 1, and the sum of the entire vector is 1.
2. The method for document location and classification based on semantic segmentation according to claim 1, characterized in that: The method also includes a processing process, which is to preprocess the input image to be detected. In the preprocessing stage, the image is converted into a floating-point tensor, the size is scaled to a specified size, and the pixel values are normalized between 0 and 1.
Citation Information
Patent Citations
Certificate positioning method
CN112017245A
Certificate positioning method and device, electronic equipment and readable storage medium
CN114240952A
Image quality evaluation method and system
CN114419008A
Object tracking method, ground object tracking method, device, system, and storage medium
WO2022152110A1