Certificate portrait segmentation method and device based on multi-scale clustering and processing system
Through the multi-scale clustering and multi-feature matrix fusion method, the dependence problem on sample data and environmental factors in image processing is solved, and the document portrait segmentation with high accuracy and stability is achieved.
Patent Information
- Application Number
- CN202411909821.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-24
- Publication Date
- 2025-05-16
AI Technical Summary
The prior art relies on a large amount of sample data in image processing and is sensitive to environmental factors such as lighting and noise, resulting in inaccurate or unstable segmentation results.
The document portrait segmentation method based on multi-scale clustering is adopted. By acquiring document images, selecting ROI areas, image enhancement processing, clustering processing, channel separation, binarization processing, morphological processing and edge contour detection, the edge contour information of portrait features is obtained, and the final segmented portrait information is obtained through the fusion of multi-feature matrix.
It does not rely on a large number of sample data, reduces sensitivity to environmental factors, improves the accuracy and stability of segmentation results, and has relatively low computational complexity.
Smart Images

Figure CN120014259A_ABST
Abstract
Description
Technical Field
[0001] The invention relates to a method, a device and a processing system for document portrait segmentation based on multi-scale clustering, and belongs to the technical field of pattern recognition and image processing. Background Art
[0002] In recent years, the field of computer pattern recognition and image processing technology has developed rapidly, and image processing has become a major hot research direction. The main methods of image processing are: 1. Using deep learning (machine learning) to train with a large number of data samples; 2. Using traditional pattern recognition image processing solutions. Through long-term practice and research, the applicant has found that there are at least the following problems in the current popular image processing technologies:
[0003] Using deep learning (or machine learning) requires a large amount of sample data for training, and the quality of the data samples and sample cleaning will have a great impact on the effect of subsequent model training, and the coverage depends on the samples themselves.
[0004] In certain industries, there is information such as users’ ID card data, portrait data, and privacy data, which is not allowed to be collected and used casually.
[0005] Traditional image processing solutions, such as threshold segmentation, edge detection, texture features and other methods, are sensitive to environmental factors such as lighting and noise, resulting in inaccurate or unstable segmentation results.
[0006] Traditional face recognition solutions also rely on a large amount of sample data for training. In actual use, what is obtained is the face recognition area information (and possibly background information), which cannot fully contain the entire portrait information. At the same time, it has high requirements for the image of the certificate. For example, it is not robust when encountering stained images, blurred light or part of the face, etc. Summary of the invention
[0007] The present invention provides a method, device and processing system for document portrait segmentation based on multi-scale clustering, which effectively solves the problem that deep learning technology solutions rely too much on data samples themselves, and effectively solves the problem that traditional image processing solutions have high sensitivity to the image itself.
[0008] In order to solve the above technical problems, the technical solution adopted by the present invention is:
[0009] A method for document portrait segmentation based on multi-scale clustering, comprising:
[0010] Get the document image;
[0011] Select the ROI area on the document image to crop the target image, and perform image enhancement processing on the cropped target image;
[0012] Perform clustering on the enhanced image and generate a new cluster image based on the cluster center and label matrix;
[0013] The newly generated cluster image is subjected to channel separation, binarization, morphological processing and edge contour detection to obtain edge contour information of portrait features;
[0014] Calculate maximum contour feature information based on edge contour information;
[0015] The maximum contour feature information is fused into a multi-feature matrix to obtain the final segmented portrait information.
[0016] Furthermore, a ROI region is selected on the document image to crop a target image, and a method for performing image enhancement processing on the cropped target image is as follows:
[0017] Suppose the document image is a function I(x,y), where (x,y) is the coordinate on the image plane, and the ROI area is a rectangular area R={(x,y)|x1≤x≤x2,y1≤y≤y2}, where x1,x2,y1,y2 are the coordinate values defining the ROI area; the cropped target area image T(x,y);
[0018] The target area image T(x, y) is subjected to histogram equalization, the RGB image is converted to the YCrCb space, and the Y channel is equalized; the processed image F(x, y) is obtained.
[0019] Furthermore, the target area image T(x, y) is subjected to histogram equalization, the RGB image is converted to the YCrCb space, and the Y channel is equalized; the method for obtaining the processed image F(x, y) is as follows:
[0020] Convert the RGB image to YCrCb space, where is the coordinate of the pixel point; the value of the pixel point is composed of the values of the three channels R, G, and B. The value of , converted to YCrCb space image , the calculation formula is:
[0021] Y = 0.299⋅R + 0.587⋅G + 0.114⋅B;
[0022] Cr = 128 + (-0.168736⋅R - 0.331264⋅G + 0.5⋅B);
[0023] Cb =128 + (0.5⋅R - 0.418688⋅G - 0.081312⋅B);
[0024] Where R, G, and B are images respectively. The three-channel value of the pixel;
[0025] The input image is G(x, y), and the grayscale range is [0, L-1], where L is the number of grayscale levels;
[0026] Calculate the probability density of each gray level k: p(k)= , in, is the number of pixels of gray level k, and N is the total number of pixels in the image;
[0027] Calculate the cumulative distribution function: C(k)= ;
[0028] Map each gray level k to a new gray level: T(k) = |(L−1)⋅C(k)|, and the output image after histogram equalization is: = T(G(x,y));
[0029] The YCrCb image after image enhancement , that is, any point in the image The value of , converted to RGB image ; The calculation formula is:
[0030] R = Y + 1.402⋅(Cr-128);
[0031] G = Y - 0.344136⋅(Cb–128) - 0.714136⋅(Cr-128);
[0032] B = Y + 1.772⋅(Cb–128).
[0033] Furthermore, the enhanced image is clustered, and a new cluster image is generated according to the cluster center and the label matrix as follows:
[0034] Create a color table colorTab, calculate the number of samples sampleCount, sample point matrix points, the number of clusters clusterCount, iteration parameter Count, and error threshold EPS;
[0035] Set the colorTab parameter value to { (0,0,255), (0,255,0), (255,0,0),(0,255,255),(255,0,255),(255,255,0)}; set sampleCount to the number of pixels in the T(x,y) image; regard each pixel as a sample point, for each pixel (col,row) in the image, set points=row*width+col, and set clusterCount to 3, indicating that the pixels in the image are to be clustered into 3 clusters. In the K-means algorithm, k is the pre-specified number of clusters.
[0036] Loop through each pixel of the F(x,y) image; for each pixel (col,row) in the image, calculate its index in the points matrix, the calculation formula is index=row×width+col; divide the RGB value of the pixel;
[0037] Don't assign them to the corresponding rows and columns in the points matrix;
[0038] Execute the K-means algorithm to cluster the sample points in the points matrix to obtain the label matrix labels; generate a new cluster image K(x,y) based on the cluster center and the label matrix.
[0039] Furthermore, the newly generated cluster image is subjected to channel separation, binarization, morphological processing and edge contour detection to obtain edge contour information of portrait features as follows:
[0040] The cluster image K(x,y) is channel-separated, and the separated image K i (x,y);
[0041] Where i is the channel index value, and the binarization calculation is:
[0042] ;
[0043] Among them, t is the threshold between feature information and background;
[0044] After binarization, the image is in black and white. is white, is black;
[0045] The morphological processing includes Perform expansion treatment, then corrosion treatment;
[0046] The dilation process uses a 3×3 matrix template to dilate the entire image. Traverse in the horizontal x and vertical y directions, the template is:
[0047] ;
[0048] Where D is the corresponding pixel value traversed; the expansion process is used to take the minimum value in the template as the value of the midpoint of the template for each traversed template; the corrosion process is used to take the maximum value of the template as the value of the midpoint of the template for each traversed template; the pixel value of the small white noise in the image is 1. Through this process, since the background around the white noise is black, that is, the pixel value is 0, when the template traverses to the white pixel during the expansion process, it will take the minimum value 0 due to the existence of the surrounding black pixels and become a black pixel. Then, the corrosion process is used to restore the expanded image to its original size. The processed image is recorded as ;
[0049] Morphologically processed image , perform Canny edge contour detection;
[0050] The edge contour detection comprises:
[0051] Perform Gaussian filtering on the image to reduce the impact of noise; the formula of the Gaussian filter is:
[0052] ;
[0053] in, is the standard deviation of the Gaussian kernel, image Convolve with the Gaussian kernel G(x,y) to get a flat
[0054] Slide image :
[0055] ;
[0056] in, is the original image, is the smoothed image;
[0057] The Sobel operator is used to calculate the gradient of the image in the x and y directions. The formula is as follows:
[0058] ;
[0059] ;
[0060] in, and Represents the gradient of the image in the x and y directions respectively; the gradient magnitude G(x,y) and the gradient direction The calculation formula is:
[0061] ;
[0062] ;
[0063] For each pixel, according to the gradient direction , compared with the two neighboring pixels in the gradient direction; if the pixel is not a local maximum, it is set to 0; set is the gradient amplitude after non-maximum suppression;
[0064] ;
[0065] Use dual threshold detection to distinguish strong edges from weak edges and suppress non-edges;
[0066] Set high threshold and low threshold , the classification rules are as follows:
[0067] Strong edge: pixels with gradient magnitude greater than the high threshold;
[0068] Weak edge: pixels with gradient magnitude between the low threshold and the high threshold;
[0069] Non-edge: pixels whose gradient magnitude is less than the low threshold;
[0070] Let E(x,y) be the edge image after double threshold detection:
[0071] ;
[0072] By connecting weak edges to strong edges, the continuity of the edges is ensured; all weak edges connected to strong edges are tracked, these edges are retained, and other weak edges are suppressed; let the final edge image be P(x,y):
[0073] ;
[0074] If E(x,y) is a strong edge or is connected to a strong edge, the value is 255, otherwise it is 0.
[0075] Furthermore, the method for calculating the maximum contour feature information according to the edge contour information is:
[0076] Image after edge contour detection , where i is the channel index value, and the area method is used to find the current maximum contour feature information Area( ), the contour area calculation formula is:
[0077] ;
[0078] Among them, C is the point set of the contour, are the coordinates of the points on the contour.
[0079] Furthermore, the feature information is subjected to multi-feature matrix fusion to obtain the final segmented portrait information as follows:
[0080] Image The corresponding maximum feature contour , where i is the channel index value, which is used as the feature contour mask FeatureMask to copy the feature portrait data in the target area image T (x, y) to the maximum feature contour in sequence to obtain the final segmented portrait information.
[0081] A second aspect of the present invention provides a device for document portrait segmentation based on multi-scale clustering, comprising:
[0082] Image preprocessing module: used to obtain the document image; select the ROI area on the document image to crop the target image, and perform image enhancement processing on the cropped target image;
[0083] Clustering processing module: used to perform clustering processing on the enhanced image and generate a new clustering image based on the cluster center and label matrix;
[0084] Feature extraction module: used to perform channel separation, binarization, morphological processing and edge contour detection on the newly generated cluster image to obtain the edge contour information of the portrait features; calculate the maximum contour feature information based on the edge contour information;
[0085] Feature fusion module: used to fuse the maximum contour feature information into multiple feature matrices to obtain the final segmented portrait information.
[0086] The third aspect of the present invention provides a system for processing ID portrait segmentation based on multi-scale clustering, comprising: an information reading module, a scanning mechanism and the above-mentioned ID portrait segmentation device based on multi-scale clustering;
[0087] The information recognition module is used to recognize and read information from the certificate;
[0088] The scanning mechanism is used to scan the image of the certificate;
[0089] The document scanned image is segmented by the portrait segmentation device to obtain the segmented portrait information.
[0090] Furthermore, it also includes: a first seat body, a second seat body and a transmission mechanism;
[0091] The first base body and the second base body are arranged to cover each other, and a transmission channel for the certificate to pass through is provided between the first base body and the second base body; the information reading module, the transmission mechanism and the scanning mechanism are all arranged in the cavity covered by the first base body and the second base body;
[0092] The certificate is inserted from one end of the transmission channel, and the transmission mechanism is used to transmit the certificate from one end of the transmission channel to the other end; the information reading module and the scanning mechanism are both arranged in the transmission direction of the transmission channel.
[0093] The beneficial effects achieved by the present invention are:
[0094] 1. The present invention does not rely on the data itself, does not require a large amount of sample data, does not require a pre-trained model, etc., and can be easily integrated into current software and hardware products.
[0095] 2. The present invention is aimed at separating the portrait of a scanned document, has low requirements on the hardware equipment for image acquisition, and saves hardware debugging time.
[0096] 3. The present invention adopts a multi-scale feature fusion solution to form a new feature fusion matrix, which ensures a high accuracy while reducing the complexity of calculation and improving the inspection accuracy.
[0097] 4. Due to the influence of thermal stability of the image's photosensitive element and the different structure of the device sensor, there will be small noise interference. Preprocessing the portrait through the feature extraction module can effectively eliminate the interference of noise on the image, highlight the portrait pattern, and improve the extraction accuracy of the feature value.
[0098] 5. The present invention extracts the contour feature information of each channel of the image, and customizes the calculation of image feature information through a multi-feature matrix fusion method, thereby ensuring the feature authenticity, integrity and accuracy of the image.
[0099] 6. The portrait segmentation method provided by the present invention has strong versatility and can be used not only in portraits of document scanning modules, but also in other scanned images that require portrait separation. BRIEF DESCRIPTION OF THE DRAWINGS
[0100] Figure 1 It is a schematic flow chart of the portrait segmentation method of the present invention;
[0101] Figure 2 This is a schematic diagram of the document recognition scanning of the present invention
[0102] Figure 3 This is a schematic diagram of the portrait of the certificate of the present invention
[0103] Figure 4 is a schematic diagram of the fusion of portrait features in the portrait segmentation method of the present invention;
[0104] Figure 5 It is a schematic diagram of the external structure of the processing system of the present invention;
[0105] Figure 6 It is a schematic diagram of the internal structure of the processing system of the present invention. DETAILED DESCRIPTION
[0106] The present invention will be further described below in conjunction with the accompanying drawings. The following embodiments are only used to more clearly illustrate the technical solution of the present invention, and cannot be used to limit the protection scope of the present invention.
[0107] Embodiment 1, as Figure 1 As shown, this embodiment provides a method for document portrait segmentation based on multi-scale clustering, comprising the following steps:
[0108] Step 1: Image preprocessing to obtain the document image;
[0109] Select the ROI area on the document image to crop the target image, and perform image enhancement processing on the cropped target image;
[0110] Step 2: Clustering processing: clustering the enhanced image, and generating a new cluster image based on the cluster center and label matrix;
[0111] Step 3: Feature extraction: the newly generated cluster image is subjected to channel separation, binarization, morphological processing, and edge contour detection to obtain edge contour information of the portrait features;
[0112] Calculate maximum contour feature information based on edge contour information;
[0113] Step 4: Feature fusion: Feature fusion fuses the maximum contour feature information into multiple feature matrices to obtain the final segmented portrait information.
[0114] The specific process of step 1 is: for the input scanned image I (x, y), Figure 2 As shown, select a suitable ROI area for cropping. Specifically, perform angle correction and portrait position positioning on the scanned document image. For the same type of document, the position of the portrait is basically fixed in a certain area. Based on the document type information obtained when the document was read previously, the position of the portrait in the current document can be determined. From the scanned image I (x, y), the image containing the portrait T (x, y) is obtained, as shown in Figure 3 As shown in the figure, the obtained portrait image T (x, y) is subjected to image enhancement processing to improve the generalization ability of the image and the subsequent feature detection effect.
[0115] Among them, the image to be processed T(x,y) converts the RGB image to YCrCb space, where is the coordinate of the pixel point. The value of the pixel point is composed of the values of the three channels R, G, and B, that is, any point in the image The value of , converted to YCrCb space image , the calculation formula is:
[0116] Y = 0.299⋅R + 0.587⋅G + 0.114⋅B;
[0117] Cr = 128 + (-0.168736⋅R - 0.331264⋅G + 0.5⋅B);
[0118] Cb =128 + (0.5⋅R - 0.418688⋅G - 0.081312⋅B);
[0119] Where R, G, and B are images respectively. The three-channel value of the pixel point. After performing histogram equalization on the Y channel for image enhancement, the processed image is used The specific steps are as follows:
[0120] 1. Input image: The input image is G(x, y) with a grayscale range of [0, L-1], where L is the number of grayscale levels.
[0121] 2. Gray level probability density function: Calculate the probability density of each gray level k: p(k)= , in, is the number of pixels of gray level k, and N is the total number of pixels in the image.
[0122] 3. Cumulative Distribution Function (CDF): Calculate the cumulative distribution function: C(k)= .
[0123] 4. Grayscale after equalization: Map each grayscale level k to a new grayscale level: T(k) = |(L−1)⋅C(k)|, the output image after histogram equalization: = T(G(x,y)) .
[0124] Finally, the YCrCb image after image enhancement is processed , that is, any point in the image The value of , converted to RGB image The calculation formula is:
[0125] R = Y + 1.402⋅(Cr-128);
[0126] G = Y - 0.344136⋅(Cb–128) - 0.714136⋅(Cr-128);
[0127] B = Y + 1.772⋅(Cb–128);
[0128] According to multiple portrait segmentation experiments, image enhancement processing can significantly improve the subsequent image feature information extraction.
[0129] The mathematical theory adopted in step 2 is as follows:
[0130] Input image after image enhancement , perform K-means clustering processing, K-means can ignore the noise and subtle changes in the image and focus on the main features. And because it is unsupervised learning, there is no need to pre-label data and create data sets. The image after K-means clustering processing is recorded as K(x,y), and the specific steps are as follows:
[0131] 1. Initialization: Randomly select z initial centers (centers of mass) , ,…, .
[0132] 2. Assign pixels: For each data point , calculate its distance from each cluster center and assign it to the cluster with the closest distance. Using Euclidean distance, the calculation formula is as follows:
[0133] = ;
[0134] in, It is a data point The index of the cluster to which it belongs, Represents the square of the Euclidean distance.
[0135] 3. Cluster center update: recalculate the center of each cluster, that is, the mean of all data points in the cluster. The new cluster center calculation formula is:
[0136]
[0137] in, is the number of pixels in the zth cluster, is the indicator function, when 1 if the value is 0, otherwise it is 0.
[0138] 4. Iteration: Repeat steps 2 and 3 until the cluster center no longer changes or the change is less than a preset threshold.
[0139] The goal of the K-means algorithm is to minimize the sum of squared errors within the cluster, and each iteration aims to reduce this objective function.
[0140] The objective function is calculated as follows:
[0141]
[0142] Based on multiple experiments on portrait segmentation, with z=3 and N=10 (maximum number of iterations), the best differentiation effect was achieved.
[0143] Based on the above mathematical theory, this embodiment provides the following processing process:
[0144] In the K-means clustering processing embodiment, a color table colorTab, a sample quantity sampleCount, a sample point matrix points, a cluster quantity clusterCount, an iteration parameter Count, an error threshold EPS, etc. are created.
[0145] Before performing clustering, set the colorTab parameter value to { (0,0,255), (0,255,0), (255,0,0),(0,255,255),(255,0,255),(255,255,0)}. Set sampleCount to the number of pixels in the T(x,y) image; treat each pixel as a sample point. For each pixel (col,row) in the image, set points=row*width+col. Set the clusterCount value to 3, indicating that the pixels in the image are to be clustered into 3 clusters. In the K-means algorithm, k (here clusterCount) is the pre-specified number of clusters.
[0146] Loop through each pixel of the F(x,y) image. For each pixel (col, row) in the image, calculate its index in the points matrix. The calculation formula is index=row×width+col. Divide the RGB value of the pixel into
[0147] Don't assign them to the corresponding rows and columns in the points matrix.
[0148] Execute K-means algorithm to cluster the sample points in the points matrix. Obtain label matrix labels. Generate new cluster image K(x,y) based on cluster center and label matrix.
[0149] The step 3 performs contour feature detection on the image after each channel is separated to obtain the maximum contour feature information of each channel. The specific steps are as follows:
[0150] Among them, the K(x,y) image is channel separated, and the image data after separation is Perform binarization processing, where i is the channel index value, and the binarization calculation formula is as follows:
[0151]
[0152] Where t is the threshold between feature information and background. According to multiple tests of portrait segmentation, the best distinction effect is obtained when t=150;
[0153] After binarization, the image is in black and white. is white, is black, and morphological processing is performed. Specifically, Perform expansion processing and then corrosion processing, where the expansion processing uses a 3×3 matrix template to corrode the entire image Traverse in the horizontal x and vertical y directions. The template is:
[0154] ;
[0155] Where is the corresponding pixel value traversed. The expansion process takes the minimum value in the template as the value of the template midpoint for each traversed template. The erosion process takes the maximum value of the template as the value of the template midpoint for each traversed template. The pixel value of the small white noise in the image is 1. Through this process, since the background around the white noise is black, that is, the pixel value is 0, when the template traverses to the white pixel during the expansion process, it will take the minimum value of 0 due to the existence of the surrounding black pixels and become a black pixel. Then, the expanded image is restored to its original size through the erosion process. The processed image is recorded as .
[0156] Morphologically processed image , perform Canny edge contour detection, the calculation formula is as follows:
[0157] 1. Gaussian filtering: First, perform Gaussian filtering on the image to reduce the impact of noise. The formula of Gaussian filter is:
[0158] ;
[0159] in, is the standard deviation of the Gaussian kernel. Convolve with the Gaussian kernel G(x,y) to get a smoothed image :
[0160] ;
[0161] in, is the original image, is the smoothed image.
[0162] 2. Gradient calculation: Use the Sobel operator to calculate the gradient of the image in the x and y directions. The formula is as follows:
[0163] ;
[0164] ;
[0165] in, and Represents the gradient of the image in the x and y directions respectively. Gradient magnitude G(x,y) and gradient direction The calculation formula is:
[0166] ;
[0167] ;
[0168] 3. Non-maximum suppression: Perform non-maximum suppression on each pixel to refine the edge. For each pixel, according to the gradient direction , and compare it with the two neighboring pixels in the gradient direction. If the pixel is not a local maximum, it is set to 0. is the gradient amplitude after non-maximum suppression.
[0169] ;
[0170] 4. Dual Threshold Detection: Use dual threshold detection to distinguish strong edges from weak edges and suppress non-edges. Set a high threshold and low threshold , the classification rules are as follows:
[0171] Strong edges: pixels with gradient magnitude greater than the high threshold.
[0172] Weak edge: pixels whose gradient magnitude is between the low threshold and the high threshold.
[0173] Non-edge: pixels whose gradient magnitude is less than the low threshold.
[0174] Let E(x,y) be the edge image after double threshold detection:
[0175] ;
[0176] 5. Edge connection:
[0177] By connecting weak edges to strong edges, edge continuity is ensured. All weak edges connected to strong edges are tracked, and these edges are retained, while other weak edges are suppressed. Let the final edge image be P(x,y):
[0178] ;
[0179] If E(x,y) is a strong edge or is connected to a strong edge, the value is 255, otherwise it is 0.
[0180] Image after Canny edge contour detection , where i is the channel index value, and the area method is used to find the current maximum contour feature information Area( ), the contour area calculation formula is:
[0181] ;
[0182] Among them, C is the point set of the contour, are the coordinates of the points on the contour.
[0183] The specific process of step 4 is as follows: Figure 4 As shown, the images of each channel after processing , the corresponding maximum feature contour , where i is the channel index value, which is used as the feature contour mask FeatureMask. The feature portrait data in the original T (x, y) image is copied to the target image in sequence to obtain the final segmented portrait information.
[0184] Through actual detection of 100 ID images, the actual portrait segmentation test accuracy is 100%. After testing, the present invention adopts clustering and multi-scale detection of contour features, and customizes the calculation of image feature information through the method of multi-feature matrix fusion, which has a higher accuracy rate. At the same time, due to the use of feature fusion matrix calculation method, the calculation complexity is not significantly increased while ensuring a high accuracy rate.
[0185] Embodiment 2: This embodiment provides a device for document portrait segmentation based on multi-scale clustering, comprising:
[0186] Image preprocessing module: used to obtain the document image; select the ROI area on the document image to crop the target image, and perform image enhancement processing on the cropped target image;
[0187] Clustering processing module: used to perform clustering processing on the enhanced image and generate a new clustering image based on the cluster center and label matrix;
[0188] Feature extraction module: used to perform channel separation, binarization, morphological processing and edge contour detection on the newly generated cluster image to obtain the edge contour information of the portrait features; calculate the maximum contour feature information based on the edge contour information;
[0189] Feature fusion module: used to fuse the maximum contour feature information into multiple feature matrices to obtain the final segmented portrait information.
[0190] Embodiment 3, as Figure 5 , Figure 6As shown, this embodiment provides a system for processing ID portrait segmentation based on multi-scale clustering, comprising: an information reading module 3, a scanning mechanism 5, and the ID portrait segmentation device based on multi-scale clustering as described in Example 2;
[0191] The information recognition module 3 is used to recognize the certificate and read the information;
[0192] The scanning mechanism 5 is used to scan the image of the document;
[0193] The document scanned image is segmented by the portrait segmentation device to obtain the segmented portrait information.
[0194] It also includes: a first base body 1, a second base body 2, a transmission mechanism 4 and a safety module 6;
[0195] The first base body 1 and the second base body 2 are arranged to cover each other, and a transmission channel A for the certificate to pass through is provided between the first base body 1 and the second base body 2; the information reading module 3, the transmission mechanism 4, the scanning mechanism 5 and the security module 6 are all arranged in the cavity covered by the first base body 1 and the second base body 2;
[0196] The document is inserted from one end of the transmission channel A, and the transmission mechanism 4 is used to transmit the document from one end of the transmission channel A to the other end; the information reading module 3, the scanning mechanism 5 and the security module 6 are all arranged in the transmission direction of the transmission channel A. The security module 6 is used to ensure information security.
[0197] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the technical principles of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.
[0198] A computer-readable storage medium storing one or more programs, wherein the one or more programs include instructions, and when the instructions are executed by a computing device, the computing device executes a document portrait segmentation method based on multi-scale clustering.
[0199] A computing device includes one or more processors, one or more memories, and one or more programs, wherein the one or more programs are stored in the one or more memories and are configured to be executed by the one or more processors, and the one or more programs include instructions for executing a document image segmentation method based on multi-scale clustering.
[0200] It will be appreciated by those skilled in the art that embodiments of the present invention may be provided as methods, systems, or computer program products. Therefore, the present invention may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0201] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0202] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0203] These computer program instructions may also be loaded onto a computer or other programmable data processing device so that a series of processing steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for executing on the computer or other programmable device to implement the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0204] The above are merely embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention are included in the scope of the claims of the present invention to be approved.
Claims
1. A method for document portrait segmentation based on multi-scale clustering, characterized in that: include: Get the document image; Select the ROI area on the document image to crop the target image, and perform image enhancement processing on the cropped target image; Perform clustering on the enhanced image and generate a new cluster image based on the cluster center and label matrix; The newly generated cluster image is subjected to channel separation, binarization, morphological processing and edge contour detection to obtain edge contour information of portrait features; Calculate maximum contour feature information based on edge contour information; The maximum contour feature information is fused into a multi-feature matrix to obtain the final segmented portrait information.
2. The method for document portrait segmentation based on multi-scale clustering according to claim 1, characterized in that: Select the ROI area on the document image to crop the target image, and perform image enhancement processing on the cropped target image as follows: Suppose the document image is a function I(x,y), where (x,y) is the coordinate on the image plane, and the ROI area is a rectangular area R={(x,y)|x1≤x≤x2,y1≤y≤y2}, where x1,x2,y1,y2 are the coordinate values defining the ROI area; the cropped target area image T(x,y); The target area image T(x, y) is subjected to histogram equalization, the RGB image is converted to the YCrCb space, and the Y channel is equalized; the processed image F(x, y) is obtained.
3. The method for document portrait segmentation based on multi-scale clustering according to claim 2, characterized in that: The target area image T(x,y) is subjected to histogram equalization, the RGB image is converted to the YCrCb space, and the Y channel is equalized; the method to obtain the processed image F(x,y) is as follows: Convert the RGB image to YCrCb space, where is the coordinate of the pixel point; the value of the pixel point is composed of the values of the three channels R, G, and B. The value of , converted to YCrCb space image , the calculation formula is: Y = 0.299⋅R + 0.587⋅G + 0.114⋅B; Cr = 128 + (-0.168736⋅R - 0.331264⋅G + 0.5⋅B); Cb =128 + (0.5⋅R - 0.418688⋅G - 0.081312⋅B); Where R, G, and B are images respectively. The three-channel value of the pixel; The input image is G(x, y), and the grayscale range is [0, L-1], where L is the number of grayscale levels; Calculate the probability density of each gray level k: p(k)= , in, is the number of pixels of gray level k, and N is the total number of pixels in the image; Calculate the cumulative distribution function: C(k)= ; Map each gray level k to a new gray level: T(k) = |(L−1)⋅C(k)|, and the output image after histogram equalization is: = T(G(x,y)); The YCrCb image after image enhancement , that is, any point in the image The value of , converted to RGB image ; The calculation formula is: R = Y + 1.402⋅(Cr-128); G = Y - 0.344136⋅(Cb–128) - 0.714136⋅(Cr-128); B = Y + 1.772⋅(Cb–128).
4. The method for document portrait segmentation based on multi-scale clustering according to claim 2, characterized in that: The enhanced image is clustered and the method of generating a new cluster image based on the cluster center and label matrix is as follows: Create a color table colorTab, calculate the number of samples sampleCount, sample point matrix points, the number of clusters clusterCount, iteration parameter Count, and error threshold EPS; Set the colorTab parameter value to { (0,0,255), (0,255,0), (255,0,0),(0,255,255),(255,0,255),(255,255,0)}; Set sampleCount to the number of pixels in the T(x,y) image; Treat each pixel as a sample point. For each pixel (col,row) in the image, set points=row*width+col. Set the clusterCount value to 3, indicating that the pixels in the image are to be clustered into 3 clusters. In the K-means algorithm, k is the pre-specified number of clusters. Loop through each pixel of the F(x,y) image; for each pixel (col,row) in the image, calculate its index in the points matrix, the calculation formula is index=row×width+col; divide the RGB value of the pixel; Don't assign them to the corresponding rows and columns in the points matrix; Execute the K-means algorithm to cluster the sample points in the points matrix to obtain the label matrix labels; generate a new cluster image K(x,y) based on the cluster center and the label matrix.
5. The method for document portrait segmentation based on multi-scale clustering according to claim 4, characterized in that: The newly generated cluster image is subjected to channel separation, binarization, morphological processing, and edge contour detection to obtain the edge contour information of the portrait features as follows: The cluster image K(x,y) is channel-separated, and the separated image K i (x,y); Where i is the channel index value, and the binarization calculation is: ; Among them, t is the threshold between feature information and background; After binarization, the image is in black and white. is white, is black; The morphological processing includes Perform expansion treatment, then corrosion treatment; The dilation process uses a 3×3 matrix template to dilate the entire image. Traverse in the horizontal x and vertical y directions, the template is: ; Where D is the corresponding pixel value traversed; the expansion process is used to take the minimum value in the template as the value of the midpoint of the template for each traversed template; the corrosion process is used to take the maximum value of the template as the value of the midpoint of the template for each traversed template; the pixel value of the small white noise in the image is 1. Through this process, since the background around the white noise is black, that is, the pixel value is 0, when the template traverses to the white pixel during the expansion process, it will take the minimum value 0 due to the existence of the surrounding black pixels and become a black pixel. Then, the corrosion process is used to restore the expanded image to its original size. The processed image is recorded as ; Morphologically processed image , perform Canny edge contour detection; The edge contour detection comprises: Perform Gaussian filtering on the image to reduce the impact of noise; the formula of the Gaussian filter is: ; in, is the standard deviation of the Gaussian kernel, image Convolve with the Gaussian kernel G(x,y) to get a flat Slide image : ; in, is the original image, is the smoothed image; The Sobel operator is used to calculate the gradient of the image in the x and y directions. The formula is as follows: ; ; in, and Represents the gradient of the image in the x and y directions respectively; the gradient magnitude G(x,y) and the gradient direction The calculation formula is: ; ; For each pixel, according to the gradient direction , compared with the two neighboring pixels in the gradient direction; if the pixel is not a local maximum, it is set to 0; set is the gradient amplitude after non-maximum suppression; ; Use dual threshold detection to distinguish strong edges from weak edges and suppress non-edges; Set high threshold and low threshold , the classification rules are as follows: Strong edge: pixels with gradient magnitude greater than the high threshold; Weak edge: pixels with gradient magnitude between the low threshold and the high threshold; Non-edge: pixels whose gradient magnitude is less than the low threshold; Let E(x,y) be the edge image after double threshold detection: ; By connecting weak edges to strong edges, the continuity of the edges is ensured; all weak edges connected to strong edges are tracked, these edges are retained, and other weak edges are suppressed; let the final edge image be P(x,y): ; If E(x,y) is a strong edge or is connected to a strong edge, the value is 255, otherwise it is 0.
6. The method for document portrait segmentation based on multi-scale clustering according to claim 5, characterized in that: The method for calculating the maximum contour feature information based on the edge contour information is: Image after edge contour detection , where i is the channel index value, and the area method is used to find the current maximum contour feature information Area( ), the contour area calculation formula is: ; Among them, C is the point set of the contour, are the coordinates of the points on the contour.
7. The method for document portrait segmentation based on multi-scale clustering according to claim 6, characterized in that: The method of fusing the feature information into a multi-feature matrix to obtain the final segmented portrait information is as follows: Image The corresponding maximum feature contour , where i is the channel index value, which is used as the feature contour mask FeatureMask to copy the feature portrait data in the target area image T (x, y) to the maximum feature contour in sequence to obtain the final segmented portrait information.
8. A device for document portrait segmentation based on multi-scale clustering, characterized in that: include: Image preprocessing module: used to obtain document images; Select the ROI area on the document image to crop the target image, and perform image enhancement processing on the cropped target image; Clustering processing module: used to perform clustering processing on the enhanced image and generate a new clustering image based on the cluster center and label matrix; Feature extraction module: used to perform channel separation, binarization, morphological processing and edge contour detection on the newly generated cluster image to obtain the edge contour information of the portrait features; calculate the maximum contour feature information based on the edge contour information; Feature fusion module: used to fuse the maximum contour feature information into multiple feature matrices to obtain the final segmented portrait information.
9. A system for processing ID card portrait segmentation based on multi-scale clustering, characterized in that: include: An information reading module (3), a scanning mechanism (5), and a device for document portrait segmentation based on multi-scale clustering as claimed in claim 8; The information recognition module (3) is used to recognize the certificate and read the information; The scanning mechanism (5) is used to perform image scanning on the certificate; The document scanned image is segmented by the portrait segmentation device to obtain the segmented portrait information.
10. The system for document portrait segmentation based on multi-scale clustering according to claim 9, characterized in that: Also includes: A first seat body (1), a second seat body (2) and a transmission mechanism (4); The first base body (1) and the second base body (2) are arranged to cover each other, and a transmission channel (A) for the certificate to pass through is provided between the first base body (1) and the second base body (2); the information reading module (3), the transmission mechanism (4) and the scanning mechanism (5) are all arranged in the cavity covered by the first base body (1) and the second base body (2); The document is inserted from one end of the transmission channel (A), and the transmission mechanism (4) is used to transmit the document from one end of the transmission channel (A) to the other end; the information reading module (3) and the scanning mechanism (5) are both arranged in the transmission direction of the transmission channel (A).