Positioning correction method and device for acquiring handheld identity card information, electronic equipment and storage medium

By combining target detection, Hough linear detection, K-Means clustering, perspective transformation and ResNet character recognition technology, the problem of unstandard boundaries of ID card image and low accuracy of information extraction is solved, and the precise positioning and high accuracy recognition of ID card information areas are achieved.

CN120147199APending Publication Date: 2025-06-13ANHUI UNIVERSITY OF TECHNOLOGY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510317774.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-18
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

In the prior art, when taking ID card images automatically identify ID card information, the boundaries of the ID card image are not standard due to the shooting angle, resulting in low accuracy of boundary detection, affecting the precise positioning of the ID card information area, and information extraction depends on the clarity and distribution of text lines. If the text lines are not clear or are unevenly distributed, it may lead to judgment errors.

Method used

The target recognition model (such as YOLOv8) is used to locate the ID card area, and combined with Hough linear detection and constrained K-Means clustering algorithm, the parallel line group that is bounded into the largest area of ​​the ID card is selected from the straight line, and the image is corrected through perspective transformation to achieve accurate positioning and standardized processing of the ID card area. Then, through character segmentation and ResNet character recognition technology, the ID card characters are extracted and identified.

Benefits of technology

It improves the accuracy of ID card image boundary detection, realizes high-precision positioning and correction of ID card information area, enhances the accuracy and robustness of character recognition, reduces manual intervention, and meets the requirements of automation, accuracy and high robustness of ID card information acquisition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120147199A_ABST
    Figure CN120147199A_ABST
Patent Text Reader

Abstract

The invention relates to a positioning correction method for acquiring handheld identity card information, which comprises the following steps: acquiring a first identity card image, and acquiring a straight line in the first identity card image; processing the straight lines to obtain a first parallel line group and a second parallel line group; processing the first parallel line group and the second parallel line group to obtain a second identity card image; processing the second identity card image to obtain identity card characters; and identifying the identity card characters to obtain identity card information. The invention further relates to a positioning correction device for acquiring the handheld identity card information, electronic equipment and a storage medium. By utilizing the positioning correction method and device for acquiring the handheld identity card information, the electronic equipment and the storage medium, the problems that the boundary of an identity card image is not standard, the accuracy of identity card boundary detection is low during information detection, the subsequent positioning of an identity card information area is influenced, and the accuracy of information detection is low can be solved. And when the information is extracted, text lines are not clear or are not uniformly distributed, so that the information is wrongly judged.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer image processing, and particularly to a positioning and calibration method, device, electronic device and storage medium for obtaining information of a handheld identity card. Background Art

[0002] In China, the identity card is the only identifier of Chinese citizens' identity information, which contains basic information of citizens and is widely used in personal and social interactions. When Chinese citizens handle business at stations, airports, hospitals, schools, government offices, banks, etc., the identity card is always an essential document. Sometimes, the entry of identity card information is completed manually, but manual entry of information is not only inefficient but also prone to errors. In order to reduce the adverse effects brought by the traditional entry method, currently, the identity card information is mostly automatically recognized by taking pictures of the identity card image, so as to realize the extraction of identity card information. However, due to the influence of shooting angles and other factors, the boundaries of the identity card image are not standard, and the accuracy rate of identity card boundary detection during information detection is relatively low, which will affect the subsequent precise positioning of the identity card information area. At the same time, the extraction of information depends on the clarity and distribution of text lines. If the text lines are not clear or unevenly distributed during shooting, it may lead to incorrect judgments.

[0003] Therefore, the prior art has defects and needs to be improved and developed. Summary of the Invention

[0004] Embodiments of the invention provide a positioning and calibration method, device, electronic device and storage medium for obtaining information of a handheld identity card, which are used to solve the problem that in the prior art, the identity card information is mostly automatically recognized by taking pictures of the identity card image, so as to realize the extraction of identity card information. However, due to the influence of shooting angles and other factors, the boundaries of the identity card image are not standard, and the accuracy rate of identity card boundary detection during information detection is relatively low, which will affect the subsequent precise positioning of the identity card information area. At the same time, the extraction of information depends on the clarity and distribution of text lines. If the text lines are not clear or unevenly distributed during shooting, it may lead to incorrect judgments.

[0005] In the first aspect of the embodiments of the present invention, a positioning correction method for obtaining handheld ID card information is provided, including: collecting a first ID card image obtained based on a target recognition model; detecting the first ID card image according to the Hough line detection method to obtain the lines in the first ID card image; processing the lines to obtain a first parallel line group and a second parallel line group for enclosing the largest area in the first ID card image, where there are two parallel lines in each of the first parallel line group and the second parallel line group, and the parallel lines in the first parallel line group are perpendicular to the parallel lines in the second parallel line group; processing the first parallel line group and the second parallel line group to obtain a second ID card image; processing the second ID card image to obtain ID card characters; and recognizing the ID card characters to obtain ID card information.

[0006] Further, the collecting a first ID card image obtained based on a target recognition model includes: obtaining an original image with an ID card; performing ID card image annotation on the original image using the labelImg annotation tool, and dividing the annotated original image into a training set, a test set, and a validation set according to a ratio of 8:1:1; importing the training set into a target recognition model using the YOLOv8 network model for training to obtain a trained target recognition model; importing the validation set into the trained target recognition model, recording the number of iterations and the mean average precision, to obtain a validated target recognition model; importing the test set into the validated target recognition model, and ending the test when either the number of iterations reaches the maximum number of iterations or the mean average precision reaches the early stopping condition, to obtain a target recognition model that passes the test; and importing the ID card image to be detected into the target recognition model that passes the test to obtain the first ID card image.

[0007] Further, the original image is an original image after data augmentation, and the data augmentation method of the original image includes at least one of rotating, moving, and flipping the original image.

[0008] Further, the processing the lines to obtain a first parallel line group and a second parallel line group for enclosing the largest area in the first ID card image, where there are two parallel lines in each of the first parallel line group and the second parallel line group, and the parallel lines in the first parallel line group are perpendicular to the parallel lines in the second parallel line group, includes: using K-Means clustering for grid search of the lines to obtain a line angle tolerance and a line distance tolerance with the highest line classification accuracy, the highest silhouette coefficient, the smallest angle standard deviation and distance standard deviation, and the highest integrity of the corrected ID card area; using constrained K-Means clustering to process the line angle tolerance and the line distance tolerance to obtain the distance D(x i ,x j), the processing formula is as follows:

[0009]

[0010] Among them, x i represents the currently judged straight line;

[0011] x j represents the currently judged cluster center;

[0012] slope i represents the angle of x i ;

[0013] distance i represents the distance of x i ;

[0014] slope j represents the angle of x j ;

[0015] distance j represents the distance of x j ;

[0016] tol_slope represents the tolerance of the straight line angle;

[0017] tol_dis represents the tolerance of the straight line distance;

[0018] Using constrained K-Means clustering, classify the straight lines with an angle difference of 5° and a distance D(x i , x j ) with a difference of 20 px; select the horizontal and vertical straight lines from the classified straight lines; select four straight lines from the horizontal and vertical straight lines to enclose the largest area in the first ID card image; divide a set of parallel straight lines among the four straight lines into the first parallel line group, and divide the other set of parallel straight lines into the second parallel line group.

[0019] Further, processing the first parallel line group and the second parallel line group to obtain a second ID card image includes: extending the straight lines of the first parallel line group and the second parallel line group to obtain four intersection points and the coordinates of the four intersection points; extending the upper, lower, left, and right boundaries of the standard ID card to obtain four vertices and the coordinates of the four vertices; using the perspective transformation method, matching the four intersection points to the four vertices to correct the first ID card image into the second ID card image that is non-tilted, non-deformed, and completely fills the entire picture.

[0020] Further, processing the second ID card image to obtain ID card characters includes: scaling the second ID card image to the standard ID card size using bilinear interpolation; cropping a positioning area from the second ID card image according to the absolute position (x, y, w, h) of the ID card characters, where x, y, w, and h respectively represent the upper left vertex coordinates and the area width and height of the positioning area on the standard ID card; cutting the positioning area to obtain a first image; performing grayscale processing on the first image to obtain a second image; using Otsu's method for binarization to process the second image so that the image of the positioning area only has two gray levels of 0 and 255, obtaining a third image with black background and white characters; calculating the sum of the gray values of all pixel points in each row of the matrix in the vertical and horizontal directions respectively as sum row and the sum of the gray values of all pixel points in each column as sum col At the same time, calculate the number of pixel points with a gray value of 255 in each row as number row and the number of pixel points with a gray value of 255 in each column as number col The calculation formula in the horizontal direction is the same as that in the vertical direction. The calculation formula in the vertical direction is as follows:

[0021]

[0022] According to the sum of the gray values of all pixel points in each row as sum row and the number of pixel points with a gray value of 255 in each row as number row obtain a horizontal frequency distribution graph; according to the sum of the gray values of all pixel points in each column as sum col and the number of pixel points with a gray value of 255 in each column as number col obtain a vertical frequency distribution graph; according to the horizontal frequency distribution graph and the vertical frequency distribution graph, segment each character of the single-line text to obtain the ID card characters.

[0023] Further, recognizing the ID card characters to obtain ID card information includes: using the PIL library in Python to generate the ID card characters into ID card character images; performing at least one of the processing methods of angle rotation, erosion, dilation, and adding noise on the ID card character images to obtain ID card character images after data augmentation, where the angle of angle rotation is one of 5° or -5°, and the dilation and erosion kernels for dilation and erosion are both 3*3, and the number of iterations is 1; using the ResNet residual neural network to train and recognize the ID card character images after data augmentation to obtain ID card information.

[0024] In the second aspect of the embodiments of the present invention, a positioning and calibration device for obtaining information of a handheld identity card is provided, including: a collection module for collecting a first identity card image obtained based on a target recognition model; a detection module for detecting the first identity card image according to the Hough line detection method to obtain the lines in the first identity card image; a first processing module for processing the lines to obtain a first parallel line group and a second parallel line group for enclosing the largest area in the first identity card image, where there are two parallel lines in each of the first parallel line group and the second parallel line group, and the parallel lines in the first parallel line group are perpendicular to the parallel lines in the second parallel line group; a second processing module for processing the first parallel line group and the second parallel line group to obtain a second identity card image; a third processing module for processing the second identity card image to obtain identity card characters; and an identification module for identifying the identity card characters to obtain identity card information.

[0025] In the third aspect of the embodiments of the present invention, an electronic device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the computer program is executed by the processor, the positioning and calibration method for obtaining information of a handheld identity card as described above is implemented.

[0026] In the fourth aspect of the embodiments of the present invention, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the positioning and calibration method for obtaining information of a handheld identity card as described above is implemented.

[0027] Beneficial effects:

[0028] As can be seen from the above technical solutions, the present invention provides a positioning and calibration method, device, electronic device, and storage medium for obtaining information of a handheld identity card:

[0029] 1. By organically combining target detection, Hough line detection, constrained K-Means clustering, perspective transformation, character segmentation, and ResNet character recognition technologies, precise positioning, calibration, and information extraction of the identity card area in the handheld identity card image are achieved. Aiming at the problems of complex background and variable shooting angles of identity card images, the Yolov8 target detection algorithm is used for positioning, and combined with the Hough transform and the constrained K-Means clustering algorithm, the four sides of the identity card needed are obtained from the cluttered lines. Finally, perspective transformation is used to correct the tilt to achieve high-precision cropping and robust positioning. At the same time, the boundary is estimated using the character position relationship and prior information to improve the segmentation accuracy, and the ResNet network is introduced to optimize character recognition, significantly improving the recognition accuracy.

[0030] 2. By adopting data augmentation techniques and a multi-level model training scheme, the robustness and generalization ability of the model are improved in the processing of the original image and the object detection link; in the straight line processing link, the straight line angle and distance tolerance parameters are optimized through grid search, making the clustering results more accurate, and then obtaining the straight line combination that encloses the largest area; in the character segmentation and recognition process, the character region is divided by using the horizontal and vertical frequency distribution maps, and the accurate recognition of characters is realized through data augmentation and the deep residual network.

[0031] 3. On the premise of maintaining the objectivity and accuracy of the processing flow, the overall method realizes the automatic correction and information extraction of ID card images under the conditions of inclination, deformation and background interference, which is beneficial to reducing manual intervention and meeting the requirements of automation, accuracy and high robustness for ID card information acquisition in practical applications.

[0032] It should be understood that all combinations of the foregoing concepts and additional concepts described in more detail below can be regarded as part of the inventive subject matter of the present disclosure as long as such concepts do not contradict each other.

[0033] The foregoing and other aspects, embodiments and features of the teachings of the present invention can be more fully understood from the following description in conjunction with the accompanying drawings. Other additional aspects of the present invention, such as the features and / or beneficial effects of exemplary embodiments, will be apparent in the following description, or will be learned through the practice of specific embodiments according to the teachings of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] The drawings are not drawn to scale in accordance with the true reference. In the drawings, each identical or nearly identical component shown in each figure may be denoted by the same reference numeral. For clarity, not every component is labeled in each figure. Now, embodiments of various aspects of the present invention will be described by way of example and with reference to the drawings, wherein:

[0035] Figure 1 is a flowchart of a positioning and correction method for obtaining information of a handheld ID card in an embodiment of the present application.

[0036] Figure 2 is an image obtained by inputting the ID card image to be detected in an embodiment of the present application into a tested model.

[0037] Figure 3 is an image obtained by positioning and cropping after Canny edge detection in an embodiment of the present application.

[0038] Figure 4 is an image obtained by detecting straight lines in the image using Hough straight line detection in an embodiment of the present application.

[0039] Figure 5This is the image obtained after applying the perspective transformation method in the embodiments of this application.

[0040] Figure 6 This is the image after obtaining the first image from the cutting and positioning area in the embodiments of this application. Detailed implementation manners

[0041] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the described embodiments of the present invention without creative efforts shall fall within the scope of protection of the present invention. Unless otherwise defined, the technical terms or scientific terms used herein shall have the ordinary meanings understood by those of ordinary skill in the art in the field to which the present invention pertains.

[0042] The terms "first", "second", and similar terms used in the specification and claims of this patent application for the present invention do not denote any order, quantity, or importance, but are only used to distinguish different components. Similarly, unless the context clearly indicates otherwise, singular forms such as "a", "an", or "the" do not denote a limitation of quantity, but rather indicate the presence of at least one. The terms "including" or "comprising" and similar terms mean that the elements or items appearing before "including" or "comprising" cover the features, wholes, steps, operations, elements, and / or components listed after "including" or "comprising", and do not exclude the existence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations. The terms "up", "down", "left", "right", etc. are only used to represent relative positional relationships, and when the absolute position of the object being described changes, the relative positional relationship may also change accordingly.

[0043] In the prior art, since most methods use the way of photographing ID card images to automatically identify ID card information, so as to extract the ID card information. However, due to the influence of factors such as the shooting angle, the boundaries of the ID card images are not standard, and the accuracy of ID card boundary detection during information detection is relatively low, which will affect the subsequent precise positioning of the ID card information area. At the same time, when extracting information, it depends on the clarity and distribution of text lines. If the text lines are not clear or unevenly distributed during shooting, it may lead to the problem of incorrect judgment.

[0044] In view of this, the embodiments of the present invention provide a positioning and correction method for obtaining handheld ID card information. Refer to Figure 1 , including:

[0045] Step S102: Collect the first ID card image obtained based on the target recognition model.

[0046] Step S104: Detect the first identity card image according to the Hough line detection method to obtain the lines in the first identity card image.

[0047] Step S106: Process the lines to obtain a first parallel line group and a second parallel line group for enclosing the largest area in the first identity card image. There are two parallel lines in each of the first parallel line group and the second parallel line group, and the parallel lines in the first parallel line group are perpendicular to the parallel lines in the second parallel line group.

[0048] Step S108: Process the first parallel line group and the second parallel line group to obtain a second identity card image.

[0049] Step S110: Process the second identity card image to obtain identity card characters.

[0050] Step S112: Identify the identity card characters to obtain identity card information.

[0051] The embodiment of the present invention provides a positioning and correction method for obtaining handheld identity card information. By using a target detection model, such as YOLOv8, the identity card area is extracted from the original image to ensure that subsequent processing is based on relatively accurate area information. The Hough line detection method is used to detect the lines in the image, and by obtaining the parameters of the lines, such as the angle and distance between the line and the horizontal axis, etc., it provides a basis for subsequent clustering and grouping. The K-Means clustering with constraints is used to process the lines, screening out two groups of lines that are parallel to each other and perpendicular to each other, and ensuring that the area enclosed by the selected lines is the largest. Using the perspective transformation technology, the four intersection points determined according to the selected lines are mapped to a preset standard identity card rectangle to achieve image correction. The corrected image is further processed, including cropping, grayscale conversion, binarization, and projection statistics, to obtain each character area; finally, the characters are recognized through a deep neural network to obtain identity card information.

[0052] Through the above method, it is ensured that in the case of image tilt, partial occlusion, or complex background, the identity card area can still be accurately located, and a standardized image can be obtained through correction. By combining line clustering and perspective transformation, the accuracy of determining the identity card boundary is improved, providing a reliable basis for character positioning. By adopting multi-level processing and data enhancement techniques, the character recognition has high accuracy and robustness.

[0053] In some embodiments, collecting the first identity card image obtained based on the target recognition model includes:

[0054] Obtain the original image with the identity card.

[0055] The ID card images are labeled in the original images using the labelImg annotation tool, and the labeled original images are divided into a training set, a test set, and a validation set according to a ratio of 8:1:1.

[0056] The training set is imported into the object recognition model using the YOLOv8 network model for training to obtain the trained object recognition model.

[0057] The validation set is imported into the trained object recognition model, the number of iterations and the mean average precision are recorded, and the validated object recognition model is obtained.

[0058] The test set is imported into the validated object recognition model. The test ends when either the number of iterations reaches the maximum number of iterations or the mean average precision reaches the early stopping condition. The purpose of ending the test when either the number of iterations reaches the maximum number of iterations or the mean average precision reaches the early stopping condition during the training process of machine learning and deep learning models is to control the termination condition of the training process to prevent overfitting or excessive training time. The maximum number of iterations is a pre-set number. For example, if the set maximum number of training rounds is 100, when the number of training rounds reaches 100, even if the model is still being optimized, the training will be forced to stop to prevent waste of computing resources or overfitting, so that the training time is controlled and infinite training is prevented. The early stopping condition is based on the early stopping mechanism, that is, when the performance metrics of the model, such as the mean average precision mAP, do not improve for multiple consecutive cycles on the validation set, the training is automatically terminated. The mean average precision is a commonly used evaluation metric in object detection tasks to measure the detection accuracy of the model for target classes. During the training process, if the mAP does not improve significantly or even decreases in multiple consecutive training rounds, it means that the model may have reached the best state, and continuing training may lead to overfitting. By setting the early stopping condition, overfitting of the model is prevented, the generalization ability is improved, waste of computing resources is avoided, the training convergence speed is accelerated, and the model is prevented from falling into local optimum or overfitting.

[0059] The ID card image to be detected is imported into the object recognition model that has passed the test to obtain the first ID card image.

[0060] The ID card image to be detected is input into the model that has passed the test, and the detected ID card area is cropped to obtain a more accurate ID card area, as Figure 2 shown. The mosaic part in the figure is the pixelation protection for privacy protection.

[0061] The captured ID card will inevitably be tilted more or less in the picture, so the ID card needs to be corrected for tilt. First, the first ID card image can be located and cropped, and Canny edge detection is used to obtain as Figure 3The shown picture is then used to detect the straight lines in the image by Hough straight line detection and record them, as Figure 4 shown.

[0062] First, obtain the original image containing the ID card, and use the annotation tool to mark the ID card area; according to the annotated data, divide the data set according to a preset ratio to provide sufficient samples for model training; use the YOLOv8 network model to train on the training set, monitor the number of iterations and the mean average precision through the validation set to achieve early stopping control; when the upper limit of the number of iterations is reached or the precision reaches the predetermined standard, use the test set to verify the model performance, and finally use the model that passes the test to process the image to be detected, so as to obtain the first ID card image. Through the data set division and model training process, it is ensured that the object detection model has high accuracy and generalization ability; the iterative control of the model, including the maximum number of iterations and the early stopping condition, avoids overfitting and improves the stability of the model. A training set is constructed by combining data augmentation and annotation tools, and a stable object detection model is obtained through a strict training, validation, and testing process; the number of iterations and the mean average precision are introduced as early stopping conditions to realize the automatic control of the model training process and ensure that the object detection model meets the requirements of practical applications.

[0063] In some embodiments, the original image is the original image after data augmentation, and the data augmentation method of the original image includes at least one of rotating, moving, and flipping the original image.

[0064] Perform rotation, translation, or flipping operations on the original image to expand the training data set by increasing sample diversity; data augmentation helps to improve the adaptability of the model to different shooting angles and image deformations, thereby enhancing the robustness of object detection and subsequent recognition. The processing of the original image improves the model's ability to recognize ID card images under different postures, illuminations, or noise interferences; increasing the diversity of training data reduces the risk of overfitting that may be caused by a single sample. Applying conventional data augmentation methods to the preprocessing of ID card images ensures that the object detection model still maintains a high recognition accuracy when facing the image diversity in the actual environment.

[0065] In some embodiments, processing the straight lines to obtain a first parallel line group and a second parallel line group for enclosing the largest area in the first ID card image, where there are two parallel lines in each of the first parallel line group and the second parallel line group, and the parallel lines in the first parallel line group are perpendicular to the parallel lines in the second parallel line group, includes:

[0066] Use K-Means clustering to perform a grid search on straight lines. When conducting the grid search, define multiple candidate values for the straight line angle tolerance and the straight line distance tolerance, and test different combinations one by one. Finally, obtain the straight line angle tolerance and the straight line distance tolerance with the highest straight line classification accuracy, the highest silhouette coefficient, the smallest angle standard deviation and distance standard deviation, and the highest integrity of the corrected ID card area. The highest straight line classification accuracy means that the four boundary straight lines of the ID card are correctly classified with the highest accuracy; the highest silhouette coefficient means that the straight lines within the same cluster are the most similar and the straight lines between different clusters are the most separated; the smallest angle standard deviation and distance standard deviation ensure that the boundary straight lines of the ID card are parallel and the positions are stable; the highest integrity of the corrected ID card area means that the ID card area after perspective transformation is not cropped, deformed, or the boundary is offset. The final optimized parameters are determined through grid search. By comparing the effects of different parameter combinations, the optimal solution is selected. Through these criteria, it can be ensured that the straight line classification result obtained by K-Means clustering is optimal, thus providing high-quality boundary information for subsequent perspective transformation and improving the accuracy of ID card positioning and correction.

[0067] Use constrained K-Means clustering to process the straight line angle tolerance and the straight line distance tolerance, and obtain the distance D(x i ,x j ) from the i-th straight line to the j-th cluster center. The processing formula is as follows:

[0068]

[0069] Among them, x i represents the currently judged straight line;

[0070] x j represents the currently judged cluster center;

[0071] slope i represents the angle of x i ;

[0072] distance i represents the distance of x i ;

[0073] slope j represents the angle of x j ;

[0074] distance j represents the distance of x j ;

[0075] tol_slope represents the straight line angle tolerance;

[0076] tol_dis represents the straight line distance tolerance;

[0077] Using constrained K-Means clustering, classify straight lines with an angular difference of 5° and a distance D(x i ,x j ) difference of 20 px; select horizontal and vertical straight lines from the classified straight lines; select four straight lines from the horizontal and vertical straight lines that enclose the largest area in the first ID card image; divide a group of parallel straight lines among the four straight lines into the first parallel line group, and the other group of parallel straight lines into the second parallel line group.

[0078] Perform constrained K-Means clustering on the straight lines detected by the Hough straight line detection, regard straight lines with similar angles and distances as one category, and at the same time find the parallel lines in each category, and then judge the obtained pairs of parallel lines to obtain two pairs of parallel lines that can be perpendicular to each other and enclose the largest area.

[0079] Through parameter optimization and constraint conditions, ensure the stability of the straight line clustering result and effectively distinguish the straight lines that form the ID card boundary; ensure that the straight lines on which the correction process is based can accurately reflect the true boundary of the ID card, thereby improving the accuracy of the perspective transformation. Combining the grid search and the constrained K-Means clustering method, a optimized classification scheme based on the straight line angle and distance is proposed; introduce specific numerical criteria to constrain the straight line classification, ensure that the obtained straight line group can be used to construct the ID card boundary of the largest area, and achieve high-precision image correction.

[0080] In some embodiments, processing the first parallel line group and the second parallel line group to obtain a second ID card image includes:

[0081] Extend the straight lines of the first parallel line group and the second parallel line group to obtain four intersection points and the coordinates of the four intersection points.

[0082] Extend the upper, lower, left, and right boundaries of the standard ID card to obtain four vertices and the coordinates of the four vertices.

[0083] Adopt the perspective transformation method to match the four intersection points to the four vertices to correct the first ID card image into a second ID card image that is non-tilted, non-deformed and completely fills the entire picture.

[0084] Through the first parallel line group and the second parallel line group, obtain the four boundary line segments of the ID card, extend them to obtain the coordinates of the four intersection points, and adopt the perspective transformation method to match the four intersection points to the four vertices to correct the first ID card image into Figure 5 as shown in the second ID card image, Figure 5The four vertices of the middle picture correspond to four intersection points. Through effective correction of the tilt and deformation caused by the shooting angle in the handheld ID card image, the corrected image completely fills the entire picture, providing an accurate basis for subsequent character positioning and recognition. A perspective transformation method that matches the intersection points with the standard vertices is adopted to automate the complex image correction process; the extension of straight lines and the calculation of intersection points are used to ensure that the selected boundary conforms to the actual ID card edge features, improving the image correction accuracy.

[0085] In some embodiments, processing the second ID card image to obtain ID card characters includes:

[0086] The second ID card image is scaled to the standard ID card size using the bilinear interpolation method.

[0087] According to the absolute position (x, y, w, h) of the ID card characters, the positioning area is cropped from the second ID card image, where x, y, w, and h respectively represent the upper left vertex coordinates and the area width and height of the positioning area on the standard ID card.

[0088] The positioning area is cut to obtain a first image, as Figure 6 shown.

[0089] The first image is grayscale processed to obtain a second image.

[0090] The second image is processed using Otsu's method for binarization so that the image of the positioning area only has two gray levels of 0 and 255, obtaining a third image with black background and white characters.

[0091] The sum sum of the gray values of all pixel points in each row of the matrix is calculated separately in the vertical and horizontal directions row , and the sum sum of the gray values of all pixel points in each column col , and at the same time, the number number of pixel points with a gray value of 255 in each row is calculated row , and the number number of pixel points with a gray value of 255 in each column col . The calculation formula in the horizontal direction is the same as that in the vertical direction. The calculation formula in the vertical direction is as follows:

[0092]

[0093] According to the sum sum of the gray values of all pixel points in each row row , and the number number of pixel points with a gray value of 255 in each row row , a horizontal frequency distribution graph is obtained; according to the sum sum of the gray values of all pixel points in each column col , and the number number of pixel points with a gray value of 255 in each column col, a vertical frequency distribution map is obtained; based on the horizontal frequency distribution map and the vertical frequency distribution map, each character of the single-line text is segmented to obtain ID card characters. The horizontal frequency distribution map and the vertical frequency distribution map are used for character segmentation, and the boundaries of the character regions are determined by statistically analyzing the pixel density changes in different directions. The horizontal frequency distribution map counts the number of pixels in the horizontal direction to form a horizontal projection curve, which is mainly used to determine the upper and lower boundaries of the characters. The vertical frequency distribution map counts the number of pixels in the vertical direction to form a vertical projection curve, which is mainly used to determine the left and right boundaries of each character.

[0094] Combining the horizontal and vertical frequency distribution maps can segment the ID card characters. First, determine the upper and lower boundaries of the single-line character region. Find the obvious peak regions in the horizontal projection curve, which are the starting row top and the ending row bottom of the ID card characters. The upper and lower boundaries can be determined by setting a threshold. Then, determine the left and right boundaries of the single-line characters. Find the peaks in the vertical projection curve at (x, y + 0.5*gap) to determine the starting column x and the ending column y of the character. Since there may be a gap between adjacent characters, a gap of half the character width is set to optimize the character boundaries. Finally, the character boundaries are: {x, y + 0.5×gap}.

[0095] Use bilinear interpolation to unify the size of the image to ensure consistent benchmarks for subsequent positioning; according to the absolute position coordinate information of the characters on the ID card, crop the positioning area containing the characters from the corrected image; use grayscale conversion and Otsu's method for binarization to process the image into an image with only two gray levels, which is convenient for subsequent statistical analysis; calculate the sum of pixel values and the number of white pixels for each row and each column in the image respectively to generate horizontal and vertical frequency distribution maps for judging the boundaries of the character regions; automatically segment each character in the single-line text according to the changes of peaks and valleys in the frequency distribution map. Through the above methods, the character regions on the ID card can be accurately extracted, and the characters can be segmented one by one to ensure the integrity of the character images; use statistical projection technology to improve the objectivity of character segmentation and reduce the risk of mis-segmentation caused by background interference or noise. Combine image preprocessing with projection statistics to realize the automatic positioning and segmentation of character regions in complex images; use grayscale statistics and frequency distribution analysis methods to make the character segmentation process highly adaptable and robust.

[0096] In some embodiments, identifying ID card characters to obtain ID card information includes:

[0097] Use the PIL library in Python to generate ID card character images from the ID card characters.

[0098] The identity card character image is processed by at least one of angle rotation, erosion, dilation, and noise addition to obtain the enhanced identity card character image through data augmentation. Erosion is used to remove small noise points and enhance the character contour; dilation is used to thicken the characters to make them more recognizable. Among them, the angle of angle rotation is one of 5° or -5°, and the dilation and erosion kernels for dilation and erosion are both 3*3, and the number of iterations is 1 time.

[0099] The enhanced identity card character image is trained and recognized using the ResNet residual neural network to obtain identity card information.

[0100] Data augmentation enables the character recognition model to maintain a high recognition rate in the face of image noise, rotation deviation, or morphological changes; using a deep residual network to classify characters can improve recognition accuracy while ensuring a moderate model complexity, ensuring the reliability of identity card information extraction. Introducing data augmentation processing means and specifying specific operation parameters ensures that the character image is still recognizable under slight deformation; using the ResNet structure for character recognition and leveraging residual connections to optimize the model training process improves the robustness and accuracy of character recognition in practical applications.

[0101] Another embodiment of the present invention also provides a positioning and correction device for obtaining identity card information held in hand, including:

[0102] An acquisition module for acquiring a first identity card image obtained based on a target recognition model;

[0103] A detection module for detecting the first identity card image according to the Hough line detection method to obtain the lines in the first identity card image;

[0104] A first processing module for processing the lines to obtain a first parallel line group and a second parallel line group for enclosing the largest area in the first identity card image. There are two parallel lines in each of the first parallel line group and the second parallel line group, and the parallel lines in the first parallel line group are perpendicular to the parallel lines in the second parallel line group;

[0105] A second processing module for processing the first parallel line group and the second parallel line group to obtain a second identity card image;

[0106] A third processing module for processing the second identity card image to obtain identity card characters;

[0107] An identification module for identifying the identity card characters to obtain identity card information.

[0108] Another embodiment of the present invention further provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the computer program is executed by the processor, it implements a positioning correction method for obtaining handheld ID card information.

[0109] The so-called processor may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), off-the-shelf programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The processor is the control center of the data processing device of the gateway, and connects various parts of the data processing device of the entire gateway through various interfaces and lines.

[0110] Another embodiment of the present invention further provides a computer-readable storage medium storing a computer program, which when executed by a processor implements a positioning correction method for obtaining handheld ID card information.

[0111] The memory may include a program storage area and a data storage area. Among them, the program storage area may store an operating system and application programs required for at least one function; the data storage area may store data created by the processor, etc. In addition, the memory is preferably but not limited to a high-speed random access memory. For example, it may also be a non-transitory memory, such as at least one disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory may also optionally include a memory remotely set relative to the processor, and these remote memories may be connected to the processor through a network. Examples of the above networks include but are not limited to the Internet, enterprise intranets, local area networks, mobile communication networks, and combinations thereof.

[0112] Those skilled in the art can understand that to implement all or part of the processes in the above-mentioned embodiment methods, a program can be used to instruct relevant hardware to complete the program, which can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the above-mentioned method embodiments. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), a random access memory (RAM), a flash memory, a hard disk drive (abbreviation: HDD), or a solid-state drive (SSD), etc.; the storage medium can also include a combination of the above-mentioned types of memories.

[0113] In summary, a positioning and correction method, device, electronic device, and storage medium for obtaining handheld ID card information provided by the present invention organically combine object detection, Hough line detection, constrained K-Means clustering, perspective transformation, character segmentation, and ResNet character recognition technologies to achieve precise positioning, correction, and information extraction of the ID card area in a handheld ID card image. Aiming at the problems of complex background and variable shooting angles of ID card images, the Yolov8 object detection algorithm is used for positioning, and combined with the Hough transform and the constrained K-Means clustering algorithm, the four sides of the ID card needed are obtained from the cluttered lines. Finally, perspective transformation is used to correct the tilt to achieve high-precision cropping and robust positioning. At the same time, the character position relationship and prior information are used to estimate the boundary to improve the segmentation accuracy, and the ResNet network is introduced to optimize character recognition, significantly improving the recognition accuracy. The data augmentation technology and the multi-level model training scheme are adopted to improve the robustness and generalization ability of the model in the processing of the original image and the object detection link; in the line processing link, the line angle and distance tolerance parameters are optimized through grid search, making the clustering result more accurate, and then obtaining the line combination that encloses the largest area; in the process of character segmentation and recognition, the horizontal and vertical frequency distribution maps are used to divide the character areas, and accurate recognition of characters is achieved through data augmentation and the deep residual network. The overall method realizes automatic correction and information extraction of ID card images under the conditions of tilt, deformation, and background interference while maintaining the objectivity and accuracy of the processing flow, which is beneficial to reducing manual intervention and meeting the requirements of automation, accuracy, and high robustness for obtaining ID card information in practical applications.

[0114] Although the present invention has been disclosed above with preferred embodiments, it is not intended to limit the present invention. Those with ordinary knowledge in the technical field to which the present invention belongs can make various modifications and refinements without departing from the spirit and scope of the present invention. Therefore, the protection scope of the present invention shall be subject to what is defined by the claims.

Claims

1. A positioning and correction method for obtaining handheld ID card information, characterized in that: include: Acquire a first ID card image obtained based on the target recognition model; Detect the first ID card image according to the Hough line detection method to obtain straight lines in the first ID card image; Processing the straight lines to obtain a first parallel line group and a second parallel line group for enclosing a maximum area in the first ID card image, wherein the first parallel line group and the second parallel line group each have two parallel lines, and the parallel lines in the first parallel line group are perpendicular to the parallel lines in the second parallel line group; Processing the first parallel line group and the second parallel line group to obtain a second ID card image; Processing the second ID card image to obtain ID card characters; Identify the ID card characters and obtain the ID card information.

2. A positioning and correction method for obtaining handheld ID card information according to claim 1, characterized in that: The collecting of the first ID card image obtained based on the target recognition model includes: Get the original image with the ID card; The original image is annotated with the ID card image using the labelImg annotation tool, and the annotated original image is divided into a training set, a test set, and a validation set in a ratio of 8:1:1; Importing the training set into a target recognition model using a YOLOv8 network model for training, thereby obtaining a trained target recognition model; Importing the verification set into the trained target recognition model, recording the number of iterations and the average precision mean, and obtaining the verified target recognition model; Importing the test set into the verified target recognition model, terminating the test when the number of iterations reaches the maximum number of iterations or the average precision reaches the early stopping condition, and obtaining a target recognition model that passes the test; The ID card image to be detected is imported into the tested target recognition model to obtain the first ID card image.

3. A positioning and correction method for obtaining handheld ID card information according to claim 2, characterized in that: The original image is an original image after data enhancement, and the data enhancement method of the original image at least includes rotating, moving, and flipping the original image.

4. A positioning and correction method for obtaining handheld ID card information according to claim 1, characterized in that: The processing of the straight lines to obtain a first parallel line group and a second parallel line group for enclosing a maximum area in the first ID card image, wherein the first parallel line group and the second parallel line group each have two parallel lines, and the parallel lines in the first parallel line group are perpendicular to the parallel lines in the second parallel line group, comprises: Using K-Means clustering to search the straight line grid, obtain the straight line angle tolerance and straight line distance tolerance with the highest straight line classification accuracy, the highest silhouette coefficient, the smallest angle standard deviation and distance standard deviation, and the highest integrity of the corrected ID card area; The straight line angle tolerance and the straight line distance tolerance are processed by constrained K-Means clustering to obtain the distance D(x i ,x j ), the processing formula is as follows: Among them, x i A straight line representing the current judgment; x j Represents the cluster center of the current judgment; slope i Represents x i Angle distance i Represents x i distance; slope j Represents x j Angle distance j Represents x j distance; tol_slope represents the straight line angle tolerance; tol_dis represents the straight-line distance tolerance; Using constrained K-Means clustering, the angle difference of 5° and the distance D(x i ,x j ) classify straight lines with a difference of 20px; select horizontal straight lines and vertical straight lines from the classified straight lines; select four straight lines for enclosing the largest area in the first ID card image from the horizontal straight lines and the vertical straight lines; classify a group of straight lines parallel to each other among the four straight lines as the first parallel line group, and classify another group of straight lines parallel to each other as the second parallel line group.

5. A positioning and correction method for obtaining handheld ID card information according to claim 1, characterized in that: The processing of the first parallel line group and the second parallel line group to obtain a second ID card image includes: Extend the straight lines of the first parallel line group and the second parallel line group to obtain four intersection points and the coordinates of the four intersection points; Extend the top, bottom, left and right boundaries of the standard ID card to obtain four vertices and their coordinates; The perspective transformation method is adopted to match the four intersection points to the four vertices so as to correct the first ID card image into the second ID card image which has no tilt or deformation and completely fills the whole picture.

6. A positioning and correction method for obtaining handheld ID card information according to claim 1, characterized in that: The step of processing the second ID card image to obtain ID card characters includes: Using bilinear interpolation method to scale the second ID card image to a standard ID card size; Cut out a positioning area from the second ID card image according to the absolute position (x, y, w, h) of the ID card characters, where x, y, w, h represent the coordinates of the upper left vertex and the width and height of the positioning area on the standard ID card respectively; Cutting the positioning area to obtain a first image; grayscale the first image to obtain a second image; The second image is processed by using Otsu's binarization method so that the image of the positioning area has only two gray levels, 0 and 255, to obtain a third image with black background and white characters; Calculate the sum of the grayscale values ​​of all pixels in each row of the matrix in the vertical and horizontal directions respectively. row , the sum of the grayscale values ​​of all pixels in each column col , and calculate the number of pixels with a gray value of 255 in each row row , the number of pixels with a grayscale value of 255 in each column number col , the calculation formula in the horizontal direction is the same as that in the vertical direction. The calculation formula in the vertical direction is as follows: According to the sum of the gray values ​​of all pixels in each row row 、The number of pixels with a grayscale value of 255 in each row number row , obtain the horizontal frequency distribution diagram; according to the sum of the gray values ​​of all pixels in each column col 、The number of pixels with a grayscale value of 255 in each column number col , obtain a vertical frequency distribution diagram; according to the horizontal frequency distribution diagram and the vertical frequency distribution diagram, segment each character of the single line of text to obtain the ID card character.

7. A positioning and correction method for obtaining handheld ID card information according to claim 1, characterized in that: The step of identifying the ID card characters and obtaining the ID card information includes: The ID card characters are generated into an ID card character image using Python's PIL library; The ID card character image is subjected to at least one of angle rotation, corrosion, expansion, and noise addition to obtain a data enhanced ID card character image, wherein the angle of angle rotation is one of 5° or -5°, the expansion and corrosion kernels of the expansion and corrosion are both 3*3, and the number of iterations is 1; The ResNet residual neural network is used to train and recognize the ID card character images after data enhancement to obtain the ID card information.

8. A positioning and correction device for obtaining handheld ID card information, characterized in that: include: An acquisition module, used for acquiring a first ID card image obtained based on a target recognition model; A detection module, used to detect the first ID card image according to the Hough line detection method to obtain the straight lines in the first ID card image; A first processing module is used to process the straight lines to obtain a first parallel line group and a second parallel line group for enclosing a maximum area in the first ID card image, wherein the first parallel line group and the second parallel line group each have two parallel lines, and the parallel lines in the first parallel line group are perpendicular to the parallel lines in the second parallel line group; A second processing module, used for processing the first parallel line group and the second parallel line group to obtain a second ID card image; A third processing module, used for processing the second ID card image to obtain ID card characters; The recognition module recognizes the ID card characters and obtains the ID card information.

9. An electronic device, characterized in that: It includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the computer program is executed by the processor, a positioning and correction method for obtaining handheld ID card information as described in any one of claims 1 to 7 is implemented.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, a positioning and correction method for obtaining handheld ID card information as described in any one of claims 1 to 7 is implemented.