Social security card data identification method based on image processing

By preprocessing and segmenting social security card images, and combining character and image recognition technology, a two-factor verification mechanism is built, the accuracy and reliability of social security card data recognition in complex scenarios is solved, and efficient and accurate recognition effect is achieved.

CN120182992AActive Publication Date: 2025-06-20TIANJIN ZHONGCHAO PAPER IND CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510662380.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-22
Publication Date
2025-06-20
Estimated Expiration
2045-05-22

AI Technical Summary

Technical Problem

The prior art is difficult to quickly and accurately identify social security card data in complex scenarios, especially when problems such as uneven lighting, tilt deformation, printing defects exist, and there is a lack of an effective verification mechanism for the identification results.

Method used

By preprocessing the original social security card image and aligning it with the standard template, segmenting it into the smallest unit sub-region block, and combining similar area blocks based on the adjacency relationship, combining character recognition and image recognition technology, a two-factor verification mechanism is built to verify the recognition results.

Benefits of technology

The social security card data recognition accuracy is improved in complex scenarios, reducing processing errors caused by image diversity, and improving the accuracy and reliability of the recognition results through the two-factor verification mechanism.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120182992A_ABST
    Figure CN120182992A_ABST
Patent Text Reader

Abstract

The invention provides a social security card data identification method based on image processing, and relates to the technical field of social security card data processing, and the method comprises the steps: preprocessing an original social security card image, and aligning the original social security card image with a standard template to obtain a structured image; segmenting the structured image according to a minimum unit to obtain a plurality of sub-region blocks, and constructing adjacent sub-region blocks; constructing adjacent sub-region blocks, and combining a plurality of adjacent sub-region blocks belonging to the character sub-region block to obtain a character region block; combining a plurality of adjacent sub-region blocks belonging to the image sub-region blocks to obtain image region blocks; and carrying out character information identification on the character region block, carrying out image identification on the image region block, verifying a character information identification result by using an image identification result, and if verification succeeds, proving that social security card data identification succeeds.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention proposes a social security card data recognition method based on image processing, which relates to the technical field of social security card data processing. Background Art

[0002] With the wide application of social security cards in government affairs, finance and other scenarios, the efficiency of manual verification is low, and automated technologies are required to achieve fast and accurate data extraction and verification. With the upgrade of the third-generation social security card, higher requirements are put forward for recognition technologies, and self-service terminals need to complete identity verification within seconds.

[0003] However, due to the complex acquisition scenarios of social security cards, the original images often have the following problems: overexposure or shadow occlusion of character regions in strong light or backlight environments; geometric distortion of images caused by shooting angle deviation or scanning device errors, affecting the accuracy of character segmentation; noise such as spots and scratches introduced by aging scanning devices or shooting jitters, covering the character edges; the social security card layout contains multiple types of information, and traditional methods are difficult to effectively parse, and elements such as text, tables, barcodes, and photos are interlaced. For example, the name and ID number may be located in different regional blocks; there are differences in social security cards in different regions, resulting in the invalidation of general templates; the fonts, font sizes, and colors vary widely, and handwritten and printed fonts coexist; relying solely on character recognition is vulnerable to image quality, and traditional segmentation methods based on thresholds or edge detection are likely to split continuous text into multiple sub-regions, or misclassify background noise as characters; the logical relationships between fields such as "gender" and "nationality" cannot be distinguished, resulting in information misalignment. Most importantly, traditional solutions lack a verification mechanism for recognition results, and the character and image information are disjoint. For example, when the photo does not match the name, it cannot be automatically detected and requires manual secondary verification, and the anti-counterfeiting ability is low. Summary of the Invention

[0004] In order to solve the above technical problems, the present invention proposes a social security card data recognition method based on image processing, including: Preprocess the original social security card image and align it with the standard template to obtain a structured image; Segment the structured image according to the smallest unit to obtain multiple sub-region blocks, and construct adjacent sub-region blocks; Construct adjacent sub-region blocks, merge multiple adjacent sub-region blocks belonging to character sub-region blocks to obtain a character region block; merge multiple adjacent sub-region blocks belonging to image sub-region blocks to obtain an image region block; Perform character information recognition on the character region block, perform image recognition on the image region block, and use the image recognition result to verify the character information recognition result. If the verification is successful, it proves that the social security card data recognition is successful.

[0005] In a preferred embodiment, for character information recognition of the character region block, an equidistant sampling method is adopted to select key points on the character region block; the stroke width from each key point to the character edge is calculated, and a set of stroke widths {Z1, Z2, …, Z k …, Z m} of m key points is constructed, the average value and median of the set of stroke widths are calculated, and the average value and median are used as the stroke features of the character, and the character information recognition result is output.

[0006] In a preferred embodiment, the average value W avg of the set of stroke widths is: ; The set of stroke widths {Z1, Z2, …, Z k …, Z m} is sorted in ascending order to obtain {Z(1), Z(2), …, Z(m’)}, and the median W med is calculated as: If m’ is odd, the median ; If m’ is even, the median .

[0007] In a preferred embodiment, for image recognition of the image region block, the recognition result of the face image region block is extracted, a face image output module and a judgment module are constructed, the face image output module and the judgment module are alternately trained, the face image output module outputs a real face image based on the recognition result of the face image region block, and the judgment module judges the authenticity of the real face image output by the face image output module and judges the gender information at the same time, and the gender information judgment result is used to verify the gender digit of the social security number in the character information recognition result.

[0008] In a preferred embodiment, a difference function L G is constructed in the face image output module:

[0009]

[0010]

[0011] where L fake is the difference function for the face image output module to guide the judgment module to judge the image authenticity, is the difference function for the judgment module to perform gender classification judgment according to the generated noisy face image G(z), is the desired generated gender label, It is the probability that the discrimination module correctly classifies the generated noisy face image G(z) as the desired generated gender label where z is the noise data sampled from the noise distribution pb(z), and E is the mathematical expectation symbol representing the probability that the generated noisy face image G(z) is judged as true by the discrimination module

[0012] In a preferred embodiment, the difference function L set in the discrimination module D is expressed as: L D =L real +L gender ; where: ; ; where L real is the difference function for the discrimination module to judge the authenticity of the real face image Q generated by the face image output module, and L gender is the difference function for the discrimination module to perform gender classification judgment based on the real face image Q, h is the gender label corresponding to the real face image is the output of the discrimination module for the probability that the real face image Q is the gender label h; pa(Q) is the real data distribution of the real face image Q, and E is the mathematical expectation symbol, D real (Q) represents the probability that the generated real face image Q is true

[0013] In a preferred embodiment, for the image region block, image recognition is performed, the recognition result of the bank logo sub-region block is extracted, and the bank name is output; the recognition result of the bank card number sub-region block of the character region block is obtained, and the issuing bank name is obtained by querying the mapping database through the IIN. If the bank name matches the issuing bank name, the verification is successful

[0014] In a preferred embodiment, an adjacent sub-region block is constructed through the adjacent weight. The adjacent weight W i,j between sub-region blocks i and j is calculated as follows: ;

[0015] where the central distance D(i,j) represents the distance between the geometric centers of sub-region blocks i and j is a hyperparameter

[0016] In a preferred embodiment, the adjacent sub-region block verification method is as follows: When the distance between the right boundary of sub-region block i and the left boundary of sub-region block j ≤ M1 pixels, it is judged that the sub-region blocks are adjacent in the horizontal direction When the distance between the lower boundary of sub-region block i and the upper boundary of sub-region block j is ≤ M2 pixels, it is determined that the sub-region blocks are vertically adjacent.

[0017] In a preferred embodiment, a detection algorithm is used to extract the four vertices of the quadrilateral of the social security card image, fix the positions of three of the vertices, and dynamically adjust the coordinates of the remaining vertex, so as to perspectively transform the social security card image into a structured image aligned with the standard template.

[0018] Compared with the prior art, the present invention has the following beneficial technical effects: 1. By preprocessing, interferences such as uneven illumination, skew distortion, and printing defects in the original image are eliminated, the character edges are made clear, and the features of the image area are stabilized. The social security card image is transformed into a structured image, reducing the processing error caused by image diversity.

[0019] 2. The image is segmented into the smallest unit sub-region blocks, and regions (character regions or image regions) with the same attributes are merged based on the adjacency relationship, which conforms to the physical structure of the social security card design (such as characters arranged in rows / columns and image logos as independent blocks). The incomplete regions are repaired by adjacency merging, improving the regional integrity.

[0020] 3. For character regions (letters, numbers), character recognition technology is adopted, and its strong parsing ability for serialized symbols is utilized to quickly extract key information such as name, card number, and expiration date; for image regions (logos, patterns), image detection methods are adopted to identify graphic information such as the card-issuing institution logo and face image; image recognition is used to verify character information, and a dual-verification mechanism is constructed to form a "character-image" dual-source evidence chain. Description of the Drawings

[0021] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0022] Figure 1 It is a flowchart of the method for identifying social security card data based on image processing according to the present invention; Figure 2 It is a schematic diagram of multiple sub-region blocks obtained by segmentation according to the present invention; Figure 3 It is a flowchart of character information recognition for character region blocks according to the present invention; Figure 4 It is a bar chart of comparison of recognition accuracy according to the present invention. Detailed Embodiments

[0023] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the following will clearly and completely describe the technical solutions in the embodiments of this application with reference to the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are part of the embodiments of this application, rather than all of them. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of this application.

[0024] Embodiment 1

[0025] As Figure 1 shown, it is a flowchart of a social security card data recognition method based on image processing according to the present invention. The social security card data recognition method based on image processing includes: To improve the efficiency and accuracy of subsequent data recognition, it is first necessary to preprocess the original social security card image to obtain a preprocessed social security card image.

[0026] Specifically, preprocessing the original social security card image includes converting a color image into a grayscale image to reduce the amount of data; using Gaussian filtering to remove noise in the image to make the image smoother; and enhancing the contrast of the image through histogram equalization to highlight text and key information.

[0027] Preferably, a thin plate spline interpolation algorithm is implemented using a grid sampler to sample the original social security card image and process local distortions in the original social security card image. Specifically, pixel values are extracted according to the positions of the sampling grid in the original social security card image, and these pixel values are interpolated using the thin plate spline interpolation algorithm to generate an interpolated image. During the interpolation process, the thin plate spline interpolation algorithm takes into account local deformations of the image, making the interpolated image more natural and accurate.

[0028] Next, align the preprocessed social security card image with a standard template to obtain a structured image.

[0029] Locate the four vertices of the quadrilateral of the preprocessed social security card image: fix the positions of three of the vertices, dynamically adjust the coordinates of the remaining vertex, and perform perspective transformation on the quadrilateral to a preset standard template to eliminate local area deformation and obtain a structured image aligned with the standard template.

[0030] Specifically, combine the contour approximation algorithm to approximate the contour of the social security card image as a quadrilateral, use a detection algorithm to extract the four vertices of the quadrilateral of the social security card image, arrange them in ascending order of the sum of the horizontal and vertical coordinates, determine the upper left, upper right, lower left, and lower right vertices, and obtain a set of four vertex coordinates P = {p1, p2, p3, p4}, where p1, p2, p3, p4 correspond to the upper left, upper right, lower left, and lower right vertices of the quadrilateral respectively.

[0031] Define a set of four vertex coordinates Q = {q1, q2, q3, q4} for the standard template, where: q1 = (x0, y0), where (x0, y0) is the coordinate of the upper-left vertex q1, and the upper-left vertex is fixed as a preset reference point; q2 = (x0 + w, y0), where (x0 + w, y0) is the coordinate of the upper-right vertex q2, and w is the width of the standard template; q3 = (x0, y0 + h), where (x0, y0 + h) is the coordinate of the lower-left vertex q3, and h is the height of the standard template; q4 = (x0 + w, y0 + h), where (x0 + w, y0 + h) is the coordinate of the lower-right vertex q4.

[0032] Match the four vertices of the quadrilateral of the detected social security card image with the four vertices of the standard template. First, fix the first three vertices of the quadrilateral of the social security card image , and only adjust the coordinate of the fourth vertex p4 to obtain the lower-right vertex of the quadrilateral of the adjusted social security card image .

[0033] Establish a mapping relationship between the vertices of the quadrilateral and the vertices of the standard template through the perspective transformation constraint conditions: ;

[0034] Among them, To satisfy the constraint of the transformation matrix H for perspective transformation, that is , H is the transformation matrix, and H is obtained by fitting through the mapping relationship of .

[0035] Through the above method, the local deformation of the social security card image can be effectively eliminated, and the precise alignment of the structured image can be achieved.

[0036] Next, divide the above structured image according to the smallest unit to obtain multiple sub-region blocks, and construct adjacent sub-region blocks.

[0037] Divide the structured image according to the smallest unit to obtain multiple sub-region blocks. For example Figure 2 Multiple sub-region blocks are schematically marked. Normalize the upper-left coordinate (x i1 , y i1 ), width and height information (w i , h i ) of each sub-region block i into the interval [0, 1].

[0038] Construct adjacent sub-region blocks through the adjacency weight W i,j . The closer the adjacency weight W i,j is to 1, the closer the two sub-region blocks are.

[0039] The adjacency weight Wi,j The calculation formula is as follows: ;

[0040] where W i,j represents the adjacency weight between sub-region blocks i and j, which is used to measure the degree of association between two sub-region blocks in terms of position. The larger the adjacency weight, the closer the spatial association between the two.

[0041] The central distance D(i,j): represents the distance between the geometric centers of sub-region blocks i and j, reflecting the spatial position relationship between the two sub-region blocks. The closer the distance, the greater the positive impact on W i,j .

[0042] is a hyperparameter. Preferably, = 10 pixels.

[0043] According to the standard structure of the social security card, the adjacency relationship of the predefined core sub-region blocks (such as the "photo" sub-region block must be adjacent to the "name" sub-region block) is used to forcibly set W i,j = 1.

[0044] The method for verifying adjacent sub-region blocks is as follows: When the distance between the right boundary of sub-region block i and the left boundary of sub-region block j is ≤ M1 pixels, it is determined that the sub-region blocks are adjacent in the horizontal direction.

[0045] When the distance between the lower boundary of sub-region block i and the upper boundary of sub-region block j is ≤ M2 pixels, it is determined that the sub-region blocks are adjacent in the vertical direction.

[0046] Through the adjacency weight W i,j and the verification of adjacent sub-region blocks, the spatial position relationship of adjacent sub-region blocks is represented in the form of parameters, which is convenient for subsequent processing using algorithms. It realizes the merging of multiple adjacent sub-region blocks that all belong to character sub-region blocks to obtain a character region block; and the merging of multiple adjacent sub-region blocks that all belong to image sub-region blocks to obtain an image region block.

[0047] In a preferred embodiment, for non-standard formats (such as social security cards in different provinces), the adjacent threshold is automatically adjusted through reinforcement learning to adapt to layout changes.

[0048] Next, for each character region block, character information recognition is performed, and the recognition result is output. The recognition process is as Figure 3 shown, including the following steps: (1) Select key points on each character region block.

[0049] Using the method of equidistant sampling, a key point is selected at a certain pixel distance d on the character region block, that is, the pixel distance for selecting the key point is d. Let the set of key point coordinates be K = {(X k , Y k )}.

[0050] (2) Calculate the stroke width from the key point to the character edge.

[0051] For the k-th key point P(X k , Y k ), starting from this key point, search horizontally to the right to find the first background pixel point (the pixel value of the background pixel point is 0), and record the stroke width Z k from this key point to the character edge.

[0052] The stroke width Z k from each key point to the character edge reflects the local stroke width near this key point. By reasonably setting the interval d, these key points can represent the overall situation of the character region to a certain extent.

[0053] Construct the set of stroke widths {Z1, Z2, …, Z k …, Z m} of m key points, and calculate the average value W avg and the median W med of the set of stroke widths, which can characterize the overall width feature of the stroke from a statistical perspective. Even if the points are selected at intervals, as long as the interval d is reasonably valued (not too large or too small), it can not only avoid excessive calculation amount but also retain sufficient information to reflect the stroke feature, thus realizing the effective extraction of the character stroke feature.

[0054] (3) Calculate the stroke feature of the character.

[0055] Use the average value and the median as the stroke feature of the character.

[0056] The formula for calculating the average value W avg of the set of stroke widths is: ; Sort the set of stroke widths {Z1, Z2, …, Z k …, Z m} in ascending order to get {Z(1), Z(2), …, Z(m')}, and calculate the median W med as follows: If m' is odd, then the median ; If m' is even, then the median ; Through the above steps, each character can be converted into stroke features including the average value and the median value.

[0057] Since the stroke features (W avg 、W med ) of different numbers or characters are different, by using these differences to train a classification model (such as a support vector machine, a neural network, etc.), the recognition of numbers and characters can be achieved.

[0058] For example, the stroke structures of "1" and "7" are different, and their stroke features will be significantly distinguishable, and the classification model can accurately classify them accordingly.

[0059] (4) Output the recognition results of each character region block.

[0060] Output the recognition text results of fields such as name and card number, and fill the recognition text into the structured database according to the spatial position of adjacent sub-region blocks, such as represented in the form of key-value pairs: {"adjacent weight": "1"; "name": "Zhang San"; "card number": "XXXX"}.

[0061] (5) Verify the data logic of the recognition results and trigger an exception flag.

[0062] Examples of data logic include whether the social security card number, bank card number, and social security number are of the specified number of digits.

[0063] Use the formulated rules to verify the data filled into the structured database. Check whether the data of each field conforms to the corresponding rules one by one. If it is found that the data of a certain field does not conform to the rules, trigger an exception flag.

[0064] For the data marked as abnormal, take corresponding handling measures. The abnormal information can be recorded, including the abnormal field, the reason for the abnormality, and the specific recognition text content. At the same time, feedback the abnormal data to the manual review process for further verification and correction by the manual.

[0065] After verification and exception handling, output the final structured social security card data. These data are presented in a clear and standardized format and can be directly used for subsequent data analysis, storage, or other business processes.

[0066] In a preferred embodiment, for the image region block, perform image recognition, extract the recognition results of the face image region block, and perform gender information judgment, and verify the gender digit of the social security number in the recognition results of the character region block.

[0067] Gender character rule of social security number (i.e., ID number): The 17th digit of the social security number represents gender, with odd numbers for males and even numbers for females. Extract the 17th digit of the recognized social security number and judge the parity and gender.

[0068] Specifically, for the image region blocks, image recognition is performed, the recognition results of the face image region blocks are extracted, and a face image output module and a judgment module are constructed.

[0069] The face image output module and the judgment module are alternately trained. The face image output module outputs a real face image based on the recognition results of the face image region blocks. The judgment module judges the authenticity of the real face image output by the face image output module, and at the same time performs gender information judgment, and uses the judgment result of the gender information to verify the gender digit of the social security number in the character information recognition result.

[0070] In order to enable the judgment module to simultaneously judge the real probability and gender probability of the generated face image, the judgment module is set to output two results: one is the judgment of the authenticity of the face image, and the other is the judgment of the gender.

[0071] Construct a difference function L in the face image output module G : ; ; ;

[0072] where L fake is the difference function for the face image output module to guide the judgment module to judge the image authenticity, is the difference function for the judgment module to perform gender classification judgment based on the generated noisy face image G(z), is the desired generated gender label, is the probability that the judgment module correctly classifies the generated noisy face image G(z) as the desired generated gender label z is the noise data sampled from the noise distribution pb(z), and E is the mathematical expectation symbol, represents the probability that the generated noisy face image G(z) is judged as true by the judgment module.

[0073] After the above training, the generated noisy face image G(z) by the face image output module is finally trained to generate a real face image Q.

[0074] Construct a difference function L in the judgment module D : L D =L real +L gender ; where: ; ; Among them, L real is the difference function for the judgment module to judge the authenticity of the generated real face image Q, L gender is the difference function for the judgment module to perform gender classification judgment based on the real face image Q, h is the gender label corresponding to the real face image, is the output of the probability that the real face image Q is the gender label h by the judgment module; pa(Q) is the true data distribution of the real face image Q, E is the mathematical expectation symbol, D real (Q) represents the probability that the generated real face image Q is true.

[0075] By maximizing the probability that the real face image is correctly recognized as true and minimizing the probability that the generated real face image is misrecognized as true, the judgment module is trained to better distinguish between real and generated real face images; at the same time, by maximizing the probability of correct gender classification, the judgment module is trained to perform well in the gender classification task.

[0076] Finally, by alternately training the face image output module and the judgment module, the face image output module can generate realistic face images with specific gender characteristics, and at the same time, the judgment module can accurately judge the authenticity and gender of the images.

[0077] Embodiment 2

[0078] Perform image recognition on the image region block, extract the recognition result of the bank logo sub-region block, and verify the recognition result of the bank card number sub-region block of the character region block.

[0079] The first 6 digits of the bank card number are the issuer identification code (IIN). The issuer name can be obtained by querying the mapping database through the IIN (for example, "622202" corresponds to the Industrial and Commercial Bank of China); at the same time, extract the recognition result of the bank logo region block, and perform text matching on the output bank name (such as "Industrial and Commercial Bank") and the IIN query result (such as "Industrial and Commercial Bank of China"), supporting fuzzy matching; if the confidence level output by the matching algorithm is lower than the threshold, it is determined as "unable to match" or manual review is prompted; if the confidence level output by the matching algorithm is not lower than the threshold, the verification is successful.

[0080] Table 1 Comparison of recognition accuracies

[0081] Through the three core technologies of adjacent merging character segmentation, mutual verification of image and character information, and structured processing, the recognition accuracy of the present invention in complex scenarios is significantly better than that of the prior art. As Figure 4 shown, in the scenarios of character adhesion, image blur, and layout change, the accuracy improvement range reaches 10%-20%.

[0082] In one embodiment, a computer device is further provided, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, the steps in the above method embodiments are implemented.

[0083] In one embodiment, a computer-readable storage medium is provided, storing a computer program, and when the computer program is executed by a processor, the steps in the above method embodiments are implemented.

[0084] In one embodiment, a computer program product or a computer program is provided. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device implements the steps in the above method embodiments.

[0085] Those skilled in the art can understand that all or part of the processes of implementing the above method embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the above method embodiments. Among them, any reference to a memory, a database, or other media used in the embodiments provided in the present application can include at least one of non-volatile and volatile memories. The databases involved in the embodiments provided in the present application can include at least one of relational databases and non-relational databases. The non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in the present application can be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, data processing logics based on quantum computing, etc., without limitation.

[0086] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.

[0087] The above-described embodiments only represent several implementation manners of the present application, and their descriptions are relatively specific and detailed, but they should not be construed as limiting the patent scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the appended claims.

Claims

1. A method for identifying social security card data based on image processing, characterized in that, Including: Preprocess the original social security card image and align it with the standard template to obtain a structured image; Segment the structured image according to the smallest unit to obtain multiple sub-region blocks, and construct adjacent sub-region blocks; Construct adjacent sub-region blocks, merge multiple adjacent sub-region blocks belonging to character sub-region blocks to obtain a character region block; merge multiple adjacent sub-region blocks belonging to image sub-region blocks to obtain an image region block; Perform character information recognition on the character region block, perform image recognition on the image region block, and use the image recognition result to verify the character information recognition result. If the verification is successful, it proves that the social security card data recognition is successful.

2. The method for identifying social security card data based on image processing according to claim 1, characterized in that, Perform character information recognition on the character region block, and use the method of equidistant sampling to select key points on the character region block; Calculate the stroke width from each key point to the edge of the character, and construct a set of stroke widths {Z1, Z2, …, Z k …, Z m} for the m key points. Calculate the average value and median of the set of stroke widths, use the average value and median as the stroke features of the character, and output the recognition result of the character information.

3. The method for identifying social security card data based on image processing according to claim 2, characterized in that, The average value W of the set of stroke widths avg is as follows: ; Sort the set of stroke widths {Z1, Z2, …, Z k …, Z m} in ascending order to obtain {Z(1), Z(2), …, Z(m’)} and calculate the median W med : If m' is odd, then the median ; If m' is even, then the median .

4. The method for identifying social security card data based on image processing according to claim 1, characterized in that, Perform image recognition on the image region block, extract the recognition result of the face image region block, construct a face image output module and a judgment module, and alternately train the face image output module and the judgment module. The face image output module outputs a real face image based on the recognition result of the face image region block. The judgment module judges the authenticity of the real face image output by the face image output module and judges the gender information at the same time, and uses the gender information judgment result to verify the gender digit of the social security number in the character information recognition result.

5. The method for identifying social security card data based on image processing according to claim 4, characterized in that, Construct the difference function L in the face image output module G : ; ; ; Among them, L fake is the difference function for the face image output module's guidance judgment module to judge the authenticity of the image, is the difference function for the judgment module to perform gender classification judgment based on the generated noisy face image G(z), is the desired generated gender label, is the probability that the judgment module correctly classifies the generated noisy face image G(z) as the desired generated gender label where z is the noise data sampled from the noise distribution pb(z), and E is the mathematical expectation symbol, represents the probability that the generated noisy face image G(z) is judged to be true by the judgment module.

6. The method for identifying social security card data based on image processing according to claim 5, characterized in that, The difference function L set in the judgment module D is expressed as: L D = L real + L gender ; Among them: ; ; Among them, L real is a difference function for the judgment module to judge the authenticity of the real face image Q generated by the face image output module. L gender is a difference function for the judgment module to perform gender classification judgment based on the real face image Q. h is the gender label corresponding to the real face image is the output of the probability that the real face image Q is the gender label h by the judgment module; pa(Q) is the true data distribution of the real face image Q, and E is the mathematical expectation symbol. D real (Q) represents the probability that the generated real face image Q is true.

7. The method for identifying social security card data based on image processing according to claim 1, characterized in that, Perform image recognition on the image region block, extract the recognition result of the bank logo sub-region block, and output the bank name; obtain the recognition result of the bank card number sub-region block of the character region block, query the mapping database through IIN, and obtain the issuing bank name. If the bank name matches the issuing bank name, the verification is successful.

8. The method for identifying social security card data based on image processing according to claim 1, characterized in that, Construct adjacent sub-region blocks through adjacent weights. The adjacent weight W between sub-region blocks i and j i,j The calculation formula is: ; Among them, the center distance D(i,j) represents the distance between the geometric centers of sub-region blocks i and j, which is a hyperparameter.

9. The method for identifying social security card data based on image processing according to claim 8, wherein, The method for verifying adjacent sub-region blocks is as follows: When the distance between the right boundary of sub-region block i and the left boundary of sub-region block j is ≤ M1 pixels, it is judged that the sub-region block is horizontally adjacent; When the distance between the lower boundary of sub-region block i and the upper boundary of sub-region block j is ≤ M2 pixels, it is judged that the sub-region block is vertically adjacent.

10. The method for identifying social security card data based on image processing according to claim 1, wherein, Use a detection algorithm to extract the four vertices of the quadrilateral of the social security card image, fix the positions of three of the vertices, dynamically adjust the coordinates of the remaining vertex, and perspective-transform the social security card image to align it with the structured image of the standard template.

Citation Information

Patent Citations

  • Identity authentication method based on identity certificate information and human face multi-feature recognition

    CN104680131A

  • Document image character identification method and device

    CN107622263A

  • Certificate identification method and device, computing device and storage medium

    CN111639648A

  • Card type file image recognition method and device

    CN113887484A

  • Identity recognition method, electronic device, and computer readable storage medium

    WO2019062080A1