Certificate character recognition method and device, electronic equipment and storage medium
By acquiring the bounding box data and face data of the ID image, the correction parameters are determined to correct the ID image, which solves the problems of low recognition efficiency and low accuracy in the existing ID word recognition methods, and achieves more efficient and accurate ID word recognition.
Patent Information
- Application Number
- CN202311689693.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-08
- Publication Date
- 2025-06-10
AI Technical Summary
The existing document text recognition methods have problems with low recognition efficiency and low text recognition accuracy, especially when there is distortion and rotation of the document image.
By acquiring bounding box data and face data of the document image to be corrected, the correction parameters are determined to correct the document image, and the recognition efficiency and accuracy are improved.
After the correction processing, the text recognition of the ID image is significantly improved.
Smart Images

Figure CN120126167A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image recognition, and in particular, to a method, apparatus, electronic device, and storage medium for recognizing document characters. Background Art
[0002] With the development of image recognition technology, people's demand for document recognition and automated processing is increasing. However, in practical applications, the document images obtained from the images to be recognized often have various problems, such as image distortion, rotation, etc. These problems will affect the recognition effect of document characters. Therefore, it is very necessary to perform correction processing on the document images. In the traditional document character recognition technology, the OCR recognition technology is used to recognize the characters on the document image. However, due to the certain degree of inclination of the provided document image and the possible presence of multiple regional images in one image, the detection and recognition rate of the OCR for the characters on the document image is extremely low, and it may also recognize other images that are not documents. Therefore, the existing document character recognition methods have the problems of low recognition efficiency and low character recognition accuracy. Summary of the Invention
[0003] An embodiment of the present invention provides a method for recognizing document characters, aiming to solve the problems of low recognition efficiency and low character recognition accuracy of the existing document character recognition methods. According to the bounding box data and face data in the document image to be corrected obtained, the correction parameters of the document image to be corrected are obtained, and the document image to be corrected is corrected through the correction parameters, and then the corrected document image is subjected to character recognition processing, thereby improving the efficiency and accuracy of document character recognition.
[0004] In a first aspect, an embodiment of the present invention provides a method for recognizing document characters, and the method includes the following steps:
[0005] Obtain the bounding box data of the document image to be corrected and the face data of the document image to be corrected from the image to be recognized;
[0006] Based on the bounding box data of the document image to be corrected and the face data of the document image to be corrected, determine the correction parameters of the document image to be corrected;
[0007] Perform correction processing on the document image to be corrected through the correction parameters to obtain the document image to be recognized;
[0008] Perform character recognition on the document image to be recognized to obtain the document character recognition result of the image to be recognized.
[0009] Optionally, before obtaining the bounding box data of the document image to be corrected and the face data of the document image to be corrected from the image to be recognized, the method further includes:
[0010] Perform edge feature recognition processing on the image to be recognized through a preset edge detection algorithm to obtain the edge information of the target object;
[0011] Based on the edge information of the target object, determine the bounding box of the target object;
[0012] If the ratio of the bounding box of the target object conforms to the ratio of the document bounding box, detect whether there is a face within the bounding box;
[0013] If there is a face within the bounding box, extract the document image from the image to be recognized based on the bounding box;
[0014] Extract features from the document image to obtain the features of the document image, and classify or recognize based on the features of the document image to obtain the document image to be corrected.
[0015] Optionally, the extracting features from the document image to obtain the features of the document image includes:
[0016] Extract features from the document image to obtain the bounding box features of the document image and the portrait features in the document image;
[0017] The classifying or recognizing based on the features of the document image to obtain the document image to be corrected includes:
[0018] Classify or recognize based on the bounding box features and the portrait features of the document image, and judge whether there is a document image that meets the expectations in the image to be processed based on the results of the classification or recognition;
[0019] If there is a document image that meets the expectations, extract the document image that meets the expectations as the document image to be corrected.
[0020] Optionally, the obtaining the bounding box data of the document image to be corrected and the face data of the document image to be corrected from the image to be recognized includes:
[0021] Determine the bounding box data of the document image to be corrected according to the edge information corresponding to the document image to be corrected;
[0022] Perform face key point detection on the face in the document image to be corrected to obtain the face data of the document image to be corrected.
[0023] Optionally, the determining the correction parameters of the document image to be corrected based on the bounding box data of the document image to be corrected and the face data of the document image to be corrected includes:
[0024] Judge whether there is face key point information in the face data;
[0025] If there is no facial key point information in the facial data, the correction parameters of the to-be-corrected certificate image are determined through the bounding box data of the to-be-corrected certificate image;
[0026] If there is facial key point information in the facial data, the correction parameters of the to-be-corrected certificate image are determined through the facial key point information of the to-be-corrected certificate image.
[0027] Optionally, if there is facial key point information in the facial data, the correction parameters of the to-be-corrected certificate image are determined through the facial key point information of the to-be-corrected certificate image, including:
[0028] If there is facial key point information in the facial data, the eye key point information is extracted from the facial key point information;
[0029] Based on the eye key point information, the correction parameters of the to-be-corrected certificate image are determined.
[0030] Optionally, the eye key point information includes the coordinates of the left eye key point and the coordinates of the right eye key point. Based on the eye key point information, determining the correction parameters of the to-be-corrected certificate image includes:
[0031] Based on the coordinates of the left eye key point and the coordinates of the right eye key point, the slope between the left eye key point and the right eye key point is determined;
[0032] Based on the slope between the left eye key point and the right eye key point, the tilt angle of the face portrait in the to-be-corrected certificate image is determined;
[0033] Based on the tilt angle of the face portrait in the to-be-corrected certificate image, the correction parameters of the to-be-corrected certificate image are determined.
[0034] In a second aspect, an embodiment of the present invention further provides a certificate text recognition device, where the certificate text recognition device includes:
[0035] A first acquisition module, configured to acquire the bounding box data of the to-be-corrected certificate image and the facial data of the to-be-corrected certificate image from the to-be-recognized image;
[0036] A first determination module, configured to determine the correction parameters of the to-be-corrected certificate image based on the bounding box data of the to-be-corrected certificate image and the facial data of the to-be-corrected certificate image;
[0037] A first processing module, configured to perform correction processing on the to-be-corrected certificate image through the correction parameters to obtain a to-be-recognized certificate image;
[0038] An identification module for performing character recognition on the image of the document to be identified, and obtaining the character recognition result of the image of the document to be identified.
[0039] In a third aspect, an embodiment of the present invention provides an electronic device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor, where when the processor executes the computer program, the steps in the document character recognition method provided by the embodiment of the present invention are implemented.
[0040] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps in the document character recognition method provided by the embodiment of the invention are implemented.
[0041] In the embodiment of the present invention, the bounding box data of the document image to be corrected and the face data of the document image to be corrected are obtained from the image to be identified. Based on the bounding box data of the document image to be corrected and the face data of the document image to be corrected, the correction parameters of the document image to be corrected are determined. The document image to be corrected is corrected through the correction parameters to obtain the document image to be identified, and character recognition is performed on the document image to be identified to obtain the character recognition result of the image to be identified. According to the obtained bounding box data and face data in the document image to be corrected, the correction parameters of the document image to be corrected are obtained, the document image to be corrected is corrected through the correction parameters, and then the corrected document image is subjected to character recognition processing, thereby improving the efficiency and accuracy of document character recognition. Description of the Drawings
[0042] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention, and those of ordinary skill in the art can obtain other drawings without creative efforts based on these drawings.
[0043] Figure 1 is a flowchart of a document character recognition method provided by an embodiment of the present invention;
[0044] Figure 2 is a distribution diagram of face key point information provided by an embodiment of the present invention;
[0045] Figure 3 is a flowchart of another document character recognition method provided by an embodiment of the present invention;
[0046] Figure 4 is a schematic structural diagram of a document character recognition device provided in an embodiment of the present invention;
[0047] Figure 5 It is a schematic structural diagram of an electronic device provided by an embodiment of the present invention. Specific embodiments
[0048] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0049] As Figure 1 shown, Figure 1 It is a flowchart of a document character recognition method provided by an embodiment of the present invention. The document character recognition method includes the steps of:
[0050] 101. Obtain the bounding box data of the document image to be corrected and the face data of the document image to be corrected from the image to be recognized.
[0051] In the embodiment of the present invention, the above document character recognition method can be applied to any image recognition platform that needs to extract document information. The above image recognition platform has functions such as data processing, data sending and receiving, and data storage, and can be built based on a server or a server cluster. The above server or server cluster can be an electronic device with functions such as image processing, image acquisition, and image recognition.
[0052] The above image to be recognized can be an image obtained by the above image recognition platform. Specifically, the above image to be recognized can contain more than one document image to be corrected. When the above image recognition platform obtains the image to be recognized, due to environmental factors or human factors during acquisition, the obtained image to be recognized is inclined and blurred to a certain extent. Therefore, it is also necessary to preprocess the above image to be recognized to obtain a clear image to be recognized, so as to improve the accuracy of subsequent processing. The above preprocessing can be operations such as resizing, noise reduction, and contrast enhancement of the above image to be recognized, making its image contour clearer.
[0053] The above image to be recognized can be an originally collected image. Specifically, it can be an image containing a document taken by a user through an image capture device. For example, during a population census, an image containing an ID card or a residence permit taken by a staff member through a mobile phone or a camera.
[0054] The above-mentioned document image to be corrected can be a document image that meets the expectations extracted from the above-mentioned image to be recognized. Since the above-mentioned image to be recognized may be tilted or not centered during shooting, the attached document image in the figure also becomes tilted along with the overall image, which will thus hinder subsequent image text recognition. Specifically, the above-mentioned image to be recognized not only has the document image to be corrected, but may also have non-expected document images and non-document images similar to the above-mentioned document image to be corrected. For example, during a census, when taking pictures and uploading the ID photos of residents, bank cards or other cards that do not conform to the document type in the same frame as the ID photo may also be taken and uploaded together. At this time, edge detection can be used to detect whether there is an edge box in the image to be recognized that is similar to the proportion of the document bounding box. If it exists, it means that the edge box may be the document bounding box, and face detection can be further performed on the edge box to determine whether there is a face in the edge box. Specifically, a face detector can be used to perform face detection on the image within the edge box that conforms to the proportion of the document border box. If no face is detected, the image to be recognized is rotated by 90° and detected again until a face is detected. After the last rotation, face detection does not need to be performed again. If a face is detected, it is determined that the image within the edge box is the document image, and feature extraction is performed on the document image to obtain the bounding box features and face features of the document image. Based on the bounding box features and face features of the document image, the document image is recognized or classified to determine whether the document image is the expected document image. If the document image is the expected document image, the document image is determined as the document image to be corrected.
[0055] If the document image is the expected document image, the bounding box data of the document image to be corrected is determined according to the edge box data corresponding to the document image. The above-mentioned bounding box data can include the length of the border length and border width of the above-mentioned document image to be corrected and the bounding box ratio data. Specifically, the image can be processed such as removing the background and binarization to obtain the coordinates and size of the bounding box.
[0056] The above-mentioned face data can include but is not limited to face key point data. The above-mentioned face key point data can refer to the coordinate positions or feature points of special parts on the human face, which are used for face recognition.
[0057] 102. Determine the correction parameters of the document image to be corrected based on the bounding box data of the document image to be corrected and the face data of the document image to be corrected.
[0058] In an embodiment of the present invention, the above correction parameters may be parameter values obtained after processing or calculation based on the bounding box data of the certificate image to be corrected and the face data of the certificate image to be corrected, that is, parameters such as rotation angle, scaling ratio, perspective transformation, and translation distance. Specifically, the above correction parameters can be used in the correction operation of the certificate image to be corrected. By correcting the certificate image to be corrected with the correction parameters, a relatively centered and horizontal certificate image can be obtained, and the text in the certificate image can also become more centered and horizontal.
[0059] In a possible embodiment, after obtaining the image to be recognized, the above image recognition platform preprocesses the image to be recognized to obtain the position of the certificate image in the image, and obtains the bounding box data and face data of the certificate image. The bounding box data and face data of the above certificate image are calculated to obtain correction parameters, and appropriate correction parameters can be selected according to the specific correction operation required. For example, if rotation is required to better obtain centered and horizontal text, the calculated rotation angle is used as the correction parameter to correct the above image to be recognized, or to correct the above certificate image to be corrected.
[0060] The above face key point data may refer to the coordinate positions or feature points of special parts on the human face, which are used for face recognition, such as Figure 2 As shown, the 68 key points in the figure are the positions of relatively special parts on the human face. These positions can all be used as face key points, and specific information about a certain part can be obtained by numbering the key points. For example, by confirming the information of key points numbered 36 to 41, the information of the right eye of the human face can be specifically determined.
[0061] The above correction parameters may include an included angle and a rotation direction. For example, it may be the included angle and rotation direction between the long side in the bounding box and the horizontal line. The above correction parameters may include the included angle and rotation direction between the line connecting face key points and the horizontal line.
[0062] 103. The certificate image to be corrected is corrected with the correction parameters to obtain the certificate image to be recognized.
[0063] In an embodiment of the present invention, the above correction process may refer to the process of horizontally correcting the image to be corrected through correction parameters. Specifically, different correction parameters can be selected for correction according to different image conditions. For example, when the face data in the image cannot be detected, the correction is performed according to the bounding box data of the image to be corrected. That is, data such as the bounding box ratio and standard shape of the bounding box data at this time can be used as correction parameters, and the above image to be corrected is horizontally corrected according to these correction parameters. In another possible embodiment, if the face data in the image to be corrected can be detected through face detection, the image can be horizontally corrected according to the eye slope among the above face key points. That is, the eye slope at this time is used as the correction parameter to horizontally correct the above image to be corrected.
[0064] Furthermore, the corresponding correction parameters can be calculated according to the long side in the bounding box. Specifically, the angle between the long side and the horizontal line can be calculated. At this time, the above correction parameters may include the angle and the rotation direction. After rotating the long side to the horizontal direction by the corresponding angle, the correction of the image of the document to be corrected can be completed. The corresponding correction parameters can be calculated according to the face key point data in the face data. At least two face key points that are horizontally distributed under normal conditions are included in the above face key point data. The two face key points are connected, and the angle between the connection line and the horizontal line is calculated. At this time, the above correction parameters may include the angle and the rotation direction. After rotating the connection line to the horizontal direction by the corresponding angle, the correction of the image of the document to be corrected can be completed.
[0065] The above image of the document to be recognized may refer to the image of the document obtained after horizontal correction through correction parameters. Compared with the image of the document to be corrected, this image of the document has the characteristics that the text is more horizontal and is more easily recognized by text recognition.
[0066] 104. Perform text recognition on the image of the document to be recognized to obtain the result of document text recognition of the image to be recognized.
[0067] In an embodiment of the present invention, the above text recognition may refer to converting the text in the image into editable and searchable text. Specifically, the above text recognition can be performed through OCR (Optical Character Recognition) technology. This technology is mainly divided into two types: optical OCR and visual OCR technology. Among them, optical OCR uses a laser or a camera to scan the text in the image, and then the text is recognized through software. Visual OCR uses computer vision technology to recognize the text in the image.
[0068] In a possible embodiment, the above image recognition platform can process the image to be recognized through the following process, such as Figure 3As shown, after the above image recognition platform obtains the image to be recognized, it preprocesses the image, then detects the object (i.e., the above-mentioned image of the certificate to be corrected), and can also extract the certificate-related characteristics of the object (i.e., the above-mentioned bounding box data) according to the algorithm model. Then, it judges the object to determine whether it has a certificate bounding box that meets the certificate bounding box ratio standard. If it meets the standard, it records the certificate position and provides the bounding box information of the certificate. Subsequently, it uses a face detector to detect the certificate. If no face is detected within the bounding box, it rotates the certificate by 90 degrees and detects it again. If no face is detected after rotation, it rotates the certificate by 90 degrees again and detects it again. When a face is detected during the rotation process, the rotation and face detection are stopped, and the face position and information are recorded. If no face is detected after rotating 3 times, the detection is stopped, and no subsequent operations are performed on the above-mentioned image to be recognized. The face key point information is returned by the face key point detector. When the eye key point information in the face key point information cannot be detected, the angle is corrected according to the image target certificate bounding box and the portrait position, and the corrected image is subjected to OCR text recognition. If the eye key point information in the face key point information can be detected, the coordinates of the left eye and the right eye are extracted and the slope between the two eyes is calculated. Then, the image is horizontally corrected according to the slope to obtain the corrected image, and the corrected image is subjected to OCR text recognition.
[0069] In the embodiment of the present invention, the bounding box data of the image of the certificate to be corrected and the face data of the image of the certificate to be corrected are obtained from the image to be recognized. Based on the bounding box data of the image of the certificate to be corrected and the face data of the image of the certificate to be corrected, the correction parameters of the image of the certificate to be corrected are determined. The image of the certificate to be corrected is corrected by the correction parameters to obtain the image of the certificate to be recognized, and the image of the certificate to be recognized is subjected to text recognition to obtain the certificate text recognition result of the image to be recognized. According to the obtained bounding box data and face data in the image of the certificate to be corrected, the correction parameters of the image of the certificate to be corrected are obtained, the image of the certificate to be corrected is corrected by the correction parameters, and then the corrected image of the certificate is subjected to text recognition processing, thereby improving the efficiency and accuracy of certificate text recognition.
[0070] Optionally, before the step of obtaining the bounding box data of the document image to be corrected and the face data of the document image to be corrected from the image to be recognized, the edge feature recognition process may be performed on the image to be recognized by a preset edge detection algorithm to obtain the edge information of the target object. Then, based on the edge information of the target object, the bounding box of the target object is determined. If the ratio of the bounding box of the target object conforms to the ratio of the document bounding box, it is detected whether there is a face within the bounding box. If there is a face within the bounding box, the document image is extracted from the image to be recognized based on the bounding box. Finally, the features of the document image are extracted to obtain the features of the document image, and classification or recognition is performed based on the features of the document image to obtain the document image to be corrected.
[0071] In the embodiment of the present invention, the above-mentioned preset edge detection algorithm can perform edge feature recognition processing on the image to be recognized, so that the document image to be corrected can be obtained from the image to be recognized. Generally speaking, the above-mentioned preset edge detection algorithm separates the edge information of each independent object in the above-mentioned image to be recognized from the background of the image to be recognized, so as to achieve the purpose of separating the image to be recognized and obtaining the document image to be corrected. The above-mentioned preset edge detection algorithm can be applied to feature recognition models such as YOLO (You Only Look Once), Faster R-CNN, and SSD (Single Shot MultiBoxDetector). These feature recognition models all have certain advantages and applicable scenarios, and can be used alone or in combination of multiple models according to specific implementation schemes to achieve the purpose of obtaining the document image to be corrected.
[0072] The above-mentioned edge information may include, but is not limited to, the edge angle and edge shape feature data. The above-mentioned edge angle may refer to the angle between the pixel points at the boundary of the target object in the above-mentioned image to be recognized and the horizontal direction. The above-mentioned edge angle can be used to describe the texture and shape of the image, as well as perform tasks such as image segmentation. Generally speaking, the angle between the pixel points at the edge of the standard document image and the horizontal direction should be 180 degrees, that is, the pixel points at the edge of the standard document image are in a horizontal state with the horizontal direction.
[0073] The above-mentioned bounding box may refer to the boundary of the above-mentioned target object, that is, the area of the target object in the above-mentioned image to be recognized can be determined according to the bounding box.
[0074] The above-mentioned ratio of the document bounding box may refer to the length and width of the bounding box of the target object, the ratio of the bounding box, and the shape of the bounding box. For example, the size of the standard ID card photo is 8.56 cm × 5.4 cm × 0.1 cm. Among them, the size of the portrait photo is 2.6 cm * 3.2 cm, 358 pixels (width) × 441 pixels (height), the resolution is 350 dpi, and the shape is rectangular.
[0075] In a possible embodiment, after obtaining the image to be recognized, the above-mentioned image recognition platform performs edge feature recognition processing on the image to be recognized through a preset edge detection algorithm in the platform, obtains the edge information of at least one target object, and based on the edge information of each target object, obtains at least one target object's edge box in the image to be recognized. Then, by calculating the ratio of the obtained target object's edge box, if the calculated edge box ratio conforms to the certificate boundary box ratio, it is detected whether there is a face within the edge box. If there is a face within the edge box, the certificate image corresponding to the target object is extracted from the image to be recognized according to the above-mentioned edge box, and then feature extraction is performed on the certificate image, and the certificate image is classified or recognized according to the features of the certificate image to obtain the certificate image to be corrected.
[0076] Optionally, in the step of performing feature extraction on the certificate image to obtain the features of the certificate image, feature extraction can also be performed on the certificate image, including the boundary box features and portrait features of the certificate image, and then classification or recognition is performed based on the features of the certificate image to obtain the certificate image to be corrected, including classifying or recognizing based on the boundary box features and portrait features of the certificate image, and judging whether there is a certificate image that meets the expectations in the image to be processed based on the results of the classification or recognition. If there is a certificate image that meets the expectations, the certificate image that meets the expectations is extracted as the certificate image to be corrected.
[0077] In the embodiment of the present invention, a preset face feature extractor can be used to perform portrait feature extraction on the face area in the above-mentioned certificate image to obtain portrait features. Specifically, the face feature extractor can be set according to specific implementation schemes. Generally, face feature extractors such as OpenCV, Dlib, and Face++ can be used. The above-mentioned face feature extractors are usually trained and learned with a large number of face samples, and they each have their own advantages and disadvantages, and a suitable face feature extractor can be selected according to specific application scenarios. The above-mentioned face area can be a portrait contour, or specifically refer to the face area within the portrait contour, that is, the curve area from the hairline to the chin. It can be considered that only face features exist in this area, so the extraction efficiency of face data in this area is high and accurate.
[0078] The boundary box features of the above-mentioned certificate image can refer to features such as the ratio of the boundary box and the shape of the boundary box. For example, different certificates need to comply with corresponding manufacturing border ratios during the manufacturing process, that is, the length-width ratio, and due to the influence of factors such as the environment of the obtained image to be recognized, the boundary box shape will be tilted, resulting in an irregular boundary box shape such as a trapezoid or a parallelogram.
[0079] In a possible embodiment, the features of the above-mentioned certificate image can also be recognized by a preset classifier, and the features corresponding to the corresponding certificate types can be queried in the database of the above-mentioned image recognition platform, and each of the above-mentioned certificate images can be classified according to the features of the corresponding certificate types. According to the classification results, it is determined whether there is a certificate image that needs to be corrected in the above-mentioned image to be recognized (i.e., the expected certificate image). If there is the above-mentioned expected certificate image, each of the above-mentioned expected certificate images contained in the above-mentioned image to be recognized is used as the certificate image to be corrected.
[0080] In a possible embodiment, after the above-mentioned image recognition platform obtains the certificate image,
[0081] The portrait features are extracted from the above-mentioned certificate image by a preset portrait feature extractor to obtain portrait features, and the bounding box features of the above-mentioned certificate image are obtained by performing recognition processing on the bounding box of the above-mentioned certificate image. Then, according to a preset classifier, classification or recognition processing is performed on the portrait features and bounding box features of the above-mentioned certificate image to obtain classification results of various certificate types, and according to the classification results, it is determined whether there is a certificate image that meets the expectations in the above-mentioned certificate image, and the certificate image that meets the expectations is extracted from the above-mentioned certificate image and saved as the certificate image to be corrected.
[0082] Optionally, in the step of obtaining the bounding box data of the certificate image to be corrected and the face data of the certificate image to be corrected from the image to be recognized, the bounding box data of the certificate image to be corrected can also be determined according to the edge information corresponding to the certificate image to be corrected, and then face key point detection is performed on the face in the certificate image to be corrected to obtain the face data of the certificate image to be corrected.
[0083] In the embodiment of the present invention, the above-mentioned bounding box data may refer to the length, width data and the border ratio data of the border of the certificate image to be corrected that needs to be horizontally corrected. The bounding box data of the above-mentioned certificate image to be corrected can be obtained by a preset recognizer. The above-mentioned preset recognizer can use a preset edge detection algorithm to recognize the length, width and ratio of the bounding box of the certificate image to be corrected. For example, a model using edge detection algorithms such as Sobel, Laplacian and Krisch is used to obtain the above-mentioned bounding box data of the certificate image to be corrected. Due to the differences in different countries, regions and certificate types, the bounding box ratios of the above-mentioned standard certificates are also different. Therefore, in this embodiment, a standard certificate database is established, which contains the bounding box ratios and requirements of the standard certificates.
[0084] In a possible embodiment, after obtaining the image of the certificate to be corrected, the above-mentioned image recognition platform uses a preset recognizer to recognize the bounding box of the image of the certificate to be corrected, obtains the bounding box data of the image of the certificate to be corrected, and then uses the above-mentioned face detector to detect the face key points in the image of the certificate to be corrected, so as to obtain the face data of the image of the certificate to be corrected.
[0085] Optionally, in the step of determining the correction parameters of the image of the certificate to be corrected based on the bounding box data of the image of the certificate to be corrected and the face data of the image of the certificate to be corrected, it is also possible to determine whether there is face key point information in the face data. If there is no face key point information in the face data, the correction parameters of the image of the certificate to be corrected are determined through the bounding box data of the image of the certificate to be corrected. If there is face key point information in the face data, the correction parameters of the image of the certificate to be corrected are determined through the face key point information of the image of the certificate to be corrected.
[0086] In the embodiment of the present invention, the above-mentioned face key point information can be detected by a face key point detector. The face key point detector is a model used to detect face key points and obtain face key point information. The face key point information can refer to the contours and feature points of parts such as eyes, nose, eyebrows, mouth, and ears. The information of these feature points plays an important role in face comparison and face recognition.
[0087] In this embodiment, based on the detection results of face key points, a suitable correction method is selected to obtain effective correction parameters to correct the image of the certificate to be corrected. The above-mentioned suitable correction method can be to adjust the bounding box shape of the image of the certificate to be corrected to the bounding box shape of a standard certificate, such as a rectangle, according to the bounding box data of the image of the certificate to be corrected, such as the bounding box shape, so as to achieve the purpose of horizontally correcting the image of the certificate to be corrected. Or it can be to perform horizontal correction calculation through the face key points in the image of the certificate to be corrected to obtain correction parameters, and correct the image of the certificate to be corrected according to the correction parameters. The above-mentioned horizontal correction calculation can refer to the process of obtaining the angle of the horizontal angle between two symmetric key points based on the coordinate information in the key point information, and obtaining the correction parameters that need to be adjusted according to this angle.
[0088] Specifically, in a possible embodiment, when the above-mentioned bounding box shape is a relatively regular quadrilateral graph, such as a trapezoid, when horizontally correcting the image of the certificate to be corrected through the bounding box data, the longest side of the bounding box can also be horizontally corrected to achieve the effect of correcting the image of the certificate to be corrected.
[0089] In another possible embodiment, when the above-mentioned bounding box is in the shape of an irregular quadrilateral, the irregular quadrilateral can also be transformed through a spatial transformation algorithm STN (Spatial Transformer Networks). The above-mentioned bounding box is spatially transformed through the STN algorithm to achieve the purpose of horizontally correcting the document image to be corrected.
[0090] In a possible embodiment, when the above-mentioned face key point detector fails to detect face key point information in the document image to be corrected, the document image to be corrected is corrected based on the bounding box data extracted by the above-mentioned preset bounding box feature extraction model. When the above-mentioned face key point detector detects face key point information in the document image to be corrected, information pairs in the current face key point information are obtained for horizontal correction calculation to obtain correction parameters, and the document image to be corrected is corrected. The above-mentioned information pairs may refer to key points in the face key points that can form a pair in face symmetry, such as the key points at the corners of the left eye and the right eye, the left and right corners of the mouth, and the left and right cheek dimples.
[0091] In another possible embodiment, for example, during a census or when registering relevant personnel information, it may be necessary to perform blurring and encryption processing on the ID photos of the above-mentioned relevant personnel. At this time, the face key point information on the ID photos of the above-mentioned relevant personnel will be covered with occlusion codes such as mosaics. If horizontal correction of the document image to be corrected is required in such cases, the document image to be corrected still needs to be horizontally corrected according to the shape of the bounding box in the above-mentioned bounding box data.
[0092] Optionally, in the step of determining the correction parameters of the document image to be corrected through the face key point information of the document image to be corrected if there is face key point information in the face data, it is also possible to extract the eye key point information from the face key point information if there is face key point information in the face data, and then determine the correction parameters of the document image to be corrected based on the eye key point information.
[0093] In the embodiments of the present invention, the above-mentioned eye key point information may refer to the information of multiple key points that specifically make up the eye part, as shown above Figure 2 where the key points from No. 36 to No. 47 form the eye key points, that is, the above-mentioned eye key point information may refer to the key point information of the above-mentioned key points from No. 36 to No. 47.
[0094] The above-mentioned face key point information may include but is not limited to eye key point information, mouth corner key point information, and dimple key point information, etc. In a certain embodiment, if there is a lack of eye key point information, the remaining key point information can be matched, and a group of key point information with symmetry is used for horizontal correction calculation to obtain correction parameters.
[0095] Specifically, in one embodiment, when the eye key point information is missing, the complete mouth corner key point information can be extracted to obtain the left mouth corner key point information and the right mouth corner key point information. By performing horizontal correction calculation on the left and right mouth corner key point information, a correction parameter can be obtained.
[0096] Optionally, in the step of determining the correction parameter of the certificate image to be corrected based on the human eye key point information, the slope between the left eye key point and the right eye key point can also be determined based on the coordinates of the left eye key point and the coordinates of the right eye key point. Then, based on the slope between the left eye key point and the right eye key point, the tilt angle of the face portrait in the certificate image to be corrected is determined. Finally, based on the tilt angle of the face portrait in the certificate image to be corrected, the correction parameter of the certificate image to be corrected is determined.
[0097] In the embodiment of the present invention, the above human eye key point information may include, but is not limited to, the coordinates of the left eye key point and the coordinates of the right eye key point. The coordinate data of the above human eyes can be obtained according to the following formula:
[0098] left_eye = (left_eye[0], left_eye[1])
[0099] right_eye = (right_eye[0], right_eye[1])
[0100] Among them, ledt_eye is the left eye coordinate, left_eye[0] is the x-axis coordinate of a certain key point of the left eye, left_eye[1] is the y-axis coordinate of the same key point of the left eye. Similarly, right_eye is the right eye coordinate, right_eye[0] is the x-axis coordinate of a certain key point of the right eye, and right_eye[1] is the y-axis coordinate of the same key point of the right eye. The number of different key points can be input in [] to obtain the coordinate information of the current key point.
[0101] The above slope data between human eyes can be obtained by the following formula:
[0102] dx = right_eye[0] - left_eye[0]
[0103] dy = right_eye[1] - left_eye[1]
[0104] angle = math.atan2(dy, dx) × 180 / math.pi
[0105] Among them, dx is the x-axis difference in the slope data between the two eyes, dy is the y-axis difference in the slope data between the two eyes, and angle is the slope calculated through a pair of symmetric key points between the two eyes. math.atan2(dy, dx) is the radian of the human eye, and the slope calculation can be obtained by multiplying the radian by 180 / math.pi.
[0106] Specifically, the slope data calculated according to the above can be used as the tilt angle of the face avatar. Therefore, after comparing the slope data with the horizontal angle, the angle that needs to be horizontally adjusted is obtained, and the tilt angle of the face avatar is horizontally corrected through the angle that needs to be horizontally adjusted, so as to achieve the purpose of correcting the image through the slope.
[0107] More specifically, in combination with Figure 2 it is described that the key points 36 of the human eye and the key points 45 of the human eye are selected for slope calculation to obtain the tilt angle of the face avatar. The coordinates of the key point 36 of the left eye and the key point 45 of the right eye can be obtained according to the following formula:
[0108] left_eye = ((36).x, (36).y)
[0109] right_eye = ((45).x, (45).y)
[0110] Then, the slope between the two eyes is calculated according to the following formula:
[0111] dx = (45).x - (36).x
[0112] dy = (45).y - (36).y
[0113] angle = math.atan2(dy, dx) × 180 / math.pi
[0114] Through the above formula, the slope between the key point 36 of the human eye and the key point 45 of the human eye can be obtained, and this slope is used as the correction parameter of the face avatar in the certificate image to be corrected, and the face avatar is horizontally corrected.
[0115] As Figure 4 shown, an embodiment of the present invention further provides a certificate text recognition device, and the certificate text recognition device includes:
[0116] The first acquisition module 401 is used to acquire the bounding box data of the certificate image to be corrected and the face data of the certificate image to be corrected from the image to be recognized;
[0117] The first determination module 402 is configured to determine the correction parameters of the image of the certificate to be corrected based on the bounding box data of the image of the certificate to be corrected and the face data of the image of the certificate to be corrected;
[0118] The first processing module 403 is configured to perform correction processing on the image of the certificate to be corrected through the correction parameters to obtain the image of the certificate to be recognized;
[0119] The recognition module 404 is configured to perform character recognition on the image of the certificate to be recognized to obtain the certificate character recognition result of the image to be recognized.
[0120] Optionally,
[0121] The above device further includes:
[0122] The second acquisition module is configured to perform edge feature recognition processing on the image to be recognized through a preset edge detection algorithm to obtain the edge information of the target object;
[0123] The second determination module is configured to determine the bounding box of the target object based on the edge information of the target object;
[0124] The detection module is configured to detect whether there is a face within the bounding box if the ratio of the bounding box of the target object conforms to the ratio of the certificate bounding box;
[0125] The extraction module is configured to extract the certificate image from the image to be recognized based on the bounding box if there is a face within the bounding box;
[0126] The classification module is configured to extract the features of the certificate image to obtain the features of the certificate image, and classify or recognize based on the features of the certificate image to obtain the image of the certificate to be corrected.
[0127] Optionally, the above classification module further includes:
[0128] The first extraction sub-module is configured to extract the features of the certificate image to obtain the bounding box features of the certificate image and the portrait features in the certificate image;
[0129] The first classification sub-module is configured to classify or recognize based on the features of the certificate image to obtain the image of the certificate to be corrected, including:
[0130] The first classification unit is configured to classify or recognize based on the bounding box features and the portrait features of the certificate image, and determine whether there is a certificate image that meets the expectation in the image to be processed based on the result of the classification or recognition;
[0131] The first extraction unit is configured to extract the certificate image that meets the expectation as the image of the certificate to be corrected if there is a certificate image that meets the expectation.
[0132] Optionally, the above first acquisition module 401 further includes:
[0133] A first determination sub-module, configured to determine the bounding box data of the to-be-corrected document image according to the edge information corresponding to the to-be-corrected document image;
[0134] A first acquisition sub-module, configured to perform face key point detection on the face in the to-be-corrected document image to obtain the face data of the to-be-corrected document image.
[0135] Optionally, the above first determination module 402 includes:
[0136] A first judgment sub-module, configured to judge whether there is face key point information in the face data;
[0137] A second determination sub-module, configured to, if there is no face key point information in the face data, determine the correction parameter of the to-be-corrected document image through the bounding box data of the to-be-corrected document image;
[0138] A third determination sub-module, configured to, if there is face key point information in the face data, determine the correction parameter of the to-be-corrected document image through the face key point information of the to-be-corrected document image. Optionally, the above third determination sub-module further includes:
[0139] A second extraction unit, configured to, if there is face key point information in the face data, extract the key point information of the human eyes from the face key point information;
[0140] A first determination unit, configured to determine the correction parameter of the to-be-corrected document image based on the key point information of the human eyes.
[0141] Optionally, the above first determination unit further includes:
[0142] A first determination sub-unit, configured to determine the slope between the left eye key point and the right eye key point based on the coordinates of the left eye key point and the coordinates of the right eye key point;
[0143] A second determination sub-unit determines the tilt angle of the face portrait in the to-be-corrected document image based on the slope between the left eye key point and the right eye key point;
[0144] A third determination sub-unit determines the correction parameter of the to-be-corrected document image based on the tilt angle of the face portrait in the to-be-corrected document image.
[0145] As Figure 5 shown, an embodiment of the present invention further provides an electronic device, including a processor, and the above processor can execute any one of the above document character recognition methods.
[0146] Specifically, it includes a processor 501, a memory 502, and a computer program stored on the memory 502 and capable of running on the processor 501 to execute the document character recognition method, where:
[0147] The processor 501 runs the calculator program of the document character recognition method stored in the memory 502 and executes the following steps:
[0148] Obtain the bounding box data of the document image to be corrected and the face data of the document image to be corrected from the image to be recognized;
[0149] Based on the bounding box data of the document image to be corrected and the face data of the document image to be corrected, determine the correction parameters of the document image to be corrected;
[0150] Perform correction processing on the document image to be corrected through the correction parameters to obtain the document image to be recognized;
[0151] Perform character recognition on the document image to be recognized to obtain the document character recognition result of the image to be recognized.
[0152] Optionally, the steps executed by the processor 501 before obtaining the bounding box data of the document image to be corrected and the face data of the document image to be corrected from the image to be recognized include:
[0153] Perform edge feature recognition processing on the image to be recognized through a preset edge detection algorithm to obtain the edge information of the target object;
[0154] Based on the edge information of the target object, determine the bounding box of the target object;
[0155] If the aspect ratio of the bounding box of the target object conforms to the aspect ratio of the document bounding box, detect whether there is a face inside the bounding box;
[0156] If there is a face inside the bounding box, extract the document image from the image to be recognized based on the bounding box;
[0157] Perform feature extraction on the document image to obtain the features of the document image, and perform classification or recognition based on the features of the document image to obtain the document image to be corrected.
[0158] Optionally, the steps executed by the processor 501 to perform feature extraction on the document image to obtain the features of the document image include:
[0159] Perform feature extraction on the document image to obtain the bounding box features of the document image and the portrait features in the document image;
[0160] Classifying or recognizing based on the features of the certificate image to obtain the certificate image to be corrected, including:
[0161] Classifying or recognizing based on the bounding box features and the portrait features of the certificate image, and judging whether there is a certificate image meeting the expectation in the image to be processed based on the classification or recognition result;
[0162] If there is a certificate image meeting the expectation, extract the certificate image meeting the expectation as the certificate image to be corrected.
[0163] Optionally, the processor 501 executes the step of obtaining the bounding box data of the certificate image to be corrected and the face data of the certificate image to be corrected from the image to be recognized, including:
[0164] Determine the bounding box data of the certificate image to be corrected according to the edge information corresponding to the certificate image to be corrected;
[0165] Perform face key point detection on the face in the certificate image to be corrected to obtain the face data of the certificate image to be corrected.
[0166] Optionally, the processor 501 executes the step of determining the correction parameters of the certificate image to be corrected based on the bounding box data of the certificate image to be corrected and the face data of the certificate image to be corrected, including:
[0167] Judge whether there is face key point information in the face data;
[0168] If there is no face key point information in the face data, determine the correction parameters of the certificate image to be corrected through the bounding box data of the certificate image to be corrected;
[0169] If there is face key point information in the face data, determine the correction parameters of the certificate image to be corrected through the face key point information of the certificate image to be corrected.
[0170] Optionally, in the above certificate text recognition method, the processor 501 executes the step of if there is face key point information in the face data, then determining the correction parameters of the certificate image to be corrected through the face key point information of the certificate image to be corrected, including:
[0171] If there is face key point information in the face data, extract the eye key point information from the face key point information;
[0172] Determine the correction parameters of the certificate image to be corrected based on the eye key point information.
[0173] Optionally, in the above-mentioned document text recognition method, the human eye key point information includes the coordinates of the left eye key point and the coordinates of the right eye key point. When the processor 501 executes the step of determining the correction parameter of the document image to be corrected based on the human eye key point information, it includes:
[0174] Based on the coordinates of the left eye key point and the coordinates of the right eye key point, determine the slope between the left eye key point and the right eye key point;
[0175] Based on the slope between the left eye key point and the right eye key point, determine the tilt angle of the face portrait in the document image to be corrected;
[0176] Based on the tilt angle of the face portrait in the document image to be corrected, determine the correction parameter of the document image to be corrected.
[0177] The embodiment of the present invention also provides a computer-readable storage medium. A computer program is stored on the computer-readable storage medium. When the computer program is executed by a processor, it realizes each process of the document text recognition method or the application-side document text recognition method provided by the embodiment of the present invention, and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.
[0178] Those of ordinary skill in the art can understand that all or part of the processes of implementing the methods in the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. Among them, the computer-readable storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM), etc.
[0179] The above-disclosed are only the preferred embodiments of the present invention. Of course, the scope of the rights of the present invention cannot be limited by this. Therefore, equivalent changes made according to the claims of the present invention still fall within the scope covered by the present invention.
Claims
1. A method for identifying document text, characterized in that, the method comprises the following steps: Obtain the bounding box data of the document image to be corrected and the face data of the document image to be corrected from the image to be recognized; Based on the bounding box data of the document image to be corrected and the face data of the document image to be corrected, determine the correction parameters of the document image to be corrected; Perform correction processing on the document image to be corrected through the correction parameters to obtain the document image to be recognized; Perform text recognition on the document image to be recognized to obtain the document text recognition result of the image to be recognized.
2. The document text recognition method according to claim 1, characterized in that, before obtaining the bounding box data of the document image to be corrected and the face data of the document image to be corrected from the image to be recognized, the method further comprises: Perform edge feature recognition processing on the image to be recognized through a preset edge detection algorithm to obtain the edge information of the target object; Based on the edge information of the target object, determine the bounding box of the target object; If the bounding box ratio of the target object conforms to the document bounding box ratio, detect whether there is a face inside the bounding box; If there is a face inside the bounding box, extract the document image from the image to be recognized based on the bounding box; Perform feature extraction on the document image to obtain the features of the document image, and classify or recognize based on the features of the document image to obtain the document image to be corrected.
3. The document text recognition method according to claim 2, characterized in that, performing feature extraction on the document image to obtain the features of the document image includes: Performing feature extraction on the document image to obtain the bounding box features of the document image and the portrait features in the document image; classifying or recognizing based on the features of the document image to obtain the document image to be corrected includes: Classifying or recognizing based on the bounding box features and the portrait features of the document image, and judging whether there is a document image meeting the expectation in the image to be processed based on the classification or recognition result; If there is a document image meeting the expectation, extract the document image meeting the expectation as the document image to be corrected.
4. The document text recognition method according to claim 2, characterized in that, obtaining the bounding box data of the document image to be corrected and the face data of the document image to be corrected from the image to be recognized includes: Determine the bounding box data of the document image to be corrected according to the corresponding edge information of the document image to be corrected; Perform face key point detection on the face in the document image to be corrected to obtain the face data of the document image to be corrected.
5. The document text recognition method according to claim 1, characterized in that, determining the correction parameters of the document image to be corrected based on the bounding box data of the document image to be corrected and the face data of the document image to be corrected includes: Judge whether there is face key point information in the face data; If there is no face key point information in the face data, determine the correction parameters of the document image to be corrected through the bounding box data of the document image to be corrected; If there is face key point information in the face data, the correction parameters of the to-be-corrected document image are determined based on the face key point information of the to-be-corrected document image.
6. The document text recognition method according to claim 5, wherein, the step of if there is face key point information in the face data, the correction parameters of the to-be-corrected document image are determined based on the face key point information of the to-be-corrected document image, includes: if there is face key point information in the face data, the eye key point information is extracted from the face key point information; Based on the eye key point information, the correction parameters of the to-be-corrected document image are determined.
7. The document text recognition method according to claim 6, wherein, the eye key point information includes the coordinates of the left eye key point and the coordinates of the right eye key point, and the step of based on the eye key point information, the correction parameters of the to-be-corrected document image are determined, includes: Based on the coordinates of the left eye key point and the coordinates of the right eye key point, the slope between the left eye key point and the right eye key point is determined; Based on the slope between the left eye key point and the right eye key point, the tilt angle of the face portrait in the to-be-corrected document image is determined; Based on the tilt angle of the face portrait in the to-be-corrected document image, the correction parameters of the to-be-corrected document image are determined.
8. A document text recognition device, wherein, the traction component monitoring device includes: A first acquisition module, configured to acquire the bounding box data of the to-be-corrected document image and the face data of the to-be-corrected document image from the to-be-recognized image; A first determination module, configured to determine the correction parameters of the to-be-corrected document image based on the bounding box data of the to-be-corrected document image and the face data of the to-be-corrected document image; A first processing module, configured to perform correction processing on the to-be-corrected document image through the correction parameters to obtain a to-be-recognized document image; A recognition module, configured to perform text recognition on the to-be-recognized document image to obtain the document text recognition result of the to-be-recognized image.
9. An electronic device, wherein, including: A memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the computer program, the steps in the document text recognition method according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium, wherein, a computer program is stored on the computer-readable storage medium, and when the computer program is executed by a processor, the steps in the document text recognition method according to any one of claims 1 to 7 are implemented.
Citation Information
Cited By
Foreign certificate identification method and system based on multi-modal fusion
CN122135187A
A method and system for foreign affairs document recognition based on multimodal fusion
CN122135187B