Passport information collection method and system based on image detection and recognition
Through the passport information acquisition method based on image detection and recognition, the deep learning model and optical character recognition function are used to extract features and build a passport recognition model, which solves the problems of low passport information acquisition efficiency and low recognition rate in the prior art, and achieves higher accuracy and reliability.
Patent Information
- Application Number
- CN202411855323.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-17
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2044-12-17
AI Technical Summary
The existing passport information collection methods are inefficient, prone to errors, and are difficult to adapt to the diversified design of passports in different countries and regions, resulting in low recognition rates.
Passport information acquisition method based on image detection and recognition is adopted, by obtaining the passport acquisition image, determining whether it is out of focus, performing pre-processing, and using deep learning models and optical character recognition functions to extract pattern features, text features and text features, constructing and training the passport recognition model to achieve automated recognition.
It improves the accuracy of passport acquisition images under different conditions, reduces manual intervention, reduces operational costs, and improves data processing speed and data reliability and security.
Smart Images

Figure CN119723602B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of image detection and recognition, and in particular to a passport information collection method and system based on image detection and recognition. Background Art
[0002] A passport is an important ID document for every international traveler. As a legal proof of citizenship, it carries the basic information of the holder, including name, gender, date of birth, nationality, passport number, date of issue and validity period. With the acceleration of globalization, the frequency of passport use has increased significantly, and the demand for entry and exit management and identity verification has become increasingly urgent. The traditional method of entering passport information mainly relies on manual input, which is not only inefficient but also prone to errors, resulting in identity recognition loopholes and security risks.
[0003] Currently, the methods for collecting passport information mainly include manual entry and simple scanning recognition. Manual entry is prone to data input errors and can easily cause congestion during peak hours. Although existing scanning technology can increase the speed of data collection, its accuracy is still limited, especially when dealing with uneven lighting, blurred images or damaged passports, the recognition rate is often lower than expected. In addition, most systems lack the ability to adapt to the diversity and complexity of passports and cannot effectively handle the diverse designs of passports from different countries and regions.
[0004] In view of this, the present invention is proposed. Summary of the invention
[0005] In order to solve the above technical problems or at least partially solve the above technical problems, the embodiments of the present disclosure provide a passport information collection method and system based on image detection and recognition, which improves the accuracy of passport captured images under different conditions, reduces manual intervention, reduces operating costs, and improves data processing speed as well as data reliability and security.
[0006] In a first aspect, an embodiment of the present disclosure provides a passport information collection method based on image detection and recognition, the method comprising:
[0007] Acquire a passport captured image, and determine whether the passport captured image is out of focus;
[0008] When it is determined that the passport captured image is not out of focus, preprocessing the passport captured image to obtain a preprocessed passport image;
[0009] Performing feature extraction on the pre-processed passport image based on a deep learning model to obtain pattern features and text features, wherein the pattern features correspond to the pattern in the passport captured image and the border of the passport, and the text features correspond to the Chinese characters in the passport captured image;
[0010] Extracting features of the pre-processed passport image based on an optical character recognition function to obtain text features, wherein the text features correspond to non-Chinese characters, letters, pinyin and numbers in the captured passport image;
[0011] Building and training a passport recognition model based on the pattern features, the character features and the text features to obtain a trained passport recognition model;
[0012] Input the new pattern features, character features and text features into the trained passport recognition model to obtain the recognition results.
[0013] In a second aspect, an embodiment of the present disclosure further provides a passport information collection system based on image detection and recognition, the system comprising:
[0014] An image acquisition unit, used to acquire a passport captured image and determine whether the passport captured image is out of focus;
[0015] A preprocessing unit, configured to preprocess the passport captured image to obtain a preprocessed passport image when determining that the passport captured image is not out of focus;
[0016] A pattern and text feature acquisition unit, used for performing feature extraction on the pre-processed passport image based on a deep learning model to obtain pattern features and text features, wherein the pattern features correspond to the pattern in the passport captured image and the border of the passport, and the text features correspond to the Chinese characters in the passport captured image;
[0017] A text feature acquisition unit, configured to extract features from the pre-processed passport image based on an optical character recognition function to obtain text features, wherein the text features correspond to non-Chinese characters, letters, pinyin and numbers in the captured passport image;
[0018] A model acquisition unit, used to construct and train a passport recognition model based on the pattern feature, the character feature and the text feature to obtain a trained passport recognition model;
[0019] The result acquisition unit is used to input new pattern features, character features and text features into the trained passport recognition model to obtain recognition results.
[0020] In a third aspect, an embodiment of the present disclosure further provides an electronic device, comprising: one or more processors; a storage device for storing one or more programs; when the one or more programs are executed by the one or more processors, the one or more processors implement the passport information collection method based on image detection and recognition as described above.
[0021] In a fourth aspect, an embodiment of the present disclosure further provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the passport information collection method based on image detection and recognition as described above.
[0022] The passport information collection method based on image detection and recognition provided by the embodiment of the present disclosure includes: obtaining a passport collection image and determining whether the passport collection image is out of focus; when determining that the passport collection image is not out of focus, preprocessing the passport collection image to obtain a preprocessed passport image; extracting features from the preprocessed passport image based on a deep learning model to obtain pattern features and text features; extracting features from the preprocessed passport image based on an optical character recognition function to obtain text features; constructing and training a passport recognition model based on pattern features, text features and text features to obtain a trained passport recognition model; inputting new pattern features, text features and text features into the trained passport recognition model to obtain a recognition result. The present disclosure avoids invalid processing, recognition and review of the out-of-focus image in subsequent processing work by judging the out-of-focus situation of the image, thereby saving computing resources and labor costs and ensuring that the quality of the collected image meets the expected standard; secondly, by using the passport recognition model to recognize the passport, the accuracy of the passport collection image under different conditions is improved, manual intervention is reduced, operating costs are reduced, and data processing speed, as well as data reliability and security are improved. In addition, the present application can implement parallel or serial feature extraction, and can adopt different strategies according to different situations (for example, different data volumes, urgency, etc.), thereby increasing selectivity and data processing efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] The above and other features, advantages and aspects of the embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. Throughout the accompanying drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic and the originals and elements are not necessarily drawn to scale.
[0024] Figure 1 The present invention is a flowchart of a passport information collection method based on image detection and recognition in an embodiment of the present invention.
[0025] Figure 2 The present invention is a schematic diagram of the structure of a passport information collection system based on image detection and recognition in an embodiment of the present invention.
[0026] Figure 3 It is a structural schematic diagram of an electronic device in an embodiment of the present disclosure. DETAILED DESCRIPTION
[0027] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as being limited to the embodiments described herein, which are instead provided for a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not intended to limit the scope of protection of the present disclosure.
[0028] It should be noted that the concepts such as "first" and "second" mentioned in the present disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.
[0029] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only used for illustrative purposes and are not used to limit the scope of these messages or information.
[0030] Traditional passport information collection technology still has certain limitations when facing complex backgrounds, different fonts and handwriting. With the rise of deep learning, especially the application of convolutional neural networks, the accuracy of image recognition has been significantly improved. Through large-scale data training, deep learning models can automatically learn image features, thereby improving the accuracy and robustness of recognition. This enables passport information collection systems based on image detection and recognition to provide more reliable solutions when facing diverse passports and complex environments. Globally, there is a growing demand for efficient collection and processing of passport information in fields such as entry and exit management, border inspection, airlines and financial institutions. The traditional manual entry method can no longer meet the rapidly growing market demand. Therefore, the present invention aims to propose a passport information collection method based on image detection and recognition.
[0031] In response to the above problems, the embodiments of the present disclosure provide a passport information collection method based on image detection and recognition, which improves the accuracy of passport collection images under different conditions, reduces manual intervention, reduces operating costs, and improves data processing speed as well as data reliability and security.
[0032] Figure 1 The flowchart of a passport full information collection method based on image detection and recognition in an embodiment of the present disclosure is shown in FIG. The method can be executed by a passport information collection system based on image detection and recognition. The device can be implemented in software and / or hardware. The system can be configured in an electronic device. Figure 1 As shown, the method may specifically include the following steps:
[0033] S110: Acquire a passport captured image, and determine whether the passport captured image is out of focus.
[0034] Specifically, a high-resolution camera or scanner can be used to automatically shoot or scan the passport to obtain the passport image. During the acquisition process, the system monitors the quality of the passport image, and needs to ensure the clarity and quality of the passport image to ensure that each passport image is clear and readable. If the passport image is found to be out of focus, the lens needs to be adjusted and re-shot. When it is determined that the passport image is not out of focus, the next step is carried out.
[0035] In an optional embodiment, whether the passport captured image is out of focus is determined based on the passport captured image and a first formula, wherein the expression of the first formula is:
[0036]
[0037] Where C represents the defocus value, A represents the passport image, μ represents the mean value of the passport image A, δ represents the pixel gradient function of the passport image A, and a 1 represents the pixel gradient adjustment factor of the passport image A, a 2 represents the pixel defocus adjustment factor of the passport image A, p j represents the probability of gray level j of passport captured image A.
[0038] It can be understood that the defocus threshold can be used to determine whether the passport captured image is out of focus. Specifically, when the defocus degree value is less than the defocus threshold, it is determined that the passport captured image is out of focus. The defocus degree value is determined by the first formula and the passport captured image, wherein the expression of the first formula is:
[0039]
[0040] Where C represents the defocus value, A represents the passport image, μ represents the mean value of the passport image A, δ(A) represents the pixel gradient function of the passport image A, and a 1 represents the pixel gradient adjustment factor of the passport image A, a 2 represents the pixel defocus adjustment factor of the passport image A, p j Represents the probability of gray level j of passport captured image A.
[0041] This embodiment can improve the accuracy of overall recognition by judging the out-of-focus and filtering unqualified images. In addition, it can judge whether the image is out-of-focus during the acquisition stage, which can avoid the subsequent processing work of invalid processing, identification and review of the out-of-focus image, thereby saving computing resources and labor costs.
[0042] S120: When it is determined that the passport captured image is not out of focus, pre-process the passport captured image to obtain a pre-processed passport image.
[0043] On the basis of the above embodiment, when it is determined that the passport captured image is not out of focus, that is, after obtaining a passport captured image that meets the definition and quality requirements, the passport captured image is preprocessed to obtain a preprocessed passport image, wherein the expression of the preprocessed passport image is:
[0044]
[0045] Where P represents the pre-processed passport image, b 1 represents the intensity adjustment factor of the passport image A, b 2 represents the intensity shift factor of the passport image A, N represents the total number of pixels of the passport image A, b 3 represents the pixel gradient weight factor, represents the pixel density probability function in the x direction, Represents the pixel density probability function in the y direction.
[0046] This embodiment improves the accuracy of subsequent image recognition by preprocessing the passport captured image.
[0047] S130: Extract features of the pre-processed passport image based on a deep learning model to obtain pattern features and text features.
[0048] Exemplarily, the deep learning model is a general model, which can be an automatically supervised deep learning model, composed of k convolutional layers, and the pre-processed passport image is input into the deep learning model for feature extraction, and pattern features and text features can be output, wherein the pattern features come from the patterns (including photos, national emblems, seals, signatures, etc.) and frames in the passport, and the text features come from the Chinese characters in the passport. Specifically, the sources of the pattern features and text features in the present application are not limited thereto, but may include any other source different from the source of the text features.
[0049] S140, performing feature extraction on the pre-processed passport image based on an optical character recognition function to obtain text features.
[0050] It should be noted that OCR (Optical Character Recognition) is a computer technology that aims to automatically recognize and extract text information from an image. The optical character recognition function F(P) of this embodiment is a general text recognition technology. Without further elaboration, this embodiment uses the optical character recognition function F(P) to recognize non-Chinese characters, letters, pinyin and / or numbers in a pre-processed passport image to obtain text features. Specifically, F(P) can extract pixel values of non-Chinese characters, letters, pinyin and / or numbers in the image as text features, but is not limited to this, and can be any parameter or data that can be characterized. In the case where a passport source includes different features (for example, pattern features, character features and text features), S130 and S140 can extract corresponding features respectively. That is, S130 and S140 are distinguished only by features, not feature sources.
[0051] It should be noted that S130 and S140 can be performed sequentially. However, S130 and S140 can be performed in parallel, and the accompanying drawings are not intended to limit the present application. In this way, different feature extraction processing methods can be selected according to different data volumes, and the efficiency of data processing can be improved.
[0052] S150: construct and train a passport recognition model based on the pattern features, the character features, and the text features to obtain a trained passport recognition model.
[0053] Based on the above embodiment, the obtained pattern features, character features and text features are used to construct a passport recognition model, wherein the expression of the passport recognition model is:
[0054]
[0055] Among them, R(P) represents the recognition result of the passport recognition model, P1 represents the image features of the preprocessed passport image P, and w 0 represents the weight parameter of the kth convolutional layer, C 0 represents the features of the hidden layer of the model after the kth convolutional layer, F(P) represents the optical character recognition function, the output is the text feature, σ represents the nonlinear activation function, T 4 represents the i-th text feature, r 4 Represents the i-th text feature T 4 The weighting coefficient of 1 Represents the i-th text feature T 4 , θ represents the hyperparameter of the passport recognition model, τ represents the regularization factor, and || represents the norm sign.
[0056] It should be noted that, based on the above embodiment, after the passport recognition model is constructed, the passport recognition model can be used to construct a loss function of the passport recognition model, wherein the loss function can be a Lagrangian function, and the loss function can be expressed as follows:
[0057] The loss function is expressed as:
[0058]
[0059] Among them, L represents the value of the loss function, γ represents the Lagrangian constraint factor, and B represents the penalty factor.
[0060] Optionally, in a specific embodiment, partial derivatives of model parameters of the passport recognition model are calculated based on the loss function to determine the partial derivative parameters; optimal model parameters are determined based on the partial derivative parameters, the second formula and the third formula; and a trained passport recognition model is determined based on the optimal model parameters.
[0061] Exemplarily, based on the above embodiment, the partial derivative of the Lagrangian function (loss function) is obtained. Specifically, the partial derivatives of the weight parameters of the kth convolutional layer and the hyperparameters of the passport recognition model are respectively obtained based on the Lagrangian function in a continuous loop to obtain the first partial derivative parameter and the second partial derivative parameter. The first partial derivative parameter and the second partial derivative parameter are then substituted into the second formula to obtain the updated weight parameters of the kth convolutional layer. The second partial derivative parameter is substituted into the third formula to obtain the updated hyperparameters of the passport recognition model, until the updated weight parameters of the kth convolutional layer and the updated hyperparameters are minimized, and the optimal weight parameters and the optimal hyperparameters are obtained. Among them, the optimal model parameters include the optimal weight parameters and the optimal hyperparameters corresponding to the optimal weight parameters. Finally, the optimal weight parameters and the optimal hyperparameters corresponding to the optimal weight parameters are substituted into the passport recognition model to obtain the trained passport recognition model, wherein the expression of the first partial derivative parameter is:
[0062]
[0063] The expression of the second partial derivative parameter is:
[0064]
[0065] Among them, w 0 1 represents the first partial derivative parameter, θ 1 Represents the second partial derivative parameter. The expression of the second formula is:
[0066]
[0067] The expression of the third formula is:
[0068]
[0069] Among them, w 0 (EFG) represents the updated weight parameter, ∈ represents w 0 The scaling factor, θ (EFG) represents the updated hyperparameters, represents the mapping factor of θ.
[0070] In this embodiment, pattern features, character features and text features are used to construct a passport recognition model, and the passport recognition model is used to efficiently and accurately extract key information in the passport. The high accuracy of key information reduces information distortion caused by human errors, ensures data reliability, and improves data processing efficiency.
[0071] S160: Input the new pattern features, character features and text features into the trained passport recognition model to obtain recognition results.
[0072] In summary, after obtaining the above-mentioned trained passport recognition model, the passport is automatically photographed or scanned by a high-resolution camera or scanner and other equipment to obtain a new passport capture image, the new passport capture image is preprocessed to obtain a new preprocessed passport image, and the new preprocessed passport image is feature extracted by using a deep learning model and an optical character recognition function to obtain new pattern features, text features and the text features, and finally the new pattern features, text features and the text features are input into the trained passport recognition model to obtain a recognition result, wherein the recognition result includes but is not limited to: name, gender, photo, date of birth, nationality, passport number, date of issuance, country of issuance and validity period and other corresponding coding feature information. For example, the coding feature information of the name Zhang Wei is: 1110100010111100, the coding feature information of the gender is: 00000001 (male), and the coding feature information of the issuing country is: 00000001 (China). The corresponding accurate name, gender, photo, date of birth, nationality, passport number, date of issue and validity period can be inferred from the recognition result (i.e., the coded feature information). At entry-exit management places such as airports and customs, the name, gender, date of birth, nationality, passport number, date of issue and validity period can be obtained based on the passport recognition model, which can improve the efficiency and accuracy of customs clearance. In combination with video surveillance, passenger behavior and their passport information can be analyzed to identify suspicious activities. For example, the passenger's historical entry-exit records and passport information can be combined to analyze whether the passenger frequently changes his passport. The passport information can be identified during each security check and stored in the database for information comparison to improve travel safety.
[0073] The passport information collection method based on image detection and recognition provided in this embodiment obtains a passport collection image and determines whether the passport collection image is out of focus; when it is determined that the passport collection image is not out of focus, the passport collection image is preprocessed to obtain a preprocessed passport image; based on the deep learning model, the preprocessed passport image is feature extracted to obtain pattern features and text features; based on the optical character recognition function, the preprocessed passport image is feature extracted to obtain text features; based on the pattern features, text features and text features, a passport recognition model is constructed and trained to obtain a trained passport recognition model; the new pattern features, text features and text features are input into the trained passport recognition model to obtain a recognition result. The present disclosure judges the out-of-focus situation of the image to avoid the subsequent processing work of invalid processing, recognition and review of the out-of-focus image, thereby saving computing resources and manpower costs, and ensuring that the quality of the collected image meets the expected standard; secondly, by using the passport recognition model to recognize the passport, the accuracy of the passport collection image under different conditions is improved, manual intervention is reduced, operating costs are reduced, and data processing speed, as well as data reliability and security are improved.
[0074] Figure 2 FIG. 1 is a schematic diagram of the structure of a passport information collection system based on image detection and recognition in an embodiment of the present disclosure. Figure 2 As shown, the system includes an image acquisition unit 210 , a preprocessing unit 220 , a pattern and text feature acquisition unit 230 , a text feature acquisition unit 240 , a model acquisition unit 250 and a result acquisition unit 260 .
[0075] The image acquisition unit 210 is used to acquire a passport captured image and determine whether the passport captured image is out of focus;
[0076] The preprocessing unit 220 is used to preprocess the passport captured image to obtain a preprocessed passport image when it is determined that the passport captured image is not out of focus;
[0077] The pattern and text feature acquisition unit 230 is used to extract features from the pre-processed passport image based on a deep learning model to obtain pattern features and text features, wherein the pattern features correspond to the pattern in the passport captured image and the border of the passport, and the text features correspond to the Chinese characters in the passport captured image;
[0078] The text feature acquisition unit 240 is used to extract features from the pre-processed passport image based on an optical character recognition function to obtain text features, wherein the text features correspond to non-Chinese characters, letters, pinyin and numbers in the passport captured image;
[0079] The model acquisition unit 250 is used to construct and train a passport recognition model based on the pattern features, the character features and the text features to obtain a trained passport recognition model;
[0080] The result acquisition unit 260 is used to input new pattern features, character features and text features into the trained passport recognition model to obtain recognition results.
[0081] It should be noted that, although the accompanying drawings illustrate that the pattern and text feature acquisition unit 230 and the text feature acquisition unit 240 can respectively perform the steps of obtaining pattern features and text features and obtaining text features in sequence, in other embodiments, the pattern and text feature acquisition unit 230 and the text feature acquisition unit 240 can respectively perform the steps of obtaining pattern features and text features and obtaining text features in parallel. That is, the pattern and text feature acquisition unit 230 and the text feature acquisition unit 240 can respectively perform the processing in sequence or in parallel. In this way, different feature extraction processing methods can be selected according to different data amounts, and the efficiency of data processing can be improved.
[0082] Optionally, the image acquisition unit 210 is further configured to determine whether the passport captured image is out of focus based on the passport captured image and a first formula, wherein the expression of the first formula is:
[0083]
[0084] Where C represents the defocus value, A represents the passport image, μ represents the mean value of the passport image A, δ(A) represents the pixel gradient function of the passport image A, and a 1 represents the pixel gradient adjustment factor of the passport image A, a 2 represents the pixel defocus adjustment factor of the passport image A, p j represents the probability of gray level j of passport captured image A.
[0085] Optionally, the expression for preprocessing the passport image is:
[0086]
[0087] Where P represents the pre-processed passport image, b 1 represents the intensity adjustment factor of the passport image A, b 2 represents the intensity shift factor of the passport image A, N represents the total number of pixels of the passport image A, b 3 represents the pixel gradient weight factor, represents the pixel density probability function in the x direction, Represents the pixel density probability function in the y direction.
[0088] Among them, the expression of the passport recognition model is:
[0089]
[0090] Among them, R(P) represents the recognition result of the passport recognition model, P1 represents the image features of the preprocessed passport image P, and w 0 represents the weight parameter of the kth convolutional layer, C 0 represents the features of the hidden layer of the model after the kth convolutional layer, F(P) represents the optical character recognition function, the output is the text feature, σ represents the nonlinear activation function, T 4 represents the i-th text feature, r 4 Represents the i-th text feature T 4 The weighting coefficient of 1 Represents the i-th text feature T 4 , θ represents the hyperparameter of the passport recognition model, τ represents the regularization factor, and || represents the norm sign.
[0091] The model acquisition unit 250 is also used to obtain partial derivatives of the model parameters of the passport recognition model based on the loss function to determine the partial derivative parameters; determine the optimal model parameters based on the partial derivative parameters, the second formula and the third formula; and determine the trained passport recognition model based on the optimal model parameters.
[0092] Wherein, the loss function is expressed as:
[0093]
[0094] Among them, L represents the value of the loss function, γ represents the Lagrangian constraint factor, and B represents the penalty factor.
[0095] Optionally, the partial derivative parameter includes a first partial derivative parameter and a second partial derivative parameter, wherein the expression of the first partial derivative parameter is:
[0096]
[0097] The expression of the second partial derivative parameter is:
[0098]
[0099] Among them, w 0 1 represents the first partial derivative parameter, θ 1 represents the second partial derivative parameter.
[0100] The expression of the second formula is:
[0101]
[0102] The expression of the third formula is:
[0103]
[0104] Among them, w 0 (EFG) represents the updated weight parameter, ∈ represents w 0 The scaling factor, θ (EFG) represents the updated hyperparameters, represents the mapping factor of θ.
[0105] The passport information collection system based on image detection and recognition provided by the embodiment of the present disclosure can execute the steps of the passport information collection method based on image detection and recognition provided by the method embodiment of the present disclosure, and the execution steps and beneficial effects are no longer repeated here.
[0106] It should be noted that this application is not limited to passports, but can be applied to any document with equivalent or similar characteristics, such as identity card, birth certificate, professional practice certificate, graduation certificate, health certificate, business license, qualification certificate, green card, permanent residence permit, visa, residence permit, temporary residence permit, social security, etc.
[0107] Figure 3 Schematic diagram of the structure of an electronic device in the embodiment of the present disclosure. Figure 3 , which shows a structural schematic diagram of an electronic device 500 suitable for implementing the embodiments of the present disclosure. Figure 3 The electronic device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present disclosure.
[0108] like Figure 3 As shown, the electronic device 500 may include a processing device 501, a ROM 502, a RAM 503, a bus 504, an input / output (I / O) interface 505, an input device 506, an output device 507, a storage device 508, and a communication device 509. The processing device (e.g., a central processing unit, a graphics processor, etc.) 501 can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 502 or the program loaded from the storage device 508 to the random access memory (RAM) 503 to implement the passport information collection method based on image detection and recognition as described in the present disclosure. In the RAM 503, various programs and data required for the operation of the electronic device 500 are also stored. The processing device 501, the ROM 502, and the RAM 503 are connected to each other through the bus 504. The input / output (I / O) interface 505 is also connected to the bus 504.
[0109] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program contains a program code for executing the method shown in the flowchart, thereby implementing the passport information collection method based on image detection and recognition as described above. In such an embodiment, the computer program can be downloaded and installed from the network through the communication device 509, or installed from the storage device 508, or installed from the ROM502. When the computer program is executed by the processing device 501, the above-mentioned functions defined in the method of the embodiment of the present disclosure are executed.
[0110] It should be noted that the computer-readable medium disclosed above may be a computer-readable signal medium or a computer-readable storage medium or any combination of the above two. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, a computer-readable storage medium may be any tangible medium containing or storing a program that may be used by or in combination with an instruction execution system, device or device. In the present disclosure, a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, in which a computer-readable program code is carried. This propagated data signal may take a variety of forms, including but not limited to an electromagnetic signal, an optical signal, or any suitable combination of the above. The computer readable signal medium may also be any computer readable medium other than a computer readable storage medium, which may send, propagate or transmit a program for use by or in conjunction with an instruction execution system, apparatus or device. The program code contained on the computer readable medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.
[0111] The computer-readable medium may be included in the electronic device; or it may exist independently without being assembled into the electronic device. The computer-readable medium carries one or more programs. When the one or more programs are executed by the electronic device, the electronic device: obtains a passport-captured image and determines whether the passport-captured image is out of focus; when it is determined that the passport-captured image is not out of focus, pre-processes the passport-captured image to obtain a pre-processed passport image; extracts features from the pre-processed passport image based on a deep learning model to obtain pattern features and text features; extracts features from the pre-processed passport image based on an optical character recognition function to obtain text features; constructs and trains a passport recognition model based on pattern features, text features and text features to obtain a trained passport recognition model; inputs new pattern features, text features and text features into the trained passport recognition model to obtain a recognition result.
[0112] Optionally, when the above one or more programs are executed by the electronic device, the electronic device may also execute other steps described in the above embodiments.
[0113] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or equipment. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0114] The above description is only a preferred embodiment of the present disclosure and an explanation of the technical principles used. Those skilled in the art should understand that the scope of disclosure involved in the present disclosure is not limited to the technical solutions formed by a specific combination of the above technical features, but should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above disclosed concept. For example, the above features are replaced with the technical features with similar functions disclosed in the present disclosure (but not limited to) by each other to form a technical solution.
Claims
1. A passport information collection method based on image detection and recognition, characterized in that: The passport information collection method based on image detection and recognition includes: Acquire a passport captured image, and determine whether the passport captured image is out of focus; When it is determined that the passport captured image is not out of focus, preprocessing the passport captured image to obtain a preprocessed passport image; Performing feature extraction on the pre-processed passport image based on a deep learning model to obtain pattern features and text features, wherein the pattern features correspond to the pattern in the passport captured image and the border of the passport, and the text features correspond to the Chinese characters in the passport captured image; Extracting features of the pre-processed passport image based on an optical character recognition function to obtain text features, wherein the text features correspond to non-Chinese characters, letters, pinyin and numbers in the captured passport image; Building and training a passport recognition model based on the pattern features, the character features and the text features to obtain a trained passport recognition model; Input the new pattern features, character features and text features into the trained passport recognition model to obtain the recognition results. Wherein, determining whether the passport captured image is out of focus includes: Determine whether the passport captured image is out of focus based on the passport captured image and a first formula, wherein the expression of the first formula is: Where C represents the defocus value, A represents the passport image, μ represents the mean value of the passport image A, δ(A) represents the pixel gradient function of the passport image A, a1 represents the pixel gradient adjustment factor of the passport image A, a2 represents the pixel defocus adjustment factor of the passport image A, p j represents the probability of gray level j of passport captured image A, Among them, the expression of the pre-processed passport image is: Where P represents the preprocessed passport image, b1 represents the intensity adjustment factor of the passport image A, b2 represents the intensity offset factor of the passport image A, N represents the total number of pixels in the passport image A, and b3 represents the pixel gradient weight factor. represents the pixel density probability function in the x direction, Represents the pixel density probability function in the y direction.
2. The passport information collection method based on image detection and recognition according to claim 1 is characterized in that: The expression of the passport recognition model is: Among them, R(P) represents the recognition result of the passport recognition model, P1 represents the image features of the preprocessed passport image P, and w k represents the weight parameter of the kth convolutional layer, C k represents the features of the hidden layer of the model after the kth convolutional layer, F(P) represents the optical character recognition function, the output is the text feature, σ represents the nonlinear activation function, T i represents the i-th text feature, r i Represents the i-th text feature T i The weighted coefficient of , d1 represents the i-th text feature T i , θ represents the hyperparameter of the passport recognition model, τ represents the regularization factor, and || represents the norm sign.
3. The passport information collection method based on image detection and recognition according to claim 2 is characterized in that: Train the passport recognition model and obtain the trained passport recognition model, including: Taking partial derivatives of model parameters of the passport recognition model based on the loss function to determine partial derivative parameters; Determine optimal model parameters based on the partial derivative parameters, the second formula and the third formula; Based on the optimal model parameters, a trained passport recognition model is determined.
4. The passport information collection method based on image detection and recognition according to claim 3 is characterized in that: The expression of the loss function is: Among them, L represents the value of the loss function, γ represents the Lagrangian constraint factor, and B represents the penalty factor.
5. The passport information collection method based on image detection and recognition according to claim 4 is characterized in that: The partial derivative parameter includes a first partial derivative parameter and a second partial derivative parameter, wherein the expression of the first partial derivative parameter is: The expression of the second partial derivative parameter is: Among them, w k 1 represents the first partial derivative parameter, θ 1 represents the second partial derivative parameter.
6. The passport information collection method based on image detection and recognition according to claim 5 is characterized in that: The expression of the second formula is: The expression of the third formula is: Among them, w k (new) represents the updated weight parameter, ∈ represents w k The scaling factor, θ (new) represents the updated hyperparameters, represents the mapping factor of θ.
7. The passport information collection system based on image detection and recognition is characterized by: The passport information collection system based on image detection and recognition includes: An image acquisition unit, used to acquire a passport captured image and determine whether the passport captured image is out of focus; A preprocessing unit, configured to preprocess the passport captured image to obtain a preprocessed passport image when determining that the passport captured image is out of focus; A pattern and text feature acquisition unit, used for performing feature extraction on the pre-processed passport image based on a deep learning model to obtain pattern features and text features, wherein the pattern features correspond to the pattern in the passport captured image and the border of the passport, and the text features correspond to the Chinese characters in the passport captured image; A text feature acquisition unit, configured to extract features from the pre-processed passport image based on an optical character recognition function to obtain text features, wherein the text features correspond to non-Chinese characters, letters, pinyin and numbers in the captured passport image; A model acquisition unit, used to construct and train a passport recognition model based on the pattern feature, the character feature and the text feature to obtain a trained passport recognition model; The result acquisition unit is used to input new pattern features, character features and text features into the trained passport recognition model to obtain recognition results. Wherein, determining whether the passport captured image is out of focus includes: Determine whether the passport captured image is out of focus based on the passport captured image and a first formula, wherein the expression of the first formula is: Where C represents the defocus value, A represents the passport image, μ represents the mean value of the passport image A, δ(A) represents the pixel gradient function of the passport image A, a1 represents the pixel gradient adjustment factor of the passport image A, a2 represents the pixel defocus adjustment factor of the passport image A, p j represents the probability of gray level j of passport captured image A, Among them, the expression of the pre-processed passport image is: Where P represents the preprocessed passport image, b1 represents the intensity adjustment factor of the passport image A, b2 represents the intensity offset factor of the passport image A, N represents the total number of pixels in the passport image A, and b3 represents the pixel gradient weight factor. represents the pixel density probability function in the x direction, Represents the pixel density probability function in the y direction.
Citation Information
Patent Citations
Text recognition method and device, electronic equipment and medium
CN114821596A
Text recognition method and device, storage medium and electronic equipment
CN116524520A