A method, an electronic device, and a storage medium for intelligent information verification of vehicle inspection forms

Through text detection and recognition model combined with VGG16 network optimization character recognition, the problem of misrecognition of near-words in vehicle inspection is solved, efficient and accurate verification of vehicle inspection form information is achieved, and vehicle inspection efficiency and safety are improved.

CN114359937BActive Publication Date: 2025-07-08DUOLUN TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210006383.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-01-05
Publication Date
2025-07-08
Estimated Expiration
2042-01-05

AI Technical Summary

Technical Problem

Traditional optical character recognition methods have a high misrecognition rate of close-to-word characters during vehicle inspection, resulting in an increase in the misrecognition rate, increasing the burden on vehicle inspection staff and the risk of fraud.

Method used

The text detection model and text recognition model are combined with the VGG16 network, and through training and preprocessing steps, accurate character recognition of vehicle inspection form images is achieved, and the cross entropy loss function and margin penalty function are used to optimize the model to reduce the misrecognition rate of similar words.

Benefits of technology

It improves the accuracy of vehicle inspection form information verification, reduces the rate of misjudgment, reduces the burden on staff, reduces the possibility of manual information fraud, and improves vehicle inspection efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114359937B_ABST
    Figure CN114359937B_ABST
Patent Text Reader

Abstract

The present invention discloses a method, an electronic device and a storage medium for intelligent information verification of vehicle inspection forms, including: obtaining a preset number of sample images of each vehicle inspection form, and a preset correspondence relationship of each pixel in each vehicle inspection form sample image with respect to a character label or a non-character label; obtaining a preset number of single-character images each containing a single character, and the character content corresponding to each single-character image; then obtaining a text detection model and a text recognition model through step A and step B respectively; and realizing the recognition of each string content in the target vehicle inspection form image based on the text detection model and the text recognition model. By using the verification method of the present invention to perform character recognition on form data such as vehicle inspection form reports, and using the recognition results for information verification and security assessment, it can not only speed up the vehicle inspection speed, reduce the burden on staff, reduce the misjudgment rate of similar-shaped characters in text recognition, but also reduce the possibility of fraud in manual information verification.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method, an electronic device and a storage medium for intelligent information verification of vehicle inspection forms, belonging to the technical field of document recognition. Background Art

[0002] With the continuous improvement of living standards, people's demand for safe transportation is increasing day by day. Vehicle annual inspection plays a role in ensuring the safety of vehicles, thus becoming an important part of safe transportation. The traditional optical character recognition method is to obtain an electronic document of a paper document through an electronic device, such as a scanner or a digital camera, split the character strings in the electronic document to form small pictures containing single characters, and then use a certain method to recognize the split characters. However, misrecognition of similar characters often occurs, resulting in a certain misjudgment rate, greatly affecting the efficiency of the vehicle inspection process. After a document misrecognition event occurs, the processes involving paper information verification in vehicle inspection, such as certificate verification and vehicle information verification, greatly increase the burden on vehicle inspection staff, and there is also a possibility of fraud and misjudgment in manually verifying these data. Summary of the Invention

[0003] The purpose of the present invention is to provide a method, an electronic device and a storage medium for intelligent information verification of vehicle inspection forms to solve the problems in the prior art.

[0004] To achieve the above purpose, the present invention provides the following technical solutions: A method for intelligent information verification of vehicle inspection forms, the steps are as follows: Obtain a preset number of sample images of each vehicle inspection form, and a preset corresponding relationship between each pixel in each vehicle inspection form sample image and a character label or a non-character label; and obtain a preset number of single-character images each containing a single character, and the character content corresponding to each single-character image; then obtain a text detection model and a text recognition model respectively through the following steps A and B; and then based on the obtained text detection model and text recognition model, execute steps I to IV to realize the recognition of each character string content in the target vehicle inspection form image;

[0005] Step A: Based on the sample images of a preset number of vehicle inspection forms as training data, and the preset corresponding relationship between each pixel in each vehicle inspection form sample image and a character label or a non-character label, use the vehicle inspection form sample image as the input and the character label or non-character label corresponding to each pixel in the vehicle inspection form sample image as the output to train the text detection model to be trained, and obtain the text detection model;

[0006] Step B: Based on a preset number of single-character images each containing a single character as training data, use the single-character image containing a single character as the input and the character content in the single-character image as the output to train the text recognition model to be trained, and obtain the text recognition model;

[0007] Step I: Process the target vehicle inspection form image through a text detection model to obtain the character labels or non-character labels corresponding to each pixel in the target vehicle inspection form image, and enter Step II; Step II: According to the character labels or non-character labels corresponding to each pixel in the target vehicle inspection form image, and in combination with the maximum distance between adjacent pixels of the preset string, divide each pixel to obtain each string image in the target vehicle inspection form image, and enter Step III; Step III: Split each string image in the target vehicle inspection form image respectively to obtain each single-character image in the string image, and further obtain each single-character image in each string image in the target vehicle inspection form image, and enter Step IV; Step IV: Apply a text recognition model to each single-character image in each string image in the target vehicle inspection form image respectively to obtain the character content in the single-character image, and further obtain the character content in each single-character image in the target vehicle inspection form image, and enter Step V; Step V: According to the character content in each single-character image in the target vehicle inspection form image, and in combination with the sorting of each single-character image in each string image in the target vehicle inspection form image, obtain the character content corresponding to each string image, that is, realize the recognition of the target vehicle inspection form image;

[0008] In the above Step B, the text recognition model uses the VGG16 network. After obtaining the output matrix of the VGG16 network, the following steps are performed on the output matrix to form the text recognition model;

[0009] Step B1: Normalize the output matrix of the VGG16 network according to the second dimension, and calculate the product of W and the output matrix to obtain the predicted vector value; where the parameter W is a matrix of 1024x num_classes, and num_classes is the number of Chinese character categories;

[0010] Step B2: Select the value corresponding to the label from the predicted vector value, and calculate its arccosine to obtain the angle θ;

[0011] Step B3: Assign cos(θ + m) to the one-hot code corresponding to the label, where m is the angular margin penalty, and m > 0;

[0012] Step B4: Multiply the predicted vector value by a fixed value s, then use the Softmax function for normalization, and finally calculate the loss using cross-entropy. The cross-entropy loss function used is:

[0013]

[0014] where LOSS is the cross-entropy loss function, N is the batch value of the training data; s is a hyperparameter of the model used to amplify the cosine value, and is set to 30; is the included angle between the i-th column vector of the parameter W matrix and the i-th class feature vector of the last layer of the output of the VGG16 network, θ j,i is the included angle between the j-th column vector and the i-th column vector of the parameter W matrix, and m is the angular margin penalty, which is used to forcibly increase the angle between the same classes.

[0015] Further, in the aforementioned step A, image preprocessing is performed on each vehicle inspection form sample image of a preset quantity, including performing median filtering on the image and adjusting the picture contrast and brightness to preset values.

[0016] Further, the sampling multiple of the aforementioned VGG16 network is 32, and the picture height is set to be scaled proportionally to 32.

[0017] Further, in the aforementioned step III, perspective transformation is performed on each single-character image in the string image to obtain a front-facing single-character picture, and then the text recognition model to be trained is trained.

[0018] Further, in the aforementioned step B, based on the regular expressions corresponding to the preset character attribute features of the single-character image, with the single-character image containing a single character as the input and the character content in the single-character image as the output, the text recognition model to be trained is trained to obtain a text recognition model.

[0019] On the other hand, the present invention provides an electronic device, including a storage device and one or more processors. The storage device is used to store one or more programs. When the one or more programs are executed by the one or more processors, the one or more processors implement the method for intelligent information verification of vehicle inspection forms described in any one of the present invention.

[0020] The present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the method for intelligent information verification of vehicle inspection forms described in any one of the present invention is implemented. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] Figure 1 is a flowchart of an exemplary method for intelligent information verification of vehicle inspection forms of the present invention;

[0022] Figure 2 is a schematic diagram of the character classification decision boundary after using the Softmax function in step B4 of the present invention;

[0023] Figure 3 is a schematic diagram of the character classification decision boundary after adding the margin penalty function in step B3 of the present invention;

[0024] Figure 4Schematic diagram of the 2D feature extraction of 10 Chinese characters by the text recognition model of the present invention using the Softmax function;

[0025] Figure 5 Schematic diagram of the 2D feature extraction of 10 Chinese characters by the text recognition model of the present invention after adding a margin penalty function. Detailed implementation manner

[0026] In the present invention, various aspects of the present invention are described with reference to the accompanying drawings, in which illustrative embodiments are shown. The embodiments of the present invention are not limited to those shown in the drawings. It should be understood that the present invention is implemented through the concepts and embodiments introduced above, as well as the concepts and implementation manners described in detail below, because the concepts and embodiments disclosed in the present invention are not limited to any implementation manner. In addition, some aspects disclosed in the present invention can be used alone, or in any suitable combination with other aspects disclosed in the present invention.

[0027] As Figure 1 shown, a method for intelligent information verification of vehicle inspection forms of the present invention is as follows:

[0028] First, obtain a preset number of sample images of each vehicle inspection form, as well as the preset correspondence relationship of each pixel in each vehicle inspection form sample image with respect to character labels or non-character labels; and obtain a preset number of single-character images each containing a single character, as well as the character content corresponding to each single-character image. Then, a text detection model and a text recognition model are obtained through the following steps A and B respectively.

[0029] Step A: Using the sample images of a preset number of each vehicle inspection form as training data, and the preset correspondence relationship of each pixel in each vehicle inspection form sample image with respect to character labels or non-character labels, taking the vehicle inspection form sample image as the input and the character label or non-character label corresponding to each pixel in the vehicle inspection form sample image as the output, train the text detection model to be trained to obtain the text detection model;

[0030] Step B: Using a preset number of single-character images each containing a single character as training data, taking the single-character image containing a single character as the input and the character content in the single-character image as the output, train the text recognition model to be trained. In step B, the text recognition model uses the VGG16 network. After obtaining the output matrix of the VGG16 network, the following steps are performed on the output matrix, including:

[0031] Step B1: Normalize the output matrix of the VGG16 network in the second dimension, and calculate the product of W and the output matrix to obtain the predicted vector value; where the parameter W is a matrix of 1024x num_classes, and num_classes is the number of categories of Chinese characters;

[0032] Step B2: Select the value corresponding to the label from the predicted vector values, calculate its arccosine to obtain the angle θ;

[0033] Step B3: Assign cos(θ + m) to the one-hot code corresponding to the label, where m is the angular margin penalty and m > 0;

[0034] Step B4: Multiply the predicted vector values by a fixed value s, then use the Softmax function for normalization, and finally calculate the loss using cross-entropy. The corresponding cross-entropy loss function is:

[0035]

[0036] where LOSS is the cross-entropy loss function, N is the batch value of the training data; s is a hyperparameter of the model used to amplify the cosine value, set to 30; is the angle between the i-th column vector of the parameter W matrix and the i-th class feature vector of the last layer of the output of the VGG16 network, θ j,i is the angle between the j-th column vector and the i-th column vector of the parameter W matrix, and m is the angular margin penalty used to forcibly increase the angle between the same classes.

[0037] Due to the addition of the margin penalty function, the angle between the same classes can be forcibly increased, making the neural network work harder to tighten the same classes during training and minimizing the within-class scatter. As Figure 2 shown, after using the Softmax function in Step B4, the decision boundary of character classification has a large scatter. As Figure 3 shown, the cross-entropy loss function used in Step B3 adds a margin penalty function. From the figure, it can be seen that the same classes will shrink tightly, and the decision boundaries between different classes are separated by a certain distance, that is, the between-class scatter of different classes increases.

[0038] After the above steps, a text recognition model is obtained; then, based on the obtained text detection model and text recognition model, the text detection model and text recognition model are applied. First, preprocess each vehicle inspection form sample image, including performing median filtering on the image and adjusting the image contrast and brightness to preset values. Execute Step I to Step IV to realize the recognition of each string content in the target vehicle inspection form image.

[0039] Step I: Apply the text detection model to process the target vehicle inspection form image to obtain the character labels or non-character labels corresponding to each pixel in the target vehicle inspection form image, and then enter Step II;

[0040] Step II: According to whether each pixel in the target vehicle inspection form image corresponds to a character label or a non-character label, and in combination with the maximum distance between adjacent pixels of the preset string, each pixel is divided to obtain each string image in the target vehicle inspection form image, and then proceed to Step III;

[0041] Step III: For each string image in the target vehicle inspection form image respectively, the string image is segmented to obtain each single-character image in the string image, and further obtain each single-character image in each string image in the target vehicle inspection form image. The perspective transformation is performed on each single-character image in the string image to obtain a front-facing single-character picture. Then, the text recognition model to be trained is trained, and then proceed to Step IV;

[0042] Step IV: For each single-character image in each string image in the target vehicle inspection form image respectively, based on the regular expression corresponding to each character attribute feature of the single-character image, the text recognition model is applied to obtain the character content in the single-character image, and further obtain the character content in each single-character image in the target vehicle inspection form image. Then, proceed to Step V;

[0043] Step V: According to the character content in each single-character image in the target vehicle inspection form image, and in combination with the sorting of each single-character image in each string image in the target vehicle inspection form image, obtain the character content corresponding to each string image respectively, that is, realize the recognition of the target vehicle inspection form image.

[0044] As Figure 4 shown, after the text recognition model uses the Softmax function to extract 10 Chinese character features and reduce the features to 2 dimensions, the neural network will also extract similar features for similar characters. If the inter-class dispersion is too small and the intra-class dispersion is too large, there will be a possibility of misidentifying similar characters. As Figure 5 shown, after the cross-entropy loss function used by the text recognition model adds the margin penalty function to extract and 2-dimension the 10 Chinese character features, the inter-class dispersion of different characters increases, and the intra-class dispersion decreases. This makes the decision boundary between similar characters increase, and thus it is difficult to misidentify.

[0045] In the second aspect of the present invention, an electronic device is proposed, including a storage device and one or more processors. The storage device is used to store one or more programs. When the one or more programs are executed by the one or more processors, the one or more processors are enabled to implement the method for intelligent information verification of vehicle inspection forms.

[0046] In the third aspect of the present invention, a computer-readable storage medium is proposed, on which a computer program is stored. When the computer program is executed by a processor, the method for intelligent information verification of vehicle inspection forms is implemented.

[0047] The method for intelligent information verification of vehicle inspection forms according to the present invention, compared with the prior art by adopting the above technical solutions, has the following technical effects: The verification method of the present invention is used to perform character recognition on form data such as motor vehicle safety technical inspection reports, and the recognition results are used for information verification and safety assessment, greatly reducing the misjudgment rate of similar characters in the text detection and text recognition processes. It can not only reduce the burden on staff, speed up the vehicle inspection process, but also reduce the possibility of fraud in manual information verification.

[0048] Although the present invention has been disclosed above with preferred embodiments, it is not intended to limit the present invention. Those with ordinary knowledge in the technical field to which the present invention pertains can make various modifications and refinements without departing from the spirit and scope of the present invention. Therefore, the protection scope of the present invention shall be subject to that defined by the claims.

Claims

1. A method for intelligent information verification of vehicle inspection forms, characterized in that, The steps are as follows: Obtain a preset number of sample images of each vehicle inspection form, and the preset corresponding relationships of each pixel in each vehicle inspection form sample image with respect to character labels or non-character labels; and obtain a preset number of single-character images each containing a single character, and the character content corresponding to each single-character image respectively; then obtain a text detection model and a text recognition model through the following Step A and Step B respectively; and then, based on the obtained text detection model and text recognition model, execute Step I to Step IV to realize the recognition of each string content in the target vehicle inspection form image. Step A: Using the sample images of a preset number of vehicle inspection forms as training data, and the preset corresponding relationships of each pixel in each vehicle inspection form sample image with respect to character labels or non-character labels, taking the vehicle inspection form sample image as the input and the character label or non-character label corresponding to each pixel in the vehicle inspection form sample image as the output, train the text detection model to be trained to obtain the text detection model. Step B: Using a preset number of single-character images each containing a single character as training data, taking the single-character image containing a single character as the input and the character content in the single-character image as the output, train the text recognition model to be trained to obtain the text recognition model. Step I: Process the target vehicle inspection form image through the text detection model to obtain the character label or non-character label corresponding to each pixel in the target vehicle inspection form image, and enter Step II; Step II: According to the character label or non-character label corresponding to each pixel in the target vehicle inspection form image, combined with the maximum distance between adjacent pixels of the preset string, divide each pixel to obtain each string image in the target vehicle inspection form image, and enter Step III; Step III: Split each string image in the target vehicle inspection form image respectively to obtain each single-character image in the string image, and further obtain each single-character image in each string image in the target vehicle inspection form image, and enter Step IV; Step IV: Apply the text recognition model to each single-character image in each string image in the target vehicle inspection form image respectively to obtain the character content in the single-character image, and further obtain the character content in each single-character image in the target vehicle inspection form image, and enter Step V. Step V: According to the character content in each single-character image in the target vehicle inspection form image, combined with the sorting of each single-character image in each string image in the target vehicle inspection form image, obtain the character content corresponding to each string image respectively, that is, realize the recognition of the target vehicle inspection form image. In the above Step B, the text recognition model uses the VGG16 network. After obtaining the output matrix of the VGG16 network, perform the following steps on the output matrix to form the text recognition model. Step B1: Normalize the output matrix of the VGG16 network in the second dimension, and calculate the product of W and the output matrix to obtain the predicted vector value; where the parameter W is a matrix of 1024 x num_classes, and num_classes is the number of categories of Chinese characters. Step B2: Select the value corresponding to the label from the predicted vector values, and calculate its arccosine to obtain the angle θ; Step B3: Assign cos(θ + m) to the one-hot code corresponding to the label, where m is the angular margin penalty and m > 0; Step B4: After multiplying the predicted vector values by a fixed value s, use the Softmax function for normalization, and finally calculate the loss using cross-entropy. The cross-entropy loss function used is: Among them, Loss is the cross-entropy loss function, N is the batch value of the training data; s is a hyperparameter of the model, used to amplify the cosine value, and is set to 30; is the angle between the i-th column vector of the parameter W matrix and the i-th class feature vector of the last layer output by the VGG16 network, θ j,i is the angle between the j-th column vector of the parameter W matrix and the i-th column vector of the parameter W matrix, and m is the angular margin penalty, which is used to forcibly widen the angles between the same classes.

2. The method for intelligent information verification of a vehicle inspection form according to claim 1, characterized in that In the said Step A, image preprocessing is performed on each vehicle inspection form sample image of a preset quantity, including performing median filtering on the image and adjusting the contrast and brightness of the image to preset values.

3. A method for intelligent information verification of vehicle inspection forms according to claim 1, characterized in that, The sampling multiple of the VGG16 network is 32, and the image height is set to be scaled proportionally to 32.

4. A method for intelligent information verification of vehicle inspection forms according to claim 1, characterized in that In the said Step III, perspective transformation is performed on each single-character image in the string image to obtain a front-facing single-character picture, and then the text recognition model to be trained is trained.

5. A method for intelligent information verification of vehicle inspection forms according to claim 1, characterized in that, In the said Step B, based on the regular expressions corresponding to the preset character attribute features of the single-character image, with the single-character image containing a single character as the input and the character content in the single-character image as the output, the text recognition model to be trained is trained to obtain a text recognition model.

6. An electronic device, characterized in that: It includes a storage device and one or more processors. The storage device is used to store one or more programs. When the one or more programs are executed by the one or more processors, the one or more processors implement the method for intelligent information verification of vehicle inspection forms as described in any one of claims 1 - 5.

7. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by the processor, it implements the method for intelligent information verification of vehicle inspection forms as described in any one of claims 1 - 5.

Citation Information

Patent Citations

  • Image processing method and device, computer equipment and storage medium

    CN112966583A