Image recognition method and device, computer equipment and storage medium

By using pre-trained image classification model and target recognition model, the images are automatically classified and recognized, and the problems of inefficient image recognition and insufficient accuracy in the prior art are solved, and efficient and accurate image recognition is achieved.

CN120020900APending Publication Date: 2025-05-20PUTIAN EASTERN COMM GRP CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202311550588.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-20
Publication Date
2025-05-20

Smart Images

  • Figure CN120020900A_ABST
    Figure CN120020900A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of image recognition, and discloses an image recognition method and device, computer equipment and a storage medium, and the method comprises the steps: obtaining a to-be-recognized image; classifying the to-be-identified image by using a pre-trained image classification model to obtain a classification result of the to-be-identified image; selecting a target recognition model corresponding to the classification result from a target recognition model set; performing to-be-recognized region positioning on the to-be-recognized image by using the target recognition model, and performing text information recognition on the to-be-recognized region to obtain a recognition result of the to-be-recognized region; and summarizing the recognition results of all the to-be-recognized areas, and determining an image recognition result of the to-be-recognized image. The to-be-recognized image can be automatically classified, manual classification is avoided, and the efficiency and accuracy of image recognition are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of image recognition technology, and particularly to an image recognition method, apparatus, computer device, and storage medium. Background Art

[0002] Although the coverage of electronic payment is becoming wider and wider, considering factors such as the independence of the financial system and the network security risks of enterprises, traditional bills and cash still occupy a place in the mainstream payment methods. This means that traditional bills and currencies are the key objects for financial personnel to sort out.

[0003] With the continuous changes in the business management requirements in the financial industry, it is required that financial personnel digitize and sort out financial materials such as bills and cash. In the actual business process, it often involves classifying and recognizing images of financial materials such as bills uploaded by users. However, considering the characteristics of text adhesion, density, or occlusion in the images, the method of manual classification and then filing and storing is often used, which is inefficient and prone to errors.

[0004] Therefore, how to improve the efficiency and accuracy of image recognition has become a technical problem that needs to be solved urgently at present. Summary of the Invention

[0005] In view of this, the present application provides an image recognition method, apparatus, computer device, and storage medium to solve the problem of how to improve the efficiency and accuracy of image recognition.

[0006] In a first aspect, the present application provides an image recognition method, which includes:

[0007] Obtain an image to be recognized;

[0008] Use a pre-trained image classification model to classify the image to be recognized, and obtain a classification result of the image to be recognized;

[0009] Select a target recognition model corresponding to the classification result from a set of target recognition models, where the set of target recognition models is pre-configured with target recognition models matching different classification results;

[0010] Use the target recognition model to locate the area to be recognized in the image to be recognized, and perform text information recognition on the area to be recognized, to obtain a recognition result of the area to be recognized;

[0011] Summarize the recognition results of all areas to be recognized, and determine an image recognition result of the image to be recognized.

[0012] In the above technical solution, the image to be recognized can be automatically classified, avoiding manual classification and improving the efficiency of image recognition. It is also possible to select a target recognition model corresponding to the classification result of the image to be recognized, perform localization and information recognition of the area to be recognized for images to be recognized of different categories, and then summarize the recognition results of all areas to be recognized of the image to be recognized to determine the image recognition result of the image to be recognized, realizing personalized information recognition of different categories of images, improving the accuracy of image recognition, and also avoiding manual recognition of images and improving the efficiency of image recognition.

[0013] In some alternative embodiments, after performing text information recognition on the area to be recognized, the method further includes:

[0014] Using a pre-trained character recognition model to recognize target characters in the area to be recognized, where the target characters include magnetic numbers and currency symbols.

[0015] Specifically, using the character recognition model to recognize target characters in the area to be recognized and comparing with the recognition result of the target recognition model for supplementation and correction, improving the recognition accuracy of the bill image.

[0016] In some alternative embodiments, after summarizing the recognition results of all areas to be recognized to determine the image recognition result of the image to be recognized, the method further includes:

[0017] Obtaining correction information corresponding to the image to be recognized;

[0018] Correcting the image recognition result according to the correction information to obtain a corrected image recognition result.

[0019] Specifically, using the correction information to correct the image recognition result further improves the accuracy of image recognition

[0020] In some alternative embodiments, obtaining correction information corresponding to the image to be recognized includes:

[0021] Searching for an information code positioning identifier in the image to be recognized to obtain an information code area;

[0022] Performing data conversion on the image in the information code area to obtain information code area data;

[0023] Determining the information code area data in the target format as the correction information corresponding to the image to be recognized.

[0024] Specifically, positioning the information code area based on the positioning identifier to ensure that valid real information is recognized and improve the accuracy of the correction information.

[0025] In some alternative embodiments, before selecting a pre-trained object recognition model corresponding to the classification result from the set of target recognition models, the method further includes:

[0026] Obtain sample images of different materials to obtain a set of sample images, and the sample images are labeled with the category to which the text information in the area to be recognized belongs, the category to which the area to be recognized belongs, and the position information of the area to be recognized;

[0027] According to the category to which the text information in the area to be recognized belongs, the category to which the area to be recognized belongs, and the position information of the area to be recognized in multiple sample images in the set of sample images, iteratively train at least one pre-constructed image recognition model until the loss function of the image recognition model meets the training stop condition, and the trained image recognition model is a target recognition model in the set of target recognition models;

[0028] The loss function is a function constructed based on the center point positioning error of the detection target, the width and height positioning error of the detection target, the confidence error of the detection target, the confidence error of the non-detection target, the classification error of the detection target, and a preset position loss weight;

[0029] In the loss function, the weight of the confidence error of the non-detection target is less than the weight of the confidence error of the detection target. In the loss function, the preset position loss weight is greater than the weights of the confidence error and the classification error in the target loss function. In the width and height positioning error of the detection target, the width is the square root of the width of the detection target, and the height is the square root of the height of the detection target.

[0030] Specifically, when training the image recognition model, make the weight of the position error in the loss function greater than the weights of the confidence error and the classification error, so as to ensure that the model pays more attention to the position of the area to be recognized in the image and ensure the accuracy of the positioning of the area to be recognized. And the confidence error of the detection target is greater than the confidence error of the non-detection target, which is beneficial to suppressing the confidence loss of the non-detection target, facilitating the convergence of the image recognition model, and accelerating the training speed of the image recognition model. Using the square root of the width and height of the detection target instead of the original width and height in the width and height positioning error of the detection target can reduce the width difference between detection targets of different sizes, so that the error accuracy of detection targets of different sizes in the target loss matches the size of the detection target, thereby improving the accuracy of the trained target recognition model and the accuracy of image recognition.

[0031] In some alternative embodiments, the image recognition model includes an input layer, a neural network layer, and an output layer. The neural network layer includes a convolutional layer and a pooling layer. The convolutional layer includes multiple convolutional kernels with a scale of 3×3, and convolutional kernels with a scale of 1×1 interspersed between the 3×3 convolutional kernels; the set of sample images includes a training sample set and a test sample set;

[0032] The input layer is used to transform the size of the sample images in the training sample set into a preset size, and input the sample images with the transformed size into the neural network layer;

[0033] The neural network layer is used to utilize the convolutional layer and the pooling layer to learn the category to which the text information in the area to be recognized in the sample image with the transformed size belongs, the category to which the area to be recognized belongs, and the position information of the area to be recognized, so as to train the parameters in the image recognition model, and perform normalization and Dropout operations on the data processed by the convolutional layer and the pooling layer;

[0034] The output layer is used to determine the value of the loss function according to the test sample set.

[0035] In some alternative embodiments, selecting the target recognition model corresponding to the classification result from the pre-trained target recognition model set includes:

[0036] If the classification result is a bill image, select the target recognition models for invoice code recognition, invoice number recognition, invoice date recognition, check code recognition, and amount recognition from the target recognition model set respectively;

[0037] If the classification result is a currency image, select the target recognition models for face value recognition, portrait recognition, watermark recognition, pattern recognition, and serial number recognition from the target recognition model set respectively.

[0038] Specifically, select the target recognition models for different recognition tasks according to different classification results, so as to perform the recognition of different contents on the image to be recognized in parallel, avoid reducing the recognition error rate by recognizing multiple contents at one time, and improve the accuracy and efficiency of image recognition.

[0039] In a second aspect, the present application provides an image recognition device, and the device includes:

[0040] The first acquisition module is used to acquire the image to be recognized;

[0041] The classification module is used to classify the image to be recognized by using the pre-trained image classification model to obtain the classification result of the image to be recognized;

[0042] The selection module is used to select the target recognition model corresponding to the classification result from the target recognition model set, wherein the target recognition model set is pre-configured with target recognition models matching different classification results;

[0043] The positioning and recognition module is used to use the target recognition model to perform positioning on the area to be recognized in the image to be recognized, and perform text information recognition on the area to be recognized to obtain the recognition result of the area to be recognized;

[0044] A summarization module for summarizing the recognition results of all regions to be recognized and determining the image recognition result of the image to be recognized.

[0045] In a third aspect, the present application provides a computer device, including: a memory and a processor, which are communicatively connected to each other. The memory stores computer instructions, and the processor executes the computer instructions to perform any of the image recognition methods in the first aspect above.

[0046] In a fourth aspect, the present application provides a computer-readable storage medium, on which computer instructions are stored, and the computer instructions are used to cause a computer to perform any of the image recognition methods in the first aspect above. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] In order to more clearly illustrate the specific embodiments of the present application or the technical solutions in the prior art, the following will briefly introduce the drawings required for use in the description of the specific embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0048] Figure 1 is a schematic structural diagram of an image recognition system according to an embodiment of the present application;

[0049] Figure 2 is a schematic flowchart of an image recognition method according to an embodiment of the present application;

[0050] Figure 3 is a schematic flowchart of another image recognition method according to an embodiment of the present application;

[0051] Figure 4 is a schematic structural diagram of an image classification model according to an embodiment of the present application;

[0052] Figure 5 is a schematic diagram of the positioning results of some check codes according to an embodiment of the present application;

[0053] Figure 6 is a schematic diagram of the positioning results of some dates according to an embodiment of the present application;

[0054] Figure 7 is a schematic diagram of the recognition results of some other check codes according to an embodiment of the present application;

[0055] Figure 8 is a schematic diagram of the recognition results of some other dates according to an embodiment of the present application;

[0056] Figure 9Schematic diagram of the recognition results of some amounts according to an embodiment of the present application;

[0057] Figure 10 Schematic diagram of the recognition results of some bill codes according to an embodiment of the present application;

[0058] Figure 11 Schematic diagram of the recognition results of some bill numbers according to an embodiment of the present application;

[0059] Figure 12 Schematic diagram of the positioning results of the face value and serial number key information of a certain banknote according to an embodiment of the present application;

[0060] Figure 13 Schematic diagram of the structure of a character recognition model according to an embodiment of the present application;

[0061] Figure 14 Schematic diagram of the structure of a traditional convolutional neural network;

[0062] Figure 15 Schematic block diagram of the structure of an image recognition device according to an embodiment of the present application;

[0063] Figure 16 Schematic diagram of the hardware structure of a computer device according to an embodiment of the present application. Detailed implementation manners

[0064] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Apparently, the described embodiments are some, but not all, of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.

[0065] The automatic recognition of financial materials such as bills and currencies is an important application field of pattern recognition and a comprehensive research topic. Although the forms of electronic payment are constantly developing and increasing, due to the unique independence and network security risks of enterprises and financial systems such as banks. In actual business, a large number of traditional machine-printed paper bills and cash are still one of the main forms that are relatively popular and recognized in non-cash payment at present. At the same time, the entry, printing, sorting, binding, and filing of various bills and business cash are not only time-consuming and laborious, but also extremely error-prone and affect the subsequent business operations. At the same time, the inefficient bill processing and the recognition and sorting of uncommon foreign currencies also restrict the work efficiency of financial enterprises such as banks and the improvement of their ability to prevent financial risks, and even affect or directly lead to the decline of the overall competitiveness of financial enterprises and the increase of the competition risk coefficient.

[0066] To address the above problems, the need for automatic recognition of bills and multi-currency currencies has emerged. Currently, there are few products on the market that can simultaneously recognize multi-currency currencies and bills. At the same time, the accuracy of existing products or algorithms for bill recognition still needs to be improved.

[0067] The key technology research of the bill and multi-currency currency recognition terminal device based on artificial intelligence in this application mainly involves research fields such as artificial intelligence, image processing, and pattern recognition. The aim is to use the deep learning method in the field of artificial intelligence to establish a neural network model for financial materials, especially for bill and multi-currency currency recognition, and solve the problems of low efficiency and accuracy in image recognition when sorting traditional financial materials, especially bills and multi-currency currencies.

[0068] Figure 1 It is a schematic structural diagram of an image recognition system according to an embodiment of this application. The system includes a terminal 110 and a server 120. The terminal 110 is communicatively connected to the server 120. The server 120 integrates an improved optical character recognition (OCR) service, that is, the image recognition method of this application. The terminal 110 may include a data processing device, a data storage device, and an image acquisition device.

[0069] The terminal 110 uses the image acquisition device to collect the original images of financial materials such as financial bills or cash, and stores the original images in the data storage device. The terminal 110 can be communicatively connected to the server 120 through a transmission network (such as a wireless communication network), and send the original images in the data storage device to the server 120 through the data processing device.

[0070] After receiving the original image, the server 120 performs processing operations such as denoising and cropping on the original image to obtain the image to be recognized. The image classification model obtained by pre-training is used to classify the image to be recognized to obtain the classification result of the image to be recognized; the target recognition model corresponding to the classification result is selected from the target recognition model set, and the target recognition model is used to locate the area to be recognized in the image to be recognized, and text information recognition is performed on the area to be recognized to obtain the recognition result of the area to be recognized; the recognition results of all areas to be recognized are summarized to determine the image recognition result of the image to be recognized. Realize personalized recognition of different types of images to be recognized, and improve the efficiency and accuracy of image recognition.

[0071] Optionally, after the server 120 recognizes the image to be recognized and obtains the classification result of the image to be recognized, it can also perform preprocessing operations such as rotation, binarization, and filtering on the image to be recognized to make the target recognition model recognize the preprocessed image to be recognized, further improving the recognition accuracy of the target recognition model.

[0072] Optionally, after the server 120 recognizes the text information in the area to be recognized, if the classification result of the image to be recognized is a bill image, the server 120 can also use a pre-trained character recognition model to recognize the target characters in the area to be recognized. This can compare and correct the recognition results of the target recognition model, improving the recognition accuracy of the bill image.

[0073] Optionally, to save the storage space of the terminal 110, the data processing device can also directly send the original image to the server 120 after the image acquisition device acquires the original image, so that the server 120 stores the original image.

[0074] Optionally, the server 120 can send the image recognition result of the image to be recognized to the terminal 110 through the transmission network, so as to facilitate the user to view the image recognition result.

[0075] Optionally, the above-mentioned server 120 can be a server, which can be a server cluster or a distributed system composed of multiple physical servers, or a cloud server that provides technical computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms.

[0076] Optionally, the system can further include a management device, which is used to manage the system (such as managing the connection status between each module and the server), and the management device is connected to the server through a communication network. Optionally, the communication network is a wired network or a wireless network.

[0077] Optionally, the above-mentioned wireless network or wired network uses standard communication technologies and / or protocols. The network is usually the Internet, but it can also be any other network, including but not limited to any combination of local area networks, metropolitan area networks, wide area networks, mobile, limited or wireless networks, private networks or virtual private networks. In some embodiments, technologies and / or formats including Hypertext Markup Language, Extensible Markup Language, etc. are used to represent the data exchanged through the network. In addition, conventional encryption technologies such as Secure Sockets Layer, Transport Layer Security, Virtual Private Network, Internet Protocol Security, etc. can be used to encrypt all or some of the links. In other embodiments, customized and / or dedicated data communication technologies can also be used to replace or supplement the above data communication technologies.

[0078] According to an embodiment of the present application, an embodiment of an image recognition method is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.

[0079] In an embodiment of the present application, an image recognition method is provided, which can be used in computer devices, such as mobile phones, tablets, computers, or Figure 1 the server 120 shown. Figure 2 is a flowchart of an image recognition method according to an embodiment of the present application, as Figure 2 shown, and the process includes the following steps:

[0080] Step 201, obtain the image to be recognized.

[0081] The image to be recognized is an image obtained based on the original image of financial materials. The financial materials can be materials filled in or declared by users, such as bills or currencies. The computer device can obtain the original image of the financial materials through an image acquisition device, and perform processing operations such as cropping and image enhancement on the original image to obtain the image to be recognized. It can also be an image of the scanned financial materials uploaded by the user, and the computer device performs processing operations such as cropping and image enhancement on the received image to obtain the image to be recognized.

[0082] Step 202, use the pre-trained image classification model to classify the image to be recognized, and obtain the classification result of the image to be recognized.

[0083] The image classification model can be a neural network model for image classification tasks constructed based on a pre-trained deep learning network, such as LeNet, AlexNet, or VGGNet. The specific type of the deep learning network can be selected by itself. In this application example, the image classification model is an AlexNet neural network. The computer device inputs the image to be recognized into the image classification model to obtain the classification result of the image to be recognized.

[0084] Step 203, select the target recognition model corresponding to the classification result from the target recognition model set.

[0085] Among them, target recognition models matching different classification results are pre-configured in the target recognition model set. The target recognition model can be a neural network model for target detection and recognition tasks constructed based on a pre-trained target detection algorithm, such as FAster-RCNN, YOLO, or SDD. In this embodiment of the present application, the target recognition model is a neural network model based on YOLOv1. The classification results corresponding to each target recognition model are recorded in the target recognition model set, so that the computer device can select at least one target recognition model from the target recognition model set according to different classification results to recognize the image to be recognized.

[0086] Step 204, using the target recognition model, locates the area to be recognized in the image to be recognized, and recognizes the text information in the area to be recognized, to obtain the recognition result of the area to be recognized.

[0087] The computer device inputs the image to be identified into each selected target recognition model. Each target recognition model will locate the area to be identified from the image to be identified and perform text information recognition on the area to be identified, and mark the file information on the image in the area to be identified. The computer device then obtains the recognition result of the area to be identified.

[0088] It can be understood that the target recognition model is used to identify a specific target. The specific recognition method of the target can be learned by the target recognition model during model training. Different target recognition models will locate different areas to be recognized. For example, when the classification result of the image to be recognized is a bill image, different target models are responsible for recognizing different fields (i.e., targets) in the bill image, and can respectively locate five areas to be recognized, namely, invoice code, invoice number, invoice date, check code, and amount. Or when the classification result of the image to be recognized is a currency image, different target models can respectively locate five areas to be recognized, namely, face value of the image, head portrait in the image, watermark, characteristic pattern, and serial number.

[0089] Step 205, summarizing the recognition results of all the areas to be recognized, and determining the image recognition result of the image to be recognized.

[0090] The computer device will aggregate the recognition results of the area to be recognized by each target recognition model in a specific format, so as to obtain the image recognition result of the image to be recognized. The specific format can be set by yourself, and the embodiment of this application does not make any specific restrictions.

[0091] In the embodiment of the present application, the images to be identified can be automatically classified, avoiding manual classification and improving the efficiency of image recognition. It is also possible to select a target recognition model corresponding to the classification result of the image to be identified, locate the area to be identified and identify information for the images to be identified of different categories, and then summarize the recognition results of all the areas to be identified of the image to be identified to determine the image recognition result of the image to be identified, thereby realizing personalized information recognition of images of different categories, improving the accuracy of image recognition, and avoiding manual recognition of images to improve the efficiency of image recognition.

[0092] In this embodiment, another image recognition method is provided, which can be used in computer devices, such as mobile phones, tablet computers, computers or Figure 1 The server 120 shown, Figure 3 is a flowchart of another image recognition method according to an embodiment of the present application, such as Figure 3 As shown in , the process includes the following steps:

[0093] Step 301: Obtain the image to be recognized.

[0094] Step 302: Use the pre-trained image classification model to classify the image to be recognized, and obtain the classification result of the image to be recognized.

[0095] Steps 301 to 302 are specifically described in Figure 2 Steps 201 to 202 of the illustrated embodiment, and will not be elaborated here.

[0096] Optionally, in order to improve the accuracy of image classification, the image classification model may include 5 convolutional layers and 3 fully connected layers. Each convolutional layer includes the activation function ReLU and local response normalization (LRN) processing, and then pooling processing is performed. The Dropout function is used in the convolutional layer to prevent the problem of overfitting during model training. The specific structure of the image classification model can be as Figure 4 shown. MAX POOL represents the pooling layer; Softmax represents the fully connected layer; S represents the stride; 55×55×96 and 13×13×384, etc. represent the size of the image. 3×3 under MAX POOL represents the size of the pooling layer. 9216, 4096, and 1000 all represent the number of neurons. The sizes of the convolutional layers adopted by the image classification model are 3×3, 5×5, and 11×11, and the depths are three types: 96, 256, and 384.

[0097] Before the image classification model classifies the image to be recognized, the image classification model needs to be trained. Specifically, the pooling layer network of the image classification model is divided into two completely identical upper and lower branches, and is trained in parallel in two graphics processing units (GPUs), and the upper and lower networks can perform information interaction in the third convolutional layer and the fully connected layer. During training, an image with an RGB (Red Green Blue) three-channel size of 227×227×3 can be input into the image classification model. The image passes through the first convolutional layer, the first pooling layer, the second convolutional layer, the second pooling layer, the third convolutional layer, the fourth convolution, the fifth convolutional layer, and the third pooling layer in sequence to reach the fully connected layer. The fully connected layer classifies the image, determines the objective function of the image classification model, such as the loss function, or the model prediction accuracy, etc., and adjusts the parameters of the image classification model until the objective function reaches the model training stop condition (for example, the objective function is minimized or the model prediction accuracy reaches a preset value, etc.) to determine that the model training is completed and obtain the trained image classification model.

[0098] When training an image classification model, the weight parameters are randomly generated during initialization and continuously optimized during the subsequent iterative process to achieve the best classification effect. These weight parameters are shared, that is, the parameter values in a filter remain unchanged during the convolution process for all image data, and the number is independent of the image size, only related to the size, depth of the filter, and the matrix depth of the current layer nodes. The Dropout function is used in the convolutional layer to prevent the model from overfitting. The Dropout function value used in the image classification model is 0.5, which means that some neurons are randomly deleted with a probability of 0.5, while keeping the number of input and output neurons in this layer unchanged, and parameter updates are performed. The above operations are repeated during iteration until the training ends.

[0099] The LRN processing in the convolutional layer of the image classification model is implemented by the following formula (1):

[0100]

[0101] Where, represents the output after applying kernel i at position (x,y) and then applying the activation function ReLU. N represents the number of channels, represents the output after normalization. n, k, α, and β are all hyperparameters of the image classification model, and are respectively set to 2, 5, e -4 and 0.75 in the embodiments of this application.

[0102] Five convolutional layers and three fully connected layers are used in the image classification model to make the model have a deeper network structure, so that the model can better learn the image classification task and improve the accuracy of image classification. In addition, the introduction of the Dropout function deletes the "inactive" neurons in the model that no longer perform forward propagation and do not participate in backward propagation from the model, reducing the complex mutual influence between neurons, preventing the model from overfitting, and improving the accuracy of the image classification model for classifying images.

[0103] Optionally, in order to further improve the efficiency and accuracy of image recognition, the image to be recognized can also be preprocessed before step 303 to eliminate irrelevant information in the image, restore useful real information, enhance the detectability of relevant information, and simplify the data to the greatest extent, thereby improving the reliability and efficiency of the target recognition model for recognizing images and the efficiency and accuracy of image recognition. That is, before selecting the target recognition model corresponding to the classification result from the set of target recognition models, the image recognition method further includes:

[0104] Preprocess the image to be recognized to obtain the preprocessed image to be recognized, so as to input the preprocessed image to be recognized into the target recognition model subsequently.

[0105] The steps of preprocessing the image to be recognized to obtain the preprocessed image to be recognized may include the following steps A1 to A4:

[0106] Step A1: Determine the tilt angle of the image to be recognized according to the straight lines of the edges in the image to be recognized and the reference straight line.

[0107] Computer design can use any current method for detecting the tilt angle or rotation of an image to extract the straight lines of the edges in the image to be recognized and the reference straight line to calculate the tilt angle of the image. In the embodiments of the present application, the image rotation method based on Opencv (Open Computer Vision Library, an open-source computer vision library) is taken as an example. Since the image to be recognized obtained by the computer device may be tilted or inverted, the computer device needs to use the image rotation algorithm of Opencv to extract the straight lines of the edges in the image to be recognized and the reference straight line, and calculate the tilt angle.

[0108] Step A2: If the tilt angle does not meet the preset condition, correct the tilt angle of the image to be recognized to obtain the corrected image to be processed.

[0109] The preset condition may be that the tilt angle of the image is 0°. The computer device adjusts the image to be recognized based on any current image transformation method so that the tilt angle of the adjusted image to be recognized meets the preset condition, and obtains the corrected image to be processed. In the embodiments of the present application, the Hough transformation method is used to adjust the image.

[0110] It should be noted that when correcting the tilt angle of the image to be recognized, the image to be recognized will also be converted into a grayscale image for subsequent operations.

[0111] Step A3: Perform binarization processing on the corrected image to be recognized to obtain a binarized image.

[0112] The binarization processing of an image is a process of changing the gray levels in the grayscale image from 256 values from 0 to 255 to 0 or 255. Binarization processing usually takes into account the overall or local features of the image to be recognized. For non-intersecting regions, connected and closed boundaries can be used to represent them. If the gray value of a certain pixel point in the image to be recognized is less than or equal to a set threshold, its gray value is set to 0; if the gray value is greater than the set threshold, its gray value is set to 255. In this process, a gray value of 0 usually represents the background, and 255 represents the object. The computer device can perform binarization processing on the corrected image to be recognized based on any current method for binarization processing, such as the bimodal method, the P-parameter method, the maximum inter-class variance method, the maximum entropy threshold method, or the optimal threshold method.

[0113] The Otsu method divides the image to be recognized into two parts: the part to be retained and the part to be removed according to the gray-scale properties of the image to be recognized. This algorithm uses the law of uniform variance distribution. If the between-class variance between the target and the background is large, it means that there is an obvious difference between the target and the background in the image to be recognized. Considering that the targets to be detected in the images of financial materials, especially bills and currencies, fluctuate in different stable gray-scale regions, the embodiments of the present application preferably adopt the Otsu method to perform binarization processing on the image to be recognized.

[0114] Step A4: Filter the noise from the binarized image. The binarized image after noise filtering is the preprocessed image to be recognized.

[0115] The computer device can remove the noise points in the binarized image based on any current image filtering algorithm, such as Sobel operator filtering, mean filtering, Gaussian filtering, and bilateral filtering, etc., without damaging the main features of the binarized image, to obtain the preprocessed image to be recognized. In the embodiments of the present application, mean filtering is taken as an example to filter the noise from the binarized image. Mean filtering is to take the mean of the gray scale of a certain point in the original image and the gray scales surrounding this point. Through this operation, the noise points in the image can be effectively removed. Mean filtering can be implemented by the following formula (2):

[0116]

[0117] Among them, M represents the size of the preset template, f(x, y) represents the pixel value at the position (x, y) in the original image, g(x, y) represents the pixel value after mean filtering, and (x, y) represents the coordinates of the pixel point in the original image.

[0118] Performing operations such as tilt angle correction, binarization, and noise filtering on the image to be recognized can remove the invalid information in the image to be recognized that may affect the image recognition result, and can also restore the useful real information in the image, enhance the detectability of relevant information, thereby improving the accuracy of image recognition.

[0119] Step 303: Select the target recognition model corresponding to the classification result from the set of target recognition models.

[0120] For details of Step 303, see Figure 2 Step 203 of the embodiment shown, which will not be elaborated here.

[0121] Optionally, step 303 may include: if the classification result is a bill image, respectively select target recognition models for invoice code recognition, invoice number recognition, invoice date recognition, check code recognition, and amount recognition from the target recognition model set; if the classification result is a currency image, respectively select target recognition models for face value recognition, portrait recognition, watermark recognition, pattern recognition, and serial number recognition from the target recognition model set. Select target recognition models for different recognition tasks according to different classification results, so as to perform parallel recognition of different contents on the image to be recognized, avoid reducing the error rate of recognition by recognizing multiple contents at one time, and improve the accuracy and efficiency of image recognition.

[0122] Optionally, before step 303, it is also necessary to train the image recognition model to obtain a target recognition model with high recognition accuracy, so as to improve the accuracy of image recognition. The image recognition method further includes the following steps B1 to step B2:

[0123] Step B1, obtain sample images of different materials to obtain a sample image set.

[0124] Among them, the sample images are labeled with the category of the text information in the area to be recognized, the category of the area to be recognized, and the position information of the area to be recognized. The position information may be coordinate values, and the text information includes numbers, amount symbols, decimal points, letters, Chinese characters, and preset special symbols (for example, $ and ¥). The materials may include bills and currencies of different denominations. The computer device can read the sample images of different materials according to the paths of the sample images in the first preset file, and read the category of the text information in the area to be recognized, the category of the area to be recognized, and the position information of the area to be recognized in the sample images from the first preset file to obtain a sample image set. For example, taking the first preset file as a txt document, the category of the text information in the area to be recognized, the category of the area to be recognized, and the position information of the area to be recognized marked in the sample images are saved in the txt document in advance, and then the paths of the sample images are extracted and saved in the txt document, so that the computer device can obtain the sample image set according to the txt document. Of course, the first preset file can also be other types of files, such as xml files or word files, etc.

[0125] It can be understood that before step B1, it is also necessary to construct a sample image set. Specifically, the original images of different materials can be collected in advance, the areas to be recognized and the categories of the areas to be recognized are marked in each original image, and the text information in each area to be recognized is divided into several categories (that is, the category of the text information is marked) to obtain an initial data set. Perform data augmentation operations on the initial data set, and use the augmented initial data set as the sample image set, and the images in the augmented initial data set as sample images.

[0126] Optionally, the specific data volume of the original images can be set by oneself. For example, it can be 4,500 images, 5,000 images, etc.

[0127] Optionally, considering the accuracy of image recognition, the difference between the numbers of the original images labeled under each type of area to be recognized is within a preset threshold, and the threshold can be 100. The numbers of the original images labeled under each type of area to be recognized range from 100 to 200.

[0128] Optionally, the specific process of labeling the area to be recognized, as well as the category to which the area to be recognized belongs, in each original image, and dividing the text information in each area to be recognized into several categories can be to use a labeling software. For example, using labelme software to label the area to be recognized, the category to which the area to be recognized belongs, and the category to which the text information in the area to be recognized belongs in the original image, to obtain a second preset file with the file type of xml. In this way, the category and location information can be recorded in the second preset file. The category and location information labeled in the sample image can be extracted from the second preset file, the extracted information can be stored in the first preset file, and then the path of the sample image can be extracted and saved in the first preset file, which is convenient for subsequent computer devices to obtain sample images of different materials according to the first preset file to obtain a sample image set.

[0129] It can be understood that the category to which the area to be recognized belongs and the location in the sample images of different materials may be different. For example, the sample image of a bill may contain 5 areas to be recognized with the categories of invoice code, invoice number, issue date, check code, and amount, and these areas to be recognized may be located in the upper left corner, upper right corner, lower right corner, etc. of the sample image. The sample image of currency may contain 5 areas to be recognized with the categories of recognized image face value, portrait in the image, watermark, feature pattern, and serial number, and these areas to be recognized may be located in the lower left corner, upper right, left, lower right, upper right, etc. of the sample image.

[0130] Optionally, the category to which the area to be recognized belongs may include date, check code, invoice code, number, amount, face value, portrait, watermark, pattern, and serial number. Among them, the number can be divided into 10 categories from 0 to 9, the amount can be divided into 12 categories including 0 to 9, decimal point, and amount symbol, and the serial number can be divided into 36 categories including A - Z and 0 to 9. It can be understood that due to the finer granularity of the category division of the number, amount, and serial number, the categories of the number, amount, and serial number can indicate both the category to which the area to be recognized belongs and the category to which the text information in the area to be recognized belongs.

[0131] Optionally, the category to which the text information belongs can be 0-9, A-Z, and, of E13B (magnetic number) characters, a total of 37 categories.

[0132] Exemplarily, taking the sample image of a bill as an example, it can be specifically labeled as follows:

[0133] Invoice code: Label 0 for the black box in the upper left corner of the image;

[0134] Invoice number: Label 1 for the number after "No" in the upper right corner of the image;

[0135] Invoice date: Label 0 for the date in the upper right corner of the image;

[0136] Check code: Label 0 for Chinese characters and 1 for numbers;

[0137] Amount: Label 0 for the amount in the lower right corner of the image.

[0138] Optionally, in order to further improve the accuracy of image recognition, the images within the regions to be recognized can also be intercepted from each of the original labeled images to obtain an initial data set composed of the images within the regions to be recognized; it can be understood that the images within the regions to be recognized are the sample images. In this way, when the initial data set after data augmentation, that is, the sample image set, is obtained later, the sample images in the sample image set are the images with invalid information further filtered out, which can make the accuracy of the image recognition model higher and improve the accuracy of image recognition during the subsequent training of the image recognition model.

[0139] Exemplarily, taking the original image of a labeled bill as an example, five images within the regions to be recognized for identifying the invoice code, invoice number, invoice date, check code, and amount can be roughly intercepted from the upper left corner, upper right corner, lower right corner, and the middle of the image. Taking the original image of a labeled currency as an example, five images within the regions to be recognized for identifying the face value, portrait, watermark, pattern, and serial number of the image can be roughly intercepted from the lower left corner, upper right, left, lower right, and upper right of the image. In this way, these five images can be input into the image recognition models for different recognition tasks during the subsequent training of the image recognition model.

[0140] It should be noted that after training the image recognition model with the labeled images within the regions to be recognized as sample images to obtain the target recognition model, when the target recognition model recognizes the image to be recognized, it will not only locate the region to be recognized in the image to be recognized, but also intercept the image within the region to be recognized from the image to be recognized, recognize the text information within the intercepted image, obtain the recognition result of the region to be recognized, and further filter out the interference information within the image to be recognized when recognizing the text information within the region to be recognized, thereby improving the accuracy of image recognition.

[0141] Furthermore, data augmentation is performed on the initial dataset because the number of images in each category of the initial dataset is mostly unequal, and some categories lack sufficient data, which may lead to overfitting of the subsequent image recognition model. To solve this problem, data augmentation is performed on the images in each category of the initial dataset to make the number of images in each category consistent. Specifically, the method of data augmentation (DataAugmentation) is adopted, which uses label-preserving transformations and the method of artificially increasing the dataset. The main data augmentation strategies include the following categories:

[0142] 1. Rotation: Slightly rotate the original image. New images can be generated by randomly rotating the image within a preset range of angles;

[0143] 2. Scaling: Scale the original image by a certain scale;

[0144] 3. Translation: Move the original image an arbitrary distance in a certain direction. Generally, the distribution of pixel points in the original image will be determined and translated according to certain rules without reducing the information of the text in the original image;

[0145] 4. Noise addition: Adding noise means adding noise to the initial dataset, and the parameters are preset within a certain range and randomly changed. Usually, the types of noise that can be added include Gaussian noise, salt-and-pepper noise, and Poisson noise, etc. The initial dataset after data augmentation is the sample image set, and the images in the initial dataset after data augmentation are the sample images. In the data augmentation strategy adopted in the embodiment of the present application, the maximum rotation angle is 2°, and of course, other rotation angles can also be set by oneself. The above strategy is used to expand the number of sample images in the initial dataset. For example, the number of sample images in each category can be expanded to about 350 images.

[0146] Optionally, considering that when training the image recognition model later, the image recognition image will unify the size of the sample images into a preset size. For example, if it is 48×48, then when performing data augmentation on the initial dataset, the data augmentation strategy of scaling can be omitted to improve the construction efficiency of the sample dataset.

[0147] Optionally, to improve the accuracy of image recognition, the number of sample images in the sample image set for the area to be recognized in each category can be 2000. All or most of the 2000 images can be used for training the image recognition model to improve the accuracy of image recognition.

[0148] Step B2: According to the category to which the text information in the region to be recognized of the sample image in the sample image set belongs, the category to which the region to be recognized belongs, and the position information of the region to be recognized, iteratively train at least one pre-constructed image recognition model until the loss function of the image recognition model meets the training stop condition. The trained image recognition model is a target recognition model in the target recognition model set.

[0149] Among them, the loss function is a function constructed based on the center point positioning error of the detection target, the width and height positioning error of the detection target, the confidence error of the detection target, the confidence error of the non-detection target, the classification error of the detection target, and a preset position loss weight.

[0150] The errors of traditional neural network models mainly come from the prediction target coordinate error, the prediction confidence error, and the prediction classification result error. Among them, the prediction target coordinate error mainly includes the center point positioning error of the detection target, the width and height positioning error of the detection target, and a preset position loss weight. The prediction confidence error mainly includes the confidence error of the detection target and the confidence error of the non-detection target. The prediction classification result error is the classification error of the detection target. Although the loss function of traditional neural network models, such as the mean squared error function, can take into account the above three aspects at the same time, there is still a large error in the model fitting degree. The reasons are as follows: The loss function of traditional neural network models only balances the losses of the above three; in addition, most grids in an image do not have targets. According to the loss function of traditional neural network models, the confidence of grids without targets in the image is close to 0. Cumulatively, the confidence of grids without targets will be relatively large, which will suppress the confidence loss of grids with target boxes, making it easy for the network to diverge and reducing the accuracy of image recognition. To solve this problem, in the embodiments of this application, the loss function pays more attention to the prediction target coordinate error and gives a relatively large loss weight to the prediction target coordinate error, that is, a relatively large value will be given to the preset position loss weight, and this value can be set by oneself. For example, 5; for the confidence loss of grids without targets in the confidence error of non-detection targets, a smaller weight will be given, and this weight can be set by oneself. For example, 0.5; for the confidence loss of grids with targets and the corresponding category loss in the confidence error of detection targets, a relatively large weight will be given, and this weight can be set by oneself. For example, 1;

[0151] In addition, for the case where the sizes of the detection targets are different, the error precision of the prediction for the smaller target must be smaller than that for the larger target. However, the final width, height, center position loss, and the offset loss for each target are the same, which still leads to a reduction in the accuracy of image recognition. To solve this problem, in the embodiments of the present application, for the width and height positioning errors of the detection targets, the square roots of the widths and heights of the targets with different sizes can be used instead of the original widths and heights. In summary, in the embodiments of the present application, the weight of the confidence error of the non-detection target in the loss function is less than the weight of the confidence error of the detection target. In the loss function, the preset position loss weight is greater than the weights of the confidence error and the classification error in the target loss function. In the width and height positioning errors of the detection target in the loss function, the width is the square root of the width of the detection target, and the height is the square root of the height of the detection target. The detection target includes the object that the image recognition model needs to recognize.

[0152] The specific loss function is shown in the following formula (3):

[0153]

[0154] where λ coord is the preset position loss weight; λ noobj is the confidence loss weight of the grid without a target; s 2 is the number of image grids (grid cells); B is the number of recognition boxes (grid boxes); takes the following value: 1 when there is a detection target in the j-th recognition box of the i-th image grid, otherwise 0; w i is the width of the recognition box; h i is the height of the recognition box; C i is the detected confidence; takes the following value: 0 when the detection target is in the j-th recognition box of the i-th image grid, otherwise 1; represents whether the center of the detection target falls within the i-th image grid. If so, it is 1; if not, it is 0; p i (c) is the confidence that the detection target belongs to class c, and classes are all the classes to be detected; is the center point positioning error of the detection target, is the width and height positioning error of the detection target, is the confidence error of the detection target, is the confidence error of the non-detection target, is the classification error of the detection target.

[0155] The computer device extracts sample images from the sample image set according to the category to which the area to be recognized belongs, and inputs the extracted sample images into the image recognition models corresponding to the category to which the area to be recognized belongs. In this way, each image recognition model can learn the category to which the text information in the area to be recognized of the sample image belongs, the category to which the area to be recognized belongs, and the position information of the area to be recognized, and calculate the loss function of the model. For each image recognition model, if the loss function does not meet the training stop condition, sample images are extracted again from the sample image set to retrain the image recognition model. If the loss function meets the training stop condition, the training is stopped, and it is determined that the image recognition training is completed. The trained image recognition model is stored in the target recognition model set as a target recognition model.

[0156] When training the image recognition model, the weight of the position error in the loss function is made greater than the weights of the confidence error and the classification error, ensuring that the model pays more attention to the position of the area to be recognized in the image and ensuring the accuracy of the positioning of the area to be recognized. And the confidence error of the detection target is greater than the confidence error of the non-detection target, which is beneficial to suppressing the confidence loss of the non-detection target, facilitating the convergence of the image recognition model, and accelerating the training speed of the image recognition model. Using the square root of the width and height of the detection target to replace the original width and height in the width and height positioning error of the detection target can reduce the width difference between detection targets of different sizes, so that the error precision of detection targets of different sizes in the target loss matches the size of the detection target, thereby improving the accuracy of the target recognition model after training and the accuracy of image recognition.

[0157] Optionally, to improve the accuracy and efficiency of image recognition, the image recognition model includes an input layer, a neural network layer, and an output layer. The neural network layer includes a convolutional layer and a pooling layer. The convolutional layer includes multiple convolutional kernels with a scale of 3×3, and convolutional kernels with a scale of 1×1 interspersed between the 3×3 convolutional kernels; the sample image set includes a test sample set and a training sample set.

[0158] Among them, the input layer is used to transform the size of the sample images in the test sample set into a preset size and input the sample images with the transformed size into the neural network layer. The preset size can be set by oneself. For example, it can be 448×448×3. The neural network layer is used to utilize the convolutional layer and the pooling layer to learn the category to which the text information in the to-be-recognized region in the sample image with the transformed size belongs, the category to which the to-be-recognized region belongs, and the position information of the to-be-recognized region, so as to train the parameters in the image recognition model and perform normalization and Dropout operations on the data processed by the convolutional layer and the pooling layer; the output layer is used to determine the value of the loss function according to the test sample set. The input layer will transform the size of the image so that the sizes of the images input into the model are all the same, which is convenient for the model to learn features and recognize images subsequently, and improves the efficiency of model training and the efficiency of image recognition. Inserting convolutional kernels with a scale of 1×1 between 3×3 convolutional kernels can deepen the network depth of the image recognition model, that is, deepen the network depth of the subsequent object recognition model, which is beneficial to improving the efficiency and accuracy of image recognition. In addition, normalization and Dropout operations can also avoid overfitting of the model when training the image recognition model, thereby improving the accuracy of the image recognition model in recognizing images, and further improving the accuracy of image recognition in the subsequent application process.

[0159] Optionally, the neural network layer may further include a fully connected layer. In the neural network layer, except for the convolutional layers with 3×3 and 1×1 convolutional kernels, the sizes of the remaining convolutional layers can be set by oneself. The stride of the convolutional kernel can be set by oneself. In the embodiments of the present application, the stride is taken as 2 for example. Except for the convolutional layers with 3×3 and 1×1 convolutional kernels, it may also include 2×2 convolutional kernels.

[0160] Optionally, the number of convolutional layers, pooling layers, and fully connected layers, as well as the pooling strategy, can all be set by oneself. In the embodiments of the present application, taking 24 convolutional layers and 2 fully connected layers as an example, global average pooling is used for the final prediction of the network.

[0161] Optionally, the number of sample images in the test sample set can be set by oneself, and preferably it is 30.

[0162] Optionally, sample images can also be extracted from the sample image set batch by batch according to the category to which the to-be-recognized region belongs. Each time the image recognition model is trained, the number of sample images grabbed can be a preset number, and the preset number can be set by oneself.

[0163] Optionally, when the number of iterations of the image recognition model reaches a threshold, the image recognition model is tested according to the test sample set to determine the loss function of the image recognition model. The threshold can be set by oneself, for example, 10000 times or 8000 times.

[0164] In an application scenario, when iteratively training an image recognition model, the regions to be recognized in the sample images of bills that belong to the category of verification codes can be divided into two categories. One category is used to indicate the Chinese character "verification code", and the other category is used to indicate the numbers after the Chinese character; the regions to be recognized that belong to the category of amount are divided into 12 categories (respectively 0-9, decimal point, and amount symbol), and the regions to be recognized of other categories are recognized and divided into 10 categories (that is, 0-9). The learning rate of the model is 0.01. After each training of all images, the learning rate is 0.96 times the original. The Dropout function is used to prevent overfitting, and the value of the Dropout function is 0.6. Test is performed every 10,000 iterations. The momentum is a constant 0.88, the weight decay coefficient is 0.0003, and the value of the batch-size (that is, the number of sample images grabbed) for inputting sample images in batches during training is 128.

[0165] Optionally, when testing the image recognition model according to the test sample set, it is also possible to output the located regions to be recognized, the regions to be recognized, and the categories to which the regions to be recognized belong (that is, the recognition results of the text information).

[0166] In yet another application scenario, in combination with the annotation in the above example where the numbers after the invoice number image "No" in the sample image of the bill are marked as 1, and the Chinese characters in the verification code are marked as 0 and the numbers are marked as 1, etc., the image recognition model is tested and the output is as Figures 5 to 12 shown in the result. Figure 5 For the positioning results of some verification codes, since the position of the verification code is not fixed, sometimes it is located in the upper left corner of the image, and sometimes it is located in the lower right corner of the image. Here, the position of the verification code is judged by judging the three Chinese characters of the verification code. Therefore, the verification code image positioning is divided into two categories: the Chinese character "verification code" and the content of the verification code. As shown by the numbers 0 and 1 above the recognition frame in the following figure, 0 indicates the Chinese character, and 1 indicates the content of the verification code. From Figure 5 it can be seen that when the verification code is penetrated by a horizontal line at the top of the image, covered by a seal at the bottom of the image, and the image is blurred, the position of the verification code can be accurately located. The "xxxx Co., Ltd." in the figure is the name of the enterprise, and the name and taxpayer identification number in the figure are the fixed information of the bill.

[0167] Figure 6 For the positioning result of the date, the date is only divided into one category. Due to standardization issues, problems such as blurred date printing and misalignment frequently occur. Figure 6 The content other than the date in the figure is the image of the QR code or invoice code, etc. around the bill date. The image recognition model used for date recognition has solved all the above problems. The 0 above the recognition frame indicates that the content in the frame represents the date.

[0168] Figure 7For the recognition results of some more check codes, the images for check code recognition are from the results obtained by check code localization, that is, the images intercepted from the regions to be recognized belonging to the category of check codes, namely the images obtained. The images of the check codes in this time are divided into 10 categories in total from 0 to 9. The main problems in check code recognition are: the check code images are covered by seals, penetrated by wire meshes, the images are blurred, and the characters in the images are inclined and cannot be accurately segmented. The image recognition model for check code recognition has well solved the above problems.

[0169] Figure 8 For the recognition results of dates, the images for date recognition are from the results obtained by date localization, that is, the images intercepted from the regions to be recognized belonging to the category of dates, namely the images obtained. The images of the dates in this time are divided into 10 categories in total from 0 to 9. The main problems in date recognition are: the date images are blurred, there are interferences of Chinese characters "year", "month", and "day", and some contents of the characters in the images are missing. For the above problems, the image recognition model for date recognition has solved them all.

[0170] Figure 9 For the recognition results of amounts, the images for amount recognition are from the results obtained by amount localization, that is, the images intercepted from the regions to be recognized belonging to the category of amounts, namely the images obtained. The images of the amounts in this time are divided into 12 categories including 0 - 9, decimal points, and amount symbols. The main problems in amount recognition are: the amount images are covered by seals, penetrated by wire meshes, and the recognition of decimal points in the images is regarded as noise. For the above problems, the image recognition model for amount recognition has solved them all. Figure 9 Among them, 10 represents the amount symbol, and 11 represents the decimal point.

[0171] Figure 10 For the recognition results of bill codes, the images for bill code recognition are from the results obtained by bill code localization, that is, the images intercepted from the regions to be recognized belonging to the category of bill codes, namely the images obtained. The images of the bill codes in this time are divided into 10 categories in total from 0 to 9. The main problems in bill code recognition are that the bill code images are blurred and covered by two-dimensional codes. For the above problems, the image recognition model for bill code recognition has solved them all.

[0172] Figure 11 For the recognition results of bill numbers, the images for bill number recognition are from the results obtained by bill number localization, that is, the images intercepted from the regions to be recognized belonging to the category of bill numbers, namely the images obtained. The images of the bill numbers in this time are divided into 10 categories in total from 0 to 9. The main problems in bill number recognition are: the tax bill number images are blurred, the fonts of the tax bill numbers are uncertain, and there are interferences of other fonts beside the tax bill numbers. For the above problems, the image recognition model for bill number recognition has solved them all.

[0173] Figure 12The positioning results of the face value and serial number key information of a certain banknote are shown. The face value and the position where the serial number is located are marked by black rectangular frames. Compared with the positioning of the area to be recognized on the bill, the positioning of the area to be recognized for the currency is relatively stable because the size and information of genuine banknotes are relatively fixed.

[0174] Step 304: Use the target recognition model to locate the area to be recognized in the image to be recognized, and recognize the text information in the area to be recognized to obtain the recognition result of the area to be recognized.

[0175] For the details of step 304, please refer to Figure 2 Step 204 of the embodiment shown, which will not be elaborated here.

[0176] Since there are various special characters on the bill and currency, such as the E13B on bank checks and irregular special characters such as the serial number in the currency, ordinary image recognition methods are difficult and error-prone. Therefore, in the embodiments of the present application, these characters are also recognized separately. Optionally, after recognizing the text information in the area to be recognized, the image recognition method further includes: using a pre-trained character recognition model to recognize the target characters in the area to be recognized.

[0177] Among them, the target characters include magnetic numbers and currency symbols. Of course, it can also include other self-set characters. Taking E13B as an example for the magnetic number, the currency symbol can be the currency symbols of different currencies in different countries. The character recognition model can be a residual neural network, and the network structure of the character recognition model can be as Figure 13 shown, including multiple weight layers and convolutional branches.

[0178] Compared with the traditional convolutional neural network that uses multiple weight layers to fit the objective function H(x), that is, directly mapping x to H(x) as shown in Figure 14 shown, the character recognition model adds a "shortcut" (that is, a convolutional branch) to the path of x mapping to H(x), and converts the fitting target to H(x) = F(x) + x. It changes the situation where the nth layer of the traditional convolutional neural network can only be connected to the n + 1th layer. The addition of the residual block (that is, the part of the neural network with a convolutional branch) allows the data to pass through the layer with the residual block structure. When the deeper network finds that the recognition error becomes larger, it will use the backpropagation algorithm to adjust the weights of each layer. At this time, the weights of the layer with the residual block structure will gradually approach 0, so the mapping relationship from x to H(x) will be transformed into H(x) ≈ x, that is, the identity mapping, which can ensure that the subsequent layers will not be affected by this layer and there will be no phenomenon of increasing error in the follow-up. Thus, it improves the accuracy of recognizing special characters, that is, target characters. It also avoids the problem of decreasing recognition accuracy caused by increasing the network depth when using the traditional convolutional neural network.

[0179] Optionally, before using the pre-trained character recognition model to recognize the target characters in the area to be recognized, it is also necessary to iteratively train the pre-constructed residual neural network, and the trained residual neural network is the character recognition model. The specific training process of the residual neural network is similar to the training process of the image recognition model, which will not be elaborated here.

[0180] Step 305: Aggregate the recognition results of all areas to be recognized to determine the image recognition result of the image to be recognized.

[0181] For details of step 305, please refer to Figure 2 Step 205 of the illustrated embodiment, which will not be elaborated here.

[0182] Step 306: Obtain the correction information corresponding to the image to be recognized.

[0183] The computer device can read the correction information in a specific area of the image to be recognized, or can recognize the image to be recognized again to obtain a secondary recognition result, and use the secondary recognition result as the correction information.

[0184] Optionally, in order to improve the accuracy of the correction information, for the image to be recognized whose classification result is a bill image, step 306 may include the following steps C1 to C3:

[0185] Step C1: Search for the information code positioning identifier in the image to be recognized to obtain the information code area.

[0186] The information code can be a two-dimensional code or a bar code. In this embodiment of the application, the two-dimensional code is taken as an example. The information codes in the image to be recognized are all set with positioning identifiers, and the relative positions are fixed. The computer device can call the information code recognition engine to search for the positioning identifiers with fixed relative positions in the pre-processed image to be recognized, and then can judge the correct direction of the information code, and connect the positioning identifiers to obtain the area as the information code area.

[0187] Step C2: Perform data conversion on the image in the information code area to obtain the information code area data.

[0188] The computer device can read the binary data from the information code area, decode the binary data to obtain the valid information and error correction information in the information code area, and use the error correction information to correct the valid information to obtain the information code area data.

[0189] Step C3: Determine the information code area data in the target format as the correction information corresponding to the image to be recognized.

[0190] The computer device obtains the correction information after organizing the obtained information code area data into the target format, and the target format can be set by itself.

[0191] In an application scenario, taking the QR code in the bill image as the information code as an example, the information code recognition engine can be ZXing or Zbar. The underlying layer of ZXing is developed in Java and is generally used for the Android platform; the underlying layer of Zbar is developed in C and can be compatible with multiple application platforms. The computer device can call the Zbar Api function to locate the position of the QR code. The key to this step is to find the three positioning identifiers of the QR code. The coordinates of the QR code image are determined by the three positioning identifiers, and the relative positions of the three positioning identifiers are fixed; as long as these three positioning identifiers are found, the correct direction of the QR code can be determined. Then, the position where the QR code stores information is determined by connecting the separators of the three positioning identifiers. The purpose of the existence of the separator is to separate the positioning identifier, the stored data, the error correction code, and the version information of the QR code.

[0192] The computer device can perform data conversion on the image within the information code area, where black and white respectively represent two binary information "0" and "1", and through a certain arrangement, the binary format is converted into its original information, that is, the valid information and the error correction information corresponding to the QR code.

[0193] The computer device arranges the obtained data in a format arranged by the bill type, bill code, bill number, tax-exclusive amount, date, check code, and randomly generated confidential information to obtain the recognition result of the QR code, that is, the calibration information. For example, 01, 04, 1200153320, 07041662, 183.49, 20221221, 83623873463907646339, 5080, where 01 represents the bill type, 04 represents the ordinary invoice, 01 represents the special invoice, 1200153320 is the bill code, 07041662 is the bill number, 183.49 is the tax-exclusive amount, 20221221 is the date, 83623873463907646339 is the check code, and 5080 is the randomly generated confidential information. Through testing, it is found that the recognition rate of the QR code for machine-printed value-added tax is as high as over 90%.

[0194] Step 307: Calibrate the image recognition result according to the calibration information to obtain the calibrated image recognition result.

[0195] The computer device will compare the calibration information with the image recognition result and modify the information in the image recognition result that is inconsistent with the calibration information to obtain the calibrated image recognition result.

[0196] In the embodiments of the present application, the input layer of the image recognition model transforms the size of the image so that the sizes of the images input into the model are all the same, which facilitates the subsequent feature learning and image recognition of the model, and improves the efficiency of model training and the efficiency of image recognition. Interspersing 1×1 convolutional kernels between 3×3 convolutional kernels can deepen the network depth of the image recognition model, that is, deepen the network depth of the subsequent target recognition model, which is beneficial to improving the efficiency and accuracy of image recognition. In addition, normalization and Dropout operations can also avoid overfitting of the model when training the image recognition model, thereby improving the accuracy of the image recognition model in recognizing images, and further improving the accuracy of image recognition in the subsequent application process. Thus, the target recognition model obtained based on the image recognition model has high image recognition efficiency and accuracy. The image recognition result is also corrected after the image is recognized, further improving the accuracy of image recognition.

[0197] In this embodiment, an image recognition device is also provided. This device is used to implement the above-mentioned embodiments and preferred implementation manners, and those that have been described will not be repeated. As used below, the term "module" can be a combination of software and / or hardware that realizes a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation in hardware, or a combination of software and hardware is also possible and contemplated.

[0198] This embodiment provides an image recognition device, as Figure 15 shown, including:

[0199] A first acquisition module 1510, configured to acquire an image to be recognized;

[0200] A classification module 1520, configured to classify the image to be recognized by using a pre-trained image classification model to obtain a classification result of the image to be recognized;

[0201] A selection module 1530, configured to select a target recognition model corresponding to the classification result from a set of target recognition models, where a set of target recognition models is pre-configured with target recognition models matching different classification results;

[0202] A localization and recognition module 1540, configured to use the target recognition model to perform localization of the area to be recognized on the image to be recognized, and perform text information recognition on the area to be recognized to obtain a recognition result of the area to be recognized;

[0203] A summarization module 1550, configured to summarize the recognition results of all areas to be recognized to determine the image recognition result of the image to be recognized.

[0204] In some alternative embodiments, the image recognition device further includes:

[0205] An identification model for identifying target characters in a region to be identified by using a pre-trained character recognition model, where the target characters include magnetic numbers and currency symbols.

[0206] In some alternative embodiments, the image recognition device further includes:

[0207] A second acquisition module for acquiring calibration information corresponding to the image to be identified;

[0208] A calibration model for calibrating the image recognition result according to the calibration information to obtain a calibrated image recognition result.

[0209] In some alternative embodiments, the second acquisition module includes:

[0210] A search unit for searching for an information code positioning identifier in the image to be identified to obtain an information code region;

[0211] A conversion unit for performing data conversion on the image in the information code region to obtain information code region data;

[0212] A determination unit for determining the target format information code region data as the calibration information corresponding to the image to be identified.

[0213] In some alternative embodiments, the image recognition device further includes:

[0214] A third acquisition module for acquiring sample images of different materials to obtain a sample image set, where the sample images are labeled with the category of the text information in the region to be identified, the category of the region to be identified, and the position information of the region to be identified;

[0215] A training module for iteratively training at least one pre-constructed image recognition model according to the category of the text information in the region to be identified, the category of the region to be identified, and the position information of the region to be identified in multiple sample images in the sample image set until the loss function of the image recognition model satisfies the training stop condition, and the trained image recognition model is a target recognition model in the target recognition model set;

[0216] The loss function is a function constructed based on the center point positioning error of the detection target, the width and height positioning error of the detection target, the confidence error of the detection target, the confidence error of the non-detection target, the classification error of the detection target, and a preset position loss weight;

[0217] The weight of the confidence error of non-detection targets in the loss function is less than the weight of the confidence error of detection targets. In the loss function, the preset position loss weight is greater than the weights of the confidence error and the classification error in the target loss function. In the loss function, the width in the width and height positioning error of the detection target is the square root of the width of the detection target, and the height in the width and height positioning error of the detection target is the square root of the height of the detection target.

[0218] In some alternative embodiments, the image recognition model includes an input layer, a neural network layer, and an output layer. The neural network layer includes a convolutional layer and a pooling layer. The convolutional layer includes a plurality of 3×3 convolutional kernels and 1×1 convolutional kernels interspersed between the 3×3 convolutional kernels. The sample image set includes a training sample set and a test sample set.

[0219] The input layer is used to transform the size of the sample images in the training sample set into a preset size and input the sample images with the transformed size into the neural network layer.

[0220] The neural network layer is used to utilize the convolutional layer and the pooling layer to learn the category to which the text information in the area to be recognized in the sample images with the transformed size belongs, the category to which the area to be recognized belongs, and the position information of the area to be recognized, so as to train the parameters in the image recognition model, and perform normalization and Dropout operations on the data processed by the convolutional layer and the pooling layer.

[0221] The output layer is used to determine the value of the loss function according to the test sample set.

[0222] In some alternative embodiments, the selection module includes:

[0223] The first selection unit is used to, if the classification result is a bill image, respectively select target recognition models for invoice code recognition, invoice number recognition, invoice date recognition, check code recognition, and amount recognition from the target recognition model set.

[0224] The second selection unit is used to, if the classification result is a currency image, respectively select target recognition models for face value recognition, portrait recognition, watermark recognition, pattern recognition, and serial number recognition from the target recognition model set.

[0225] The further function descriptions of the above-mentioned various modules and units are the same as those in the corresponding above embodiments and will not be elaborated here.

[0226] The image recognition device in this embodiment is presented in the form of functional units. Here, the unit refers to an ASIC (Application Specific Integrated Circuit) circuit, a processor and a memory that execute one or more software or fixed programs, and / or other devices that can provide the above functions.

[0227] The embodiments of the present application further provide a computer device having the above-mentioned Figure 15 image recognition device shown.

[0228] Please refer to Figure 16 , Figure 16 which is a schematic structural diagram of a computer device provided by an alternative embodiment of the present application. As Figure 16 shown, the computer device includes: one or more processors 10, a memory 20, and an interface for connecting each component, including a high-speed interface and a low-speed interface. Each component communicates with each other using different buses and can be installed on a common motherboard or installed in other ways as needed. The processor can process instructions executed within the computer device, including instructions stored in the memory or on the memory to display graphical information of the GUI on an external input / output device (such as a display device coupled to the interface). In some alternative embodiments, if necessary, multiple processors and / or multiple buses can be used together with multiple memories and multiple memories. Similarly, multiple computer devices can be connected, and each device provides some necessary operations (for example, as a server array, a set of blade servers, or a multi-processor system). Figure 16 One processor 10 is taken as an example in

[0229] The processor 10 can be a central processing unit, a network processor, or a combination thereof. Among them, the processor 10 can further include a hardware chip. The above-mentioned hardware chip can be an application-specific integrated circuit, a programmable logic device, or a combination thereof. The above-mentioned programmable logic device can be a complex programmable logic device, a field programmable gate array, a general array logic, or any combination thereof.

[0230] Among them, the memory 20 stores instructions executable by at least one processor 10, so that the at least one processor 10 executes the method shown in the above embodiments.

[0231] The memory 20 can include a program storage area and a data storage area. Among them, the program storage area can store an operating system and application programs required for at least one function; the data storage area can store data created according to the use of the computer device, etc. In addition, the memory 20 can include a high-speed random access memory, and can also include a non-transitory memory, such as at least one disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some alternative embodiments, the memory 20 can optionally include a memory remotely set relative to the processor 10, and these remote memories can be connected to the computer device through a network. Examples of the above-mentioned network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.

[0232] The memory 20 may include volatile memory, such as random access memory; the memory may also include non-volatile memory, such as flash memory, hard disk or solid state drive; the memory 20 may further include a combination of the above types of memory.

[0233] The computer device further includes a communication interface 30 for the computer device to communicate with other devices or communication networks.

[0234] The embodiments of the present application also provide a computer-readable storage medium. The methods according to the embodiments of the present application can be implemented in hardware, firmware, or be implemented as computer code that can be recorded on a storage medium, or be implemented by downloading through a network the original computer code stored in a remote storage medium or a non-transitory machine-readable storage medium and to be stored in a local storage medium, so that the methods described herein can be stored in such software processes on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only memory, a random access memory, a flash memory, a hard disk or a solid state drive, etc.; further, the storage medium may further include a combination of the above types of memory. It can be understood that a computer, a processor, a microprocessor controller or programmable hardware includes a storage component that can store or receive software or computer code, and when the software or computer code is accessed and executed by the computer, the processor or the hardware, the methods shown in the above embodiments are implemented.

[0235] Although the embodiments of the present application are described in conjunction with the accompanying drawings, those skilled in the art can make various modifications and variations without departing from the spirit and scope of the present application, and such modifications and variations all fall within the scope defined by the appended claims.

Claims

1. An image recognition method, characterized in that: The method comprises: Obtain an image to be recognized; Using a pre-trained image classification model to classify the image to be identified, and obtaining a classification result of the image to be identified; Selecting a target recognition model corresponding to the classification result from a target recognition model set, wherein the target recognition model set is pre-configured with target recognition models matching different classification results; Using the target recognition model, positioning the area to be recognized in the image to be recognized, and recognizing text information in the area to be recognized to obtain a recognition result of the area to be recognized; The recognition results of all the areas to be recognized are summarized to determine the image recognition result of the image to be recognized.

2. The method according to claim 1, characterized in that After performing text information recognition on the to-be-recognized area, the method further includes: The target characters in the to-be-recognized area are recognized using a pre-trained character recognition model, wherein the target characters include magnetic numbers and currency symbols.

3. The method according to claim 1 or 2, characterized in that: After summarizing the recognition results of all the to-be-recognized areas to determine the image recognition result of the to-be-recognized image, the method further includes: Acquire correction information corresponding to the image to be recognized; The image recognition result is corrected according to the correction information to obtain a corrected image recognition result.

4. The method according to claim 3, characterized in that The obtaining of correction information corresponding to the image to be recognized includes: Searching for the information code location mark in the image to be identified to obtain the information code area; Performing data conversion on the image in the information code area to obtain information code area data; The information code area data in the target format is determined as the correction information corresponding to the image to be identified.

5. The method according to claim 4, characterized in that Before selecting a pre-trained target recognition model corresponding to the classification result from the target recognition model set, the method further includes: Acquire sample images of different materials to obtain a sample image set, wherein the sample images are annotated with the category of the text information in the to-be-recognized area, the category of the to-be-recognized area, and the location information of the to-be-recognized area; Iteratively training at least one pre-built image recognition model according to the category of text information in the to-be-recognized area of ​​a plurality of sample images in the sample image set, the category of the to-be-recognized area, and the position information of the to-be-recognized area, until the loss function of the image recognition model meets the training stop condition, and the trained image recognition model is a target recognition model in the target recognition model set; The loss function is a function constructed based on the center point positioning error of the detection target, the width and height positioning error of the detection target, the confidence error of the detection target, the confidence error of the non-detection target, the classification error of the detection target and the preset position loss weight; In the loss function, the weight of the confidence error of the non-detected target is smaller than the weight of the confidence error of the detected target, in the loss function, the weight of the preset position loss is larger than the weight of the confidence error and the weight of the classification error in the target loss function, and in the loss function, the width in the width-height positioning error of the detected target is the square root of the width of the detected target, and the height in the width-height positioning error of the detected target is the square root of the height of the detected target.

6. The method according to claim 5, characterized in that The image recognition model includes an input layer, a neural network layer and an output layer, the neural network layer includes a convolution layer and a pooling layer, the convolution layer includes a plurality of convolution kernels with a scale of 3×3, and convolution kernels with a scale of 1×1 interspersed between the 3×3 convolution kernels; the sample image set includes a training sample set and a test sample set; The input layer is used to transform the size of the sample image in the training sample set into a preset size, and input the sample image after the size transformation into the neural network layer; The neural network layer is used to use the convolution layer and the pooling layer to learn the category of the text information in the area to be identified in the sample image after the size transformation, the category of the area to be identified and the position information of the area to be identified, so as to train the parameters in the image recognition model, and to perform normalization and Dropout operations on the data processed by the convolution layer and the pooling layer; The output layer is used to determine the value of the loss function according to the test sample set.

7. The method according to any one of claims 1 or 2 or 4 to 6, characterized in that The selecting a target recognition model corresponding to the classification result from a set of pre-trained target recognition models comprises: If the classification result is a bill image, target recognition models for invoice code recognition, invoice number recognition, invoice date recognition, check code recognition and amount recognition are selected from the target recognition model set respectively; If the classification result is a currency image, target recognition models for face value recognition, portrait recognition, watermark recognition, pattern recognition and serial number recognition are selected from the target recognition model set.

8. An image recognition device, characterized in that: The device comprises: A first acquisition module, used to acquire an image to be identified; A classification module, used to classify the image to be identified using a pre-trained image classification model to obtain a classification result of the image to be identified; A selection module, used to select a target recognition model corresponding to the classification result from a target recognition model set, wherein the target recognition model set is pre-configured with target recognition models matching different classification results; A positioning and recognition module, used to use the target recognition model to locate the area to be recognized in the image to be recognized, and to recognize text information in the area to be recognized, so as to obtain a recognition result of the area to be recognized; The summarizing module is used to summarize the recognition results of all the areas to be recognized and determine the image recognition result of the image to be recognized.

9. A computer device, characterized in that: include: A memory and a processor, wherein the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the image recognition method according to any one of claims 1 to 7 by executing the computer instructions.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a computer to execute the image recognition method according to any one of claims 1 to 7.

Citation Information

Cited By

  • Intelligent multi-medium image recognition and automatic processing method and system

    CN121459364A