Method, apparatus, and server for identifying text strings
Through the improved recognition model, the problem of small size and low resolution in character recognition is solved, and high-precision and efficient character recognition are achieved, which is suitable for complex scenarios such as banknote crown font size.
Patent Information
- Application Number
- CN202110370697.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-04-07
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2041-04-07
AI Technical Summary
The prior art is difficult to effectively recognize that the character size and resolution in images are small, resulting in poor recognition accuracy.
The preset hollow convolution layer is used instead of the combination of the convolution network layer and the pooling layer to build an improved recognition model, combine the positioning sub-model and the classification sub-model, extract image features through the hollow convolution layer, and filter candidate boxes using anchor regression and softening non-maximum suppression algorithms to achieve end-to-end character recognition.
It improves the accuracy and efficiency of character recognition, reduces recognition errors, and is suitable for images with small character size and low resolution, especially in scenes with high recognition difficulty such as banknote crown font size.
Smart Images

Figure CN113095313B_ABST
Abstract
Description
Technical Field
[0001] This specification belongs to the technical field of artificial intelligence, and particularly relates to a method, apparatus, and server for identifying text strings. Background Art
[0002] In many data processing scenarios, it is often necessary to first identify and extract the text strings contained in an image; and then use the obtained text strings for the next step of business data processing.
[0003] However, sometimes the character size of the string in the image to be recognized and processed is small and the resolution is low, and the image features related to the string that can be extracted by a conventional recognition model are relatively few. In view of the above situation, when identifying and extracting the string based on the existing method, errors are likely to occur, and the accuracy of the obtained string is relatively poor.
[0004] In view of the above problems, no effective solution has been proposed yet. Summary of the Invention
[0005] This specification provides a method, apparatus, and server for identifying text strings, which can be well applicable to the situation where the character size of the string in the image is small, the resolution is low, and the relevant image features that can be extracted are relatively few, and accurately identify and determine the target string contained in the target image.
[0006] This specification provides a method for identifying text strings, including:
[0007] Obtain a target image containing a target string to be recognized;
[0008] Preprocess the target image to obtain a preprocessed target image;
[0009] Call a preset recognition model to process the preprocessed target image to obtain a corresponding target processing result; wherein, the preset recognition model at least includes a preset dilated convolutional layer; the preset dilated convolutional layer is used to extract target image features related to the target string from the preprocessed target image instead of a combination of a convolutional network layer and a pooling layer;
[0010] Determine the target string in the target image according to the target processing result.
[0011] In one embodiment, the preset recognition model further includes a positioning sub-model; wherein, the positioning sub-model is connected to the preset dilated convolutional layer, and the positioning sub-model is configured to generate a plurality of corresponding candidate boxes for each text character in the target string through anchor regression according to the target image features and the preset anchor box parameters; and screen out a candidate box that meets the requirements from the corresponding plurality of candidate boxes for each text character as the bounding box of the text character; the bounding box carries the position information of the text character it contains.
[0012] In one embodiment, the preset anchor box parameters are obtained in the following manner:
[0013] Obtain a sample image containing a sample string;
[0014] According to the preset annotation rules, in the sample image, label corresponding bounding boxes for each sample character in the sample string; and collect the bounding box parameters of the sample characters; wherein, the overlapping area range between the bounding boxes of two adjacent sample characters is less than the preset area range threshold;
[0015] Perform clustering processing on the bounding box parameters of the sample characters to obtain the preset anchor box parameters.
[0016] In one embodiment, screening out a candidate box that meets the requirements from the corresponding plurality of candidate boxes for each text character as the bounding box of the text character includes:
[0017] Screen out a candidate box that meets the requirements from the corresponding plurality of candidate boxes for the current text character in the target string as the bounding box in the following manner:
[0018] Call the preset soft non-maximum suppression algorithm to process the plurality of candidate boxes to screen out a candidate box with a confidence level that meets the requirements as the bounding box of the current text character; and filter out other candidate boxes except the bounding box from the plurality of candidate boxes.
[0019] In one embodiment, the preset recognition model further includes a classification sub-model; wherein, the classification sub-model is connected to the dilated convolutional layer, and the classification sub-model is configured to identify and determine the category values of each text character in the target string to be recognized according to the target image features.
[0020] In one embodiment, preprocessing the target image includes:
[0021] Detect the target image and determine a target image region in the target image that contains the target string to be recognized;
[0022] Crop out the target image area from the target image as the preprocessed target image.
[0023] In one embodiment, preprocessing the target image further includes:
[0024] Performing image correction processing on the target image; and / or, performing noise reduction processing on the target image.
[0025] In one embodiment, the target string to be recognized includes at least one of the following: the serial number on the target currency, the drawer account number on the target check, and the logistics number on the target express waybill.
[0026] In one embodiment, when the target string to be recognized includes the serial number on the target currency, after determining the target string in the target image, the method further includes:
[0027] Determine the target string as the serial number on the target currency;
[0028] Track and determine the transaction flow path of the target currency according to the serial number on the target currency;
[0029] Determine whether there is a transaction risk according to the transaction flow path of the target currency.
[0030] In one embodiment, the method further includes:
[0031] Use a preset dilated convolutional layer to replace the combination of the convolutional network layer and the pooling layer as the extraction structure of the image features in the network model to construct an initial recognition model;
[0032] Obtain a sample image containing a sample string to be recognized; and perform annotation processing on the sample image to obtain an annotated sample image;
[0033] Train the initial recognition model using the annotated sample image to obtain a preset recognition model.
[0034] In one embodiment, performing annotation processing on the sample image to obtain an annotated sample image includes:
[0035] According to a preset annotation rule, in the sample image, respectively mark corresponding bounding boxes for each sample character in the sample string; wherein, the overlapping area range between the bounding boxes of two adjacent sample characters is less than a preset area range threshold;
[0036] Mark the corresponding character category value according to the sample character included in the bounding box to obtain the annotated sample image.
[0037] The present specification also provides a recognition device for text strings, including:
[0038] An acquisition module, configured to acquire a target image including a target string to be recognized;
[0039] A preprocessing module, configured to preprocess the target image to obtain a preprocessed target image;
[0040] An invocation module, configured to invoke a preset recognition model to process the preprocessed target image to obtain a corresponding target processing result; wherein, the preset recognition model at least includes a preset dilated convolutional layer; the preset dilated convolutional layer is used to extract target image features related to the target string from the preprocessed target image instead of a combination of a convolutional network layer and a pooling layer;
[0041] A determination module, configured to determine the target string in the target image according to the target processing result.
[0042] The present specification also provides a recognition method for text strings, including:
[0043] Acquire a target image including a target string to be recognized;
[0044] Invoke a preset recognition model to process the target image to obtain a corresponding target processing result; wherein, the preset recognition model at least includes a preset dilated convolutional layer; the preset dilated convolutional layer is used to extract target image features related to the target string from the target image instead of a combination of a convolutional network layer and a pooling layer;
[0045] Determine the target string in the target image according to the target processing result.
[0046] The present specification also provides a server, including a processor and a memory for storing processor-executable instructions, and when the processor executes the instructions, the related steps of the recognition method for text strings are implemented.
[0047] The present specification also provides a computer-readable storage medium, on which computer instructions are stored, and when the instructions are executed, the related steps of the recognition method for text strings are implemented.
[0048] The text string recognition method, device, and server provided in this specification have made targeted improvements to the model structure of the recognition model used to recognize and extract text strings in images before specific implementation: using a preset dilated convolutional layer to replace the combination of a convolutional network layer and a pooling layer as the feature extraction structure for extracting image features related to text strings, and obtaining an improved preset recognition model with better effects. During specific implementation, after preprocessing the acquired target image to obtain the preprocessed target image, the above-mentioned preset recognition model can be called to process the preprocessed target image, so that it can be better applied to the situation with higher recognition difficulty where the character size of the string in the image is small, the resolution is low, and relatively few relevant image features can be extracted, accurately recognize and determine the target string contained in the target image, improve the recognition accuracy and efficiency of the text string, and reduce the recognition error. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] To more clearly illustrate the embodiments of this specification, the accompanying drawings required for use in the embodiments will be briefly introduced below. The accompanying drawings in the following description are only some embodiments recorded in this specification. For those of ordinary skill in the art, other accompanying drawings can be obtained based on these drawings without creative efforts.
[0050] Figure 1 FIG. is a schematic diagram of an embodiment of the structure of a system applying the text string recognition method provided by the embodiments of this specification;
[0051] Figure 2 FIG. is a schematic diagram of an embodiment of applying the text string recognition method provided by the embodiments of this specification in a scenario example;
[0052] Figure 3 FIG. is a schematic diagram of an embodiment of applying the text string recognition method provided by the embodiments of this specification in a scenario example;
[0053] Figure 4 FIG. is a flowchart of the text string recognition method provided by an embodiment of this specification;
[0054] Figure 5 FIG. is a schematic diagram of an embodiment of applying the text string recognition method provided by the embodiments of this specification in a scenario example;
[0055] Figure 6 FIG. is a flowchart of the text string recognition method provided by an embodiment of this specification;
[0056] Figure 7 FIG. is a schematic diagram of the structure of a server provided by an embodiment of this specification;
[0057] Figure 8 It is a schematic structural composition diagram of an identification device for text strings provided by an embodiment of this specification;
[0058] Figure 9 It is a schematic diagram of an embodiment of applying the text string identification method provided by the embodiments of this specification in a scenario example. Detailed implementation manners
[0059] In order to enable those skilled in the art to better understand the technical solutions in this specification, the following will clearly and completely describe the technical solutions in the embodiments of this specification with reference to the accompanying drawings in the embodiments of this specification. Obviously, the described embodiments are only a part of the embodiments of this specification, rather than all the embodiments. Based on the embodiments in this specification, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of this specification.
[0060] The embodiments of this specification provide a text string identification method, and the text string identification method can be specifically applied to a system including a server and a terminal device. Specifically, reference can be made to Figure 1 As shown, the server and the terminal device can be connected by wired or wireless means for specific data interaction.
[0061] In this embodiment, the server can specifically include a background cloud server applied on one side of the network platform, which can implement functions such as data transmission and data processing. Specifically, the server can be, for example, an electronic device with data operation, storage functions, and network interaction functions. Or, the server can also be a software program running in the electronic device, providing support for data processing, storage, and network interaction. In this embodiment, the number of the servers is not specifically limited. The server can specifically be one server, or several servers, or a server cluster formed by several servers.
[0062] In this embodiment, the terminal device can specifically include a front-end electronic device disposed on the user side, configured or connected with a camera, and capable of implementing functions such as picture data acquisition and data transmission. Specifically, the terminal device can be, for example, a surveillance camera, a desktop computer, a tablet computer, a laptop computer, a smart phone, etc. Or, the terminal device can also be a software application that can run in the above-mentioned electronic devices. For example, it can be a certain APP running on a smart phone.
[0063] In specific implementation, the user can use a smart phone as a terminal device to capture a target string to be recognized on a target object (for example, the serial number on a target currency), and collect a photo containing the target string as a target image. Reference can be made to Figure 2 as shown. Among them, the above-mentioned target string may specifically include one or more text characters.
[0064] After the terminal device collects the above-mentioned target image, it can send the target image to the server in a wired or wireless manner. Correspondingly, the server receives and obtains the target image from the terminal device.
[0065] The server can first preprocess the target image to obtain a preprocessed target image that is more suitable for subsequent recognition and extraction of the text string.
[0066] When specifically performing preprocessing, the server can first perform noise reduction processing on the target image to initially filter out the image noise in the target image and obtain a target image with less noise and relatively purer.
[0067] Next, the server can call a pre-trained preset text character region recognition model to process the target image, so as to first find an image region containing the target string to be recognized in the target image as the target image region. The server then crops out the target image region from the above-mentioned target image to obtain a preprocessed target image with a relatively smaller data volume. Reference can be made to Figure 2 as shown.
[0068] It should be noted that for target strings such as the serial number on a target currency, compared with conventional strings, the text characters in the above-mentioned target string are often smaller in size, the feature information that can be extracted is relatively less, and the resolution is poor; and because target objects such as target banknotes are often used more frequently, most of the image regions where the target strings are located in the collected target images will also have image noise formed by factors such as creases and stains, which interfere with the recognition of the target strings.
[0069] If a conventional recognition model is directly used to recognize and extract the target string in the above situation, the obtained target string often has a large error and relatively poor accuracy.
[0070] Noting the above problems, when specifically performing the recognition and extraction of the target string, the server uses a preset recognition model with a different model structure from the conventional recognition model and improved.
[0071] The server can use the pre - processed target image as the model input, input it into the above - mentioned preset recognition model, and run the preset recognition model to obtain the corresponding model output as the corresponding target processing result.
[0072] Among them, the above - mentioned preset recognition model is different from the conventional recognition model. It uses a preset dilated convolutional layer to replace the combination of the convolutional network layer and the pooling layer used in the conventional recognition model.
[0073] When the preset recognition model runs specifically, it can finely and comprehensively extract target image features related to the target string and with a large receptive field from the target image through the above - mentioned preset dilated convolutional layer, avoiding feature loss caused by the pooling effect of the pooling layer.
[0074] The above - mentioned preset recognition model can also be integrated with a localization sub - model and a classification sub - model at the same time. Among them, the localization sub - model is connected to the above - mentioned preset dilated convolutional layer, and the classification sub - model is connected to the above - mentioned preset dilated convolutional layer.
[0075] When the preset recognition model runs specifically, it can execute the localization process for each text character in the target string through the above - mentioned localization sub - model. Specifically, the localization sub - model can receive the target image features output from the preset dilated convolutional layer; and according to the above - mentioned target image features, combined with the preset anchor box parameters, through anchor regression, generate multiple corresponding candidate boxes for each text character in the target string; further, the localization sub - model can screen out a candidate box that meets the requirements from the multiple candidate boxes corresponding to each text character as the bounding box of the text character, and delete the remaining redundant candidate boxes. Among them, the above - mentioned bounding box can carry the position information of the contained text character. For example, position information such as the arrangement serial number of the contained text character in the target string.
[0076] While the preset recognition model executes the localization process through the localization sub - model in the above - mentioned manner, it can also execute the classification process for each text character in the target string through the classification sub - model. Specifically, the classification sub - model can receive the target image features output from the preset dilated convolutional layer; and according to the above - mentioned target image features, through logistic regression, identify and determine the class values of each text character in the target string.
[0077] Since the above - mentioned localization sub - model and classification sub - model are both integrated in the same preset recognition model and are both connected to the same preset dilated convolutional layer. Therefore, when the preset recognition model runs specifically, the above - mentioned localization process and classification process involved are executed simultaneously.
[0078] On the one hand, this can avoid the increase in processing time and the decrease in processing efficiency caused by separately and sequentially executing the localization process (including splitting the bounding box) and the classification process as in existing methods and models. On the other hand, it can also avoid the cumulative loss of precision layer by layer during the execution of different processes and the impact on the precision of the final result due to the separation and sequential execution of the localization process and the classification process as in existing methods and models.
[0079] The preset recognition model operates in the above manner. By utilizing the above-mentioned preset dilated convolutional layer, localization sub-model, and classification sub-model, it can finally output the category value list of text characters arranged in order based on the position information carried by the bounding box as the target processing result. Refer to Figure 3 as shown.
[0080] The server can obtain a target string with high precision and small error recognized and extracted from the target image according to the above target processing result.
[0081] Furthermore, the server can perform further data processing on the target string according to the extracted target string in combination with the specific application scenario.
[0082] For example, in the scenario of transaction risk detection, the server can track and determine the transaction flow path of the target currency according to the serial number on the target currency recognized and extracted. Subsequently, it can analyze whether there are transaction risks such as money laundering and gambling in the transaction behavior involved in the target banknote according to the transaction flow path of the target currency. Thus, it can detect the transaction risks of transaction behaviors more efficiently and intelligently.
[0083] Through the above system, by using the improved preset recognition model, it can effectively be applicable to the situation where the string characters in the image are small in size and low in resolution, and relatively few relevant image features can be extracted, accurately recognize and determine the target string contained in the target image, improve the recognition precision and efficiency of the text string, and reduce the recognition error.
[0084] Refer to Figure 4 As shown, the embodiment of the present specification provides a method for recognizing a text string. Among them, this method is specifically applied to the server side. Specifically in implementation, this method may include the following content.
[0085] S401: Obtain a target image containing a target string to be recognized;
[0086] S402: Preprocess the target image to obtain a preprocessed target image;
[0087] S403: calling a preset recognition model to process the preprocessed target image to obtain a corresponding target processing result; wherein the preset recognition model at least includes a preset hole convolution layer; the preset hole convolution layer is used to replace the combination of the convolution network layer and the pooling layer to extract target image features related to the target string from the preprocessed target image;
[0088] S404: Determine a target character string in the target image according to the target processing result.
[0089] Through the above embodiments, a preset recognition model that uses a preset hole convolution layer to replace the combination of a convolutional network layer and a pooling layer in a conventional recognition model can be utilized, which is effectively applicable to situations where the character string in the image is small in size, the resolution is low, and relatively few relevant image features can be extracted. The target character string contained in the target image can be accurately identified and determined, thereby improving the recognition accuracy and efficiency of the text character string and reducing recognition errors.
[0090] In some embodiments, the target image may be an image containing a target string to be identified. Specifically, the target image may be obtained by taking a photo containing the target string, or by capturing a screenshot containing the target string from a video.
[0091] The target character string may specifically be a text character string to be identified and extracted, wherein the target character string may contain only one text character, or may contain multiple text characters arranged in sequence.
[0092] In some embodiments, the above-mentioned target character string can specifically be a text character string with low recognition difficulty, and the extracted text character string (which can be recorded as a first-class text character string) can be more accurately recognized based on a conventional recognition model, for example, a text character string with a larger image Chinese character size, higher resolution, and larger character spacing.
[0093] The target character string mentioned above may also be a text character string that is difficult to recognize, and the conventional recognition model often cannot accurately recognize the extracted text character string (which can be recorded as the second type of character string), for example, a text character string with a small image Chinese character size, low resolution, and small character spacing. For this type of target character string, since the image features that can be extracted based on the conventional recognition model are relatively few, there is feature loss, and the receptive field is relatively limited, resulting in poor recognition accuracy. In addition, the character spacing between adjacent characters in the target character string is small, which further increases the difficulty of recognition, resulting in errors such as missing text characters in the character string and misaligned character recognition when using a conventional recognition model for recognition.
[0094] In some embodiments, the target string to be recognized may specifically include at least one of the following: the serial number on the target currency, the drawer account number on the target check, the logistics number on the target express bill, etc. Among them, the above serial number may specifically refer to a string composed of multiple numbers and letters set on a currency (for example, a banknote). Usually, one serial number corresponds to one currency with this serial number set on it.
[0095] The serial number on the target currency, the drawer account number on the target check, and the logistics number on the target express bill listed above all belong to the second type of strings with relatively high recognition difficulty. Usually, the conventional recognition models used are often difficult to accurately and quickly recognize and extract the above target strings from the target image.
[0096] Of course, the above-listed target strings are only illustrative. In specific implementation, according to specific application scenarios and processing requirements, other types of text strings can also be introduced as the target strings to be recognized. This specification does not limit this.
[0097] Through the above embodiments, the recognition method of the text string provided in this specification can be applied to a variety of different business scenarios to accurately recognize text strings with relatively high recognition difficulty, such as the serial number on the currency, the drawer account number on the target check, and the logistics number on the target express bill.
[0098] In some embodiments, the above preset recognition model can be specifically understood as a pre-trained neural network model that can relatively accurately recognize and extract the target string from an image. Among them, the above preset recognition model is different from the conventional recognition and has an improved model structure.
[0099] Specifically, the above preset recognition model at least includes a preset dilated convolutional layer. In the preset recognition model, the above preset dilated convolutional layer is used to replace the combination of the conventional convolutional network layer and the pooling layer. Through the above preset dilated convolutional layer, while extracting the target image features with a larger receptive field, the loss of feature information can be avoided, so that relatively comprehensive, complete, and better-performing target image features can be obtained.
[0100] It should be noted that based on the conventional recognition model, in order to extract relatively better-performing image features, after extracting the corresponding image features from the image using the convolutional network layer, a pooling layer is also used to perform a pooling operation on the extracted image features to achieve the effect of increasing the receptive field.
[0101] However, in the process of performing pooling operations on image features using a pooling layer, some feature information in the image state will be filtered out simultaneously, resulting in the loss of detailed features between characters, etc., making the finally obtained image features incomplete and missing. For the case where the string characters in the image are small in size and low in resolution, and relatively few relevant image features can be extracted, using a conventional recognition model for the above processing will make the finally obtained image features even fewer, thereby resulting in a worse accuracy in recognizing the string.
[0102] In this embodiment, by introducing a preset dilated convolution layer in the preset recognition model to replace the combination of a conventional convolutional network layer and a pooling layer, it is possible to extract image features with a larger receptive field and better performance while avoiding the loss of features and ensuring the integrity of the extracted image features.
[0103] In some embodiments, the above-mentioned dilated convolution layer (Dilated Convolution) specifically refers to injecting holes into the standard Convolution Map to increase the reception field. Compared with the original normal Convolution (for example, the convolutional network layer), the dilated convolution layer has an additional hyper-parameter, which can be called the dilation rate, specifically referring to the number of intervals of the kernel.
[0104] In some embodiments, the above-mentioned preset dilated convolution layer can be specifically configured with a preset dilation coefficient and a corresponding convolution kernel. Among them, the above-mentioned preset dilation coefficient and the size parameters of the convolution kernel can be flexibly set according to information such as the proportion and resolution of the target string in the target image.
[0105] When specifically running the preset recognition model to process the above-mentioned preprocessed target image, the preset dilation coefficient can be used to perform dilation processing on the initial image feature matrix obtained after the convolution kernel performs normal convolution operations on the preprocessed target image, and use the data value 0 to fill the newly added matrix area after dilation, so as to obtain a dilated image feature matrix containing feature information with a relatively large visual field range as the target image feature. This can effectively increase the receptive field of the obtained target image features while not losing the feature information originally extracted by the convolution kernel.
[0106] Specifically, reference can be made to Figure 5As shown, the preset dilated convolutional layer used is configured with a preset dilation coefficient with a data value of 1 and a 3×3 convolutional kernel. When processing the preprocessed target image using the preset dilated convolutional layer, first, a convolutional operation can be performed on the preprocessed target image through the 3×3 convolutional kernel to extract the initial 3×3 image feature matrix shown on the left. Then, the above initial image feature matrix can be dilated using the preset dilation coefficient (1) to obtain an expanded 5×5 matrix; and the newly added matrix area after expansion can be filled with a data value of 0, so that the dilated 5×5 image feature matrix shown on the right can be obtained as the target image feature.
[0107] In some embodiments, the preset recognition model may specifically further include a localization sub-model; wherein, the localization sub-model is connected to the preset dilated convolutional layer, and the localization sub-model is used to generate a corresponding plurality of candidate boxes for each text character in the target string through anchor regression according to the target image feature and the preset anchor box parameters; and for each text character, a candidate box that meets the requirements is selected from the corresponding plurality of candidate boxes as the bounding box of the text character; the bounding box carries the position information of the contained text character.
[0108] Through the above embodiments, when running the preset recognition model, the bounding boxes of each text character in the target string can be accurately located first by using the internally integrated localization sub-model, and each text character can be accurately cut based on the bounding boxes.
[0109] In some embodiments, the preset anchor box parameters can be specifically obtained in the following manner:
[0110] S1: Obtain a sample image containing a sample string;
[0111] S2: According to the preset annotation rules, in the sample image, corresponding bounding boxes are respectively annotated for each sample character in the sample string; and the bounding box parameters of the sample characters are collected; wherein, the overlapping area range between the bounding boxes of two adjacent sample characters is less than the preset area range threshold;
[0112] S3: Perform clustering processing on the bounding box parameters of the sample characters to obtain the preset anchor box parameters.
[0113] Through the above embodiments, preset anchor box parameters with better effects and greater suitability can be obtained and utilized for anchor regression, thereby accelerating the convergence speed of the network model, improving the calculation efficiency of the model, and generating corresponding multiple candidate boxes for each text character more accurately and efficiently.
[0114] In some embodiments, the above-mentioned preset anchor box parameters can specifically be understood as the parameter values of an anchor (Anchor).
[0115] In some embodiments, the clustering process for the bounding box parameters of the sample characters may specifically include: performing a clustering process on the bounding box parameters of the sample characters based on the K-means clustering algorithm.
[0116] It should be noted that based on the existing method, usually the above-mentioned anchor box parameters are manually set fixedly. As a result, the anchor box parameters used in anchor regression do not match the actual target string, thus affecting the result of anchor regression.
[0117] In some embodiments, the above-mentioned process of screening out a qualified candidate box from the corresponding multiple candidate boxes for each text character as the bounding box of the text character may specifically include the following when implemented: screening out a qualified candidate box from the corresponding multiple candidate boxes for the current text character in the target string as the bounding box in the following manner: calling a preset softened non-maximum suppression algorithm to process the multiple candidate boxes to screen out a candidate box with a confidence level meeting the requirements as the bounding box of the current text character; and filtering out other candidate boxes among the multiple candidate boxes except the bounding box.
[0118] Through the above embodiments, the softened non-maximum suppression algorithm can be well applied to the recognition scenario of text strings with small character intervals, and can effectively avoid the problem that the finally extracted target string is incomplete and missing due to the misdeletion of candidate boxes of adjacent other text characters.
[0119] In some embodiments, it should be noted that based on the existing method, usually the non-maximum suppression algorithm is called to determine the corresponding bounding box from multiple candidate boxes. Specifically, when implementing based on the non-maximum suppression algorithm, the candidate box of the character with the highest confidence is first selected as the reference box. If there is a candidate box overlapping with it, the ratio of the overlapping area of the two to the total area is calculated. If this ratio is greater than the set threshold, the confidence of this candidate box is directly set to 0.
[0120] In this embodiment, the above-mentioned process of calling a preset softened non-maximum suppression algorithm to process the multiple candidate boxes may specifically include: first selecting the candidate box of the character with the highest confidence as the reference box. If there is a candidate box overlapping with it, the ratio of the overlapping area of the two to the total area is calculated. If this ratio is greater than the set threshold, a preset linear function is used to modify and adjust the confidence of this candidate box, rather than directly setting it to 0 as in the non-maximum suppression algorithm. This can effectively avoid the misdeletion of candidate boxes of adjacent other text characters when the intervals between text characters are relatively close.
[0121] The above-mentioned preset linear function can be specifically expressed in the following form:
[0122]
[0123] where S i represents the confidence of the i-th candidate box, and IoU i represents the proportion of the overlapping area between the candidate box and the reference box in the total area, and T represents the set threshold. The specific value of the above-mentioned set threshold can be determined according to the minimum character interval between adjacent characters in the target string.
[0124] In some embodiments, the preset recognition model may specifically further include a classification sub-model; wherein, the classification sub-model is connected to the dilated convolutional layer, and the classification sub-model is used to identify and determine the class values of each text character in the target string to be recognized according to the target image features.
[0125] Through the above embodiments, when running the preset recognition model, the class values of each text character in the target string can be accurately recognized by using the internally integrated classification sub-model.
[0126] In some embodiments, in specific implementation, the above classification sub-model can identify and determine the class values of each text character through logistic regression according to the target image features.
[0127] In some embodiments, when the target string includes the serial number on the target currency, the character class values may specifically include: 0-9 and / or A-Z, etc.
[0128] In some embodiments, the above-mentioned preset recognition model can be integrated with a localization sub-model and a classification sub-model at the same time. Correspondingly, when the preset recognition model is specifically running, the above-mentioned localization sub-model and classification sub-model can be used to simultaneously execute the localization process and the classification process according to the target image features extracted by the preset dilated convolutional layer, so as to be able to efficiently segment the target string into multiple bounding boxes that are sequentially connected and each contains a text character, and at the same time, identify the class values of the text characters in each bounding box, so that the class values of multiple text characters arranged in order based on the position information carried by the bounding boxes can be obtained as the target processing result output by the preset recognition model.
[0129] By utilizing the preset recognition model that integrates the positioning sub-model and the classification sub-model simultaneously as described above, the positioning process and the classification process can be executed simultaneously, avoiding the cumulative loss of accuracy that occurs when the above two processes are executed separately in sequence, and improving the accuracy of the obtained target processing result. At the same time, since the above two processes are executed simultaneously, the processing efficiency of the model is also improved.
[0130] In some embodiments, when specifically implementing the above preprocessing of the target image, it may include the following: Detect the target image and determine a target image area in the target image that contains the target string to be recognized; Crop out the target image area from the target image as the preprocessed target image.
[0131] Through the above embodiments, a relatively small target image area containing the target string can be initially located from the target image, and then an image containing only the above target image area can be cropped out from the target image as the input value of the preprocessed target image into the preset recognition model for subsequent recognition processing. Thereby, the data processing amount of the subsequent preset recognition model can be reduced, and the recognition efficiency and recognition accuracy can be improved.
[0132] In some embodiments, when specifically implementing, the server can call a pre-trained preset text character area recognition model to process the target image, so as to be able to detect and find out the target image area containing the target string in the target image more quickly and accurately.
[0133] In some embodiments, when specifically implementing the above preprocessing of the target image, it may further include: performing image correction processing on the target image; and / or, performing noise reduction processing on the target image.
[0134] Through the above embodiments, the influence of interference factors such as image noise in the target image on subsequent string recognition can be reduced, which helps to improve the accuracy of subsequent string recognition.
[0135] In some embodiments, when the target string to be recognized includes the serial number on the target currency, after determining the target string in the target image, when specifically implementing the method, it may further include the following:
[0136] S1: Determine the target string as the serial number on the target currency;
[0137] S2: According to the serial number on the target currency, track and determine the transaction flow path of the target currency;
[0138] S3: According to the transaction flow path of the target currency, determine whether there is a transaction risk.
[0139] Through the above embodiments, subsequent specific data processing can be performed using the identified and extracted target string according to specific situations and processing requirements.
[0140] In some embodiments, after determining the target string as the serial number on the target currency, the method may further include: determining the authenticity of the target banknote according to the serial number.
[0141] In some embodiments, when the target string to be recognized includes the logistics number on the target express bill, after determining the target string in the target image, the specific implementation of the method may further include the following: determining the target string as the logistics number on the target express bill; tracking the logistics of the package or mail with the target express bill set according to the logistics number, and timely feedback the latest logistics information to the user.
[0142] In some embodiments, before the specific implementation of the method, the following may further be included:
[0143] S1: Use a preset dilated convolutional layer to replace the combination of the convolutional network layer and the pooling layer as the extraction structure of the image features in the network model to construct an initial recognition model;
[0144] S2: Obtain a sample image containing the sample string to be recognized; and perform annotation processing on the sample image to obtain the annotated sample image;
[0145] S3: Use the annotated sample image to train the initial recognition model to obtain a preset recognition model.
[0146] Through the above embodiments, a preset recognition model for string recognition with high recognition difficulty in the case where the character size of the string in the image is small, the resolution is low, and relatively few relevant image features can be extracted can be pre-constructed and trained.
[0147] In some embodiments, in specific implementation, on the basis of the network structure of Tiny-YOLOv2, a preset dilated convolutional layer can be used to replace the combination of the convolutional network layer and the pooling layer as the feature extraction structure in the model to obtain an initial recognition model.
[0148] In some embodiments, the above-mentioned annotation processing of the sample image to obtain the annotated sample image may specifically include the following when implemented: according to a preset annotation rule, in the sample image, boundary boxes corresponding to each sample character in the sample string are respectively annotated; wherein, the overlapping area range between the boundary boxes of two adjacent sample characters is less than a preset area range threshold; according to the sample characters included in the boundary boxes, corresponding character category values are annotated to obtain the annotated sample image.
[0149] Through the above embodiments, the sample image can be annotated more effectively and accurately, and an annotated sample image with relatively good training effect can be obtained.
[0150] As can be seen from the above, before the specific implementation of the text string recognition method provided by the embodiments of this specification, the model structure of the recognition model used to recognize and extract the text string in the image is specifically improved: a preset dilated convolutional layer is used to replace the combination of the convolutional network layer and the pooling layer as the feature extraction structure for extracting image features related to the text string, and a preset improved recognition model with better effect is obtained. When specifically implemented, after preprocessing the obtained target image, the above-mentioned preset recognition model can be called to process the preprocessed target image, so that it can be effectively applied to the situation where the string characters in the image are small in size, low in resolution, and relatively few relevant image features can be extracted, accurately recognize and determine the target string included in the target image, improve the recognition accuracy and recognition efficiency of the text string, and reduce the recognition error. Also, by using the preset anchor box parameters obtained based on clustering processing in advance for specific anchor regression, the generated candidate boxes are relatively more accurate and reasonable, and the processing accuracy when determining the boundary boxes subsequently is improved. Also, by introducing and using the preset soft non-maximum suppression algorithm to process the multiple candidate boxes corresponding to each text character to screen out the corresponding boundary boxes, it effectively avoids the problem that when the intervals between text characters are relatively close, the candidate boxes of adjacent other text characters are mistakenly deleted, resulting in the incomplete and missing target string finally extracted.
[0151] Refer to Figure 6 As shown, the embodiments of this specification also provide another text string recognition method. When specifically implemented, the method may include the following:
[0152] S601: Obtain a target image containing a target string to be recognized;
[0153] S602: Process the target image using a preset recognition model to obtain a corresponding target processing result; wherein, the preset recognition model at least includes a preset dilated convolutional layer; the preset dilated convolutional layer is used to extract target image features related to the target string from the target image instead of the combination of a convolutional network layer and a pooling layer;
[0154] S603: Determine the target string in the target image according to the target processing result.
[0155] Through the above embodiments, a target string can be directly recognized from a target image more efficiently using a preset recognition model.
[0156] This specification also provides a method for recognizing text characters, including: obtaining a target image containing a target character to be recognized; processing the target image using a preset recognition model to obtain a corresponding target processing result; wherein, the preset recognition model at least includes a preset dilated convolutional layer; the preset dilated convolutional layer is used to extract target image features related to the target character from the target image instead of the combination of a convolutional network layer and a pooling layer; determining the target character in the target image according to the target processing result.
[0157] An embodiment of this specification also provides a server, including a processor and a memory for storing instructions executable by the processor. When specifically implemented, the processor may execute the following steps according to the instructions: obtaining a target image containing a target string to be recognized; preprocessing the target image to obtain a preprocessed target image; processing the preprocessed target image using a preset recognition model to obtain a corresponding target processing result; wherein, the preset recognition model at least includes a preset dilated convolutional layer; the preset dilated convolutional layer is used to extract target image features related to the target string from the preprocessed target image instead of the combination of a convolutional network layer and a pooling layer; determining the target string in the target image according to the target processing result.
[0158] To be able to complete the above instructions more accurately, refer to Figure 7 As shown, an embodiment of this specification also provides another specific server. The server includes a network communication port 701, a processor 702, and a memory 703. The above structures are connected by internal cables so that each structure can perform specific data interactions.
[0159] Among them, the network communication port 701 can specifically be used to obtain a target image containing a target string to be recognized.
[0160] The processor 702 can be specifically used to preprocess the target image to obtain a preprocessed target image; call a preset recognition model to process the preprocessed target image to obtain a corresponding target processing result; wherein, the preset recognition model at least includes a preset dilated convolutional layer; the preset dilated convolutional layer is used to extract target image features related to the target string from the preprocessed target image instead of a combination of a convolutional network layer and a pooling layer; and determine the target string in the target image according to the target processing result.
[0161] The memory 703 can be specifically used to store corresponding instruction programs.
[0162] In this embodiment, the network communication port 701 can be bound to different communication protocols, so as to send or receive different data. For example, the network communication port can be a port responsible for web data communication, can also be a port responsible for FTP data communication, and can also be a port responsible for mail data communication. In addition, the network communication port can also be a physical communication interface or a communication chip. For example, it can be a wireless mobile network communication chip, such as GSM, CDMA, etc.; it can also be a Wifi chip; it can also be a Bluetooth chip.
[0163] In this embodiment, the processor 702 can be implemented in any suitable manner. For example, the processor can take the form of, for example, a microprocessor or a processor and a computer-readable medium storing computer-readable program code (such as software or firmware) executable by the (micro)processor, logic gates, switches, an application specific integrated circuit (ASIC), a programmable logic controller, and an embedded microcontroller, etc. This specification does not make a limitation.
[0164] In this embodiment, the memory 703 can include multiple levels. In a digital system, anything that can store binary data can be a memory; in an integrated circuit, a circuit with a storage function without a physical form is also called a memory, such as RAM, FIFO, etc.; in a system, a storage device with a physical form is also called a memory, such as a memory stick, a TF card, etc.
[0165] The embodiments of this specification also provide a computer storage medium for an identification method based on the above text string. The computer storage medium stores computer program instructions, which when executed implement the following: obtaining a target image containing a target string to be identified; preprocessing the target image to obtain a preprocessed target image; calling a preset identification model to process the preprocessed target image to obtain a corresponding target processing result; wherein the preset identification model at least includes a preset dilated convolutional layer; the preset dilated convolutional layer is used to extract target image features related to the target string from the preprocessed target image instead of a combination of a convolutional network layer and a pooling layer; determining the target string in the target image according to the target processing result.
[0166] In this embodiment, the above storage medium includes but is not limited to a random access memory (RAM), a read-only memory (ROM), a cache, a hard disk drive (HDD), or a memory card. The memory can be used to store computer program instructions. The network communication unit can be set according to the standards specified by the communication protocol and is used for the interface of network connection communication.
[0167] In this embodiment, the functions and effects specifically implemented by the program instructions stored in this computer storage medium can be explained by comparison with other embodiments and will not be elaborated here.
[0168] The embodiments of this specification also provide another computer storage medium for an identification method based on the above text string. The computer storage medium stores computer program instructions, which when executed implement the following: obtaining a target image containing a target string to be identified; calling a preset identification model to process the target image to obtain a corresponding target processing result; wherein the preset identification model at least includes a preset dilated convolutional layer; the preset dilated convolutional layer is used to extract target image features related to the target string from the target image instead of a combination of a convolutional network layer and a pooling layer; determining the target string in the target image according to the target processing result.
[0169] Refer to Figure 8 As shown, at the software level, the embodiments of this specification also provide an identification device for a text string. The device can specifically include the following structural modules:
[0170] An obtaining module 801, which can specifically be used to obtain a target image containing a target string to be identified;
[0171] A preprocessing module 802 can be specifically used to preprocess the target image to obtain a preprocessed target image;
[0172] An invocation module 803 can be specifically used to invoke a preset recognition model to process the preprocessed target image to obtain a corresponding target processing result; wherein, the preset recognition model at least includes a preset dilated convolutional layer; the preset dilated convolutional layer is used to extract target image features related to the target string from the preprocessed target image instead of a combination of a convolutional network layer and a pooling layer;
[0173] A determination module 804 can be specifically used to determine the target string in the target image according to the target processing result.
[0174] It should be noted that the units, devices, or modules, etc. described in the above embodiments can be specifically implemented by computer chips or entities, or by products with certain functions. For the convenience of description, when describing the above devices, they are divided into various modules according to functions and described separately. Of course, when implementing this specification, the functions of each module can be implemented in the same or multiple software and / or hardware, or the modules implementing the same function can be implemented by a combination of multiple sub-modules or sub-units, etc. The device embodiments described above are only illustrative. For example, the division of the units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of the devices or units can be in electrical, mechanical or other forms.
[0175] As can be seen from the above, based on the text string recognition device provided in the embodiments of this specification, it can be well applicable to the situation where the character size of the string in the image is small, the resolution is low, and relatively few relevant image features can be extracted, accurately recognize and determine the target string contained in the target image, improve the recognition accuracy and efficiency of the text string, and reduce the recognition error.
[0176] In a specific scenario example, the text string recognition method provided in this specification can be applied to accurately recognize the serial number on banknotes. The specific implementation process can refer to the following content.
[0177] In this scenario example, considering that the area of the serial number on the banknote (the target string to be recognized) is roughly fixed on the banknote, the area where the serial number of the banknote is located can be roughly determined according to the known prior information. If the existing recognition method is adopted, generally three steps are required: precise positioning, single-character segmentation, and character recognition (equivalent to the positioning process and the classification process). And each step is an independent process, resulting in isolation between different steps. Inevitably, there will be a certain loss of accuracy in the process of sequential execution of each step. The errors introduced by the accuracy loss in the three steps will accumulate layer by layer in the whole process and finally act on the recognition accuracy, resulting in poor recognition accuracy. Therefore, it is considered that the error influence between different tasks can be reduced, the three steps can be integrated, and an end-to-end serial number recognition network (the corresponding preset recognition model) is proposed to be constructed and trained.
[0178] Furthermore, considering that most of the existing text detection network models (i.e., conventional recognition models) use training data sets that are mostly sample images with many picture pixels, less noise, or relatively regular high-definition and less noisy ones. However, during the circulation of banknotes, many creases and stains will occur, and these noises will be randomly distributed in the serial number area, which will have a certain impact on the accuracy of serial number recognition. At the same time, due to the small area of the serial number characters, the serial number character pictures have the characteristics of high noise and low resolution, making the recognition difficult.
[0179] In addition, if a deep learning model with a deeper network layer is used, although it can meet the accuracy requirements, it often requires sacrificing time or depends on the performance of the device. However, in the scenario of banknote serial number recognition, there is no powerful performance device, and at the same time, there are high requirements for time. Therefore, most of the existing network structures cannot meet the actual business needs.
[0180] Based on the above considerations, in order to solve the limitations brought by the recognition method that the traditional character recognition process is divided into three steps, and the problem that the accuracy and timeliness of the existing deep learning network structure cannot be satisfied at the same time, and obtain the correct serial number character sequence, specifically, the recognition method of the text string provided in this specification can be combined, and an end-to-end serial number character recognition method based on Tiny-YOLOv2 for banknote serial numbers is further proposed.
[0181] The specific implementation of this method can include: First, the image with a relatively large area and rough positioning of the serial number area obtained after preprocessing (for example, the preprocessed target image) is used as the input and sent into the serial number recognition network for feature extraction. The pooling operation of the traditional convolution layer of this part of the recognition network is removed and replaced with dilated convolution (for example, the preset dilated convolution layer).
[0182] Next, according to the extracted feature properties, multiple different possible candidate character boxes (e.g., candidate boxes) can be generated, and the characters can be initially classified into specific categories. Then, redundant boxes are removed through the Soft-NMS (Soft-Non-Maximum Suppression) algorithm, so that each character finally outputs only one prediction box (e.g., bounding box) with the highest confidence. Then, the specific coordinate information of a single character is determined through anchor regression.
[0183] Finally, according to the arrangement characteristics of the serial number and the coordinate information of each obtained character, the serial number is arranged in sequence to obtain the final recognition result (e.g., the target processing result).
[0184] Considering that in the image containing the banknote serial number, the character resolution of the serial number is relatively low, and fewer features of the serial number can be extracted. At the same time, relatively high requirements are usually placed on the recognition speed and accuracy. Therefore, the entire recognition network structure needs to adopt a relatively concise network as the basic network of the algorithm, and make algorithm improvements based on this to achieve a balance between the recognition speed and accuracy.
[0185] First, the image containing the banknote can be preprocessed. Specifically, it includes steps such as image rectification, image cropping, and resizing to obtain a roughly located image of the area containing the serial number (e.g., the preprocessed target image); then, the roughly located image is used as the input image and sent into the serial number recognition network for feature extraction; the specific category of each character is predicted, that is, a character between 0-9 or A-Z; the specific coordinates of each character are obtained through anchor regression; redundant boxes are removed through the non-maximum suppression algorithm, and finally each character only outputs one prediction box with the highest confidence; finally, according to the arrangement characteristics of the serial number and the coordinate information of each predicted character, the serial number is arranged in sequence, and the final output is the serial number recognition result. Reference can be made to Figure 9 as shown.
[0186] During specific processing, the following steps can be included.
[0187] Step 101: Extract image features through dilated convolution operations.
[0188] In the structure of a general convolutional neural network, the convolutional layer is used to extract features. After passing through the pooling layer, some features are selected, and at the same time, the effect of increasing the receptive field is achieved. However, because there are few license plate number characters in the image and few features can be extracted, using the pooling layer not only fails to achieve the effect of feature filtering but also loses the detailed features between characters. But if the pooling layer is directly removed, it cannot guarantee the same receptive field. Therefore, dilated convolution is used to maintain the receptive field equivalent to that of pooling without reducing the character feature information.
[0189] In this scenario example, based on dilated convolution, through a dilation coefficient of 1, the convolutional kernel can be dilated to the scale set by the dilation coefficient, and the extra areas after dilation are filled with 0. Therefore, each convolutional kernel can extract a larger range of feature information compared to before. Please refer to Figure 5 As shown, it represents the convolutional kernel of 3*3 after the dilation operation with a dilation coefficient of 1, which is the actual convolutional kernel for convolution operation.
[0190] There is no difference in time consumption between dilated convolution and ordinary convolution. And since dilated convolution does not increase the number of parameters, a smaller convolutional kernel can be used to achieve the previous effect. At the same time, due to the increased receptive field, the pooling layer can be correspondingly reduced, thereby reducing information loss.
[0191] The operations of general convolution plus pooling can be replaced by dilated convolution, and dilated convolution can accelerate the computational efficiency of the convolutional neural network without additional time consumption. For the scenario of license plate number recognition, pooling is a process of reducing features, but due to the low resolution of license plate number characters themselves.
[0192] If pooling is performed multiple times, it will further lose the already scarce feature information of license plate number characters, resulting in insufficient extraction of local feature information of characters, and thus leading to unclear discrimination of similar images.
[0193] The accuracy of license plate number recognition decreases. Therefore, by reducing the pooling layer, the accuracy of license plate number recognition can be improved.
[0194] Step 102: Preset anchor points through k-means clustering.
[0195] The framework of object detection usually presets boxes of different sizes and aspect ratios on the image in advance, and these boxes are called anchors. For a network framework of object detection, setting the anchor boxes to a reasonable value will not only accelerate the network convergence speed but also ensure the final detection effect.
[0196] The values of the anchor points are generally set manually. For example, in the well-known Fast-RCNN, 9 different anchor points are designed. However, there is a problem with these manually designed anchor points, that is, they are not well-suited for our dataset of banknote serial number characters. Therefore, we propose to use the K-means clustering algorithm to automatically generate anchor points suitable for the dataset.
[0197] The clustering algorithm adopted draws on the K-means clustering algorithm, converting the original method of clustering by Euclidean distance into clustering by Intersection-over-Union (IoU), thereby generating initial anchor points and ensuring that the size of the error is independent of the size of the true box.
[0198] Step 103: Use soft non-maximum suppression to filter out redundant bounding boxes.
[0199] After obtaining the bounding boxes through anchor point regression, multiple targets may be detected for the same character category, and there may be multiple overlapping bounding boxes for each character. In the object detection network, non-maximum suppression is usually used to screen out a bounding box with a relatively high confidence. The core idea of the non-maximum suppression algorithm is: only one optimal box is retained for each character. First, select the box of the character with the highest confidence as the benchmark. If there are candidate boxes overlapping with it, calculate the proportion of the overlapping area between the two to the total area. If the proportion is greater than the set threshold, it is considered a redundant candidate box for this character, and the confidence of this box is set to 0, which is equivalent to removing this candidate box. If it is less, it is considered a candidate box for other characters and its confidence is not changed, which is equivalent to retaining this candidate box. This method has a good effect when there are multiple targets in the picture and the targets are relatively far apart.
[0200] However, due to the relatively dense distribution of characters on the banknote serial number and the small interval between characters, when there is noise interference, it is very easy to cause a large overlap between the candidate boxes of two characters. When one character is selected as the benchmark, another character with a high overlap with it is very likely to cause the loss of character detection. It can be seen from Figure 3 As shown, when the character 4 (the 4th character from left to right) is selected as the benchmark, the candidate box of the character 0 (the 5th character from left to right) will be deleted.
[0201] To address this situation, in this scenario example, the method of soft non-maximum suppression is adopted. For candidate boxes with a proportion greater than the threshold, the confidence is not directly set to 0, but is reduced using a linear function. This is equivalent to smoothing the attenuation function of the ordinary non-maximum suppression algorithm, as shown in the following formulas. Formula (1) represents non-maximum suppression, and formula (2) represents the softened one.
[0202]
[0203]
[0204] Among them, S i represents the confidence of the i-th candidate box, and IoU i represents the ratio of the overlapping area of the candidate box and a certain reference box to the total area, and T represents the set threshold. The method of soft non-maximum suppression is adopted, and the candidate boxes of other characters with a relatively high degree of overlap are not directly deleted. It is applicable to images of banknote serial number characters with fewer pixel points and small object spacing, increasing the detection rate and improving the final recognition accuracy at the same time.
[0205] Through the above scenario example, it is verified that the character recognition method provided in this specification can integrate the problems of character segmentation and classification in a network model, and well solve the problem of feature loss caused by separating the two before; by using dilated convolution to extract the features of banknote serial number character images, it can better solve the problem of feature loss caused by pooling process for images of banknote serial number characters with fewer pixels; by using the soft non-maximum suppression algorithm to smooth the original evaluation function, it can better solve the problem that the distribution of banknote serial number characters is relatively dense and the candidate boxes of character detection frames overlap, which is prone to misdeleting the candidate boxes of adjacent characters.
[0206] Although this specification provides method operation steps as described in the embodiments or flowcharts, more or fewer operation steps may be included based on conventional or non-creative means. The step order listed in the embodiments is only one way among the execution orders of numerous steps and does not represent the only execution order. When the actual device or client product is executed, it can be executed in the order of the method shown in the embodiments or the drawings or executed in parallel (for example, in a parallel processor or multi-threaded processing environment, or even in a distributed data processing environment). The term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, product or device comprising a series of elements not only includes those elements but also includes other elements not explicitly listed, or also includes elements inherent to such process, method, product or device. Without further limitation, it does not exclude the existence of additional identical or equivalent elements in the process, method, product or device comprising the said elements. The words such as first, second, etc. are used to represent names and do not represent any specific order.
[0207] Those skilled in the art also know that in addition to implementing the controller in the form of pure computer-readable program code, the method steps can be logically programmed to enable the controller to implement the same functions in the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, embedded microcontrollers, etc. Therefore, such a controller can be regarded as a hardware component, and the devices included therein for implementing various functions can also be regarded as the structures within the hardware component. Or even, the devices for implementing various functions can be regarded as either software modules for implementing the method or structures within the hardware component.
[0208] This specification can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, classes, etc. that perform specific tasks or implement specific abstract data types. This specification can also be practiced in a distributed computing environment where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.
[0209] From the description of the above embodiments, those skilled in the art can clearly understand that this specification can be implemented by means of software plus a necessary general hardware platform. Based on such an understanding, the technical solution of this specification can essentially be embodied in the form of a software product, and this computer software product can be stored in a storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to enable a computer device (which can be a personal computer, mobile terminal, server, or network device, etc.) to execute the methods described in various embodiments or some parts of the embodiments of this specification.
[0210] The various embodiments in this specification are described in a progressive manner. For the same or similar parts among the various embodiments, reference can be made to each other, and the key point of each embodiment is to illustrate the differences from other embodiments. This specification can be used in numerous general or special computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multi-processor systems, microprocessor-based systems, set-top boxes, programmable electronic devices, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and so on.
[0211] Although this specification is depicted through embodiments, those of ordinary skill in the art know that this specification has many variations and changes without departing from the spirit of this specification, and it is hoped that the appended claims will include these variations and changes without departing from the spirit of this specification.
Claims
1. A method for identifying a text string, characterized in that, Including: Obtain a target image including a target string to be recognized; wherein, the character size of the target string in the target image is small, the resolution is low, and there are few relevant image features extracted based on the target image; Preprocess the target image to obtain a preprocessed target image; Call the preset recognition model to process the preprocessed target image to obtain the corresponding target processing result; wherein, the preset recognition model at least includes a preset dilated convolutional layer; the preset dilated convolutional layer is used to extract target image features related to the target string from the preprocessed target image instead of the combination of a convolutional network layer and a pooling layer; the preset recognition model also integrates a localization sub-model and a classification sub-model that are respectively connected to the preset dilated convolutional layer. When running the preset recognition model, the localization process and the classification process are simultaneously executed according to the target image features to obtain the target processing result; during the execution of the localization process, the method further includes: selecting the candidate box of the character with the highest confidence as the reference box, and if there is a candidate box overlapping with the reference box, calculating the proportion of the overlapping area of the two to the total area; if this proportion is greater than the set threshold, modifying and adjusting the confidence of the candidate box using the following preset linear function: wherein, S i represents the confidence of the i-th candidate box, and IoU i represents the proportion of the overlapping area of the candidate box and the reference box to the total area, and T represents the set threshold; Determine the target string in the target image according to the target processing result; Wherein, the preset dilated convolutional layer is configured with a preset dilation coefficient and a corresponding convolutional kernel, and the preset dilation coefficient and the size parameter of the convolutional kernel are set according to the proportion and resolution of the target string in the target image; Correspondingly, use the preset dilated convolutional layer to process the preprocessed target image to extract target image features related to the target string, including: using the preset dilation coefficient to dilate the initial image feature matrix obtained after the convolutional kernel performs a normal convolution operation on the preprocessed target image, and using the data value 0 to fill the newly added matrix area dilated, to obtain a dilated image feature matrix containing feature information in a large field of view and avoiding loss of feature information, as the target image feature.
2. The method according to claim 1, wherein The preset recognition model further includes a localization sub-model; wherein, the localization sub-model is connected to the preset dilated convolutional layer, and the localization sub-model is used to generate a corresponding plurality of candidate boxes for each text character in the target string according to the target image features and the preset anchor box parameters; and screen out a candidate box that meets the requirements from the corresponding plurality of candidate boxes for each text character as the bounding box of the text character; the bounding box carries the position information of the contained text character.
3. The method according to claim 2, wherein The preset anchor box parameters are obtained in the following manner: Obtain a sample image including a sample string; According to the preset annotation rules, in the sample image, label corresponding bounding boxes for each sample character in the sample string; And collect the bounding box parameters of the sample characters; wherein, the overlapping area range between the bounding boxes of two adjacent sample characters is less than a preset area range threshold; Perform clustering processing on the bounding box parameters of the sample characters to obtain the preset anchor box parameters.
4. The method according to claim 2, wherein Screen out a candidate box that meets the requirements from the corresponding plurality of candidate boxes for each text character as the bounding box of the text character, including: Screen out a candidate box that meets the requirements from the corresponding plurality of candidate boxes for the current text character in the target string as the bounding box in the following manner: Call the preset soft non-maximum suppression algorithm to process the plurality of candidate boxes to screen out a candidate box with a confidence level that meets the requirements from the plurality of candidate boxes as the bounding box of the current text character; and filter out other candidate boxes except the bounding box from the plurality of candidate boxes.
5. The method according to claim 2, characterized in that, The preset recognition model further includes a classification sub-model; wherein, the classification sub-model is connected to the dilated convolutional layer, and the classification sub-model is used to identify and determine the class values of each text character in the target string to be recognized according to the target image features.
6. The method according to claim 1, wherein Preprocessing the target image includes: Detect the target image and determine the target image region in the target image that contains the target string to be recognized; Crop out the target image region from the target image as the preprocessed target image.
7. The method according to claim 6, characterized in that, The preprocessing of the target image further includes: Performing image correction processing on the target image; and / or, performing noise reduction processing on the target image.
8. The method according to claim 1, wherein The target string to be recognized includes at least one of the following: the serial number on the target currency, the drawer account number on the target check, and the logistics number on the target express bill.
9. The method according to claim 8, characterized in that, When the target string to be recognized includes the serial number on the target currency, after determining the target string in the target image, the method further includes: Determining the target string as the serial number on the target currency; Tracking and determining the transaction flow path of the target currency according to the serial number on the target currency; Determining whether there is a transaction risk according to the transaction flow path of the target currency.
10. The method according to claim 1, wherein The method further includes: Using a preset dilated convolutional layer to replace the combination of the convolutional network layer and the pooling layer as the extraction structure of the image features in the network model to construct an initial recognition model; Obtaining a sample image containing a sample string to be recognized; and performing annotation processing on the sample image to obtain an annotated sample image; Training the initial recognition model with the annotated sample image to obtain a preset recognition model.
11. The method according to claim 10, wherein Performing annotation processing on the sample image to obtain an annotated sample image, including: According to a preset annotation rule, in the sample image, respectively annotating corresponding bounding boxes for each sample character in the sample string; wherein, the overlapping area range between the bounding boxes of two adjacent sample characters is less than a preset area range threshold; Annotating the corresponding class value according to the sample character included in the bounding box to obtain the annotated sample image.
12. An apparatus for recognizing a text string, characterized in that, Including: An acquisition module, configured to acquire a target image containing a target string to be recognized; wherein, the target string characters in the target image are small in size, low in resolution, and there are few relevant image features extracted based on the target image; A preprocessing module, configured to preprocess the target image to obtain a preprocessed target image; A calling module, which is used to call a preset recognition model to process the preprocessed target image and obtain a corresponding target processing result; wherein, the preset recognition model at least includes a preset dilated convolutional layer; the preset dilated convolutional layer is used to extract target image features related to the target string from the preprocessed target image instead of the combination of a convolutional network layer and a pooling layer; the preset recognition model also simultaneously integrates a localization sub-model and a classification sub-model respectively connected to the preset dilated convolutional layer. When running the preset recognition model, a localization process and a classification process are simultaneously executed according to the target image features to obtain the target processing result; during the execution of the localization process, the calling module is also used to select the candidate box of the character with the highest confidence as the reference box. If there is a candidate box overlapping with the reference box, calculate the proportion of the overlapping area between the two in the total area; if this proportion is greater than the set threshold, modify and adjust the confidence of this candidate box using the following preset linear function: where S i represents the confidence of the i-th candidate box, and IoU i represents the proportion of the overlapping area of this candidate box and the reference box in the total area, and T represents the set threshold; A determination module, configured to determine the target string in the target image according to the target processing result; Wherein, the preset dilated convolutional layer is configured with a preset dilation coefficient and a corresponding convolutional kernel, and the preset dilation coefficient and the size parameters of the convolutional kernel are set according to the proportion and resolution of the target string in the target image; Correspondingly, using the preset dilated convolutional layer to process the preprocessed target image to extract target image features related to the target string, including: using the preset dilation coefficient to perform dilation processing on the initial image feature matrix obtained after the convolutional kernel performs a normal convolution operation on the preprocessed target image, and using the data value 0 to fill the newly added matrix area dilated out to obtain a dilated image feature matrix containing feature information in a large field of view and avoiding loss of feature information as the target image feature.
13. A method for identifying a text string, characterized in that, Including: Obtain a target image containing a target string to be recognized; wherein, the characters of the target string in the target image are small in size, low in resolution, and there are few relevant image features extracted based on the target image; Process the target image using a preset recognition model to obtain a corresponding target processing result; wherein, the preset recognition model at least includes a preset dilated convolutional layer; the preset dilated convolutional layer is used to extract target image features related to the target string from the target image instead of a combination of a convolutional network layer and a pooling layer; the preset recognition model also integrates a localization sub-model and a classification sub-model respectively connected to the preset dilated convolutional layer, and when running the preset recognition model, a localization process and a classification process are simultaneously executed according to the target image features to obtain the target processing result; during the execution of the localization process, the method further includes: selecting the candidate box of the character with the highest confidence as the reference box, if there is a candidate box overlapping with the reference box, calculate the proportion of the overlapping area between the two to the total area; if this proportion is greater than a set threshold, modify and adjust the confidence of this candidate box using the following preset linear function: where S i represents the confidence of the i-th candidate box, and IoU i represents the proportion of the overlapping area of this candidate box and the reference box to the total area, and T represents the set threshold; Determine the target string in the target image according to the target processing result; Wherein, the preset dilated convolutional layer is configured with a preset dilation coefficient and a corresponding convolutional kernel, and the preset dilation coefficient and the size parameter of the convolutional kernel are set according to the proportion and resolution of the target string in the target image; Correspondingly, using the preset dilated convolutional layer to process the target image to extract target image features related to the target string, including: using the preset dilation coefficient to perform dilation processing on the initial image feature matrix obtained after the convolutional kernel performs normal convolution operation on the target image, and using the data value 0 to fill the newly added matrix area after dilation to obtain a dilated image feature matrix containing feature information in a large field of view and avoiding loss of feature information, as the target image feature.
14. A server, characterized in that, It includes a processor and a memory for storing processor-executable instructions, and when the processor executes the instructions, it implements the steps of the method according to any one of claims 1 to 11.
15. A computer-readable storage medium, characterized in that, Stored thereon are computer instructions, and when the instructions are executed, they implement the steps of the method according to any one of claims 1 to 11.
Citation Information
Patent Citations
Printed and handwritten mixed text line extraction system
CN108537146A
Text image detection method and device, computer device and storage medium
CN110674804A