A Multi-Task Enhanced Scene Text Recognition Method and System

The multi-task enhanced scene text recognition method improves accuracy by using a unified architecture with correction, feature generation, and branch tasks, ensuring performance enhancement without increasing model scale or speed, addressing the limitations of existing text recognition technologies.

CN114821559BActive Publication Date: 2025-07-15XIDIAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202210339990.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-01
Publication Date
2025-07-15
Estimated Expiration
2042-04-01

AI Technical Summary

Technical Problem

In the prior art, the recognition rate is difficult to improve under the condition that the speed and scale of the text detection and recognition in the prior art, which affects the accuracy of text recognition.

Method used

Multi-task enhanced scene text recognition model is adopted, including correction network module, feature generation module, context modeling module, prediction module and branch task module. Through branch length prediction tasks and character statistics tasks, the model is trained under the multi-task framework, the supervision information dimension is increased, and the model performance is improved.

Benefits of technology

Without changing the model scale and speed, the text recognition rate is significantly improved, adapted to most scene text recognition algorithms, and stably improved model performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114821559B_ABST
    Figure CN114821559B_ABST
Patent Text Reader

Abstract

The present invention provides a multi-task enhanced scene text recognition method and system. By inputting the original information of the first scene into a correction network module, generating a first corrected result information and inputting it into a feature generation module to obtain first text feature information; inputting it into a context modeling module and a branch task module to obtain a first context modeling result and a second text recognition result; inputting the first context modeling result into a prediction module to obtain a first text recognition result; training a multi-task enhanced scene text recognition model according to the first text recognition result and the second text recognition result; and inputting the original information of the second scene to obtain a first multi-task text recognition result. It solves the technical problem in the prior art that it is difficult to improve the recognition rate of text detection and recognition under the condition of keeping the speed and scale unchanged, which affects the accuracy of text recognition. It achieves the technical effect of stably improving the model performance by using a unified architecture of multi-task and multi-model, and there is no loss in scale and speed during the inference stage.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of text recognition, and particularly to a multi-task enhanced scene text recognition method and system. Background Art

[0002] With the development of the intelligent trend in industries such as industry and service, the text detection and recognition technology in natural scenes has been widely applied to various industries in society, such as the fields of finance, education, and healthcare. Specific applications of text detection and recognition, including document entry, invoice recognition, license plate recognition, and document recognition, are used to improve the work efficiency of all walks of life and simplify the operation process of users. Although the recognition rate of some high-performance text detection and recognition methods in general scenarios has reached over 90%, these methods are difficult to implement due to the actual requirements for algorithm speed and scale. Currently, widely implemented algorithms, such as CRNN, RARE, Rosetta, etc., although meeting the requirements for algorithm speed and scale, it is still relatively difficult to improve the model recognition rate while maintaining the speed and scale unchanged.

[0003] There are at least the following technical problems in the prior art:

[0004] In the prior art, it is difficult to improve the recognition rate of text detection and recognition under the condition of maintaining the speed and scale unchanged, which affects the technical problem of the accuracy of text recognition. Summary of the Invention

[0005] The present application provides a multi-task enhanced scene text recognition method and system for solving the technical problem that in the prior art, it is difficult to improve the recognition rate of text detection and recognition under the condition of maintaining the speed and scale unchanged, which affects the accuracy of text recognition. It achieves the technical effect of being able to adapt to the text recognition requirements of most scenarios by using a unified architecture of multi-task and multi-model, and by proposing a branch length prediction task and a character statistics task, the model performance can be stably improved without any loss of scale and speed in the inference stage.

[0006] In view of the above problems, the present application provides a multi-task enhanced scene text recognition method and system.

[0007] In the first aspect of the present application, a multi-task enhanced scene text recognition method is provided. The method applies a multi-task enhanced scene text recognition model, which includes a correction network module, a feature generation module, a context modeling module, a prediction module, and a branch task module. The method includes: inputting the first scene original information into the correction network module to generate the first corrected result information; inputting the first corrected result information into the feature generation module to obtain the first text feature information; inputting the first text feature information into the context modeling module to obtain the first context modeling result; inputting the first context modeling result into the prediction module to obtain the first text recognition result; inputting the first text feature information into the branch task module to generate the second text recognition result; training the multi-task enhanced scene text recognition model according to the first text recognition result and the second text recognition result; and inputting the second scene original information into the multi-task enhanced scene text recognition model to obtain the first multi-task text recognition result.

[0008] In the second aspect of the present application, a multi-task enhanced scene text recognition system is provided. The system includes:

[0009] A first execution unit, which is configured to input the first scene original information into the correction network module to generate the first corrected result information;

[0010] A first obtaining unit, which is configured to input the first corrected result information into the feature generation module to obtain the first text feature information;

[0011] A second obtaining unit, which is configured to input the first text feature information into the context modeling module to obtain the first context modeling result;

[0012] A third obtaining unit, which is configured to input the first context modeling result into the prediction module to obtain the first text recognition result;

[0013] A second execution unit, which is configured to input the first text feature information into the branch task module to generate the second text recognition result;

[0014] A third execution unit, which is configured to train the multi-task enhanced scene text recognition model according to the first text recognition result and the second text recognition result;

[0015] A first processing unit, which is configured to input the second scene original information into the multi-task enhanced scene text recognition model to obtain the first multi-task text recognition result.

[0016] In a third aspect of the present application, a multi-task enhanced scene text recognition system is provided, including: a processor, the processor is coupled to a memory, and the memory is used to store a program, when the program is executed by the processor, the device is caused to execute the steps of the method described in the first aspect.

[0017] In a fourth aspect of the present application, a computer-readable storage medium is provided, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the method described in the first aspect are implemented.

[0018] One or more technical solutions provided in the present application have at least the following technical effects or advantages:

[0019] In the present application, the original information of the first scene is input into the correction network module to generate the first corrected result information; the obtained first corrected result information is used as the input information and input into the feature generation module to obtain the first text feature information; the first text feature information is used as the input information and enters the main task and the branch task respectively, and is processed according to the set branch task situation. The main task is to input the first text feature information into the context modeling module to obtain the first context modeling result; the first context modeling result is input into the prediction module to obtain the first text recognition result; the branch task is to input the first text feature information into the branch task module to generate the second text recognition result; finally, the multi-task enhanced scene text recognition model is trained according to the first text recognition result and the second text recognition result; the original information of the second scene is input into the multi-task enhanced scene text recognition model to obtain the first multi-task text recognition result. The unified architecture of the multi-task enhanced scene text recognition model constructed by using the correction network module, the feature generation module, the context modeling module, the prediction module and the branch task module can adapt to most scene text recognition algorithms, and a branch recognition task is proposed, which can stably improve the model performance in this model framework. The model trained in the multi-task framework improves the model performance through different tasks, achieving the technical effect of being able to adapt to most scene text recognition requirements by using the multi-task multi-model unified architecture, and proposing the branch length prediction task and the character statistics task, which can stably improve the model performance in the framework and have no scale and speed loss in the inference stage. Thus, the technical problem in the prior art that it is difficult to improve the recognition rate while maintaining the speed and scale unchanged in text detection and recognition, which affects the accuracy of text recognition, is solved.

[0020] The above description is only an overview of the technical solutions of the present application. In order to be able to understand the technical means of the present application more clearly, it can be implemented according to the content of the specification. And in order to make the above and other purposes, features and advantages of the present application more obvious and understandable, the specific embodiments of the present application are hereinafter specifically exemplified. Brief Description of the Drawings

[0021] Figure 1 Flow diagram of a multi-task enhanced scene text recognition method provided for this application;

[0022] Figure 2 Flow diagram of another multi-task enhanced scene text recognition method in this application;

[0023] Figure 3 Flow diagram of another multi-task enhanced scene text recognition method in this application;

[0024] Figure 4 Flow diagram of another multi-task enhanced scene text recognition method in this application;

[0025] Figure 5 This application provides a structural schematic diagram of a multi-task enhanced scene text recognition system;

[0026] Figure 6 Structural schematic diagram of an exemplary electronic device for this application;

[0027] Figure 7 Framework diagram of the multi-task enhanced scene text recognition model in this application.

[0028] Explanation of reference numerals: First execution unit 11, First acquisition unit 12, Second acquisition unit 13, Third acquisition unit 14, Second execution unit 15, Third execution unit 16, First processing unit 17, Electronic device 300, Memory 301, Processor 302, Communication interface 303, Bus architecture 304. Detailed implementation manners

[0029] This application provides a multi-task enhanced scene text recognition method and system for solving the technical problem that in the prior art, it is difficult to improve the recognition rate of text detection and recognition under the conditions of maintaining the speed and scale unchanged, which affects the accuracy of text recognition. Summary of the invention

[0031] Existing text recognition methods such as CRNN, RARE, Rosetta, and ASTER are widely used in the industry. However, due to the constraints on the scale and speed of the model in the application, it is difficult to further improve the recognition rate of these methods.

[0032] ACE introduces additional supervision for the model by modifying the loss function and optimization objective, and can improve the model without changing the original model. However, since this method is essentially a weak supervision and removes the strong supervision constraint, it often shows instability in convergence and requires multiple trainings to obtain a model with good performance.

[0033] PIMNet improves the model performance by changing the attention mechanism of the prediction network and introduces the idea of distillation to add additional supervision to the model. However, this method can only solve the problem of improving the recognition rate of the parallel attention mechanism model and cannot be extended to most scene text recognition models. In addition, although this method does not increase the model size, due to its iterative prediction method, it will still reduce the model speed.

[0034] Therefore, the disadvantages of the prior art include:

[0035] 1. The prior art has a small scope of application. These technologies only improve the performance in a certain type and can only be used for that type, making it difficult to be extended to other types of algorithms for scene text recognition;

[0036] 2. The enhancement effect of the prior art is unstable. For example, although the ACE algorithm can enhance the model, due to the lack of strong supervision information, the model needs to be trained multiple times to obtain a satisfactory enhancement effect;

[0037] 3. The prior art still has losses in scale or speed. The existing technologies for improving performance often add an incremental model, which will lead to an increase in the model size or speed.

[0038] In view of the above technical problems, the general idea of the technical solution provided by this application is as follows:

[0039] The present invention proposes a unified architecture for a multi-task enhanced scene text recognition model, which can be adapted to most scene text recognition algorithms, and proposes branch recognition tasks: length prediction task and character statistics task, which can stably improve the model performance under this model framework. The model trained under the multi-task framework has no loss in scale and speed during the inference stage.

[0040] After introducing the basic principle of this application, hereinafter, the technical solutions in this application will be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments of this application. It should be understood that this application is not limited by the example embodiments described here. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of this application. Additionally, it should be noted that for the sake of description, only the parts related to this application are shown in the accompanying drawings rather than all of them.

[0041] Embodiment 1

[0042] Such as Figure 1As shown in the figure, an embodiment of the present application provides a multi-task enhanced scene text recognition method. The method applies a multi-task enhanced scene text recognition model, which includes a correction network module, a feature generation module, a context modeling module, a prediction module, and a branch task module. The method includes:

[0043] Specifically, the multi-task enhanced scene text recognition model in the embodiment of the present application is a unified architecture of multiple models, which includes five modules, namely, a correction network module, a feature generation module, a context modeling module, a prediction module, and a branch task module. Different models are selected in each module, and the several modules work together to achieve character recognition.

[0044] For example, CRNN selects None (correction module) + VGG (feature generation module) + BiLSTM (context modeling module) + CTC (prediction module); another example is that RARE selects TPS (correction module) + VGG (feature generation module) + BiLSTM (context modeling module) + serial attention mechanism (prediction module).

[0045] Such as Figure 7 As shown in the framework diagram of the multi-task enhanced scene text recognition model, the output end of the correction network model is connected to the input end of the feature generation module. The output end of the feature generation module is connected to the input end of the context modeling module on the one hand and the input end of the branch task module on the other hand. The output end of the context modeling module is connected to the input end of the prediction module. The branch task module is relatively independent of the main task composed of the other four modules, but there is a gradient propagation association with the main task during the training stage. To ensure that the branch task is available in any scene text recognition algorithm, the input of the branch task module comes from the output of the feature generation module. Among them, the correction network module, the feature generation module, the context modeling module, and the prediction module are used to complete the main task, and the branch task module is used to complete the branch task. The multi-model is reflected in that it is not limited to a specific scene text recognition model, and any scene text recognition algorithm that can be incorporated into the main task is available. During the training stage, the five modules work simultaneously, and different supervision information is added to the main task and the branch task to enhance the model.

[0046] S10: Input the original information of the first scene into the correction network module to generate the first corrected result information.

[0047] Specifically, the original scene information that needs to be text-recognized is input into the correction module. The first original scene information is the acquisition result of the information to be text-recognized for any scene, including forms such as pictures and videos, and is used as the original scene information in the training stage. After using the model in the correction network module to perform correction operation processing on the input first original scene information, its correction result is output as the first corrected result information. Among them, the correction network module is a functional module used to correct the text content in the scene, and the correction content includes but is not limited to all existing known text correction methods.

[0048] S20: Input the first corrected result information into the feature generation module to obtain the first text feature information.

[0049] Specifically, take the result output by the correction module, that is, the first corrected result information, as the input information of the feature generation module. Use the model in the feature generation module to perform operation processing on the input first corrected result information, output the text features in the first corrected result, and the first text feature information is the text feature information determined by feature analysis through the feature generation model. Among them, the feature generation module is a functional module used to extract text features from text content, and the types of text features extracted and the extraction algorithms can be selected according to the actual required scene.

[0050] S30: Input the first text feature information into the context modeling module to obtain the first context modeling result.

[0051] S40: Input the first context modeling result into the prediction module to obtain the first text recognition result.

[0052] Specifically, the first text feature information will enter the main task and the branch task respectively for subsequent text recognition processing. For the main task, use the first text feature information as the input information and input it into the context modeling module. Use the model algorithm in the context modeling module to perform operation processing on the first text feature information. After obtaining the first context modeling result, input it into the prediction module for operation processing to obtain the character classification prediction result and add corresponding supervision to obtain the first text recognition result. Among them, the context modeling module is a functional module used to concatenate the context of the text; the prediction module is a functional module used to process according to the text context information and text feature information to perform text content prediction.

[0053] S50: Input the first text feature information into the branch task module to generate the second text recognition result.

[0054] Furthermore, as Figure 2 shown, the step of inputting the first text feature information into the branch task module to generate the second text recognition result, S50 includes:

[0055] S501: Obtain the first branch task module and / or the second branch task module according to the branch task module;

[0056] S502: Input the first text feature information into the first branch task module to obtain a first task text recognition result;

[0057] S503: Input the first text feature information into the second branch task module to obtain a second task text recognition result.

[0058] Specifically, for branch tasks, the first text feature information is used as input information and input into the branch task module. The algorithm model in the branch task module is used to perform arithmetic processing on the first text feature information, which includes two branch tasks (the first / second branch task modules): length prediction task and character statistics task. The first text feature information is input into the two branch tasks respectively, and the corresponding supervision information is used to obtain the second task text recognition result. It should be understood that one or two branch tasks can be selected. For the case where two are set, they enter the first branch task module and the second branch task module respectively. For the case where only one branch task is set, it is correspondingly input into the branch task module, that is, the first or second branch task module. Among them, the branch task module is used to perform different text feature processing tasks for the execution and prediction modules, so as to achieve the technical effect of improving the scene text recognition rate.

[0059] S60: Train a multi-task enhanced scene text recognition model according to the first text recognition result and the second text recognition result;

[0060] Specifically, the first text recognition result obtained from the main task, and the two branch recognition results obtained from the branch module, the length prediction result and the character recognition result, are jointly used as the output information in the training stage of the multi-task enhanced scene text recognition model. The supervision signals corresponding to the recognition results of the main task and the supervision signals corresponding to the branch recognition results of the branch tasks are used to supervise the output information. When the output of the multi-task enhanced scene text recognition model converges, a multi-task enhanced scene text recognition model is generated.

[0061] Since in the training stage, the supervised training of the branch task module is increased, the dimension of the supervised information during the training of the multi-task enhanced scene text recognition model is increased, thereby ensuring the information dimension of the recognition result, achieving the technical effect of improving the performance of the backbone network of the multi-task enhanced scene text recognition model, and thus improving the text recognition rate. After the model is constructed, the next step enters the inference stage.

[0062] S70: Input the second scene original information into the multi-task enhanced scene text recognition model to obtain a first multi-task text recognition result.

[0063] Specifically, after the multi-task enhanced scene text recognition model is trained, it can enter the inference stage for scene text recognition. The second scene original information is the information to be text recognized with the same scene type as the first scene original information but different other information, including but not limited to: data forms such as pictures, videos, etc.; the first multi-task text recognition result refers to the text recognition result obtained after processing the second scene original information by the multi-task enhanced scene text recognition model.

[0064] Since the multi-task enhanced scene text recognition model is trained in the training stage by means of a multi-task module, the performance of the backbone network of the multi-task enhanced scene text recognition model is strengthened. In the inference stage, that is, the scene text recognition stage, only a few modules of the main task need to work to perform multi-task text recognition on the scene text, thereby ensuring that the speed and scale of the inserted model remain unchanged, and the recognition rate is greatly improved. Thus, it solves the technical problem that in the prior art, it is difficult to improve the recognition rate under the condition of maintaining the speed and scale in text detection and recognition, which affects the accuracy of text recognition. It achieves the technical effect that by using a unified architecture of multi-task and multi-model, it can adapt to the text recognition requirements of most scenes, and proposes a branch length prediction task and a character statistics task, which can stably improve the model performance under the framework without any loss of scale and speed in the inference stage.

[0065] Further, as Figure 3 shown, the step of inputting the first text feature information into the first branch task module to obtain the first task text recognition result, S502 includes:

[0066] S5021: Generate a first position encoding according to the first text feature information;

[0067] S5022: Perform feature fusion on the first text feature and the first position encoding through the attention mechanism to obtain a first position fusion feature;

[0068] S5023: Input the first position fusion feature into the first fully connected layer for existence feature extraction, and then process it through the Sigmoid activation function to obtain a first existence feature;

[0069] S5024: Input the first existence feature into the length prediction formula to obtain a first string length prediction result;

[0070] Among them, the length prediction formula is:

[0071]

[0072] Among them, Y LEN represents the string length prediction result of the current position. Indicates the existence feature of the length of the string characterized by T at the corresponding position, W LEN Is the weight corresponding to the length of the string characterized by T, R1 T×T Is a value matrix composed of T×T string lengths, T represents the longest string length, F e Indicates the existence feature, R1 T×1 Represents the value matrix of the string length.

[0073] Specifically, the first branch task module is a length prediction task, which takes the output feature of the feature generation module, the first text feature information, and the position encoding P LEN ∈R T×D As the input of this branch task together, and then uses the attention mechanism to fuse the relevant position features to obtain the features at each position, that is, the first position fusion feature. The first position fusion feature is the fusion feature of any position. Through a linear layer and an activation function, the existence feature of each position character is obtained, that is, the feature information of whether there is a string at the corresponding position, as the first existence feature. Finally, the first existence feature is used for length prediction to obtain the string length prediction result as the first substring length prediction result. Since the string length is a discrete integer value, a classification strategy is used to predict the string length. When setting the labels, the string lengths are counted, and then the cross-entropy loss function is used to constrain the prediction results.

[0074] Further, generating the first position encoding according to the first text feature information, S5021 includes:

[0075] S50211: Obtain the first position encoding calculation formula:

[0076] P LEN =P origin W pe ,P LEN 、W pe ∈R2 T×D

[0077] Wherein, P LEN Represents the text position encoding, P origin =I T×T Represents the one-hot encoding of the character position, I represents the corrected scene text information, T represents the longest string length, D represents the dimension number of the first text feature information, W pe Is the weight matrix of the learnable P origin , R2 T×D Is a value matrix composed of weights corresponding to T×D string lengths;

[0078] S50212: Input the first text feature into the first position encoding calculation formula to generate the first position encoding.

[0079] Specifically, the position encoding of the characters in the first text feature information is processed through the first position encoding calculation formula. Through the formula: P LEN = P origin W pe , P LEN , W pe ∈R2 T×D , after operating on the text feature in the first text feature information, the position encoding of this feature is obtained.

[0080] Furthermore, the feature fusion is performed based on the attention mechanism through the first text feature and the first position encoding to obtain the first position fusion feature. S5022 includes:

[0081] S50221: Obtain the first position fusion feature calculation formula:

[0082]

[0083] Among them, F I represents the text feature information, represents the string length corresponding to the text feature, F LEN represents the position fusion feature;

[0084] S50222: Input the first text feature information and the first position encoding into the first position feature calculation formula to obtain the first position fusion feature.

[0085] Furthermore, as Figure 4 shown, the step of inputting the first text feature information into the second branch task module to obtain the second task text recognition result, S503 includes:

[0086] S5031: Traverse the first text feature information to generate the first query vector;

[0087] S5032: Based on the attention mechanism, input the first query vector into the first character fusion feature calculation formula for feature fusion to obtain the first character fusion feature;

[0088] Among them, the first character fusion feature calculation formula is:

[0089]

[0090] Among them, F CS represents the character fusion feature, W cp represents the weight occupied by the text feature F I R4D×D Denote a value matrix composed of text feature weights of D×D dimensions, where D represents the number of text feature dimensions, and F I Denote text features, Denote the character features corresponding to the text features;

[0091] S5033: Input the first character fusion feature into the second fully connected layer to obtain a first classification result.

[0092] Specifically, in order to introduce the task of counting the number of each type of character, that is, the second branch task, the network structure is designed according to the attention mechanism idea of the string length branch. Similarly, the first text feature information is used as the input information. At the same time, for the first text feature information, its first query vector is determined, and the first query vector and the first text feature information are used as the input information of this branch task at the same time. The first query vector is the query vector in the attention mechanism, representing the encoding of all characters in the character set in the main task. Use the attention mechanism to take the first query vector, the image features and their transformations as key-value pairs, perform feature fusion, calculate the fusion features, and classify the fusion features obtained for each character through a fully connected layer, that is, the second fully connected layer.

[0093] It can be obtained through the first character fusion feature calculation formula that, different from the branch length task, the character counting task first calculates the correlation between the character features and the image features, and then performs weighted summation on the transformed image features according to the correlation. Since the image features directly calculate the correlation with the character features, compared with the string length task, it more enhances the ability of the character feature growth in the image feature generation network.

[0094] Further, the traversing the first text feature information to generate the first query vector, S5031 includes:

[0095] S50311: Obtain the first query vector calculation formula:

[0096] C CS =SW ce ,

[0097] where, C CS denotes the query vector, S denotes the one-hot encoding set of the character set, and W ce is the weight matrix corresponding to the learnable character set, N c denotes the number of characters in the character set, is N c ×N c a value matrix composed of the number of characters in the character set, is a value matrix composed of the number of characters corresponding to D feature dimensions;

[0098] S50312: Traverse the first text feature for one-hot encoding extraction to obtain the first character encoding set;

[0099] S50313: Input the first character encoding set into the first query vector calculation formula to obtain the first query vector.

[0100] Specifically, the first query vector is obtained through the operation of the above first query vector calculation formula. Traverse and encode the first text feature from beginning to end to obtain the one-hot encoding of all characters contained therein. Use the one-hot encoding set of all characters as the first character encoding set, and perform operations based on all character encodings through the first query vector calculation formula to obtain the query vectors of all characters.

[0101] Furthermore, the method further includes:

[0102] S701: Based on the prediction module and the branch task module, match the first weight parameter;

[0103] S702: Generate the first cross-entropy loss function according to the first weight parameter:

[0104] L = αL C + βL LEN + γL CS

[0105] where α, β, and γ are weight parameters determined according to the usage of the task module, L is the total loss function, L C is the loss function of the prediction module, L LEN is the loss function of the first branch task module, L CS is the loss function of the second branch task module;

[0106] S703: Evaluate the recognition rate of the multi-task text recognition result according to the first cross-entropy loss function.

[0107] Specifically, the loss function of the entire model consists of three parts, namely the loss Lc of the backbone task, the length classification cross-entropy loss LLEN of the text length, and the cross-entropy loss LCS of the classification of the number of each type of character. Among them, Lc has two optional loss functions, namely applying CTC Loss to the model without using the attention mechanism to decode features and applying the character classification cross-entropy loss function to the model using the attention mechanism to decode features. Since the quantity itself is a discrete value within a certain range, and introducing MSE may lead to problems such as gradient disappearance, we model both branch tasks as classification tasks of quantities.

[0108] Among them, in the string length prediction task, the length of each string is counted during training to generate labels. In the character counting task, the types of characters in each word are counted during training to generate labels. The cross-entropy loss function is applied to both tasks, and the total loss is: L = αLc + βLLEN + γLCS, where α, β, and γ are optional weight parameters. For a model with only the string length constraint branch added, α = 1.0, β = 0.5, γ = 0.0; for a model with only the character counting task branch added, α = 1.0, β = 0.0, γ = 0.5; for a model using both branches, α = 1.0, β = 0.25, γ = 0.25. After the model framework is applied according to the first cross-entropy loss function, the recognition accuracy of the processing result can be evaluated. The larger the loss function value, the lower the recognition accuracy.

[0109] In summary, the beneficial effects achieved by this application at least include the following:

[0110] 1. Through the unified architecture of multi-task and multi-model for scene text recognition, it can adapt to most scene text recognition algorithms, and at the same time, based on the multi-task learning strategy, it can enhance any text recognition algorithm incorporated into the framework.

[0111] 2. By training the text recognition model through the unified framework based on multi-task and multi-model, any branch task can be added to enhance the original model without changing the original model, strengthening the model performance of the backbone network, and thus ensuring that the recognition rate of the text recognition model can be improved without changing the scale and speed.

[0112] 3. By adding the length prediction task and the character counting task, the model performance can be stably improved under the framework, and there is no loss in scale and speed during the inference stage.

[0113] Embodiment 2

[0114] Based on the same inventive concept as a multi-task enhanced scene text recognition method in the foregoing embodiment, as Figure 5 shown, this application provides a multi-task enhanced scene text recognition system, wherein the system includes:

[0115] The first execution unit 11 is used to input the first scene original information into the correction network module to generate the first corrected result information;

[0116] The first obtaining unit 12 is used to input the first corrected result information into the feature generation module to obtain the first text feature information;

[0117] The second obtaining unit 13 is used to input the first text feature information into the context modeling module to obtain the first context modeling result;

[0118] A third acquisition unit 14, configured to input the first context modeling result into a prediction module to obtain a first text recognition result;

[0119] A second execution unit 15, configured to input the first text feature information into a branch task module to generate a second text recognition result;

[0120] A third execution unit 16, configured to train a multi-task enhanced scene text recognition model according to the first text recognition result and the second text recognition result;

[0121] A first processing unit 17, configured to input second scene original information into the multi-task enhanced scene text recognition model to obtain a first multi-task text recognition result.

[0122] Further, the system further includes:

[0123] A fourth acquisition unit, configured to obtain a first branch task module and / or a second branch task module according to the branch task module;

[0124] A fifth acquisition unit, configured to input the first text feature information into the first branch task module to obtain a first task text recognition result;

[0125] A sixth acquisition unit, configured to input the first text feature information into the second branch task module to obtain a second task text recognition result.

[0126] Further, the system further includes:

[0127] A second processing unit, configured to generate a first position encoding according to the first text feature information;

[0128] A seventh acquisition unit, configured to perform feature fusion on the first text feature and the first position encoding based on an attention mechanism to obtain a first position fusion feature;

[0129] A third processing unit, configured to input the first position fusion feature into a first fully connected layer to perform existence feature extraction, and then process it through a Sigmoid activation function to obtain a first existence feature;

[0130] An eighth acquisition unit, configured to input the first existence feature into a length prediction formula to obtain a first string length prediction result;

[0131] Among them, the length prediction formula is as follows:

[0132]

[0133] Among them, Y LEN represents the prediction result of the string length at the current position, represents the existence feature of the string length characterized by T at the corresponding position, and W LEN is the weight corresponding to the string length characterized by T, and R1 T×T is a value matrix composed of T×T string lengths, where T represents the longest string length, and F e represents the existence feature, and R1 T×1 represents the value matrix of the string length.

[0134] Furthermore, the system further includes:

[0135] A ninth acquisition unit, which is used to acquire the first position encoding calculation formula:

[0136] P LEN = P origin W pe , P LEN , W pe ∈R2 T×D

[0137] Among them, P LEN represents the text position encoding, and P origin = I T×T represents the one-hot encoding of the character position, I represents the corrected scene text information, T represents the longest string length, D represents the dimension number of the first text feature information, and W pe is the weight matrix of the learnable P origin , and R2 T×D is a value matrix composed of weights corresponding to T×D string lengths;

[0138] A fourth processing unit, which is used to input the first text feature into the first position encoding calculation formula to generate the first position encoding.

[0139] Furthermore, the system further includes:

[0140] A tenth acquisition unit, which is used to acquire the first position fusion feature calculation formula:

[0141]

[0142] Among them, F I represents the text feature information, represents the string length corresponding to the text feature, and FLEN Represents the positional fusion feature;

[0143] A fifth processing unit, which is configured to input the first text feature information and the first positional encoding into the first positional feature calculation formula to obtain the first positional fusion feature.

[0144] Furthermore, the system further includes:

[0145] A fourth execution unit, which is configured to traverse the first text feature information to generate a first query vector;

[0146] An eleventh obtaining unit, which is configured to input the first query vector into the first character fusion feature calculation formula for feature fusion based on the attention mechanism to obtain the first character fusion feature;

[0147] Wherein, the first character fusion feature calculation formula is:

[0148]

[0149] Wherein, F CS represents the character fusion feature, W cp represents the text feature F I occupies the weight, R4 D×D represents a value matrix composed of text feature weights of D×D dimensions, D represents the number of dimensions of the text feature, F I represents the text feature, represents the character feature corresponding to the text feature;

[0150] A twelfth obtaining unit, which is configured to input the first character fusion feature into the second fully-connected layer to obtain a first classification result.

[0151] Furthermore, the system further includes:

[0152] A thirteenth obtaining unit, which is configured to obtain a first query vector calculation formula:

[0153] C CS =SW ce ,

[0154] Wherein, C CS represents the query vector, S represents the one-hot encoding set of the character set, W ce is the weight matrix corresponding to the learnable character set, N c represents the number of characters in the character set, is N c ×N cA value matrix composed of the number of characters in a character set, is a value matrix composed of the number of characters corresponding to D feature dimensions;

[0155] A fourteenth acquisition unit, configured to traverse the first text feature for one-hot encoding extraction to obtain a first character encoding set;

[0156] A fifteenth acquisition unit, configured to input the first character encoding set into the first query vector calculation formula to obtain the first query vector.

[0157] Further, the system further includes:

[0158] A first matching unit, configured to match a first weight parameter based on the prediction module and the branch task module;

[0159] A fifth execution unit, configured to generate a first cross-entropy loss function according to the first weight parameter:

[0160] L = αL C + βL LEN + γL CS

[0161] Where α, β, and γ are weight parameters determined according to the usage of the task module, L is the total loss function, and L C is the loss function of the prediction module, and L LEN is the loss function of the first branch task module, and L CS is the loss function of the second branch task module;

[0162] A first evaluation unit, configured to evaluate the recognition rate of the multi-task text recognition result according to the first cross-entropy loss function.

[0163] Embodiment III

[0164] Based on the same inventive concept as a multi-task enhanced scenario text recognition method in the foregoing embodiment, the present application further provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the method in Embodiment I.

[0165] Exemplary electronic device

[0166] Next, refer to Figure 6 to describe the electronic device of the present application.

[0167] Based on the same inventive concept as a multi-task enhanced scenario text recognition method in the foregoing embodiments, the present application further provides a multi-task enhanced scenario text recognition system, including: a processor, the processor is coupled to a memory, and the memory is used to store a program. When the program is executed by the processor, the device is caused to execute the steps of the method described in Embodiment 1.

[0168] The electronic device 300 includes: a processor 302, a communication interface 303, and a memory 301. Optionally, the electronic device 300 may further include a bus architecture 304. Among them, the communication interface 303, the processor 302, and the memory 301 may be interconnected through the bus architecture 304; the bus architecture 304 may be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus, etc. The bus architecture 304 may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 6 only a thick line is used to represent it in the figure, but it does not mean that there is only one bus or one type of bus.

[0169] The processor 302 may be a CPU, a microprocessor, an ASIC, or one or more integrated circuits for controlling the execution of the program of the solution of the present application.

[0170] The communication interface 303 uses any device of the transceiver type for communicating with other devices or communication networks, such as Ethernet, radio access network (RAN), wireless local area networks (WLAN), wired access networks, etc.

[0171] The memory 301 can be a ROM or other types of static storage devices that can store static information and instructions, a RAM or other types of dynamic storage devices that can store information and instructions, or can also be an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM), or other optical disc storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media, or other magnetic storage devices, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory can exist independently and be connected to the processor through the bus architecture 304. The memory can also be integrated with the processor.

[0172] Among them, the memory 301 is used to store computer execution instructions for executing the solution of this application, and is controlled by the processor 302 for execution. The processor 302 is used to execute the computer execution instructions stored in the memory 301, so as to implement a multi-task enhanced scenario text recognition method provided by the above embodiments of this application.

[0173] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in this application are generated in whole or in part. The computer can be a general-purpose computer, a dedicated computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from a website, computer, server, or data center to another website, computer, server, or data center in a wired manner (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or a wireless manner (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer, or a data storage device such as a server or data center that includes one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid-state disk (SSD)), etc.

[0174] Although the present application has been described in connection with specific features and their embodiments, it will be apparent that various modifications and combinations can be made without departing from the spirit and scope of the present application. Accordingly, the present specification and the drawings are merely exemplary illustrations of the present application and are considered to have covered any and all modifications, variations, combinations, or equivalents within the scope of the present application. Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the present application and its equivalent technologies, the present application is intended to include these changes and modifications.

Claims

1. A multi-task enhanced scene text recognition method, characterized in that The method applies a multi-task enhanced scene text recognition model, which includes a correction network module, a feature generation module, a context modeling module, a prediction module, and a branch task module. The method includes: Input the first scene original information into the correction network module to generate the first corrected result information; Input the first corrected result information into the feature generation module to obtain the first text feature information; Input the first text feature information into the context modeling module to obtain the first context modeling result; Input the first context modeling result into the prediction module to obtain the first text recognition result; Input the first text feature information into the branch task module to generate the second text recognition result; Train the multi-task enhanced scene text recognition model according to the first text recognition result and the second text recognition result; Input the second scene original information into the multi-task enhanced scene text recognition model to obtain the first multi-task text recognition result; The step of inputting the first text feature information into the branch task module to generate the second text recognition result includes: Obtain the first branch task module and / or the second branch task module according to the branch task module; Input the first text feature information into the first branch task module to obtain the first task text recognition result; Input the first text feature information into the second branch task module to obtain the second task text recognition result; The step of inputting the first text feature information into the first branch task module to obtain the first task text recognition result includes: Generate the first position encoding according to the first text feature information; Perform feature fusion based on the attention mechanism through the first text feature and the first position encoding to obtain the first position fusion feature; Input the first position fusion feature into the first fully connected layer for existence feature extraction, and then process it through the Sigmoid activation function to obtain the first existence feature; Input the first existence feature into the length prediction formula to obtain the first string length prediction result; Wherein, the length prediction formula is: Among them, represents the prediction result of the string length at the current position, represents the existence feature of the string length characterized at the corresponding position, is the weight corresponding to the string length characterized, is a value matrix composed of several string lengths, represents the longest string length, represents the existence feature, represents the value matrix of the string length; The step of inputting the first text feature information into the second branch task module to obtain the second task text recognition result includes: The second branch task module is a task of introducing the statistics of the number of each type of character; Traverse the first text feature information to generate the first query vector; Based on the attention mechanism, input the first query vector into the first character fusion feature calculation formula for feature fusion to obtain the first character fusion feature; Wherein, the first character fusion feature calculation formula is: Among them, represents the character fusion feature, represents the text feature occupies the weight, represents a value matrix composed of the weights of text features in dimensions, represents the number of dimensions of text features, represents the text feature, represents the character feature corresponding to the text feature; Input the first character fusion feature into the second fully connected layer to obtain the first classification result, including using the attention mechanism to perform feature fusion on the first query vector, image features and their transformations as key-value pairs, calculate the fusion features, and classify the fusion features obtained for each character through a fully connected layer, i.e., the second fully connected layer; The step of traversing the first text feature information to generate the first query vector includes: Obtain the first query vector calculation formula; Among them, represents the query vector, represents the encoding set of the character set, is the weight matrix corresponding to the learnable character set, represents the number of characters in the character set, is a value matrix composed of the number of characters in the character set, is the value matrix composed of the number of characters corresponding to D feature dimensions; Traverse the first text feature to perform encoding extraction to obtain a first character encoding set; Input the first character encoding set into the first query vector calculation formula to obtain the first query vector.

2. The method according to claim 1, wherein Generating a first positional encoding according to the first text feature information includes: Obtaining a first positional encoding calculation formula: Among them, represents the text position encoding, represents the encoding of the character position, represents the corrected scene text information, represents the length of the longest string, represents the number of dimensions of the first text feature information, is a learnable weight matrix, is a value matrix composed of weights corresponding to the string lengths; Inputting the first text feature into the first positional encoding calculation formula to generate the first positional encoding.

3. The method according to claim 2, wherein Performing feature fusion based on the attention mechanism through the first text feature and the first positional encoding to obtain a first position fusion feature, including: Obtaining a first position fusion feature calculation formula: Among them, represents text feature information, represents the string length corresponding to the text feature, represents the position fusion feature; Inputting the first text feature information and the first positional encoding into the first position fusion feature calculation formula to obtain the first position fusion feature.

4. The method according to claim 1, wherein The method further includes: Matching a first weight parameter based on the prediction module and the branch task module; Generating a first cross-entropy loss function according to the first weight parameter; Among them, , and are weight parameters determined according to the usage of the task modules, is the total loss function, is the loss function of the prediction module, is the loss function of the first branch task module, is the loss function of the second branch task module; Evaluating the recognition rate of the multi-task text recognition result according to the first cross-entropy loss function.

5. A multi-task enhanced scene text recognition system, characterized in that, The system is applied to any of the methods in claims 1-4. The system includes: A first execution unit for inputting first scene original information into a correction network module to generate first corrected result information; A first obtaining unit for inputting the first corrected result information into a feature generation module to obtain first text feature information; A second obtaining unit for inputting the first text feature information into a context modeling module to obtain a first context modeling result; A third obtaining unit for inputting the first context modeling result into a prediction module to obtain a first text recognition result; A second execution unit for inputting the first text feature information into a branch task module to generate a second text recognition result; A third execution unit for training a multi-task enhanced scene text recognition model according to the first text recognition result and the second text recognition result; A first processing unit for inputting second scene original information into the multi-task enhanced scene text recognition model to obtain a first multi-task text recognition result; Further, the system further includes: A fourth obtaining unit for obtaining a first branch task module and / or a second branch task module according to the branch task module; A fifth obtaining unit for inputting the first text feature information into the first branch task module to obtain a first task text recognition result; A sixth obtaining unit for inputting the first text feature information into the second branch task module to obtain a second task text recognition result; A second processing unit for generating a first positional encoding according to the first text feature information; A seventh obtaining unit for performing feature fusion based on the attention mechanism through the first text feature and the first positional encoding to obtain a first position fusion feature; A third processing unit for inputting the first position fusion feature into a first fully-connected layer for existence feature extraction, and then processing through a Sigmoid activation function to obtain a first existence feature; An eighth acquisition unit, configured to input the first existence feature into a length prediction formula to obtain a first string length prediction result; wherein, the length prediction formula is: Among them, represents the prediction result of the string length at the current position, represents the existence feature of the string length represented at the corresponding position, is the weight corresponding to the string length represented, is a value matrix composed of string lengths, represents the longest string length, represents the existence feature, represents the value matrix of the string length; A fourth execution unit, configured to traverse the first text feature information to generate a first query vector; An eleventh acquisition unit, configured to input the first query vector into a first character fusion feature calculation formula based on an attention mechanism for feature fusion to obtain a first character fusion feature; wherein, the first character fusion feature calculation formula is: Among them, represents the character fusion feature, represents the text feature occupies the weight, represents a value matrix composed of the weights of the text features in represents the number of dimensions of the text feature, represents the text feature, represents the character feature corresponding to the text feature; A twelfth acquisition unit, configured to input the first character fusion feature into a second fully-connected layer to obtain a first classification result; A thirteenth acquisition unit, configured to obtain a first query vector calculation formula: Among them, represents the query vector, represents the encoding set of the character set, is the weight matrix corresponding to the learnable character set, represents the number of characters in the character set, is a value matrix composed of the number of characters in the character set, is the value matrix composed of the number of characters corresponding to D feature dimensions; The fourteenth acquisition unit is configured to traverse the first text feature to perform encoding extraction to obtain a first character encoding set; A fifteenth acquisition unit, configured to input the first character encoding set into the first query vector calculation formula to obtain the first query vector.

6. A multi-task enhanced scene text recognition system, characterized in that Including: A processor, the processor is coupled with a memory, the memory is configured to store a program, and when the program is executed by the processor, the device is caused to execute the steps of the method according to any one of claims 1 to 4.