Training method of text recognition model, text recognition method and device, electronic device and medium

By introducing supervised joint semi-supervised training and random character perturbation enhancement in the text recognition model, combined with three-branch residual blocks and LSTM networks, the accuracy and robustness of text recognition in common scenarios are solved, and more efficient text recognition capabilities are achieved.

CN114462489BActive Publication Date: 2025-08-22ZHEJIANG DAHUA TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202111633893.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-29
Publication Date
2025-08-22
Estimated Expiration
2041-12-29

AI Technical Summary

Technical Problem

When the prior art recognizes text in general scenarios, the recognition accuracy and robustness are insufficient, and the recognition of scene text is limited, and the generalization ability is poor, especially under the condition of unlabeled data.

Method used

A text recognition model is designed to obtain feedback joint losses of labeled data and labelless data, random character perturbation enhancement is performed, and supervised joint semi-supervised training is carried out. The model training process is optimized by using three-branch residual blocks and LSTM network structure, combining the CTC loss function and mean square error loss.

Benefits of technology

It improves the recognition accuracy and robustness of the text recognition model in general scenarios, effectively utilizes labelless data, enhances the anti-noise ability of the model in complex scenarios, and improves the recognition effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114462489B_ABST
    Figure CN114462489B_ABST
Patent Text Reader

Abstract

The present application discloses a training method for a text recognition model, a text recognition method and device, an electronic device and a medium, the method comprising: obtaining labeled data, unlabeled data and a feedback joint loss of the labeled data and the unlabeled data, the feedback joint loss being calculated based on the labeled data, the unlabeled data and the loss function; performing random character perturbation enhancement on the unlabeled data to obtain the perturbed unlabeled data; using the labeled data, the feedback joint loss and the perturbed unlabeled data to perform supervised and semi-supervised training on the text recognition model in training until the loss function converges, thereby obtaining the trained text recognition model. In the above manner, the present application uses the labeled data, the feedback joint loss and the perturbed unlabeled data to implement supervised and semi-supervised training on the text recognition model in training, thereby improving the text recognition capability of the text recognition model in general scenarios.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of text recognition technology, and in particular to a text recognition model training method, a text recognition method and device, an electronic device, and a medium. Background Art

[0002] Generally, as people's demands for the use of products and equipment increase, when using products and equipment for text recognition, users often hope to maintain both the timeliness and accuracy of text recognition by the products and equipment.

[0003] Optical character recognition (OCR) has become one of the most important technologies in the field of artificial intelligence. Recognizing text in common scenarios based on OCR technology is of great significance. The accuracy of text recognition in common scenarios is related to the size of the data sample. Most scenes are captured using devices such as smartphones and cameras.

[0004] At present, in the process of recognizing text on the acquired image, due to the screening of scene text, some difficult text samples are filtered out, which further reduces the acquired text samples. In addition, there is often no regular semantic information between the text or characters in the text, and it is impossible to perform semantic modeling on the numbers in the scene text. In addition, a large amount of manual annotation is often required in real scenes, resulting in limited text recognition in various scenes, prone to misrecognition, poor robustness, and poor generalization ability. Summary of the Invention

[0005] In order to solve the above technical problems, the technical solution adopted in the first aspect of this application is to provide a training method for a text recognition model, which method includes: obtaining labeled data, unlabeled data, and feedback joint loss of labeled data and unlabeled data, and the feedback joint loss is calculated based on the labeled data, the unlabeled data and the loss function; performing random character perturbation enhancement on the unlabeled data to obtain perturbed unlabeled data; using the labeled data, the feedback joint loss, and the perturbed unlabeled data to perform supervised and semi-supervised training on the text recognition model in training until the loss function converges, thereby obtaining the trained text recognition model.

[0006] In order to solve the above technical problems, the technical solution adopted in the second aspect of this application is to provide an identification device, which includes: an acquisition module for acquiring labeled data, unlabeled data and the feedback joint loss of labeled data and unlabeled data, and the feedback joint loss is calculated based on the labeled data, the unlabeled data and the loss function; a perturbation enhancement module for performing random character perturbation enhancement on the unlabeled data to obtain the perturbed unlabeled data; a supervised training module for jointly performing supervised and semi-supervised training on the text recognition model in training with the labeled data, the feedback joint loss and the perturbed unlabeled data with the semi-supervised training module until the loss function converges to obtain the trained text recognition model.

[0007] In order to solve the above technical problems, the technical solution adopted in the third aspect of this application is to provide an electronic device, which includes: a processor and a memory, the memory stores a computer program, and the processor is used to execute the computer program to implement the method as described in the first aspect of this application.

[0008] In order to solve the above technical problems, the technical solution adopted in the fourth aspect of this application is to provide a computer-readable storage medium, which stores a computer program, and the computer program can implement the method of the first aspect of this application when executed by a processor.

[0009] The beneficial effects of the present application are: in order to accurately identify text content in general scenarios, the present application designs a text recognition model, performs perturbation enhancement on unlabeled data, and uses labeled data, feedback joint loss and perturbation unlabeled data to implement supervised and semi-supervised training of the text recognition model in training, thereby improving the text recognition ability of the text recognition model in general scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0011] Figure 1 It is a flowchart of the training method of the text recognition model of this application;

[0012] Figure 2 This is a schematic diagram of a specific framework flow of the training method of the text recognition model of this application;

[0013] Figure 3 This application Figure 2 A specific structural flow chart of the middle school student model;

[0014] Figure 4 This application Figure 3 Schematic diagram of the structural flow of the three-branch residual block;

[0015] Figure 5 This application Figure 1 A flow chart of a specific embodiment of step S12;

[0016] Figure 6 yes Figure 5 A flowchart of a specific random character perturbation enhancement embodiment;

[0017] Figure 7 This application Figure 1 A flow chart of a specific embodiment of step S13;

[0018] Figure 8 This is a schematic block diagram of the structure of an embodiment of a text recognition device of the present application;

[0019] Figure 9 This is a schematic block diagram of the structure of an electronic device embodiment of the present application;

[0020] Figure 10 This is a schematic block diagram of a circuit of an embodiment of a computer-readable storage medium of the present application. DETAILED DESCRIPTION

[0021] In the following description, specific details such as specific system structures and techniques are provided for purposes of illustration rather than limitation to facilitate a thorough understanding of the embodiments of the present application. However, it will be apparent to those skilled in the art that the present application may be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid obscuring the description of the present application with unnecessary detail.

[0022] It will be understood that when used in this specification and the appended claims, the term "comprising" indicates the presence of described features, integers, steps, operations, elements and / or components, but does not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups thereof.

[0023] It should also be understood that the terms used in this specification are for the purpose of describing specific embodiments only and are not intended to limit the present application. As used in this specification and the appended claims, the singular forms "a," "an," and "the" are intended to include the plural forms unless the context clearly indicates otherwise.

[0024] It should be further understood that the term "and / or" used in this specification and the appended claims refers to and includes any and all possible combinations of one or more of the associated listed items.

[0025] As used in this specification and the appended claims, the term "if" can be interpreted as "when" or "upon" or "in response to determining" or "in response to detecting," depending on the context. Similarly, the phrase "if it is determined" or "if [described condition or event] is detected" can be interpreted as meaning "upon determination" or "in response to determining" or "upon detection of [described condition or event]" or "in response to detecting [described condition or event]," depending on the context.

[0026] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0027] In order to illustrate the technical solution of this application, the following is a specific embodiment to illustrate that this application provides a text recognition method, which is applied to the text recognition model. Figure 1 , Figure 1 This is a flow chart of the training method of the text recognition model of the present application, which specifically includes the following steps:

[0028] S11: Obtain the joint loss of labeled data, unlabeled data, and feedback of labeled and unlabeled data;

[0029] Typically, the input to a text recognition model is text images from various scenes. If someone uses labels to annotate text images as samples, then the text images and the labels corresponding to the text images are collectively referred to as labeled data, while only text images are determined to be unlabeled data.

[0030] Real-world scenarios contain a large number of unlabeled samples. For text recognition models, the input can be both labeled and unlabeled data. This expands the sample range of the input text recognition model and effectively utilizes the large amount of unlabeled sample data, achieving better recognition results compared to other scene text recognition methods.

[0031] In order to accurately identify text content in common scenarios, this application designs a text recognition model. The text recognition model can process labeled data and unlabeled data to obtain a feedback joint loss, wherein the feedback joint loss is obtained by a backpropagation process. Specifically, the feedback joint loss is calculated based on labeled data, unlabeled data and a loss function, and is used to perform feedback adjustment on the text recognition model.

[0032] S12: Perform random character perturbation enhancement on the unlabeled data to obtain the perturbed unlabeled data;

[0033] In general scenarios, the accuracy of text recognition is related to the data sample capacity. When the text recognition model is running, sometimes the amount of data is too small and the training results are inaccurate. In this case, the number of characters can be randomly perturbed to expand the capacity, such as by floating the characters up and down within a range or Gaussian perturbation, to increase the data volume of the characters.

[0034] For unlabeled samples, random character perturbation is added to increase the diversity of each character in the text string in the text image. This can make the text recognition model more resistant to noise and improve its robustness.

[0035] S13: Use labeled data, feedback joint loss, and perturbed unlabeled data to perform supervised and semi-supervised training on the text recognition model in training until the loss function converges to obtain the trained text recognition model.

[0036] Specifically, the labeled data, the feedback joint loss and the perturbed unlabeled data are input into the text recognition model under training to perform supervised and semi-supervised training on the text recognition model under training, and finally the trained text recognition model can be obtained when the loss function converges.

[0037] In this way, under the condition of small sample sizes with a limited number of labeled samples, we can efficiently utilize unlabeled samples to design a general scene text recognition method based on semi-supervised learning, and use the corresponding text recognition model to run the text recognition method to obtain accurate predicted text sequences.

[0038] Therefore, in order to accurately identify text content in common scenarios, this application designs a text recognition model, performs perturbation enhancement on unlabeled data, and uses labeled data, feedback joint loss, and perturbation-free unlabeled data to implement supervised and semi-supervised training of the text recognition model in training, thereby improving the text recognition ability of the text recognition model in common scenarios.

[0039] Furthermore, the training text recognition model includes the training student model and the training teacher model, see Figure 2 , Figure 2 This is a specific framework flow diagram of the training method of the text recognition model of this application. Its overall architecture includes two networks, the student model in training and the teacher model in training. The network structure framework of these two models is the same. The parameters of the teacher network are calculated through the student network, and the parameters of the student network are obtained through loss function gradient descent and back propagation update.

[0040] like Figure 2 As shown in the figure, during the training phase, the labeled data and the unlabeled data are combined and input into the training text recognition model, including the training student model and the training teacher model. The training student model is supervised and the labeled data is trained with connectionist temporal classification loss (CTC). The CTC loss function can solve the problem of whether the input and output are aligned, avoiding the labeling of each character, and only requires labeling samples line by line.

[0041] The method uses mean squared error loss training for unlabeled samples, and then uses an appropriate coefficient to balance the importance of the two. The appropriate coefficient here refers to the supervisory weight determined by the mean squared error and can be preset manually. At the beginning of training, the weighted coefficient of the mean squared error loss of the student model and the teacher model is 0. Labeled data is required for supervised training of the student model to obtain a student model with good recognition performance. As the number of training times increases, the weighted coefficient of the mean squared error loss increases, making full and effective use of the large amount of unlabeled sample data. Compared with other scene text recognition methods, it can achieve better recognition results with only limited labeled sample data.

[0042] Furthermore, the text recognition model in training is trained in a supervised and semi-supervised manner using labeled data, feedback joint loss, and perturbed unlabeled data. Specifically, the following steps are performed:

[0043] The feedback joint loss, labeled data, and perturbed unlabeled data are input into the student model in training for supervised and semi-supervised training to obtain a first prediction value; the perturbed unlabeled data are input into the teacher model in training for semi-supervised training to obtain a second prediction value.

[0044] Among them, the network structure of the teacher model is copied through the student model. The student model participates in supervised training and semi-supervised training, and the teacher model only participates in semi-supervised training.

[0045] For the student model, see Figure 3 and Figure 4 , Figure 3 This application Figure 2A specific structural flow chart of the middle school student model; Figure 4 This application Figure 3 Schematic diagram of the structural flow of the three-branch residual block;

[0046] Furthermore, the network structure of the student model includes a three-branch residual block, a pooling layer, and a recurrent neural network (LSTM).

[0047] Among them, the three-branch residual block is a 3*3 convolutional neural network with a 1*1 convolutional neural network residual structure and a first residual structure of cross-layer connection added. The three-branch residual block is used to independently represent the text sequence features, and the convolutional neural network is used to obtain the image information of the scene text; Figure 4 As shown in the figure, the input is divided into three branches. The left side is the first residual structure formed by directly connecting weighted points across layers, the middle is the input 3*3 convolutional neural network, and the right side is the input 1*1 convolutional neural network. The three are weighted and then output uniformly.

[0048] Recurrent neural networks model sequential information of text sequence features to learn the relationship between characters.

[0049] Since the residual structure has multiple branches, it is equivalent to adding multiple gradient flow paths in the text recognition model. Its function can learn more unique feature representations in complex general text scenes, solve problems such as gradient vanishing and gradient explosion in deep text recognition models, and accelerate the convergence of text recognition models.

[0050] Furthermore, the image information includes at least one of texture, spatial information, and local details of the scene text image. That is, in scene text recognition, the convolutional neural network can obtain information such as texture, spatial information, and local details of the scene text image. Spatial information refers to positional information, and local details refers to the receptive field. Local details refer to the local feature information of the text image captured by the convolution window.

[0051] Furthermore, in order to further optimize the feature learning ability of the scene text recognition model, a second residual structure is added between the last layer of the convolutional neural network and the first layer of the recurrent neural network in the student model to combine the text sequence features extracted by the convolutional neural network, strengthen the recurrent neural network to learn the semantic information between texts, and improve the text recognition ability of general scenarios.

[0052] Going further, random character perturbation is performed on unlabeled data, see Figure 5 and Figure 6 , Figure 5 This application Figure 1 The flowchart of step S12 in a specific embodiment is as follows: Figure 6 yes Figure 5A flowchart of a specific random character perturbation enhancement embodiment includes the following steps:

[0053] S21: Obtain text images input to the student model as unlabeled data;

[0054] For unlabeled samples, random character perturbation enhancement is added to increase the diversity of each character in the text string in the image. Specifically, a text image is input, where the text image + label becomes labeled data, and only the text image is unlabeled data.

[0055] S22: Divide the text image into N image sub-blocks, where N is a positive integer greater than or equal to 1;

[0056] like Figure 6 As shown, the text image is divided into N image sub-blocks on average. Specifically, when N=4, it means there are 4 image sub-blocks.

[0057] S23: Initialize N image sub-blocks along the boundary of the text image to form 2(N+1) reference points, wherein each reference point is set to a range circle with a radius of R, with the center of the circle as the initial origin;

[0058] Then, 2(N+1) reference points are initialized along the top, bottom, left, and right edges of the image. Each reference point is assigned a range circle with a radius of R, with the center of the circle as the initial origin. Specifically, when there are N=4 image sub-blocks, there are 2(N+1)=10 reference points, where R can be manually preset and can be set to greater than or equal to 0. The specific setting is based on the needs and is not limited here.

[0059] S24: Randomly perturb the pixel points within the range circle according to Gaussian distribution to change the shape and / or distortion of each character in the unlabeled data.

[0060] Specifically, the Gaussian distribution formula used is as follows:

[0061]

[0062] Where μ is the mean, σ is the variance, x is the horizontal coordinate of the pixel, and p(x) is the vertical coordinate of the pixel. This application adopts the standard Gaussian distribution, that is, μ = 0, σ 2 =1, the formula is as follows (2):

[0063]

[0064] for example Figure 6 Each character in the network presents different degrees of distortion and different shapes, but it does not affect the semantic information it represents. Therefore, by adding random character perturbation enhancement, the network can be made more resistant to noise and more robust.

[0065] Furthermore, before performing random character perturbation enhancement on the unlabeled data, the method further includes: inserting preset characters between repeated characters when performing encoding preprocessing on the label, wherein the preset characters are different from the label.

[0066] Specifically, since there may usually be repeated characters in text images, in the process of training the text recognition model, it is often possible to insert preset characters between repeated characters when encoding the labels before performing random character perturbation enhancement on the unlabeled data. For example, spaces are added between repeated characters. Of course, other characters are also acceptable as long as they are different from the labels.

[0067] Thus, during decoding, the best path is calculated by selecting the most likely character at each time step, removing duplicate characters, and then removing all special characters from the path. What remains is the recognized text.

[0068] Furthermore, after the perturbed unlabeled data, labeled data and feedback joint loss are trained in a supervised and semi-supervised manner, see Figure 7 , Figure 7 This application Figure 1 The method further includes:

[0069] S31: extracting labels from labeled data;

[0070] In real-world scenarios, characters are often manually labeled one by one. These sample annotations are labels of labeled data. By extracting these labels, a comparison standard can be provided for labeled data in the student model, allowing the student model to perform supervised training on labeled data.

[0071] S32: Input the label, the first predicted value, and the second predicted value into the loss function to calculate the loss;

[0072] A loss function is provided in the text recognition model. By calculating the loss on the input label, the first prediction value and the second prediction value, a preset loss result can be obtained to make the loss function converge and obtain the trained text recognition model.

[0073] S33: Determine the first prediction value or the second prediction value corresponding to the obtained preset loss result as a predicted text sequence.

[0074] Specifically, the first prediction value or the second prediction value corresponding to the preset loss result can be determined as the predicted text sequence through comparison and selection, wherein the preset loss result can be the minimum loss result, so that the predicted text sequence is closer to the real text sequence.

[0075] The preset loss result may correspond to the first predicted value or the second predicted value.

[0076] Furthermore, the loss function includes a supervised loss function and an unsupervised loss function; the supervised loss function at least includes a connection time classification loss function.

[0077] The loss function is input into the label, the first predicted value, and the second predicted value to calculate the loss, including:

[0078] Call the supervised loss function, fit the label and the first predicted value, and obtain the connection time classification loss value, so that the connection time classification loss function converges and obtains the trained student model.

[0079] The unsupervised loss function is called to process the second prediction value and the first prediction value to obtain the mean square error loss value, so that the unsupervised loss function converges and the trained teacher model is obtained. The difference between the second prediction value and the first prediction value is less than the preset difference, wherein the loss value is close to 0, mainly to ensure that the prediction result of the teacher model is as similar as possible to the prediction result of the student network.

[0080] Among them, such as Figure 2 As shown, the feedback joint loss is determined based on the sum of the joint temporal classification loss value and the mean squared error loss value.

[0081] Furthermore, the method also includes: using the loss function gradient descent and the optimizer to update the parameters of the student model, wherein the student model parameters are obtained by the loss function gradient descent and back propagation, and in addition, the parameters of the teacher model are obtained by updating the parameters of the student model by the sliding average function, and there is no need to store all historical training results.

[0082] Specifically, the sliding average function formula is as follows:

[0083] θ′ t =αθ′ t-1 +(1-α)θ t (3)

[0084] Among them, α is the sliding average weight balance coefficient, θ t is the parameter of the student model at time t, θ′ t-1 is the parameter of the teacher model at time t-1, θ′ t are the parameters of the teacher model at time t.

[0085] Furthermore, the parameters of the student model are updated using a loss function gradient descent and an optimizer. During backpropagation, the optimizer's algorithm adjusts the weights and biases of the convolutional neural network and the recurrent neural network within the student model's network structure to update the student model's parameters. The smaller the CTC loss, the closer the text sequence predicted by the student model is to the true text sequence. In other words, the magnitude of the joint temporal classification loss is negatively correlated with the accuracy of the predicted text sequence.

[0086] In order to illustrate the technical solution of the present application, the present application also provides a text recognition method, which includes: obtaining a text image; and calling the above-mentioned trained text recognition model to recognize the text image to obtain a predicted text sequence.

[0087] Therefore, this application provides a general scene text recognition method based on semi-supervised learning; designs a student and teacher scene text recognition model based on three-branch residual block and residual LSTM, and uses labeled data to perform supervised training on the student model, and unlabeled data to perform joint semi-supervised training on the student model and the teacher model.

[0088] Compared to using supervised learning alone, using joint semi-supervised learning to effectively utilize unlabeled samples, given the same number of labeled samples, can significantly improve text recognition rates in general scenarios. Furthermore, by adding random character perturbations to unlabeled samples, the general-scenario text recognition model's ability to resist interference is enhanced, significantly improving its robustness.

[0089] In order to illustrate the technical solution of this application, this application also provides a text recognition device, please refer to Figure 8 , Figure 8 It is a structural schematic block diagram of an embodiment of the text recognition device of the present application. The text recognition device 60 includes: an acquisition module 61, a disturbance enhancement module 62, a supervised training module 63 and a semi-supervised training module 64.

[0090] An acquisition module 61 is configured to acquire labeled data, unlabeled data, and a feedback joint loss of the labeled data and the unlabeled data, where the feedback joint loss is calculated based on the labeled data, the label data, and a loss function;

[0091] The disturbance enhancement module 62 is used to perform random character disturbance enhancement on the unlabeled data to obtain disturbed unlabeled data;

[0092] The supervised training module 63 is used to jointly conduct semi-supervised training on the training text recognition model using labeled data, feedback joint loss and perturbed unlabeled data until the loss function converges, thereby obtaining the trained text recognition model.

[0093] Therefore, in order to accurately identify the text content in general scenarios, this application designs a text recognition model, performs perturbation enhancement on the unlabeled data through the perturbation enhancement module 62, and combines the supervised training module 63 with the semi-supervised training module 64, and uses labeled data, feedback joint loss and perturbation of the unlabeled data to implement supervised and semi-supervised training on the text recognition model in training, thereby improving the text recognition ability of the trained text recognition model in general scenarios.

[0094] In order to illustrate the technical solution of the present application, the present application also provides an electronic device, which can be a computer or a mobile phone, etc., without specific limitation, please refer to Figure 9 , Figure 9 This is a schematic block diagram of the structure of an embodiment of an electronic device of the present application. The electronic device 7 includes: a processor 71 and a memory 72. The memory 72 stores a computer program 721. The processor 71 is used to execute the computer program 721 to implement the method of the embodiment of the present application, which will not be repeated here.

[0095] In addition, this application also provides a computer-readable storage medium, see Figure 10 , Figure 10 This is a circuit schematic block diagram of an embodiment of a computer-readable storage medium of the present application. The computer-readable storage medium 8 stores a computer program 81. When the computer program 81 is executed by a processor, it can implement the method of the embodiment of the present application, which will not be repeated here.

[0096] If it is implemented in the form of a software functional unit and sold or used as an independent product, it can also be stored in a device with a storage function. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage device, including a number of instructions (program data) to enable a computer device (which can be a personal computer, server, or network device, etc.) or a processor (processor) to execute all or part of the steps of the various embodiments of the present invention. The aforementioned storage device includes various media such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and electronic devices such as computers, mobile phones, laptops, tablet computers, cameras, etc. having the above-mentioned storage media.

[0097] The description of the execution process of program data in the device with storage function can be referred to the description in the above-mentioned method embodiment of the present application, which will not be repeated here.

[0098] The above description is merely an embodiment of the present application and does not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made using the contents of the present application specification and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present application.

Claims

1. A method for training a text recognition model, characterized in that: The text recognition model in training includes a student model in training and a teacher model in training, and the network structure of the teacher model is the same as the network structure of the student model; the method includes: Obtaining labeled data, unlabeled data, and a feedback joint loss of the labeled data and the unlabeled data, where the feedback joint loss is calculated based on the labeled data, the unlabeled data, and a loss function, wherein the loss function includes a supervised loss function and an unsupervised loss function, and the supervised loss function includes at least a joint temporal classification loss function; Performing character perturbation enhancement on the unlabeled data to obtain perturbed unlabeled data; Performing supervised and semi-supervised training on the text recognition model in training using the labeled data, the feedback joint loss, and the perturbed unlabeled data until the loss function converges, thereby obtaining a trained text recognition model; The method of using the labeled data, the feedback joint loss, and the perturbed unlabeled data to perform supervised and semi-supervised training on the text recognition model in training comprises: inputting the feedback joint loss, the labeled data, and the perturbed unlabeled data into the student model in training to perform supervised and semi-supervised training to obtain a first prediction value, and inputting the perturbed unlabeled data into the teacher model in training to perform semi-supervised training to obtain a second prediction value; When the loss function converges, a trained text recognition model is obtained, including: extracting the label of the labeled data; calling the supervised loss function, fitting the label and the first prediction value, obtaining the connection time classification loss value so that the connection time classification loss function converges, and obtaining the trained student model; calling the unsupervised loss function, processing the second prediction value and the first prediction value, obtaining the mean square error loss value so that the unsupervised loss function converges, and obtaining the trained teacher model, the difference between the second prediction value and the first prediction value is less than the preset difference, so that the trained student model and the trained teacher model constitute the trained text recognition model; wherein, the feedback joint loss is determined based on the sum of the connection time classification loss value and the mean square error loss value.

2. The method according to any one of claim 1, characterized in that The character perturbation enhancement of the unlabeled data includes: Obtaining a text image input to the student model as the unlabeled data; Divide the text image into N image sub-blocks, where N is a positive integer greater than or equal to 1; Initializing the N image sub-blocks to form 2(N+1) reference points along the boundary of the text image, wherein each reference point is set to a range circle with a radius of R, with the center of the circle as the initial origin, wherein R is greater than or equal to 0; The pixels within the range circle are randomly perturbed according to Gaussian distribution to change the shape and / or distortion of each character in the unlabeled data.

3. The method according to claim 2, characterized in that Before performing character perturbation enhancement on the unlabeled data, the method further includes: When performing encoding preprocessing on the label, preset characters are inserted between repeated characters, wherein the preset characters are different from the label.

4. The method according to claim 1, wherein The network structure of the student model includes a three-branch residual block, a pooling layer and a recurrent neural network; The three-branch residual block is a 3*3 convolutional neural network with a 1*1 convolutional neural network residual structure and a first residual structure with cross-layer connections added. The three-branch residual block is used to independently represent text sequence features, and the convolutional neural network is used to obtain image information of scene text. The recurrent neural network performs sequential information modeling on the text sequence features to learn the association relationship between characters.

5. The method according to claim 4, characterized in that A second residual structure is provided between the last layer of the convolutional neural network and the first layer of the recurrent neural network in the student model. The second residual structure is used to combine the text sequence features extracted by the convolutional neural network to strengthen the semantic information between the characters learned by the recurrent neural network.

6. The method according to claim 4, characterized in that The image information includes at least one of texture, space, and local details of the scene text image.

7. The method according to claim 1, characterized in that The method further comprises: The parameters of the student model are updated using the loss function gradient descent and the optimizer, wherein the parameters of the teacher model are obtained by updating the parameters of the student model using a sliding average function.

8. The method according to claim 7, characterized in that The loss function gradient descent and the use of an optimizer to update the parameters of the student model include: During the back-propagation process, the weights and bias items of the convolutional neural network and the weights and bias items of the recurrent neural network in the network structure of the student model are adjusted by the optimizer algorithm to update the parameters of the student model.

9. A method for character recognition, characterized in that: The method comprises: Get text image; Calling the trained text recognition model according to any one of claims 1 to 8 to recognize the text image to obtain a predicted text sequence.

10. A text recognition device, characterized in that: The character recognition device comprises: an acquisition module, configured to acquire labeled data, unlabeled data, and a feedback joint loss of the labeled data and the unlabeled data, wherein the feedback joint loss is calculated based on the labeled data, the unlabeled data, and a loss function, wherein the loss function includes a supervised loss function and an unsupervised loss function, and the supervised loss function includes at least a joint temporal classification loss function; The disturbance enhancement module is used to perform character disturbance enhancement on the unlabeled data to obtain the disturbed unlabeled data; A supervised training module is used for a joint semi-supervised training module, which uses the labeled data, the feedback joint loss and the perturbed unlabeled data to perform supervised joint semi-supervised training on the text recognition model in training, until the loss function converges to obtain the trained text recognition model; wherein, the text recognition model in training includes a student model in training and a teacher model in training, and the network structure of the teacher model is the same as the network structure of the student model; the use of the labeled data, the feedback joint loss and the perturbed unlabeled data to perform supervised joint semi-supervised training on the text recognition model in training includes: inputting the feedback joint loss, the labeled data and the perturbed unlabeled data into the student model in training to perform supervised joint semi-supervised training to obtain a first prediction value, and inputting the perturbed unlabeled data into the teacher model in training. The teacher model performs the semi-supervised training to obtain a second prediction value; when the loss function converges, a trained text recognition model is obtained, including: extracting the label of the labeled data, calling the supervised loss function, fitting the label and the first prediction value, obtaining the connection time classification loss value so that the connection time classification loss function converges, and obtaining the trained student model; calling the unsupervised loss function, processing the second prediction value and the first prediction value, obtaining the mean square error loss value so that the unsupervised loss function converges, and obtaining the trained teacher model, the difference between the second prediction value and the first prediction value is less than the preset difference, so that the trained student model and the trained teacher model constitute the trained text recognition model; wherein, the feedback joint loss is determined based on the sum of the connection time classification loss value and the mean square error loss value.

11. An electronic device, characterized in that: include: A processor and a memory, wherein the memory stores a computer program, and the processor is configured to execute the computer program to implement the method according to any one of claims 1 to 9.

12. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 9 can be implemented.