Model generation system and method
The model generation system addresses the challenge of recognizing text lines in diverse styles by training a text line recognition model with a small number of labeled samples and fine-tuning it with unlabeled data, achieving effective recognition across different styles with reduced data requirements.
Patent Information
- Application Number
- JP2022045991
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2022-03-22
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2042-03-22
AI Technical Summary
Existing text line recognition methods based on deep learning achieve high accuracy for training data but perform poorly when faced with data of different styles, requiring large amounts of labeled data for acceptable recognition rates.
A model generation system that generates a text line recognition model using a visual feature extractor and a linguistic context-related network, where the model is trained with a small number of labeled data samples and fine-tuned using a large amount of unlabeled data to learn invariant features across different styles.
The system enables the generation of a text line recognition model that can accurately recognize text lines in images of different styles even with a small number of training samples, improving recognition accuracy and reducing the need for extensive labeled data.
Smart Images

Figure 0007674295000004 
Figure 0007674295000005 
Figure 0007674295000006
Abstract
Description
[Technical field]
[0001] The present invention relates to a technique for recognizing text in a text line image. [Background technology]
[0002] Document recognition brings many benefits to various fields such as retail, finance, education, logistics, and healthcare. Typically, document recognition begins with detecting text lines, and then recognizing the text lines. Current technologies struggle to recognize characters in text line images, since text lines can have multiple styles, such as different handwriting, complex backgrounds, and different fonts.
[0003] In the past decades, research on text line recognition has mostly focused on the method of segmenting a string into single characters and recognizing them individually, as shown in Non-Patent Document 1. In this method, a text line is segmented into character patterns by a projection profile (projection histogram) and a number of heuristic hypotheses. The segmented character patterns are recognized by a feature matching model, and the recognition candidates are combined with a language context model in a lattice diagram. The optimal path in the diagram is searched and output as the recognition result.
[0004] In recent years, with the rapid progress of deep learning, the approach of text line recognition has mostly shifted to segmentation-free methods of convolutional neural networks (CNNs) and recurrent neural networks (RNNs). These are superior to the methods based on the heuristic assumptions mentioned above. First, an end-to-end trainable method based on CNNs, bidirectional long short-term memory (BLSTM), and connectionist temporal classification (CTC) was proposed to recognize scene text images as shown in Non-Patent Document 2. One problem with this method is that it assumes the independence of features of the time steps of BLMST when predicting the output label of the CTC layer. This is known as the hard alignment problem, which reduces the accuracy of the model.
[0005] To solve this problem, a method based on the RNN encoder-decoder attention mechanism was then proposed, as shown in Non-Patent Document 3. In this method, features are extracted using CNN. The features are encoded by the RNN. The attention mechanism aligns the encoded features to the output labels. The RNN decoder then learns to decode the encoded features into the corresponding labels in sequence, from the beginning to the end of the text line. The decoded results of the previous time steps are also used to decode the later time steps, so this method can solve the problems of the CTC-based method. However, the problem with this method is that it is sequential learning, and the errors caused by decoding will spread to later time steps.
[0006] Recently, a self-attention method without RNNs has been proposed, as shown in Non-Patent Document 4. In this method, CNNs are used to extract features. A dot-production self-attention based encoder encodes the CNN features. Then, a dot-production self-attention based decoder learns to decode the encoded features into output labels. In the training phase, this learning process is done in parallel for every character of the output label. Therefore, it overcomes the problems of the RNN encoder-decoder attention based method.
[0007] In the above deep learning based methods, the model is generally composed of a CNN feature extractor (called FEX) and a linguistic contextual relation network (called RN). FEX extracts deep visual features of text line images, which downsamples the input image to reduce the computational cost of the next layer, and RN learns the relationship between character patterns in text line images. The above methods achieve high accuracy when the training data and test data are of the same style, but the accuracy decreases when testing on new styles of data. To generalize the above models to data of various styles, a lot of labeled data is needed to obtain an acceptable recognition rate.
[0008] A solution to save labeling costs for preparing labeled data is to apply transferable feature learning methods such as dropout and batch normalization to build models, diversify training data by data augmentation methods, and apply domain adaptation methods, as shown in [5]. Currently, transferable feature learning methods and data augmentation are often applied to deep learning-based models, but these methods are not robust enough for documents of various styles. Recently, a method of domain adaptation, as shown in [6], has produced promising results for recognizing documents of various styles. This method utilizes unlabeled data to make a model learn invariant features of text line images of various styles. One of the drawbacks of this method is that a large amount of unlabeled data is not always available. Therefore, a domain adaptation method with a small number of samples is required. [Prior art documents] [Non-patent literature]
[0009] [Non-Patent Document 1] LIU, Cheng-Lin, et al. “Online and offline handwritten Chinese character recognition: benchmarking on new databases.” Pattern Recognition, 2013, 46.1: 155-162. [Non-Patent Document 2] Shi, Baoguang, Xiang Bai, and Cong Yao. (2016) “An end to-end trainable neural network for image-based sequence recognition and its application to scene text recognition.” IEEE transactions on pattern analysis and machine intelligence 39, no. 11: 2298-2304. [Non-Patent Document 3] Kang, Lei, J. Ignacio Toledo, Pau Riba, Mauricio Villegas, Alicia Fornes, and Marcal Rusinol. (2018) “Convolve, attend and spell: An attention-based sequence-to-sequence model for handwritten word recognition.” In German Conference on Pattern Recognition, pp. 459-472. Springer, Cham. [Non-Patent Document 4] Lee, Junyeop, Sungrae Park, Jeonghun Baek, Seong Joon Oh, Seonghyeon Kim, and Hwalsuk Lee. “On recognizing texts of arbitrary shapes with 2D self-attention.” In Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition Workshops, pp. 546-547. 2020. [Non-Patent Document 5] Wang, Mei, and Weihong Deng. (2018) “Deep visual domain adaptation: A survey.” Neurocomputing 312: 135-153. [Non-Patent Document 6] Zhang, Yaping, Shuai Nie, Wenju Liu, Xing Xu, Dongxiang Zhang, and Heng Tao Shen. “Sequence-to-sequence domain adaptation network for robust text-image recognition.” In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 2740-2749. 2019. Summary of the Invention [Problem to be solved by the invention]
[0010] The main problem with deep learning based text line recognition methods is overfitting, which means that the model can achieve high recognition accuracy for data that is the same as the training data, but the recognition accuracy decreases for data with a different style from the training data. There are various types of text line data, such as printed text, scene text, handwritten text, etc. In each data type, the text lines have different handwriting styles, fonts, and backgrounds. In addition, the contents of the text lines are rich. Therefore, a large amount of data is required to train the model to obtain acceptable accuracy.
[0011] The present invention has been made in consideration of the above circumstances, and an object of the present invention is to provide a technology that can appropriately generate a text line recognition model that can be adapted to a text line image of a desired style even if the number of training data samples is small. [Means for solving the problem]
[0012] In order to achieve the above object, a model generation system according to one aspect is a model generation system for generating a text line recognition model for recognizing text lines included in a text line image, the model generation system including a processor unit, the text line recognition model including a visual feature extractor that, when executed by the processor unit, outputs image features from a text line image, and a linguistic context relation network that, when executed by the processor unit, inputs the features output from the visual feature extractor and outputs text lines, the processor unit obtains text data for training, and trains the linguistic context relation network using the obtained text data to determine variables of the linguistic context relation network, and executes the text line recognition model with the variables of the linguistic context relation network fixed to the determined variables, Existing The variables of the visual feature extractor are determined by training using labeled text line images, and the text line recognition model is generated with the variables of the linguistic contextual relation network being the determined variables of the linguistic contextual relation network and the variables of the visual feature extractor being the determined variables of the visual feature extractor. Effect of the Invention
[0013] According to the present invention, even with a small number of training data samples, a text line recognition model that can be adapted to a text line image of a desired style can be appropriately generated. [Brief description of the drawings]
[0014] [Figure 1] FIG. 1 is a diagram illustrating a text line recognition model generated in a model generation system according to an embodiment. [Diagram 2] FIG. 2 is a hardware configuration diagram of a model generation system according to an embodiment. [Diagram 3] FIG. 3 is a diagram showing an example of a GUI screen related to a training process for training an RN according to an embodiment. [Figure 4]FIG. 4 is a diagram illustrating a first example of a training process for training an RN according to an embodiment. [Diagram 5] FIG. 5 is a diagram illustrating a second example of a training process for training an RN according to an embodiment. [Figure 6] FIG. 6 is a diagram showing an example of a GUI screen related to a generation process of a prototype model according to an embodiment. [Figure 7] FIG. 7 is a flowchart of a prototype model generation process according to an embodiment. [Figure 8] FIG. 8 is a diagram showing an example of a GUI screen related to a retraining process for retraining a text line recognition model according to an embodiment. [Figure 9] FIG. 9 is a flow chart of a process for retraining a text line recognition model according to one embodiment. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0015] The following embodiments will be described with reference to the drawings. Note that the following embodiments do not limit the scope of the invention, and not all of the elements and combinations thereof described in the embodiments are necessarily essential to the solution of the invention.
[0016] In the following description, detailed explanations of deep learning such as convolutional neural networks (CNN) and recurrent neural networks (RNN) may be omitted since they are understood by those skilled in the art.
[0017] In the following description, a "processor unit" includes one or more processors. The at least one processor is typically a microprocessor such as a CPU (Central Processing Unit) or a GPU (Graphical Processing Unit). Each of the one or more processors may be a single-core or multi-core. The processor may include a hardware circuit that performs part or all of the processing.
[0018] FIG. 1 is a diagram illustrating a text line recognition model generated in a model generation system according to an embodiment.
[0019] The model generation system 10 includes a text line recognition model 100. The text line recognition model 100 includes a visual feature extractor (FEX) 101 and a linguistic-context relation network (RN) 102.
[0020] The FEX 101 inputs a text line image and outputs a feature amount in the text line image. The FEX 101 includes a shallow layer of a CNN such as VGGNet or RestNet. Note that since VGGNet and RestNet are well-known technologies, detailed explanations are omitted. This FEX 101 can reduce the calculation cost of subsequent processing by downsampling the input image.
[0021] The RN 102 receives the feature and outputs the text contained in the text line image. The RN 102 includes, for example, an encoder that encodes the input feature and a decoder that receives the encoded data and restores each character. The RN 102 may be, for example, networks 103, 104, 105, and 106. The network 103 is a BLSTM encoder The network 104 includes an RNN encoder 104a that encodes input features, an attention unit 104b that infers which part of the features to focus on, and an RNN decoder 104c that receives the data inferred by the attention unit 104b and restores each character. The network 105 includes a dot production self-attention encoder 105a and a dot production self-attention decoder 105b. The network 106 includes a natural language processing model 106a.
[0022] Next, an example of the hardware configuration of the model generation system 10 will be described.
[0023] FIG. 2 is a hardware configuration diagram of a model generation system according to an embodiment.
[0024] The model generation system 10 is configured by a computer such as a PC (Personal Computer) or a general-purpose server, and includes a communication interface (communication I / F) 11, a CPU 12, an input device 13, a storage device 14, a memory 15, a display device 16, a GPU 17, and a bus 18. The communication I / F 11, the CPU 12, the input device 13, the storage device 14, the memory 15, the display device 16, and the GPU 17 are connected via the bus 18. The model generation system 10 may be configured by a plurality of computers.
[0025] The communication I / F 11 is connected to a network such as the Internet (not shown) and transmits and receives data to and from other devices connected to the network. The CPU 12 executes various processes by executing programs stored in the memory 15. In this embodiment, the CPU 12 executes a process for executing the text line recognition model 100, but has the GPU 17 execute some of the processes.
[0026] The storage device 14 is, for example, a non-temporary storage device (non-volatile storage device) such as a Hard Disk Drive (HDD) or a Solid State Drive (SSD), and stores programs executed by the CPU 12 and various information. The memory 15 is, for example, a RAM (Random Access Memory), and stores programs executed by the CPU 12 and various information.
[0027] The GPU 17 is a processor suitable for executing specific processes such as image processing and execution of a neural network model, and is suitable for executing processes performed in parallel. In this embodiment, the GPU 17 executes a predetermined process according to an instruction from the CPU 12.
[0028] The input device 13 is, for example, a mouse, a keyboard, etc., and receives various inputs from an operator. The display device 16 is, for example, a display, and displays and outputs a screen including various information using a GUI (Graphical User Interface).
[0029] Here, we show how to generalize the text line recognition model to recognize text line images of different styles. First, we present a lemma for generalizing the text line recognition model.
[0030] Lemma:φ e and φ r Let FEX and RN be the weights of the text line recognition model, respectively. Let I∈R be the text image with various styles. (W×h×c) (where R denotes a set of images, W denotes the image width, h denotes the image height, and c denotes the image channel (e.g., RGB)). If FEX can be generalized to φ e and φ r The prototype model f with is to be generalized.
[0031] According to the lemma, the process of the model generation system 10 to generate a generalized text line recognition model and fine-tune the text line recognition model to newly emerging data is as follows.
[0032] Step 1: The model generation system 10 obtains a large amount of copyright-free text from pages published on the Internet via the Internet, and uses this text to train the RN 102. Details of step 1 will be described later with reference to Figures 3 to 5.
[0033] Step 2: The model generation system 10 generalizes the FEX 101 to text line images of various styles by training the text line recognition model 100 using existing labeled text line images while freezing (fixing) the weights (variables) of the RN 102 obtained by training in step 1. That is, the variables of the FEX 101 are adjusted. Details of step 2 will be described later with reference to Figs. 6 and 7.
[0034] Step 3: The model generation system 10 fine-tunes the text line recognition model 100, i.e., fine-tunes the variables of the text line recognition model 100, by training the text line recognition model 100 trained up to step 2 using data of several labeled text line images (called bootstrap data) that are samples of the desired style to be recognized. The text line recognition model 100 thus fine-tuned can achieve high recognition accuracy in text recognition in text line images containing the style to be recognized. Details of step 3 will be described later with reference to Figs. 8 and 9.
[0035] Next, the process (step 1) for training the RN 102 will be described with reference to FIGS.
[0036] FIG. 3 is a diagram showing an example of a GUI screen related to a training process for training an RN according to an embodiment.
[0037] A GUI screen 200 related to processing for training the RN 102 includes a text box 201 , an operation panel 202 , and a status display window 207 .
[0038] A text box 201 is an area for inputting a link to a text resource obtained from the Internet. The link may be, for example, a copyright-free resource or a resource for which permission has been obtained from the operator.
[0039] The status display window 207 is an area in which various status information is displayed.
[0040] The operation panel 202 includes an acquire button 203 , a training button 204 , a stop button 205 , and a close button 206 .
[0041] When the operator presses (clicks) the acquire button 203, the model generation system 10 executes a text acquisition process to acquire text data via the Internet from the resource of the link entered in the text box 201, and when this process is completed, displays a completion message in the status display window 207. After this, the operator can press the train button 204 to train the RN 102 using the acquired text data.
[0042] When the operator presses the training button 204, the model generation system 10 executes a training process (see FIGS. 4 and 5) for training the RN 102 using the acquired text data. The model generation system 10 displays the execution status of the training process in a status display window 207.
[0043] When the operator presses the stop button 205 after starting the training process of RN 102, the model generation system 10 stops the training process and saves the weights (variables) of RN 102 at the time the training process was stopped in the storage device 14. After that, when the operator presses the training button 204, the model generation system 10 reloads the RN 102 in the state at the time the training process was stopped in the memory 15, and resumes the processing after the time the training process was stopped.
[0044] When the operator presses the close button 206 after the training process is completed, the model generation system 10 saves the variables of the RN 102 after training in the storage device 14.
[0045] Next, a first example of a training process for training the RN 102 by the model generation system 10 will be described.
[0046] FIG. 4 is a diagram illustrating a first example of a training process for training an RN according to an embodiment.
[0047] In the first training process 300, the model generation system 10 inputs acquired text data to the embedding layer 301. The model generation system 10 performs an embedding process to convert the text into a numerical value using the embedding layer 301, and a convolution process on the converted numerical value to convert it into a convolution feature. The model generation system 10 performs a linear convolution on the convolution feature using the projection layer 302 to adjust the size of the data. The model generation system 10 trains the RN 102 using the data output from the projection layer 302.
[0048] Next, a second example of a training process for training the RN 102 by the model generation system 10 will be described.
[0049] FIG. 5 is a diagram illustrating a second example of a training process for training an RN according to an embodiment.
[0050] In the second training process 303, the model generation system 10 inputs the acquired text data to the text line image generation unit 304. The text line image generation unit 304 converts the text data into a text line image using a predetermined available digital font (for example, Arial, MS Gothic, etc.). The model generation system 10 extracts features of the text line image using the FEX 305. Note that the FEX 305 may have the same structure as the FEX 101 but different variables set therein. The variables of the FEX 305 may be determined in advance by training. The model generation system 10 trains the RN 102 using the features output from the FEX 305.
[0051] Next, a generation process for generating a generalized prototype model of a text line recognition model will be described.
[0052] In this embodiment, the model generation system 10 generates a text line recognition model 100 to be trained for generating a prototype model by combining the RN 102 trained by the above-mentioned training process with the FEX 101 that has not been trained. In this text line recognition model 100, the weights (variables) of the RN 102 are frozen (fixed), and training is performed using existing labeled text line images (training text line data). Here, the training text line data is managed by classifying them into domains (an example of an image group by style) for each text line image of the same style. For example, text line images by the same writer are classified into the same domain. Also, for example, in the case of scene text, bank forms, invoices, receipts, and other printed text line images, if they are created with the same font and similar background or texture, they are classified into the same domain.
[0053] FIG. 6 is a diagram showing an example of a GUI screen related to a generation process of a prototype model according to an embodiment.
[0054] A GUI screen 400 for generating a prototype model includes an operation panel 401 and a training status display window 407 .
[0055] The training status display window 407 is an area where training status information is displayed.
[0056] The operation panel 401 includes an input box 402 , an input box 403 , a training button 404 , a stop button 405 , and a close button 406 .
[0057] Input box 402 is an area where the operator inputs the number of tasks (t), which is the number of domains to be used for training. Input box 403 is an area where the operator inputs the number of training text line data samples to be used for training for each domain.
[0058] When the operator presses the training button 404, the model generation system 10 executes a prototype model generation process (see FIG. 7) for training the text line recognition model 100 to be trained and generating prototype data. The model generation system 10 displays the training status in the prototype model generation process in a training status display window 407.
[0059] When the operator presses the stop button 405 after starting the prototype model generation process, the model generation system 10 stops the prototype model generation process and saves the weights (variables) of the text line recognition model 100 at the time the process was stopped in the storage device 14. After that, when the operator presses the train button 404, the model generation system 10 reloads the text line recognition model 100 in the state it was in at the time the process was stopped into the memory 15, and resumes the process after the prototype model generation process was stopped.
[0060] When the operator presses the close button 406 after the prototype model generation process is completed, the model generation system 10 saves the variables of the trained text line recognition model 100 in the storage device 14 .
[0061] Next, a prototype model generation process for generating a prototype model by the model generation system 10 will be described.
[0062] FIG. 7 is a flowchart of a prototype model generation process according to an embodiment.
[0063] In this description, the weight in the prototype model of the text line recognition model 100 is denoted as φ, and the weight in the model (clone model) created as a clone of the prototype model is denoted as φ'.
[0064] The model generation system 10 initializes the weights of the FEX 101 of the internal training rate α, the meta-training rate β, and the text line recognition model 100 to be trained (step 502). Note that the weights of the RN 102 of the text line recognition model 100 to be trained are copied from the RN 102 trained by the training process and frozen in the prototype model generation process.
[0065] Next, the model generation system 10 executes the iterative process 500 to generate (train) a prototype model.
[0066] In the iterative process 500, the model generation system 10 first defines a task (step 503). Specifically, the model generation system 10 randomly selects t domains (the value input to the input box 402) from n domains D = {D1, D2, ···, D n} of the training text line data. Here, t << n. Next, the model generation system 10 randomly extracts two sets T i = {D i sp , D i qr} from the selected domain i. Here, T i represents the data of the i-th domain, D i sp is called the support set and is the set used for training, and D i qr is called the query set and is the set used for evaluating the model. Each set contains s samples (the value input to the input box 403).
[0067] Next, the model generation system 10 creates a clone model of the prototype model (step 504).
[0068] Next, the model generation system 10 repeatedly executes the process 501 for each task using the data of each domain.
[0069] The model generation system 10 is the task TAi About the clone model FEX101 support set D i sp ={I i sp ,L i sp} (step 505), where I i sp is the text line image in the support set, and L i sp are the labels corresponding to the text line images in the support set.
[0070] In the training step 505, the weights φ′ of the clone models are updated as shown in equation (1).
[0071]
number
[0072] where L is the loss function for the model's output and input labels, and ∇ is the gradient of the loss function. f i ∧sp I i sp This is the output of the clone model that inputs
[0073] The model generation system 10 then performs the task TA i For the clone model FEX101, query set D i qr ={I i qr ,L i qr} (step 506), where I i qr are the text line images in the query set, and L i qr are the labels corresponding to the text line images in the query set.
[0074] In the evaluation of step 506, the total evaluation loss L itis updated as shown in equation (2).
[0075]
number
[0076] where f i ∧qr I i qr This is the output of the clone model that inputs
[0077] Next, the model generation system 10 determines whether all tasks have been completed (step 507), and if not all tasks have been completed (step 507: No), the process proceeds to step 505 to process other tasks.
[0078] On the other hand, if all tasks have been completed, i.e., if training and evaluation of the clone models have been completed for all tasks (step 507: Yes), the model generation system 10 updates the weights of the prototype model using the total evaluation loss as shown in equation (3) (step 508).
[0079]
number
[0080] Next, the model generation system 10 determines whether the predetermined number of iterations have been completed (step 509), and if the predetermined number of iterations have not been completed (step 509: No), the process proceeds to step 503 to further execute the iterative process 500. As a result, in each iterative process 500, the prototype model is trained to improve the recognition accuracy of the query set using the support set. Note that by performing the iterative process a sufficient number of times, the prototype model can acquire generalization properties, and high recognition accuracy can be obtained with a small number of training samples.
[0081] On the other hand, if the predetermined number of iterations have been completed (step 509: Yes), the model generation system 10 ends the prototype model generation process.
[0082] Next, a retraining process for retraining the text line recognition model will be described.
[0083] FIG. 8 is a diagram showing an example of a GUI screen related to a retraining process for retraining a text line recognition model according to an embodiment.
[0084] A GUI screen 600 for retraining a text line recognition model includes an operation panel 610 and a window 609 .
[0085] The operation panel 610 includes a new button 601 , an open button 602 , a start application button 603 , a stop button 604 , a recognize button 605 , and a close button 606 .
[0086] When the New button 601 is pressed, the model generation system 10 displays in a window 609 a predetermined number (for example, S) of input areas 607 in which text lines can be written by hand, and S text boxes 608 in which the operator can input labels corresponding to the text lines input in the input areas 607. Note that S may be a number less than 5.
[0087] In addition, when the Open button 602 is pressed, the model generation system 10 displays a window (not shown) in which the operator can select S text line images to be used from the storage device 14, and displays the S text line images selected by the operator in window 609, as well as S text boxes 608 in which the operator can input labels corresponding to the displayed text line images.
[0088] When the operator presses the Start Adaptation button 603, the model generation system 10 starts a retraining process (see FIG. 9) to fine-tune the prototype model using S input samples (pairs of text line images and their corresponding labels) entered in window 609.
[0089] When the operator presses the stop button 604 after starting the retraining process, the model generation system 10 stops the retraining process and saves the weights (variables) of the prototype model at the time the process was stopped in the storage device 14. After that, when the operator presses the start adaptation button 603, the model generation system 10 reloads the prototype model 0 in the state at the time the process was stopped into the memory 15, and resumes the process after the time the retraining process was stopped.
[0090] Furthermore, when the operator inputs a handwritten or selected text line image and then presses the recognize button 605, the model generation system 10 performs text recognition on the input text line image using the prototype model at that point in time, and displays the recognition result in window 609. This allows the operator to test the text recognition of the retrained prototype model.
[0091] When the operator presses the close button 606 after the retraining process is completed, the model generation system 10 stores the variables of the prototype model after the retraining process in the storage device 14. Thereafter, when recognizing text in a text line image, the text line recognition model 100 with these variables set will be used.
[0092] Next, the retraining process of the text line recognition model by the model generation system 10 will be described.
[0093] FIG. 9 is a flow chart of a process for retraining a text line recognition model according to one embodiment.
[0094] The model generation system 10 sets the number of adaptation steps 700 to be executed (the number of adaptation steps) (step 701). The number of adaptation steps may be any number. The model generation system 10 then retrains (fine-tunes) the prototype model using the input samples (step 702).
[0095] The model generation system 10 then determines whether the adaptive number of steps have been completed (step 703), and if not (step P7 If the step is completed (step 03: No), the next adaptation step 700 is executed. P7 In case of (03: Yes), the application completion is displayed in the window 609, and the retraining process is terminated.
[0096] The present invention is not limited to the above-described embodiment, and can be modified as appropriate without departing from the spirit of the present invention.
[0097] For example, in the above-described embodiment, a part or all of the processing performed by the processor may be performed by a hardware circuit. Also, the program in the above-described embodiment may be installed from a program source. The program source may be a program distribution server or a storage medium (e.g., a portable storage medium). [Explanation of symbols]
[0098] 10...model generation system, 11...CPU, 100...text line recognition model, 101...visual feature extractor, 102...linguistic context relational network
Claims
1. A model generation system for generating a text line recognition model for recognizing text lines included in a text line image, comprising: The model generation system includes a processor unit; the text line recognition model includes a visual feature extractor that, when executed by the processor, outputs image features from a text line image; and a linguistic context relational network that, when executed by the processor, inputs the features output from the visual feature extractor and outputs a text line; The processor unit includes: obtaining training text data and training the linguistic contextual relation network using the obtained text data to determine parameters of the linguistic contextual relation network; determining the variables of the visual feature extractor by training the text line recognition model using existing labeled text line images with the variables of the linguistic contextual relation network fixed to the determined variables; Generate the text line recognition model with the variables of the linguistic contextual relation network as the determined variables of the linguistic contextual relation network and the variables of the visual feature extractor as the determined variables of the visual feature extractor. Model generation system.
2. The processor unit adjusts variables of the text line recognition model by training the text line recognition model with less than a predetermined number of labeled text line images. The model generation system of claim 1 .
3. the model generation system is connected to the Internet; The processor unit acquires the training text data via the Internet. The model generation system of claim 1 .
4. The training text data is copyright-free text data available on the Internet. The model generation system of claim 3 .
5. The processor unit includes: Accepting input of a text line image and a label to be attached to the text line image from a user; Training the text line recognition model using the received text line images and labels to adjust the parameters of the text line recognition model. The model generation system of claim 2 .
6. The processor unit includes: In training the language context relation network, text line data for training is acquired, word embedding is performed to convert the acquired text line data into a numerical value, and the converted data is convoluted and then inputted to the language context relation network, thereby training the language context relation network. The model generation system of claim 1 .
7. The processor unit includes: In training the language context relation network, the training text line data is acquired, the acquired text line data is converted into a text line image using a predetermined font, the converted text line image is input into a predetermined visual feature extractor, and an output from the visual feature extractor is input into the language context relation network, thereby training the language context relation network. The model generation system of claim 1 .
8. The existing labeled text line images are managed as a plurality of style-specific image groups each including text line images of the same style; The processor section includes: determining the variables of the visual feature extractor by training the text line recognition model using labeled text line images from each of the plurality of style-specific image sets while fixing the variables of the linguistic contextual relation network to the determined variables; The model generation system of claim 1 .
9. A model generation method for a model generation system that generates a text line recognition model for recognizing text lines included in a text line image, comprising the steps of: the text line recognition model includes a visual feature extractor that, when executed by the model generation system, outputs image features from a text line image; and a linguistic context relational network that, when executed by the model generation system, inputs the features output from the visual feature extractor and outputs a text line; The model generation system comprises: obtaining training text data and training the linguistic contextual relation network using the obtained text data to determine parameters of the linguistic contextual relation network; determining the variables of the visual feature extractor by training the text line recognition model using existing labeled text line images with the variables of the linguistic contextual relation network fixed to the determined variables; Generate the text line recognition model with the variables of the linguistic contextual relation network as the determined variables of the linguistic contextual relation network and the variables of the visual feature extractor as the determined variables of the visual feature extractor. Model generation method.
Citation Information
Patent Citations
Text Recognition
JP2021520561A