Model training method, text line determination method and device
Through the two-stage model training method, combined with simulation and real sample data, the misrecognition problem of the text recognition model in delimiter recognition is solved, and the accurate distinction between line text lines and non-line text lines is achieved, which improves the recognition effect.
Patent Information
- Application Number
- CN202210738482.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-27
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2042-06-27
AI Technical Summary
During the training process, the existing text recognition model cannot effectively recognize the separator because all characters in the text line have the same font, which causes the model to mistakenly recognize the separator as a font, and cannot accurately distinguish between line text lines and non-line text lines, and the recognition effect is poor.
The two-stage model training method is adopted. First, the pre-trained model is trained through simulated sample images, and the separator annotation is added, and then the separator between the real sample images is combined for secondary training to identify the separator between different fonts to improve the recognition accuracy of the model.
By combining two-stage training with simulation and real sample data, the separator between characters in different fonts is effectively recognized, which improves the recognition effect of the model and can accurately distinguish between line text lines and non-line text lines.
Smart Images

Figure CN115019295B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of reinforcement learning, and particularly to a model training method, a text line determination method, and an apparatus. Background Art
[0002] With the continuous improvement of the economic level, the types of entertainment videos are also increasing. People can watch entertainment videos through electronic devices (such as computers, mobile phones, etc.) to enrich their spare time. For the platform that provides entertainment videos, in the process of generating corresponding lines for the entertainment videos within the platform, it is possible to effectively distinguish between the line text of the lines and the non-line text based on the font differences used between different text lines, which plays an important role in filtering all text lines.
[0003] Currently, it is usually adopted to use an optical character recognition network to recognize the line text and non-line text in the video image. In the training of the text line font recognition model, if the training samples used are all from the text lines in the real scene, then almost every text line contains the same font. During the training process of character font attribute recognition, the font attribute corresponding to each character at each position can be obtained through training. However, since all the characters in the text line have the same font, it is impossible to effectively obtain the separator between characters through model training, which will cause the model to misrecognize the separator in the predicted font attribute sequence as a font, resulting in a decrease in the model loss function and a poor recognition effect of the trained model, making it impossible to accurately distinguish between the line text and non-line text in the image. Summary of the Invention
[0004] The purpose of the embodiments of this application is to provide a model training method, a text line determination method, an apparatus, an electronic device, and a storage medium, so as to perform two-stage model training by combining simulation sample data and real sample data, improve the recognition effect of the trained font recognition model, and accurately distinguish between the line text and non-line text in the image. The specific technical solutions are as follows:
[0005] In the first aspect implemented in this application, first, a model training method is provided, including:
[0006] Obtain a first sample image and a second sample image, where both the first sample image and the second sample image are sample images containing text lines, and the first sample image is a simulated image containing multiple text fonts;
[0007] Train a font recognition model to be trained based on the first sample image to obtain a pre-trained font recognition model;
[0008] Train the pre-trained font recognition model based on the first sample image and the second sample image to obtain a target font recognition model.
[0009] Optionally, the obtaining of the first sample image and the second sample image includes:
[0010] Obtaining a second sample image containing text lines from a preset image library;
[0011] Obtaining an initial image without text lines from the preset image library;
[0012] Adding text lines to the initial image to generate a first sample image; each text line in the first sample image contains multiple text fonts.
[0013] Optionally, each character in the first sample image is labeled with a first font label, and a separator is labeled between two adjacent fonts in the same text line;
[0014] The training of the font recognition model to be trained based on the first sample image to obtain a pre-trained font recognition model includes:
[0015] Inputting the first sample image into the font recognition model to be trained;
[0016] Processing the first sample image based on the font recognition model to be trained to obtain predicted font labels of the first sample image containing separators;
[0017] Calculating a first loss value of the font recognition model to be trained according to the first font label and the predicted font label;
[0018] When the first loss value is within a first preset range, determining the trained font recognition model to be trained as the pre-trained font recognition model.
[0019] Optionally, the calculating of the first loss value of the font recognition model to be trained according to the first font label and the predicted font label includes:
[0020] Determining multiple font paths corresponding to each text line in the first sample image according to the predicted font label;
[0021] Determining the font probability of the font to which each character belongs according to the first font label and the predicted font label;
[0022] Calculating the font path probabilities corresponding to the multiple font paths according to the font probability of the font to which each character belongs;
[0023] Calculating the first loss value of the font recognition model to be trained according to the maximum font path probability among the font path probabilities.
[0024] Optionally, after calculating the first loss value of the font recognition model to be trained according to the first font label and the predicted font label, the following steps are further included:
[0025] In the case where the first loss value is outside the first preset range, the font recognition model to be trained after training is trained according to the first sample image until the calculated first loss value is within the first preset range.
[0026] Optionally, each character in the first sample image is labeled with a second font label, each character in the second sample image is labeled with a third font label, and a separator is labeled between two adjacent fonts in the same text line;
[0027] Training the pre-trained font recognition model based on the first sample image and the second sample image to obtain a target font recognition model includes:
[0028] Inputting the first sample image and the second sample image into the pre-trained font recognition model;
[0029] Processing the first sample image and the second sample image based on the pre-trained font recognition model to obtain the first predicted font label of the first sample image and the second predicted font label of the second sample image including the separator;
[0030] Calculating the second loss value of the pre-trained font recognition model according to the second font label and the first predicted font label, and the third font label and the second predicted font label;
[0031] In the case where the second loss value is within the second preset range, the pre-trained font recognition model after training is determined as the target font recognition model.
[0032] Optionally, calculating the second loss value of the pre-trained font recognition model according to the second font label and the first predicted font label, and the third font label and the second predicted font label includes:
[0033] Determining multiple first font paths corresponding to each text line in the first sample image according to the first predicted font label, and determining second font paths corresponding to each text line in the second sample image according to the second predicted font label;
[0034] Determine the first font probability of each character in the first sample image belonging to a font according to the second font tag and the first predicted font tag, and determine the second font probability of each character in the second sample image belonging to a font according to the third font tag and the second predicted font tag;
[0035] Calculate the first font path probability corresponding to the first font path according to the first font probability, and calculate the second font path probability corresponding to the second font path according to the second font path probability;
[0036] Calculate the second loss value according to the largest first font path probability among the first font path probabilities and the largest second font path probability among the second font path probabilities.
[0037] Optionally, after calculating the second loss value of the pre-trained font recognition model according to the second font tag and the first predicted font tag, and the third font tag and the second predicted font tag, it further includes:
[0038] In the case where the second loss value is outside the second preset range, train the trained font recognition model according to the first sample image and the second sample image until the calculated second loss value is within the second preset range.
[0039] In the second aspect of the implementation of the present application, a text line determination method is provided, including:
[0040] Obtain an image to be recognized, where the image to be recognized is an image containing a text line;
[0041] Input the image to be recognized into the target font recognition model;
[0042] Perform recognition processing on the image to be recognized based on the target font recognition model to obtain a text attribute sequence corresponding to the text line in the image to be recognized;
[0043] Determine the dialogue text line and non-dialogue text line in the text line in the image to be recognized according to the text attribute sequence.
[0044] In the third aspect of the implementation of the present application, a model training device is provided, including:
[0045] A sample image acquisition module, configured to acquire a first sample image and a second sample image, both the first sample image and the second sample image are sample images containing text lines, and the first sample image is a simulated image containing multiple text fonts;
[0046] A pre-training model acquisition module, configured to train a font recognition model to be trained based on the first sample image, and obtain a pre-trained font recognition model;
[0047] A target recognition model acquisition module, configured to train the pre-trained font recognition model based on the first sample image and the second sample image, and obtain a target font recognition model.
[0048] Optionally, the sample image acquisition module includes:
[0049] A sample image acquisition unit, configured to obtain a second sample image including text lines from a preset image library;
[0050] An initial image acquisition unit, configured to obtain an initial image without text lines from the preset image library;
[0051] A sample image generation unit, configured to add text lines to the initial image to generate a first sample image; each text line in the first sample image includes multiple text fonts.
[0052] Optionally, each character in the first sample image is labeled with a first font label, and a separator is labeled between two adjacent fonts in the same text line;
[0053] The pre-training model acquisition module includes:
[0054] A first sample image input unit, configured to input the first sample image into the font recognition model to be trained;
[0055] A first predicted font label acquisition unit, configured to process the first sample image based on the font recognition model to be trained, and obtain a predicted font label including a separator of the first sample image;
[0056] A first loss value calculation unit, configured to calculate a first loss value of the font recognition model to be trained according to the first font label and the predicted font label;
[0057] A pre-training model determination unit, configured to determine the trained font recognition model to be trained as the pre-trained font recognition model when the first loss value is within a first preset range.
[0058] Optionally, the first loss value calculation unit includes:
[0059] A first font path determination subunit, configured to determine multiple font paths corresponding to each text line in the first sample image according to the predicted font label;
[0060] A first font probability determination subunit, configured to determine the font probability of each character belonging to a font according to the first font label and the predicted font label;
[0061] A first font path probability calculation subunit, configured to calculate the font path probabilities corresponding to the multiple font paths according to the font probabilities of each character belonging to a font;
[0062] A first loss value calculation subunit, configured to calculate a first loss value of the to-be-trained font recognition model according to the maximum font path probability among the font path probabilities.
[0063] Optionally, the apparatus further includes:
[0064] A first model training module, configured to train the to-be-trained font recognition model according to the first sample image when the first loss value is outside a first preset range until the calculated first loss value is within the first preset range.
[0065] Optionally, each character in the first sample image is labeled with a second font label, each character in the second sample image is labeled with a third font label, and a separator is labeled between two adjacent fonts in the same text line;
[0066] The target recognition model acquisition module includes:
[0067] A second sample image input unit, configured to input the first sample image and the second sample image into the pre-trained font recognition model;
[0068] A second predicted font label acquisition unit, configured to process the first sample image and the second sample image based on the pre-trained font recognition model to obtain a first predicted font label of the first sample image and a second predicted font label of the second sample image including a separator;
[0069] A second loss value calculation unit, configured to calculate a second loss value of the pre-trained font recognition model according to the second font label and the first predicted font label, and the third font label and the second predicted font label;
[0070] A target recognition model determination unit, configured to determine the trained pre-trained font recognition model as the target font recognition model when the second loss value is within a second preset range.
[0071] Optionally, the second loss value calculation unit includes:
[0072] A second font path determination subunit, configured to determine multiple first font paths corresponding to each text line in the first sample image according to the first predicted font label, and determine a second font path corresponding to each text line in the second sample image according to the second predicted font label;
[0073] A second font probability determination subunit, configured to determine a first font probability of the font to which each character in the first sample image belongs according to the second font label and the first predicted font label, and determine a second font probability of the font to which each character in the second sample image belongs according to the third font label and the second predicted font label;
[0074] A second font path probability calculation subunit, configured to calculate a first font path probability corresponding to the first font path according to the first font probability, and calculate a second font path probability corresponding to the second font path according to the second font path probability;
[0075] A second loss value calculation subunit, configured to calculate the second loss value according to the maximum first font path probability among the first font path probabilities and the maximum second font path probability among the second font path probabilities.
[0076] Optionally, the apparatus further includes:
[0077] A second model training module, configured to, when the second loss value is outside a second preset range, train the trained font recognition model according to the first sample image and the second sample image until the calculated second loss value is within the second preset range.
[0078] In a fourth aspect of the implementation of the present application, a text line determination apparatus is provided, including:
[0079] An image to be recognized acquisition module, configured to acquire an image to be recognized, where the image to be recognized is an image containing text lines;
[0080] An image to be recognized input module, configured to input the image to be recognized into a target font recognition model;
[0081] A text attribute sequence acquisition module, configured to perform recognition processing on the image to be recognized based on the target font recognition model to obtain a text attribute sequence corresponding to the text lines in the image to be recognized;
[0082] A dialogue text line determination module, configured to determine dialogue text lines and non-dialogue text lines in the text lines in the image to be recognized according to the text attribute sequence.
[0083] In a fifth aspect of the implementation of this application, an electronic device is provided, including:
[0084] At least one processor; and
[0085] A memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the model training method described in any one of the above, or the above text line determination method.
[0086] In a sixth aspect of the implementation of this application, a non-transitory computer-readable storage medium storing computer instructions is provided, and the computer instructions are used to cause the computer to execute the model training method described in any one of the above, or the above text line determination method.
[0087] In a seventh aspect of the implementation of this application, a computer program product is provided, including a computer program, and the computer program realizes the model training method described in any one of the above, or the above text line determination method when executed by a processor.
[0088] The model training method, text line determination method, device, electronic device and storage medium provided by the embodiments of this application, by obtaining a first sample image and a second sample image, both the first sample image and the second sample image are sample images containing text lines, and the first sample image is an image simulated with multiple text fonts. The font recognition model to be trained is trained based on the first sample image to obtain a pre-trained font recognition model. The pre-trained font recognition model is trained based on the first sample image and the second sample image to obtain a target font recognition model. The embodiments of this application perform two-stage model training by combining simulation sample data and real sample data, adding texts of different fonts in the simulation sample data, and effectively identifying the delimiters between characters of different fonts during the model training process, so as to avoid the model misidentifying the delimiters in the predicted font attribute sequence as fonts, improve the recognition effect of the trained model, and further accurately distinguish the dialogue text lines and non-dialogue text lines in the image. BRIEF DESCRIPTION OF THE DRAWINGS
[0089] In order to more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art.
[0090] Figure 1 It is a flowchart of the steps of a model training method provided by an embodiment of this application;
[0091] Figure 2 It is a flowchart of the steps of a sample image acquisition method provided by an embodiment of this application;
[0092] Figure 3 It is a flowchart of the steps of a method for pre-training a font recognition model provided by an embodiment of the present application;
[0093] Figure 4 It is a flowchart of the steps of a method for calculating a first loss value provided by an embodiment of the present application;
[0094] Figure 5 It is a schematic diagram of the calculation path of a simulation sample provided by an embodiment of the present application;
[0095] Figure 6 It is a flowchart of the steps of a method for training a target font recognition model provided by an embodiment of the present application;
[0096] Figure 7 It is a flowchart of the steps of a method for calculating a second loss value provided by an embodiment of the present application;
[0097] Figure 8 It is a flowchart of the steps of a method for determining a text line provided by an embodiment of the present application;
[0098] Figure 9 It is a schematic diagram of the structure of a model training device provided by an embodiment of the present application;
[0099] Figure 10 It is a schematic diagram of the structure of a text line determination device provided by an embodiment of the present application;
[0100] Figure 11 It is a schematic diagram of the structure of an electronic device provided by an embodiment of the present application. Detailed implementation manners
[0101] Next, the technical solutions in the embodiments of the present application will be described with reference to the accompanying drawings in the embodiments of the present application.
[0102] Referring to Figure 1 , a flowchart of the steps of a model training method provided by an embodiment of the present application is shown. As Figure 1 shown, the model training method may include the following steps:
[0103] Step 101: Obtain a first sample image and a second sample image. Both the first sample image and the second sample image are sample images containing text lines, and the first sample image is an image simulating multiple text fonts.
[0104] The embodiments of the present application can be applied to a scenario of training a font recognition model by combining simulation sample images and real sample images.
[0105] In this example, both the first sample image and the second sample image are sample images containing text lines. Among them, the first sample image is a simulated image containing various text fonts, that is, the first sample image is a simulation sample image.
[0106] The second sample image is an image obtained by combining real source data, that is, an image containing line text randomly selected within a video playback platform.
[0107] When training a font recognition model, a simulation sample image (i.e., the first sample image) and a real sample image (i.e., the second sample image) can be obtained. The acquisition process of the simulation sample image and the real sample image can be combined Figure 2 and described in detail as follows.
[0108] Refer to Figure 2 , which shows the flowchart of the steps of a sample image acquisition method provided by an embodiment of the present application. As Figure 2 shown, the sample image acquisition method may include: step 201, step 202, and step 203.
[0109] Step 201: Obtain a second sample image containing text lines from a preset image library.
[0110] In this embodiment, when training a font recognition model, an image containing text lines can be obtained from a preset image library as the second sample image. Specifically, the OCR (Optical Character Recognition) technology can be used to identify each image in the preset image library one by one to identify an image containing text lines as the second sample image.
[0111] Step 202: Obtain an initial image that does not contain text lines.
[0112] The initial image refers to an image that does not contain text lines.
[0113] When making a simulation image, an initial image that does not contain text lines can be obtained. In a specific implementation, the initial image can be an image extracted from a preset image library. For example, when screening the second sample image from the preset image library, the unselected image can be used as the initial image. The initial image can also be an image taken currently, or an image downloaded from the Internet, etc. Specifically, the acquisition method of the initial image can be determined according to usage requirements, and this embodiment does not limit this.
[0114] After obtaining the initial image that does not contain text lines, step 203 is executed.
[0115] Step 203: Add text lines to the initial image to generate a first sample image; each text line in the first sample image contains multiple text fonts.
[0116] After obtaining the initial image that does not contain text lines, text lines (i.e., dialogue text lines) can be added to the initial image to generate a first sample image. Specifically, an image containing text lines can be pre-generated, and then this image is subjected to image synthesis processing with the initial image to generate the first sample image. Alternatively, a preset image editing tool can be used to process the initial image to edit text lines in the initial image to obtain the first sample image, etc.
[0117] In this example, each text line in the dialogue text lines added to the initial image contains multiple text fonts, that is, each line contains multiple (two or more) text fonts. For example, the text line added to the initial image is one line, and the fonts in this text line include: Song typeface, boldface, regular script, etc.
[0118] It can be understood that the above examples are only examples listed for better understanding of the technical solutions of the embodiments of the present application and do not serve as the sole limitation of this embodiment.
[0119] After obtaining the first sample image and the second sample image, step 102 is executed.
[0120] Step 102: Train the font recognition model to be trained based on the first sample image to obtain a pre-trained font recognition model.
[0121] After obtaining the first sample image, the font recognition model to be trained can be trained based on the first sample image first to obtain a pre-trained font recognition model. Specifically, the model can be trained with the simulated first sample image until the model converges, so as to obtain a pre-trained font recognition model.
[0122] For the pre-training process of the model, it can be combined with Figure 3 and described in detail as follows.
[0123] Refer to Figure 3 , which shows the step flowchart of a font recognition model pre-training method provided by an embodiment of the present application. As Figure 3 shown, the font recognition model pre-training method may include: step 301, step 302, step 303, step 304, and step 305.
[0124] Step 301: Input the first sample image into the font recognition model to be trained.
[0125] In this embodiment, the first sample image corresponds to a first font attribute label. Each character in the first sample image is labeled with a first font label, and a separator is labeled between two adjacent fonts in the same text line.
[0126] After obtaining the first sample image, the first sample image can be input into the font recognition model to be trained.
[0127] After inputting the first sample image into the font recognition model to be trained, step 302 is executed.
[0128] Step 302: Process the first sample image based on the font recognition model to be trained to obtain the predicted font label of the first sample image that contains the separator.
[0129] The predicted font label refers to the label of each character attribute in the text line of the first sample image predicted by the font recognition model to be trained.
[0130] After inputting the first sample image into the font recognition model to be trained, the first sample image can be processed based on the font recognition model to be trained to obtain the predicted font label of the first sample image that contains the separator. The predicted font label includes the font of each character in the text line of the predicted first sample image (i.e., Song typeface, Kai typeface, boldface, etc.) and the separator between two fonts.
[0131] After processing the first sample image based on the font recognition model to be trained to obtain the predicted font label of the first sample image that contains the separator, step 303 is executed.
[0132] Step 303: Calculate the first loss value of the font recognition model to be trained according to the first font label and the predicted attribute label.
[0133] The first loss value refers to the loss value of the font recognition model to be trained calculated when using the simulated sample image to train the font recognition model to be trained.
[0134] After processing the first sample image based on the font recognition model to be trained to obtain the predicted font attribute label of the first sample image, the first loss value of the font recognition model to be trained can be calculated according to the first font attribute label and the predicted font attribute label. The calculation process of the first loss value can be described in detail as follows in combination with Figure 4 as follows.
[0135] Referring to Figure 4 , a step flowchart of a first loss value calculation method provided by an embodiment of the present application is shown. As Figure 4As shown, the first loss calculation method may include: step 401, step 402, step 403, and step 404.
[0136] Step 401: Determine multiple font paths corresponding to each text line in the first sample image according to the predicted font label.
[0137] In this embodiment, after obtaining the predicted font label of the first sample image containing delimiters, multiple font paths corresponding to each text line in the first sample image can be determined according to the predicted font label. As Figure 5 shown, Figure 5 each circle in the figure represents a predicted font label. The circles in the first row represent the predicted font labels of the delimiter "-", the circles in the second row represent the predicted font labels of "Xie",..., the circles in the 7th row represent the predicted font labels of the delimiter "-". The font labels in the same column are the same, and the font labels in different columns are different. The text contained in the simulation sample image is "Thank you", and the calculation paths are as Figure 5 shown. There are a total of 7 positions (i.e., four delimiters and three characters) with predicted font labels. Then, multiple font paths can be generated according to the predicted font labels, as Figure 5 shown by the connection lines.
[0138] It can be understood that the above examples are only examples listed for better understanding the technical solutions of the embodiments of the present application and do not serve as the sole limitation of this embodiment.
[0139] Step 402: Determine the font probability of each character belonging to a font according to the first font label and the predicted font label.
[0140] After obtaining the predicted font label, the font probability of each character belonging to a font can be determined according to the first font label of each character marked in the first sample image and the predicted font label of each character, that is, the probability of each character belonging to a font. In this example, the font probability of each character belonging to a font is output by the model. During the actual training process, the true font label (i.e., the first font label) can guide the font recognition model to be trained to predict multiple fonts to which each character belongs (i.e., each character corresponds to multiple predicted font labels) and output the font probability of each character belonging to a font.
[0141] After calculating the font probability of each character belonging to a font, step 403 is executed.
[0142] Step 403: Calculate the font path probability corresponding to the multiple font paths according to the font probability of each character belonging to a font.
[0143] After calculating the font probabilities of each character belonging to a font, based on the font probabilities of each character belonging to a font, the font path probabilities corresponding to multiple font paths can be calculated, that is, the font probabilities of each character belonging to a font on each font path are added together, and the resulting probability sum value is used as the font path probability of the font path.
[0144] After calculating the font path probabilities corresponding to multiple font paths based on the font probabilities of each character belonging to a font, step 404 is executed.
[0145] Step 404: Calculate a first loss value of the to-be-trained font recognition model according to the maximum font path probability among the font path probabilities.
[0146] After calculating the font path probabilities corresponding to multiple font paths based on the font probabilities of each character belonging to a font, a first loss value of the to-be-trained font recognition model can be calculated according to the maximum font path probability among the font path probabilities. Specifically, this maximum font path probability can be used as the first loss value of the to-be-trained font recognition model. As Figure 5 shown, when calculating the first loss value, the corresponding probability of this path can be calculated respectively from Figure 5 all the paths shown, then select a path with the maximum probability as the optimal path from all the paths, and calculate the sum value of the probabilities of each prediction position of this optimal path as the first loss value of the to-be-trained font recognition model, etc.
[0147] After calculating the first loss value of the to-be-trained font recognition model, step 304 is executed.
[0148] Step 304: In the case where the first loss value is within a first preset range, determine the trained to-be-trained font recognition model as the pre-trained font recognition model.
[0149] The first preset range refers to the range of loss values when the to-be-trained font recognition model converges, which is preset.
[0150] After calculating the first loss value of the to-be-trained font recognition model, it can be determined whether the first loss value is within the first preset range.
[0151] When the first loss value is within the first preset range, it indicates that the to-be-trained font recognition model has converged. At this time, the trained to-be-trained font recognition model can be used as the pre-trained font recognition model.
[0152] Step 305: In the case where the first loss value is outside the first preset range, train the trained to-be-trained font recognition model according to the first sample image until the calculated first loss value is within the first preset range.
[0153] When the first loss value is not within the first preset range, the simulation sample images can be continuously used to train the font recognition model to be trained until the model converges (that is, until the calculated first loss value is within the first preset range).
[0154] After training the pre-trained font recognition model, step 103 is executed.
[0155] Step 103: Train the pre-trained font recognition model based on the first sample image and the second sample image to obtain a target font recognition model.
[0156] After training the pre-trained font recognition model, the pre-trained font recognition model can be trained based on the first sample image and the second sample image to obtain a target font recognition model, that is, the pre-trained font recognition model is retrained by combining the simulation sample images and the real sample images until the pre-trained font recognition model converges, so that the target font recognition model can be obtained.
[0157] The training process of the pre-trained font recognition model can be combined with Figure 6 to be described in detail as follows.
[0158] Referring to Figure 6 , a step flowchart of a method for training a target font recognition model provided by an embodiment of the present application is shown. As Figure 6 shown, the method for training the target font recognition model may include: step 601, step 602, and step 603.
[0159] Step 601: Input the first sample image and the second sample image into the pre-trained font recognition model.
[0160] In this embodiment, each character in the first sample image is labeled with a second font label, each character in the second sample image is labeled with a third font label, and a separator is labeled between two adjacent fonts in the same text line.
[0161] After training the pre-trained font recognition model, the first sample image and the second sample image can be input into the pre-trained font recognition model to retrain the pre-trained font recognition model through the first sample image and the second sample image.
[0162] After inputting the first sample image and the second sample image into the pre-trained font recognition model, step 602 is executed.
[0163] Step 602: Process the first sample image and the second sample image based on the pre-trained font recognition model to obtain a first predicted font label for the first sample image and a second predicted font label containing delimiters for the second sample image.
[0164] The first predicted label refers to the label of the font to which each character in the text line of the first sample image belongs, predicted by the pre-trained font recognition model.
[0165] The second predicted label refers to the label of the font to which each character in the text line of the second sample image belongs, predicted by the pre-trained font recognition model.
[0166] After inputting the first sample image and the second sample image into the pre-trained font recognition model, the first sample image and the second sample image can be processed based on the pre-trained font recognition model to obtain a first predicted font label for the first sample image and a second predicted font label containing delimiters for the second sample image.
[0167] After obtaining the first predicted font label and the second predicted font label, step 603 is executed.
[0168] Step 603: Calculate a second loss value of the pre-trained font recognition model according to the second font label and the first predicted font label, and the third font label and the second predicted font label.
[0169] The second loss value refers to the loss value of the pre-trained font recognition model calculated when training the pre-trained font recognition model with simulated sample images and real sample images.
[0170] After obtaining the first predicted font label and the second predicted font label, a second loss value of the pre-trained font recognition model can be calculated according to the second font label and the first predicted label, and the third font label and the second predicted font label. The calculation process of the second loss value can be combined with Figure 6 and is described in detail as follows.
[0171] Refer to Figure 7 , which shows a flowchart of the steps of a method for calculating a second loss value provided by an embodiment of the present application. As Figure 7 shown, the method for calculating the second loss value may include: step 701, step 702, step 703, step 704, step 705, and step 706.
[0172] Step 701: Determine multiple first font paths corresponding to each text line in the first sample image according to the first predicted font label, and determine second font paths corresponding to each text line in the second sample image according to the second predicted font label.
[0173] In this embodiment, after obtaining the first predicted font label and the second predicted font label, multiple first font paths corresponding to each text line in the first sample image can be determined according to the first predicted font label, and the second font path corresponding to each text line in the second sample image can be determined according to the second predicted font label. The implementation manner of this step 701 is similar to that of the above step 401, and will not be elaborated herein in this embodiment.
[0174] Step 702: Determine the first font probability of the font to which each character in the first sample image belongs according to the second font label and the first predicted font label, and determine the second font probability of the font to which each character in the second sample image belongs according to the third font label and the second predicted font label.
[0175] After obtaining the first predicted font label and the second predicted font label, the first font probability of the font to which each character in the first sample image belongs can be determined according to the second font label and the first predicted font label. And the second font probability of the font to which each character in the second sample image belongs can be determined according to the third font label and the second predicted font label. The specific implementation process can refer to the description of the above step 402, and will not be elaborated herein in this example.
[0176] After calculating the first font probability and the second font probability, step 703 is executed.
[0177] Step 703: Calculate the first font path probability corresponding to the first font path according to the first font probability, and calculate the second font path probability corresponding to the second font path according to the second font path probability.
[0178] After calculating the first font probability of the font to which each character in the first sample image belongs, the first font path probability of multiple first font paths can be calculated according to the first font probability of the font to which each character in the first sample image belongs, that is, adding the first font probabilities of the fonts to which each character of each first font path belongs, and taking the obtained probability sum value as the first font path probability of this first font path.
[0179] After calculating the second font probability of the font to which each character in the second sample image belongs, the second font path probability of multiple second font paths can be calculated according to the second font probability of the font to which each character in the second sample image belongs, that is, adding the second font probabilities of the fonts to which each character of each second font path belongs, and taking the obtained probability sum value as the second font path probability of this second font path.
[0180] After obtaining the first font path probability and the second font path probability, step 704 is executed.
[0181] Step 704: Calculate the second loss value according to the largest first font path probability among the first font path probabilities and the largest second font path probability among the second font path probabilities.
[0182] After obtaining the first font path probability and the second font path probability, the largest first font path probability among the first font path probabilities and the largest second font path probability among the second font path probabilities can be obtained, and the second loss value of the pre-trained font recognition model is calculated according to the largest first font path probability and the largest second font path probability. Specifically, the sum value of the largest first font path probability and the largest second font path probability can be calculated and used as the second loss value.
[0183] After calculating the second loss value of the pre-trained font recognition model, step 705 is executed.
[0184] Step 705: When the second loss value is within the second preset range, determine the trained pre-trained font recognition model as the target font recognition model.
[0185] The second preset range refers to the range of loss values when the pre-trained font recognition model converges, which is set in advance.
[0186] After calculating the second loss value of the pre-trained font recognition model, it can be determined whether the second loss value is within the second preset range.
[0187] If the second loss value is within the second preset range, it indicates that the pre-trained font recognition model converges. At this time, the trained pre-trained font recognition model can be used as the target font recognition model.
[0188] Step 706: When the second loss value is outside the second preset range, train the trained font recognition model according to the first sample image and the second sample image until the calculated second loss value is within the second preset range.
[0189] If the second loss value is not within the second preset range, at this time, the pre-trained font recognition model can be continuously trained in combination with the first sample image and the second sample image until the model converges (that is, until the calculated second loss value is within the second preset range).
[0190] In the prior art, when training a font recognition model using sample images with the same font, during the training process, multiple font paths can still be generated. However, due to using the same font, in all output positions, it is more likely to predict this position as the label of the font, resulting in the inability to predict the empty label " ". During the process of predicting the font as the corresponding font label, the convergence direction of the loss is correct, but it will cause the positions where the label " " should appear to be mispredicted as font labels. And as the training process progresses and the loss gradually decreases, the probability that all positions are predicted as font labels becomes greater. Although the model predicted using the same font also has an optimal path, the probability corresponding to the position where " " should appear in this path and the label " " will be much lower than the labels corresponding to the fonts existing in the text line. However, if a font recognition model already trained in C1 is used and real - world training samples are added to the highly simulated data, at this time, since the model already has good recognition ability for the font attributes of text line characters, after adding some real - world samples (only one font in the same text line) to the training samples, the text recognition effect of the trained font recognition model can be improved.
[0191] The model training method provided by the embodiments of the present application obtains a first sample image and a second sample image. Both the first sample image and the second sample image are sample images containing text lines, and the first sample image is a simulated image containing multiple text fonts. The to - be - trained font recognition model is trained based on the first sample image to obtain a pre - trained font recognition model. The pre - trained font recognition model is trained based on the first sample image and the second sample image to obtain a target font recognition model. By combining simulation sample data and real sample data for two - stage model training, adding texts with different fonts to the simulation sample data, effectively identifying the delimiters between characters of different fonts during the model training process, the embodiments of the present application can avoid the model misidentifying the delimiters in the predicted font attribute sequence as fonts, improve the recognition effect of the trained model, and further accurately distinguish the dialogue text lines and non - dialogue text lines in the image.
[0192] Referring to Figure 8 , a flowchart showing the steps of a text line determination method provided by the embodiments of the present application is shown. As Figure 8 shown, the text line determination method may include the following steps:
[0193] Step 801: Obtain an image to be recognized, where the image to be recognized is an image containing a text line.
[0194] The embodiments of the present application can be applied to the scenario where the target font recognition model trained in combination with the above embodiments is used to recognize the line text and non-line text in an image.
[0195] This embodiment can be applied to the process of producing lines in a video image on a video website. According to the font differences used between different text lines in the video image, the line text and non-line text can be effectively distinguished, which plays an important role in filtering all text lines.
[0196] The image to be recognized refers to an image used to distinguish the line text and non-line text in an image. The image to be recognized is an image containing text lines. In a specific implementation, the image to be recognized can be a video frame image containing text lines in a video played on a video playback platform.
[0197] After the target font recognition model is trained, during the application process of the target font recognition model, an image to be recognized containing text lines can be obtained.
[0198] After the image to be recognized is obtained, step 802 is executed.
[0199] Step 802: Input the image to be recognized into the target font recognition model.
[0200] After the image to be recognized is obtained, the image to be recognized can be input into the target font recognition model so that the target font recognition model can recognize the font of the text in the image to be recognized.
[0201] After the image to be recognized is input into the target font recognition model, step 803 is executed.
[0202] Step 803: Perform recognition processing on the image to be recognized based on the target font recognition model to obtain a text attribute sequence corresponding to the text lines in the image to be recognized.
[0203] After the image to be recognized is input into the target font recognition model, recognition processing can be performed on the image to be recognized based on the target font recognition model to obtain a text attribute sequence corresponding to the text lines in the image to be recognized.
[0204] After the text attribute sequence corresponding to the text lines in the image to be recognized is obtained, step 204 is executed.
[0205] Step 804: Determine the line text and non-line text in the text lines in the image to be recognized according to the text attribute sequence.
[0206] After obtaining the text attribute sequence corresponding to the text line in the image to be recognized, the line of dialogue text and the non-dialogue text line in the text line within the image to be recognized can be determined according to the text attribute sequence. Specifically, when it is recognized that only one font is included in a certain text line, it is determined that the text line is a line of dialogue text. When it is recognized that two or more fonts are included in a certain text line, it is determined that the text line is a non-dialogue text line, such as billboard text or bullet screen text, etc. By this method, non-dialogue text lines can be effectively filtered out to produce effective lines of dialogue text.
[0207] The method for determining a text line provided by an embodiment of the present application includes: obtaining an image to be recognized, where the image to be recognized is an image containing a text line; inputting the image to be recognized into a target font recognition model; performing recognition processing on the image to be recognized based on the target font recognition model to obtain a text attribute sequence corresponding to the text line in the image to be recognized; and determining the line of dialogue text and the non-dialogue text line in the text line within the image to be recognized according to the text attribute sequence. The target font recognition model obtained by the embodiment of the present application through two-stage model training by combining simulated sample data and real sample data can accurately distinguish the line of dialogue text and the non-dialogue text line in the image.
[0208] Referring to Figure 9 , a schematic structural diagram of a model training device provided by an embodiment of the present application is shown. As Figure 9 shown, the model training device 900 may include the following modules:
[0209] A sample image acquisition module 910, configured to acquire a first sample image and a second sample image, where both the first sample image and the second sample image are sample images containing text lines, and the first sample image is a simulated image containing multiple text fonts;
[0210] A pre-trained model acquisition module 920, configured to train a font recognition model to be trained based on the first sample image to obtain a pre-trained font recognition model;
[0211] A target recognition model acquisition module 930, configured to train the pre-trained font recognition model based on the first sample image and the second sample image to obtain a target font recognition model.
[0212] Optionally, the sample image acquisition module 910 includes:
[0213] A sample image acquisition unit, configured to acquire a second sample image containing a text line from a preset image library;
[0214] An initial image acquisition unit, configured to acquire an initial image that does not contain a text line from the preset image library;
[0215] A sample image generation unit, configured to add text lines to the initial image to generate a first sample image; each text line in the first sample image contains multiple text fonts.
[0216] Optionally, each character in the first sample image is labeled with a first font label, and a separator is labeled between two adjacent fonts in the same text line;
[0217] The pre-trained model acquisition module 920 includes:
[0218] A first sample image input unit, configured to input the first sample image into the font recognition model to be trained;
[0219] A first predicted font label acquisition unit, configured to process the first sample image based on the font recognition model to be trained, and obtain the predicted font label containing the separator of the first sample image;
[0220] A first loss value calculation unit, configured to calculate a first loss value of the font recognition model to be trained according to the first font label and the predicted font label;
[0221] A pre-trained model determination unit, configured to determine the trained font recognition model to be trained as the pre-trained font recognition model when the first loss value is within a first preset range.
[0222] Optionally, the first loss value calculation unit includes:
[0223] A first font path determination subunit, configured to determine multiple font paths corresponding to each text line in the first sample image according to the predicted font label;
[0224] A first font probability determination subunit, configured to determine the font probability of the font to which each character belongs according to the first font label and the predicted font label;
[0225] A first font path probability calculation subunit, configured to calculate the font path probability corresponding to the multiple font paths according to the font probability of the font to which each character belongs;
[0226] A first loss value calculation subunit, configured to calculate a first loss value of the font recognition model to be trained according to the maximum font path probability in the font path probability.
[0227] Optionally, the device further includes:
[0228] A first model training module, configured to train the to-be-trained font recognition model after training according to the first sample image when the first loss value is outside a first preset range until the calculated first loss value is within the first preset range.
[0229] Optionally, each character in the first sample image is labeled with a second font label, each character in the second sample image is labeled with a third font label, and a separator is labeled between two adjacent fonts in the same text line;
[0230] The target recognition model acquisition module 930 includes:
[0231] A second sample image input unit, configured to input the first sample image and the second sample image into the pre-trained font recognition model;
[0232] A second predicted font label acquisition unit, configured to process the first sample image and the second sample image based on the pre-trained font recognition model to obtain a first predicted font label of the first sample image and a second predicted font label containing a separator of the second sample image;
[0233] A second loss value calculation unit, configured to calculate a second loss value of the pre-trained font recognition model according to the second font label and the first predicted font label, and the third font label and the second predicted font label;
[0234] A target recognition model determination unit, configured to determine the trained pre-trained font recognition model as the target font recognition model when the second loss value is within a second preset range.
[0235] Optionally, the second loss value calculation unit includes:
[0236] A second font path determination subunit, configured to determine multiple first font paths corresponding to each text line in the first sample image according to the first predicted font label, and determine a second font path corresponding to each text line in the second sample image according to the second predicted font label;
[0237] A second font probability determination subunit, configured to determine a first font probability of the font to which each character in the first sample image belongs according to the second font label and the first predicted font label, and determine a second font probability of the font to which each character in the second sample image belongs according to the third font label and the second predicted font label;
[0238] A second font path probability calculation subunit, configured to calculate a first font path probability corresponding to the first font path according to the first font probability, and calculate a second font path probability corresponding to the second font path according to the second font path probability;
[0239] A second loss value calculation subunit, configured to calculate the second loss value according to the maximum first font path probability in the first font path probabilities and the maximum second font path probability in the second font path probabilities.
[0240] Optionally, the apparatus further includes:
[0241] A second model training module, configured to, when the second loss value is outside a second preset range, train the trained font recognition model according to the first sample image and the second sample image until the calculated second loss value is within the second preset range.
[0242] The model training apparatus provided by an embodiment of the present application obtains a first sample image and a second sample image. Both the first sample image and the second sample image are sample images containing text lines, and the first sample image is a simulated image containing multiple text fonts. The to-be-trained font recognition model is trained based on the first sample image to obtain a pre-trained font recognition model. The pre-trained font recognition model is trained based on the first sample image and the second sample image to obtain a target font recognition model. By combining simulation sample data and real sample data for two-stage model training, adding texts of different fonts to the simulation sample data, and effectively identifying the delimiters between characters of different fonts during the model training process, the model provided by an embodiment of the present application can avoid misidentifying the delimiters in the predicted font attribute sequence as fonts, improve the recognition effect of the trained model, and further accurately distinguish the dialogue text lines and non-dialogue text lines in the image.
[0243] Referring to Figure 10 , a schematic structural diagram of a text line determination apparatus provided by an embodiment of the present application is shown. As Figure 10 shown, the text line determination apparatus 1000 may include the following modules:
[0244] An image to be recognized acquisition module 1010, configured to acquire an image to be recognized, where the image to be recognized is an image containing text lines;
[0245] An image to be recognized input module 1020, configured to input the image to be recognized into the target font recognition model;
[0246] The text attribute sequence acquisition module 1030 is configured to perform recognition processing on the image to be recognized based on the target font recognition model, and obtain a text attribute sequence corresponding to the text line in the image to be recognized;
[0247] The line text determination module is configured to determine the line text and non-line text in the text lines in the image to be recognized according to the text attribute sequence.
[0248] The text line determination device provided by the embodiment of the present application obtains an image to be recognized, where the image to be recognized is an image containing text lines, inputs the image to be recognized into a target font recognition model, performs recognition processing on the image to be recognized based on the target font recognition model, obtains a text attribute sequence corresponding to the text line in the image to be recognized, and determines the line text and non-line text in the text lines in the image to be recognized according to the text attribute sequence. The target font recognition model obtained by the embodiment of the present application through two-stage model training by combining simulation sample data and real sample data can accurately distinguish line text and non-line text in the image.
[0249] The embodiment of the present application also provides an electronic device, as Figure 11 shown, including a processor 1101, a communication interface 1102, a memory 1103, and a communication bus 1104. Among them, the processor 1101, the communication interface 1102, and the memory 1103 complete mutual communication through the communication bus 1104.
[0250] The memory 1103 is used to store a computer program;
[0251] When the processor 1101 is configured to execute the program stored on the memory 1103, the following steps are implemented:
[0252] Obtain a first sample image and a second sample image, where both the first sample image and the second sample image are sample images containing text lines, and the first sample image is an image simulated with multiple text fonts;
[0253] Train the font recognition model to be trained based on the first sample image to obtain a pre-trained font recognition model;
[0254] Train the pre-trained font recognition model based on the first sample image and the second sample image to obtain a target font recognition model.
[0255] Optionally, the obtaining of the first sample image and the second sample image includes:
[0256] Obtain a second sample image containing text lines from a preset image library;
[0257] Obtain an initial image without text lines from the preset image library;
[0258] Add text lines to the initial image to generate a first sample image; each text line in the first sample image contains multiple text fonts.
[0259] Optionally, each character in the first sample image is labeled with a first font label, and a separator is labeled between two adjacent fonts in the same text line;
[0260] Training the font recognition model to be trained based on the first sample image to obtain a pre-trained font recognition model, including:
[0261] Input the first sample image into the font recognition model to be trained;
[0262] Process the first sample image based on the font recognition model to be trained to obtain the predicted font label of the first sample image containing the separator;
[0263] Calculate the first loss value of the font recognition model to be trained according to the first font label and the predicted font label;
[0264] When the first loss value is within the first preset range, determine the trained font recognition model to be trained as the pre-trained font recognition model.
[0265] Optionally, the calculating the first loss value of the font recognition model to be trained according to the first font label and the predicted font label includes:
[0266] Determine multiple font paths corresponding to each text line in the first sample image according to the predicted font label;
[0267] Determine the font probability of the font to which each character belongs according to the first font label and the predicted font label;
[0268] Calculate the font path probability corresponding to the multiple font paths according to the font probability of the font to which each character belongs;
[0269] Calculate the first loss value of the font recognition model to be trained according to the maximum font path probability among the font path probabilities.
[0270] Optionally, after calculating the first loss value of the font recognition model to be trained according to the first font label and the predicted font label, it further includes:
[0271] In the case where the first loss value is outside the first preset range, the to-be-trained font recognition model after training is trained according to the first sample image until the calculated first loss value is within the first preset range.
[0272] Optionally, each character in the first sample image is labeled with a second font label, each character in the second sample image is labeled with a third font label, and a separator is labeled between two adjacent fonts in the same text line;
[0273] The training of the pre-trained font recognition model based on the first sample image and the second sample image to obtain a target font recognition model includes:
[0274] Input the first sample image and the second sample image into the pre-trained font recognition model;
[0275] Based on the pre-trained font recognition model, process the first sample image and the second sample image to obtain a first predicted font label of the first sample image and a second predicted font label of the second sample image including a separator;
[0276] According to the second font label and the first predicted font label, and the third font label and the second predicted font label, calculate a second loss value of the pre-trained font recognition model;
[0277] In the case where the second loss value is within the second preset range, determine the trained pre-trained font recognition model as the target font recognition model.
[0278] Optionally, the calculating the second loss value of the pre-trained font recognition model according to the second font label and the first predicted font label, and the third font label and the second predicted font label includes:
[0279] According to the first predicted font label, determine multiple first font paths corresponding to each text line in the first sample image, and according to the second predicted font label, determine second font paths corresponding to each text line in the second sample image;
[0280] According to the second font label and the first predicted font label, determine a first font probability of the font to which each character in the first sample image belongs, and according to the third font label and the second predicted font label, determine a second font probability of the font to which each character in the second sample image belongs;
[0281] Based on the first font probability, calculate the first font path probability corresponding to the first font path, and based on the second font path probability, calculate the second font path probability corresponding to the second font path;
[0282] Calculate the second loss value according to the largest first font path probability among the first font path probabilities and the largest second font path probability among the second font path probabilities.
[0283] Optionally, after calculating the second loss value of the pre-trained font recognition model according to the second font label and the first predicted font label, and the third font label and the second predicted font label, it further includes:
[0284] In the case where the second loss value is outside the second preset range, train the trained font recognition model according to the first sample image and the second sample image until the calculated second loss value is within the second preset range.
[0285] When the processor 1101 executes the program stored on the memory 1103, it can also implement the following steps:
[0286] Obtain an image to be recognized, where the image to be recognized is an image containing a text line;
[0287] Input the image to be recognized into the target font recognition model;
[0288] Perform recognition processing on the image to be recognized based on the target font recognition model to obtain a text attribute sequence corresponding to the text line in the image to be recognized;
[0289] Determine the dialogue text line and non-dialogue text line in the text line in the image to be recognized according to the text attribute sequence.
[0290] The communication bus mentioned in the above terminal may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of simplicity, only a thick line is shown in the figure, but it does not mean that there is only one bus or one type of bus.
[0291] The communication interface is used for communication between the above terminal and other devices.
[0292] The memory may include a Random Access Memory (RAM), or may also include a non-volatile memory, such as at least one disk memory. Optionally, the memory may also be at least one storage device located away from the aforementioned processor.
[0293] The aforementioned processor may be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it may also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.
[0294] In another embodiment provided by the present application, a computer-readable storage medium is also provided. Instructions are stored in the computer-readable storage medium. When it runs on a computer, it causes the computer to execute the model training method described in any one of the above embodiments, or the text line determination method in the above embodiments.
[0295] In another embodiment provided by the present application, a computer program product containing instructions is also provided. When it runs on a computer, it causes the computer to execute the model training method described in any one of the above embodiments, or the text line determination method in the above embodiments.
[0296] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that can be accessed by a computer, or a data storage device such as a server or data center that includes one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid state disk (SSD)).
[0297] It should be noted that, in this document, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise", or any other variation thereof are intended to cover non-exclusive inclusion, so that a process, method, article, or device including a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article, or device including the element.
[0298] Each embodiment in this specification is described in a related manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the description of the method embodiment.
[0299] The above are only the preferred embodiments of the present application and are not intended to limit the protection scope of the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application are all included in the protection scope of the present application.
Claims
1. A model training method, characterized in that, Including: Obtain a first sample image and a second sample image, both the first sample image and the second sample image being sample images containing text lines, the first sample image being an image simulating multiple text fonts; each character in the first sample image is labeled with a first font label, and a separator is labeled between two adjacent fonts in the same text line; Train a font recognition model to be trained based on the first sample image to obtain a pre-trained font recognition model; Train the pre-trained font recognition model based on the first sample image and the second sample image to obtain a target font recognition model; The training of the font recognition model to be trained based on the first sample image to obtain a pre-trained font recognition model includes: Input the first sample image into the font recognition model to be trained; Process the first sample image based on the font recognition model to be trained to obtain a predicted font label of the first sample image containing a separator; Calculate a first loss value of the font recognition model to be trained according to the first font label and the predicted font label; When the first loss value is within a first preset range, determine the trained font recognition model to be trained as the pre-trained font recognition model.
2. The method according to claim 1, wherein The obtaining of the first sample image and the second sample image includes: Obtain a second sample image containing text lines from a preset image library; Obtain an initial image that does not contain text lines from the preset image library; Add text lines to the initial image to generate a first sample image; each text line in the first sample image contains multiple text fonts.
3. The method according to claim 1, wherein The calculating of the first loss value of the font recognition model to be trained according to the first font label and the predicted font label includes: Determine multiple font paths corresponding to each text line in the first sample image according to the predicted font label; Determine the font probability of the font to which each character belongs according to the first font label and the predicted font label; Calculate the font path probability corresponding to the multiple font paths according to the font probability of the font to which each character belongs; Calculate the first loss value of the font recognition model to be trained according to the maximum font path probability in the font path probability.
4. The method according to claim 1, wherein After calculating the first loss value of the font recognition model to be trained according to the first font label and the predicted font label, it further includes: When the first loss value is outside the first preset range, train the trained font recognition model to be trained according to the first sample image until the calculated first loss value is within the first preset range.
5. The method according to claim 1, wherein Each character in the first sample image is labeled with a second font label, each character in the second sample image is labeled with a third font label, and a separator is labeled between two adjacent fonts in the same text line; The training of the pre-trained font recognition model based on the first sample image and the second sample image to obtain a target font recognition model includes: Input the first sample image and the second sample image into the pre-trained font recognition model; Based on the pre-trained font recognition model, process the first sample image and the second sample image to obtain a first predicted font label of the first sample image and a second predicted font label of the second sample image that contains a delimiter; According to the second font label and the first predicted font label, and the third font label and the second predicted font label, calculate a second loss value of the pre-trained font recognition model; When the second loss value is within a second preset range, determine the trained pre-trained font recognition model as the target font recognition model.
6. The method according to claim 5, wherein The calculating the second loss value of the pre-trained font recognition model according to the second font label and the first predicted font label, and the third font label and the second predicted font label includes: According to the first predicted font label, determine multiple first font paths corresponding to each text line in the first sample image, and according to the second predicted font label, determine second font paths corresponding to each text line in the second sample image; According to the second font label and the first predicted font label, determine a first font probability of the font to which each character in the first sample image belongs, and according to the third font label and the second predicted font label, determine a second font probability of the font to which each character in the second sample image belongs; According to the first font probability, calculate a first font path probability corresponding to the first font path, and according to the second font path probability, calculate a second font path probability corresponding to the second font path; According to the largest first font path probability among the first font path probabilities and the largest second font path probability among the second font path probabilities, calculate the second loss value.
7. The method according to claim 5, characterized in that After calculating the second loss value of the pre-trained font recognition model according to the second font label and the first predicted font label, and the third font label and the second predicted font label, it further includes: When the second loss value is outside the second preset range, train the trained font recognition model according to the first sample image and the second sample image until the calculated second loss value is within the second preset range.
8. A method for determining a text line, characterized in that, Includes: Obtain an image to be recognized, where the image to be recognized is an image containing text lines; Input the image to be recognized into the target font recognition model; The target font recognition model is trained by the method described in claim 1; Based on the target font recognition model, perform recognition processing on the image to be recognized to obtain a text attribute sequence corresponding to the text lines in the image to be recognized; According to the text attribute sequence, determine the dialogue text lines and non-dialogue text lines in the text lines in the image to be recognized.
9. A model training device, characterized in that, Includes: A sample image acquisition module, configured to acquire a first sample image and a second sample image. Both the first sample image and the second sample image are sample images containing text lines. The first sample image is an image simulating various text fonts. In the first sample image, each character is labeled with a first font label, and a separator is labeled between two adjacent fonts within the same text line. A pre-trained model acquisition module, configured to train a font recognition model to be trained based on the first sample image to obtain a pre-trained font recognition model. A target recognition model acquisition module, configured to train the pre-trained font recognition model based on the first sample image and the second sample image to obtain a target font recognition model. The pre-trained model acquisition module includes: A first sample image input unit, configured to input the first sample image into the font recognition model to be trained. A first predicted font label acquisition unit, configured to process the first sample image based on the font recognition model to be trained to obtain a predicted font label of the first sample image containing a separator. A first loss value calculation unit, configured to calculate a first loss value of the font recognition model to be trained according to the first font label and the predicted font label. A pre-trained model determination unit, configured to determine the trained font recognition model to be trained as the pre-trained font recognition model when the first loss value is within a first preset range.
10. A text line determination device, characterized in that, It includes: An image to be recognized acquisition module, configured to acquire an image to be recognized, where the image to be recognized is an image containing a text line. An image to be recognized input module, configured to input the image to be recognized into the target font recognition model. The target font recognition model is trained by the method described in claim 1. A text attribute sequence acquisition module, configured to perform recognition processing on the image to be recognized based on the target font recognition model to obtain a text attribute sequence corresponding to the text line in the image to be recognized. A dialogue text line determination module, configured to determine dialogue text lines and non-dialogue text lines in the text lines in the image to be recognized according to the text attribute sequence.
11. An electronic device, characterized in that, It includes: At least one processor; And A memory communicatively connected to the at least one processor. Wherein, the memory stores instructions executable by the at least one processor. The instructions are executed by the at least one processor so that the at least one processor can execute the model training method described in any one of claims 1-7, or the text line determination method described in claim 8.
12. A non-transitory computer-readable storage medium storing computer instructions, characterized in that, The computer instructions are used to cause the computer to execute the model training method described in any one of claims 1-7, or the text line determination method described in claim 8.
13. A computer program product, including a computer program, where the computer program, when executed by a processor, implements the model training method described in any one of claims 1-7, or the text line determination method described in claim 8.
Citation Information
Patent Citations
Cutting force neural network prediction model training method based on transfer learning
CN112699550A
Deep learning of electrical properties tomography
CN114080184A
Text processing method and device, electronic equipment and storage medium
CN114596522A