Model Training and Language Classification Methods, Devices, Equipment, and Storage Media
By training the language classification model for two types of sample pictures to identify the characteristics of characters and text-free areas, the problem of misdetecting text boxes in the prior art is solved, and more accurate text recognition is achieved.
Patent Information
- Application Number
- CN202210306757.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-25
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2042-03-25
AI Technical Summary
In the prior art, if there is only background information but no text content in the text box obtained by the text detection algorithm CTPN, if there is only background information but no text content, it is easily detected as containing text.
The first type of sample pictures and the second type of sample pictures are used to train the language classification model respectively. The ratio of the character area occupied by the first type of sample pictures to the picture area is greater than the threshold, which is used to train the model to recognize character features; the ratio of the character area occupied by the second type of sample pictures to the picture area is less than the threshold, which is used to train the model to recognize textless areas.
The language classification model obtained through training can effectively identify text images with only background information but no text content, solving the problem of error detection.
Smart Images

Figure CN114707588B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computers, and in particular, to a method, device, equipment, and storage medium for model training and language classification. Background Art
[0002] In the process of line production, there will be a situation where the lines of a film and television drama with mixed Chinese and English subtitles are produced. To ensure the accuracy of line production, after detecting the text box of the line, a language classification algorithm is used to identify the language of the detected text box, and according to the identified result, the text box specifically uses Chinese OCR (Optical Character Recognition) or English OCR.
[0003] In the application of the language classification algorithm, the input text box is obtained through the text detection algorithm CTPN. Using CTPN will mis-detect a text box that only has background information but no text content (such as Figure 1 as shown), so how to identify such text boxes has become a technical problem that needs to be solved urgently by those skilled in the art. Summary of the Invention
[0004] This application provides a method, device, equipment, and storage medium for model training and language classification, which is used to solve the problem of mis-detecting a text box that only has background information but no text content in the related art.
[0005] In a first aspect, a model training method is provided, including:
[0006] Obtain a first type of sample picture and a second type of sample picture; the first type of sample picture includes N first sample pictures; for any one of the first sample pictures, the ratio of the area occupied by the characters in the first sample picture to the area of the first sample picture is greater than a first threshold; the second type of sample picture includes M second sample pictures, and for any one of the second sample pictures, the ratio of the area occupied by the characters in the second sample picture to the area of the second sample picture is less than a second threshold;
[0007] Use the first type of sample picture to train the language classification model until the language classification model converges, and obtain a stage language classification model;
[0008] Use the second type of sample picture to train the stage language classification model until the stage language classification model converges, and obtain a final language classification model.
[0009] Optionally, using the first type of sample picture to train the language classification model to obtain a stage language classification model includes:
[0010] For any of the first sample pictures in the first type of sample pictures, perform the following processing:
[0011] Process the first sample picture through the language classification model to obtain an S*T scoring matrix, where S is the number of sub-regions divided from the first sample picture, T is the number of pre-set languages, and each element in the scoring matrix indicates the score of a sub-region belonging to one of the T languages;
[0012] Generate an optimal path based on the scoring matrix; the optimal path includes S languages; for any one of the S languages, the score in the row corresponding to the language in the scoring matrix is the highest;
[0013] Calculate a loss function based on the optimal path and the annotation information of the first sample picture;
[0014] Optimize the parameters of the language classification model using the loss function.
[0015] Optionally, train the stage language classification model using the second type of sample pictures to obtain a final language classification model, including:
[0016] For any of the second sample pictures in the second type of sample pictures, perform the following processing:
[0017] Divide the region of the second sample picture to obtain P sub-regions;
[0018] Use the stage language classification model to identify the target sub-regions with character features and non-target sub-regions without character features in the P sub-regions;
[0019] Based on the target sub-regions and the character features, predict the scores of the language categories corresponding to the target sub-regions; and predict the scores of the language categories corresponding to the non-target sub-regions;
[0020] Obtain an optimal path based on the scores of the language categories of the target sub-regions and the scores of the language categories of the non-target sub-regions;
[0021] Calculate a loss function based on the optimal path and the annotation information of the second sample picture;
[0022] Optimize the parameters of the stage language classification model using the loss function.
[0023] Optionally, the first type of samples include two types of pictures: pictures intercepted from the captions of audio-visual videos and pictures made based on script stickers;
[0024] The second type of samples includes two types of pictures, namely pictures intercepted from the subtitle text of audio-visual videos and pictures made based on script stickers.
[0025] Optionally, the language classification model at least includes:
[0026] A convolutional neural network for extracting the hidden layer feature vectors of the sample pictures, and a recurrent neural network for extracting the semantic information of the sample pictures based on the hidden layer feature vectors, where the semantic information is used to generate the scoring matrix.
[0027] In a second aspect, a language classification method is provided, including:
[0028] Obtain a text picture to be recognized, where there are no characters in the text picture to be recognized;
[0029] Use the final language classification model trained by the model training method described in the first aspect to recognize the text picture to be recognized, and obtain a scoring matrix, where the language category with the highest score in each row of data in the scoring matrix is empty;
[0030] Obtain an optimal path based on the scoring matrix;
[0031] Normalize the optimal path to obtain a recognition result indicating that the language category of the text picture to be recognized is empty.
[0032] In a third aspect, a model training device is provided, including:
[0033] A first acquisition unit for acquiring the first type of sample pictures and the second type of sample pictures; the first type of sample pictures includes N first sample pictures; for any one of the first sample pictures, the area ratio of the area occupied by the characters in the first sample picture to the area of the first sample picture is greater than a first threshold; the second type of sample pictures includes M second sample pictures, and for any one of the second sample pictures, the area ratio of the area occupied by the characters in the second sample picture to the area of the second sample picture is less than a second threshold;
[0034] A first training unit for training the language classification model using the first type of sample pictures until the language classification model converges to obtain a stage language classification model;
[0035] A second training unit for training the stage language classification model using the second type of sample pictures until the stage language classification model converges to obtain the final language classification model.
[0036] In a fourth aspect, a language classification device is provided, including:
[0037] A second acquisition unit for acquiring a text picture to be recognized, where there are no characters in the text picture to be recognized;
[0038] An obtaining unit, configured to use the final language classification model trained by the model training method described in the first aspect to identify the text image to be recognized, and obtain a scoring matrix, where the language category with the highest score in each row of data in the scoring matrix is empty;
[0039] A second obtaining unit, configured to obtain an optimal path based on the scoring matrix;
[0040] A normalization unit, configured to normalize the optimal path to obtain an identification result indicating that the language category of the text image to be recognized is empty.
[0041] In a fifth aspect, there is provided an electronic device, including: a processor, a memory, and a communication bus, where the processor and the memory complete communication with each other through the communication bus;
[0042] The memory is configured to store a computer program;
[0043] The processor is configured to execute the program stored in the memory to implement the model training method in the first aspect or the language classification method described in the second aspect.
[0044] In a sixth aspect, there is provided a computer-readable storage medium storing a computer program, where when the computer program is executed by a processor, the model training method in the first aspect or the language classification method described in the second aspect is implemented.
[0045] The above technical solutions provided in the embodiments of the present application have the following advantages compared with the prior art: In the method provided in the embodiments of the present application, since the language classification model is trained using the first type of sample images and the second type of sample images respectively, the obtained language classification model can identify text images that only have background information but no text content, thereby solving the problem of false detection of text images that only have background information but no text content in the related art. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] The accompanying drawings herein are incorporated into the specification and constitute a part of the specification, showing embodiments consistent with the present invention, and are used together with the specification to explain the principles of the present invention.
[0047] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0048] Figure 1 It is a schematic diagram of a text box that only has background information but no text content shown in the related art;
[0049] Figure 2 It is a schematic flowchart of the model training method shown in the embodiments of the present application;
[0050] Figure 3 It is a schematic diagram of a first sample picture shown in the embodiments of the present application;
[0051] Figure 4 It is a schematic diagram of a second sample picture shown in the embodiments of the present application;
[0052] Figure 5 It is another schematic diagram of the first sample picture shown in the embodiments of the present application;
[0053] Figure 6 corresponding to the embodiments of the present application Figure 5 It is a schematic diagram of the path of the first sample picture shown;
[0054] Figure 7 It is another schematic diagram of the second sample picture shown in the embodiments of the present application;
[0055] Figure 8 corresponding to the embodiments of the present application Figure 7 It is a schematic diagram of the path of the second sample picture shown;
[0056] Figure 9 It is a schematic diagram of the language classification model finding the optimal path shown in the embodiments of the present application;
[0057] Figure 10 It is a schematic flowchart of the language classification method shown in the embodiments of the present application;
[0058] Figure 11 It is a schematic structural diagram of the model training device shown in the embodiments of the present application;
[0059] Figure 12 It is a schematic structural diagram of the language classification device shown in the embodiments of the present application;
[0060] Figure 13 It is a schematic structural diagram of the electronic device shown in the embodiments of the present application. Detailed implementation manners
[0061] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are some, rather than all, of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the scope of protection of the present application.
[0062] It should be noted that the terms "first", "second", etc. in the description, claims and the above-mentioned drawings of the present application are used to distinguish similar objects, and do not necessarily have to be used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present application described here can be implemented in an order other than those illustrated or described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0063] An embodiment of the present application provides a model training method, which can be applied to any electronic device;
[0064] The electronic device described in the embodiments of the present application may include a terminal or a server, and the embodiments of the present application do not make specific limitations. Among them, the terminal includes various handheld devices, vehicle-mounted devices, wearable devices (such as smart watches, smart bracelets, pedometers, etc.), and computing devices with wireless communication functions.
[0065] As Figure 2 shown, the method may include the following steps:
[0066] Step 201, obtain a first type of sample picture and a second type of sample picture.
[0067] The first type of sample pictures includes N first sample pictures; the area ratio of the area occupied by the characters in any first sample picture to the area of the first sample picture is greater than a first threshold; the second type of sample pictures includes M second sample pictures, and the area ratio of the area occupied by the characters in any second sample picture to the area of the second sample picture is less than a second threshold.
[0068] Here, the first sample picture refers to a picture in which the background area is evenly filled with characters. It should be understood that the so-called "filled" does not mean that there is no area in the picture that is not filled with characters, but means that the area ratio of the area occupied by the characters to the area of the picture is within a range, usually greater than the first threshold. In applications, the first threshold can be set manually according to experience or based on actual needs. The so-called "evenly" means that the characters are not filled in a pile when filling the first sample picture. Among them, filling in a pile means that all the characters are concentrated in a certain area of the first sample picture, and the other areas of the first sample picture only have background information but are not filled with characters.
[0069] In this embodiment, both the area occupied by the characters and the area of the sample image can be indicated by area, that is, the area occupied by the characters in the sample image indicates the area occupied by the characters, and the area of the sample image indicates the area of the sample image.
[0070] Please refer to Figure 3 , Figure 3 which is a schematic diagram of the first sample image shown in this embodiment. For Figure 3 any one of the characters in, the height of the character is basically the same as the width of the image, and the sum of the lengths of all the characters is basically the same as the length of the image. It should be noted that the width of the image refers to the shorter side of the image, and the length of the image refers to the longer side of the image. Correspondingly, the height of the character refers to the length of the character in the direction same as the width of the image, and the length of the character refers to the length of the character in the direction same as the length of the image.
[0071] In this embodiment, the second sample image refers to a situation where there are some areas in the background area that are completely not filled with characters. It should be understood that the width of these partial areas is the same as the width of the second sample image, and the length of the distinguishing area is less than the length of the second sample image.
[0072] Please refer to Figure 4 , Figure 4 which is a schematic diagram of the second sample image shown in this embodiment. In Figure 4 , the area marked by the dashed box is the area in the second sample image that is completely not filled with characters.
[0073] Step 202: Train the language classification model using the first type of sample images until the language classification model converges to obtain a stage language classification model.
[0074] In this embodiment, since each first sample image in the first type of sample images is uniformly filled with characters, when training the language classification model using the first type of sample images, the language classification model can quickly learn the features of the characters and the delimiters. The delimiter refers to the area between two adjacent characters.
[0075] In this embodiment, T language categories are preset in advance. When training the language classification model using the first sample images in the first type of sample images, a loss function is generated based on the category scores of the characters in the first sample image, and the parameters of the language classification model are optimized using the loss function until the language classification model converges to obtain a stage language classification model.
[0076] In an alternative embodiment, for any one of the first sample images in the first type of sample images, the following processing is performed:
[0077] The first sample image is processed by a language classification model to obtain an S*T scoring matrix, where S is the number of sub-regions divided by the first sample image, T is the number of pre-set languages, and each element in the scoring matrix indicates that the language category of a sub-region is a score of one of the T languages; based on the scoring matrix, an optimal path is generated; the optimal path includes S languages; any one of the S languages has the highest score in the row where the same language is located in the scoring matrix; based on the optimal path and the annotation information of the first sample image, a loss function is calculated; and the loss function is used to optimize the parameters of the language classification model.
[0078] It should be understood that for different first sample images, the number of sub-areas divided by the first sample image is different. Usually, the calculation formula of the sub-area is as follows:
[0079]
[0080] Wherein S is the number of sub-regions divided by the first sample image; L is the length of the first sample image; δ is a preset division coefficient, which can be set manually based on experience, such as setting δ to 16.
[0081] It should be understood that the characters in the first sample image may correspond to one or more sub-areas depending on the size of the characters. When a character corresponds to multiple sub-areas, the language categories of the multiple sub-areas are the same. For example, when the character is the Chinese character "国", according to the above-mentioned sub-area division method, the character corresponds to 8 sub-areas, and the language category of the 8 sub-areas is "C", where "C" represents Chinese.
[0082] In this embodiment, the languages that can be set include but are not limited to Chinese, English, Korean, Japanese, Spanish, French, and German.
[0083] For example, if T is 3 and S is 16, that is, the language categories are Chinese, English, and empty, then the scoring matrix is a 3*16 matrix, and the first element in the scoring matrix (i.e., the first row and the first column) represents the score of the first area in the sample image being Chinese. The language category corresponding to the element with the highest score in each row of the scoring matrix is the language category of the sub-area corresponding to the row.
[0084] Please refer to Figure 5 as well as Figure 6 , Figure 5 Another example of the first sample picture is: Figure 6 For the corresponding Figure 5 The path diagram of the first sample picture shown is shown in FIG. Figure 5 The first sample image shown is divided into 8 sub-regions, that is, there are 8 locations that need to predict the language category. All the paths for generating language categories are as follows Figure 6Paths of two colors are shown. The path indicated by the thicker line is the obtained optimal path. The output result is "∈C∈CLC∈L", where "C" represents Chinese, "L" represents English, and "∈" represents the separator.
[0085] In this embodiment, the language classification model can specifically be composed of a CNN network (Convolutional Neural Network) and an RNN network (Recurrent Neural Network). Among them, the CNN network is used to extract the hidden layer feature vectors of the first sample pictures, and the RNN network is used to perform temporal feature extraction on the semantic information in the hidden layer feature vectors. When training with real samples, the semantic information in the text line can be extracted. When testing with real samples, the actual semantic information of the test samples can be effectively obtained, increasing the accuracy of language prediction. It should be understood that the speech information here is used to generate the scoring matrix.
[0086] In the application, taking the convolutional neural network having N network levels as an example, the process of obtaining the hidden layer feature vectors of the first sample pictures based on the CNN network is as follows:
[0087] Use the 1st network level to perform convolution and pooling on the sample image to obtain the features of the 1st network level of the sample image;
[0088] Use the i-th network level to perform convolution and pooling on the features of the (i - 1)-th network level of the sample image to obtain the features of the i-th network level of the sample image, where the value of i is greater than 1 and less than N;
[0089] Use the N-th network level to perform convolution on the features of the (N - 1)-th network level of the sample image to obtain the features of the N-th network level of the sample image;
[0090] Perform L2 (two-norm) normalization operation on the features of the N-th network level, and use the normalized features as the hidden layer feature vectors of the sample image.
[0091] In this embodiment, in order to be as close as possible to the actual application and improve the recognition accuracy of the language classification model, the first type of samples used include two types of samples: real samples and highly simulated samples. Among them, the real samples can be pictures intercepted from the subtitles of audio-visual videos, and the highly simulated samples can be pictures made based on script textures.
[0092] Step 203: Use the second type of sample pictures to train the stage language classification model until the stage language classification model converges to obtain the final language classification model.
[0093] In this embodiment, since the stage language classification model can recognize the features of characters and delimiters, when the second type of samples are used to train the stage language classification model, the stage language classification model can quickly recognize the characters and delimiters in the second sample picture, and quickly locate the sub-regions of the characters and delimiters based on the recognition of the characters and delimiters. Further, based on the positioning of the sub-regions of the characters and delimiters, the positions and features of the regions in the second sample picture that only have background information and are not filled with characters are determined, and finally, the loss function is calculated based on the recognition results.
[0094] In an alternative embodiment, for any second sample picture in the second type of sample pictures, the following processing is performed: the region of the second sample picture is divided to obtain P sub-regions; the stage language classification model is used to recognize the target sub-regions with character features and the non-target sub-regions without character features in the P sub-regions; based on the target sub-regions and character features, the language category corresponding to the target sub-regions is predicted; and the language category corresponding to the non-target sub-regions is predicted; based on the language category of the target sub-regions and the language category of the non-target sub-regions, the optimal path is obtained; based on the optimal path and the annotation information of the second sample picture, the loss function is calculated; and the parameters of the stage language classification model are optimized using the loss function.
[0095] It should be understood that when predicting the scores of the corresponding language categories of the target sub-regions based on the target sub-regions and character features, a score is predicted for each preset language category, and the scores corresponding to different language categories are usually different, and the score of the actual language category closest to the target sub-region is the highest. Similarly, there are also multiple scores for the language categories of the non-target sub-regions, that is, for each preset language category, the non-target sub-region has a predicted score, and the score of the actual language category closest to the non-target sub-region is the highest.
[0096] It should be understood that when generating the optimal path based on the scores of the language categories of the target sub-regions and the scores of the language categories of the non-target sub-regions, the highest scores of the language categories of the target sub-regions and the highest scores of the language categories of the non-target sub-regions are obtained; the first language category corresponding to the highest score of the language category of the target sub-regions and the second language category corresponding to the highest score of the language category of the non-target sub-regions are obtained; and the optimal path is generated based on the first language category and the second language category.
[0097] Please refer to Figure 7 and Figure 8 , Figure 7 which is another example of the second sample picture shown in the embodiments of the present application, Figure 8 is the path schematic diagram corresponding to Figure 7 the second sample picture shown. Among them, it is assumed that Figure 7The first sample picture shown is divided into 10 sub-regions in total, that is, there are 10 positions where the language category needs to be predicted. All the paths generating the language categories are as Figure 8 the paths of the three colors shown. The gray path is the optimal path obtained. The output result is "∈∈C∈CLC∈L∈". Among them, "C" represents Chinese, "L" represents English, and "∈" represents the separator. Due to the training basis in the first stage, when using dynamic programming to calculate the loss and obtain the optimal path, it is easier to determine the optimal path. At this time, the obtained optimal path is more consistent with the actual sample.
[0098] In this embodiment, in order to be as close as possible to the actual application and improve the recognition accuracy of the language classification model, the second type of samples adopted include two types of samples: real samples and highly simulated samples. Among them, the real samples can be pictures intercepted from the captions of audio-visual videos, and the highly simulated samples can be pictures made based on script textures.
[0099] In the technical solution provided in this embodiment, since the first type of sample pictures and the second type of sample pictures are used to train the language classification model respectively, the obtained language classification model can identify text pictures with only background information but no text content, thus solving the problem of misdetection of text pictures with only background information but no text content in the related art.
[0100] For the convenience of comparison, a comparison is given in which the first type of samples and the second type of samples are mixed to train the language classification model without dividing into stages. The following still takes the language classification model to Figure 7 train the sample pictures shown as an example. Please refer to Figure 9 , Figure 9 as a schematic diagram for finding the optimal path of the language classification model. Among them, it is still assumed that Figure 7 the first sample picture shown is divided into 10 sub-regions in total, that is, there are 10 positions where the language category needs to be predicted. All the paths generating the language categories are as Figure 9 the paths of the three colors shown. Just because there are two more positions for predicting the language category, the number of paths increases significantly. It can be seen that if the number of the second type of samples in the sample is very large, and each position needs to predict and output a label indicating the language category, the number of labels increases linearly, and the amount of path data will increase exponentially, resulting in a significant increase in the amount of calculation.
[0101] The reason for the occurrence of Figure 9The numerous paths shown are because in the actual training process, since the initial model has not converged, assuming that the number of samples of the second type in the training samples is very large, when calculating the optimal path, there is a high probability that the calculated path is incorrect, resulting in the failure to match the language type result of the character with the actual position of the character, that is, the wrong text or background is misrecognized as a certain language type.
[0102] An embodiment of the present application provides a language classification method, which can be applied to any electronic device in which a final language classification model is deployed, and the final language classification model is verified using a text image to be recognized. As Figure 10 shown, the method may include the following steps:
[0103] Step 1001, obtain a text image to be recognized, and there are no characters in the text image to be recognized;
[0104] Step 1002, use the final language classification model to recognize the text image to be recognized, and obtain a scoring matrix, where the language category with the highest score in each row of data in the scoring matrix is empty;
[0105] Step 1003, obtain an optimal path based on the scoring matrix;
[0106] Step 1004, normalize the optimal path to obtain a recognition result indicating that the language category of the text image to be recognized is empty.
[0107] Based on the same concept, an embodiment of the present application provides a model training device. For the specific implementation of the device, reference can be made to the description in the method embodiment part, and the repeated parts will not be elaborated. As Figure 11 shown, the device mainly includes:
[0108] A first acquisition unit 1101, configured to acquire a first type of sample image and a second type of sample image; the first type of sample image includes N first sample images; the area ratio of the area occupied by the characters in any first sample image to the area of the first sample image is greater than a first threshold; the second type of sample image includes M second sample images, and the area ratio of the area occupied by the characters in any second sample image to the area of the second sample image is less than a second threshold;
[0109] A first training unit 1102, configured to train the language classification model using the first type of sample image until the language classification model converges, and obtain a stage language classification model;
[0110] A second training unit 1103, configured to train the stage language classification model using the second type of sample image until the stage language classification model converges, and obtain a final language classification model.
[0111] The first training unit 1102 is used for:
[0112] For any first sample picture in the first type of sample pictures, perform the following processing:
[0113] Process the first sample picture through a language classification model to obtain an S*T scoring matrix, where S is the number of sub-regions divided from the first sample picture, T is the number of preset languages, and each element in the scoring matrix indicates the score of a sub-region for one of the T languages;
[0114] Based on the scoring matrix, generate an optimal path; the optimal path includes S languages; for any one of the S languages, the score in the row corresponding to any language in the scoring matrix is the highest;
[0115] Based on the optimal path and the annotation information of the first sample picture, calculate the loss function;
[0116] Use the loss function to optimize the parameters of the language classification model.
[0117] The second training unit 1103 is used for:
[0118] For any second sample picture in the second type of sample pictures, perform the following processing:
[0119] Divide the region of the second sample picture to obtain P sub-regions;
[0120] Use a stage language classification model to identify target sub-regions with character features and non-target sub-regions without character features in the P sub-regions;
[0121] Based on the target sub-regions and character features, predict the scores of the language categories corresponding to the target sub-regions; and predict the scores of the language categories corresponding to the non-target sub-regions;
[0122] Based on the scores of the language categories of the target sub-regions and the scores of the language categories of the non-target sub-regions, obtain an optimal path;
[0123] Based on the optimal path and the annotation information of the second sample picture, calculate the loss function;
[0124] Use the loss function to optimize the parameters of the stage language classification model.
[0125] The first type of samples includes two types of pictures: pictures intercepted from the subtitles of audio-visual videos and pictures made based on script stickers.
[0126] The second type of samples includes two types of pictures: pictures intercepted from the subtitle text of audio-visual videos and pictures made based on script stickers.
[0127] The language classification model at least includes:
[0128] A convolutional neural network for extracting the hidden layer feature vectors of sample pictures, and a recurrent neural network for extracting the semantic information of sample pictures based on the hidden layer feature vectors, where the semantic information is used to generate a scoring matrix.
[0129] Based on the same concept, an embodiment of the present application provides a language classification device. For the specific implementation of this device, reference may be made to the description in the method embodiment part, and repeated parts will not be elaborated. As Figure 12 shown, this device mainly includes:
[0130] A second acquisition unit 1201, for a text picture to be recognized, where there are no characters in the text picture to be recognized;
[0131] A first acquisition unit 1202, configured to use the final language classification model obtained by the model training method to recognize the text picture to be recognized, and obtain a scoring matrix, where the language category with the highest score in each row of data in the scoring matrix is empty;
[0132] A second acquisition unit 1203, configured to obtain an optimal path based on the scoring matrix;
[0133] A normalization unit 1204, configured to normalize the optimal path to obtain an identification result indicating that the language category of the text picture to be recognized is empty.
[0134] Based on the same concept, an embodiment of the present application also provides an electronic device. As Figure 13 shown, this electronic device mainly includes: a processor 1301, a memory 1302, and a communication bus 1303. Among them, the processor 1301 and the memory 1302 communicate with each other through the communication bus 1303. Among them, a program executable by the processor 1301 is stored in the memory 1302, and the processor 1301 executes the program stored in the memory 1302 to implement the following steps:
[0135] Obtain a first type of sample pictures and a second type of sample pictures; the first type of sample pictures includes N first sample pictures; for any first sample picture, the area ratio of the area occupied by characters in the first sample picture to the area of the first sample picture is greater than a first threshold; the second type of sample pictures includes M second sample pictures, and for any second sample picture, the area ratio of the area occupied by characters in the second sample picture to the area of the second sample picture is less than a second threshold; use the first type of sample pictures to train the language classification model until the language classification model converges to obtain a stage language classification model; use the second type of sample pictures to train the stage language classification model until the stage language classification model converges to obtain a final language classification model;
[0136] Alternatively, obtain an image of the text to be recognized, where the image of the text to be recognized does not have characters; use the final language classification model trained by the model training method to recognize the image of the text to be recognized, obtain a scoring matrix, and the language category with the highest score in each row of data in the scoring matrix is empty; obtain the optimal path based on the scoring matrix; normalize the optimal path to obtain a recognition result indicating that the language category of the image of the text to be recognized is empty.
[0137] The communication bus 1303 mentioned in the above electronic device can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. The communication bus 1303 can be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience of representation, Figure 13 only a thick line is used to represent it in the figure, but it does not mean that there is only one bus or one type of bus.
[0138] The memory 1302 can include a Random Access Memory (RAM), and can also include a non-volatile memory, such as at least one disk memory. Optionally, the memory can also be at least one storage device located far from the aforementioned processor 1301.
[0139] The aforementioned processor 1301 can be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc., and can also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.
[0140] In another embodiment of the present application, a computer-readable storage medium is further provided. A computer program is stored in the computer-readable storage medium. When the computer program runs on a computer, the computer is caused to execute the model training method or the language classification method described in the above embodiment.
[0141] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions are transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wire (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, microwave, etc.). The computer-readable storage medium can be any available medium that the computer can access, or a data storage device such as a server or data center that includes one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape, etc.), an optical medium (such as a DVD), or a semiconductor medium (such as a solid-state drive), etc.
[0142] It should be noted that in this document, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover a non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not expressly listed, or elements inherent to such process, method, article, or device. Without further limitation, an element defined by the statement "comprising a..." does not exclude the presence of additional identical elements in the process, method, article, or device comprising the element.
[0143] The above are only specific embodiments of the present invention, enabling those skilled in the art to understand or implement the present invention. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to these embodiments shown herein, but rather to the broadest scope consistent with the principles and novel features claimed herein.
Claims
1. A model training method, characterized in that, Including: Obtaining first - type sample pictures and second - type sample pictures; The first - type sample pictures include N first - sample pictures; For any one of the first - sample pictures, the area ratio of the area occupied by the characters in the first - sample picture to the area of the first - sample picture is greater than a first threshold; The second - type sample pictures include M second - sample pictures. For any one of the second - sample pictures, the area ratio of the area occupied by the characters in the second - sample picture to the area of the second - sample picture is less than a second threshold; Training the language classification model with the first - type sample pictures until the language classification model converges to obtain a stage language classification model; Training the stage language classification model with the second - type sample pictures until the stage language classification model converges to obtain a final language classification model.
2. The method according to claim 1, wherein Training the language classification model with the first - type sample pictures to obtain a stage language classification model, including: For any one of the first - sample pictures in the first - type sample pictures, perform the following processing: Processing the first - sample picture through the language classification model to obtain an S*T scoring matrix, where S is the number of sub - regions divided from the first - sample picture, T is the number of pre - set languages, and each element in the scoring matrix indicates the score that the language category of a sub - region is one of the T languages; Generating an optimal path based on the scoring matrix; the optimal path includes S languages; for any one of the S languages, the score in the row corresponding to the any one of the languages in the scoring matrix is the highest; Calculating a loss function based on the optimal path and the annotation information of the first - sample picture; Optimizing the parameters of the language classification model using the loss function.
3. The method according to claim 1, characterized in that, Training the stage language classification model with the second - type sample pictures to obtain a final language classification model, including: For any one of the second - sample pictures in the second - type sample pictures, perform the following processing: Dividing the area of the second - sample picture to obtain P sub - regions; Identifying target sub - regions with character features and non - target sub - regions without character features in the P sub - regions using the stage language classification model; Predicting the scores of the language categories corresponding to the target sub - regions based on the target sub - regions and the character features; and predicting the scores of the language categories corresponding to the non - target sub - regions; Obtaining an optimal path based on the scores of the language categories of the target sub - regions and the scores of the language categories of the non - target sub - regions; Calculating a loss function based on the optimal path and the annotation information of the second - sample picture; Optimizing the parameters of the stage language classification model using the loss function.
4. The method according to claim 1, wherein The first - type samples include two types of pictures: pictures intercepted from the captions of audio - visual videos and pictures made based on script stickers; The second - type samples include two types of pictures: pictures intercepted from the caption texts of audio - visual videos and pictures made based on script stickers.
5. The method according to claim 2, characterized in that The language classification model at least includes: A convolutional neural network for extracting the hidden layer feature vectors of sample images, and a recurrent neural network for extracting the semantic information of the sample images based on the hidden layer feature vectors, where the semantic information is used to generate the scoring matrix.
6. A language classification method, characterized in that, It includes: Obtain a text image to be recognized, where there are no characters in the text image to be recognized; Use the final language classification model trained by the model training method according to any one of claims 1-5 to recognize the text image to be recognized, and obtain a scoring matrix, where the language category with the highest score in each row of data in the scoring matrix is empty; Obtain an optimal path based on the scoring matrix; Normalize the optimal path to obtain a recognition result indicating that the language category of the text image to be recognized is empty.
7. A model training device, characterized in that, It includes: A first acquisition unit for acquiring a first type of sample image and a second type of sample image; The first type of sample image includes N first sample images; For any one of the first sample images, the area ratio of the area occupied by the characters in the first sample image to the area of the first sample image is greater than a first threshold; The second type of sample image includes M second sample images, and for any one of the second sample images, the area ratio of the area occupied by the characters in the second sample image to the area of the second sample image is less than a second threshold; A first training unit for training the language classification model with the first type of sample image until the language classification model converges to obtain a stage language classification model; A second training unit for training the stage language classification model with the second type of sample image until the stage language classification model converges to obtain a final language classification model.
8. A language classification device, characterized in that, It includes: A second acquisition unit, a text image to be recognized, where there are no characters in the text image to be recognized; A first obtaining unit for using the final language classification model trained by the model training method according to any one of claims 1-5 to recognize the text image to be recognized, and obtaining a scoring matrix, where the language category with the highest score in each row of data in the scoring matrix is empty; A second obtaining unit for obtaining an optimal path based on the scoring matrix; A normalization unit for normalizing the optimal path to obtain a recognition result indicating that the language category of the text image to be recognized is empty.
9. An electronic device, characterized in that, It includes: A processor, a memory, and a communication bus. Among them, the processor and the memory complete mutual communication through the communication bus; The memory is used to store a computer program; The processor is used to execute the program stored in the memory to implement the model training method according to any one of claims 1-5 or the language classification method according to claim 6.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the model training method according to any one of claims 1-5 or the language classification method according to claim 6.
Citation Information
Patent Citations
Document type picture recognition method and device and storage medium
CN112131957A
Optical character recognition method and apparatus, electronic device and storage medium
US20210390296A1