Text recognition method, device, electronic device and storage medium
By sorting text content in language before text recognition and selecting matching text recognition algorithms, the problems of omissions and errors in multilingual text recognition are solved, and the recognition accuracy is improved.
Patent Information
- Application Number
- CN202111425615.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-26
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2041-11-26
AI Technical Summary
Existing text recognition algorithms can easily lead to text omissions or recognition errors in some languages when processing multilingual text, resulting in inaccurate recognition results.
By extracting the text image feature vectors in the video frame image, classifying the text content in the text box using preset text arrangement rules, and selecting a text recognition algorithm that matches the language for recognition.
Improve the accuracy of text recognition, especially when dealing with mixed subtitles in Chinese and English, and improve the accuracy of line recognition.
Smart Images

Figure CN114140782B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and in particular to a method and apparatus for text recognition, an electronic device, and a computer-readable storage medium. Background Art
[0002] Currently, recognizing text in an image needs to be achieved through two steps. First, detect the text position in the image through a detector, and then recognize the specific text through a text recognition algorithm.
[0003] When using a text recognition algorithm to recognize text including different languages, since the text recognition algorithm is usually only used to recognize text in a single language, there will be problems of omission or incorrect recognition of some languages in the text of some languages, resulting in inaccurate recognition results. Summary of the Invention
[0004] The purpose of the embodiments of the present invention is to provide a method and apparatus for text recognition, an electronic device, and a computer-readable storage medium, which solve the problem of inaccurate text recognition. The specific technical solutions are as follows:
[0005] In the first aspect of the embodiments of the present invention, first, a method for text recognition is provided, including: obtaining a video frame image to be processed; extracting a text image from the video frame image and recognizing a feature vector of the text image, where the text image includes at least one text box; classifying the text content in each text box according to the feature vector and a preset text arrangement rule to obtain corresponding language information, where the text arrangement rule represents the character information occupied by the text content corresponding to each language information when displayed; for the text content in each text box, respectively select a text recognition algorithm corresponding to each language information, and recognize the text content in the corresponding text box according to the selected text recognition algorithm, where the language information corresponding to the text content is the same as the language information corresponding to the selected text recognition algorithm.
[0006] Optionally, the classifying the text content in each text box according to the feature vector and a preset text arrangement rule to obtain corresponding language information includes: classifying each character included in the text content in each text box according to the feature vector and the text arrangement rule to obtain a classification result; mapping the classification result corresponding to each text box to the language information corresponding to each text box respectively.
[0007] Optionally, classifying each character included in the text content within each of the text boxes according to the feature vector and the text arrangement rule to obtain a classification result includes: inputting the text image into a convolutional neural network to obtain an image feature vector of each pixel point of the text image, and inputting the image feature vector into a recurrent neural network to obtain a text feature vector of the pixel point; counting the number of associated pixel points or independent pixel points occupied by each character according to the text feature vector, where the associated pixel points are a group of adjacent pixel points and the group of adjacent pixel points has an associated text feature vector, and the independent pixel points are pixel points whose adjacent pixel points do not have an associated text feature vector; classifying each character according to the number and the text arrangement rule to obtain the classification result.
[0008] Optionally, the text arrangement rule includes the correspondence between the quantity range where the quantity is located and the classification result; classifying each character according to the quantity and the text arrangement rule to obtain the classification result includes: for each character, using the classification result corresponding to the quantity range where the quantity is located as the classification result of each character.
[0009] Optionally, mapping the classification result corresponding to each text box to the language information corresponding to each text box respectively includes: for each text box, if the classification results of the characters within the same text box are the same, using the classification results of the characters as the language information corresponding to the same text box; if the classification results of the characters within the same text box are different, using the classification results of the characters together as the language information corresponding to the same text box.
[0010] Optionally, the language information includes Chinese language and English language; respectively selecting text recognition algorithms corresponding to each language information includes: when the language information corresponding to the text box only includes the Chinese language, selecting a Chinese recognition algorithm corresponding to the Chinese language; when the language information corresponding to the text box only includes the English language, selecting an English recognition algorithm corresponding to the English language; when the language information corresponding to the text box includes the Chinese language and the English language, selecting the Chinese recognition algorithm corresponding to the Chinese language.
[0011] In the second aspect of the implementation of the present invention, there is also provided a text recognition device, including: an acquisition module for acquiring a video frame image to be processed; a processing module for extracting a text image from the video frame image and identifying a feature vector of the text image, where the text image includes at least one text box; a classification module for classifying the text content in each text box according to the feature vector and a preset text arrangement rule to obtain corresponding language information, and the text arrangement rule represents the character information occupied by the text content corresponding to each language information when displayed; an identification module for respectively selecting a text recognition algorithm corresponding to each language information for the text content in each text box, and identifying the text content in the corresponding text box according to the selected text recognition algorithm, and the language information corresponding to the text content is the same as the language information corresponding to the selected text recognition algorithm.
[0012] Optionally, the classification module includes: a character classification module for classifying each character included in the text content in each text box according to the feature vector and the text arrangement rule to obtain a classification result; a language mapping module for mapping the classification result corresponding to each text box to the language information corresponding to each text box respectively.
[0013] Optionally, the character classification module includes: a feature extraction module for inputting the text image into a convolutional neural network to obtain an image feature vector of each pixel point of the text image, and inputting the image feature vector into a recurrent neural network to obtain a text feature vector of the pixel point; a quantity statistics module for statistically calculating the number of associated pixel points or independent pixel points occupied by each character according to the text feature vector, where the associated pixel points are a group of adjacent pixel points, and the group of adjacent pixel points has an associated text feature vector, and the independent pixel points are pixel points whose adjacent pixel points do not have an associated text feature vector; a result classification module for classifying each character according to the quantity and the text arrangement rule to obtain the classification result.
[0014] Optionally, the text arrangement rule includes the corresponding relationship between the quantity range where the quantity is located and the classification result; the result classification module is used for, for each character, taking the classification result having the corresponding relationship with the quantity range where the quantity is located as the classification result of each character.
[0015] Optionally, the language mapping module is configured to, for each of the text boxes, if the classification results of the characters in the same text box are the same, use the classification results of the characters as the language information corresponding to the same text box; if the classification results of the characters in the same text box are different, use the classification results of the characters together as the language information corresponding to the same text box.
[0016] Optionally, the language information includes Chinese and English; the recognition module is configured to, when the language information corresponding to the text box only includes Chinese, select a Chinese recognition algorithm corresponding to Chinese; when the language information corresponding to the text box only includes English, select an English recognition algorithm corresponding to English; when the language information corresponding to the text box includes Chinese and English, select the Chinese recognition algorithm corresponding to Chinese.
[0017] In another aspect of the implementation of the present invention, there is also provided a computer-readable storage medium, in which instructions are stored, and when it runs on a computer, it causes the computer to execute the text recognition method described in any one of the above.
[0018] In another aspect of the implementation of the present invention, there is also provided a computer program product containing instructions, and when it runs on a computer, it causes the computer to execute the text recognition method described in any one of the above.
[0019] The text recognition solution provided by the embodiments of the present invention adopts the technical means of obtaining a video frame image to be processed, extracting a text image from the video frame image, and recognizing a feature vector of the text image, where the text image includes at least one text box. According to a preset text arrangement rule and the feature vector, the text content in each text box is classified to obtain corresponding language information. The arrangement rule represents the character information occupied by the text content corresponding to each language information when displayed. Furthermore, for the text content in each text box, a text recognition algorithm corresponding to each language information is selected, and the text content in the corresponding text box is recognized according to the selected text recognition algorithm. Before recognizing the text content, the text content is classified to obtain corresponding language information, and then a text recognition algorithm corresponding to the language information is selected for the text content, and the text content is recognized by using the selected text recognition algorithm. In the embodiments of the present invention, by selecting a text recognition algorithm that matches the language in the text content to recognize the text content, it can, to a certain extent, solve the technical problems of inaccurate recognition such as omission of text recognition and text recognition errors of some languages when using one text recognition algorithm to recognize text content containing multiple languages, and achieve the effect of improving the text recognition accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art.
[0021] Figure 1 It is a flowchart of the steps of a text recognition method according to an embodiment of the present invention.
[0022] Figure 2 It is a flowchart of the steps of classifying each character included in the text content in each text box to obtain a classification result according to an embodiment of the present invention.
[0023] Figure 3 It is a schematic diagram of the training process of a text classification algorithm according to an embodiment of the present invention.
[0024] Figure 4 It is a schematic diagram of the classification process of a text line according to an embodiment of the present invention.
[0025] Figure 5 It is a schematic diagram of the correct text recognition result according to an embodiment of the present invention.
[0026] Figure 6 It is a schematic diagram of the incorrect text recognition result according to an embodiment of the present invention.
[0027] Figure 7It is a step flowchart of a method for identifying Chinese-English subtitles according to an embodiment of the present invention.
[0028] Figure 8 It is a schematic structural diagram of an apparatus for identifying text according to an embodiment of the present invention.
[0029] Figure 9 It is a schematic working flowchart of a system for identifying dialogue subtitles in an audio-visual file according to an embodiment of the present invention.
[0030] Figure 10 It is a schematic structural diagram of an electronic device according to an embodiment of the present invention. Detailed implementation manners
[0031] Next, the technical solutions in the embodiments of the present invention will be described in conjunction with the accompanying drawings in the embodiments of the present invention.
[0032] An embodiment of the present invention proposes a text recognition solution. Before recognizing the text content, the language information corresponding to the text box where the text content is located can be determined through a text classification algorithm, and then a text recognition algorithm corresponding to the language information is selected to recognize the text content.
[0033] As Figure 1 shown, it shows a step flowchart of a method for identifying text according to an embodiment of the present invention. This text recognition method can be applied to a terminal or a server. The text recognition method specifically may include the following steps.
[0034] Step 101, obtain a video frame image to be processed.
[0035] In an embodiment of the present invention, the video frame image may be any frame of image containing text content in video data or audio-visual data, and the present invention does not make specific limitations on the resolution, size, format, etc. of the video frame image.
[0036] Step 102, extract a text image from the video frame image and recognize the feature vector of the text image.
[0037] In an embodiment of the present invention, a text image can be extracted from a video frame image based on any image segmentation algorithm. Moreover, the text image can contain at least one text box. For example, for a certain video frame image of a certain film and television work, a text image is extracted from the video frame image, and the text image contains two text boxes. One text box contains the Chinese dialogue subtitles of the characters in the video frame image, and the other text box contains the English dialogue subtitles of the characters in the video frame image. A text box can be understood as a rectangular box where the text content in the text image is located. In practical applications, a text image is extracted from a video frame image, and then each text box is recognized from the text image. Specifically, the position of a text box can be represented by a pair of pixel point coordinates. For example, the position of a text box is represented by the coordinates of the lower left pixel point and the upper right pixel point of the text box.
[0038] In an embodiment of the present invention, the text image can be input into a neural network model to output a feature vector of the text image. The neural network model can be any network model with an image feature extraction function. The present invention does not specifically limit the network structure, network parameters, training process, etc. of the neural network model.
[0039] Step 103: Classify the text content in each text box according to the feature vector and a preset text arrangement rule to obtain the corresponding language information.
[0040] In an embodiment of the present invention, the text arrangement rule can represent the character information occupied by the text content corresponding to various language information when being displayed. In practical applications, for example, a Chinese character corresponding to the language information of the Chinese language occupies two characters when being displayed. An English word corresponding to the language information of the English language occupies one character when being displayed.
[0041] Step 104: For the text content in each text box, respectively select a text recognition algorithm corresponding to each language information, and recognize the text content in the corresponding text box according to the selected text recognition algorithm.
[0042] In an embodiment of the present invention, one language information can correspond to one text recognition algorithm. If multiple text boxes in the text image correspond to different language information, multiple text recognition algorithms can be selected, and the text content in each text box is recognized respectively by using the selected text recognition algorithms.
[0043] The text recognition solution provided by the embodiments of the present invention adopts the technical means of obtaining a video frame image to be processed, extracting a text image from the video frame image, and recognizing a feature vector of the text image, where the text image contains at least one text box. According to a preset text arrangement rule and the feature vector, the text content in each text box is classified to obtain corresponding language information. The arrangement rule represents the character information occupied by the text content corresponding to each language information when displayed. Furthermore, for the text content in each text box, a text recognition algorithm corresponding to each language information is selected, and the text content in the corresponding text box is recognized according to the selected text recognition algorithm. Before recognizing the text content, the text content is classified to obtain corresponding language information, and then a text recognition algorithm corresponding to the language information is selected for the text content, and the text content is recognized by using the selected text recognition algorithm. In the embodiments of the present invention, by selecting a text recognition algorithm that matches the language in the text content to recognize the text content, it can to a certain extent solve the technical problems of inaccurate recognition such as omission of text recognition for some languages and incorrect text recognition when using one text recognition algorithm to recognize text content containing multiple languages, and achieve the effect of improving the text recognition accuracy.
[0044] In an exemplary embodiment of the present invention, an implementation manner of classifying the text content in each text box according to the feature vector and the preset text arrangement rule to obtain corresponding language information is as follows: According to the feature vector and the text arrangement rule, each character included in the text content in each text box is classified to obtain a classification result. The classification result corresponding to each text box is mapped to the language information corresponding to each text box respectively.
[0045] For example, the text image P contains text boxes T1 and T2. The text content in the text box T1 contains characters h1 and h2, and the text content in the text box T2 contains characters y1 and y2. Then, according to the feature vector of the text image P and the text arrangement rule, each character h1, h2, y1, and y2 included in the text content in the text boxes T1 and T2 can be classified to obtain the respective classification results of the characters h1, h2, y1, and y2. The classification result may include a classification identifier and a confidence level corresponding to the classification identifier.
[0046] In an exemplary embodiment of the present invention, as Figure 2 shown, Figure 2 shows a flowchart of the steps of classifying each character included in the text content in each text box to obtain a classification result.
[0047] An implementation manner of classifying each character included in the text content in each text box according to the feature vector and the text arrangement rule to obtain a classification result includes:
[0048] Step 201: Input the text image into a Convolutional Neural Network (CNN for short) to obtain the image feature vectors of each pixel point of the text image, and then input the image feature vectors into a Recurrent Neural Network (RNN for short) to obtain the text feature vectors of the pixel points.
[0049] Step 202: According to the text feature vectors, count the number of associated pixel points or independent pixel points occupied by each character.
[0050] In the embodiment of the present invention, the associated pixel points are a group of adjacent pixel points, and the group of adjacent pixel points has associated text feature vectors. The independent pixel points refer to the pixel points whose adjacent pixel points do not have associated text feature vectors.
[0051] First, the pixel points occupied by each character can be recorded, and then among the pixel points occupied by each character, it is identified which pixel points are associated pixel points or independent pixel points. Before identifying which pixel points are associated pixel points or independent pixel points, it is necessary to understand the associated text feature vectors first. The associated text feature vectors can represent the same text feature vectors or similar text feature vectors. Among them, the similar text feature vectors can be understood as the vector differences between multiple text feature vectors are within a preset difference range, and this difference range can be preset according to the actual situation.
[0052] For example, if the adjacent pixel points x1 and x2 have the same text feature vector, then the adjacent pixel points x1 and x2 are associated pixel points. If the pixel point x3 does not have associated text feature vectors with the adjacent pixel points x2 and x4, then the pixel point x3 is an independent pixel point.
[0053] The above-mentioned associated pixel points or independent pixel points can represent the text feature vector association relationship between each pixel point and its adjacent pixel points among several adjacent pixel points. If the number of associated pixel points occupied by a certain character is large, it means that the number of pixel points with strong text feature vector association occupied by this character is large. If the number of independent pixel points occupied by a certain character is large, it means that the number of pixel points with weak text feature vector association occupied by this character is large.
[0054] For another example, if the character Z1 occupies the associated pixel points x1 and x2, then the number of associated pixel points occupied by the character Z1 is 2. If the character Z2 occupies the independent pixel point x3, then the number of independent pixel points occupied by the character Z2 is 1.
[0055] Step 203: Classify each character according to the quantity and text arrangement rules to obtain a classification result.
[0056] The text arrangement rule in the embodiment of the present invention includes the correspondence between the quantity range where the above quantity is located and the classification result. In practical applications, for each character, the classification result corresponding to the quantity range where the above quantity is located can be used as the classification result of each character. Specifically, the classification result of a character can be Chinese, English, symbol, etc. Further, symbols can be classified into Chinese or English.
[0057] For example, the text arrangement rule includes the correspondence between the quantity range of associated pixel points and the classification result, the correspondence between the quantity range of independent pixel points and the classification result, and the correspondence between the quantity range of associated pixel points and the quantity range of independent pixel points and the classification result. If a certain character includes associated pixel points and does not include independent pixel points, and the quantity of associated pixel points included in this character is within the quantity range F1, and the quantity range F1 of associated pixel points corresponds to the classification result L1, then the classification result of this character is L1. If a certain character does not include associated pixel points but includes independent pixel points, and the quantity of independent pixel points included in this character is within the quantity range F2, and the quantity range F2 of independent pixel points corresponds to the classification result L2, then the classification result of this character is L2. If a certain character includes both associated pixel points and independent pixel points, and the quantity of associated pixel points included in this character is within the quantity range F3, and the quantity of independent pixel points included in this character is within the quantity range F4, and the quantity range F3 of associated pixel points corresponds to the classification result L3, and the quantity range F4 of independent pixel points corresponds to the classification result L4. When the quantity of associated pixel points is greater than or equal to the quantity of independent pixel points, the classification result of this character is L3; when the quantity of associated pixel points is less than the quantity of independent pixel points, the classification result of this character is L4. It should be noted that the above quantity ranges F1, F2, F3, and F4 can be preset according to the actual situation.
[0058] In an exemplary embodiment of the present invention, an implementation manner of mapping the classification result corresponding to each text box to the language information corresponding to each text box respectively is that for each text box, if the classification results of the characters within the same text box are the same, then the classification results of the characters are used as the language information corresponding to the same text box; if the classification results of the characters within the same text box are different, then the classification results of the characters are jointly used as the language information corresponding to the same text box. For example, the text box T1 is composed of characters Z3 and Z4, and the classification results of the adjacent characters Z3 and Z4 are both Chinese classifications, then the language information of the text box T1 is mapped to the Chinese classification. Another example, the text box T2 is composed of characters Z1 and Z2, the classification result of the character Z1 is the Chinese classification, and the classification result of the character Z2 is the English classification, then the language information of the text box T2 is mapped to the Chinese classification and the English classification.
[0059] In an exemplary embodiment of the present invention, the above language information may include Chinese, English, etc. One implementation of selecting a text recognition algorithm corresponding to the language information is that when the language information corresponding to the text box includes Chinese, the Chinese recognition algorithm corresponding to Chinese is selected. It can be understood that when the language information corresponding to the text box only includes Chinese, or when the language information corresponding to the text box includes Chinese and English, the Chinese recognition algorithm corresponding to Chinese is selected. When the language information corresponding to the text box is English, the English recognition algorithm corresponding to English is selected. It can be understood that when the language information corresponding to the text box is only English, the English recognition algorithm corresponding to English is selected.
[0060] Based on the above relevant description of the embodiment of a text recognition method, the following introduces a Chinese-English text recognition solution. When extracting the line subtitles in a video, there are usually cases of mixed Chinese-English subtitles. English subtitles are often missed during the recognition process. Therefore, two text recognition algorithms can be used to separately recognize Chinese / English characters. In practical applications, the text detection algorithm can be used to obtain the text area first, but it is not possible to determine the language to which the text in each area belongs, and it is not possible to judge which text recognition algorithm should be used for that area to obtain the specific text content. The embodiment of the present invention can effectively connect the text detection algorithm and the text recognition algorithm. By selecting a better text recognition algorithm, the effect of further improving the text recognition accuracy rate in the lines is achieved.
[0061] In order to distinguish which text recognition algorithm should be used for the text area specifically, the embodiment of the present invention proposes a text classification algorithm, which selects a more suitable text recognition algorithm for text recognition by judging the specific language information to which the text area belongs.
[0062] The text classification algorithm proposed in the embodiment of the present invention uses a scheme similar to the text recognition algorithm. First, the feature vector of the image containing the text is extracted. For the extracted feature vector, combined with the text arrangement rule, each character in the text is classified, and then the language information to which each character belongs is obtained through the decoder.
[0063] Different from the text recognition algorithm, in the process of classifying characters, the number of categories required for classification by the text classification algorithm is less. For the application scenario of the embodiment of the present invention, only two categories need to be divided from the language perspective: Chinese and English. In a conventional text recognition algorithm, taking English as an example, there are 26 uppercase letters and 26 lowercase letters. In the text recognition algorithm, English characters need to be divided into at least 52 categories. And Chinese characters need to be divided into 7,000 - 10,000 categories according to common usage. The embodiment of the present invention greatly reduces the number of classification categories and improves the accuracy of single-character classification at the same time.
[0064] Moreover, during the classification process, the language corresponding to each character is obtained first, and then the language corresponding to the entire text line is obtained. When obtaining the language corresponding to each character in the text line, if there are omissions of individual characters due to different occupancy amounts of single characters (mainly for English characters), it will not affect the classification of the language of the text line. The reason is that if the classification results of all characters in the text line are English classifications, even if a certain character is omitted, the language information corresponding to the text line is still the English language; if the classification results of all characters in the text line include English classifications and Chinese classifications, even if a certain character is omitted, the language information corresponding to the text line is still the Chinese language. Therefore, as a preprocessing algorithm for the text recognition algorithm, the text classification algorithm provided by the embodiments of the present invention can achieve an accuracy rate of more than 99.5% in the language classification of movie and TV subtitles.
[0065] The above-mentioned text classification algorithm proposed by the embodiments of the present invention can be obtained by training a neural network model. As Figure 3 shown, a schematic diagram of the training process of a text classification algorithm according to an embodiment of the present invention is shown. The training sample image is first input into a convolutional neural network, and the convolutional neural network is used to extract the image feature vector of the training sample image. Then, the image feature vector is input into a recurrent neural network, and the recurrent neural network is used to extract the text feature vector. After obtaining the text feature vector, the online real classification loss function can be used as the loss function of the initial network model of the text classification algorithm. The training sample image and the text feature vector are used as the input items of the initial network model. By adjusting the network parameters of the initial network model until the training result of the network model meets the expected requirements, the text classification algorithm is finally obtained after training.
[0066] As Figure 4 shown, a schematic diagram of the classification process of a text line according to an embodiment of the present invention is shown. The text line in the embodiment of the present invention can be understood as the above-mentioned text box. Classify the text line "Okay I will", and the classification result is "C-CSESE-E-E-E". The classification result is sorted into "CCSESEEEE". Among them, "C" represents Chinese characters, "S" represents symbol characters, and "E" represents English characters. It should be noted that symbol characters can be further divided into Chinese symbol characters and English symbol characters. That is to say, "S" can belong to "C" or "E". In the actual text recognition process, due to the different number of characters occupied by Chinese and English, there will be problems such as character omission in English. Figure 5 It is a schematic diagram of the correct text recognition result. Figure 6 It is a schematic diagram of the wrong text recognition result. Figure 6The situation of missing characters (l) also occurs in the text classification algorithm. However, as an intermediate step in the entire text recognition process, even if some English characters are missing, the language information corresponding to this text line is still Chinese classification and English classification. Moreover, subsequently, a Chinese recognition algorithm corresponding to the Chinese language is selected for this text line. Therefore, the situation of character omission does not affect the text recognition result.
[0067] In one embodiment, if a text line contains both Chinese text and English text, the language information corresponding to this text line includes the Chinese language and the English language. Furthermore, a Chinese recognition algorithm corresponding to the Chinese language can be used to recognize the Chinese text content of this text line, and an English recognition algorithm corresponding to the English language can be used to recognize the English text content of this text line.
[0068] As Figure 7 shown, a step flowchart of a method for recognizing Chinese and English subtitles according to an embodiment of the present invention is shown. An image containing subtitles is input into a detector. The detector is used to detect the text box where the subtitles are located from the image. A text classification algorithm is used to classify the subtitles in the text box to obtain a classification result. It is determined whether the classification result indicates that each character in the subtitles is English. If the classification result indicates that each character in the subtitles is English, an English recognition algorithm is selected to recognize the subtitles in the text box to obtain English line text. If the classification result indicates that not every character in the subtitles is English, that is, the subtitles also contain Chinese, a Chinese recognition algorithm is selected to recognize the subtitles in the text box to obtain Chinese line text or Chinese and English line text. Sorting operations such as storing, outputting, packing, and displaying the recognized English line text, Chinese line text, or Chinese and English line text are performed.
[0069] As Figure 8 shown, a schematic structural diagram of a text recognition device according to an embodiment of the present invention is shown. The text recognition device may include the following modules.
[0070] An acquisition module 91, configured to acquire a video frame image to be processed;
[0071] A processing module 92, configured to extract a text image from the video frame image and recognize a feature vector of the text image, where the text image includes at least one text box;
[0072] A classification module 93, configured to classify the text content in each text box according to the feature vector and a preset text arrangement rule to obtain corresponding language information, where the text arrangement rule represents the character information occupied by the text content corresponding to each language information when displayed;
[0073] An identification module 94, configured to respectively select a text recognition algorithm corresponding to each language information for the text content in each of the text boxes, and recognize the text content in the corresponding text box according to the selected text recognition algorithm, where the language information corresponding to the text content is the same as the language information corresponding to the selected text recognition algorithm.
[0074] In an exemplary embodiment of the present invention, the classification module 93 includes:
[0075] A character classification module, configured to classify each character included in the text content in each text box according to the feature vector and the text arrangement rule to obtain a classification result;
[0076] A language mapping module, configured to map the classification result corresponding to each text box to the language information corresponding to each text box respectively.
[0077] In an exemplary embodiment of the present invention, the character classification module includes:
[0078] A feature extraction module, configured to input the text image into a convolutional neural network to obtain an image feature vector of each pixel point of the text image, and input the image feature vector into a recurrent neural network to obtain a text feature vector of the pixel point;
[0079] A quantity statistics module, configured to count the number of associated pixel points or independent pixel points occupied by each character according to the text feature vector, where the associated pixel points are a group of adjacent pixel points, and the group of adjacent pixel points has an associated text feature vector, and the independent pixel point is a pixel point whose adjacent pixel points do not have an associated text feature vector;
[0080] A result classification module, configured to classify each character according to the quantity and the text arrangement rule to obtain the classification result.
[0081] In an exemplary embodiment of the present invention, the text arrangement rule includes a correspondence relationship between the quantity range where the quantity is located and the classification result;
[0082] The result classification module is configured to, for each character, use the classification result having the correspondence relationship with the quantity range where the quantity is located as the classification result of each character.
[0083] In an exemplary embodiment of the present invention, the language mapping module is configured to, for each of the text boxes, if the classification results of the characters within the same text box are the same, use the classification results of the characters as the language information corresponding to the same text box; if the classification results of the characters within the same text box are different, use the classification results of the characters together as the language information corresponding to the same text box.
[0084] In an exemplary embodiment of the present invention, the language information includes Chinese and English.
[0085] The recognition module 94 is configured to, when the language information corresponding to the text box only includes the Chinese language, select a Chinese recognition algorithm corresponding to the Chinese language; when the language information corresponding to the text box only includes the English language, select an English recognition algorithm corresponding to the English language; when the language information corresponding to the text box includes the Chinese language and the English language, select the Chinese recognition algorithm corresponding to the Chinese language.
[0086] Based on the description of the above embodiments of the recognition method and device for text, etc., a recognition system for dialogue subtitles in an audio-visual file will be introduced below. The recognition system can be composed of hardware devices such as a personal computer, a server, or a server cluster. A dialogue subtitle detection framework is deployed in the recognition system, and the dialogue subtitle detection framework mainly includes a dialogue subtitle detection model, a dialogue subtitle filtering model, a dialogue subtitle tracking model, and a text classification model.
[0087] Referring to Figure 9 , a schematic diagram of the working process of a recognition system for dialogue subtitles in an audio-visual file according to an embodiment of the present invention is shown. In the actual application process, a complete audio-visual file is input into the recognition system, and a video frame result is extracted from the audio-visual file by using the graphics processing unit (GPU) of the hardware device. The specific video frame result can be partial frame images and full-frame images. Among them, the partial frame images can be three frame images extracted from a 1-second time period of the audio-visual file. In order to reduce the computational amount of extracting the video frame result, the extraction process can be performed on partial video frame images of the audio-visual file. For example, only the lower 1 / 3 part of the video frame image is extracted. The extracted partial frame images can be continuously written into a memory queue. The extracted full-frame images can be stored on the hard disk.
[0088] Based on multi-threading, multiple subtitle detection models are started to continuously read partial frame images from the memory queue and splice the read partial frame images. For example, 3 continuously read partial frame images are spliced into 1 frame image. The subtitle detection model locates the positions of all text boxes from the spliced frame image.
[0089] The subtitle filtering model counts the frequencies of all text boxes appearing in the time axis, determines a high-frequency heat map area according to the frequencies, and then uses this high-frequency heat map area as the filtering standard area for subtitles and non-subtitles, and filters all text boxes into subtitle text boxes and non-subtitle text boxes using this filtering standard area. Among them, non-subtitle text boxes are discarded and do not participate in subsequent processing.
[0090] The subtitle tracking model extracts the deep features of Optical Character Recognition (OCR) from the subtitle text boxes and performs tracking processing based on the full-frame images stored on the hard disk to obtain the start and end time information of each subtitle text box appearing in the audio-visual file.
[0091] The text classification model determines the language information corresponding to each subtitle text box, and transmits the deep features to the corresponding OCR prediction network according to the language information. The OCR prediction network uses their respective text recognition algorithms to recognize the subtitles in the subtitle text boxes to obtain subtitle results. For example, the language information can be Chinese language, English language. The Chinese language corresponds to the Chinese OCR prediction network, and the English language corresponds to the English OCR prediction network.
[0092] The recognition system provided by the embodiments of the present invention constructs a subtitle frame-level recognition Software Development Kit (SDK). For the same audio-visual file, the ratio of the time for subtitle recognition on the GPU to the time for subtitle recognition through this SDK is about 1:0.11, which greatly improves the speed and accuracy of subtitle recognition. According to the recognized subtitle results, external subtitles for the audio-visual file can be generated, and the embedded subtitles of the audio-visual file can be quickly converted into external subtitles.
[0093] When the recognition system provided by the embodiments of the present invention filters text boxes, it uses the time domain information of the text boxes to filter the text boxes into subtitle text boxes and non-subtitle text boxes, and the accuracy of text box filtering is very high. When the audio-visual file is a movie or TV drama work, the filtering accuracy exceeds 99%; when the audio-visual file is a variety show work, the filtering accuracy exceeds 98.5%.
[0094] The recognition system provided by the embodiments of the present invention can also extract video frames for specified areas such as name strips, lyric versions, and lyrics in variety shows, and then perform processes such as positioning, filtering, tracking, classification, and recognition of names, lyrics, etc., to achieve text recognition of names, lyrics, etc.
[0095] In addition to being able to recognize the line subtitles in Chinese and English languages, the recognition system provided by the embodiments of the present invention can also achieve intelligent recognition of the line subtitles in Chinese-English, Chinese-Japanese, Chinese-Korean, etc. in bilingual video files, and support the recognition of the line subtitles in multi-language audio-video files.
[0096] The embodiments of the present invention also provide an electronic device, as Figure 10 shown, including a processor 111, a communication interface 112, a memory 113, and a communication bus 114. Among them, the processor 111, the communication interface 112, and the memory 113 complete mutual communication through the communication bus 114.
[0097] The memory 113 is used to store a computer program.
[0098] When the processor 111 is used to execute the program stored in the memory 113, the following steps are implemented:
[0099] Obtain a video frame image to be processed; extract a text image from the video frame image, and identify a feature vector of the text image, where the text image includes at least one text box; according to the feature vector and a preset text arrangement rule, classify the text content in each text box to obtain corresponding language information, and the text arrangement rule represents the character information occupied by the text content corresponding to each language information when displayed; for the text content in each text box, respectively select a text recognition algorithm corresponding to each language information, and recognize the text content in the corresponding text box according to the selected text recognition algorithm, and the language information corresponding to the text content is the same as the language information corresponding to the selected text recognition algorithm.
[0100] The classifying the text content in each text box according to the feature vector and the preset text arrangement rule to obtain corresponding language information includes: classifying each character included in the text content in each text box according to the feature vector and the text arrangement rule to obtain a classification result; mapping the classification result corresponding to each text box to the language information corresponding to each text box respectively.
[0101] Classifying each character included in the text content within each of the text boxes according to the feature vector and the text arrangement rule to obtain a classification result, including: inputting the text image into a convolutional neural network to obtain an image feature vector of each pixel point of the text image, and inputting the image feature vector into a recurrent neural network to obtain a text feature vector of the pixel point; counting the number of associated pixel points or independent pixel points occupied by each character according to the text feature vector, where the associated pixel points are a group of adjacent pixel points and the group of adjacent pixel points has an associated text feature vector, and the independent pixel points are pixel points whose adjacent pixel points do not have an associated text feature vector; classifying each character according to the number and the text arrangement rule to obtain the classification result.
[0102] The text arrangement rule includes the correspondence between the quantity range where the quantity is located and the classification result; the classifying each character according to the quantity and the text arrangement rule to obtain the classification result includes: for each character, taking the classification result corresponding to the quantity range where the quantity is located as the classification result of each character.
[0103] Mapping the classification result corresponding to each text box to the language information corresponding to each text box respectively, including: for each text box, if the classification results of the characters within the same text box are the same, taking the classification results of the characters as the language information corresponding to the same text box; if the classification results of the characters within the same text box are different, taking the classification results of the characters together as the language information corresponding to the same text box.
[0104] The language information includes Chinese language and English language; respectively selecting a text recognition algorithm corresponding to each language information, including: when the language information corresponding to the text box only includes the Chinese language, selecting a Chinese recognition algorithm corresponding to the Chinese language; when the language information corresponding to the text box only includes the English language, selecting an English recognition algorithm corresponding to the English language; when the language information corresponding to the text box includes the Chinese language and the English language, selecting the Chinese recognition algorithm corresponding to the Chinese language.
[0105] The communication bus mentioned in the above terminal may be a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into an address bus, a data bus, a control bus, etc. For the convenience of representation, only a thick line is used in the figure, but it does not mean that there is only one bus or one type of bus.
[0106] The communication interface is used for communication between the above terminal and other devices.
[0107] The memory may include a Random Access Memory (RAM), or may also include a non-volatile memory, such as at least one disk memory. Optionally, the memory may also be at least one storage device located far from the aforementioned processor.
[0108] The above-mentioned processor may be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it may also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.
[0109] In another embodiment provided by the present invention, there is also provided a computer-readable storage medium, in which instructions are stored. When it runs on a computer, it causes the computer to execute the text recognition method described in any one of the above embodiments.
[0110] In another embodiment provided by the present invention, there is also provided a computer program product containing instructions. When it runs on a computer, it causes the computer to execute the text recognition method described in any one of the above embodiments.
[0111] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present invention are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wire (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that includes one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid state disk (SSD)).
[0112] It should be noted that, in this document, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including", or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article, or device that includes a series of elements includes not only those elements but also other elements that are not explicitly listed, or also includes elements that are inherent to such process, method, article, or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article, or device that includes the element.
[0113] Each embodiment in this specification is described in a related manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and reference can be made to the corresponding part of the method embodiment for the relevant content.
[0114] The above are only the preferred embodiments of the present invention and are not intended to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention are all included in the protection scope of the present invention.
Claims
1. A method for recognizing text, characterized in that, Including: Obtain a video frame image to be processed; Extract a text image from the video frame image, and identify a feature vector of the text image, where the text image includes at least one text box; According to the feature vector and a preset text arrangement rule, classify the text content in each text box to obtain corresponding language information. The text arrangement rule represents the character information occupied by the text content corresponding to each language information when displayed. Different language information occupies different numbers of characters. The characters include associated pixel points and / or independent pixel points. The text arrangement rule also includes the corresponding relationship between the number range of associated pixel points and / or independent pixel points and the classification result. The associated pixel points are a group of adjacent pixel points, and this group of adjacent pixel points has an associated text feature vector. The independent pixel points are pixel points whose adjacent pixel points do not have the associated text feature vector; For the text content in each text box, respectively select a text recognition algorithm corresponding to each language information, and recognize the text content in the corresponding text box according to the selected text recognition algorithm. The language information corresponding to the text content is the same as the language information corresponding to the selected text recognition algorithm.
2. The method according to claim 1, wherein The step of classifying the text content in each text box according to the feature vector and the preset text arrangement rule to obtain corresponding language information includes: Classify each character included in the text content in each text box according to the feature vector and the text arrangement rule to obtain a classification result; Map the classification result corresponding to each text box to the language information corresponding to each text box respectively.
3. The method according to claim 2, wherein The step of classifying each character included in the text content in each text box according to the feature vector and the text arrangement rule to obtain a classification result includes: Input the text image into a convolutional neural network to obtain an image feature vector of each pixel point of the text image, and input the image feature vector into a recurrent neural network to obtain a text feature vector of the pixel point; According to the text feature vector, count the number of associated pixel points or independent pixel points occupied by each character; Classify each character according to the number and the text arrangement rule to obtain the classification result.
4. The method according to claim 3, characterized in that The text arrangement rule includes the corresponding relationship between the number range where the number is located and the classification result; The step of classifying each character according to the number and the text arrangement rule to obtain the classification result includes: For each character, use the classification result having the corresponding relationship with the number range where the number is located as the classification result of each character.
5. The method according to claim 2, wherein The step of mapping the classification result corresponding to each text box to the language information corresponding to each text box respectively includes: For each of the text boxes, if the classification results of all the characters within the same text box are the same, then use the classification results of all the characters as the language information corresponding to the same text box; if the classification results of all the characters within the same text box are different, then use the classification results of all the characters together as the language information corresponding to the same text box.
6. The method according to any one of claims 1 to 5, characterized in that, The language information includes Chinese and English. The method of respectively selecting text recognition algorithms corresponding to each of the language information includes: When the language information corresponding to the text box only includes the Chinese language, select the Chinese recognition algorithm corresponding to the Chinese language. When the language information corresponding to the text box only includes the English language, select the English recognition algorithm corresponding to the English language. When the language information corresponding to the text box includes the Chinese language and the English language, select the Chinese recognition algorithm corresponding to the Chinese language.
7. An apparatus for recognizing text, characterized in that, It includes: An acquisition module, configured to acquire a video frame image to be processed. A processing module, configured to extract a text image from the video frame image and identify a feature vector of the text image, where the text image includes at least one text box. A classification module, configured to classify the text content within each text box according to the feature vector and a preset text arrangement rule to obtain corresponding language information. The text arrangement rule represents the character information occupied by the text content corresponding to each language information during display. Different language information occupies different numbers of characters. The characters include associated pixel points and / or independent pixel points. The text arrangement rule also includes the corresponding relationship between the quantity range of associated pixel points and / or independent pixel points and the classification result. The associated pixel points are a group of adjacent pixel points, and this group of adjacent pixel points has an associated text feature vector. The independent pixel points are pixel points whose adjacent pixel points do not have an associated text feature vector. An identification module, configured to respectively select text recognition algorithms corresponding to each of the language information for the text content within each text box, and identify the text content in the corresponding text box according to the selected text recognition algorithm. The language information corresponding to the text content is the same as the language information corresponding to the selected text recognition algorithm.
8. The device according to claim 7, characterized in that, The classification module includes: A character classification module, configured to classify each character included in the text content within each text box according to the feature vector and the text arrangement rule to obtain a classification result. A language mapping module, configured to map the classification result corresponding to each text box to the language information corresponding to each text box respectively.
9. An electronic device, characterized in that, It includes a processor, a communication interface, a memory, and a communication bus. Among them, the processor, the communication interface, and the memory complete communication with each other through the communication bus. The memory is used to store a computer program. The processor is configured to implement the method according to any one of claims 1 to 6 when executing the program stored on the memory.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Text recognition method and device, computer equipment and storage medium
CN111832657A