Text processing method, device, electronic device and storage medium

By performing text detection and clustering network on video frame sequences, the problems of misidentification and misfiltering in text recognition and translation in the prior art are solved, and higher recognition accuracy and translation efficiency are achieved.

CN114596522BActive Publication Date: 2025-05-30BEIJING IQIYI TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210134995.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-02-14
Publication Date
2025-05-30
Estimated Expiration
2042-02-14

AI Technical Summary

Technical Problem

In text search, recognition and translation, the prior art relies on the position information of text lines for judgment, which can easily lead to misidentification of non-target text and position error of target text, resulting in poor results, manual positioning and extraction, which is inefficient and costly.

Method used

By performing text detection on the video frame sequence, multiple text lines are determined, initial classification is performed based on position information, font feature information is obtained, text lines are clustered using a pre-constructed clustering network, clustering the clustering results are adjusted for secondary classification, and the final line text is determined.

Benefits of technology

It improves the accuracy of line text recognition, reduces the need for manual verification, reduces the cost, and improves the efficiency and accuracy of text translation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114596522B_ABST
    Figure CN114596522B_ABST
Patent Text Reader

Abstract

An embodiment of the present invention provides a text processing method, apparatus, electronic device, and storage medium, which relate to the technical field of image processing. The method includes: performing text detection on a video frame sequence of a current video to determine a plurality of text lines; performing initial classification on the text lines according to the position information of the text lines to obtain a first set and a second set, where the first set is recognized as dialogue text and the second set is recognized as non-dialogue text; performing clustering on the text lines according to the font feature information of the text lines and a preset clustering network to obtain a plurality of clustering results, and the text lines in the same clustering result have the same font; performing secondary classification on the text lines according to the plurality of clustering results to determine the final dialogue text of the current video. This method performs secondary filtering on the initially classified text lines through font feature information, and can filter out misdetected non-dialogue text in the dialogue text and recall the missed-recognized dialogue text, improving the accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and particularly to a text processing method, apparatus, electronic device, and storage medium. Background Art

[0002] In the fields of text search, text recognition, text translation, etc., it is first necessary to locate the target text in the target area, and then the located target text can be recognized and translated. The currently common method is to determine whether a text line belongs to the target text in the target area based on the position information of the text line. However, using only the position information of the text line to make a judgment will result in the situation that when non-target text appears in the target area, due to the successful matching of the position information, the non-target text is misrecognized as the target text. Similarly, there will also be cases where the position of the target text has a large error or just reaches the critical threshold, resulting in matching failure and misfiltering. As a result, the effects of text search, recognition, or translation are poor, and it may be necessary to manually perform further positioning and extraction of the target text, which is inefficient and costly. For example, in the process of searching for and translating the lines of a variety show, it is necessary to locate the line text in the variety show video and filter out non-line text, such as the promotional slogans of sponsors and advertisements. However, using only the position information of the text line to make a judgment will result in the situation that when other text lines appear in the position of the line text, due to the successful matching of the position information, the non-line text line may be misrecognized as the line text line. Similarly, there will also be cases where the position of the line text has a large error or just reaches the critical threshold, resulting in matching failure and misfiltering. Summary of the Invention

[0003] To solve the above technical problems or at least partially solve the above technical problems, embodiments of the present invention provide a text processing method, apparatus, electronic device, and storage medium.

[0004] In the first aspect of the embodiments of the present invention, a text processing method is first provided, including: performing text detection on a video frame sequence of a current video to determine a plurality of text lines in the video frame sequence; performing initial classification on the text lines of the current video according to the position information of the plurality of text lines to obtain a first set and a second set, where the text lines in the first set are recognized as line text, and the text lines in the second set are recognized as non-line text; performing clustering on the text lines of the current video according to the font feature information corresponding to the plurality of text lines and a pre-constructed clustering network to obtain a plurality of clustering results, and the text lines in the same clustering result have the same font; adjusting the text lines in the first set and the second set according to the plurality of clustering results to perform secondary classification on the text lines of the current video and determine the final line text of the current video.

[0005] Optionally, the pre-constructed clustering network is obtained according to the following process: Obtain a training sample set, where the training samples in the training sample set include real line images and simulated line images; According to a pre-constructed font feature extraction model, obtain the font feature information of each training sample in the training sample set; Train the clustering network according to the font feature information of each training sample.

[0006] Optionally, according to the multiple clustering results, adjust the text lines in the first set and the second set, including: For each clustering result, determine the identifier of the clustering result according to the first set and the second set; Wherein, the identifier of the clustering result includes a line identifier and a non-line identifier; Adjust the text lines in the first set and the second set according to the identifier of the clustering result.

[0007] Optionally, determining the identifier of the clustering result according to the first set and the second set includes: For each clustering result, determine the first proportion of the text lines belonging to the first set in the clustering result; In the case where the first proportion is greater than a preset first threshold, determine the identifier of the clustering result as a line identifier; In the case where the first proportion is less than or equal to the preset first threshold, determine the identifier of the clustering result as a non-line identifier.

[0008] Optionally, adjusting the text lines in the first set and the second set according to the multiple clustering results includes: For the clustering result with a line identifier, determine the first text line to be recognized belonging to the second set in the clustering result, and perform secondary classification on the first text line to be recognized according to the position information of the first text line to be recognized to determine whether the first text line to be recognized is a line text; If so, migrate the first text line to be recognized from the second set to the first set; For the clustering result with a non-line identifier, determine the second text line to be recognized belonging to the first set in the clustering result; Determine that the second text line to be recognized is non-line text and migrate the second text line to be recognized from the second set to the first set.

[0009] Optionally, the initial classification of the text lines of the current video according to the position information of the multiple text lines includes: determining the frequency of occurrence of text lines at each pixel point on the video frame of the current video according to the position information of the multiple text lines; determining the area where the pixel points with the frequency of occurrence of text lines greater than a preset second threshold are located as the dialogue area; determining the width information and height information of the text area corresponding to the text line according to the position information of the text line; calculating the area intersection-over-union ratio of the text area and the dialogue area according to the width information and height information of the text area corresponding to the text line; calculating the height intersection-over-union ratio of the text area and the dialogue area according to the height information of the text area corresponding to the text line; calculating the width intersection-over-union ratio of the text area and the dialogue area according to the width information of the text area corresponding to the text line; determining that the text line is dialogue text when the area intersection-over-union ratio of the text line is greater than a preset third threshold, the height intersection-over-union ratio is greater than a preset fourth threshold, and the width intersection-over-union ratio is greater than a preset fifth threshold; determining that the text line is dialogue text when the area intersection-over-union ratio of the text line is greater than a preset third threshold, the height intersection-over-union ratio is greater than a preset fourth threshold, the width intersection-over-union ratio is not greater than a preset fifth threshold, but the text line falls within the range of the dialogue area in the width direction.

[0010] Optionally, the secondary classification of the text line to be recognized according to the position information of the first text line to be recognized includes: updating the preset third threshold to a preset sixth threshold and updating the preset fourth threshold to a preset seventh threshold, where the preset sixth threshold is less than the preset third threshold, and the preset seventh threshold is less than the preset fourth threshold; determining that the text line is dialogue text when the area intersection-over-union ratio of the text line is greater than a preset sixth threshold, the height intersection-over-union ratio is greater than a preset seventh threshold, and the width intersection-over-union ratio is greater than a preset fifth threshold; determining that the text line is dialogue text when the area intersection-over-union ratio of the text line is greater than a preset sixth threshold, the height intersection-over-union ratio is greater than a preset seventh threshold, the width intersection-over-union ratio is not greater than a preset fifth threshold, but the text line falls within the range of the dialogue area in the width direction.

[0011] In a second aspect of the embodiments of the present invention, a text processing device is provided, including: a text detection module for detecting text in a video frame sequence of a current video to determine a plurality of text lines in the video frame sequence; an identification module for initially classifying the text lines of the current video according to the position information of the plurality of text lines to obtain a first set and a second set, wherein the text lines in the first set are identified as line text, and the text lines in the second set are identified as non-line text; a clustering module for clustering the text lines of the current video according to the font feature information corresponding to the plurality of text lines and a pre-constructed clustering network to obtain a plurality of clustering results, and the text lines in the same clustering result have the same font; an updating module for adjusting the text lines in the first set and the second set according to the plurality of clustering results to perform secondary classification on the text lines of the current video and determine the final line text of the current video.

[0012] Optionally, the device further includes a training module for: obtaining a training sample set, where the training samples in the training sample set include real line images and simulated line images; obtaining the font feature information of each training sample in the training sample set according to a pre-constructed font feature extraction model; and training the clustering network according to the font feature information of each training sample.

[0013] Optionally, the updating module is further configured to: for each clustering result, determine the identifier of the clustering result according to the first set and the second set; wherein the identifier of the clustering result includes a line identifier and a non-line identifier; and adjust the text lines in the first set and the second set according to the identifier of the clustering result.

[0014] Optionally, the updating module is further configured to: for each clustering result, determine the first proportion of the text lines belonging to the first set in the clustering result; in a case where the first proportion is greater than a preset first threshold, determine the identifier of the clustering result as a line identifier; and in a case where the first proportion is less than or equal to the preset first threshold, determine the identifier of the clustering result as a non-line identifier.

[0015] Optionally, the updating module is further configured to: for the clustering result marked with a line identifier, determine the first text line to be recognized that belongs to the second set in the clustering result, and perform secondary classification on the first text line to be recognized according to the position information of the first text line to be recognized, so as to determine whether the first text line to be recognized is a line text; if so, migrate the first text line to be recognized from the second set to the first set; for the clustering result marked with a non-line identifier, determine the second text line to be recognized that belongs to the first set in the clustering result; determine that the second text line to be recognized is non-line text, and migrate the second text line to be recognized from the second set to the first set.

[0016] Optionally, the recognition module is further configured to: according to the position information of the multiple text lines, determine the frequency of occurrence of text lines at each pixel point on the video frame of the current video; determine the area where the pixel points with the frequency of occurrence of text lines greater than a preset second threshold form as the line area; according to the position information of the text line, determine the width information and height information of the text area corresponding to the text line; calculate the area intersection-over-union ratio of the text area and the line area according to the width information and height information of the text area corresponding to the text line; calculate the height intersection-over-union ratio of the text area and the line area according to the height information of the text area corresponding to the text line; calculate the width intersection-over-union ratio of the text area and the line area according to the width information of the text area corresponding to the text line; when the area intersection-over-union ratio of the text line is greater than a preset third threshold, the height intersection-over-union ratio is greater than a preset fourth threshold, and the width intersection-over-union ratio is greater than a preset fifth threshold, determine that the text line is line text; when the area intersection-over-union ratio of the text line is greater than a preset third threshold, the height intersection-over-union ratio is greater than a preset fourth threshold, and the width intersection-over-union ratio is not greater than a preset fifth threshold but the text line falls within the range of the line area in the width direction, determine that the text line is line text.

[0017] Optionally, the updating module is further configured to: update the preset third threshold to a preset sixth threshold and update the preset fourth threshold to a preset seventh threshold, where the preset sixth threshold is less than the preset third threshold, and the preset seventh threshold is less than the preset fourth threshold; when the area intersection-over-union ratio of the text line is greater than the preset sixth threshold, the height intersection-over-union ratio is greater than the preset seventh threshold, and the width intersection-over-union ratio is greater than the preset fifth threshold, determine that the text line is line text; when the area intersection-over-union ratio of the text line is greater than the preset sixth threshold, the height intersection-over-union ratio is greater than the preset seventh threshold, and the width intersection-over-union ratio is not greater than the preset fifth threshold but the text line falls within the range of the line area in the width direction, determine that the text line is line text.

[0018] In a third aspect of the embodiments of the present invention, an electronic device is provided, including a processor, a communication interface, a memory, and a communication bus. Among them, the processor, the communication interface, and the memory complete communication with each other through the communication bus; the memory is used to store a computer program; when the processor executes the program stored on the memory, the text processing method of any embodiment of the present invention is implemented.

[0019] In a fourth aspect of the embodiments of the present invention, a computer-readable storage medium is provided, on which a computer program is stored, and when the program is executed by a processor, the text processing method of any embodiment of the present invention is implemented.

[0020] The text processing method, device, electronic device, and computer-readable storage medium provided by the embodiments of the present invention detect text in the video frame sequence of the current video, determine multiple text lines in the video frame sequence, and perform an initial classification on the text lines of the current video according to the position information of the multiple text lines to preliminarily determine whether the text line is a line text line, so as to divide the text lines of the current video into two sets. Then, the font feature information of each text line is obtained, and all text lines are clustered according to the font feature information to obtain multiple clustering results, where the text lines in each clustering result have the same font; then, according to the obtained clustering results, the two obtained sets are filtered again to filter out the mis-identified non-line text lines in the line text line set and recall the missed-identified line text lines in the non-line text line set, thereby improving the accuracy of the line recognition result. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art.

[0022] Figure 1 Schematically shows a schematic diagram of the main process of the text processing method of the embodiments of the present invention;

[0023] Figure 2 Schematically shows a schematic diagram of the font feature information extracted by the text processing method of the embodiments of the present invention;

[0024] Figure 3 Schematically shows a schematic diagram of the sub-process of the text processing method of the embodiments of the present invention;

[0025] Figures 4 - 7 Schematically shows schematic diagrams of the clustering results of the text processing method of the embodiments of the present invention respectively;

[0026] Figures 8 - 9 Schematically shows a schematic diagram of adjusting the classification result by the text processing method of the embodiments of the present invention;

[0027] Figure 10 Schematically shows a structural diagram of a text processing device according to an embodiment of the present invention;

[0028] Figure 11 Schematically shows a structural diagram of an electronic device applicable to a text processing method according to an embodiment of the present invention. Detailed implementation manners

[0029] The technical solutions in the embodiments of the present invention will be described below with reference to the accompanying drawings in the embodiments of the present invention.

[0030] In application scenarios such as scene text recognition, text search, and text translation, it is first necessary to accurately locate the target text, and then recognize or translate the target text. For example, the target text is the line text of the current video, where the line text is the dialogue and monologue between each character in the current video, that is, the language spoken by each character in the current video. In the production translation scenario of variety show lines, it is necessary to locate the line text rows in the variety show video and filter out non-line text rows, such as filtering out the promotional slogans of sponsors and advertisements. The line text rows in the variety show video include the dialogue between the program host and guests, the lyrics of the songs sung, and the name bars, where the name bars are used to indicate the names of the characters who speak the current dialogue or sing the current song. Currently, the common method is to determine whether it is a line text row based on the position information of the text row. However, only using the position information of the text row to judge it will cause that when other text rows appear in the position of the line text row, it may be misrecognized as a line text row due to the successful matching of the position information. Similarly, there will also be cases where the position of the line text row has a large error or just at the critical threshold, resulting in the failure of the matching and misfiltering.

[0031] In view of the above technical problems, considering that in a variety show video, usually the same font is used to display the same type of logo or the same category of text, and different fonts are used to display different logos or different categories of text. For example, different fonts are used to display the lines of the variety show video, the advertisement text, and the scrolling subtitles at the end of the video. Therefore, in the embodiments of the present invention, by clustering the font features of the text rows appearing in the video, the text rows with the same font features are clustered into the same clustering result, so as to cluster the text belonging to the same logo or the same category into the same clustering result. Then, the clustering result is merged with the classification result obtained according to the position information of the text row, and the text rows of other fonts that are misretained in the current classification result are filtered out. At the same time, the lines with a large difference between the position information of the text detection and the line area are recalled through the clustering result of the line identification information, thereby improving the accuracy of the text row recognition. For text translation, it can effectively improve the accuracy and efficiency of the line translation result, avoid further manual verification, and effectively reduce the cost.

[0032] To more clearly understand the technical solutions of the embodiments of the present invention, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings.

[0033] Figure 1 A schematic diagram schematically showing the main process of the text processing method according to the embodiment of the present invention is as Figure 1 shown, and the method includes:

[0034] Step 101: Perform text detection on the video frame sequence of the current video to determine a plurality of text lines in the video frame sequence.

[0035] In this embodiment, the current video can be various types of variety shows (a variety show is a variety show that combines multiple art forms and has entertainment), such as interview shows, song shows, etc. The video frames in the video frame sequence can be all the video frames of the current video, or the video frames obtained by extracting frames from all the video frames of the current video. This step performs text detection on each video frame in the video frame sequence to determine whether there is a text line in the video frame. If there is a text line in the video frame, the position information of the text line can be determined, such as the coordinates of the text line in the video frame. Through the coordinates, the width and height of the text area corresponding to the text line can be determined, that is, the size of the text area can be determined. If there is no text line in the video frame, the next video frame is detected.

[0036] As an example, a CTPN text detector can be used to detect text lines in video frames. Among them, CTPN (Detecting Text in Natural Image with Connectionist Text Proposal Network, text detection based on a connection network) can detect horizontal or slightly inclined text lines, and a text line is regarded as a character sequence. The front end of the CTPN text detector (which can also be called the CTPN model) can use VGGNet-16 (VGGNet is a deep convolutional neural network jointly developed by the Computer Vision Group of the University of Oxford and Google, and VGGNe-16 is a deep convolutional neural network with 16 layers) as the basic network to extract local image features of each character. The middle uses a BLSTM (Bi-directional Long Short-Term Memory) layer to extract character sequence context features, and then through a fully connected layer, the coordinate values of each text block are output at the end, and adjacent small text blocks are merged into text lines.

[0037] Step 102: According to the position information of the multiple text lines, initially classify the text lines of the current video to obtain a first set and a second set, where the text lines in the first set are identified as dialogue texts, and the text lines in the second set are identified as non-dialogue texts.

[0038] In this embodiment, the detected text lines are divided into two categories, one is dialogue text, and the other is non-dialogue text. Dialogue texts may include the dialogues between the host and the guests, the lyrics of the sung songs, and name bars. Among them, the name bars are used to indicate the names of the characters who speak the current dialogue or sing the current song.

[0039] In this step, for each text line, according to the position information of this text line, it can be determined whether this text line is located in the dialogue area. If this text line is located in the dialogue area, it can be determined that this text line is a dialogue text. If this text line is located in other areas outside the dialogue area, it can be determined that this text line is a non-dialogue text. Among them, the dialogue area can be determined according to the position information of all text lines in the video sequence.

[0040] Step 103: According to the font feature information corresponding to the multiple text lines and a pre-constructed clustering network, cluster the text lines of the current video to obtain multiple clustering results, where the text lines in the same clustering result have the same font.

[0041] Clustering is a process of dividing a set of physical or abstract objects into multiple classes composed of similar objects. The clusters generated by clustering are a set of data objects, and these objects are similar to each other within the same cluster and different from the objects in other clusters. In this step, through the pre-constructed clustering network and the font feature information of the text lines, the text lines of the current video are clustered to cluster the text lines with the same font into one clustering result and the text lines with different fonts into different clustering results. Among them, the font feature information is used to characterize the characteristics of a font that are different from other fonts. The font feature information corresponding to the text lines can be extracted by a pre-constructed font recognition model, and the pre-constructed font recognition model can be a convolutional neural network model. The pre-constructed clustering model can also be a neural network model. The pre-constructed clustering model can be a centerless clustering model. As an example, the pre-constructed clustering model can be trained through a GCN network. Among them, GCN (Graph Convolutional networks) is a neural network framework for machine learning on graphs.

[0042] Step 104: According to the multiple clustering results, adjust the text lines in the first set and the second set to perform a secondary classification on the text lines of the current video and determine the final dialogue text of the current video.

[0043] In this step, based on the clustering results, the classification results of text lines obtained according to the text line position information are filtered and recalled. The mis-identified non-line text in the first set is filtered out, and the missed line text in the second set is recalled, realizing the secondary classification of the text lines in the current video frame, effectively filtering out the mis-identified non-line text in the initial classification results, and recalling the missed line text, thereby improving the accuracy of line text recognition.

[0044] The text processing method of the embodiment of the present invention performs text detection on the video frame sequence of the current video to determine multiple text lines in the video frame sequence. According to the position information of the multiple text lines, the text lines of the current video are initially classified to preliminarily determine whether the text line is a line text line, thereby dividing the text lines of the current video into two sets. Then, the font feature information of each text line is obtained, and all text lines are clustered according to the font feature information to obtain multiple clustering results, where the text lines in each clustering result have the same font. Then, according to the obtained clustering results, the two obtained sets are filtered again, filtering out the mis-identified non-line text lines in the line text line set, and recalling the missed line text lines in the non-line text line set, thereby improving the accuracy of the line recognition result.

[0045] In an optional embodiment, the pre-constructed clustering network in step 103 can be obtained through the following process:

[0046] First, a training sample set is obtained. The training samples in the training sample set include real line images and simulated line images. Among them, the real line images in the training sample set can be obtained by taking screenshots from variety shows. The text lines in the real line images can cover lines in various fonts. The simulated line images can be obtained by pasting texts in different fonts on pictures.

[0047] Secondly, according to the pre-constructed font feature extraction model, the font feature information of each training sample in the training sample set is obtained. Among them, the pre-constructed font recognition model can be a convolutional neural network model, such as Figure 2 shown, the extracted font feature information is 2048-dimensional float data. Float refers to the floating-point data type.

[0048] Finally, according to the font feature information of each training sample, the clustering network is trained. Among them, the font feature information of the training samples can be trained through a preset cross-entropy loss function and a preset GCN network to obtain the clustering network.

[0049] The real line image and the simulated line image of this embodiment can cover a relatively large number of fonts. By using the real line image and the simulated line image as training samples for training, the pre-constructed clustering network obtained can accurately cluster text lines, making the fonts within the same clustering result the same and ensuring the purity within the clustering result.

[0050] In an alternative embodiment, the process of adjusting the text lines in the first set and the second set according to the multiple clustering results includes:

[0051] For each clustering result, determine the identifier of the clustering result according to the first set and the second set, where the identifier of the clustering result includes a line identifier and a non-line identifier;

[0052] Adjust the text lines in the first set and the second set according to the identifier of the clustering result.

[0053] Since the text lines with the same font shown in the video are clustered into the same clustering result, and there will be no text lines with multiple fonts mixed in this clustering result, the purity within the clustering result is relatively high. Therefore, if more text lines in this category are recognized as line text, that is, more text lines in this category exist in the first set, then it can be determined that the identifier of this clustering result is a line identifier; if more text lines in this category are recognized as non-line text, that is, more text lines in this category exist in the second set, then it can be determined that the identifier of this clustering result is a non-line identifier. Therefore, in this embodiment, the identifier of the clustering result can be determined by the proportion of line text in the clustering result.

[0054] On the other hand, if more text lines in the clustering result are recognized as line text, then other text lines in this category (that is, the text lines that do not exist in the first set but exist in the second set) may be misrecognized as non-line text; if more text lines in the clustering result are recognized as non-line text, then other text lines in this category (that is, the text lines that do not exist in the second set but exist in the first set) may be misrecognized as line text. Therefore, in the embodiments of the present invention, the text lines in the first set and the second set can be classified again according to the identifier of the clustering result, that is, through the clustering result, the classification result of the text lines obtained according to the text line position information is filtered and recalled, the misrecognized non-line text in the first set is filtered, and the missed line text lines in the second set are recalled, so as to improve the accuracy of line text recognition.

[0055] In an alternative embodiment, for each clustering result, the process of determining the identifier of the clustering result according to the first set and the second set includes: for each clustering result, determining a first ratio of the text lines belonging to the first set in the clustering result; in the case where the first ratio is greater than a preset first threshold, determining the identifier of the clustering result as a line-of-dialogue identifier; in the case where the first ratio is less than or equal to the preset first threshold, determining the identifier of the clustering result as a non-line-of-dialogue identifier. Wherein, the preset first threshold can be flexibly set according to the application scenario, and the present invention does not limit this here. Preferably, compared with filtering out mis-identified non-line-of-dialogue texts, recalling missed line-of-dialogue texts is more important. Therefore, the first preset threshold can be set to be relatively small. For example, the preset first threshold can be set to 3%, 4%, etc.

[0056] In an alternative embodiment, according to multiple clustering results, the process of adjusting the text lines in the first set and the second set to perform secondary classification on the text lines of the current video includes:

[0057] For a clustering result with a line-of-dialogue identifier, determining first to-be-identified text lines belonging to the second set in the clustering result, and performing secondary classification on the first to-be-identified text lines according to the position information of the first to-be-identified text lines to determine whether the first to-be-identified text lines are line-of-dialogue texts; if so, migrating the first to-be-identified text lines from the second set to the first set;

[0058] For a clustering result with a non-line-of-dialogue identifier, determining second to-be-identified text lines belonging to the first set in the clustering result; determining that the second to-be-identified text lines are non-line-of-dialogue texts, and migrating the second to-be-identified text lines from the second set to the first set.

[0059] As an example, assume that the first threshold is 4%. If the first ratio of the text lines belonging to the first set in a certain clustering result A is 60%, then the identifier of the clustering result is determined as a line-of-dialogue identifier. 40% of the text lines in the clustering result A do not belong to the first set but belong to the second set. The text lines belonging to the second set in the clustering result A are the first to-be-identified text lines. For the first to-be-identified text lines, secondary classification can be performed according to their position information. When performing secondary classification on them, the classification conditions can be relaxed.

[0060] If the first proportion of the text lines belonging to the first set in a clustering result B is 2%, then determine that the identifier of this clustering result is a non-line identifier. 98% of the text lines in this clustering result B do not belong to the second set but belong to the first set. The text lines belonging to the first set in this clustering result B are the second text lines to be recognized. For the second text lines to be recognized, determine them as non-line text and migrate the second text lines to be recognized from the second set to the first set.

[0061] In this embodiment, through the clustering results, the existing classification results of text lines obtained based on the position information of text lines are filtered and recalled. The mis-recognized non-line text in the first set is filtered, and the missed-recognized line text in the second set is recalled, realizing the secondary classification of the text lines in the current video frame, effectively filtering out the mis-recognized non-line text in the initial classification results, and recalling the missed-recognized line text, thereby improving the accuracy of line text recognition.

[0062] In an alternative embodiment, as Figure 3 shown, the initial classification of the text lines of the current video according to the position information of the multiple text lines includes:

[0063] Step 301: According to the position information of the multiple text lines, determine the frequency of text lines appearing at each pixel point on the video frame of the current video.

[0064] Step 302: Determine the area composed of pixel points with the frequency of text lines appearing greater than a preset second threshold as the line area.

[0065] Among them, the second threshold can be flexibly set according to the application scenario, and the present invention does not limit it here. After determining the line area, the position information of the line area can be determined. The position information of the line area can include the coordinates of the four endpoints of the line area: (x1, y1), (x1, y2), (x2, y1), and (x2, y2). According to x1 and x2, the width information of the line area can be determined, and according to y1 and y2, the height information of the line area can be determined.

[0066] Step 303: According to the position information of the text line, determine the width information and height information of the text area corresponding to the text line. Among them, the position information of the text line can include the coordinates of the four endpoints of the text line: (x3, y3), (x3, y4), (x4, y3), and (x4, y4). According to x3 and x4, the width information of the text area can be determined, and according to y3 and y4, the height information of the text area can be determined.

[0067] Step 304: Calculate the area intersection-over-union (IoU) of the text region and the dialogue region according to the width information and height information of the text region corresponding to the text line. Here, the intersection-over-union (IoU) is the overlap rate of the generated candidate box and the original marked box (ground truth bound), that is, the ratio of their intersection to their union. In this step, the area intersection-over-union is the ratio of the overlapping region of the text region and the dialogue region to the combined region of the text region and the dialogue region.

[0068] Step 305: Calculate the height intersection-over-union of the text region and the dialogue region according to the height information of the text region corresponding to the text line. The height intersection-over-union is the ratio of the intersection to the union between the height information [y3, y4] of the text region and the height information [y1, y2] of the dialogue region. For example, if y3 = 100, y4 = 200, y1 = 110, and y2 = 210, then the height intersection-over-union = 90 / 110 = 9 / 11.

[0069] Step 306: Calculate the width intersection-over-union of the text region and the dialogue region according to the width information of the text region corresponding to the text line. The width intersection-over-union is the ratio of the intersection to the union between the width information [x3, x4] of the text region and the width information [x1, x2] of the dialogue region. For example, if x3 = 90, x4 = 200, x1 = 110, and x2 = 210, then the width intersection-over-union = 90 / 120 = 9 / 12.

[0070] Step 307: When the area intersection-over-union of the text line is greater than a preset third threshold, the height intersection-over-union is greater than a preset fourth threshold, and the width intersection-over-union is greater than a preset fifth threshold, determine that the text line is dialogue text; or, when the area intersection-over-union of the text line is greater than a preset third threshold, the height intersection-over-union is greater than a preset fourth threshold, and the width intersection-over-union is not greater than a preset fifth threshold but the text line falls within the range of the dialogue region in the width direction, determine that the text line is dialogue text.

[0071] In this embodiment, when determining whether a certain text line is dialogue text, it is necessary to separately determine whether the area, height, and width of the text line all meet the requirements. Only when the area, height, and width of the text line all meet the requirements can it be determined that the text line is dialogue text. The requirement for the area of the text line is that the area intersection-over-union of the text line and the dialogue region is greater than the third threshold. The requirement for the height of the text line is that the height intersection-over-union of the text line and the dialogue region is greater than the fourth threshold. For the width of the text line, the judgment is divided into 3 cases:

[0072] (1) The width intersection-over-union of the text line and the dialogue region is greater than the fifth threshold;

[0073] (2) The width intersection ratio of the text line and the line area is not greater than the fifth threshold, and the text line falls within the text area in the width direction;

[0074] (3) The width intersection ratio of the text line and the line area is not greater than the fifth threshold, and the text line does not fall within the text area in the width direction.

[0075] For the above cases (1) and (2), it can be determined that the width of the text line meets the requirements. For case (3), it can be determined that the requirements are not met. When the area intersection ratio and the height intersection ratio are greater than the corresponding thresholds, for the above cases (1) and (2), it can be determined that the text line is the line text.

[0076] Among them, the third threshold, the fourth threshold, and the fifth threshold can be flexibly set according to the application scenario, and the present invention does not limit this. Exemplarily, the third threshold is 0.9, the fourth threshold is 0.9, and the fifth threshold is 0.2. The situation where the width intersection ratio is not greater than the preset fifth threshold but the text line falls within the line area in the width direction means that the text area is shorter in the width direction. Although it falls within the line area, its width intersection ratio with the line area is not greater than the preset fifth threshold. For example, the width information of the text area is: x3 = 90, x4 = 110. The width information of the line area is: x1 = 80, x2 = 210, then the width intersection ratio = 20 / 130 = 0.15. Although the width intersection ratio of 0.15 between this text area and the line area is less than the preset fifth threshold of 0.2, this text area completely falls within this line area. Therefore, if the area intersection ratio and the height intersection ratio of this text area meet the requirements, it can be determined that this text line is the line text.

[0077] In an alternative embodiment, when performing secondary classification on the first text line to be recognized, the judgment conditions can be relaxed according to the position information of the first text line to be recognized. For example, without changing the requirement for judging the width of the text line, the requirements for judging the area and height of the text line can be reduced. Exemplarily, the requirements for judging the area and height of the text line can be reduced by decreasing the corresponding thresholds, that is, updating the preset third threshold to the preset sixth threshold and updating the preset fourth threshold to the preset seventh threshold, where the preset sixth threshold is less than the preset third threshold, and the preset seventh threshold is less than the preset fourth threshold. Then, for the first text line to be recognized, when the intersection over union of the area of the first text line to be recognized is greater than the preset sixth threshold, the intersection over union of the height is greater than the preset seventh threshold, and the intersection over union of the width is greater than the preset fifth threshold, it is determined that the text line is a line of dialogue text; or when the intersection over union of the area of the first text line to be recognized is greater than the preset sixth threshold, the intersection over union of the height is greater than the preset seventh threshold, and the intersection over union of the width is not greater than the preset fifth threshold but the text line falls within the range of the dialogue area in the width direction, it is determined that the text line is a line of dialogue text.

[0078] As an example, assume that the third threshold is 0.9, then the sixth threshold is a value less than 0.9, such as 0.85, 0.8, 0.7, etc. Assume that the fourth threshold is 0.9, and the seventh threshold can be a value less than 0.9, such as 0.85, 0.8, 0.7, etc. When performing secondary judgment on the first text line to be recognized, the third threshold and the fourth threshold are decreased, thereby relaxing the judgment conditions to recall the text lines of dialogue text that have been misrecognized as non-dialogue text, that is, migrating the text lines misrecognized as non-dialogue text in the second set to the first set.

[0079] In this embodiment, secondary classification is performed on the first text line to be recognized in the clustering result marked with the dialogue label. When performing secondary classification on it, the judgment conditions can be relaxed, so that the text lines of dialogue text that have been misrecognized or missed can be effectively recalled.

[0080] The following takes the current video as a variety show video as an example to illustrate the process of the text processing method of the embodiment of the present invention:

[0081] First, each video frame in the video frame sequence of the variety show video is detected for text by a CTPN text detector, all text lines in the variety show video are detected, and the position information of all text lines is determined.

[0082] Secondly, based on the position information of all text lines, determine the position information of the dialogue area. For each text line, according to the position information of this text line and the position information of the dialogue area, determine whether this text line is dialogue text. If it is determined that this text line is dialogue text, write this text line into the first set. If it is determined that this text line is non-dialogue text, write this text line into the second set. Thus, the initial classification of all text lines of this variety show video is achieved.

[0083] Then, extract the font feature information of each text line according to the pre-constructed font recognition model. Among them, the extracted font feature information is 2048-dimensional float data. Furthermore, according to the font feature information of all text lines and the pre-constructed clustering network, cluster all text lines to obtain multiple clustering results. Among them, the text lines within the same clustering result have the same font. Among them, some clustering results are as Figures 4 - 7 shown, Figure 4 shown, the text lines in the clustering result have the same font, and the text lines within this clustering result are all irrelevant advertisement information. Figure 5 shown, the text lines in the clustering result have the same font, and the text lines within this clustering result are all name bars. Figure 6 shown, the text lines in the clustering result have the same font, and the text lines within this clustering result are all variety show dialogues. Figure 7 shown, the text lines in the clustering result have the same font, and the text lines within this clustering result are all variety show dialogues. Because there are many text lines in a variety show video, the text lines with the same font may be clustered into multiple clustering results. For example, the text lines with the same font are respectively clustered into Figure 6 and Figure 7 shown clustering results.

[0084] Finally, according to multiple clustering results, adjust the text lines in the first set and the second set to perform secondary classification on the text lines of the current video, filter out the mis-identified non-dialogue text in the first set, and recall the missed-identified dialogue text in the second set, so as to determine the final dialogue text of this variety show video. As Figure 8 and Figure 9 shown, recall the missed-identified dialogue text "Okay", "After several years", "What should I do" in the second set, and filter out the non-dialogue text "Guiding unit:", "Li Moumou:" in the first set.

[0085] In an embodiment of the present invention, by clustering all text lines according to the font feature information of the text lines appearing in the variety show video, the text lines with the same font are clustered into the same clustering result, so that the text lines belonging to the same identifier or the same category are clustered into the same clustering result, and then the clustering result is merged with the classification result obtained according to the text line position information, filtering out the mis-identified non-line text in the first set, and recalling the missed-identified line text in the second set, thereby improving the accuracy of line text recognition.

[0086] Figure 10 FIG. schematically shows a structural diagram of a text processing device 1000 according to an embodiment of the present invention, as Figure 10 shown, the text processing device 1000 includes:

[0087] A text detection module 1001, configured to perform text detection on a video frame sequence of a current video to determine a plurality of text lines in the video frame sequence;

[0088] An identification module 1002, configured to perform an initial classification on the text lines of the current video according to the position information of the plurality of text lines, to obtain a first set and a second set, wherein the text lines in the first set are identified as line text, and the text lines in the second set are identified as non-line text;

[0089] A clustering module 1003, configured to cluster the text lines of the current video according to the font feature information corresponding to the plurality of text lines and a pre-constructed clustering network, to obtain a plurality of clustering results, and the text lines in the same clustering result have the same font;

[0090] An update module 1004, configured to adjust the text lines in the first set and the second set according to the plurality of clustering results, to perform a secondary classification on the text lines of the current video, and determine the final line text of the current video.

[0091] Optionally, the device further includes a training module, configured to: obtain a training sample set, where the training samples in the training sample set include real line images and simulated line images; obtain the font feature information of each training sample in the training sample set according to a pre-constructed font feature extraction model; and train the clustering network according to the font feature information of each training sample.

[0092] Optionally, the update module 1004 is further configured to: for each clustering result, determine the identifier of the clustering result according to the first set and the second set; wherein the identifier of the clustering result includes a line identifier and a non-line identifier; and adjust the text lines in the first set and the second set according to the identifier of the clustering result.

[0093] Optionally, the update module 1004 is further configured to: for each clustering result, determine a first ratio of the text lines belonging to the first set in the clustering result; in the case where the first ratio is greater than a preset first threshold, determine that the identifier of the clustering result is a line identifier; in the case where the first ratio is less than or equal to the preset first threshold, determine that the identifier of the clustering result is a non-line identifier.

[0094] Optionally, the update module 1004 is further configured to: for the clustering result with the line identifier, determine a first text line to be recognized belonging to the second set in the clustering result, and perform secondary classification on the first text line to be recognized according to the position information of the first text line to be recognized, so as to determine whether the first text line to be recognized is a line text; if so, migrate the first text line to be recognized from the second set to the first set; for the clustering result with the non-line identifier, determine a second text line to be recognized belonging to the first set in the clustering result; determine that the second text line to be recognized is a non-line text, and migrate the second text line to be recognized from the second set to the first set.

[0095] Optionally, the recognition module 1002 is further configured to: determine a line region composed of pixel points with the frequency of the text line appearance greater than a preset second threshold; determine the width information and height information of the text region corresponding to the text line according to the position information of the text line; calculate the area intersection-over-union ratio of the text region and the line region according to the width information and height information of the text region corresponding to the text line; calculate the height intersection-over-union ratio of the text region and the line region according to the height information of the text region corresponding to the text line; calculate the width intersection-over-union ratio of the text region and the line region according to the width information of the text region corresponding to the text line; in the case where the area intersection-over-union ratio of the text line is greater than a preset third threshold, the height intersection-over-union ratio is greater than a preset fourth threshold, and the width intersection-over-union ratio is greater than a preset fifth threshold, determine that the text line is a line text; or in the case where the area intersection-over-union ratio of the text line is greater than a preset third threshold, the height intersection-over-union ratio is greater than a preset fourth threshold, the width intersection-over-union ratio is not greater than a preset fifth threshold but the text line falls within the range of the line region in the width direction, determine that the text line is a line text.

[0096] Optionally, the updating module 1004 is further configured to: update the preset third threshold to a preset sixth threshold and update the preset fourth threshold to a preset seventh threshold, where the preset sixth threshold is less than the preset third threshold, and the preset seventh threshold is less than the preset fourth threshold; determine that the text line is a line of dialogue text when the intersection over union of the area of the text line is greater than the preset sixth threshold, the intersection over union of the height is greater than the preset seventh threshold, and the intersection over union of the width is greater than the preset fifth threshold; and determine that the text line is a line of dialogue text when the intersection over union of the area of the text line is greater than the preset sixth threshold, the intersection over union of the height is greater than the preset seventh threshold, the intersection over union of the width is not greater than the preset fifth threshold, but the text line falls within the range of the dialogue area in the width direction.

[0097] The text processing device according to the embodiment of the present invention determines a plurality of text lines in a video frame sequence of a current video by performing text detection on the video frame sequence of the current video, and performs initial classification on the text lines of the current video according to the position information of the plurality of text lines to preliminarily determine whether a text line is a line of dialogue text line, so as to divide the text lines of the current video into two sets. Then, font feature information of each text line is obtained, and all text lines are clustered according to the font feature information to obtain a plurality of clustering results, where the text lines in each clustering result have the same font. Then, according to the obtained clustering results, the two obtained sets are filtered twice, filtering out non-dialogue text lines mis-identified in the set of dialogue text lines and recalling missed-identified dialogue text lines in the set of non-dialogue text lines, thereby improving the accuracy of the dialogue recognition result. For text translation, the accuracy of the dialogue translation result can be effectively improved, and the cost is effectively reduced.

[0098] The above device can execute the method provided by the embodiment of the present invention and has corresponding functional modules and beneficial effects for executing the method. For technical details not described in detail in this embodiment, reference can be made to the method provided by the embodiment of the present invention.

[0099] Figure 11 The structural diagram of the electronic device according to the embodiment of the present invention is schematically shown. As Figure 11 shown, the electronic device includes: a processor 1101, a communication interface 1102, a memory 1103, and a communication bus 1104. Among them, the processor 1101, the communication interface 1102, and the memory 1103 communicate with each other through the communication bus 1104.

[0100] The memory 1103 is used to store a computer program.

[0101] When the processor 1101 executes the program stored in the memory 1103, the following steps are implemented:

[0102] Perform text detection on the video frame sequence of the current video to determine multiple text lines in the video frame sequence; according to the position information of the multiple text lines, perform initial classification on the text lines of the current video to obtain a first set and a second set, where the text lines in the first set are identified as dialogue text, and the text lines in the second set are identified as non-dialogue text; according to the font feature information corresponding to the multiple text lines and a pre-constructed clustering network, perform clustering on the text lines of the current video to obtain multiple clustering results, and the text lines in the same clustering result have the same font; according to the multiple clustering results, adjust the text lines in the first set and the second set to perform secondary classification on the text lines of the current video to determine the final dialogue text of the current video.

[0103] The communication bus mentioned in the above terminal may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience of representation, only a thick line is used in the figure, but it does not mean that there is only one bus or one type of bus.

[0104] The communication interface is used for communication between the above terminal and other devices.

[0105] The memory may include a Random Access Memory (RAM), and may also include a non-volatile memory, such as at least one disk memory. Optionally, the memory may also be at least one storage device located far from the aforementioned processor.

[0106] The above-mentioned processor may be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it may also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.

[0107] In another embodiment provided by the present invention, there is also provided a computer-readable storage medium storing instructions which, when run on a computer, cause the computer to execute the text processing method described in any one of the above embodiments.

[0108] In another embodiment provided by the present invention, there is also provided a computer program product containing instructions which, when run on a computer, cause the computer to execute the text processing method described in any one of the above embodiments.

[0109] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present invention are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from a website, computer, server, or data center to another website, computer, server, or data center by wire (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that includes one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid state disk (SSD)).

[0110] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including", or any other variation thereof is intended to cover non-exclusive inclusion, so that a process, method, article, or device including a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article, or device including the element.

[0111] Each embodiment in this specification is described in a related manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and for the related parts, reference can be made to the partial description of the method embodiment.

[0112] The above are only the preferred embodiments of the present invention, and are not intended to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention are all included in the protection scope of the present invention.

Claims

1. A text processing method, characterized in that, it includes: performing text detection on the video frame sequence of the current video to determine multiple text lines in the video frame sequence; performing initial classification on the text lines of the current video according to the position information of the multiple text lines to obtain a first set and a second set, wherein the text lines in the first set are identified as dialogue text, and the text lines in the second set are identified as non-dialogue text; performing clustering on the text lines of the current video according to the font feature information corresponding to the multiple text lines and a pre-constructed clustering network to obtain multiple clustering results, and the text lines in the same clustering result have the same font; adjusting the text lines in the first set and the second set according to the multiple clustering results to perform secondary classification on the text lines of the current video and determine the final dialogue text of the current video.

2. The method according to claim 1, characterized in that, the pre-constructed clustering network is obtained according to the following process: acquiring a training sample set, where the training samples in the training sample set include real dialogue images and simulated dialogue images; acquiring the font feature information of each training sample in the training sample set according to a pre-constructed font feature extraction model; training the clustering network according to the font feature information of each training sample.

3. The method according to claim 1, characterized in that, adjusting the text lines in the first set and the second set according to the multiple clustering results includes: for each clustering result, determining the identifier of the clustering result according to the first set and the second set; wherein the identifier of the clustering result includes a dialogue identifier and a non-dialogue identifier; adjusting the text lines in the first set and the second set according to the identifier of the clustering result.

4. The method according to claim 3, characterized in that, determining the identifier of the clustering result according to the first set and the second set includes: for each clustering result, determining the first proportion of the text lines belonging to the first set in the clustering result; when the first proportion is greater than a preset first threshold, determining the identifier of the clustering result as a dialogue identifier; when the first proportion is less than or equal to the preset first threshold, determining the identifier of the clustering result as a non-dialogue identifier.

5. The method according to claim 3 or 4, characterized in that, adjusting the text lines in the first set and the second set according to the multiple clustering results to perform secondary classification on the text lines of the current video includes: for the clustering result with a dialogue identifier, determining the first text line to be recognized belonging to the second set in the clustering result, and performing secondary classification on the first text line to be recognized according to the position information of the first text line to determine whether the first text line to be recognized is dialogue text; if so, migrating the first text line to be recognized from the second set to the first set; For the clustering results labeled as non-line-of-dialogue labels, determine the second text lines to be recognized in the clustering results that belong to the first set; determine that the second text lines to be recognized are non-line-of-dialogue texts, and migrate the second text lines to the first set from the second set.

6. The method according to claim 5, wherein, the initial classification of the text lines of the current video according to the position information of the multiple text lines includes: determining the frequency of occurrence of text lines at each pixel point on the video frame of the current video according to the position information of the multiple text lines; determining the area where the pixel points with the frequency of occurrence of text lines greater than a preset second threshold form as the line-of-dialogue area; determining the width information and height information of the text area corresponding to the text line according to the position information of the text line; calculating the area intersection-over-union ratio of the text area and the line-of-dialogue area according to the width information and height information of the text area corresponding to the text line; calculating the height intersection-over-union ratio of the text area and the line-of-dialogue area according to the height information of the text area corresponding to the text line; calculating the width intersection-over-union ratio of the text area and the line-of-dialogue area according to the width information of the text area corresponding to the text line; when the area intersection-over-union ratio of the text line is greater than a preset third threshold, the height intersection-over-union ratio is greater than a preset fourth threshold, and the width intersection-over-union ratio is greater than a preset fifth threshold, determining that the text line is a line-of-dialogue text; when the area intersection-over-union ratio of the text line is greater than a preset third threshold, the height intersection-over-union ratio is greater than a preset fourth threshold, the width intersection-over-union ratio is not greater than a preset fifth threshold but the text line falls within the range of the line-of-dialogue area in the width direction, determining that the text line is a line-of-dialogue text.

7. The method according to claim 6, wherein, the secondary classification of the text line to be recognized according to the position information of the first text line to be recognized includes: updating the preset third threshold to a preset sixth threshold and updating the preset fourth threshold to a preset seventh threshold, where the preset sixth threshold is less than the preset third threshold and the preset seventh threshold is less than the preset fourth threshold; when the area intersection-over-union ratio of the text line is greater than a preset sixth threshold, the height intersection-over-union ratio is greater than a preset seventh threshold, and the width intersection-over-union ratio is greater than a preset fifth threshold, determining that the text line is a line-of-dialogue text; when the area intersection-over-union ratio of the text line is greater than a preset sixth threshold, the height intersection-over-union ratio is greater than a preset seventh threshold, the width intersection-over-union ratio is not greater than a preset fifth threshold but the text line falls within the range of the line-of-dialogue area in the width direction, determining that the text line is a line-of-dialogue text.

8. A text processing device, wherein, comprising: a text detection module, configured to perform text detection on a video frame sequence of a current video to determine multiple text lines in the video frame sequence; An identification module, configured to perform an initial classification on the text lines of the current video according to the position information of the multiple text lines, to obtain a first set and a second set, wherein the text lines in the first set are identified as dialogue texts, and the text lines in the second set are identified as non-dialogue texts; A clustering module, configured to perform clustering on the text lines of the current video according to the font feature information corresponding to the multiple text lines and a pre-constructed clustering network, to obtain multiple clustering results, and the text lines in the same clustering result have the same font; An updating module, configured to adjust the text lines in the first set and the second set according to the multiple clustering results, so as to perform a secondary classification on the text lines of the current video, and determine the final dialogue texts of the current video.

9. An electronic device, characterized in that, it includes a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other through the communication bus; the memory is used for storing a computer program; the processor is configured to implement the method steps of any one of claims 1-7 when executing the program stored on the memory.

10. A computer-readable storage medium, on which a computer program is stored, characterized in that, when the program is executed by a processor, it implements the method as described in any one of claims 1-7.

Citation Information

Patent Citations

  • Method for training video text classification model and video text classification method and device

    CN112036373A

  • Method and device for extracting video subtitles, electronic equipment and storage medium

    CN112925905A