Method and device for video subtitle recognition, electronic device, storage medium

By generating a set of text boxes in video frames and calculating the mean square error of their widths, the problem of distinguishing subtitles from other text in video subtitle recognition is solved, achieving efficient and accurate subtitle recognition.

CN114581900BActive Publication Date: 2026-04-21BEIJING XUEZHITU NETWORK TECH
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
BEIJING XUEZHITU NETWORK TECH
Filing Date
2022-03-09
Publication Date
2026-04-21

AI Technical Summary

Technical Problem

Existing technologies struggle to effectively distinguish subtitle text from other text when recognizing video subtitles, leading to misrecognition, especially when the position, type, and font settings of subtitles vary across different video content, resulting in low recognition accuracy.

Method used

By performing text recognition on video frames, generating a set based on the frequency of occurrence of text box heights, and calculating the mean square error of text box widths, subtitles can be recognized.

Benefits of technology

It improves the accuracy of video subtitle recognition, reduces computational complexity, speeds up computation, and effectively eliminates interference information in the video.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114581900B_ABST
    Figure CN114581900B_ABST
Patent Text Reader

Abstract

This application relates to the field of video processing technology and discloses a method for video subtitle recognition, comprising: performing text recognition on multiple video frames to obtain all text boxes in each video frame; determining multiple text box sets based on the frequency of occurrence of text box heights; calculating the root mean square error of the text box width for each text box set; and determining subtitles based on the root mean square error of the width of each text box. Since the text boxes in the video have the same height, but the text box widths differ in different video frames, the text boxes of video subtitles can be accurately identified by their height and width. When recognizing different categories of video subtitles, it is not necessary to consider information such as video size and subtitle position, nor is manual annotation for subtitle classification required. This application also discloses an apparatus, electronic device, and storage medium for video subtitle recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of video processing technology, such as a method and apparatus, electronic device, and storage medium for video subtitle recognition. Background Technology

[0002] Currently, video has become an increasingly popular medium for conveying information, and it can be categorized into long videos, short videos, and live videos. Regardless of the type, subtitles can be added to ensure the video content is clearly conveyed to the user, and to facilitate the extraction or recognition of subtitles for content comprehension and other purposes.

[0003] In related technologies, methods for identifying subtitles include: using specific location methods, such as when the video content is a movie clip or TV series clip, and the subtitles are placed below the video, with a fixed height for subtitle positioning; or using classification methods, such as when the video content contains text that is not subtitles, and a classification model is trained using a neural network to locate the subtitles by labeling a large number of common subtitles and non-subtitles. Related technologies also provide a video subtitle positioning method, which includes: acquiring all image frames of the video and performing text detection on all image frames to obtain a first set of text boxes for all image frames; traversing the first set of text boxes and obtaining a first similarity in a first direction for text boxes between every two image frames; constructing a first graph network about the first set of text boxes based on multiple first similarities; clustering the first graph network to obtain multiple first subnetworks, and extracting the text boxes of the video subtitles from the first subnetworks whose number of nodes meets a first preset condition.

[0004] In the process of implementing the embodiments of this disclosure, at least the following problems were found in the related art:

[0005] Because different videos contain different content, the position, type, and font of subtitles vary from video to video. Furthermore, some videos may include not only subtitle text but also background information, bullet comments, and special effects text. Therefore, under the influence of numerous factors, subtitle recognition cannot effectively distinguish subtitle text from other text, easily leading to misidentification. Summary of the Invention

[0006] To provide a basic understanding of some aspects of the disclosed embodiments, a brief summary is given below. This summary is not intended as a general commentary, nor is it intended to identify key / important components or describe the scope of protection of these embodiments, but rather as a prelude to the detailed description that follows.

[0007] This disclosure provides a method, apparatus, electronic device, and storage medium for video subtitle recognition, which can improve the accuracy of subtitle recognition for different videos.

[0008] In some embodiments, the video caption recognition method includes: performing text recognition on multiple video frames to obtain all text boxes in each video frame; wherein each text box includes a line of text; determining multiple text box sets based on the frequency of occurrence of the text box height; wherein the text boxes in each text box set have the same height; calculating the root mean square error of the text box width of each text box set; and determining captions based on the root mean square error of the width of each text box.

[0009] In some embodiments, the apparatus for video subtitle recognition includes: a text recognition module configured to perform text recognition on multiple video frames to obtain all text boxes in each video frame; wherein each text box includes a line of text; a determination module configured to determine multiple sets of text boxes based on the frequency of occurrence of the text box height; wherein the text boxes in each set of text boxes have the same height; a calculation module configured to calculate the root mean square error of the text box width for each set of text boxes; and a subtitle recognition module to determine subtitles based on the root mean square error of the width of each text box.

[0010] In some embodiments, the apparatus for video caption recognition includes a processor and a memory storing program instructions, the processor being configured to execute the method for video caption recognition as described above when the program instructions are executed.

[0011] In some embodiments, the electronic device includes the means for video caption recognition as described above.

[0012] In some embodiments, the storage medium stores program instructions that, when executed, perform the method for video subtitle recognition as described above.

[0013] The method, apparatus, electronic device, and storage medium for video subtitle recognition provided in this disclosure can achieve the following technical effects: When obtaining text boxes in all video frames, since the text boxes in the video have the same height but different widths in different video frames, multiple sets of text boxes are obtained through the text box heights. By calculating the mean square error of the text box widths of each set of text boxes, the text boxes of video subtitles can be accurately identified. When recognizing different categories of video subtitles, it is not necessary to consider information such as video size and subtitle position, nor is it necessary to manually label and classify subtitles. The method for video subtitle localization in this disclosure is based on the field of deep learning technology. Through computer vision methods, it reduces computational complexity and speeds up computation, thereby improving the accuracy of recognizing different video subtitles and effectively eliminating other interference information in the video.

[0014] The above general description and the description below are exemplary and illustrative only and are not intended to limit this application. Attached Figure Description

[0015] One or more embodiments are illustrated by way of example with reference to the accompanying drawings. These illustrations and drawings do not constitute a limitation on the embodiments. Elements having the same reference numerals in the drawings are shown as similar elements. The drawings are not to be scaled. And wherein:

[0016] Figure 1 This is a schematic diagram of a method for video subtitle recognition provided in an embodiment of this disclosure;

[0017] Figure 2 This is a schematic diagram of another method for video caption recognition provided in an embodiment of this disclosure;

[0018] Figure 3 This is a schematic diagram of another method for video caption recognition provided in an embodiment of this disclosure;

[0019] Figure 4 This is a schematic diagram of another method for video caption recognition provided in an embodiment of this disclosure;

[0020] Figure 5 This is a schematic diagram of a device for video subtitle recognition provided in an embodiment of this disclosure;

[0021] Figure 6 This is a schematic diagram of another device for video caption recognition provided in an embodiment of this disclosure. Detailed Implementation

[0022] To provide a more detailed understanding of the features and technical content of the embodiments of this disclosure, the implementation of the embodiments of this disclosure will be described in detail below with reference to the accompanying drawings. The accompanying drawings are for illustrative purposes only and are not intended to limit the embodiments of this disclosure. In the following technical description, for ease of explanation, several details are used to provide a full understanding of the disclosed embodiments. However, one or more embodiments may still be implemented without these details. In other cases, well-known structures and devices may be simplified in their depiction to simplify the drawings.

[0023] The terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this disclosure are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate for the embodiments of this disclosure described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion.

[0024] Unless otherwise stated, the term "multiple" means two or more.

[0025] In this embodiment of the disclosure, the character " / " indicates that the objects before and after it are in an "or" relationship. For example, A / B means: A or B.

[0026] The term "and / or" describes an association between objects, indicating that three relationships can exist. For example, A and / or B means: A or B, or A and B.

[0027] The term "correspondence" can refer to an association or binding relationship. The correspondence between A and B means that there is an association or binding relationship between A and B.

[0028] Combination Figure 1 As shown, this disclosure provides a method for video subtitle recognition, including:

[0029] Step S101: The processor performs text recognition on multiple video frames to obtain all text boxes in each video frame; wherein, each text box includes a line of text;

[0030] Step S102: The processor determines multiple text box sets based on the frequency of occurrence of the text box height; wherein the text boxes in each text box set have the same height.

[0031] Step S103: The processor calculates the mean squared error of the text box width for each set of text boxes;

[0032] Step S104: The processor determines the subtitles based on the mean squared error of the width of each text box.

[0033] The method for video subtitle recognition provided in this disclosure, when acquiring text boxes in all video frames, since the text boxes in the video have the same height but different widths in different video frames, obtains multiple sets of text boxes based on their heights. By calculating the mean squared error of the text box widths in each set, the text boxes of the video subtitles can be accurately identified. When recognizing different categories of video subtitles, there is no need to consider information such as video size and subtitle position, nor is manual annotation for subtitle classification required. Thus, by using computer vision methods, computational complexity is reduced, computation speed is increased, and the accuracy of recognizing different video subtitles is improved, effectively eliminating other interfering information in the video.

[0034] The system performs text recognition on multiple video frames to obtain all text boxes in each video frame. Optionally, the video can be of any type, such as long video, short video, live video, TV series, movie, variety show, etc. The video includes subtitle text, such as bullet screen text, background text, advertising text, etc. If it is a short video, it may also contain other text information, such as watermark text, user nickname, and video name.

[0035] Here, the video is segmented into frames to obtain all the frames that make up the video. Each frame is then processed using OCR (Optical Character Recognition) to recognize text, thus generating text boxes containing all the aforementioned text. OCR acquires text image information from paper using optical input methods such as scanning and photography. Various pattern recognition algorithms analyze the morphological features of the text to convert invoices, newspapers, books, manuscripts, and other printed materials into image information. Finally, text recognition technology converts this image information into usable computer input.

[0036] In this embodiment, multiple text box sets are determined based on the frequency of occurrence of the text box height. Different texts possess different characteristics. For example, background text in a video will have essentially the same characteristics within the same video segment, provided the scene remains unchanged. However, for video subtitles, the height of the text may only change slightly with variations in the text itself, while the width of the entire line of text will change significantly. Furthermore, different types of subtitles have different heights and frequencies of occurrence. Therefore, the frequency of occurrence of the text box height allows for the classification of all text boxes.

[0037] In multiple text box sets, each text box set has multiple text boxes of the same height, that is, each text box set has multiple text boxes of the same height. This makes it easy to distinguish which text box of which height is the text box for the subtitle text.

[0038] In this embodiment, the mean squared deviation of the text box widths for each set of text boxes is calculated. The mean squared deviation generally refers to the standard deviation, which is the square root of the arithmetic mean (i.e., variance) of the squared deviations from the mean. The mean squared deviation reflects the dispersion of a dataset. Since the standard deviation is usually defined relative to the mean of the sample data, it indicates how far a particular data observation is from the mean. Therefore, the standard deviation is greatly affected by extreme values. In this application, it can be represented as the dispersion of the text box widths within the set of text boxes.

[0039] In this embodiment, since the width of the entire line of subtitle text varies significantly, it is possible to obtain the width of the text boxes in each text box set. That is, the smaller the standard deviation, the closer the widths of the text boxes in the text box set are; the larger the standard deviation, the more dispersed the widths of the text boxes in the text box set are. Therefore, the subtitle is determined based on the mean square deviation of the widths of each text box.

[0040] Optionally, the method for video subtitle recognition provided in this disclosure can also recognize text in multiple image frames or screen frames. The operation steps are the same as those described above, and will not be repeated here.

[0041] Combination Figure 2 As shown, this disclosure provides another method for video caption recognition, including:

[0042] Step S201: The processor performs text recognition on multiple video frames to obtain all text boxes in each video frame; wherein, each text box includes a line of text;

[0043] Step S202: The processor clusters text boxes with the same height into multiple text box sets; wherein the text boxes in each text box set have the same height.

[0044] Step S203: The processor sorts the text box sets in descending order according to the number of text boxes in each set;

[0045] Step S204: The processor selects the first N text boxes, where N is an integer greater than 1;

[0046] Step S205: The processor calculates the mean squared error of the text box width for each set of text boxes;

[0047] Step S206: The processor determines the subtitles based on the mean squared error of the width of each text box.

[0048] In this embodiment, to further improve the accuracy of recognizing different video subtitles, a sufficient number of video frames and a sufficient set of text boxes are required, and the text box set needs to contain a certain number of text boxes. This is necessary for more accurate recognition of video subtitles.

[0049] In this embodiment, since the text boxes in the video have the same height, multiple text box sets are generated by clustering text boxes with the same height. These text box sets are then sorted in descending order based on the number of text boxes in each set. This ensures that the text box set with the largest number of text boxes is at the front of the text box set queue, providing better data for calculating the mean squared error of the text box width, thereby improving the accuracy of subtitle recognition for different videos.

[0050] In this embodiment, when there is a lot of dialogue text in the video, the frequency of video frame cutting can be reduced appropriately; when there is a little dialogue text in the video, the frequency of video frame cutting can be increased appropriately.

[0051] In this embodiment, when there are many video frames, sorting the text box sets in descending order can also filter out some non-subtitle text. For example, if the video contains bullet comments, these bullet comments only exist in one video frame during frame cutting. Thus, the bullet comment text sets are at the back of the text box set queue. Selecting the preceding text box sets directly filters out the bullet comment text sets, effectively eliminating interference information in the video, reducing computational complexity, and speeding up the computation.

[0052] Optionally, in some embodiments, the text box sets can be sorted in ascending order according to the number of text boxes in each text box set; then N text box sets are selected, where N is an integer greater than 1.

[0053] To further calculate the mean squared error of the text box width for each set of text boxes, in some embodiments, the mean squared error is calculated according to the following formula:

[0054]

[0055] Where, σ I Let w be the mean squared error of the text box widths in the i-th text box set, where i = 1, ..., N; n is the number of text boxes in the text box set; i Let be the width of the i-th text box, where i = 1, ..., n; This represents the average width of the text boxes in the text box collection.

[0056] In some specific embodiments, the video is divided into 100 video frames, and text recognition is performed using OCR, resulting in 544 text boxes from the 100 video frames. The correspondence between the height of the text boxes and the text box height is as follows:

[0057] Height h 10mm 11mm 12mm 13mm 14mm 15mm 16mm Quantity n 102 82 73 53 46 100 88

[0058] In this embodiment, the text boxes are sorted by the number of text boxes, and the first 5 text box sets are selected, namely the 10mm text box set, the 15mm text box set, the 16mm text box set, the 11mm text box set, and the 12mm text box set.

[0059] First, calculate the average width of the text boxes corresponding to each set of text boxes. Calculate the average width of the 102 text boxes in the 10mm text box set, the average width of the 100 text boxes in the 15mm text box set, the average width of the 88 text boxes in the 16mm text box set, the average width of the 82 text boxes in the 11mm text box set, and the average width of the 73 text boxes in the 12mm text box set.

[0060] By calculating the mean squared error of the width of each text box, the average width of the text boxes, and the number of text boxes using the mean squared error formula, the mean squared error of the width of the text boxes in the five text box sets is finally obtained. The subtitles are then determined based on the mean squared error of the width of the text boxes. In this way, interference information in the video is effectively eliminated, which reduces the computational complexity and speeds up the calculation.

[0061] Combination Figure 3 As shown, this disclosure provides another method for video caption recognition, including:

[0062] Step S301: The processor performs text recognition on multiple video frames to obtain all text boxes in each video frame; wherein, each text box includes a line of text;

[0063] Step S302: The processor clusters text boxes with the same height into multiple text box sets; wherein the text boxes in each text box set have the same height.

[0064] Step S303: The processor sorts the text box sets in descending order according to the number of text boxes in each set;

[0065] Step S304: The processor selects the first N text boxes, where N is an integer greater than 1;

[0066] Step S305: The processor calculates the mean squared error of the text box width for each set of text boxes;

[0067] Step S306: The processor determines the maximum value among the mean squared deviations of the widths of each text box;

[0068] Step S307: The processor uses each text box in the set of text boxes corresponding to the maximum value as the caption.

[0069] In this embodiment, to further improve the accuracy of recognizing different video subtitles, the width of each text box, the average width of the text boxes, and the number of text boxes are calculated using the mean squared error formula. Finally, the mean squared error of the text box widths of multiple text box sets is obtained. The text boxes in the text box set corresponding to the maximum value of the mean squared error of the width of each text box can be directly compared and used as subtitles. Alternatively, the corresponding text box set can be selected by sorting the mean squared error values ​​of the widths of each text box, and the text boxes in that text box set can be used as subtitles.

[0070] In some embodiments, when determining multiple text box sets based on the frequency of height occurrence, there can be some error in detecting the height of the same text in different video frames. Furthermore, the height of the text also varies slightly with the text itself. Therefore, in some cases, it is necessary to calculate the average height of similar text boxes and determine the text box set based on this average height. For example, the average height of the text boxes... Will By clustering the text boxes into sets, the accuracy of recognizing subtitles for different videos can be improved.

[0071] In this embodiment, when there are many text boxes, the number of selected text boxes can be appropriately reduced to reduce computational complexity and computational load, and speed up the computation.

[0072] Optionally, in some embodiments, the mean squared error of the width of each text box can be sorted in ascending order according to the value of the mean squared error of the width of each text box; and the mean squared error of the width of the text box with the largest value can be selected.

[0073] Combination Figure 4 As shown, this disclosure provides another method for video caption recognition, including:

[0074] Step S401: Processor 100 performs text recognition on multiple consecutive video frames;

[0075] Step S402: The processor 100 identifies all text boxes in each video frame and obtains the height and width of each text box; wherein, each text box includes a line of text;

[0076] Step S403: The processor 100 clusters text boxes with the same height into multiple text box sets; wherein the text boxes in each text box set have the same height.

[0077] Step S404: The processor sorts the text box sets in descending order according to the number of text boxes in each set;

[0078] Step S405: The processor selects the first N text boxes, where N is an integer greater than 1;

[0079] Step S406: The processor calculates the mean squared error of the text box width for each set of text boxes;

[0080] Step S407: The processor determines the maximum value among the mean squared deviations of the widths of each text box;

[0081] Step S408: The processor uses each text box in the set of text boxes corresponding to the maximum value as the caption.

[0082] In this embodiment, text recognition is performed on multiple consecutive video frames to identify all text boxes in each video frame, and the height and width of each text box are obtained. Text recognition is performed on multiple consecutive video frames at preset time intervals. Optionally, the interval within the preset time period can be 1 second, 2 seconds, 3 seconds, 4 seconds, etc. Here, taking a preset time period of 2 minutes and an interval of 1 second as an example, the video is sliced ​​every second within 2 minutes, resulting in 120 video frames. All text boxes in the 120 video frames are identified, and the height and width of each text box are recorded. Text boxes with the same height are clustered into five sets of text boxes with different heights: h1 has 3 text boxes, h2 has 1 text box, h3 has 10 text boxes, h4 has 8 text boxes, and h5 has 15 text boxes. The text box sets are sorted in descending order according to the number of text boxes: h5 > h3 > h4 > h1 > h2. The text boxes of h5, h3, and h4 are selected, and the mean squared deviation of the text box widths of h5, h3, and h4 is calculated. The final ranking of the mean squared deviation of the text box widths is h5 > h3 > h4. That is, the text boxes in the h5 text box set are subtitles. Therefore, different subtitles can be classified by the height of the text box. At the same time, by calculating the mean square error of the width of each text box set, the text box of the video subtitle can be accurately identified, effectively eliminating other interfering information in the video and improving the accuracy of recognizing different video subtitles.

[0083] Combination Figure 5 As shown, this disclosure provides an apparatus for video subtitle recognition, including a text recognition module 501, a determination module 502, a calculation module 503, and a subtitle recognition module 504. The text recognition module 501 is configured to perform text recognition on multiple video frames to obtain all text boxes in each video frame; wherein each text box includes a line of text. The determination module 502 is configured to determine multiple sets of text boxes based on the frequency of occurrence of the text box height; wherein the text boxes in each set of text boxes have the same height. The calculation module 503 is configured to calculate the root mean square error of the text box width in each set of text boxes. The subtitle recognition module 504 determines the subtitles based on the root mean square error of the width of each text box.

[0084] The device for video subtitle recognition provided in this disclosure, when acquiring text boxes in all video frames, obtains multiple sets of text boxes based on their heights, since the text boxes in the video have the same height but different widths in different video frames. By calculating the mean square error of the text box widths in each set, the text boxes of the video subtitles can be accurately identified. When recognizing different categories of video subtitles, there is no need to consider information such as video size and subtitle position, nor is manual annotation for subtitle classification required. This reduces computational complexity and speeds up calculations, thereby improving the accuracy of recognizing different video subtitles and effectively eliminating other interfering information in the video.

[0085] Optionally, the determining module 502 further includes an aggregation unit, a sorting unit, and a selection unit. The aggregation unit clusters text boxes with the same height into multiple text box sets; the sorting unit is configured to sort the text box sets in descending order according to the number of text boxes in each set; and the selection unit is configured to select the first N text box sets, where N is an integer greater than 1.

[0086] Optionally, the text recognition module 501 is specifically configured to perform text recognition on multiple consecutive video frames; identify all text boxes in each video frame, and obtain the height and width of each text box.

[0087] Optionally, the subtitle recognition module 504 is specifically configured to determine the maximum value among the mean squared deviations of the widths of each text box; and to use each text box in the set of text boxes corresponding to the maximum value as the subtitle.

[0088] Combination Figure 6 As shown, this disclosure provides an apparatus for video subtitle recognition, including a processor 100 and a memory 101. Optionally, the apparatus may further include a communication interface 102 and a bus 103. The processor 100, communication interface 102, and memory 101 can communicate with each other via the bus 103. The communication interface 102 can be used for information transmission. The processor 100 can call logical instructions in the memory 101 to execute the video subtitle recognition method of the above embodiment.

[0089] Furthermore, the logic instructions in the aforementioned memory 101 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium.

[0090] The memory 101, as a computer-readable storage medium, can be used to store software programs and computer-executable programs, such as program instructions / modules corresponding to the methods in the embodiments of this disclosure. The processor 100 executes functional applications and data processing by running the program instructions / modules stored in the memory 101, that is, it implements the method for video subtitle recognition in the above embodiments.

[0091] The memory 101 may include a program storage area and a data storage area. The program storage area may store the operating system and applications required for at least one function; the data storage area may store data created based on the use of the terminal device. Furthermore, the memory 101 may include high-speed random access memory and may also include non-volatile memory.

[0092] This disclosure provides an electronic device that includes the above-described apparatus for video subtitle recognition.

[0093] This disclosure provides a computer-readable storage medium storing computer-executable instructions configured to perform the above-described method for video subtitle recognition.

[0094] This disclosure provides a computer program product, which includes a computer program stored on a computer-readable storage medium. The computer program includes program instructions that, when executed by a computer, cause the computer to perform the above-described method for video subtitle recognition.

[0095] The aforementioned computer-readable storage medium may be a transient computer-readable storage medium or a non-transitory computer-readable storage medium.

[0096] The technical solutions of this disclosure can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes one or more instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in this disclosure. The aforementioned storage medium can be a non-transitory storage medium, including: a USB flash drive, a portable hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk, and other media capable of storing program code; it can also be a transient storage medium.

[0097] The foregoing description and accompanying drawings fully illustrate embodiments of this disclosure to enable those skilled in the art to practice them. Other embodiments may include structural, logical, electrical, procedural, and other changes. The embodiments represent only possible variations. Individual components and functions are optional unless explicitly required, and the order of operation may vary. Parts and features of some embodiments may be included in or replace parts and features of other embodiments. Moreover, the terminology used in this application is for describing embodiments only and is not intended to limit the claims. As used in the description of embodiments and claims, the singular forms “a,” “an,” and “the” are intended to equally include the plural forms unless the context clearly indicates otherwise. Similarly, the term “and / or” as used in this application means including one or more of the associated listed items and all possible combinations thereof. Additionally, when used in this application, the term "comprise" and its variations "comprises" and / or "comprising" refer to the presence of stated features, integrals, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components, and / or groups thereof. Without further limitations, an element defined by the phrase "comprises a..." does not exclude the presence of other identical elements in the process, method, or apparatus that includes said element. In this document, each embodiment may focus on the differences from other embodiments, and similar or identical parts between embodiments can be referred to mutually. For methods, products, etc., disclosed in the embodiments, if they correspond to the method section disclosed in the embodiments, the relevant parts can be referred to the description of the method section.

[0098] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the embodiments of this disclosure. Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0099] The methods and products (including but not limited to devices and equipment) disclosed in the embodiments herein can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For instance, the division of units may be merely a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the shown or discussed units may be through some interfaces, and the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the units may be selected to implement this embodiment according to actual needs. Furthermore, the functional units in the embodiments of this disclosure may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.

[0100] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions marked in the blocks may occur in a different order than that shown in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. In the descriptions corresponding to the flowcharts and block diagrams in the accompanying drawings, the operations or steps corresponding to different blocks may also occur in a different order than disclosed in the description, and sometimes there is no specific order between different operations or steps. For example, two consecutive operations or steps may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. Each block in a block diagram and / or flowchart, and combinations of blocks in a block diagram and / or flowchart, can be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.

Claims

1. A method for video subtitle recognition, characterized in that, include: Text recognition is performed on multiple video frames to obtain all text boxes in each video frame; wherein each text box includes a line of text; Multiple text box sets are determined based on the frequency of occurrence of the text box height; the text boxes in each text box set have the same height. Calculate the mean squared error of the text box width for each set of text boxes; where the mean squared error is used to represent the degree of dispersion of the text box width in the set of text boxes. The subtitles are determined based on the mean squared error of the width of each text box; Based on the frequency of occurrence of text box height, determine multiple text box sets, including: clustering text boxes with the same height into multiple text box sets; sorting each text box set in descending order according to the number of text boxes in each set; and selecting the first N text box sets, where N is an integer greater than 1. The captions are determined based on the mean squared deviation of the widths of each text box, including: determining the maximum value among the mean squared deviations of the widths of each text box; and using each text box in the set corresponding to the maximum value as the caption.

2. The method according to claim 1, characterized in that, The mean squared error of the text box width for each set of text boxes is calculated using the following formula: Where, σ I Let w be the mean squared error of the text box widths of the i-th text box set, where i = 1, ..., N, N is the number of text box sets; n is the number of text boxes in each text box set; w i Let be the width of the i-th text box, where i = 1, ..., n; This represents the average width of the text boxes in the text box collection.

3. The method according to claim 1, characterized in that, Perform text recognition on multiple video frames to obtain all text boxes in each video frame, including: Perform text recognition on multiple consecutive video frames; Identify all text boxes in each video frame and obtain the height and width of each text box.

4. A device for video subtitle recognition, characterized in that, include: The text recognition module is configured to perform text recognition on multiple video frames to obtain all text boxes in each video frame; wherein each text box includes a line of text. The determination module is configured to determine multiple sets of text boxes based on the number of occurrences of the text box height; wherein the text boxes in each set of text boxes have the same height; The calculation module is configured to calculate the mean squared error of the text box width for each set of text boxes; wherein the mean squared error of the width is used to represent the degree of dispersion of the text box width in the set of text boxes. The subtitle recognition module determines the subtitles based on the mean squared error of the width of each text box. The determining module includes: an aggregation unit, which clusters text boxes with the same height into multiple text box sets; a sorting unit, configured to sort the text box sets in descending order according to the number of text boxes in each set; and a selection unit, configured to select the first N text box sets, where N is an integer greater than 1. The subtitle recognition module is configured to: determine the maximum value among the mean squared deviations of the widths of each text box, and use each text box in the set of text boxes corresponding to the maximum value as the subtitle.

5. An apparatus for video subtitle recognition, comprising a processor and a memory storing program instructions, characterized in that, The processor is configured to perform the method for video subtitle recognition as described in any one of claims 1 to 3 when executing the program instructions.

6. An electronic device, characterized in that, Includes the device for video caption recognition as described in claim 4 or 5.

7. A storage medium storing program instructions, characterized in that, When the program instructions are executed, they perform the method for video subtitle recognition as described in any one of claims 1 to 3.

Citation Information

Patent Citations

  • Video caption positioning method, electronic equipment and computer storage medium

    CN110598622A