Video subtitle detection method and related products

By analyzing the position and duration of the text in the video, and classifying the subtitle class, the problem of low subtitle detection efficiency in the existing technology is solved, and efficient and accurate subtitle detection is achieved.

CN117750147BActive Publication Date: 2025-08-26SHUXING TECH (BEIJING) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310125520.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-02-14
Publication Date
2025-08-26
Estimated Expiration
2043-02-14

AI Technical Summary

Technical Problem

The existing video subtitle detection methods are inefficient and cannot efficiently extract subtitle text in videos.

Method used

By obtaining the text in the video and classifying it according to the location and duration, it is determined that the text class containing the number greater than the threshold is a subtitle class, which improves the accuracy and efficiency of subtitle detection.

Benefits of technology

It realizes efficient detection of video subtitles, improving the accuracy and efficiency of subtitle detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117750147B_ABST
    Figure CN117750147B_ABST
Patent Text Reader

Abstract

This application discloses a method for detecting video subtitles and related products. The method includes: obtaining at least one first text in a video to be detected and the position of the at least one first text in the video to be detected; dividing the at least one first text into at least one text class by grouping first texts located in the same position in the video to be detected; determining a text class in which the number of first texts contained in the at least one text class is greater than or equal to a first threshold as a target text class; and using the first text in the target text class as the subtitle of the video to be detected.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence technology, and in particular to a method for detecting video subtitles and related products. Background Art

[0002] With the explosive growth of various videos, video management is becoming increasingly difficult. When subtitles are present in a video, they are often related to the video's content. Therefore, subtitles can be used as a basis for video management, reducing the difficulty of video management. For example, subtitles can be used to recommend videos or search for videos based on them. Therefore, extracting subtitle text from videos with subtitles is extremely important. Current methods involve performing text detection (optical character recognition, or OCR) on video frames to obtain text within the video frames, and then manually determining whether the obtained text contains subtitles. However, this method is inefficient in detecting subtitles. Summary of the Invention

[0003] The present application provides a video subtitle detection method and related products.

[0004] In a first aspect, a method for detecting video subtitles is provided, the method comprising:

[0005] Obtaining at least one first text in a video to be detected and a position of the at least one first text in the video to be detected;

[0006] The at least one first text is divided into at least one text class by classifying first texts located at the same position in the video to be detected into one class;

[0007] determining a text class in which the number of first texts contained in the at least one text class is greater than or equal to a first threshold as a target text class;

[0008] The first text in the target text class is used as the subtitle of the video to be detected.

[0009] In conjunction with any embodiment of the present application, after obtaining at least one first text in a video to be detected and the position of the at least one first text in the video to be detected, and before dividing the at least one first text into one category by classifying the first texts located at the same position in the video to be detected to obtain at least one text category, the method further includes:

[0010] When the number of the first texts is greater than 1, determining, according to a position of the at least one first text in the video to be detected, an intersection-over-union ratio of areas of any two first texts in the at least one first text;

[0011] When the IoU ratio is greater than or equal to a second threshold, it is determined that the two first texts corresponding to the IoU ratio are at the same position in the video to be detected.

[0012] In combination with any embodiment of the present application, the pixel coordinate system of the video to be detected includes a horizontal axis and a vertical axis, the direction of the horizontal axis is horizontal, and the direction of the vertical axis is vertical;

[0013] After determining the intersection-over-union ratio of areas of any two first texts in the at least one first text, the method further includes:

[0014] When the intersection-and-union ratio is less than the second threshold value, determining, based on positions of the two first texts corresponding to the intersection-and-union ratio in the video to be detected, a horizontal distance between the centers of the two first texts corresponding to the intersection-and-union ratio in the horizontal direction, a vertical distance between the centers of the two first texts corresponding to the intersection-and-union ratio in the vertical direction, a maximum horizontal distance between the two first texts corresponding to the intersection-and-union ratio in the horizontal direction, and a maximum vertical distance between the two first texts corresponding to the intersection-and-union ratio in the vertical direction;

[0015] When the ratio of the center horizontal distance to the maximum horizontal distance is greater than or equal to a third threshold, and the ratio of the center vertical distance to the maximum vertical axis distance is greater than or equal to the third threshold, it is determined that the two first texts corresponding to the intersection-union ratio have the same position in the video to be detected.

[0016] In combination with any embodiment of the present application, obtaining at least one first text in the video to be detected includes:

[0017] Acquire at least one second text in the video to be detected and a duration of the at least one second text in the video to be detected;

[0018] A text having a duration within a preset interval is selected from the at least one second text to obtain the at least one first text.

[0019] In combination with any embodiment of the present application, obtaining at least one second text in the video to be detected includes:

[0020] Obtaining the video to be detected;

[0021] Performing text detection on the video to be detected to obtain at least one third text;

[0022] The same third text in the at least one third text is merged to obtain the at least one second text.

[0023] In combination with any embodiment of the present application, obtaining the duration of the at least one second text in the video to be detected includes:

[0024] Using the timestamp of the video frame corresponding to the third text as the timestamp of the third text;

[0025] Determine a minimum timestamp of a third text corresponding to the second text as the start time of the second text;

[0026] Determine a maximum timestamp of a third text corresponding to the second text as an end time of the second text;

[0027] According to the start time of the at least one second text and the end time of the at least one second text, a duration of the at least one second text in the video to be detected is obtained.

[0028] In combination with any embodiment of the present application, the merging of the identical third texts in the at least one third text to obtain the at least one second text includes:

[0029] merging identical third texts in the at least one third text to obtain the at least one fourth text;

[0030] When the number of fourth texts is greater than 1 and a fifth text and a sixth text having a similarity greater than or equal to a fourth threshold exist in the at least one fourth text, merging the fifth text and the sixth text in the at least one fourth text to obtain the at least one second text;

[0031] If there are no two fourth texts with a similarity greater than or equal to a fourth threshold value in the at least one fourth text, the at least one fourth text is used as the at least one second text.

[0032] In combination with any embodiment of the present application, the merging of the fifth text and the sixth text in the at least one fourth text to obtain the at least one second text includes:

[0033] Obtaining the start time of the fifth text, the end time of the fifth text, the start time of the sixth text, and the end time of the sixth text, where the start time of the fifth text is earlier than the start time of the sixth text;

[0034] Determining a first time difference between a start time of the fifth text and an end time of the sixth text;

[0035] Determine a median of the start time and the end time of the fifth text to obtain a time median of the fifth text;

[0036] Determine a median of the start time and the end time of the sixth text to obtain a time median of the sixth text;

[0037] determining a second time difference between a time median of the fifth text and a time median of the sixth text;

[0038] When the ratio of the first time difference to the second time difference is less than or equal to a fifth threshold, the fifth text and the sixth text in the at least one fourth text are merged to obtain the at least one second text.

[0039] In a second aspect, a video subtitle detection device is provided, the detection device comprising:

[0040] an acquiring unit, configured to acquire at least one first text in a video to be detected and a position of the at least one first text in the video to be detected;

[0041] A dividing unit, configured to divide the at least one first text into one category by dividing the first texts located at the same position in the video to be detected into one category, thereby obtaining at least one text category;

[0042] a determining unit, configured to determine a text class in which the number of first texts contained in the at least one text class is greater than or equal to a first threshold as a target text class;

[0043] A processing unit is configured to use the first text in the target text class as the subtitle of the video to be detected.

[0044] The determining unit is further configured to:

[0045] When the number of the first texts is greater than 1, determining, according to a position of the at least one first text in the video to be detected, an intersection-over-union ratio of areas of any two first texts in the at least one first text;

[0046] When the IoU ratio is greater than or equal to a second threshold, it is determined that the two first texts corresponding to the IoU ratio are at the same position in the video to be detected.

[0047] In combination with any embodiment of the present application, the pixel coordinate system of the video to be detected includes a horizontal axis and a vertical axis, the direction of the horizontal axis is horizontal, and the direction of the vertical axis is vertical;

[0048] The determining unit is further configured to:

[0049] When the intersection-and-union ratio is less than the second threshold value, determining, based on positions of the two first texts corresponding to the intersection-and-union ratio in the video to be detected, a horizontal distance between the centers of the two first texts corresponding to the intersection-and-union ratio in the horizontal direction, a vertical distance between the centers of the two first texts corresponding to the intersection-and-union ratio in the vertical direction, a maximum horizontal distance between the two first texts corresponding to the intersection-and-union ratio in the horizontal direction, and a maximum vertical distance between the two first texts corresponding to the intersection-and-union ratio in the vertical direction;

[0050] When the ratio of the center horizontal distance to the maximum horizontal distance is greater than or equal to a third threshold, and the ratio of the center vertical distance to the maximum vertical axis distance is greater than or equal to the third threshold, it is determined that the two first texts corresponding to the intersection-union ratio have the same position in the video to be detected.

[0051] In combination with any embodiment of the present application, the acquiring unit is configured to:

[0052] Acquire at least one second text in the video to be detected and a duration of the at least one second text in the video to be detected;

[0053] A text having a duration within a preset interval is selected from the at least one second text to obtain the at least one first text.

[0054] In combination with any embodiment of the present application, the acquiring unit is configured to:

[0055] Obtaining the video to be detected;

[0056] Performing text detection on the video to be detected to obtain at least one third text;

[0057] The same third text in the at least one third text is merged to obtain the at least one second text.

[0058] In combination with any embodiment of the present application, the acquiring unit is configured to:

[0059] Using the timestamp of the video frame corresponding to the third text as the timestamp of the third text;

[0060] Determine a minimum timestamp of a third text corresponding to the second text as the start time of the second text;

[0061] Determine a maximum timestamp of a third text corresponding to the second text as an end time of the second text;

[0062] According to the start time of the at least one second text and the end time of the at least one second text, a duration of the at least one second text in the video to be detected is obtained.

[0063] In combination with any embodiment of the present application, the acquiring unit is configured to:

[0064] merging identical third texts in the at least one third text to obtain the at least one fourth text;

[0065] When the number of fourth texts is greater than 1 and a fifth text and a sixth text having a similarity greater than or equal to a fourth threshold exist in the at least one fourth text, merging the fifth text and the sixth text in the at least one fourth text to obtain the at least one second text;

[0066] If there are no two fourth texts with a similarity greater than or equal to a fourth threshold value in the at least one fourth text, the at least one fourth text is used as the at least one second text.

[0067] In combination with any embodiment of the present application, the acquiring unit is configured to:

[0068] Obtaining the start time of the fifth text, the end time of the fifth text, the start time of the sixth text, and the end time of the sixth text, where the start time of the fifth text is earlier than the start time of the sixth text;

[0069] Determining a first time difference between a start time of the fifth text and an end time of the sixth text;

[0070] Determine a median of the start time and the end time of the fifth text to obtain a time median of the fifth text;

[0071] Determine a median of the start time and the end time of the sixth text to obtain a time median of the sixth text;

[0072] determining a second time difference between a time median of the fifth text and a time median of the sixth text;

[0073] When the ratio of the first time difference to the second time difference is less than or equal to a fifth threshold, the fifth text and the sixth text in the at least one fourth text are merged to obtain the at least one second text.

[0074] In a third aspect, an electronic device is provided, characterized in that it includes: a processor and a memory, the memory is used to store computer program code, the computer program code includes computer instructions, and when the processor executes the computer instructions, the electronic device executes the method as described in the first aspect above and any possible implementation method thereof.

[0075] In a fourth aspect, another electronic device is provided, comprising: a processor, a sending device, an input device, an output device and a memory, wherein the memory is used to store computer program code, and the computer program code includes computer instructions. When the processor executes the computer instructions, the electronic device executes the first aspect and any one of its embodiments described above.

[0076] In a fifth aspect, a computer-readable storage medium is provided, in which a computer program is stored. The computer program includes program instructions, and when the program instructions are executed by a processor, the processor is caused to execute the first aspect and any of its embodiments described above.

[0077] In a sixth aspect, a computer program product is provided, which includes a computer program or instructions, and when the computer program or instructions are run on a computer, the computer is enabled to execute the above-mentioned first aspect and any embodiment thereof.

[0078] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present application.

[0079] In the present application, after obtaining at least one first text in a video to be detected and the position of at least one first text in the video to be detected, the detection device divides the at least one first text by classifying the first texts located at the same position in the video to be detected into one category, thereby obtaining at least one text class. Then, by determining from the at least one text class a target text class containing a number of first texts greater than or equal to a first threshold, the target text class corresponding to the position where the subtitles appear is determined. Finally, the first text in the target text class is used as the subtitle of the video to be detected, thereby detecting the subtitles in the video to be detected and improving the efficiency of detecting subtitles. BRIEF DESCRIPTION OF THE DRAWINGS

[0080] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the background technology, the drawings required for use in the embodiments of the present application or the background technology will be described below.

[0081] The drawings herein are incorporated into and constitute a part of the specification. These drawings illustrate embodiments consistent with the present application and, together with the specification, are used to illustrate the technical solutions of the present application.

[0082] Figure 1 A flowchart of a method for detecting video subtitles provided in an embodiment of the present application;

[0083] Figure 2 A flowchart of another method for detecting video subtitles provided in an embodiment of the present application;

[0084] Figure 3 A schematic diagram of a process for removing duplicate fields provided in an embodiment of the present application;

[0085] Figure 4 A schematic diagram of a clustering process provided in an embodiment of the present application;

[0086] Figure 5 A schematic diagram of the structure of a video subtitle detection device provided in an embodiment of the present application;

[0087] Figure 6 A schematic diagram of the hardware structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0088] In order to enable those skilled in the art to better understand the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0089] The terms "first," "second," and the like in the specification and claims of this application and the accompanying drawings are used to distinguish between different objects, not to describe a particular order. Furthermore, the terms "including," "having," and any variations thereof, are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or apparatus comprising a series of steps or elements is not limited to the listed steps or elements but may optionally include steps or elements not listed, or may optionally include other steps or elements inherent to the process, method, product, or apparatus.

[0090] References herein to "embodiments" mean that a particular feature, structure, or characteristic described in connection with the embodiments may be included in at least one embodiment of the present application. The appearance of this phrase in various places in the specification does not necessarily refer to the same embodiment, nor does it constitute an independent or alternative embodiment that is mutually exclusive of other embodiments. It is understood, both explicitly and implicitly, by those skilled in the art that the embodiments described herein may be combined with other embodiments.

[0091] The embodiment of the present application is implemented by a video subtitle detection device (hereinafter referred to as the detection device), wherein the detection device can be any electronic device capable of executing the technical solution disclosed in the embodiment of the present application. Optionally, the detection device can be one of the following: a computer or a server.

[0092] It should be understood that the method embodiment of the present application can also be implemented by a processor executing computer program code. The following describes the embodiment of the present application in conjunction with the drawings in the embodiment of the present application. Figure 1 , Figure 1 This is a flowchart of a method for detecting video subtitles provided in an embodiment of the present application.

[0093] 101. Obtain at least one first text in a video to be detected and a position of the at least one first text in the video to be detected.

[0094] In the embodiments of the present application, the video to be detected can be any video. The video to be detected can be an offline video or an online video. An offline video can be a video captured by a camera or a mobile smart device. An online video can be a video captured in real time by a camera. The video to be detected can be a video containing any content, for example, a movie clip, an instructional video, or a sports game video.

[0095] In the embodiment of the present application, the number of first texts is one or more, and the first texts are all texts in the video to be detected. For example, the video to be detected includes image a and image b, where image a includes text c and image b includes text d. Then, text c and text d are both first texts, that is, at least one first text includes text c and text d.

[0096] The position of at least one first text in the video to be detected includes the position of each first text in the video to be detected. In one possible implementation, the position of the first text in the video to be detected is the position of a text box containing the first text in the video to be detected. For example, by performing optical character recognition (OCR) on the video to be detected, the text box containing the first text can be determined from the video to be detected. In this case, the position of the first text in the video to be detected is the position of the text box containing the first text in the video to be detected.

[0097] Optionally, the position in the video to be detected is the position in the pixel coordinate system of the video to be detected. In this case, the position of the first text in the video to be detected is the position of the first text in the pixel coordinate system to be detected.

[0098] In one implementation of obtaining at least one first text, a detection device receives at least one first text input by a user through an input component, wherein the input component includes at least one of the following: a keyboard, a mouse, a touch screen, a touch pad, and an audio input device.

[0099] In another implementation of obtaining at least one first text, the detection device receives at least one first text sent by a terminal. The terminal may be any one of the following: a mobile phone, a computer, a tablet computer, or a server.

[0100] In another implementation of obtaining at least one first text, the detection device, upon obtaining the video to be detected, extracts at least one first text from the video to be detected by performing OCR on the video to be detected.

[0101] In an implementation of obtaining a position of at least one first text in a video to be detected, a detection device receives the position of at least one first text in the video to be detected input by a user through an input component.

[0102] In another implementation manner of obtaining the position of at least one first text in the video to be detected, the detection device receives the position of the at least one first text in the video to be detected sent by the terminal.

[0103] In another implementation method of obtaining the position of at least one first text in a video to be detected, the detection device, when obtaining the video to be detected, extracts at least one first text from the video to be detected by performing OCR on the video to be detected, and determines the position of the at least one first text in the video to be detected.

[0104] It should be understood that the step of the detection device obtaining at least one first text in the video to be detected and the step of obtaining the position of the at least one first text in the video to be detected can be performed separately or simultaneously, and this application does not limit this.

[0105] 102. Classify the first texts located at the same position in the video to be detected into one category, and divide the at least one first text to obtain at least one text category.

[0106] In step 102, the detection device classifies at least one text based on the position of the first text in the video to be detected. Specifically, first texts with the same position in the video to be detected are classified into the same class. Thus, by classifying at least one first text, at least one text class is obtained. The number of text classes is one or more, and each text class includes one or more first texts.

[0107] Optionally, the positions of the two first texts in the video to be detected are the same, which means that the distance between the two first texts in the video to be detected is less than or equal to a sixth threshold value. For example, the sixth threshold value is 2, and based on the position of the first text a in the video to be detected and the position of the first text b in the video to be detected, the distance between the first text a and the first text b in the video to be detected is determined to be d. Then, when d is less than or equal to 2, it is determined that the position of the first text a in the video to be detected is the same as the position of the first text b in the video to be detected. When d is greater than 2, it is determined that the position of the first text a in the video to be detected is different from the position of the first text b in the video to be detected.

[0108] 103. Determine a text class in which the number of first texts included in the at least one text class is greater than or equal to a first threshold as a target text class.

[0109] Subtitles in a video usually appear in the same position, or different subtitles appear in similar positions. The time interval between two adjacent frames in the video is short, and the subtitles usually appear for a period of time (such as 3 seconds to 15 seconds). That is, subtitles usually appear in multiple video frames. Therefore, the number of texts located at the position where subtitles appear in the video should be large, and conversely, the number of texts not located at the position where subtitles appear in the video should be small.

[0110] By executing step 102, the detection device has classified at least one first text located in the video to be detected according to position to obtain at least one text class. Therefore, the detection device can determine whether the position corresponding to the text class is the position where subtitles appear based on the number of first texts in the text class.

[0111] In an embodiment of the present application, the detection device uses the first threshold as a basis to determine whether the number of first texts in the text class is large or small, and thus can determine whether the position corresponding to the text class is the position where subtitles appear. Specifically, if the number of first texts in the text class is less than the first threshold, it means that the number of first texts in the text class is small, and thus it is determined that the position corresponding to the text class is not the position where subtitles appear. Conversely, if the number of first texts in the text class is greater than or equal to the first threshold, it means that the number of first texts in the text class is large, and thus it is determined that the position corresponding to the text class is the position where subtitles appear. In an embodiment of the present application, the text class whose corresponding position is the position where subtitles appear is referred to as the target text class.

[0112] For example, at least one text class includes text class a and text class b, wherein the number of first texts in text class a is 1, and the number of first texts in text class b is 40. If the first threshold is 5, then the number of first texts in text class a is less than the first threshold, that is, text class a is not the target text class, and the number of first texts in text class b is greater than the first threshold, that is, text class b is the target text class.

[0113] 104. Use the first text in the target text class as the subtitle of the video to be detected.

[0114] Since the position corresponding to the target text class is the position where the subtitle appears, and the first text in the target text class is a subtitle, the detection device can use the first text in the target text class as the subtitle of the video to be detected.

[0115] In an embodiment of the present application, after obtaining at least one first text in a video to be detected and the position of at least one first text in the video to be detected, the detection device divides the at least one first text into one category by classifying the first texts located at the same position in the video to be detected, thereby obtaining at least one text class. Then, by determining from the at least one text class a target text class containing a number of first texts greater than or equal to a first threshold, the target text class corresponding to the position where the subtitles appear is determined. Finally, the first text in the target text class is used as the subtitle of the video to be detected, thereby detecting the subtitles in the video to be detected and improving the efficiency of detecting subtitles.

[0116] As an optional implementation, the detection device further performs the following steps after performing step 101 and before performing step 102:

[0117] 201. When the number of first texts is greater than 1, determine, based on a position of the at least one first text in the video to be detected, an intersection over union (IoU) of areas of any two first texts in the at least one first text.

[0118] In the embodiment of the present application, the IoU of the areas of the two first texts is calculated as follows: the area of ​​the intersection of the pixel areas covered by the two first texts / the area of ​​the union of the pixel areas covered by the two first texts.

[0119] When the number of first texts is greater than one, the detection device can determine the pixel areas covered by the first texts in the video to be detected based on the positions of the first texts in the video to be detected. Thus, the pixel areas covered by each first text in the video to be detected can be determined based on the positions of each first text in the video to be detected. After determining the pixel areas covered by each first text in the video to be detected, the IoU of the areas of any two first texts can be determined.

[0120] In one possible implementation, the position of the first text in the video to be detected is the position of the text box containing the first text in the video to be detected. Based on the positions of any two first texts in the video to be detected, the detection device can determine the intersection of the text boxes corresponding to the two first texts and the union of the text boxes corresponding to the two first texts. Furthermore, based on the area of ​​the intersection of the text boxes corresponding to the two first texts and the area of ​​the union of the text boxes corresponding to the two first texts, the IoU of the areas of the two first texts can be obtained.

[0121] Optionally, when the number of the second texts is 1, the detection device uses the second text as the first text.

[0122] 202. When the IoU is greater than or equal to a second threshold, determine that the two first texts corresponding to the IoU are at the same position in the video to be detected.

[0123] A large IoU between the two first texts indicates a high degree of overlap between the two first texts in the video to be detected. Conversely, a small IoU between the two first texts indicates a low degree of overlap between the two first texts in the video to be detected. If the two first texts have a high degree of overlap in the video to be detected, it can be determined that the two first texts are located in the same position in the video to be detected.

[0124] In the embodiment of the present application, the detection device determines whether the IoU is large or small based on the second threshold, thereby determining whether the two first texts corresponding to the IoU are in the same position in the video to be detected. Specifically, if the IoU is greater than or equal to the second threshold, it indicates that the IoU is large, and the two first texts corresponding to the IoU are in the same position in the video to be detected. Conversely, if the IoU is less than the second threshold, it indicates that the IoU is small, and the overlap of the two first texts corresponding to the IoU is low.

[0125] In this embodiment, when the number of first texts is greater than 1, the detection device determines the IoU of the areas of any two first texts in the at least one first text based on the position of the at least one first text in the video to be detected. Then, based on the IoU of the two first texts, it can be determined whether the positions of the two first texts in the video to be detected are the same. Specifically, when the IoU is greater than or equal to a second threshold, it is determined that the two first texts corresponding to the IoU are in the same position in the video to be detected, thereby improving the accuracy of determining whether the two first texts are in the same position in the video to be detected.

[0126] When a subtitle contains a large number of characters, the subtitle may be displayed in at least two parts. For example, in a video to be detected, if the subtitle contains a large number of characters, the subtitle may be displayed in two lines, one above the other. Obviously, when a subtitle is displayed in at least two parts, different contents of the subtitle will be displayed in different areas. In other words, the text in different areas all belong to the same subtitle, that is, the different areas together constitute the area where the subtitle is displayed. Therefore, the text in different areas should be classified into the same category. In other words, the positions of different texts in the same subtitle in the video to be detected should be considered to be the same.

[0127] However, when the text of the subtitle is divided into at least two parts for display, there may be no intersection between the texts located in different areas, which leads to a low degree of overlap of the texts located in different areas, that is, the IoU of the texts located in different areas is small. That is to say, when the IoU of the two first texts is small, it may be misjudged to determine that the positions of the two first texts in the video to be detected are different. Therefore, when the IoU of the two first texts is less than the second threshold, the detection device can determine whether the positions of the two first texts in the video to be detected are the same by executing the following optional implementation method.

[0128] As an optional embodiment, the pixel coordinate system of the video to be detected includes a horizontal axis and a vertical axis, wherein the horizontal axis is in the horizontal direction and the vertical axis is in the vertical direction. After determining the intersection-and-union ratio of the areas of any two first texts in the at least one first text, the detection device further performs the following steps:

[0129] 301. When the IoU is less than the second threshold, determine the horizontal distance between the centers of the two first texts corresponding to the IoU in the horizontal direction, and determine the vertical distance between the centers of the two first texts corresponding to the IoU in the vertical direction, and determine the maximum horizontal distance between the two first texts corresponding to the IoU in the horizontal direction, and determine the maximum vertical distance between the two first texts corresponding to the IoU in the vertical direction, according to the positions of the two first texts corresponding to the IoU in the video to be detected.

[0130] In the embodiment of the present application, the center of the first text may be the geometric center of the pixel area covered by the first text. Alternatively, the center of the first text may be the geometric center of the text box containing the first text.

[0131] In the embodiments of the present application, the horizontal distance between the centers of two first texts is referred to as the horizontal center distance. Specifically, the horizontal distance between the centers of two first texts is the distance difference between the horizontal coordinates of the centers of the two first texts. For example, if the horizontal coordinate of the center of first text a is 2 and the horizontal coordinate of the center of first text b is 5, then the distance difference between the horizontal coordinates of the centers of first text a and first text b is (5-2)=3, that is, the horizontal center distance between first text a and first text b is 3.

[0132] In the embodiments of the present application, the vertical distance between the centers of two first texts is referred to as the center longitudinal distance. Specifically, the vertical distance between the centers of two first texts is the difference in the distances between the longitudinal coordinates of the centers of the two first texts. For example, if the longitudinal coordinate of the center of first text a is 3 and the longitudinal coordinate of the center of first text b is 8, then the difference in the distances between the longitudinal coordinates of the centers of first text a and first text b is (8-3)=5, i.e., the vertical distance between the centers of first text a and first text b is 5.

[0133] In an embodiment of the present application, the maximum horizontal distance between two first texts is referred to as the maximum horizontal distance. Specifically, the maximum horizontal distance between two first texts is the distance difference between the maximum horizontal coordinate and the minimum horizontal coordinate in the two first texts. For example, the minimum horizontal coordinate of the pixel area covered by the first text a is 2, the maximum horizontal coordinate of the pixel area covered by the first text a is 10, the minimum horizontal coordinate of the pixel area covered by the first text b is 6, and the maximum horizontal coordinate of the pixel area covered by the first text a is 13. At this time, the minimum horizontal coordinate in the first text a and the first text b is 2, the maximum horizontal coordinate in the first text a and the first text b is 13, and the maximum horizontal distance between the first text a and the first text b is (13-2)=9.

[0134] In an embodiment of the present application, the maximum vertical distance between two first texts is referred to as the maximum vertical distance. Specifically, the maximum vertical distance between the two first texts is the difference between the maximum vertical coordinate and the minimum vertical coordinate in the two first texts. For example, the minimum vertical coordinate of the pixel area covered by the first text a is 7, the maximum vertical coordinate of the pixel area covered by the first text a is 19, the minimum vertical coordinate of the pixel area covered by the first text b is 13, and the maximum vertical coordinate of the pixel area covered by the first text a is 20. At this time, the minimum vertical coordinate in the first text a and the first text b is 7, the maximum vertical coordinate in the first text a and the first text b is 20, and the maximum vertical distance between the first text a and the first text b is (20-7)=13.

[0135] 302. When the ratio of the above-mentioned central horizontal distance to the above-mentioned maximum horizontal distance is greater than or equal to the third threshold, and the ratio of the above-mentioned central vertical distance to the above-mentioned maximum vertical axis distance is greater than or equal to the above-mentioned third threshold, it is determined that the two first texts corresponding to the above-mentioned IoU are in the same position in the above-mentioned video to be detected.

[0136] When subtitle text is displayed in at least two parts, the distance between text in different regions is typically small. This means that the text in different regions has a high degree of horizontal and vertical overlap. Therefore, based on the horizontal and vertical overlap of two first texts, it can be determined whether the two first texts are located in the same position in the video to be detected.

[0137] The larger the ratio of the horizontal distance between the centers of the two first texts to the maximum horizontal distance between the two first texts, the higher the horizontal overlap of the two first texts. Similarly, the larger the ratio of the vertical distance between the centers of the two first texts to the maximum vertical distance between the two first texts, the higher the vertical overlap of the two first texts. Therefore, when the ratio of the horizontal distance between the centers of the two first texts to the maximum horizontal distance between the two first texts is large, and the vertical distance between the centers of the two first texts is large compared to the maximum vertical distance between the two first texts, the detection device can determine that the two first texts are in the same position in the video to be detected.

[0138] In an embodiment of the present application, the detection device uses the third threshold as a basis to determine whether the ratio of the central horizontal distance of the two first texts to the horizontal maximum distance of the two first texts is large or small, and can also use the third threshold as a basis to determine whether the ratio of the central vertical distance of the two first texts to the vertical maximum distance of the two first texts is large or small, thereby determining whether the positions of the two first texts in the video to be detected are the same. Specifically, when the ratio of the central horizontal distance of the two first texts to the horizontal maximum distance of the two first texts is greater than or equal to the third threshold, and the ratio of the central vertical distance of the two first texts to the vertical maximum distance of the two first texts is greater than or equal to the third threshold, the detection device determines that the ratio of the central vertical distance of the two first texts to the vertical maximum distance of the two first texts is large, and the ratio of the central vertical distance of the two first texts to the vertical maximum distance of the two first texts is large, thereby determining that the positions of the two first texts in the video to be detected are the same.

[0139] In this embodiment, when the intersection-and-union ratio is less than a second threshold value, the detection device determines the center horizontal distance, center vertical distance, horizontal maximum distance, and vertical maximum distance of the two first texts corresponding to the intersection-and-union ratio based on the positions of the two first texts corresponding to the intersection-and-union ratio in the video to be detected, and then determines that the two first texts corresponding to the intersection-and-union ratio have the same position in the video to be detected based on the center horizontal distance, center vertical distance, horizontal maximum distance, and vertical maximum distance of the two first texts corresponding to the intersection-and-union ratio, thereby improving the accuracy of judging that the two first texts have the same position in the video to be detected.

[0140] As an optional implementation manner, the detection device obtains at least one first text in the video to be detected by performing the following steps:

[0141] 401. Obtain at least one second text in the video to be detected and a duration of the at least one second text in the video to be detected.

[0142] In the embodiment of the present application, the number of second texts is one or more, and each second text is text in the video to be detected. For example, if the video to be detected includes image a and image b, where image a includes text c and image b includes text d and text e, then text c, text d, and text e are all second texts, i.e., at least one second text includes text c, text d, and text e.

[0143] In an embodiment of the present application, the duration of the second text in the video to be detected refers to the duration of the second text appearing in the video to be detected. For example, if the second text first appears in the tenth frame of the video to be detected and last appears in the thirtieth frame of the video to be detected, then the duration of the second text is the duration from the timestamp of the tenth frame to the timestamp of the thirtieth frame. The duration of at least one second text includes the duration of each second text.

[0144] In an implementation manner of acquiring at least one second text, a detection device receives at least one second text input by a user through an input component.

[0145] In another implementation manner of obtaining at least one second text, the detection device receives at least one second text sent by the terminal.

[0146] In yet another implementation of obtaining at least one second text, the detection device, upon obtaining the video to be detected, extracts at least one second text from the video to be detected by performing OCR on the video to be detected.

[0147] In an implementation manner of obtaining the duration of at least one second text, a detection device receives the duration of at least one second text input by a user through an input component.

[0148] In another implementation manner of obtaining the duration of at least one second text, the detection device receives the duration of at least one second text sent by the terminal.

[0149] It should be understood that the step of the detection device obtaining at least one second text in the video to be detected and the step of obtaining the duration of at least one second text can be performed separately or simultaneously, and this application does not limit this.

[0150] 402. Select texts whose duration is within a preset range from the at least one second text to obtain the at least one first text.

[0151] The text in the video includes both subtitles and non-subtitles. For example, the watermark in the video also belongs to the text in the video. Therefore, the text in the video needs to be filtered to filter out the text belonging to the subtitles from the text in the video.

[0152] As described in step 103, the appearance of subtitles in a video usually lasts for a period of time. Therefore, the duration of the subtitles can be used as a basis to filter out text belonging to the subtitles from at least one second text in the video to be detected. Specifically, the detection device selects text belonging to the subtitles (i.e., the at least one first text) whose duration is within a preset interval from the at least one second text. Optionally, the preset interval is 3 seconds to 15 seconds.

[0153] In this embodiment, after obtaining at least one second text in the video to be detected and the duration of at least one second text in the video to be detected, the detection device selects a text whose duration is within a preset interval from the at least one second text as at least one first text, which can increase the probability that the at least one first text is a subtitle text.

[0154] As an optional implementation manner, the detection device obtains at least one second text in the video to be detected by performing the following steps:

[0155] 501. Obtain the video to be detected.

[0156] In one implementation of obtaining a video to be detected, a detection device receives the video to be detected input by a user through an input component. In another implementation of obtaining a video to be detected, the detection device receives the video to be detected sent by a terminal. In yet another implementation of obtaining a video to be detected, the detection device obtains the video to be detected by downloading the video from the internet. In yet another implementation of obtaining a video to be detected, a communication connection is established between the detection device and a camera, and the camera obtains the video captured by the camera as the video to be detected via the communication connection.

[0157] 502. Perform OCR on the video to be detected to obtain at least one third text.

[0158] In an embodiment of the present application, the number of third texts is one or more. The detection device obtains at least one third text by performing OCR on the video frames in the video to be detected, that is, at least one third text includes the text in each video frame. For example, the video to be detected includes a first frame and a second frame. The detection device obtains text a and text b by performing OCR on the first frame. The detection device obtains text b and text c by performing OCR on the second frame. At this time, at least one third text includes text a, text b in the first frame, text b in the second frame, and text c. In other words, when the text content in different video frames is the same, two different third texts can also be obtained by performing OCR on the video to be detected.

[0159] It should be understood that by performing OCR on the video to be detected, not only at least one third text can be obtained, but also the timestamp of the third text and the position of the third text in the video to be detected can be obtained, wherein the timestamp of the third text is the timestamp of the video frame containing the third text.

[0160] 503. Merge identical third texts in the at least one third text to obtain the at least one second text.

[0161] When the number of the third texts is one, the detection device uses the third text as the second text. When the number of the third texts is greater than one, the detection device merges the same texts in all the third texts to obtain at least one second text.

[0162] In an embodiment of the present application, merging the same third text means retaining text information of two or more same third texts, and determining the duration of the retained text information based on the timestamps of the two or more same third texts to obtain a second text, wherein the duration of the retained text information is the time difference between the minimum timestamp and the maximum timestamp of the two or more same third texts.

[0163] For example, third text a, third text b, and third text c are three identical third texts, where the timestamp of third text a is t1, the timestamp of third text b is t2, and the timestamp of third text c is t3. If t1 is less than t2, and t2 is less than t3, the text obtained by merging third text a, third text b, and third text c is second text d. Then, the duration of second text d is (t3-t1), where the text information of second text d is the same as the text information of any of the text information of third text a, third text b, and third text c.

[0164] In this embodiment, after acquiring the video to be detected, the detection device performs text detection on the video to be detected to obtain at least one third text, and then merges the same third text in the at least one third text to obtain at least one second text.

[0165] It should be understood that in this embodiment, by merging the same third text in at least one third text to obtain at least one second text, the duration of the second text can be determined, that is, the duration of the second text can also be determined through this embodiment.

[0166] As an optional implementation manner, the detection device obtains the duration of at least one second text in the video to be detected by performing the following steps:

[0167] 601. Use the timestamp of the video frame corresponding to the third text as the timestamp of the third text.

[0168] In other words, the timestamp of the video frame containing the third text is used as the timestamp of the third text.

[0169] 602. Determine a minimum timestamp of a third text corresponding to the second text as the start time of the second text.

[0170] In this embodiment, the second text is obtained by merging the same third text, so the third text corresponding to the second text is the merged third text. For example, the detection device obtains the second text c by merging the third text a and the third text b, then the third text corresponding to the second text c includes the third text a and the third text b.

[0171] There are two or more third texts corresponding to a second text, and different third texts have different timestamps. Therefore, the minimum timestamp of the third texts corresponding to the second text is the minimum timestamp among all the third texts corresponding to the second text. For example, the third texts corresponding to second text a include third text b and third text c, where third text b has a timestamp of t1 and third text c has a timestamp of t2. If t1 is less than t2, then the minimum timestamp of the third texts corresponding to second text a is t1.

[0172] 603. Determine a maximum timestamp of a third text corresponding to the second text as the end time of the second text.

[0173] The maximum timestamp of the third text corresponding to the second text is the maximum timestamp among all the third texts corresponding to the second text. For example, the third texts corresponding to the second text a include third texts b and c, where the timestamp of third text b is t1 and the timestamp of third text c is t2. If t1 is less than t2, then the maximum timestamp of the third text corresponding to the second text a is t2.

[0174] Based on steps 601 to 603 , the detection device can determine the start time and end time of any second text, that is, through steps 601 to 603 , the start time and end time of at least one second text can be determined.

[0175] 604. Obtain a duration of the at least one second text in the video to be detected according to the start time and the end time of the at least one second text.

[0176] The detection device determines the time difference between the end time of the second text and the start time of the second text, and can obtain the duration of the second text in the video to be detected.

[0177] By executing this embodiment, the detection device can determine the duration of at least one second text in the video to be detected.

[0178] As an optional implementation manner, the detection device performs the following steps during the execution of step 503:

[0179] 701. Merge identical third texts in the at least one third text to obtain the at least one fourth text.

[0180] In this step, the "identical third text" refers to a third text having the same text information as determined by OCR recognition. The implementation method for merging identical third texts in this step can be found in step 503 and will not be described in detail here. It should be understood that in step 503, the detection device obtains at least one second text by merging identical third texts in at least one third text. However, in this step, the result is not at least one second text, but rather at least one fourth text, where the number of fourth texts is one or more.

[0181] 702. When the number of fourth texts is greater than 1 and a fifth text and a sixth text having a similarity greater than or equal to a fourth threshold exist in the at least one fourth text, merge the fifth text and the sixth text in the at least one fourth text to obtain the at least one second text.

[0182] In the embodiment of the present application, the similarity between the two fourth texts is large, indicating that the semantic matching degree of the two fourth texts, that is, the matching degree of the text information carried by the two fourth texts is high. In this case, the two fourth texts can be regarded as the same fourth text.

[0183] In an embodiment of the present application, the detection device determines whether the similarity between the two fourth texts is large or small based on the fourth threshold. Specifically, if the similarity between the two fourth texts is greater than or equal to the fourth threshold, it means that the similarity between the two fourth texts is large. Conversely, if the similarity between the two fourth texts is less than the fourth threshold, it means that the similarity between the two fourth texts is small.

[0184] In this embodiment of the present application, if the number of fourth texts is greater than one, the fifth and sixth texts are two different fourth texts among all fourth texts, and the similarity between the fifth and sixth texts is greater than or equal to the fourth threshold, i.e., the fifth and sixth texts are the same text. Therefore, by executing step 702 to obtain at least one second text, the detection device can improve the accuracy of the at least one second text.

[0185] It should be understood that the fifth and sixth texts are merely description objects selected to concisely describe the technical solutions provided by the embodiments of the present application. It should not be understood that when the number of fourth texts is greater than 1, only the fifth and sixth texts in at least one fourth text have a similarity greater than or equal to the fourth threshold, nor should it be understood that only the fifth and sixth texts in at least one fourth text are merged. In actual applications, at least one fourth text may contain one or more fourth text pairs with a similarity greater than or equal to the fourth threshold. The detection device can obtain at least one second text by merging all fourth text pairs in at least one fourth text with a similarity greater than or equal to the fourth threshold.

[0186] 703. If there are no two fourth texts with a similarity greater than or equal to a fourth threshold in the at least one fourth text, use the at least one fourth text as the at least one second text.

[0187] There are no two fourth texts with a similarity greater than or equal to the fourth threshold in the at least one fourth text, indicating that there are no identical two fourth texts in the at least two fourth texts. Therefore, the detection device uses the at least one fourth text as the at least one second text.

[0188] In this embodiment, after the detection device merges the identical third texts in at least one third text to obtain at least one fourth text, it further determines whether the identical fourth text exists in the at least one fourth text based on the similarity between the fourth texts. If it is determined that the identical fourth text exists in the at least one fourth text, at least one second text is obtained by merging the identical fourth texts in the at least one fourth text. If it is determined that the identical fourth text does not exist in the at least one fourth text, the at least one fourth text is used as the at least one second text, thereby improving the accuracy of the at least one second text.

[0189] As an optional implementation, when the number of fourth texts is greater than 1 and at least one fourth text contains a fifth text and a sixth text whose similarity is greater than or equal to a fourth threshold, the detection device merges the fifth text and the sixth text in the at least one fourth text by performing the following steps to obtain at least one second text:

[0190] 801. Obtain the start time of the fifth text, the end time of the fifth text, the start time of the sixth text, and the end time of the sixth text. The start time of the fifth text is earlier than the start time of the sixth text.

[0191] 802. Determine a first time difference between the start time of the fifth text and the end time of the sixth text.

[0192] In the embodiment of the present application, the time difference between the start time of the fifth text and the end time of the sixth text is referred to as the first time difference.

[0193] 803. Determine a median of the start time of the fifth text and the end time of the fifth text to obtain a time median of the fifth text.

[0194] For example, the start time of the fifth text is t1, and the end time of the fifth text is t2, then the time median of the fifth text is (t2-t1) / 2.

[0195] 804. Determine a median of the start time of the sixth text and the end time of the sixth text to obtain a time median of the sixth text.

[0196] For example, the start time of the sixth text is t1, and the end time of the sixth text is t2, then the time median of the sixth text is (t2-t1) / 2.

[0197] 805. Determine a second time difference between a time median of the fifth text and a time median of the sixth text.

[0198] In the embodiment of the present application, the time difference between the time median of the fifth text and the time median of the sixth text is referred to as the second time difference.

[0199] 806. When the ratio of the first time difference to the second time difference is less than or equal to a fifth threshold, merge the fifth text and the sixth text in the at least one fourth text to obtain the at least one second text.

[0200] A video may have the same subtitle appear repeatedly at different times. If the time interval between the appearances of the same subtitle is large, it is unreasonable to merge the two identical subtitles into one. For example, subtitle A appears between 1 minute and 1 minute 20 seconds into the video being tested, and also between 15 minutes and 15 minutes 20 seconds into the video being tested. In this case, merging the two subtitles into one subtitle is unreasonable.

[0201] Therefore, when the detection device determines that two fourth texts are the same fourth text, it can further determine whether the time interval between the two fourth texts is large or small, and merge the two fourth texts if it is determined that the time interval between the two fourth texts is large, and not merge the two fourth texts if it is determined that the time interval between the two fourth texts is small.

[0202] In an embodiment of the present application, if the ratio of the first time difference to the second time difference is less than or equal to the fifth threshold, it means that the interval between the time when the fifth text appears in the video to be detected and the time when the sixth text appears in the video to be detected is small. In this case, the fifth text and the sixth text can be determined to be the same text, and then at least one second text can be obtained by merging the fifth text and the sixth text in at least one fourth text. Conversely, if the ratio of the first time difference to the second time difference is large, it means that the interval between the time when the fifth text appears in the video to be detected and the time when the sixth text appears in the video to be detected is large. In this case, the fifth text and the sixth text can be determined to be different texts, and then the fifth text and the sixth text are not merged.

[0203] In this embodiment, when the number of fourth texts is greater than 1 and there is a fifth text and a sixth text in at least one fourth text whose similarity is greater than or equal to a fourth threshold, the detection device obtains the start time of the fifth text, the end time of the fifth text, the start time of the sixth text, and the end time of the sixth text, wherein the start time of the fifth text is earlier than the start time of the sixth text. Then, a first time difference between the start time of the fifth text and the end time of the sixth text is determined, a median of the start time of the fifth text and the end time of the fifth text is determined to obtain the time median of the fifth text, a median of the start time of the sixth text and the end time of the sixth text is determined to obtain the time median of the sixth text, and a second time difference between the time median of the fifth text and the time median of the sixth text is determined.

[0204] Finally, based on the ratio of the first time difference to the second time difference, it is determined whether the interval between the time when the fifth text appears in the video to be detected and the time when the sixth text appears in the video to be detected is larger or smaller. Specifically, if the ratio of the first time difference to the second time difference is less than or equal to a fifth threshold, it is determined that the interval between the time when the fifth text appears in the video to be detected and the time when the sixth text appears in the video to be detected is smaller, and the fifth text and the sixth text in the at least one fourth text are merged to obtain at least one second text. This can improve the accuracy of the at least one second text.

[0205] Based on the technical solution provided in the embodiment of this application, the embodiment of this application also provides a possible implementation method. Figure 2 , Figure 2 This is a flow chart of another method for detecting video subtitles provided in an embodiment of the present application. Figure 2 As shown, the detection device first extracts frames from the video to be processed to obtain the video to be detected. Optionally, the detection device extracts frames from the video to be processed every 0.2 seconds. Then, OCR recognition is performed on the extracted frame result (i.e., the video to be detected) to obtain at least one third text. The text information in the third text is used as a key, and the position of the third text in the video to be detected and the timestamp of the third text are stored as values ​​in the dictionary.

[0206] At least one fourth text is obtained by merging the same third text in at least one third text. Repeated fields (i.e., fields generated by OCR recognition errors) in at least one fourth text are removed. Figure 2 Removing duplicate fields caused by OCR recognition errors). For example, if there are cursive characters or handwritten characters in the video to be detected, OCR may easily misdetect them.

[0207] Optional, see Figure 3 , Figure 3 A schematic diagram of a process for removing duplicate fields provided in an embodiment of the present application. Figure 3 As shown, the detection device first compares the field similarity (i.e., determines the similarity of any two fourth texts). When it is determined that there is a field pair whose similarity exceeds the threshold value, it is judged that the field pair has an intersection on the time axis, and the implementation process of determining that there is a field pair whose similarity exceeds the threshold value can be referred to as the implementation process of determining that at least one fourth text has a fifth text and a sixth text whose similarity is greater than or equal to the fourth threshold value as described above. It is judged that the field pair has an intersection on the time axis, and the interval between the time when the fifth text appears in the video to be detected and the time when the sixth text appears in the video to be detected is large or small as described above. Specifically, when the interval between the time when the fifth text appears in the video to be detected and the time when the sixth text appears in the video to be detected is small, it is determined that the fifth text and the sixth text have an intersection on the time axis, and then it is judged that the fifth text and the sixth text are repeated fields (i.e., the fifth text and the sixth text are identical texts), and then the fifth text and the sixth text are merged. When the interval between the time when the fifth text appears in the video to be detected and the time when the sixth text appears in the video to be detected is large, it is determined that the fifth text and the sixth text have no intersection on the time axis, and then it is judged that the fifth text and the sixth text are two sentences at different times (that is, the fifth text and the sixth text are different texts), and the fifth text and the sixth text are not merged.

[0208] After obtaining at least one second text by removing repeated fields caused by OCR recognition errors, remove fields with a duration that is too long or too short, that is, remove second texts with a duration outside the preset range, to obtain at least one first text. The implementation of this process can be found in the above step 402.

[0209] After obtaining at least one first text, the text positions identified by OCR are clustered to obtain at least one text class. That is, based on the position of the first text in the video to be detected, the at least one text is classified to obtain at least one text class. The implementation of this process can be seen in step 102. Clustering results with fewer than a threshold number of samples are removed to obtain the OCR results of the subtitles. The implementation of this process can be seen in steps 103 and 104.

[0210] Optional, Figure 4 A schematic diagram of a clustering process provided in an embodiment of the present application, wherein the detection device removes clustering results with a sample number less than a threshold value and obtains an OCR result of a subtitle by executing Figure 4 Clustering is performed using the clustering process shown in Figure 4 As shown, the detection device first compares the horizontal intersection-and-union ratio of the field text box, that is, compares the horizontal intersection-and-union ratio of the text boxes of any two first texts. When the horizontal intersection-and-union ratio is less than the preset value, the vertical intersection-and-union ratio of the field text box is further compared, that is, compares the vertical intersection-and-union ratio of the text boxes of the two first texts. When the vertical intersection-and-union ratio is also less than the preset value, and the intersection-and-union ratio of the area of ​​the field text box is determined to be less than or equal to the second threshold by comparing the intersection-and-union ratio of the area of ​​the field text box, the text boxes are clustered, and the clusters with fewer samples are deleted according to the clustering results to obtain subtitle clustering, wherein the intersection-and-union ratio of the area of ​​the field text box is less than or equal to the second threshold, that is, the intersection-and-union ratio of the area of ​​the text boxes of the two first texts is less than or equal to the second threshold, clustering the text boxes determines that the two first texts have the same position in the video to be detected, and then the two first texts are divided into the same text class. The clusters with fewer samples are deleted according to the clustering results, and the implementation of obtaining subtitle clustering can be seen in step 104.

[0211] Based on the technical solution provided by the embodiment of the present application, the embodiment of the present application also provides several possible application scenarios. With the rapid development of user generated content (UGC) communities, the application of UGC communities has become more and more extensive, including users sharing videos through UGC communities. Specifically, users upload videos to the UGC community so that all users in the UGC community can watch the video. Since different users are interested in different videos, it is necessary to manage the videos uploaded to UGC (hereinafter referred to as uploaded videos for short) and present the videos that users are interested in to users. If the server running the UGC community is used as the above-mentioned detection device, then when the server stores the uploaded video, the server can determine the subtitles in the uploaded video by executing the video subtitle detection method described above, and then the subtitles can be used as tags for the uploaded video. In this way, the server can determine the videos that users are interested in based on the tags of the uploaded subtitles.

[0212] In one possible implementation scenario, after obtaining user information (such as the user's age, interests, gender, etc.), the server identifies uploaded videos with tags that match the user information as videos of interest to the user, where the user information can be used to identify the user. In another possible implementation scenario, the user sends a video search instruction to the server via a terminal, where the video search instruction includes a search term for searching for videos. After obtaining the video search instruction, the server identifies uploaded videos with tags that match the search term in the video search instruction as videos of interest to the user.

[0213] Those skilled in the art will understand that in the above-mentioned method of the specific implementation method, the writing order of each step does not mean a strict execution order and does not constitute any limitation on the implementation process. The specific execution order of each step should be determined by its function and possible internal logic.

[0214] If the technical solution of this application involves personal information, the product that applies the technical solution of this application has clearly informed the personal information processing rules and obtained the individual's voluntary consent before processing personal information. If the technical solution of this application involves sensitive personal information, the product that applies the technical solution of this application has obtained the individual's separate consent before processing sensitive personal information, and at the same time meets the "explicit consent" requirement. For example, on personal information collection devices such as cameras, a clear and prominent sign is set to inform that the personal information collection scope has been entered and personal information will be collected. If the individual voluntarily enters the collection scope, it is deemed that they agree to the collection of their personal information; or on the personal information processing device, when the personal information processing rules are notified by obvious signs / information, the individual's authorization is obtained through pop-up information or by asking the individual to upload their personal information; among which, personal information processing may include information such as the personal information processor, the purpose of personal information processing, the processing method, and the type of personal information processed.

[0215] The above describes in detail the method of the embodiment of the present application, and the following provides an apparatus of the embodiment of the present application.

[0216] See also Figure 5 , Figure 5 This is a structural diagram of a video subtitle detection device 1 provided in an embodiment of the present application. The detection device 1 includes: an acquisition unit 11, a division unit 12, a determination unit 13, and a processing unit 14. Specifically:

[0217] An acquiring unit 11 is configured to acquire at least one first text in a video to be detected and a position of the at least one first text in the video to be detected;

[0218] A dividing unit 12 is configured to divide the at least one first text into one category by dividing the first texts located at the same position in the video to be detected into one category, thereby obtaining at least one text category;

[0219] A determining unit 13 is configured to determine a text class in which the number of first texts contained in the at least one text class is greater than or equal to a first threshold as a target text class;

[0220] The processing unit 14 is configured to use the first text in the target text class as the subtitle of the video to be detected.

[0221] The determining unit 13 is further configured to:

[0222] When the number of the first texts is greater than 1, determining, according to a position of the at least one first text in the video to be detected, an intersection-over-union ratio of areas of any two first texts in the at least one first text;

[0223] When the IoU ratio is greater than or equal to a second threshold, it is determined that the two first texts corresponding to the IoU ratio are at the same position in the video to be detected.

[0224] In combination with any embodiment of the present application, the pixel coordinate system of the video to be detected includes a horizontal axis and a vertical axis, the direction of the horizontal axis is horizontal, and the direction of the vertical axis is vertical;

[0225] The determining unit 13 is further configured to:

[0226] When the intersection-and-union ratio is less than the second threshold value, determining, based on positions of the two first texts corresponding to the intersection-and-union ratio in the video to be detected, a horizontal distance between the centers of the two first texts corresponding to the intersection-and-union ratio in the horizontal direction, a vertical distance between the centers of the two first texts corresponding to the intersection-and-union ratio in the vertical direction, a maximum horizontal distance between the two first texts corresponding to the intersection-and-union ratio in the horizontal direction, and a maximum vertical distance between the two first texts corresponding to the intersection-and-union ratio in the vertical direction;

[0227] When the ratio of the center horizontal distance to the maximum horizontal distance is greater than or equal to a third threshold, and the ratio of the center vertical distance to the maximum vertical axis distance is greater than or equal to the third threshold, it is determined that the two first texts corresponding to the intersection-union ratio have the same position in the video to be detected.

[0228] In combination with any embodiment of the present application, the acquiring unit 11 is configured to:

[0229] Acquire at least one second text in the video to be detected and a duration of the at least one second text in the video to be detected;

[0230] A text having a duration within a preset interval is selected from the at least one second text to obtain the at least one first text.

[0231] In combination with any embodiment of the present application, the acquiring unit 11 is configured to:

[0232] Obtaining the video to be detected;

[0233] Performing text detection on the video to be detected to obtain at least one third text;

[0234] The same third text in the at least one third text is merged to obtain the at least one second text.

[0235] In combination with any embodiment of the present application, the acquiring unit 11 is configured to:

[0236] Using the timestamp of the video frame corresponding to the third text as the timestamp of the third text;

[0237] Determine a minimum timestamp of a third text corresponding to the second text as the start time of the second text;

[0238] Determine a maximum timestamp of a third text corresponding to the second text as an end time of the second text;

[0239] According to the start time of the at least one second text and the end time of the at least one second text, a duration of the at least one second text in the video to be detected is obtained.

[0240] In combination with any embodiment of the present application, the acquiring unit 11 is configured to:

[0241] merging identical third texts in the at least one third text to obtain the at least one fourth text;

[0242] When the number of fourth texts is greater than 1 and a fifth text and a sixth text having a similarity greater than or equal to a fourth threshold exist in the at least one fourth text, merging the fifth text and the sixth text in the at least one fourth text to obtain the at least one second text;

[0243] If there are no two fourth texts with a similarity greater than or equal to a fourth threshold value in the at least one fourth text, the at least one fourth text is used as the at least one second text.

[0244] In combination with any embodiment of the present application, the acquiring unit 11 is configured to:

[0245] Obtaining the start time of the fifth text, the end time of the fifth text, the start time of the sixth text, and the end time of the sixth text, where the start time of the fifth text is earlier than the start time of the sixth text;

[0246] Determining a first time difference between a start time of the fifth text and an end time of the sixth text;

[0247] Determine a median of the start time and the end time of the fifth text to obtain a time median of the fifth text;

[0248] Determine a median of the start time and the end time of the sixth text to obtain a time median of the sixth text;

[0249] determining a second time difference between a time median of the fifth text and a time median of the sixth text;

[0250] When the ratio of the first time difference to the second time difference is less than or equal to a fifth threshold, the fifth text and the sixth text in the at least one fourth text are merged to obtain the at least one second text.

[0251] In an embodiment of the present application, after obtaining at least one first text in a video to be detected and the position of at least one first text in the video to be detected, the detection device divides the at least one first text into one category by classifying the first texts located at the same position in the video to be detected, thereby obtaining at least one text class. Then, by determining from the at least one text class a target text class containing a number of first texts greater than or equal to a first threshold, the target text class corresponding to the position where the subtitles appear is determined. Finally, the first text in the target text class is used as the subtitle of the video to be detected, thereby detecting the subtitles in the video to be detected and improving the efficiency of detecting subtitles.

[0252] In some embodiments, the functions or modules included in the device provided in the embodiments of the present application can be used to execute the method described in the above method embodiments. The specific implementation can refer to the description of the above method embodiments. For the sake of brevity, it will not be repeated here.

[0253] Figure 6A schematic diagram of the hardware structure of an electronic device provided in an embodiment of the present application. The electronic device 2 includes a processor 21 and a memory 22. Optionally, the electronic device 2 also includes an input device 23 and an output device 24. The processor 21, the memory 22, the input device 23 and the output device 24 are coupled via a connector, and the connector includes various interfaces, transmission lines or buses, etc., which are not limited in the embodiments of the present application. It should be understood that in each embodiment of the present application, coupling refers to mutual connection in a specific manner, including direct connection or indirect connection through other devices, for example, connection through various interfaces, transmission lines, buses, etc.

[0254] The processor 21 may include one or more processors, for example, one or more central processing units (CPUs). In the case where the processor is a CPU, the CPU may be a single-core CPU or a multi-core CPU. Alternatively, the processor 21 may be a processor group consisting of multiple CPUs, wherein the multiple processors are coupled to each other via one or more buses. Alternatively, the processor may also be other types of processors, etc., which are not limited in the embodiments of the present application.

[0255] The memory 22 can be used to store computer program instructions and various computer program codes, including program codes for executing the solution of the present application. Optionally, the memory includes, but is not limited to, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), or portable compact disc read-only memory (CD-ROM), which is used for related instructions and data.

[0256] The input device 23 is used to input data and / or signals, and the output device 24 is used to output data and / or signals. The input device 23 and the output device 24 can be independent devices or an integrated device.

[0257] It can be understood that in the embodiment of the present application, the memory 22 can be used not only to store relevant instructions, but also to store relevant data. The embodiment of the present application does not limit the specific data stored in the memory.

[0258] It is understandable that Figure 6Only a simplified design of an electronic device is shown. In actual applications, the electronic device may further include other necessary components, including but not limited to any number of input / output devices, processors, memories, etc., and all electronic devices that can implement the embodiments of the present application are within the scope of protection of the present application.

[0259] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0260] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here. Those skilled in the art will also clearly understand that the descriptions of the various embodiments of this application have different focuses. For the convenience and brevity of description, the same or similar parts may not be repeated in different embodiments. Therefore, for parts not described or not described in detail in a certain embodiment, reference can be made to the descriptions of other embodiments.

[0261] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0262] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0263] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.

[0264] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted via the computer-readable storage medium. The computer instructions can be transmitted from one website, computer, server or data center to another website, computer, server or data center via wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that includes one or more available media integrated therein. The available medium may be a magnetic medium (eg, a floppy disk, a hard disk, a magnetic tape), an optical medium (eg, a digital versatile disc (DVD)), or a semiconductor medium (eg, a solid state disk (SSD)).

[0265] Those skilled in the art will appreciate that all or part of the processes in the above-described method embodiments can be implemented by a computer program instructing related hardware to perform the processes. The program can be stored in a computer-readable storage medium, and when executed, the program can include the processes in the above-described method embodiments. The aforementioned storage medium includes various media capable of storing program code, such as read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

Claims

1. A method for detecting video subtitles, characterized in that: The method comprises: Obtaining at least one first text in a video to be detected and a position of the at least one first text in the video to be detected; When the number of first texts is greater than 1, determining, based on positions of the at least one first text in the video to be detected, an intersection-and-union ratio of areas of any two first texts in the at least one first text; a pixel coordinate system of the video to be detected includes a horizontal axis and a vertical axis, the horizontal axis is in a horizontal direction, and the vertical axis is in a vertical direction; When the intersection-and-union ratio is less than a second threshold value, determining, based on positions of the two first texts corresponding to the intersection-and-union ratio in the video to be detected, a horizontal distance between the centers of the two first texts corresponding to the intersection-and-union ratio in the horizontal direction, a vertical distance between the centers of the two first texts corresponding to the intersection-and-union ratio in the vertical direction, a maximum horizontal distance between the two first texts corresponding to the intersection-and-union ratio in the horizontal direction, and a maximum vertical distance between the two first texts corresponding to the intersection-and-union ratio in the vertical direction; When the ratio of the central horizontal distance to the maximum horizontal distance is greater than or equal to a third threshold, and the ratio of the central longitudinal distance to the maximum longitudinal axis distance is greater than or equal to the third threshold, determining that the two first texts corresponding to the intersection-over-union ratio are at the same position in the video to be detected; The at least one first text is divided into at least one text class by classifying first texts located at the same position in the video to be detected into one class; determining a text class in which the number of first texts contained in the at least one text class is greater than or equal to a first threshold as a target text class; The first text in the target text class is used as the subtitle of the video to be detected.

2. The method according to claim 1, characterized in that When the number of first texts is greater than 1, after determining the intersection-over-union ratio of areas of any two first texts in the at least one first text according to a position of the at least one first text in the video to be detected, the method further includes: When the IoU ratio is greater than or equal to a second threshold, it is determined that the two first texts corresponding to the IoU ratio are at the same position in the video to be detected.

3. The method according to claim 1, characterized in that The obtaining of at least one first text in the video to be detected includes: Acquire at least one second text in the video to be detected and a duration of the at least one second text in the video to be detected; A text having a duration within a preset interval is selected from the at least one second text to obtain the at least one first text.

4. The method according to claim 3, characterized in that The obtaining of at least one second text in the video to be detected includes: Obtaining the video to be detected; Performing text detection on the video to be detected to obtain at least one third text; The same third text in the at least one third text is merged to obtain the at least one second text.

5. The method according to claim 4, characterized in that The obtaining of the duration of the at least one second text in the video to be detected includes: Using the timestamp of the video frame corresponding to the third text as the timestamp of the third text; Determine a minimum timestamp of a third text corresponding to the second text as the start time of the second text; Determine a maximum timestamp of a third text corresponding to the second text as an end time of the second text; According to the start time of the at least one second text and the end time of the at least one second text, a duration of the at least one second text in the video to be detected is obtained.

6. The method according to claim 4, characterized in that The merging of the identical third texts in the at least one third text to obtain the at least one second text includes: merging identical third texts in the at least one third text to obtain the at least one fourth text; When the number of fourth texts is greater than 1 and a fifth text and a sixth text having a similarity greater than or equal to a fourth threshold exist in the at least one fourth text, merging the fifth text and the sixth text in the at least one fourth text to obtain the at least one second text; If there are no two fourth texts with a similarity greater than or equal to a fourth threshold value in the at least one fourth text, the at least one fourth text is used as the at least one second text.

7. The method according to claim 6, characterized in that The merging of the fifth text and the sixth text in the at least one fourth text to obtain the at least one second text includes: Obtaining the start time of the fifth text, the end time of the fifth text, the start time of the sixth text, and the end time of the sixth text, where the start time of the fifth text is earlier than the start time of the sixth text; Determining a first time difference between a start time of the fifth text and an end time of the sixth text; Determine a median of the start time and the end time of the fifth text to obtain a time median of the fifth text; Determine a median of the start time and the end time of the sixth text to obtain a time median of the sixth text; determining a second time difference between a time median of the fifth text and a time median of the sixth text; When the ratio of the first time difference to the second time difference is less than or equal to a fifth threshold, the fifth text and the sixth text in the at least one fourth text are merged to obtain the at least one second text.

8. A video subtitle detection device, characterized in that: The detection device comprises: an acquiring unit, configured to acquire at least one first text in a video to be detected and a position of the at least one first text in the video to be detected; a determining unit configured to determine, when the number of first texts is greater than one, an intersection-and-union ratio of areas of any two first texts in the at least one first text based on positions of the at least one first text in the video to be detected; wherein the pixel coordinate system of the video to be detected includes a horizontal axis and a vertical axis, wherein the horizontal axis is in a horizontal direction and the vertical axis is in a vertical direction; The determining unit is configured to determine, based on positions of the two first texts corresponding to the intersection-and-union ratio in the video to be detected, a horizontal distance between the centers of the two first texts corresponding to the intersection-and-union ratio in the horizontal direction, a vertical distance between the centers of the two first texts corresponding to the intersection-and-union ratio in the vertical direction, a maximum horizontal distance between the two first texts corresponding to the intersection-and-union ratio in the horizontal direction, and a maximum vertical distance between the two first texts corresponding to the intersection-and-union ratio in the vertical direction when the intersection-and-union ratio is less than a second threshold; The determining unit is configured to determine that the two first texts corresponding to the intersection-over-union ratios are at the same position in the video to be detected when a ratio of the central horizontal distance to the maximum horizontal distance is greater than or equal to a third threshold, and a ratio of the central longitudinal distance to the maximum longitudinal axis distance is greater than or equal to the third threshold; A dividing unit, configured to divide the at least one first text into one category by dividing the first texts located at the same position in the video to be detected into one category, thereby obtaining at least one text category; The determining unit is configured to determine a text class in which the number of first texts contained in the at least one text class is greater than or equal to a first threshold as a target text class; A processing unit is configured to use the first text in the target text class as the subtitle of the video to be detected.

9. An electronic device, characterized in that: include: A processor and a memory, the memory is used to store computer program code, the computer program code includes computer instructions, and when the processor executes the computer instructions, the electronic device executes the method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, wherein the computer program includes program instructions. When the program instructions are executed by a processor, the processor is caused to execute the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Subtitle area identification method, device and equipment and storage medium

    CN112232260A

  • Text recognition method and electronic equipment

    CN115063800A