Video subtitle detection method, device and terminal equipment
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-10-28
- Publication Date
- 2026-08-11
AI Technical Summary
[0004]本申请实施例提供了一种视频字幕检测方法、装置及终端设备,可以解决字幕检测时需要大量的训练数据集与计算资源且泛化性不足的问题
[0017]第五方面,本申请实施例提供了一种计算机程序产品,当计算机程序产品在终端设备上运行时,使得终端设备执行上述第一方面中任一项所述的视频字幕检测方法。
Smart Images

Figure CN114419082B_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of image processing technology, and in particular relates to a video subtitle detection method, apparatus and terminal equipment. Background Technology
[0002] Video subtitles generally refer to the addition of text descriptions or narration to videos to display dialogue, voice-over, and other elements. Video subtitles are often closely related to the video content, thus playing a crucial role in video content analysis and comprehension. Therefore, video subtitles can enhance the understanding and memory of ordinary viewers, introducing plot points and scenes. Subtitles also help deaf and hard-of-hearing individuals watch videos. Furthermore, subtitles make video viewing more flexible; whether in a noisy subway, train station, or quiet library, you can completely turn off the video sound and still enjoy the content through subtitles. Subtitles can also attract foreign audiences.
[0003] Existing video caption detection methods generally use deep learning methods. However, deep learning methods have many shortcomings. They require the collection of sample sets and computational resources, and model training is necessary to obtain a suitable model. If the sample set is not well selected, or if the difference between the sample to be detected and the training sample is too large, the trained model may not be suitable for various real-world situations, causing the model to fail on the data to be detected and affecting the accuracy of video caption detection. Summary of the Invention
[0004] This application provides a video subtitle detection method, apparatus, and terminal device, which can solve the problems of requiring a large training dataset and computing resources and insufficient generalization in subtitle detection.
[0005] In a first aspect, embodiments of this application provide a video subtitle detection method, including:
[0006] Obtain frame images of the subtitles in the video to be detected;
[0007] Obtain the first edge extraction features of the frame image, and obtain the first edge information of the frame image based on the first edge extraction features;
[0008] Obtain the adaptive Gaussian filter features of the frame image, and obtain the second edge information of the frame image based on the adaptive Gaussian filter features;
[0009] The intersection edge information of the first edge information and the second edge information is obtained to obtain the video subtitle region.
[0010] Secondly, embodiments of this application provide a video subtitle detection device, comprising:
[0011] The image acquisition module is used to acquire frame images of the subtitles in the video to be detected;
[0012] The first edge information module is used to obtain the first edge extraction features of the frame image and obtain the first edge information of the frame image based on the first edge extraction features.
[0013] The second edge information module is used to obtain the adaptive Gaussian filter features of the frame image and obtain the second edge information of the frame image based on the adaptive Gaussian filter features.
[0014] The cross-edge information module is used to obtain cross-edge information based on the first edge information and the second edge information to obtain the video subtitle area.
[0015] Thirdly, embodiments of this application provide a terminal device, a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the video subtitle detection method described in any one of the first aspects above.
[0016] Fourthly, embodiments of this application provide a computer-readable storage medium that stores a program that, when executed by a processor, implements the video subtitle detection method described in any one of the first aspects.
[0017] Fifthly, embodiments of this application provide a computer program product that, when run on a terminal device, causes the terminal device to execute the video subtitle detection method described in any one of the first aspects.
[0018] It is understood that the beneficial effects of the second to fifth aspects mentioned above can be found in the relevant descriptions in the first aspect mentioned above, and will not be repeated here.
[0019] The beneficial effects of this application embodiment compared with the prior art are as follows: This application embodiment provides a video subtitle detection method, which involves acquiring a frame image of the video subtitle to be detected; acquiring first edge extraction features of the frame image, and acquiring first edge information of the frame image based on the first edge extraction features; acquiring adaptive Gaussian filter features of the frame image, and acquiring second edge information of the frame image based on the adaptive Gaussian filter features; and acquiring cross-edge information based on the first edge information and the second edge information to obtain the video subtitle region. The video subtitle detection method provided by this application embodiment does not require a large number of training samples, has low complexity, consumes little computational resources, and can effectively find subtitle regions in videos. Attached Figure Description
[0020] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0021] Figure 1 This is a schematic diagram of a video subtitle detection method provided in an embodiment of this application;
[0022] Figure 2 The image of the video subtitles to be detected in this embodiment of the application;
[0023] Figure 3 This application embodiment illustrates the acquisition of first edge information in an image based on first edge extraction features;
[0024] Figure 4 This application's embodiment illustrates the acquisition of second edge information in an image based on adaptive Gaussian filtering features;
[0025] Figure 5 This application's embodiment illustrates the acquisition of cross information in an image;
[0026] Figure 6 This is a schematic diagram of the video subtitle detection device provided in the embodiments of this application;
[0027] Figure 7 This is a schematic diagram of the structure of the terminal device provided in the embodiments of this application. Detailed Implementation
[0028] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of this application. However, those skilled in the art will understand that this application may also be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods have been omitted so as not to obscure the description of this application with unnecessary detail.
[0029] It should be understood that, when used in this application specification and the appended claims, the term "comprising" indicates the presence of the described features, integrals, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or a collection thereof.
[0030] It should also be understood that the term “and / or” as used in this application specification and the appended claims means any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.
[0031] As used in this application specification and the appended claims, the term "if" may be interpreted, depending on the context, as "when," "once," "in response to determination," or "in response to detection." Similarly, the phrase "if determination" or "if the described condition or event is detected" may be interpreted, depending on the context, as "once determination," "in response to determination," "once the described condition or event is detected," or "in response to the detection of the described condition or event."
[0032] Furthermore, in the description of this application and the appended claims, the terms "first," "second," "third," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.
[0033] References to "one embodiment" or "some embodiments" as described in this specification mean that one or more embodiments of this application include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "comprising," "including," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.
[0034] Figure 1 This is a schematic diagram of a video subtitle detection method provided in an embodiment of this application, as shown below. Figure 1 As shown, the method includes:
[0035] Step S101: Obtain frame images of the video subtitles to be detected, for example... Figure 2 As shown;
[0036] Since video subtitles often remain on the screen for more than one second, and each second of video contains 24 or more frames, it is unnecessary to process every single frame. To ensure real-time detection while maintaining detection effectiveness, it is not necessary to select all frames; instead, a subset of frames can be chosen. The number of selected frames is denoted as L, and the range can be 5 to 15 frames, with 5 to 10 frames typically yielding better results.
[0037] Step S102: Obtain the first edge extraction features of the frame image, and obtain the first edge information of the frame image based on the first edge extraction features, specifically including:
[0038] Edge extraction features can include first-order edge extraction features and second-order edge extraction features. First-order edge extraction features include Roberts features, Sobel features, Prewitt features, Kirsch features, Robinson features, etc., while second-order edge extraction features include LOG (Laplacian of Gaussian) features, Canny features, Marr-Hildreth features, etc.
[0039] Step S1021: Obtain the first edge extraction feature from all frame images, denoted as Edge1_feature_image_k. Edge1_feature_image_k = f(image_k)
[0040] Where f(.) represents the algorithm for obtaining the first edge extraction feature. This application does not limit the specific acquisition algorithm, such as using Laplacian of Gaussian, k represents the frame number of the frame image, k is a natural number, and k≤L.
[0041] By obtaining the first edge extraction feature image of a single frame through step S1021 in the embodiments of this application, the first edge extraction feature of each pixel (i, j) in each first edge extraction feature image of a single frame can be determined, where (i, j) are the coordinates of the pixel, i represents the horizontal coordinate and j represents the vertical coordinate.
[0042] Step S1022: Accumulate the first edge extraction features of all frame images to obtain the first edge extraction feature accumulation image.
[0043] The sum of the edge extraction features of each pixel in the cumulative image corresponds to the sum of the edge extraction features of pixels at the same position in all frames, denoted as SUM_edge1(i,j). The formula for calculating SUM_edge1(i,j) is as follows:
[0044]
[0045] By obtaining the cumulative image of the first edge extraction features of all frames through step S1022 in the embodiments of this application, the sum of the first edge extraction features of each pixel (i, j) can be determined.
[0046] Step S1023: Perform a binarization operation on the multi-frame first edge extraction feature accumulation image to obtain a first binarized image, such as... Figure 3 As shown.
[0047] The first edge extraction feature can express the edge information in the image. Subtitles will last for a period of time in the video. Subtitles usually appear in the same position in multiple frames of the image. Therefore, the first edge extraction feature value corresponding to the subtitle and the sum of the first edge extraction feature values will be relatively large.
[0048] If the pixel value in the accumulated image of the first edge extraction feature is greater than or equal to the first feature threshold T log If pixel (i, j) is considered to belong to a stationary edge, its value is updated to 1, and this pixel corresponds to the video subtitle area. Otherwise, if pixel (i, j) is considered not to belong to a stationary edge, its value is updated to 0, where (i, j) are the coordinates of the pixel. It can be understood that if pixel (i, j) belongs to a stationary edge, its value can be updated to 0; if pixel (i, j) does not belong to a stationary edge, its value can be updated to 1. Whether each pixel in the first binarized image belongs to a stationary edge can be represented by a matrix, denoted as Edge1(i, j), where the information that a pixel in the first binarized image belongs to a stationary edge is denoted as the first edge information, and the stationary edge can optionally be a text edge.
[0049]
[0050] The above SUM_edge1(i,j) and Edge(i,j) are both matrices with M rows and N columns, where M and N are the resolution of the image (M*N), and L is the total number of frames of the image.
[0051] The area initially identified as subtitles in the first binarized image can be denoted as the first video subtitle area, which may also include some noise.
[0052] In one possible implementation, morphological operations, such as erosion, dilation, and hole filling, are performed on the first binarized image to achieve edge smoothing and fine line removal, thereby removing the defects caused by T. log To mitigate misjudgments caused by inaccuracies, we can reduce interference from complex backgrounds in caption region detection. For example, we can use 5x5 rectangular, cross-shaped, or elliptical structuring elements on the binarized image and perform morphological operations.
[0053] Through the embodiments of this application, edge extraction features in frame images can be used to effectively identify edge information in an image.
[0054] Step S103: Obtain the adaptive Gaussian filter features of the frame image, and obtain the second edge information of the frame image based on the adaptive Gaussian filter features, specifically including:
[0055] Step S1031: Obtain the adaptive Gaussian filter features (ADA) of all frame images. gau ;
[0056] Step S1031 in this embodiment of the application can obtain a single-frame adaptive Gaussian filter feature image, and determine the adaptive Gaussian filter feature ADA of each pixel (i, j) in the adaptive Gaussian filter feature image. gau .
[0057] Step S1032: Accumulate the adaptive Gaussian filter features (ADA) of all frames. gau The adaptive Gaussian filter feature accumulation image is obtained.
[0058] The value of each pixel in the adaptive Gaussian filter feature accumulation image corresponds to the sum of the adaptive Gaussian filter features of the pixels at the same position in all frames of the image, denoted as Sum_ada_gau(i,j).
[0059] Step S1022 in this embodiment of the application can be used to obtain an adaptive Gaussian filter feature accumulation image and determine the sum of adaptive Gaussian filter features for each pixel (i, j).
[0060] Step S1033: Perform a binarization operation on the adaptive Gaussian filter feature accumulation image to obtain a second binarized image, such as... Figure 4 As shown.
[0061] If the pixel value in the adaptive Gaussian filter feature accumulation image is greater than or equal to the second feature threshold T ada_gau If the pixel is considered to be on a stationary edge, its value is updated to 1, and this pixel corresponds to the video subtitle area. Otherwise, if the pixel is not on a stationary edge, its value is updated to 0. It is understandable that if a pixel is on a stationary edge, its value can be updated to 0; if a pixel is not on a stationary edge, its value can be updated to 1.
[0062] Whether each pixel in the second binarized image belongs to a stationary edge can be represented by a matrix, denoted as Edge_ada_gau(i,j), where the information that a pixel in the second binarized image belongs to a stationary edge is denoted as the second edge information.
[0063]
[0064] Similarly, ADA gau Both Edge_ada_gau(i,j) are matrices with M rows and N columns, where M and N are the resolutions of the images (M*N).
[0065] The area initially identified as subtitles in the second binarized image can be denoted as the second video subtitle area, which may also include some noise.
[0066] In one possible implementation, morphological operations, such as erosion, dilation, and hole filling, are performed on the second binarized image obtained in step S1033 to achieve edge smoothing and fine line removal, thereby removing the artifacts caused by the threshold T. ada_gau Some misjudgments were caused by inaccurate reasons.
[0067] Through the embodiments of this application, edge information in an image can be effectively identified based on the adaptive Gaussian filtering features in the frame image.
[0068] There is no requirement for the order of steps S102 and S103.
[0069] Step S104: Obtain the intersection edge information based on the first edge information and the second edge information to obtain the video subtitle region.
[0070] The intersection edge information is obtained by performing a bitwise AND operation on the first edge information and the second edge information, specifically including:
[0071] Perform an AND operation on the first and second binarized images, specifically on Edge1(i,j) and Edge_ada_gau(i,j), to obtain the third binarized image, as follows. Figure 5 As shown, the third binarized image includes a set of pixels that are detected as stationary edges in both the first edge extraction feature and the adaptive Gaussian filter feature, that is, a set of pixels that are detected as stationary edges in both the first edge information and the second edge information, denoted as the cross edge information.
[0072] The detected video subtitle region is obtained based on the cross-edge information. The cross-edge information corresponds to the common area of the first video subtitle region and the second video subtitle region, and the common area is used as the detected video subtitle region. This application does not specifically limit how the detected video subtitle region is obtained based on the cross-edge information Cross_edge(i, j).
[0073] Where Cross_edge(i,j) represents the cross edge information corresponding to pixel (i,j), and Cross_edge(i,j) is a matrix.
[0074]
[0075] In one possible implementation, to remove some noise and make the detected video subtitle region more accurate, falsely detected cross-edge information can be further removed. For example, connected component area information can be extracted from the third binarized image; connected components with excessively large or small areas are considered falsely detected. For example, if the connected component area is less than a first threshold T... area1 Or the area of the connected region is greater than the second threshold T area2 The connected component is considered a falsely detected connected component. The falsely detected connected component corresponds to the falsely detected intersection edge information. The falsely detected connected component is removed to obtain the final video subtitle region, where the first threshold is less than the second threshold. Alternatively, morphological operations are performed on the third binarized image obtained in step S104 to remove the falsely detected intersection edge information, such as erosion, dilation, and hole filling, to achieve edge smoothing and fine line removal, thereby eliminating some false judgments.
[0076] The video subtitle detection method provided in this application does not require a large number of training samples, has low complexity and low computational resource consumption, and can effectively find subtitle regions in videos.
[0077] See Figure 6 This is a schematic diagram of a video subtitle detection device according to an embodiment of this application. For ease of explanation, only the parts related to the embodiment of this application are shown, including:
[0078] Image acquisition module 61 is used to acquire frame images of the subtitles in the video to be detected;
[0079] The first edge information module 62 is used to obtain the first edge extraction features of the frame image and obtain the first edge information of the frame image based on the first edge extraction features;
[0080] The second edge information module 63 is used to obtain the adaptive Gaussian filter features of the frame image and obtain the second edge information of the frame image based on the adaptive Gaussian filter features.
[0081] The cross-edge information module 64 is used to obtain cross-edge information based on the first edge information and the second edge information to obtain the video subtitle area.
[0082] The noise reduction module 65 is used to remove falsely detected cross-edge information.
[0083] The first edge information module 62 also includes:
[0084] Obtain the first edge extraction features of the frame image, and determine the first edge extraction features of each pixel (i, j) in the first edge extraction feature image of a single frame;
[0085] The first edge extraction features of all frames are summed to obtain the first edge extraction feature summed image;
[0086] The multi-frame first edge extraction feature accumulation image is binarized to obtain a first binarized image. Whether each pixel in the first binarized image belongs to a stationary edge can be represented by a matrix, denoted as Edge1(i,j). The information that the pixel in the first binarized image belongs to a stationary edge is denoted as the first edge information.
[0087] If the value of a pixel in the accumulated image of the first edge extraction features is greater than or equal to the first feature threshold, then the pixel belongs to a stationary edge, and the value of the pixel is updated to 1; otherwise, the value of the pixel is updated to 0. It can be understood that if the pixel belongs to a stationary edge, the value of the pixel can also be updated to 0; if the pixel does not belong to a stationary edge, the value of the pixel can also be updated to 1.
[0088] The first binarized image is initially identified as the subtitle region, which can be denoted as the first video subtitle region. Of course, it may also include some noise.
[0089] In one possible implementation, morphological operations, such as erosion, dilation, and hole filling, are performed on the first binarized image to achieve edge smoothing and fine line removal, thereby removing the defects caused by T. log To mitigate misjudgments caused by inaccuracies, we can reduce interference from complex backgrounds in caption region detection. For example, we can use 5x5 rectangular, cross-shaped, or elliptical structuring elements on the binarized image and perform morphological operations.
[0090] The second edge information module 63 also includes:
[0091] Obtain the adaptive Gaussian filter features of all frame images;
[0092] The adaptive Gaussian filter features of all frames are summed to obtain the adaptive Gaussian filter feature summed image;
[0093] The adaptive Gaussian filter feature accumulation image is binarized to obtain a second binarized image. Whether each pixel in the second binarized image belongs to a stationary edge can be represented by a matrix, denoted as Edge_ada_gau(i,j). The information that a pixel in the second binarized image belongs to a stationary edge is denoted as the second edge information.
[0094] If the value of a pixel in the adaptive Gaussian filter feature accumulation image is greater than or equal to the second feature threshold, then the pixel belongs to a stationary edge, and its value is updated to 1; otherwise, the pixel value is updated to 0. It can be understood that if a pixel belongs to a stationary edge, its value can also be updated to 0; if a pixel does not belong to a stationary edge, its value can also be updated to 1.
[0095] The area initially identified as subtitles in the second binarized image can be denoted as the second video subtitle area, which may also include some noise.
[0096] In one possible implementation, morphological operations, such as erosion, dilation, and hole filling, are performed on the second binarized image obtained in step S1033 to achieve edge smoothing and fine line removal, thereby removing the artifacts caused by the threshold T. ada_gau Some misjudgments were caused by inaccurate reasons.
[0097] The cross-edge information module 64 also includes:
[0098] The intersection edge information is obtained by performing an AND operation on the first edge information and the second edge information. Specifically, this includes performing an AND operation on the first binarized image and the second binarized image, that is, performing an AND operation on Edge1(i,j) and Edge_ada_gau(i,j), to obtain a third binarized image. The third binarized image includes a set of pixels that are detected as stationary edges by both the first edge extraction feature and the adaptive Gaussian filter feature, that is, a set of pixels that are detected as stationary edges by both the first edge information and the second edge information, denoted as the intersection edge information.
[0099] The detected video subtitle region is obtained based on the cross-edge information. The cross-edge information corresponds to the common area of the first video subtitle region and the second video subtitle region, and the common area is used as the detected video subtitle region. This application does not specifically limit how the detected video subtitle region is obtained based on the cross-edge information Cross_edge(i,j).
[0100] In one possible implementation, a denoising module may also be included. This module further extracts connected component area information from the third binarized image, classifying connected components with excessively large or small areas as falsely detected connected components. For example, if the connected component area is less than a first threshold T... area1 Or the area of the connected region is greater than the second threshold T area2 The connected component is considered a falsely detected connected component. The falsely detected connected component corresponds to the falsely detected intersection edge information. The falsely detected connected component is removed to obtain the final video subtitle region, where the first threshold is less than the second threshold.
[0101] Alternatively, morphological operations can be performed on the third binarized image to remove falsely detected intersecting edge information, such as erosion, dilation, and hole filling, to achieve edge smoothing and fine line removal, thereby eliminating some false positives. For example, morphological operations can be performed on the binarized image using 5x5 rectangular, cross-shaped, or elliptical structuring elements.
[0102] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional modules is merely an example. In practical applications, the above functions can be assigned to different functional units or modules as needed, that is, the internal structure of the mobile terminal can be divided into different functional units or modules to complete all or part of the functions described above. The functional modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. Furthermore, the specific names of the functional modules are only for easy differentiation and are not intended to limit the scope of protection of this application. The specific working process of the modules in the mobile terminal can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0103] Figure 7 This is a schematic diagram of a terminal device provided in an embodiment of the present invention, as shown below. Figure 7 As shown, the terminal device 7 in this embodiment includes: a processor 70, a memory 71, and a computer program 72 stored in the memory 71 and executable on the processor 70. When the processor 70 executes the computer program 72, it implements the steps in the above-described video subtitle detection method embodiment, for example... Figure 1 Steps S101 to S105 are shown. Alternatively, when the processor 70 executes the computer program 72, it implements the functions of each module / unit in the above-described video subtitle detection device embodiment, for example... Figure 6 The functions of modules 61 to 65 are shown.
[0104] For example, the computer program 72 may be divided into one or more modules / units, which are stored in the memory 71 and executed by the processor 70 to complete the present invention. The one or more modules / units may be a series of computer program instruction segments capable of performing a specific function, which describe the execution process of the computer program 72 in the terminal device 7.
[0105] The terminal device 7 can be a desktop computer, laptop, handheld computer, or cloud server, etc. The terminal device may include, but is not limited to, a processor 70 and a memory 71. Those skilled in the art will understand that... Figure 7 This is merely an example of terminal device 7 and does not constitute a limitation on terminal device 7. It may include more or fewer components than shown, or combine certain components, or different components. For example, the terminal device may also include input / output devices, network access devices, buses, etc.
[0106] The processor 70 may be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor.
[0107] The memory 71 can be an internal storage unit of the terminal device 7, such as a hard disk or memory of the terminal device 7. The memory 71 can also be an external storage device of the terminal device 7, such as a plug-in hard disk, smart media card (SMC), secure digital card (SD), flash card, etc., equipped on the terminal device 7. Furthermore, the memory 71 can include both internal and external storage units of the terminal device 7. The memory 71 is used to store the computer program and other programs and data required by the terminal device. The memory 71 can also be used to temporarily store data that has been output or will be output.
[0108] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps described in the various method embodiments above.
[0109] This application provides a computer program product that, when run on a mobile terminal, enables the mobile terminal to implement the steps described in the above-described method embodiments.
[0110] Those skilled in the art will understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is merely an example. In practical applications, the functions described above can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. Furthermore, the specific names of the functional units and modules are only for easy differentiation and are not intended to limit the scope of protection of this invention. The specific working process of the units and modules in the above system can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0111] In the above embodiments, the descriptions of each embodiment have different focuses. For parts that are not described in detail or recorded in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0112] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.
[0113] In the embodiments provided by this invention, it should be understood that the disclosed apparatus / terminal devices and methods can be implemented in other ways. For example, the apparatus / terminal device embodiments described above are merely illustrative. For instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.
[0114] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0115] Furthermore, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0116] If the integrated module / unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments can also be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording media, USB flash drives, portable hard drives, magnetic disks, optical disks, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc. It should be noted that the content included in the computer-readable medium can be appropriately added or removed according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, computer-readable media do not include electrical carrier signals and telecommunication signals.
[0117] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.
Claims
1. A video subtitle detection method, characterized in that, include: Obtain frame images of the subtitles in the video to be detected; Obtain the first edge extraction features of the frame image, and obtain the first edge information of the frame image based on the first edge extraction features; Obtain the adaptive Gaussian filter features of the frame image, and obtain the second edge information of the frame image based on the adaptive Gaussian filter features; Based on the first edge information and the second edge information, the cross edge information is obtained to determine the common area, and the video subtitle area is obtained; The process of obtaining first edge extraction features of frame images and obtaining first edge information of frame images based on the first edge extraction features includes: obtaining first edge extraction features of all frame images; accumulating the first edge extraction features of all frame images to obtain a first edge extraction feature accumulation image; performing a binarization operation on the first edge extraction feature accumulation image to obtain a first binarized image, wherein the information of pixels in the first binarized image belonging to stationary edges is recorded as the first edge information.
2. The video subtitle detection method as described in claim 1, characterized in that, Binarization is performed on the first edge-extracted feature accumulation image to obtain a first binarized image, including: If the value of a pixel in the first edge extraction feature accumulation image is greater than or equal to the first feature threshold, then the pixel belongs to a stationary edge, and the value of the pixel is updated to 1; otherwise, the value of the pixel is updated to 0.
3. The video subtitle detection method as described in claim 1 or 2, characterized in that, Obtaining the adaptive Gaussian filter features of the frame image, and obtaining the second edge information of the frame image based on the adaptive Gaussian filter features, including: Obtain the adaptive Gaussian filter features of all frame images; The adaptive Gaussian filter features of all frames are summed to obtain the adaptive Gaussian filter feature summed image; The adaptive Gaussian filter feature accumulation image is binarized to obtain a second binarized image. The information of pixels belonging to stationary edges in the second binarized image is recorded as the second edge information.
4. The video subtitle detection method as described in claim 3, characterized in that, Binarizing the adaptive Gaussian filter feature accumulation image to obtain a second binarized image includes: If the value of a pixel in the adaptive Gaussian filter feature accumulation image is greater than or equal to the second feature threshold, then the pixel belongs to a stationary edge and the value of the pixel is updated to 1; otherwise, the value of the pixel is updated to 0.
5. The video subtitle detection method as described in claim 4, characterized in that, Based on the first edge information and the second edge information, the intersection edge information is obtained to obtain the video subtitle region, including: A third binarized image is obtained by performing an AND operation on the first edge information and the second edge information. The third binarized image includes a set of pixels in the first edge information and the second edge information that are detected as stationary edges, denoted as the intersection edge information. The video subtitle region is obtained based on the cross-edge information.
6. The video subtitle detection method as described in claim 5, characterized in that, After obtaining the intersection edge information based on the first edge information and the second edge information, the method further includes: Remove falsely detected cross-edge information.
7. The video subtitle detection method as described in claim 6, characterized in that, Remove falsely detected cross-edge information, including: Extract connected components from the third binarized image; If the area of the connected component is less than a first threshold or the area of the connected component is greater than a second threshold, then the connected component is identified as falsely detected cross edge information, wherein the first threshold is less than the second threshold. Remove the connected components; or, Morphological operations are performed on the third binarized image to remove falsely detected cross-edge information.
8. A video subtitle detection device, characterized in that, include: The image acquisition module is used to acquire frame images of the subtitles in the video to be detected; The first edge information module is used to obtain the first edge extraction features of the frame image and obtain the first edge information of the frame image based on the first edge extraction features. The second edge information module is used to obtain the adaptive Gaussian filter features of the frame image and obtain the second edge information of the frame image based on the adaptive Gaussian filter features. The cross-edge information module is used to obtain cross-edge information based on the first edge information and the second edge information to determine the common area and obtain the video subtitle area; Specifically, the first edge information module is used to: acquire the first edge extraction features of all frame images; The first edge extraction features of all frames are accumulated to obtain a first edge extraction feature accumulation image; the first edge extraction feature accumulation image is binarized to obtain a first binarized image, and the information of pixels in the first binarized image belonging to stationary edges is recorded as the first edge information.
9. A terminal device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the method as described in any one of claims 1 to 7.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Method and device for extracting video subtitles
CN102915438A
Automatic detecting and recognizing method of scroll captions in videos
CN104244073A
Method and device for eliminating image subtitles
CN110942420A
Method and apparatus to detect subtitle
KR101390561B1