A subtitle detection method and related device

By obtaining the subtitle characteristics of the target video and sample video frames, and automatically detecting whether the subtitle display is abnormal by character structure comparison, the problem of time-consuming and labor-consuming manual detection in the prior art is solved, and efficient automatic subtitle detection is achieved.

CN113822273BActive Publication Date: 2025-07-11TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110713663.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-06-25
Publication Date
2025-07-11
Estimated Expiration
2041-06-25

AI Technical Summary

Technical Problem

In the prior art, subtitle inspection requires manual inspection, which consumes a lot of manpower and material resources and has a long detection cycle.

Method used

By obtaining the subtitle characteristics of the target video and sample video frames, using the character structure feature ratio to measure the subtitles and sample subtitles, it will automatically detect whether the subtitle display is abnormal.

Benefits of technology

It realizes automated subtitle detection without manual participation, improves detection efficiency and reduces resource utilization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113822273B_ABST
    Figure CN113822273B_ABST
Patent Text Reader

Abstract

Embodiments of the present application disclose a subtitle detection method and related devices. When performing anomaly detection on the subtitle display function of a program under test, subtitle detection can be performed through the character structure of the subtitle. Since the feature expression of the character structure is simple, and it can clearly reflect the subtitle structure form in the video frame, and there are obvious structural differences between different characters in the subtitle, therefore, by comparing the sample subtitle features with the features of the subtitle under test, it can be accurately and quickly determined whether the subtitle content displayed in the video frame under test is correct, whether the position is offset, etc., enabling the anomaly detection of the subtitle display of the program under test to be automated and no longer requiring manual participation. Moreover, during the detection process, there is no need to identify and process the complex semantic information of the subtitle, and the detection can be completed through simplified subtitle features, reducing the resource occupancy of automated detection and improving the detection efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of video technology, and in particular, to a subtitle detection method and related device. Background Art

[0002] Subtitles refer to the character content used to reflect the audio information in a video, which can help users watching the video better understand the information described in the video.

[0003] In order to improve the user's viewing experience, the subtitle content corresponding to the audio information of the video can be automatically displayed through a related program in the video for users to watch. In order to ensure the accuracy of subtitle display in the video by the program, in the related art, it is necessary to manually check the subtitle content displayed by the program, resulting in a large consumption of human and material resources and a long detection cycle. Summary of the Invention

[0004] To solve the above technical problems, the present application provides a subtitle detection method, which can more accurately determine whether there is an abnormality in the subtitle display result of the program to be tested, thereby eliminating the need for manual detection and improving the detection efficiency.

[0005] The embodiments of the present application disclose the following technical solutions:

[0006] In a first aspect, the embodiments of the present application disclose a subtitle detection method, the method comprising:

[0007] Obtain a target video and sample subtitle features corresponding to a sample video frame, where the sample subtitle features are used to identify the character structure of the sample subtitle in the sample video frame, the sample video frame is a video frame in a sample video, the sample video is the target video that displays the sample subtitle, and the sample video frame corresponds to a target video frame in the target video;

[0008] According to the audio information in the target video, display the subtitle to be tested corresponding to the audio information in the target video through the program to be tested to obtain a video to be tested;

[0009] Determine subtitle features to be tested according to the video frame to be tested corresponding to the target video frame in the video to be tested, where the subtitle features to be tested are used to identify the character structure of the subtitle to be tested in the video frame to be tested;

[0010] Determine whether there is an abnormality in the subtitle display of the program to be tested in the video frame to be tested according to the sample subtitle features and the subtitle features to be tested.

[0011] In a second aspect, the embodiments of the present application disclose a subtitle detection device, the device comprising a first acquisition unit, a display unit, a first determination unit and a second determination unit:

[0012] The first acquisition unit is configured to acquire a target video and sample subtitle features corresponding to a sample video frame, where the sample subtitle features are used to identify the character structure of the sample subtitle in the sample video frame, the sample video frame is a video frame in a sample video, the sample video is the target video that shows the sample subtitle, and the sample video frame corresponds to a target video frame in the target video;

[0013] The display unit is configured to, according to the audio information in the target video, display a to-be-tested subtitle corresponding to the audio information in the target video through a to-be-tested program, so as to obtain a to-be-tested video;

[0014] The first determination unit is configured to determine to-be-tested subtitle features according to a to-be-tested video frame corresponding to the target video frame in the to-be-tested video, where the to-be-tested subtitle features are used to identify the character structure of the to-be-tested subtitle in the to-be-tested video frame;

[0015] The second determination unit is configured to determine whether an abnormality occurs in the subtitle display of the to-be-tested program in the to-be-tested video frame according to the sample subtitle features and the to-be-tested subtitle features.

[0016] In a third aspect, an embodiment of the present application discloses a computer device, where the device includes a processor and a memory:

[0017] The memory is configured to store program code and transmit the program code to the processor;

[0018] The processor is configured to execute the subtitle detection method according to any one of the items in the first aspect according to the instructions in the program code.

[0019] In a fourth aspect, an embodiment of the present application discloses a computer-readable storage medium, where the computer-readable storage medium is configured to store a computer program, and the computer program is configured to execute the subtitle detection method according to any one of the items in the first aspect.

[0020] As can be seen from the above technical solution, when performing anomaly detection on the subtitle display function of the program under test, the sample subtitle features corresponding to the sample video frames in the target video without displayed subtitles and the sample video can be obtained. The sample video is the target video with accurate subtitles displayed, and the sample subtitle features identify the character structures of the sample subtitles in the corresponding sample video frames. According to the audio information in the target video, the program under test displays the subtitles under test corresponding to the audio information in the target video to obtain the video under test, and determines the subtitles under test features of the video frames under test from the video under test. The video frames under test and the sample video frames both correspond to the same target video frame in the target video. The subtitles under test features can identify the character structures of the subtitles under test in the video frames under test. Since the feature expression of the character structure is simple, and it can clearly reflect the subtitle structure form in the video frame, and there are obvious structural differences between different characters in the subtitle, the comparison between the sample subtitle features and the subtitles under test features can accurately and quickly determine whether the subtitle content displayed in the video frames under test is correct and whether the position is offset. This enables the automatic implementation of anomaly detection for the subtitle display of the program under test, eliminating the need for manual participation. Moreover, during the detection process, there is no need to identify and process the complex semantic information of the subtitles, and the detection can be completed through the simplified subtitle features, reducing the resource occupancy of automatic detection and improving the detection efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0022] Figure 1 Schematic diagram of a subtitle detection method in an actual application scenario provided by an embodiment of the present application;

[0023] Figure 2 Flowchart of a subtitle detection method provided by an embodiment of the present application;

[0024] Figure 3 Schematic diagram of a subtitle detection method provided by an embodiment of the present application;

[0025] Figure 4 Flowchart of a subtitle detection method in an actual application scenario provided by an embodiment of the present application;

[0026] Figure 5 Schematic diagram of a subtitle detection method provided by an embodiment of the present application;

[0027] Figure 6Schematic diagram of a subtitle detection method provided by an embodiment of the present application;

[0028] Figure 7 Structural block diagram of a subtitle detection device provided by an embodiment of the present application;

[0029] Figure 8 Structural diagram of a computer device provided by an embodiment of the present application;

[0030] Figure 9 Structural diagram of a server provided by an embodiment of the present application. Detailed implementation manners

[0031] The embodiments of the present application will be described below with reference to the accompanying drawings.

[0032] To facilitate users to better view the video content, some video applications can provide the function of identifying the audio information contained in the video and automatically displaying the subtitles corresponding to the audio information. In the related art, in order to ensure the normal operation of the subtitle recognition function, relevant technicians need to conduct a detailed inspection of the background code of the program that implements this function, and need to visually verify the subtitle recognition results. On the one hand, it is difficult to perform relatively intuitive anomaly detection by checking the background code. On the other hand, visual recognition requires a large amount of human and material resources, so the detection efficiency is low.

[0033] To solve the above technical problems, the present application provides a subtitle detection method. The processing device can compare the character structure of the sample subtitle corresponding to the accurate subtitle display result with the character structure of the subtitle to be tested displayed by the program to be tested. Based on the fact that the character structure can relatively accurately represent the characters, this detection method can relatively accurately determine whether there is an anomaly in the subtitle display result of the program to be tested, thus eliminating the need for manual detection and improving the detection efficiency.

[0034] It can be understood that this method can be applied to a processing device, which is a processing device capable of performing subtitle detection, such as a terminal device or a server with subtitle detection function. This method can be independently executed by the terminal device or the server, or can be applied to a network scenario where the terminal device and the server communicate, and is executed in cooperation with the terminal device and the server. Among them, the terminal device can be a device such as a computer or a mobile phone. The server can be understood as an application server or a Web server. In actual deployment, the server can be an independent server or a cluster server.

[0035] In addition, this application also relates to Artificial Intelligence (AI) technology. Artificial Intelligence is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, Artificial Intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. Artificial Intelligence also studies the design principles and implementation methods of various intelligent machines to enable machines to have the functions of perception, reasoning, and decision-making.

[0036] Artificial Intelligence technology is an interdisciplinary subject with a wide range of fields involved, including both hardware-level and software-level technologies. The basic technologies of Artificial Intelligence generally include technologies such as sensors, dedicated Artificial Intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. The software technologies of Artificial Intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning, autonomous driving, and intelligent transportation. This application mainly relates to natural language processing technology, speech technology, and computer vision technology among them.

[0037] Natural Language Processing (NLP) is an important direction in the fields of computer science and Artificial Intelligence. It studies various theories and methods that can achieve effective communication between humans and computers in natural language. Natural Language Processing is a science that integrates linguistics, computer science, and mathematics. Therefore, the research in this field will involve natural language, that is, the language people use in daily life, so it has a close connection with the research of linguistics. Natural Language Processing technology usually includes technologies such as text processing, semantic understanding, machine translation, robot question answering, and knowledge graphs.

[0038] The key technologies of Speech Technology include automatic speech recognition technology, speech synthesis technology, and voiceprint recognition technology. Enabling computers to listen, see, speak, and feel is the future development direction of human-computer interaction, and among them, speech has become one of the most promising human-computer interaction methods in the future.

[0039] Computer Vision Technology (CV) Computer vision is a science that studies how to enable machines to "see". More specifically, it refers to machine vision that uses cameras and computers to replace human eyes for tasks such as object recognition, tracking, and measurement, and further performs image processing to make the computer-processed images more suitable for human eye observation or transmission to instrument detection. As a scientific discipline, computer vision researches related theories and technologies and attempts to build artificial intelligence systems that can obtain information from images or multi-dimensional data. Computer vision technology usually includes image processing, image recognition, image semantic understanding, image retrieval, OCR, video processing, video semantic understanding, video content / behavior recognition, 3D object reconstruction, 3D technology, virtual reality, augmented reality, simultaneous localization and mapping, autonomous driving, intelligent transportation, etc. technologies, and also includes common biometric recognition technologies such as face recognition and fingerprint recognition.

[0040] For example, in the embodiments of the present application, the subtitle content corresponding to the audio information can be recognized through speech technology and natural language processing technology, and the subtitles in the video frames can be recognized and feature-extracted through computer vision technology, etc.

[0041] To facilitate the understanding of the technical solution of the present application, the subtitle detection method provided in the embodiments of the present application will be introduced below in combination with an actual application scenario.

[0042] See Figure 1 , Figure 1 which is a schematic diagram of a subtitle detection method in an actual application scenario provided in the embodiments of the present application. In this actual application scenario, the processing device can be the terminal device 101, and the terminal device 101 can be a mobile phone or the like used by a tester when testing a program under test.

[0043] The program under test in the terminal device 101 can automatically display corresponding subtitles in the video based on the audio information of the video. To detect whether the subtitle display function of the program under test is abnormal, the terminal device 101 can first obtain the sample subtitle features corresponding to the target video and the sample video frame. The sample video frame is a video frame in the sample video, the sample video is the target video that shows the sample subtitles, the sample subtitles are the correct subtitle content corresponding to the audio information in the target video, the sample video frame corresponds to the target video frame in the target video, and the target video is a video that does not include subtitle content but has audio information.

[0044] The terminal device can, according to the audio information in the target video, display the corresponding subtitle to be tested in the target video through the program to be tested, so as to obtain the video to be tested. The terminal device 101 can determine the video frame to be tested corresponding to the target video frame from the video frames of the video to be tested. That is, relative to the target video frame, the sample video frame and the video frame to be tested are respectively the target video frame with accurate subtitle content displayed and the target video frame with subtitle content displayed through the program to be tested. The processing device can determine the subtitle feature to be tested according to the video frame to be tested, and the subtitle feature to be tested is used to identify the character structure of the subtitle to be tested in the video frame to be tested. Thus, by comparing the sample subtitle feature and the subtitle feature to be tested, the processing device can detect whether the sample subtitle and the subtitle to be tested are consistent from the dimension of the character structure, and obtain the detection result of the program to be tested. If the matching degree between the sample subtitle feature and the subtitle feature to be tested is relatively high, it indicates that the sample subtitle and the subtitle to be tested are consistent, and the program to be tested can display the subtitle accurately.

[0045] As Figure 1 shown, since the subtitle content of the sample subtitle is "This is the sea view we see", and the subtitle content of the subtitle to be tested is "At this time, the sea view we see", there are two different characters in the sample subtitle and the subtitle to be tested. At this time, there will be a relatively obvious parameter gap between the sample subtitle feature and the subtitle feature to be tested. Therefore, the terminal device 101 can determine that there is an abnormality in the subtitle display of the program to be tested in the video frame to be tested according to the parameter gap. Since the character structure is a relatively distinct character feature of the character, and the character structure gap between different characters is relatively obvious, the credibility of the detection result of the program to be tested determined by this method is relatively high. Based on this, the terminal device 101 can perform accurate program detection through a reasonable subtitle detection method without manual participation, reducing the demand for manpower. At the same time, in the above detection process, it is not necessary to identify and process the complex semantic information of the subtitle, and the character structure is a relatively easy-to-obtain character feature, so the detection efficiency is improved.

[0046] Next, a subtitle detection method provided by an embodiment of the present application will be introduced in conjunction with the accompanying drawings.

[0047] See Figure 2 , Figure 2 which is a flowchart of a subtitle detection method provided by an embodiment of the present application. The method includes:

[0048] S201: Obtain a target video and a sample subtitle feature corresponding to a sample video frame.

[0049] In order to be able to automatically detect the subtitle display function, the processing device can first determine a feature that can identify the subtitle content. It can be understood that each character has a relatively unique character structure. For example, each Chinese character has a unique stroke composition, writing method, etc. Through the character structure, it is possible to accurately distinguish whether two characters are the same. Based on this, in the embodiments of the present application, the processing device can detect whether subtitles match based on the character structure of the subtitles. When the character structures corresponding to two subtitles are different, it can be determined that the two subtitles do not match.

[0050] First, the processing device can obtain a target video and sample subtitle features corresponding to a sample video frame. The target video is a video with audio information and no subtitle content. The sample video frame is a video frame in the sample video, and the sample video is the target video that displays the sample subtitles. The sample subtitle is the accurate subtitle corresponding to the audio information in the target video. Among them, the sample video frame corresponds to the target video frame in the target video, that is, the sample video frame is obtained by accurately displaying subtitles in the target video frame. The sample subtitle feature is used to identify the character structure of the sample subtitles in the sample video frame. It can be understood that the accurate subtitle display means that if the target video frame has corresponding audio information, there are sample subtitles corresponding to the audio information in the sample video frame; if the target video frame does not have audio information that can be used for subtitle display in the target video, the sample subtitles in the sample video frame can be empty subtitles.

[0051] S202: According to the audio information in the target video, the program under test displays the subtitle under test corresponding to the audio information in the target video to obtain a test video.

[0052] After obtaining the sample subtitle features corresponding to the sample video frame, the processing device can detect whether the program under test can accurately display subtitles for the target video frame. The processing device can, according to the audio information in the target video, display the subtitle under test corresponding to the audio information in the target video through the program under test to obtain a test video. Among them, the program under test is a program with a subtitle display function, that is, it can display the corresponding subtitle content based on the audio information, and the subtitle under test is the subtitle content displayed by the program under test.

[0053] S203: Determine the subtitle feature under test according to the test video frame corresponding to the target video frame in the test video.

[0054] To determine whether there is an abnormality in the subtitle display function of the program under test, the processing device can match the subtitle under test displayed by the program under test with the sample subtitle. If, for the same video frame in the target video, the subtitle under test displayed by the program under test in this video frame is the same as the sample subtitle corresponding to this video frame in the sample video, it indicates that the program under test can accurately display subtitles in this video frame.

[0055] Since the above sample video frame corresponds to the target video frame in the target video, the processing device can determine the subtitle feature under test according to the video frame under test corresponding to the target video frame in the video under test. The subtitle feature under test is used to identify the character structure of the subtitle under test in the video frame under test.

[0056] S204: Determine whether there is an abnormality in the subtitle display of the program under test in the video frame under test according to the sample subtitle feature and the subtitle feature under test.

[0057] Since both the video frame under test and the sample video frame correspond to the target video frame, if there is no abnormality in the subtitle display of the program under test in the video frame under test, the sample subtitle in the sample video frame should be the same as the subtitle under test in the video frame under test, that is, the sample subtitle and the subtitle under test have the same character structure, and the sample subtitle feature and the subtitle feature under test are consistent; on the contrary, if there is an abnormality in the subtitle display of the program under test, there will be obvious feature differences between the sample subtitle feature and the subtitle feature under test. Therefore, through the sample subtitle feature and the subtitle feature under test, the processing device can determine whether there is an abnormality in the subtitle display of the program under test in the video frame under test.

[0058] It can be seen from the above technical solution that when performing abnormality detection on the subtitle display function of the program under test, since the feature expression of the character structure is simple, and it can clearly reflect the subtitle structure form in the video frame, and there are obvious structural differences between different characters in the subtitle, the comparison between the sample subtitle feature and the subtitle feature under test can accurately and quickly determine whether the subtitle content displayed in the video frame under test is correct and whether the position is offset. This enables the automatic detection of abnormalities in the subtitle display of the program under test, eliminating the need for manual participation. Moreover, during the detection process, there is no need to identify and process the complex semantic information of the subtitle, and the detection can be completed through the simplified subtitle feature, reducing the resource occupancy of automatic detection and improving the detection efficiency.

[0059] It can be understood that the image content in a video is composed of individual pixel points. When displaying subtitles in a video, it is also achieved by setting the colors of some pixel points in the video frame to the subtitle colors corresponding to the subtitles. Therefore, when the character structures corresponding to the subtitles are different, the display situations of the subtitles in the video frame are also different, resulting in different pixel points corresponding to the subtitles in the video frame. Based on this, by analyzing the pixel points with subtitle colors in the video frame, the processing device can know the display situation of the subtitle in the video frame, and thus can determine the character structure corresponding to the subtitle to a certain extent.

[0060] In a possible implementation manner, when determining the characteristics of the subtitle to be measured, the processing device can determine the subtitle pixel points among the pixel points included in the video frame to be measured, and the subtitle pixel points are the pixel points with the colors corresponding to the subtitle to be measured. Through the subtitle pixel points, the processing device can analyze the display situation of the subtitle to be measured in the video frame to be measured, and then can determine the characteristics of the subtitle to be measured corresponding to the video frame to be measured. By analyzing the different subtitle pixel points corresponding to different subtitles in the same video frame, the processing device can determine the differences between the character structures corresponding to the subtitles.

[0061] It can be understood that when determining the subtitle characteristics based on the subtitle pixel points, the higher the accuracy of the determined subtitle pixel points, the more accurately the finally obtained subtitle characteristics can identify the character structure of the subtitle. Therefore, in order to improve the accuracy of the characteristics of the subtitle to be measured, the processing device can set the colors of the pixel points irrelevant to the subtitle to be measured in the video frame to other colors different from the subtitle color to improve the distinguishability of the subtitle pixel points in the video frame to be measured.

[0062] In a possible implementation manner, the processing device can set that the video content in the target video has a single content color, and the content color is different from the color corresponding to the subtitle to be measured. The video content refers to all the content displayed in the target video. Thus, when the subtitle to be measured is displayed in the target video through the program to be measured, the processing device can accurately determine the subtitle pixel points in each video frame. For example, the processing device can construct a video with a pure black screen and audio information as the target video, set the color of the subtitle to be measured to white, and the processing device can determine the subtitle pixel points by determining the pixel points whose colors are not black or the pixel points whose colors are white in the video frame to be measured, as Figure 5 shown. At the same time, since the parts other than the subtitle to be measured in the video frame to be measured all have a single content color, even if the subtitle to be measured is a color subtitle including multiple character colors, the processing device can accurately determine the subtitle pixel points through the non-content color. For example, when the video content of the target video is all black, the processing device can determine the subtitle pixel points corresponding to the colorful subtitle through the non-black pixel points.

[0063] There are also multiple ways to determine the features of the subtitles to be tested based on subtitle pixels. The following will mainly introduce the methods based on subtitle projection and subtitle pixel distribution relationship:

[0064] (1) Subtitle projection-based method

[0065] As mentioned above, subtitles with different character structures have different corresponding subtitle pixels in the video frame, and the video frame is usually obtained based on the combination of multiple rows and columns of pixels. Therefore, when the subtitles do not match, the number of subtitle pixels corresponding to the subtitles in each column or row of pixels will also be different. Figure 3 As shown, the number of character pixels in each column of pixels of the two characters "日" and "田" is different.

[0066] Based on this, in a possible implementation method, when analyzing the differences between the subtitle pixel points corresponding to the subtitles, the processing device can project the subtitles in columns or rows, that is, the number of subtitle pixel points can be counted in rows or columns to determine the distribution of the subtitle pixel points in the video frame, and then the character structure corresponding to the subtitles can be reflected through the distribution.

[0067] On the one hand, the video frame to be tested may include N columns of pixel points, and the processing device may determine a first number of subtitle pixel points respectively included in the N columns of pixel points, and then determine the subtitle features to be tested corresponding to the video frame to be tested based on the first number and the arrangement relationship between the N columns of pixel points. Through the first number, the processing device can determine the distribution of subtitle pixel points corresponding to the subtitles to be tested in each column of pixel points, that is, it can determine the display of the character structure corresponding to the subtitles to be tested in each column of pixel points; through the arrangement relationship, the processing device can accurately combine the display of the character structure in each column of pixel points, so as to accurately reflect the character structure of the subtitles to be tested on the N columns of pixel points. Finally, the subtitle features to be tested determined in this way can accurately identify the character structure of the subtitles to be tested. Figure 6 As shown, by projecting the subtitle to be tested “This is the sea view we see” in the column direction, the processing device can obtain the subtitle features to be tested in the form of a one-dimensional array.

[0068] Similarly, to some extent, the distribution of subtitle pixel points corresponding to the subtitles among the pixel points in each row can also reflect the character structure of the subtitles. In one possible implementation, the video frame to be tested may include M rows of pixel points. The processing device can determine the second number of subtitle pixel points included in each of the M rows of pixel points, and then determine the subtitle feature to be tested corresponding to the video frame to be tested according to the second number and the arrangement relationship among the M rows of pixel points. Through this second number, the processing device can determine the distribution of the subtitle pixel points corresponding to the subtitle to be tested among the pixel points in each row, that is, it can determine the display of the character structure corresponding to the subtitle to be tested among the pixel points in each row; through this arrangement relationship, the processing device can accurately combine the display of the character structure in each row of pixel points, so as to accurately reflect the character structure of the subtitle to be tested on the M rows of pixel points. Finally, the subtitle feature to be tested determined by this method can also accurately identify the character structure of the subtitle to be tested.

[0069] It can be understood that, in order to further improve the accuracy of the subtitle feature to be tested, in addition to determining the subtitle feature to be tested based on the distribution of subtitle pixel points in each column and the distribution of subtitle pixel points in each row separately, the processing device can also comprehensively determine by combining the distribution of subtitle pixel points in each column and each row.

[0070] For example, in one possible implementation, after determining the video frame to be tested corresponding to the target video frame in the video to be tested, the processing device can perform subtitle projection on the subtitle to be tested in the video frame to be tested in the column direction and the row direction, and comprehensively determine the subtitle feature to be tested corresponding to the subtitle to be tested. Among them, assuming that the video frame to be tested is the nth frame in the video to be tested, the processing device can determine the projection result of the subtitle to be tested in the column direction and the projection result of the subtitle to be tested in the row direction s n′ and v n′ as the subtitle feature to be tested corresponding to the subtitle to be tested. Among them, is the first number of subtitle pixel points in the i-th column of pixel points, is the second number of subtitle pixel points in the i-th row of pixel points, The arrangement order between them corresponds to the arrangement relationship between the pixel points in each column, The arrangement order between them corresponds to the arrangement relationship between the pixel points in each row. The processing device can compare the s n′ and v n′ with the sample subtitle feature and corresponding to the sample video frame to determine whether there is an abnormality in the subtitle display of the program to be tested in the video frame to be tested.

[0071] (2) Method based on the distribution relationship of subtitle pixel points

[0072] It is understandable that subtitles are displayed through the distribution between subtitle pixels and non-subtitle pixels in a video frame, that is, a specific character in the subtitle is drawn through a certain characteristic distribution relationship between subtitle pixels and non-subtitle pixels. Therefore, to a certain extent, the unique character structure of the subtitle in the video frame can also be reflected through the distribution relationship between subtitle pixels and non-subtitle pixels in the video frame.

[0073] In a possible implementation manner, the processing device can determine the subtitle feature to be measured corresponding to the video frame to be measured according to the distribution relationship between subtitle pixels and non-subtitle pixels in the video frame to be measured. Non-subtitle pixels refer to the pixels in the video frame to be measured other than subtitle pixels. Through this distribution relationship, the structural composition of the characters in the subtitle to be measured in the video frame to be measured can be reflected, so that the character structure corresponding to the subtitle to be measured can be identified.

[0074] Among them, there can also be various ways to determine the subtitle feature to be measured based on this distribution relationship. For example, in a possible implementation manner, the processing device can determine the identifier corresponding to the subtitle pixels in the video frame to be measured as the first identifier, for example, it can be determined as 1; and determine the identifier corresponding to the non-subtitle pixels as the second identifier, for example, it can be determined as 0. The processing device can generate an identifier sequence corresponding to the video frame to be measured based on the identifiers respectively corresponding to each pixel in the video frame to be measured and the distribution relationship between the pixels, for example, it can be a 0, 1 sequence. Thus, through this 0, 1 sequence, the distribution of subtitle pixels in the video frame to be measured can be reflected, and further the character structure corresponding to the subtitle to be measured can be identified. When two subtitles are different, the subtitle pixels corresponding to the characters included in the subtitle in the video frame to be measured are also different, so the obtained identifier sequences are also different. Based on this, through the identifier sequence corresponding to the video frame to be measured, the processing device can determine whether the subtitle to be measured matches the sample subtitle.

[0075] In another possible implementation manner, the processing device can determine the hash value corresponding to the video frame to be measured based on this distribution relationship through the MD5 Message-Digest Algorithm. This hash value is used to identify the integrity of the video frame to be measured. That is, this hash value can identify the unique distribution relationship between subtitle pixels and non-subtitle pixels in the sample video frame. When this distribution relationship changes, the corresponding hash value will also change.

[0076] Therefore, when the hash value corresponding to the sample video frame is different from the hash value corresponding to the video frame to be tested, it can be indicated that there is a different distribution relationship between the sample video frame and the video frame to be tested, that is, the character structures corresponding to the subtitles are different, and the processing device can determine that the subtitle to be tested and the sample subtitle are different subtitles.

[0077] Specifically, in a possible implementation manner, after determining the subtitle feature to be tested corresponding to the video frame to be tested through the above multiple methods, the processing device can determine the matching degree between the sample subtitle and the subtitle to be tested according to the sample subtitle feature and the subtitle feature to be tested. Since the sample subtitle feature can identify the character structure of the sample subtitle, and the subtitle feature to be tested can identify the character structure of the subtitle to be tested, the difference in the character structures between the sample subtitle and the subtitle to be tested can be determined through the difference between the features. On the basis that the character structure can accurately reflect the corresponding characters, the processing device can determine the matching degree between the sample subtitle and the subtitle to be tested. The processing device can set a matching threshold, and this matching threshold is used to determine whether the subtitle display of the program to be tested is abnormal. If the matching degree meets the matching threshold, it indicates that the subtitle to be tested and the sample subtitle have the same character structure, the subtitle to be tested and the sample subtitle match, and the processing device can determine that the subtitle display of the program to be tested in the video frame to be tested does not appear abnormal. If the matching degree does not meet the matching threshold, it indicates that there is a difference in the character structures between the subtitle to be tested and the sample subtitle, the subtitle to be tested and the sample subtitle do not match, and the processing device can determine that the subtitle display of the program to be tested in the video frame to be tested appears abnormal.

[0078] For example, after determining the subtitle feature to be tested based on the subtitle projection method, the processing device can respectively obtain the subtitle feature corresponding to the video frame to be tested and as well as the sample subtitle feature and This target video frame is the nth video frame in the target video frame, s n is the number of subtitle pixels in each column of pixels in the sample video frame, and v n is the number of subtitle pixels in each row of pixels in the sample video frame. The processing device can use formula (1) and formula (2) to determine whether the subtitle display appears abnormal:

[0079]

[0080] where k s and k vBoth are matching thresholds, which can be set to 1% for example. If the above formula is satisfied, it indicates that the sample subtitle feature is consistent with the subtitle feature to be measured, the sample subtitle and the subtitle to be measured have the same character structure, and there is no abnormality in the subtitle display of the program to be measured in the video frame to be measured.

[0081] It can be understood that in order to improve the diversity of subtitle display and thus enhance the viewing experience of users when watching videos, in addition to being able to display corresponding subtitle content based on audio information, the program to be measured can also display subtitles through a variety of rich display forms. This display form refers to the form of subtitle content display, which can include, for example, display position, display zoom level, display font, etc.

[0082] Among them, the display form of subtitles will also affect the subtitle pixel points corresponding to the subtitles in the video frame to a certain extent. For example, when the display zoom level of the subtitles is different, the subtitles with a larger zoom factor may occupy more subtitle pixel points in the target video frame; at the same time, when the same subtitle is in different positions in the video frame, the distribution of the corresponding subtitle pixel points in the video frame to be measured is also different. Since subtitles are displayed in the video frame through subtitle pixel points, when the subtitle pixel points corresponding to the subtitles are different, the manifestation of the character structure of the subtitles may also be different, which may affect the subtitle detection based on the character structure. Based on this, in order to further improve the accuracy of subtitle detection, the processing device can first reduce the interference of the subtitle display form on subtitle feature matching before determining the subtitle feature to be measured, and convert the subtitle to be measured and the sample subtitle to the same subtitle display form dimension for subtitle feature matching.

[0083] In a possible implementation, the processing device can also determine the subtitle display parameter to be measured of the program to be measured, and this subtitle display parameter to be measured is used to identify the display form of subtitle display of the program to be measured. Subsequently, in order to convert the subtitle to be measured and the sample subtitle to the same display form, the processing device can obtain the sample subtitle display parameter for displaying the sample subtitle in the sample video, and this sample subtitle display parameter is used to identify the display form of displaying the sample subtitle.

[0084] When determining whether there is an abnormality in subtitle display, the processing device can first convert the subtitle feature to be measured into a conversion feature that conforms to the sample subtitle display parameter according to the mapping relationship of subtitle pixels between the subtitle display parameter to be measured and the sample subtitle display parameter. Since the subtitle display parameter and the subtitle display parameter to be measured respectively identify the display forms of the sample subtitle and the subtitle to be measured, the processing device can, through this mapping relationship, know the conversion method between the two display forms, and then can convert the subtitle feature to be measured and the sample subtitle feature into the dimension of the same display form for comparison, reducing the interference of different display forms on subtitle feature comparison. The processing device can determine whether there is an abnormality in the subtitle display of the program to be measured in the video frame to be measured according to the sample subtitle feature and the conversion feature, so that the detection result can accurately reflect the difference in subtitle content between the sample subtitle and the subtitle to be measured.

[0085] For example, when displaying the sample subtitle, the point (0, 0) can be used as the center point of subtitle display to determine the display position of the sample subtitle, and no scaling process is used as the display scaling degree of the sample subtitle for display; when displaying the subtitle to be measured, the point (x0, y0) can be used as the center point of subtitle display to determine the display position of the sample subtitle, and a 3-fold magnification scaling process is used as the display scaling degree of the sample subtitle for display. Before determining whether the subtitle display is abnormal, the processing device can first move the subtitle to be measured to the subtitle position with the point (0, 0) as the center point based on the mapping relationship of subtitle pixels between the point (0, 0) and the point (x0, y0), and scale it down by 3 times in the same scaling manner as the sample subtitle, so that the sample subtitle and the converted subtitle to be measured are in the same display form. Subsequently, the processing device can determine the conversion feature that conforms to the sample subtitle display parameter based on the converted subtitle to be measured. Or, when the fonts of the sample subtitle and the subtitle to be measured are different, through the mapping relationship of subtitle pixels between different fonts, the processing device can also convert the subtitle to be measured and the sample subtitle into the same subtitle font, so that the obtained conversion feature and the sample subtitle feature are compared in the same font dimension.

[0086] As mentioned above, to a certain extent, the display form of the subtitle will determine the subtitle pixels corresponding to the subtitle in the video frame, so whether the subtitle display form is accurate will also affect whether the subtitle pixels corresponding to the subtitle in the video frame are accurate. Since the subtitle pixels can reflect the character structure of the subtitle, if the display form is abnormal, the processing device may not be able to accurately identify the character structure of the subtitle, and thus may not be able to accurately determine whether the sample subtitle and the subtitle to be measured have the same subtitle content based on the character structure of the subtitle.

[0087] Based on this, when detecting the subtitle display function of the program to be tested, in order to further improve the accuracy of the detection result, in addition to detecting the character content in the subtitle, the display form of the subtitle to be tested by the program to be tested can also be detected. In a possible implementation, the display form may include the display position. Before determining whether the program to be tested is abnormal based on the sample subtitle features and the subtitle features to be tested, the processing device may determine the actual subtitle position parameter corresponding to the subtitle to be tested in the video frame to be tested, and the actual subtitle position parameter is used to determine the actual display position corresponding to the subtitle to be tested. The processing device may determine whether the display form of the subtitle display in the video frame to be tested by the program to be tested is abnormal according to the target display position identified by the subtitle display parameter to be tested and the actual display position, and the target display position is the accurate display position corresponding to the program to be tested when displaying the subtitle to be tested. If the target display position is different from the actual display position, it indicates that the subtitle position is abnormal when the program to be tested displays the subtitle.

[0088] For example, the processing device may determine the left and right endpoints of the subtitle to be tested displayed in the video frame to be tested according to the column positions of the leftmost column and the rightmost column among the columns of pixel points including subtitle pixel points in the video frame to be tested, use the left and right endpoints as the actual subtitle position parameter, and determine the abscissa of the center point of the subtitle to be tested according to the actual subtitle position parameter, and the abscissa is the actual display position corresponding to the subtitle to be tested. The processing device may compare the abscissa with the abscissa of the target display position identified by the subtitle display parameter to be tested to determine whether the display position is abnormal when the subtitle to be tested displays the subtitle.

[0089] In another possible implementation, the display form may include the display scaling degree. Before determining whether the subtitle display is abnormal, the processing device may determine the actual subtitle scaling parameter corresponding to the subtitle to be tested in the video frame to be tested, and the actual subtitle scaling parameter is used to determine the actual display scaling degree corresponding to the subtitle to be tested. For example, the actual subtitle scaling parameter may be the height and width corresponding to the subtitle to be tested in the video frame to be tested. By comparing the height and width with the default height and width, the display scaling degree of the subtitle to be tested can be determined.

[0090] The processing device may determine whether the display form of the subtitle display in the video frame to be tested by the program to be tested is abnormal according to the target display scaling degree identified by the subtitle display parameter to be tested and the actual display scaling degree.

[0091] For example, after determining the subtitle features to be tested based on the subtitle projection method, the processing device may respectively obtain the subtitle features to be tested corresponding to the video frame to be tested and and the sample subtitle features and Let s n′ If the subscript value of the first non - zero value in is n1 and the subscript value of the last non - zero value is n2, then the abscissa of the actual display position determined can be The processing device can compare it with the abscissa x0 of the target display position through formula (3):

[0092]

[0093] where k x is the determination threshold, for example, it can be 1%. Similarly, the processing device can determine the first non - zero subscript value l1 and the last non - zero subscript value l2 in v n′ and compare them with the ordinate y0 of the target display position.

[0094] Next, the processing device can detect whether the display scaling degree of the subtitle to be measured is abnormal. When the display scaling degree of the sample subtitle is not scaled, the processing device can set s n If the subscript value of the first non - zero value in is n3 and the subscript value of the last non - zero value is n4, then the subtitle width corresponding to the sample subtitle is n4 - n3, and the corresponding length of the subtitle to be measured is n2 - n1; similarly, the processing device can determine the first non - zero subscript value l3 and the last non - zero subscript value l4 in v n Then the subtitle height corresponding to the sample subtitle is l4 - l3, and the subtitle height corresponding to the subtitle to be measured is l2 - l1. The processing device can first detect whether the horizontal and vertical scaling degrees of the subtitle to be measured are consistent through formula (4):

[0095]

[0096] where k b is the determination threshold, which can be set to 1%. If the formula is satisfied, it means that the horizontal and vertical scaling ratios of the subtitle to be measured are the same when performing subtitle scaling. Subsequently, if the target display scaling degree is to be magnified by m times, the processing device can detect whether the display scaling degree of the subtitle to be measured is abnormal through formula (5):

[0097]

[0098] where k b is the determination threshold, which can be set to 1%. If the formula is satisfied, it means that the display scaling degree corresponding to the subtitle to be measured is to be magnified by m times when performing subtitle display.

[0099] It can be understood that when the display form of the program under test is abnormal during subtitle display, to a certain extent, it will affect the subtitle pixel points corresponding to the subtitle to be tested in the video frame to be tested, and further affect the subtitle detection based on the character structure. Therefore, in a possible implementation, if it is determined through the above method that the display form of the subtitle display of the program under test in the video frame to be tested is abnormal, it can be considered that the subtitle detection based on the character structure lacks detection accuracy at this time. The processing device can directly determine that the subtitle display of the program under test in the video frame to be tested is abnormal, and there is no need to perform the operation of detecting whether the subtitle matches based on the subtitle features, so as to be able to report an error in a timely manner and save the time required for subtitle detection.

[0100] In addition to including various methods when detecting the display form and determining subtitle features, the method for the processing device to select sample video frames can also be different based on different requirements. In a possible implementation, the sample video frame can include the video frames corresponding to the start display and / or end display moments of the sample subtitle in the sample video, and can also include the video frames corresponding to the middle moment of the sample subtitle display in the sample video, so as to detect whether there is an abnormality in the display time interval when the program under test displays the sample subtitle, and whether there are abnormalities in the subtitle display content and display form during the display time interval. Or, the sample video frame can be collected from the sample video based on the video frame sampling interval, and the video frame sampling interval can be determined based on the audio information in the target video. Thus, subtitle detection based on the sample video frame can determine whether the subtitle to be tested is displayed at the wrong time, such as whether the subtitle to be tested is already displayed in the video to be tested when the sample subtitle has not been displayed yet. By determining representative and targeted sample video frames in the sample video for detection, it is possible to ensure the accuracy of subtitle detection without detecting all video frames in the sample video, reducing the detection time required and further improving the detection efficiency.

[0101] Next, a subtitle detection method provided by an embodiment of the present application will be introduced in combination with an actual application scenario.

[0102] See Figure 4 , Figure 4 which is a flowchart of a subtitle detection method in an actual application scenario provided by an embodiment of the present application. In this actual application scenario, the processing device can be a terminal device used for subtitle detection, and the program under test can be a to-be-tested application with a subtitle display function in the terminal device. The method includes:

[0103] S401: Construct a target video.

[0104] The tester can first construct a video V with a pure black screen and audio information t , and determine it as the target video. Set its frequency f = 30 frames per second, and the total video duration t = 10 seconds.

[0105] S402: Generate a sample video.

[0106] The tester can use this target video as the input video of a video editing application, and add sample subtitles to this target video to generate a sample video V s , and manually detect the accuracy of the sample subtitles in the sample video.

[0107] S403: Calculate and store the sample subtitle features corresponding to each sample video frame in the sample video.

[0108] The terminal device can determine the features of each frame in the sample video through the above subtitle projection method to obtain the sample subtitle features corresponding to this sample video. For example, for a certain sample video frame, the terminal device can first perform vertical projection through formula (6), that is, determine the number of subtitle pixel points included in each column of pixel points in this sample video frame. The subtitle pixel point is a "non-all-black" pixel point (that is, a pixel point where R≠0, G≠0, B≠0):

[0109]

[0110] where s x is the vertical projection of the x-th column, n is the n-th frame of the sample video, x represents the abscissa, y represents the ordinate, and the expression of b(x, y) is shown in formula (7) as follows:

[0111]

[0112] That is, the non-black dot count is 1, and the count of other pixel points is 0. Thus, for the n-th frame, the terminal device can simplify the sample subtitle features in this sample video frame into a one-dimensional array Similarly, a one-dimensional array can be obtained through horizontal projection The terminal device can store the sample subtitle features s = [s 0 , s 1 , s 2 ……] and v = [v 0 , v 1 , v 2 ……] corresponding to all frames of the sample video for subsequent automatic subtitle detection. Since only the one-dimensional array as the sample subtitle feature needs to be stored and there is no need to store this sample video, the storage space is greatly saved.

[0113] S404: According to the audio information in the target video, display the corresponding subtitle to be tested in the target video through the application to be tested, and obtain the video to be tested.

[0114] The terminal device can obtain the target video, and display subtitles in the target video through the application to be tested, so as to obtain the video to be tested V s′ . When displaying subtitles, the center point of the subtitle display position can be manually set as (x0, y0), and the display zoom level is enlarged by m times.

[0115] S405: Take each frame of the video to be tested frame by frame, and determine the corresponding subtitle feature to be tested.

[0116] The terminal device can determine the subtitle feature s corresponding to each frame in the video to be tested through the subtitle projection method ′ =[s 0′ , s 1′ , s 2′ ……] and v ′ =[v 0′ , v 1′ , v 2′ ……]. For example, for the nth frame, the horizontal projection result is Vertical projection result

[0117] S406: Perform automatic subtitle detection frame by frame.

[0118] Among them, for the nth frame, the steps for the terminal device to perform automatic subtitle detection can be as follows:

[0119] S4061: Determine whether the display time is abnormal.

[0120] If All the data in are 0, and are all 0, it means that the target video frame corresponding to the video frame to be tested does not have corresponding audio information, and the display time is normal; if one of them is 0 and the other is not 0, it means that there may be a situation where the subtitle to be tested is missing or the subtitle to be tested is displayed in advance, and it is determined that the subtitle display of the application to be tested is abnormal and an error is reported.

[0121] If they are not all 0, then execute step S4062.

[0122] S4062: Determine whether the display position is abnormal.

[0123] The terminal device can detect whether the center point of the subtitle to be tested is (x0, y0). If it is abnormal, the process is terminated and an abnormal error is reported.

[0124] S4063: Determine whether the display zoom level is abnormal.

[0125] The terminal device can detect whether the display zoom level of the subtitle to be tested is magnified by m times. If it is abnormal, the process is terminated and an abnormal error is reported.

[0126] S4064: Determine the subtitle display parameters to be tested.

[0127] In this actual application scenario, the subtitle display parameters to be tested are the display position (x0, y0) and the display zoom level magnified by m times.

[0128] S4065: Obtain the subtitle display parameters of the sample.

[0129] S4066: According to the mapping relationship of subtitle pixel points between the subtitle display parameters to be tested and the subtitle display parameters of the sample, convert the subtitle feature to be tested into a conversion feature that conforms to the subtitle display parameters of the sample.

[0130] The terminal device can first perform feature conversion, convert the sample subtitle feature and the subtitle feature to be tested into the dimension of the same display form for comparison. For example, when the sample subtitle is in the default display form, since the subtitle to be tested has been scaled and displaced, the subtitle to be tested is first moved back to the default center point (0, 0) and reduced by m times, and the conversion feature is obtained based on the restored subtitle to be tested by performing subtitle projection. and

[0131] S4067: According to the conversion feature and the sample subtitle feature, determine whether there is an abnormality in the subtitle display of the program to be tested in the nth frame of the video frame to be tested.

[0132] The terminal device can perform abnormal detection through the following formulas (8) and (9):

[0133]

[0134]

[0135] If the formula is satisfied, it means that the subtitle to be tested in the nth frame of the video frame to be tested matches the sample subtitle, and there is no abnormality in the display form and display time of the subtitle display of the application to be tested, and the automatic detection passes.

[0136] Based on the subtitle detection method provided in the above embodiments, the embodiments of the present application also provide a subtitle detection device. Refer to Figure 7 , Figure 7 which is a structural block diagram of a subtitle detection device 700 provided by the embodiments of the present application. The device 700 includes a first acquisition unit 701, a display unit 702, a first determination unit 703, and a second determination unit 704:

[0137] A first acquisition unit 701 is configured to acquire a target video and sample subtitle features corresponding to sample video frames, where the sample subtitle features are used to identify the character structure of sample subtitles in the sample video frames, the sample video frames are video frames in a sample video, the sample video is the target video that shows the sample subtitles, and the sample video frames correspond to target video frames in the target video;

[0138] A display unit 702 is configured to, according to audio information in the target video, display a to-be-tested subtitle corresponding to the audio information in the target video through a to-be-tested program to obtain a to-be-tested video;

[0139] A first determination unit 703 is configured to determine to-be-tested subtitle features according to to-be-tested video frames corresponding to the target video frames in the to-be-tested video, where the to-be-tested subtitle features are used to identify the character structure of to-be-tested subtitles in the to-be-tested video frames;

[0140] A second determination unit 704 is configured to determine whether there is an abnormality in subtitle display of the to-be-tested program in the to-be-tested video frames according to the sample subtitle features and the to-be-tested subtitle features.

[0141] In a possible implementation manner, the first determination unit 703 is specifically configured to:

[0142] Determine subtitle pixel points among pixel points included in the to-be-tested video frame, where the subtitle pixel points are pixel points having a color corresponding to the to-be-tested subtitle;

[0143] According to the subtitle pixel points, determine the to-be-tested subtitle features corresponding to the to-be-tested video frame.

[0144] In a possible implementation manner, the to-be-tested video frame includes N columns of pixel points, and the first determination unit 703 is specifically configured to:

[0145] Determine first quantities of the subtitle pixel points respectively included in the N columns of pixel points;

[0146] According to the first quantities and the arrangement relationship among the N columns of pixel points, determine the to-be-tested subtitle features corresponding to the to-be-tested video frame.

[0147] In a possible implementation manner, the to-be-tested video frame includes M rows of pixel points, and the first determination unit 703 is specifically configured to:

[0148] Determine second quantities of the subtitle pixel points respectively included in the M rows of pixel points;

[0149] According to the second quantities and the arrangement relationship among the M rows of pixel points, determine the to-be-tested subtitle features corresponding to the to-be-tested video frame.

[0150] In a possible implementation, the first determination unit 703 is specifically configured to:

[0151] Determine the subtitle feature to be measured corresponding to the video frame to be measured according to the distribution relationship between the subtitle pixel points and non-subtitle pixel points in the video frame to be measured.

[0152] In a possible implementation, the second determination unit 704 is specifically configured to:

[0153] Determine the matching degree between the sample subtitle feature and the subtitle feature to be measured;

[0154] If the matching degree meets the matching threshold, determine that there is no abnormality in the subtitle display of the program to be measured in the video frame to be measured;

[0155] If the matching degree does not meet the matching threshold, determine that there is an abnormality in the subtitle display of the program to be measured in the video frame to be measured.

[0156] In a possible implementation, the apparatus 700 further includes a third determination unit and a second acquisition unit:

[0157] The third determination unit is configured to determine the subtitle display parameter to be measured of the program to be measured, and the subtitle display parameter to be measured is used to identify the display form of the subtitle display of the program to be measured;

[0158] The second acquisition unit is configured to acquire the sample subtitle display parameter for displaying the sample subtitle in the sample video;

[0159] The second determination unit 704 is specifically configured to:

[0160] Convert the subtitle feature to be measured into a conversion feature that conforms to the sample subtitle display parameter according to the mapping relationship of subtitle pixel points between the subtitle display parameter to be measured and the sample subtitle display parameter;

[0161] Determine whether there is an abnormality in the subtitle display of the program to be measured in the video frame to be measured according to the sample subtitle feature and the conversion feature.

[0162] In a possible implementation, the display form includes a display position, and the apparatus 700 further includes a fourth determination unit and a fifth determination unit:

[0163] The fourth determination unit is configured to determine the actual subtitle position parameter corresponding to the subtitle to be measured in the video frame to be measured, and the actual subtitle position parameter is used to determine the actual display position corresponding to the subtitle to be measured;

[0164] A fifth determination unit, configured to determine whether an abnormality occurs in the display form of the subtitle display of the program to be tested in the video frame to be tested according to the target display position identified by the subtitle display parameter to be tested and the actual display position.

[0165] In a possible implementation, the display form includes a display zoom level, and the apparatus 700 further includes a sixth determination unit and a seventh determination unit:

[0166] A sixth determination unit, configured to determine an actual subtitle zoom parameter corresponding to the subtitle to be tested in the video frame to be tested, where the actual subtitle zoom parameter is used to determine an actual display zoom level corresponding to the subtitle to be tested;

[0167] A seventh determination unit, configured to determine whether an abnormality occurs in the display form of the subtitle display of the program to be tested in the video frame to be tested according to the target display zoom level identified by the subtitle display parameter to be tested and the actual display zoom level.

[0168] In a possible implementation, the apparatus 700 further includes an eighth determination unit:

[0169] An eighth determination unit, configured to determine that an abnormality occurs in the subtitle display of the program to be tested in the video frame to be tested if it is determined that an abnormality occurs in the display form of the subtitle display of the program to be tested in the video frame to be tested.

[0170] In a possible implementation, the sample video frame includes a video frame corresponding to a start display and / or an end display moment of the sample subtitle in the sample video;

[0171] Or, the sample video frame is collected from the sample video based on a video frame sampling interval.

[0172] In a possible implementation, the video content in the target video has a single content color, and the content color is different from the corresponding color of the subtitle to be tested.

[0173] The embodiments of the present application further provide a computer device, which will be introduced below with reference to the accompanying drawings. Please refer to Figure 8 As shown, the embodiments of the present application provide a device, and the device may also be a terminal device. The terminal device may be any intelligent terminal including a mobile phone, a tablet computer, a personal digital assistant (Personal Digital Assistant, abbreviated as PDA), a point of sales (Point of Sales, abbreviated as POS), an in-vehicle computer, etc. Taking the terminal device as a mobile phone as an example:

[0174] Figure 8Shown is a block diagram of a partial structure of a mobile phone related to the terminal device provided in an embodiment of the present application. Refer to Figure 8 , the mobile phone includes: a Radio Frequency (RF) circuit 710, a memory 720, an input unit 730, a display unit 740, a sensor 750, an audio circuit 760, a wireless fidelity (WiFi) module 770, a processor 780, and a power supply 790 and other components. Those skilled in the art can understand that Figure 8 the structure of the mobile phone shown in

[0175] does not limit the mobile phone, and may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements. Figure 8 The following specifically introduces each component of the mobile phone:

[0176] The RF circuit 710 can be used for receiving and sending signals during information reception or call processes. Specifically, after receiving the downlink information of the base station, it is given to the processor 780 for processing; in addition, the designed uplink data is sent to the base station. Generally, the RF circuit 710 includes, but is not limited to, antennas, at least one amplifier, a transceiver, a coupler, a Low Noise Amplifier (LNA), a duplexer, etc. In addition, the RF circuit 710 can also communicate with the network and other devices through wireless communication. The above wireless communication can use any communication standard or protocol, including but not limited to the Global System of Mobile communication (GSM), General Packet Radio Service (GPRS), Code Division Multiple Access (CDMA), Wideband Code Division Multiple Access (WCDMA), Long Term Evolution (LTE), email, Short Messaging Service (SMS), etc.

[0177] The memory 720 can be used to store software programs and modules. The processor 780 executes various functional applications and data processing of the mobile phone by running the software programs and modules stored in the memory 720. The memory 720 mainly includes a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for at least one function (such as a sound playback function, an image playback function, etc.); the data storage area can store data created according to the use of the mobile phone (such as audio data, phone book, etc.). In addition, the memory 720 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other volatile solid-state storage devices.

[0178] The input unit 730 can be used to receive input digital or character information, and generate key signal inputs related to the user settings and function controls of the mobile phone. Specifically, the input unit 730 may include a touch panel 731 and other input devices 732. The touch panel 731, also known as a touch screen, can collect touch operations of the user on or near it (such as operations of the user using a finger, a stylus, or any suitable object or accessory on or near the touch panel 731), and drive corresponding connection devices according to a preset program. Optionally, the touch panel 731 may include two parts: a touch detection device and a touch controller. Among them, the touch detection device detects the touch position of the user, detects the signal brought by the touch operation, and transmits the signal to the touch controller; the touch controller receives the touch information from the touch detection device, converts it into contact coordinates, and then sends it to the processor 780, and can receive and execute commands sent by the processor 780. In addition, various types such as resistive, capacitive, infrared, and surface acoustic wave can be used to implement the touch panel 731. In addition to the touch panel 731, the input unit 730 may also include other input devices 732. Specifically, the other input devices 732 may include, but are not limited to, one or more of a physical keyboard, function keys (such as volume control keys, power on / off keys, etc.), a trackball, a mouse, a joystick, etc.

[0179] The display unit 740 can be used to display information input by the user or information provided to the user, as well as various menus of the mobile phone. The display unit 740 may include a display panel 741. Optionally, the display panel 741 can be configured in the form of, for example, a liquid crystal display (LCD), an organic light-emitting diode (OLED), etc. Further, the touch panel 731 can cover the display panel 741. When the touch panel 731 detects a touch operation on or near it, it is transmitted to the processor 780 to determine the type of touch event. Subsequently, the processor 780 provides a corresponding visual output on the display panel 741 according to the type of touch event. Although in Figure 8 the touch panel 731 and the display panel 741 are implemented as two independent components to realize the input and input functions of the mobile phone, in some embodiments, the touch panel 731 and the display panel 741 can be integrated to realize the input and output functions of the mobile phone.

[0180] The mobile phone may further include at least one sensor 750, such as a light sensor, a motion sensor, and other sensors. Specifically, the light sensor may include an ambient light sensor and a proximity sensor. Among them, the ambient light sensor can adjust the brightness of the display panel 741 according to the brightness of the ambient light, and the proximity sensor can turn off the display panel 741 and / or the backlight when the mobile phone is moved to the ear. As a kind of motion sensor, the accelerometer sensor can detect the magnitude of acceleration in all directions (generally three axes). When stationary, it can detect the magnitude and direction of gravity, and can be used for applications that identify the posture of the mobile phone (such as horizontal and vertical screen switching, related games, magnetometer posture calibration), vibration recognition related functions (such as pedometer, tapping), etc.; as for other sensors that the mobile phone can also be configured with, such as gyroscopes, barometers, hygrometers, thermometers, infrared sensors, etc., they will not be elaborated here.

[0181] The audio circuit 760, the speaker 761, and the microphone 762 can provide an audio interface between the user and the mobile phone. The audio circuit 760 can transmit the electrical signal converted from the received audio data to the speaker 761, and the speaker 761 converts it into a sound signal for output; on the other hand, the microphone 762 converts the collected sound signal into an electrical signal, which is received by the audio circuit 760 and then converted into audio data. After the audio data is output to the processor 780 for processing, it is sent to, for example, another mobile phone through the RF circuit 710, or the audio data is output to the memory 720 for further processing.

[0182] WiFi belongs to short-range wireless transmission technology. The mobile phone can help users send and receive emails, browse the web, and access streaming media through the WiFi module 770. It provides users with wireless broadband Internet access. Although Figure 8A WiFi module 770 is shown, but it is understandable that it is not an essential component of the mobile phone and can be omitted as needed without changing the essence of the invention.

[0183] The processor 780 is the control center of the mobile phone. It uses various interfaces and lines to connect various parts of the entire mobile phone. It executes various functions of the mobile phone and processes data by running or executing software programs and / or modules stored in the memory 720, and calling data stored in the memory 720. Optionally, the processor 780 may include one or more processing units; preferably, the processor 780 may integrate an application processor and a modem processor, wherein the application processor mainly processes the operating system, user interface, and application programs, and the modem processor mainly processes wireless communications. It is understandable that the above-mentioned modem processor may not be integrated into the processor 780.

[0184] The mobile phone also includes a power supply 790 (such as a battery) for supplying power to various components. Preferably, the power supply can be logically connected to the processor 780 through a power management system, so that the power management system can manage charging, discharging, power consumption and other functions.

[0185] Although not shown, the mobile phone may also include a camera, a Bluetooth module, etc., which will not be described in detail here.

[0186] In this embodiment, the processor 780 included in the terminal device also has the following functions:

[0187] Acquire a target video and a sample subtitle feature corresponding to a sample video frame, wherein the sample subtitle feature is used to identify a character structure of a sample subtitle in the sample video frame, the sample video frame is a video frame in a sample video, the sample video is the target video showing the sample subtitle, and the sample video frame corresponds to a target video frame in the target video;

[0188] According to the audio information in the target video, displaying the subtitles to be tested corresponding to the audio information in the target video through the program to be tested, so as to obtain the video to be tested;

[0189] Determining a subtitle feature to be tested according to a video frame to be tested corresponding to the target video frame in the video to be tested, wherein the subtitle feature to be tested is used to identify a character structure of the subtitle to be tested in the video frame to be tested;

[0190] According to the sample subtitle features and the subtitle features to be tested, it is determined whether the subtitle display of the program to be tested in the video frame to be tested is abnormal.

[0191] The present application also provides a server. Figure 9 As shown, Figure 9The following is a structural diagram of the server 800 provided by an embodiment of this application. The server 800 may vary significantly due to different configurations or performances, and may include one or more central processing units (CPUs) 822 (for example, one or more processors) and a memory 832, and one or more storage media 830 (for example, one or more mass storage devices) for storing application programs 842 or data 844. Among them, the memory 832 and the storage media 830 may be transient storage or persistent storage. The program stored in the storage media 830 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations on the server. Further, the central processing unit 822 may be configured to communicate with the storage media 830 and execute a series of instruction operations in the storage media 830 on the server 800.

[0192] The server 800 may further include one or more power supplies 826, one or more wired or wireless network interfaces 850, one or more input / output interfaces 858, and / or one or more operating systems 841, such as Windows ServerTM, Mac OS XTM, UnixTM, LinuxTM, FreeBSDTM, and so on.

[0193] In the above embodiment, the steps executed by the server may be based on Figure 9 the server structure shown.

[0194] Those of ordinary skill in the art can understand that all or part of the steps for implementing the above method embodiments can be completed by hardware related to program instructions. The foregoing program can be stored in a computer-readable storage medium. When the program is executed, it executes the steps including the above method embodiments; and the foregoing storage medium may be at least one of the following media: read-only memory (abbreviation: ROM), RAM, magnetic disk, or optical disc, etc., which can store program codes.

[0195] It should be noted that the embodiments in this specification are all described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the device and system embodiments, since they are basically similar to the method embodiments, the description is relatively simple. For the relevant parts, reference can be made to the description of the method embodiments. The device and system embodiments described above are only illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative efforts.

[0196] As described above, this is only a specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed in the present application should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A subtitle detection method, characterized in that, The method includes: Obtaining a target video and sample subtitle features corresponding to a sample video frame, where the sample subtitle features are used to identify the character structure of the sample subtitle in the sample video frame, the sample video frame is a video frame in a sample video, the sample video is the target video that shows the sample subtitle, and the sample video frame corresponds to a target video frame in the target video; According to the audio information in the target video, the program under test displays the subtitle under test corresponding to the audio information in the target video to obtain a video under test; Determining subtitle features under test according to the video frame under test corresponding to the target video frame in the video under test, where the subtitle features under test are used to identify the character structure of the subtitle under test in the video frame under test; Determining whether there is an abnormality in the subtitle display of the program under test in the video frame under test according to the sample subtitle features and the subtitle features under test.

2. The method according to claim 1, wherein The determining subtitle features under test according to the video frame under test corresponding to the target video frame in the video under test includes: Determining subtitle pixels among the pixels included in the video frame under test, where the subtitle pixels are pixels having the color corresponding to the subtitle under test; Determining subtitle features under test corresponding to the video frame under test according to the subtitle pixels.

3. The method according to claim 2, wherein The video frame under test includes N columns of pixels, and the determining subtitle features under test corresponding to the video frame under test according to the subtitle pixels includes: Determining the first quantity of the N columns of pixels that respectively include the subtitle pixels; Determining subtitle features under test corresponding to the video frame under test according to the first quantity and the arrangement relationship among the N columns of pixels.

4. The method according to claim 2, wherein The video frame under test includes M rows of pixels, and the determining subtitle features under test corresponding to the video frame under test according to the subtitle pixels includes: Determining the second quantity of the M rows of pixels that respectively include the subtitle pixels; Determining subtitle features under test corresponding to the video frame under test according to the second quantity and the arrangement relationship among the M rows of pixels.

5. The method according to claim 2, wherein The determining subtitle features under test corresponding to the video frame under test according to the subtitle pixels includes: Determining subtitle features under test corresponding to the video frame under test according to the distribution relationship between the subtitle pixels and non-subtitle pixels in the video frame under test.

6. The method according to claim 1, wherein The determining whether there is an abnormality in the subtitle display of the program under test in the video frame under test according to the sample subtitle features and the subtitle features under test includes: Determining the matching degree between the sample subtitle and the subtitle under test according to the sample subtitle features and the subtitle features under test; If the matching degree meets the matching threshold, determining that there is no abnormality in the subtitle display of the program under test in the video frame under test; If the matching degree does not meet the matching threshold, determining that there is an abnormality in the subtitle display of the program under test in the video frame under test.

7. The method according to claim 1, wherein The method further includes: Determining subtitle display parameters under test of the program under test, where the subtitle display parameters under test are used to identify the display form of the subtitle display by the program under test; Obtaining sample subtitle display parameters for displaying the sample subtitle in the sample video; Determining whether there is an abnormality in the subtitle display of the program under test in the video frame under test according to the sample subtitle feature and the subtitle feature to be tested includes: Converting the subtitle feature to be tested into a conversion feature that conforms to the sample subtitle display parameter according to the mapping relationship of subtitle pixel points between the subtitle display parameter to be tested and the sample subtitle display parameter; Determining whether there is an abnormality in the subtitle display of the program under test in the video frame under test according to the sample subtitle feature and the conversion feature.

8. The method according to claim 7, wherein The display form includes a display position. Before determining whether there is an abnormality in the subtitle display of the program under test in the video frame under test according to the sample subtitle feature and the subtitle feature to be tested, the method further includes: Determining an actual subtitle position parameter corresponding to the subtitle to be tested in the video frame under test, where the actual subtitle position parameter is used to determine an actual display position corresponding to the subtitle to be tested; Determining whether there is an abnormality in the display form of the subtitle display of the program under test in the video frame under test according to the target display position identified by the subtitle display parameter to be tested and the actual display position.

9. The method according to claim 7, characterized in that The display form includes a display scaling degree. Before determining whether there is an abnormality in the subtitle display of the program under test in the video frame under test according to the sample subtitle feature and the subtitle feature to be tested, the method further includes: Determining an actual subtitle scaling parameter corresponding to the subtitle to be tested in the video frame under test, where the actual subtitle scaling parameter is used to determine an actual display scaling degree corresponding to the subtitle to be tested; Determining whether there is an abnormality in the display form of the subtitle display of the program under test in the video frame under test according to the target display scaling degree identified by the subtitle display parameter to be tested and the actual display scaling degree.

10. The method according to claim 8 or 9, characterized in that The method further includes: If it is determined that there is an abnormality in the display form of the subtitle display of the program under test in the video frame under test, determining that there is an abnormality in the subtitle display of the program under test in the video frame under test.

11. The method according to claim 1, wherein The sample video frame includes a video frame corresponding to the start display and / or end display moment of the sample subtitle in the sample video; Or, the sample video frame is obtained by sampling video frames in the sample video based on a video frame sampling interval.

12. The method according to claim 1, wherein The video content in the target video has a single content color, and the content color is different from the color corresponding to the subtitle to be tested.

13. A subtitle detection device, characterized in that, The device includes a first acquisition unit, a display unit, a first determination unit, and a second determination unit: The first acquisition unit is configured to acquire a target video and a sample subtitle feature corresponding to a sample video frame, where the sample subtitle feature is used to identify a character structure of a sample subtitle in the sample video frame, the sample video frame is a video frame in a sample video, the sample video is the target video that displays the sample subtitle, and the sample video frame corresponds to a target video frame in the target video; The display unit is configured to display a subtitle to be tested corresponding to the audio information in the target video through a program under test according to the audio information in the target video, to obtain a video under test; The first determination unit is configured to determine a to-be-tested subtitle feature according to a to-be-tested video frame corresponding to the target video frame in the to-be-tested video, where the to-be-tested subtitle feature is used to identify a character structure of a to-be-tested subtitle in the to-be-tested video frame; The second determination unit is configured to determine whether there is an abnormality in the subtitle display of the to-be-tested program in the to-be-tested video frame according to the sample subtitle feature and the to-be-tested subtitle feature.

14. The device according to claim 13, wherein The first determination unit is specifically configured to: Determine subtitle pixel points among the pixel points included in the to-be-tested video frame, where the subtitle pixel points are pixel points having a color corresponding to the to-be-tested subtitle; Determine a to-be-tested subtitle feature corresponding to the to-be-tested video frame according to the subtitle pixel points.

15. The device according to claim 14, characterized in that, The to-be-tested video frame includes N columns of pixel points, and the first determination unit is specifically configured to: Determine a first quantity of the subtitle pixel points respectively included in the N columns of pixel points; Determine a to-be-tested subtitle feature corresponding to the to-be-tested video frame according to the first quantity and the arrangement relationship among the N columns of pixel points.

16. The device according to claim 14, characterized in that, The to-be-tested video frame includes M rows of pixel points, and the first determination unit is specifically configured to: Determine a second quantity of the subtitle pixel points respectively included in the M rows of pixel points; Determine a to-be-tested subtitle feature corresponding to the to-be-tested video frame according to the second quantity and the arrangement relationship among the M rows of pixel points.

17. The device according to claim 14, characterized in that, The first determination unit is specifically configured to: Determine a to-be-tested subtitle feature corresponding to the to-be-tested video frame according to the distribution relationship between the subtitle pixel points and non-subtitle pixel points in the to-be-tested video frame.

18. The device according to claim 13, characterized in that, The second determination unit is specifically configured to: Determine the matching degree between the sample subtitle and the to-be-tested subtitle according to the sample subtitle feature and the to-be-tested subtitle feature; If the matching degree meets a matching threshold, determine that there is no abnormality in the subtitle display of the to-be-tested program in the to-be-tested video frame; If the matching degree does not meet the matching threshold, determine that there is an abnormality in the subtitle display of the to-be-tested program in the to-be-tested video frame.

19. The device according to claim 13, characterized in that, The apparatus further includes: A third determination unit, configured to determine a to-be-tested subtitle display parameter of the to-be-tested program, where the to-be-tested subtitle display parameter is used to identify a display form of the subtitle display by the to-be-tested program; A second acquisition unit, configured to acquire a sample subtitle display parameter for displaying the sample subtitle in the sample video; The second determination unit is specifically configured to: Convert the to-be-tested subtitle feature into a conversion feature conforming to the sample subtitle display parameter according to the mapping relationship of subtitle pixel points between the to-be-tested subtitle display parameter and the sample subtitle display parameter; Determine whether there is an abnormality in the subtitle display of the to-be-tested program in the to-be-tested video frame according to the sample subtitle feature and the conversion feature.

20. The device according to claim 19, characterized in that, The display form includes a display position, and the apparatus further includes: A fourth determination unit, configured to determine an actual subtitle position parameter corresponding to the to-be-tested subtitle in the to-be-tested video frame, where the actual subtitle position parameter is used to determine an actual display position corresponding to the to-be-tested subtitle; A fifth determination unit, configured to determine whether an abnormality occurs in the display form of the subtitle display of the program to be tested in the video frame to be tested according to the target display position identified by the subtitle display parameter to be tested and the actual display position.

21. The device according to claim 19, characterized in that, The display form includes a display zoom level, and the apparatus further includes: A sixth determination unit, configured to determine an actual subtitle zoom parameter corresponding to the subtitle to be tested in the video frame to be tested, where the actual subtitle zoom parameter is used to determine an actual display zoom level corresponding to the subtitle to be tested; A seventh determination unit, configured to determine whether an abnormality occurs in the display form of the subtitle display of the program to be tested in the video frame to be tested according to the target display zoom level identified by the subtitle display parameter to be tested and the actual display zoom level.

22. The device according to claim 20 or 21, characterized in that, The apparatus further includes: An eighth determination unit, configured to determine that an abnormality occurs in the subtitle display of the program to be tested in the video frame to be tested if it is determined that an abnormality occurs in the display form of the subtitle display of the program to be tested in the video frame to be tested.

23. The device according to claim 13, characterized in that, The sample video frame includes a video frame corresponding to the start display and / or end display moment of the sample subtitle in the sample video; Alternatively, the sample video frame is collected in the sample video based on a video frame sampling interval.

24. The device according to claim 13, characterized in that, The video content in the target video has a single content color, and the content color is different from the corresponding color of the subtitle to be tested.

25. A computer device, characterized in that, The device includes a processor and a memory: The memory is configured to store program code and transmit the program code to the processor; The processor is configured to execute the subtitle detection method according to any one of claims 1-12 based on the instructions in the program code.

26. A computer-readable storage medium, characterized in that, The computer-readable storage medium is configured to store a computer program, and the computer program is configured to implement the subtitle detection method according to any one of claims 1-12 when being executed by a processor.

Citation Information

Patent Citations

  • Text detection method and device

    CN112749696A

  • Caption checking apparatus and caption checking method

    JP2009260823A