Subtitle Area Recognition Method, Device, Equipment and Storage Medium

By identifying text in the video and filtering candidate subtitle areas, automated subtitle extraction is realized, solving the problem of inefficient subtitle extraction in the prior art, saving human resources and improving efficiency.

CN112232260BActive Publication Date: 2025-06-13TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202011165751.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-10-27
Publication Date
2025-06-13
Estimated Expiration
2040-10-27

AI Technical Summary

Technical Problem

In the prior art, subtitle extraction in short videos requires a lot of manpower to perform manual annotation and OCR recognition, which is inefficient.

Method used

By identifying the text in the video, the text list is obtained, and the subtitle area is sorted into candidate subtitle areas, and the subtitle area is filtered out from it according to the subtitle area filtering strategy to achieve automatic subtitle extraction.

Benefits of technology

Save human resources, improve the speed and efficiency of subtitle recognition, and realize the automated subtitle extraction process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112232260B_ABST
    Figure CN112232260B_ABST
Patent Text Reader

Abstract

The present application discloses a subtitle area recognition method, device, equipment and storage medium, which relates to the computer vision technology of artificial intelligence. The method includes: recognizing the text in the video to obtain a text list, the text list includes at least one piece of text data, the text data includes text content, text area and display duration, and the text content includes at least one text located on the text area; regularizing the text areas into n candidate subtitle areas, and the position deviation between the text areas belonging to the i-th candidate subtitle area and the i-th candidate subtitle area is less than the deviation threshold; screening the subtitle area from the n candidate subtitle areas according to the subtitle area screening strategy; the subtitle area screening strategy is used to determine the candidate subtitle area with the lowest repetition rate of the text content and the longest total display duration among the n candidate subtitle areas as the subtitle area. This method can save the human resources required for subtitle area recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of computer vision in artificial intelligence, and particularly to a method, apparatus, device, and storage medium for subtitle area recognition. Background Art

[0002] With the popularity of short videos, subtitle extraction technology in videos is required in various scenarios. For example, during the training process of a speech-to-text model, subtitles in the video need to be used as training samples.

[0003] In related technologies, since the text information in short videos may not necessarily be subtitle text, it may also include brand watermark text, video title text, etc. Therefore, for subtitle extraction in short videos, manual subtitle area annotation is performed, and then OCR (Optical Character Recognition) technology is used to recognize the text at the annotated position to obtain subtitles. For example, a video is manually screenshot, and then the screenshot is opened with an image viewing software. By moving the mouse to the upper left and lower right positions of the subtitle, the coordinates of the two positions can be obtained, and thus the position of the subtitle can be obtained.

[0004] The method in related technologies requires a large amount of manpower for subtitle extraction. Summary of the Invention

[0005] Embodiments of this application provide a method, apparatus, device, and storage medium for subtitle area recognition, which can automatically perform subtitle extraction and save human resources. The technical solutions are as follows.

[0006] According to one aspect of this application, a method for subtitle area recognition is provided. The method includes:

[0007] Recognize the text in the video to obtain a text list. The text list includes at least one piece of text data. The text data includes text content, text area, and display duration. The text content includes at least one character located on the text area;

[0008] Regularize the text areas into n candidate subtitle areas. The position deviation between the text area belonging to the i-th candidate subtitle area and the i-th candidate subtitle area is less than a deviation threshold. n is a positive integer, and i is a positive integer less than or equal to n;

[0009] The subtitle area is screened from the n candidate subtitle areas according to the subtitle area screening strategy; the subtitle area screening strategy is used to determine, as the subtitle area, the candidate subtitle area among the n candidate subtitle areas with a repetition rate of the text content lower than a repetition rate threshold and the longest total display duration, where the total display duration is the sum of the display durations of all the text content belonging to the candidate subtitle area.

[0010] According to another aspect of the present application, there is provided a subtitle recognition device, the device includes:

[0011] An identification module, configured to identify text in a video to obtain a text list, the text list includes at least one piece of text data, the text data includes text content, a text area, and a display duration, and the text content includes at least one character located on the text area;

[0012] A candidate module, configured to regularize the text areas into n candidate subtitle areas, where the position deviation between the text area belonging to the i-th candidate subtitle area and the i-th candidate subtitle area is less than a deviation threshold, n is a positive integer, and i is a positive integer less than or equal to n;

[0013] A screening module, configured to screen the subtitle area from the n candidate subtitle areas according to the subtitle area screening strategy; the subtitle area screening strategy is used to determine, as the subtitle area, the candidate subtitle area among the n candidate subtitle areas with a repetition rate of the text content lower than a repetition rate threshold and the longest total display duration, where the total display duration is the sum of the display durations of all the text content belonging to the candidate subtitle area.

[0014] According to another aspect of the present application, there is provided a computer device, the computer device includes: a processor and a memory, and at least one instruction, at least one program, a code set, or an instruction set is stored in the memory, and the at least one instruction, the at least one program, the code set, or the instruction set is loaded and executed by the processor to implement the subtitle area recognition method described in the above aspect.

[0015] According to another aspect of the present application, there is provided a computer-readable storage medium, and at least one instruction, at least one program, a code set, or an instruction set is stored in the storage medium, and the at least one instruction, the at least one program, the code set, or the instruction set is loaded and executed by a processor to implement the subtitle area recognition method described in the above aspect.

[0016] According to another aspect of the embodiments of the present disclosure, there is provided a computer program product or a computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the subtitle area recognition method provided in the above optional implementation manners.

[0017] The beneficial effects brought by the technical solutions provided in the embodiments of this application at least include the following beneficial effects.

[0018] By using a subtitle area screening strategy, the text areas in the text list recognized from the video are screened to obtain the subtitle area, so that the subtitles of the video to be recognized can be extracted according to the subtitle area. Compared with the method of manually annotating the subtitle area, this method saves the human resources required for subtitle recognition and speeds up the subtitle recognition speed and efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0020] Figure 1 is a block diagram of a computer system provided by an exemplary embodiment of the present application;

[0021] Figure 2 is a flowchart of a subtitle area recognition method provided by an exemplary embodiment of the present application;

[0022] Figure 3 is a schematic diagram of a video frame image of a subtitle area recognition method provided by another exemplary embodiment of the present application;

[0023] Figure 4 is a schematic diagram of a video frame image of a subtitle area recognition method provided by another exemplary embodiment of the present application;

[0024] Figure 5 is a flowchart of a subtitle area recognition method provided by another exemplary embodiment of the present application;

[0025] Figure 6 is a schematic diagram of a video frame image of a subtitle area recognition method provided by another exemplary embodiment of the present application;

[0026] Figure 7 is a schematic diagram of a text area of a subtitle area recognition method provided by another exemplary embodiment of the present application;

[0027] Figure 8 is a flowchart of a subtitle area recognition method provided by another exemplary embodiment of the present application;

[0028] Figure 9 is a flowchart of a subtitle area recognition method provided by another exemplary embodiment of the present application;

[0029] Figure 10 is a block diagram of a subtitle recognition device provided by another exemplary embodiment of the present application;

[0030] Figure 11 is a schematic structural diagram of a server provided by another exemplary embodiment of the present application;

[0031] Figure 12 is a block diagram of a terminal provided by another exemplary embodiment of the present application. Detailed implementation manners

[0032] To make the objectives, technical solutions, and advantages of the present application clearer, the following will further describe the embodiments of the present application in detail with reference to the accompanying drawings.

[0033] First, a brief introduction to several nouns related to the embodiments of the present application is given.

[0034] Artificial Intelligence (AI) is to use a digital computer or a machine controlled by a digital computer to simulate, extend, and expand human intelligence, and is a theory, method, technology, and application system for perceiving the environment, acquiring knowledge, and using knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science, which attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines, enabling the machines to have the functions of perception, reasoning, and decision-making.

[0035] Artificial intelligence technology is an interdisciplinary subject, involving a wide range of fields, including both hardware-level technologies and software-level technologies. Artificial intelligence basic technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technologies, operation / interaction systems, and mechatronics. Artificial intelligence software technologies mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0036] Computer Vision Technology (CV) is a science that studies how to enable machines to "see". Further speaking, it refers to using cameras and computers to replace the human eye for object recognition and measurement, etc., which is machine vision, and further performing graphic processing to make the computer-processed images more suitable for human eye observation or transmission to instrument detection. As a scientific discipline, computer vision researches related theories and technologies, and attempts to establish an artificial intelligence system that can obtain information from images or multi-dimensional data. Computer vision technology usually includes image processing, image recognition, image semantic understanding, image retrieval, OCR (Optical Character Recognition), video processing, video semantic understanding, video content / behavior recognition, three-dimensional object reconstruction, 3D (Three Dimensional) technology, virtual reality, augmented reality and other technologies.

[0037] OCR is the abbreviation of the English Optical Character Recognition, which means optical character recognition, and can also be simply called character recognition. It is a method of automatic text input. It obtains the text image information on paper through optical input methods such as scanning and photography. By using various pattern recognition algorithms to analyze the morphological characteristics of the text, it can convert bills, newspapers, books, manuscripts and other printed materials into image information, and then use character recognition technology to convert the image information into computer input technology that can be used.

[0038] Figure 1 The structural schematic diagram of a computer system provided by an exemplary embodiment of the present application is shown. The computer system includes a terminal 120 and a server 140.

[0039] The terminal 120 and the server 140 are connected to each other through a wired or wireless network.

[0040] The terminal 120 includes at least one of a smart phone, a laptop computer, a desktop computer, a tablet computer, a smart speaker, and a smart robot. In an optional implementation manner, the terminal uploads the video that needs to be subtitle-recognized to the server, and the server performs subtitle recognition on the video uploaded by the terminal. In another optional manner, the server can also perform subtitle recognition on the locally stored video. In another optional manner, the terminal can also perform subtitle recognition on the locally stored video. In another optional manner, the terminal can also download the video through the network and perform subtitle recognition on the downloaded video.

[0041] Exemplarily, the terminal 120 further includes a display; the display is used to display the picture of the video.

[0042] The terminal 120 includes a first memory and a first processor. A first program is stored in the first memory; the first program is called and executed by the first processor to implement the subtitle area recognition method provided in this application. The first memory may include, but is not limited to, the following: Random Access Memory (RAM), Read-Only Memory (ROM), Programmable Read-Only Memory (PROM), Erasable Programmable Read-Only Memory (EPROM), and Electric Erasable Programmable Read-Only Memory (EEPROM).

[0043] The first processor may be composed of one or more integrated circuit chips. Optionally, the first processor may be a general-purpose processor, for example, a Central Processing Unit (CPU) or a Network Processor (NP). Optionally, the first processor may implement the subtitle area recognition method provided in this application by calling a subtitle recognition algorithm.

[0044] The server 140 includes a second memory and a second processor. A second program is stored in the second memory, and the second program is called by the second processor to implement the subtitle area recognition method provided in this application. Exemplarily, a subtitle recognition algorithm is stored in the second memory. In an optional implementation manner, the server receives a video sent by the terminal and uses the subtitle recognition algorithm to perform subtitle recognition. Optionally, the second memory may include, but is not limited to, the following: RAM, ROM, PROM, EPROM, EEPROM. Optionally, the second processor may be a general-purpose processor, for example, a CPU or an NP.

[0045] The server 140 may be an independent physical server, or a server cluster or distributed system composed of multiple physical servers. It may also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The terminal may be a smart phone, a tablet computer, a notebook computer, a desktop computer, a smart speaker, a smart watch, etc., but is not limited thereto. The terminal and the server may be directly or indirectly connected through wired or wireless communication methods, and this application does not make any restrictions here.

[0046] Schematically, the subtitle area recognition method provided by the present application can be applied to scenarios such as video subtitle extraction and obtaining training samples for speech-to-text models. Taking the example of using the subtitle area recognition method provided by the present application to obtain training samples for a speech-to-text model, after obtaining the subtitle area of a video, obtain the text area belonging to the subtitle area and the corresponding text data. The text content in the text data is the text part of the training sample. According to the display duration (start time and end time) in the text data, intercept the audio corresponding to the time from the video. This audio is the speech part of the training sample, and the text part and the speech part are stored corresponding to each other as training samples.

[0047] Figure 2 The flowchart of the subtitle area recognition method provided by an exemplary embodiment of the present application is shown. This method can be executed by a computer device. For example, it can be executed by a terminal or a server as shown. The method includes the following steps. Figure 1 As shown.

[0048] Step 201, recognize the text in the video to obtain a text list. The text list includes at least one piece of text data. The text data includes text content, text area, and display duration. The text content includes at least one text located on the text area.

[0049] Exemplarily, the video can be any type of video file. For example, short videos, TV dramas, movies, variety shows, etc. Exemplarily, the video includes subtitles. Taking a short video as an example, the text in the short video screen not only includes subtitles but may also include other text information. For example, the watermark text of the short video application, the user nickname of the short video publisher, the video name of the short video, and so on. Therefore, it is impossible to accurately obtain the subtitles of a short video only by using OCR technology for text recognition. And the method of manually annotating the subtitle area and then performing text recognition on the annotated position to obtain subtitles requires a lot of manpower. Therefore, the present application provides a subtitle recognition method that can accurately identify subtitles from multiple text information in the video, saving the step of manually annotating the subtitle area and improving the efficiency of subtitle extraction.

[0050] Exemplarily, the video can be obtained in any way. The video can be a video file stored locally on a computer device or a video file obtained from other computer devices. For example, when the computer device is a server, the server can receive the video file uploaded by the terminal; when the computer device is a terminal, the terminal can also download the video file stored on the server through the network. Taking the computer device as a server as an example, a client with a font model extraction function can be installed on the terminal. The user can select the video file stored locally on the user interface of the client and click the upload control to upload the video file to the server. The server performs subsequent subtitle area recognition processing on the video file.

[0051] Exemplarily, the computer device performs character recognition on the video to obtain a text list. Exemplarily, the text list can be a data table, where each row represents a piece of text data, and each column represents the specific content of the text data: the text content, the text area, and the display duration. For a video frame image of the video, different areas on the image may contain different text contents. For multiple video frame images of the video, the same area on the image may also display different text contents at different times. Therefore, by extracting multiple text contents with different text areas and different display times in the video, multiple pieces of text data can be obtained, which form a text list. Exemplarily, if the same text content is displayed in the same text area at different time periods in the video, then these two text contents belong to two pieces of text data respectively. That is, if the same text content is displayed in the same text area on consecutive video frame images, then this text content belongs to one piece of text data, and the duration of the consecutive video frame images is the display duration (the display duration of the text content) in this piece of text data. For example, the first text content is displayed in the first area on the video frame images from the 1st to the 3rd second (s), no text is displayed in the first area on the video frame images from the 3rd to the 4th second, and the first text content is displayed again in the first area on the video frame images from the 4th to the 5th second. Then these two first text contents correspond to two pieces of text data respectively, and the display durations in the two pieces of text data are 2 s and 1 s respectively.

[0052] Exemplarily, the text list can also be a data set, a database, a document file, etc. composed of multiple pieces of text data.

[0053] Exemplarily, the text area includes the position of a text box for enclosing the text. Exemplarily, the text box is a rectangular box, and the position of the text box can be expressed by the positions of four lines (the upper line, the lower line, the left line, and the right line), or by the coordinates of the four vertices of the text box, or by the coordinates of two diagonal vertices of the text box.

[0054] Step 202: Regularize the text areas into n candidate subtitle areas. The position deviation between the text area belonging to the i-th candidate subtitle area and the position of the i-th candidate subtitle area is less than the deviation threshold, where n is a positive integer and i is a positive integer less than or equal to n.

[0055] Exemplarily, regularization means classifying the text areas according to the position distribution of the text areas, and classifying multiple text areas with a position deviation less than the deviation threshold into the same type of text areas, that is, the same candidate subtitle area.

[0056] Exemplarily, after obtaining the text list, the text list includes multiple text regions. Since the subtitles of a video are usually displayed in the same region position, these text regions are thus consolidated to obtain multiple candidate subtitle regions. Exemplarily, since the content of different subtitle texts is different, the displayed region ranges may also have some differences. For example, as Figure 3 in (1) and (2) in [reference] are two video frame images of the video, with the first text content located in the first text region 501 and the second text content located in the second text region 502 on the two video frame images respectively. Both of these text contents are subtitles. However, due to the different numbers of characters and lines of the text contents, there are some differences in the text regions of these two text contents. However, both of these text regions are subtitle regions. Therefore, when consolidating the candidate subtitle regions, a deviation threshold needs to be set. If the position deviation between two text regions is less than the deviation threshold, it should be considered that these two text regions belong to the same candidate subtitle region. In this way, the multiple text regions in the text list can be consolidated, and finally several candidate subtitle regions can be obtained.

[0057] Exemplarily, taking the calculation of the position deviation between the first text region and the second text region as an example, the first text region includes a first upper line, a first lower line, a first left line, and a first right line, and the second text region includes a second upper line, a second lower line, a second left line, and a second right line. The position deviation includes at least one of the deviation between the first upper line and the second upper line, the deviation between the first lower line and the second lower line, the deviation between the first left line and the second left line, and the deviation between the first right line and the second right line. Exemplarily, since subtitles are usually horizontally displayed subtitles, due to the different numbers of characters in the text content, the position differences in the left-right direction of the text regions are relatively large, and the position differences in the up-down direction are relatively small. Then the position deviation can include the deviation between the two upper lines and the deviation between the two lower lines of the two text regions, that is, the text regions with relatively close vertical positions are grouped into the same candidate subtitle region. Exemplarily, since some subtitles are vertically displayed subtitles, the position deviation can also include the deviation between the two left lines and the deviation between the two right lines of the two text regions, that is, the text regions with relatively close horizontal positions are grouped into the same candidate subtitle region.

[0058] Exemplarily, the specific value of the deviation threshold can be arbitrary. Exemplarily, after repeated experiments, it is found that it is better to take the deviation threshold as 30 pixels - 50 pixels. For example, if the deviation threshold is set to 40 pixels, then the two text regions with the deviation between the two upper lines less than 40 pixels and the deviation between the two lower lines also less than 40 pixels are grouped into the same candidate subtitle region.

[0059] Exemplarily, a candidate subtitle region has a region position, that is, where the candidate subtitle region is located. Exemplarily, the region position of the candidate subtitle region is the largest text region belonging to the candidate subtitle region. Exemplarily, the region position of the candidate subtitle region is the text region with the largest height belonging to the candidate subtitle region (corresponding to horizontally displayed subtitles), or the region position of the candidate subtitle region is the text region with the largest width belonging to the candidate subtitle region (corresponding to vertically displayed subtitles).

[0060] Exemplarily, after regularizing the text regions into multiple candidate subtitle regions, a column of data for the candidate subtitle regions can be added to the text list. Then, a piece of data for the candidate subtitle region to which each text content belongs is added to each piece of text data. Thus, each text content corresponds to a text region, a display duration, and a candidate subtitle region.

[0061] Step 203: Screen out subtitle regions from the n candidate subtitle regions according to the subtitle region screening strategy; the subtitle region screening strategy is used to determine, as the subtitle region, the candidate subtitle region with the lowest repetition rate of the text content and the longest total display duration among the n candidate subtitle regions. The total display duration is the sum of the display durations of all text contents belonging to the candidate subtitle region.

[0062] Exemplarily, after obtaining the candidate subtitle regions, the computer device can call the algorithm of the subtitle region screening strategy to identify the subtitle regions of the video from the candidate subtitle regions. Exemplarily, since some interfering text (non-subtitle text) that may appear in the video includes video titles, application watermarks, user nicknames, etc., and these interfering text has the characteristics of long display time and single and unchanging displayed text, the subtitle regions can be screened out from the text data according to these characteristics of the interfering text.

[0063] Exemplarily, the subtitle region screening strategy is set according to the display characteristics of the interfering text and the display characteristics of the subtitles. Subtitles have characteristics such as long display time, fixed position, and diverse text contents.

[0064] For the subtitle region screening strategy provided in this application, first, it is determined whether a single text content is displayed on each candidate subtitle region. If it is a single text content, then the candidate subtitle region is not a subtitle region. Then, the candidate subtitle region with the longest total display duration is selected from the remaining candidate subtitle regions as the subtitle region. Since some interfering text, such as the title text of a TV drama, will only be displayed in the first few seconds of the video and will not be displayed later. For example, Figure 4As shown in the figure, a video title 401 and subtitles 402 are displayed on the video frame image. The video title 401 will disappear after being displayed for a while, and no text will be displayed at this position anymore, while text will be displayed at the position of the subtitles 402 for a long time. Therefore, the candidate subtitle area with the longest total display duration is selected from the remaining candidate subtitle areas as the subtitle area.

[0065] In summary, the method provided in this embodiment screens the text areas in the text list recognized from the video through the subtitle area screening strategy to obtain the subtitle area, so that the subtitles of the video can be extracted according to the subtitle area. Compared with the method of manually annotating the subtitle area, this method saves the human resources required for subtitle recognition and speeds up the subtitle recognition speed and efficiency.

[0066] Exemplarily, an exemplary embodiment of subtitle area screening according to the subtitle area screening strategy is given.

[0067] Figure 5 The flowchart of the subtitle area recognition method provided by an exemplary embodiment of the present application is shown. This method can be executed by a computer device. For example, it can be executed by a terminal or a server as shown. Figure 1 As shown. Figure 2 Based on the exemplary embodiment shown, step 201 further includes steps 2011 to 2012, step 202 further includes steps 2021 to 2025, and step 203 further includes steps 2031 to 2034.

[0068] Step 2011, periodically intercept the video frame images of the video.

[0069] Exemplarily, first, frame extraction processing needs to be performed on the video. Frame extraction processing is to periodically intercept video frame images from the video and store them sequentially. Exemplarily, the time interval (period) for intercepting video frame images from the video can be arbitrary. For example, 2 video frame images are intercepted per second. Exemplarily, each frame of the video can also be intercepted as a video frame image. Exemplarily, multiple video frame images can be intercepted from a video.

[0070] Step 2011, recognize the text in the video frame image to obtain a text list.

[0071] Exemplarily, the computer device performs text recognition on each frame of the video frame image to obtain a text list.

[0072] Exemplarily, an optical character recognition (OCR) model is called to recognize a video frame image, obtaining candidate text content in the video frame image and the text region of the candidate text content, and obtaining the display time of the candidate text content according to the display time of the video frame image; the candidate text content is de-duplicated to obtain text content; the de-duplication includes determining, as the text content, the candidate text content with the earliest display time among multiple candidate text contents with consecutive display times, the same text region, and the same candidate text content, and calculating the display duration of the text content according to the display times of the multiple candidate text contents; a text list is generated according to the text content, the text region of the text content, and the display duration.

[0073] Exemplarily, an OCR model is called to recognize the text in a video frame image, and the OCR model outputs the candidate text content in the video frame image and the text region of the candidate text content. In this way, a data table can be obtained that includes: candidate text content, text region, and display time.

[0074] Among them, the display time of the video frame image refers to the time when the video frame image is displayed in the video. The display time of the candidate text content extracted from the video frame image is the same as the display time of the video frame image.

[0075] The OCR model is used to perform text recognition on the video frame image, recognize the text in the video frame image, and output the text and the text region. Exemplarily, the OCR model is a neural network model, and any known OCR model can be used.

[0076] For example, as Figure 6 shown, in a video frame image of a video, three pieces of text are displayed: the first text 301, the second text 302, and the third text 303. The OCR model recognizes these three pieces of text and outputs: for the candidate text content of the first text 301: "《Thirty-something》 Mom can do anything for her child", text region: the left boundary position x1 = 2, right boundary position x2 = 8, upper boundary position y1 = 10, lower boundary position y2 = 8 of the first text box 304; for the candidate text content of the second text 302: "Why are you drinking", text region: the left boundary position x3 = 3, right boundary position x4 = 7, upper boundary position y3 = 6, lower boundary position y4 = 5 of the second text box 305; for the candidate text content of the third text 303: "WS TV series", text region: the left boundary position x5 = 4, right boundary position x6 = 6, upper boundary position y5 = 3, lower boundary position y6 = 2 of the third text box 306.

[0077] Exemplarily, a video frame image corresponds to a display moment in the video. When intercepting a video frame image, the video frame images are stored in chronological order, and the display moment corresponding to the video frame image in the video is stored. For example, when intercepting the video frame at the 1st second in the video to obtain the video frame image at the 1st second, the video frame image is stored corresponding to the 1st second.

[0078] Therefore, the candidate text content recognized from each video frame image can also correspond to the display moment of the video frame image in the video. For a candidate text content, it is possible to sequentially search in subsequent video frame images to find whether there is a candidate text content that is the same as the candidate text content and has the same text area. If there is, it is determined that these candidate text contents are the same text content, and the display duration of the text content can be obtained according to the display moment of the video frame image corresponding to the first occurrence of the candidate text content and the display moment of the video frame image corresponding to the last occurrence. Exemplarily, this search is continuous. When the candidate text content is not found in the next video frame image, the search stops. That is, multiple candidate text contents that are continuous in time, have the same text area, and have the same candidate text content are merged into one text content.

[0079] For example, as shown in Table 1, after text recognition by the OCR model, 7 candidate text contents are recognized from 7 video frame images from the 1st second to the 7th second. Among them, the first "Hello" appears in the text areas (1,1) and (2,2) from the 1st second to the 4th second. Then it is determined that these four candidate text contents "Hello" are the same text content. According to the first moment 1s and the last moment 4s when it appears, the display duration of this text content can be calculated as 3s; similarly, the display duration of the second "Hello" can be obtained as 1s. For the candidate text content displayed on only one video frame image, it is directly used as the text content, and its display duration can be set as the time interval for intercepting the video frame image. For example: 1s. Therefore, after merging the candidate text contents, the text content shown in Table 2 can be obtained.

[0080] Table 1

[0081]

[0082]

[0083] Table 2

[0084] Text content Text area Display duration Hello (1,1),(2,2) 3s Hi (1,1),(2,2) 1s Hello (1,1),(2,2) 1s

[0085] Exemplarily, the text list includes at least one piece of text data of at least one text content, and one text content corresponds to one text area and one display duration.

[0086] Exemplarily, the display duration in the text list also needs to include the start time and end time of the display, that is, store the start time and end time as the display duration, and the display duration can be calculated based on the start time and end time. For example, after the computer device obtains a video, it generates a video link for the video, and then recognizes the text in the video to obtain the text list shown in Table 3. Among them, the text area is described by the left line x1, right line x2, upper line y1, and lower line y2 of the rectangle, and the display duration is described by the start time "startTime" and end time "endTime".

[0087] Table 3

[0088]

[0089] Step 2021: Extract one text area from the m text areas as the first text area, determine the first text area as the first candidate subtitle area, and add the first candidate subtitle area to the candidate subtitle area list.

[0090] Step 2022: Loop through steps 2022 to 2023 until the remaining number of the m text areas is 0: Extract one text area from the remaining m - k + 1 text areas as the kth text area.

[0091] Step 2023: Determine whether the position deviation between the kth text area and the candidate subtitle area is greater than the deviation threshold. If it is greater than (or equal to), perform step 2025; if it is less than (or equal to), perform step 2024.

[0092] Step 2024: In response to the first position deviation between the kth text area and the wth candidate subtitle area in the candidate subtitle area list being less than the deviation threshold, classify the kth text area into the wth candidate subtitle area.

[0093] Exemplarily, after classifying the kth text area into the wth candidate subtitle area, calculate the first height of the kth text area, where the first height is the difference between the upper line and the lower line of the kth text area;

[0094] Calculate the second height of the wth candidate subtitle area, where the second height is the difference between the upper line and the lower line of the wth candidate subtitle area; In response to the first height being greater than the second height, determine the kth text area as the wth candidate subtitle area; where k is a positive integer less than or equal to m, w is a positive integer less than or equal to n, and n and m are positive integers.

[0095] Step 2025: In response to the second position deviation between the k-th text region and all candidate subtitle regions in the candidate subtitle region list being greater than the deviation threshold, determine the k-th text region as the y-th candidate subtitle region, and add the y-th candidate subtitle region to the candidate subtitle region list.

[0096] Among them, the first position deviation includes the difference between two upper lines and the difference between two lower lines. The second position deviation includes the difference between two upper lines or the difference between two lower lines. y is a positive integer less than or equal to n, k is a positive integer less than or equal to m, w is a positive integer less than or equal to n, and m and n are positive integers.

[0097] Exemplarily, steps 2021 to 2025 are method steps for regularizing text regions to obtain candidate subtitle regions. Taking the example where the text list includes m text data and the text regions are described by the positions of the upper and lower lines of a rectangle.

[0098] Exemplarily, it is possible to sequentially read from the first text region according to the arrangement order of the text data in the text list (which can be any sorting method), directly use the first text region as a candidate subtitle region and put it into the candidate subtitle region list, and then start comparing the second text region with the existing candidate subtitle regions in the candidate subtitle region list to see if it can match the existing candidate subtitle regions (the difference between the upper lines of the two regions should be less than the deviation threshold and the deviation of the lower lines should also be less than the deviation threshold). If there is a matching candidate subtitle region, then assign this text region to this candidate subtitle region; if there is no matching candidate subtitle region, then use this text region as a new candidate subtitle region and store it in the candidate subtitle region list; in this way, traverse each text region in the text list to obtain the candidate subtitle regions stored in the candidate subtitle region list.

[0099] Exemplarily, a candidate subtitle region may contain multiple text regions, but the region position of the candidate subtitle region (including the upper and lower lines) is only one. The region position of the candidate subtitle region is the text region with the highest height (upper and lower lines) among the text regions belonging to this candidate subtitle region.

[0100] Therefore, after assigning a text region to a candidate subtitle region, it is necessary to determine whether the height of the newly added text region is greater than the height of the current region position of the candidate subtitle region. If the height of the newly added text region is greater, then update the newly added text region as the region position of the candidate subtitle region. If the height difference of the newly added text region is less than the current region position of the candidate subtitle region, then keep the current region position of the candidate subtitle region unchanged.

[0101] Exemplarily, in another alternative implementation, first calculate the height difference of each text region, then sort the text regions in ascending order of the height difference to obtain a text region order list, and read and determine the candidate subtitle region starting from the first text region according to the order of the text region order list. This method can solve the problem of inaccurate determination of candidate subtitle regions. For example, as Figure 7 shown, taking the first text region 701, the second text region 702, and the third text region 703 as examples, where the first text region 701 is smaller than the third text region 703 which is smaller than the second text region 702, and the position deviation between the first text region 701 and the second text region 702 is greater than the deviation threshold, the position deviation between the second text region 702 and the third text region 703 is less than the deviation threshold, and the position deviation between the first text region 701 and the third text region 703 is less than the deviation threshold. If the text regions are extracted in the order of the first text region 701, the second text region 702, and the third text region 703, when the second text region 702 is extracted, since the position deviation between the second text region 702 and the first text region 701 is greater than the deviation threshold, the second text region 702 will be used as a new candidate subtitle region, resulting in inaccurate recognition results of the candidate subtitle region; however, if the text regions are sorted according to the height difference, the third text region 703 will be extracted first after the first text region 701 is extracted. The position deviation between the third text region 703 and the first text region 701 is less than the deviation threshold, and the height difference of the third text region 703 is greater than that of the first text region 701. Then the region position of the candidate subtitle region will be updated to the third sub-region 703. When the second text region 702 is extracted later, since the position deviation between the second text region 702 and the third text region 703 is less than the deviation threshold, the second text region 702 will also be classified into the candidate subtitle region, and the second text region 702 will be updated to the region position of the candidate subtitle region.

[0102] Exemplarily, due to the habitual reading order, most subtitles are horizontal subtitles. Steps 2021 to 2025 take horizontal subtitles as an example, and use the upper line and the lower line as the text region; similarly, if vertical subtitles are to be recognized, the above-mentioned upper line and lower line are changed to the left line and the right line, that is, the text region is the left line and the right line.

[0103] Step 2031, calculate the repetition rate of the candidate subtitle region. The repetition rate is the ratio of the cumulative duration to the total video duration of the video, and the cumulative duration is the sum of the display durations of the same text content.

[0104] Exemplarily, a method for calculating the repetition rate is given: Obtain the j-th group of text data corresponding to the j-th candidate subtitle region. The text regions in the j-th group of text data belong to the j-th candidate subtitle region, where j is a positive integer less than or equal to n, and n is a positive integer; Group the text data with the same text content in the j-th group of text data into the same text data set to obtain at least one text data set; Calculate the sum of the display durations in each text data set to obtain at least one cumulative duration; Calculate the ratio of the maximum cumulative duration to the total video duration of the video to obtain the repetition rate; Repeat the above four steps to calculate the repetition rate of each candidate subtitle region.

[0105] That is, all the text data belonging to the candidate subtitle region will be obtained, and then the text data with the same text content will be merged: Only one text content is retained, and the display durations are accumulated to obtain the cumulative duration. Since the text positions are not needed here, they can be removed; There is no duplicate text content in the merged text data. Divide the maximum cumulative duration in the merged text data by the total video duration of the video to obtain the repetition rate.

[0106] The repetition rate is the proportion of the cumulative display duration of the same text content displayed in the candidate subtitle region to the total video duration. If the same text content is always displayed at a certain position, it is very likely that the text at that position is interfering text (such as video titles, watermarks, etc.).

[0107] Step 2032: Determine the candidate subtitle regions with a repetition rate of the text content lower than the repetition rate threshold as the preliminary screening subtitle regions.

[0108] Exemplarily, the repetition rate threshold can be set arbitrarily. Exemplarily, the repetition rate threshold can be taken as 10%.

[0109] Exemplarily, the candidate subtitle regions with a repetition rate higher than the repetition rate threshold may be the text regions where watermarks are located, the text regions where video titles are located, or the subtitle regions where other text content in the video remains fixed (with very few changes).

[0110] Step 2033: Calculate the total display duration of the preliminary screening subtitle regions.

[0111] Exemplarily, a method for calculating the total display duration is given: Calculate the sum of the display durations of the text data corresponding to the preliminary screening subtitle regions to obtain the total display duration of the preliminary screening subtitle regions.

[0112] Exemplarily, after initially screening the candidate subtitle regions to obtain the initially screened subtitle regions, calculate the total display duration of each initially screened subtitle region. The total display duration is the total duration during which the text content is displayed on the initially screened subtitle region. Since in a video, text may be briefly displayed at certain positions. For example, at the beginning of a TV drama, the episode number may be displayed in the middle of the screen, or some scenes with text may be briefly captured in the video. The regions where these texts are located are not subtitle regions. Subtitle regions will have text content displayed for a long time. Therefore, select the initially screened subtitle region with the longest total display duration among the initially screened subtitle regions as the subtitle region.

[0113] For example, in the first initially screened subtitle region, the first text content is displayed for 1s, the second text content is displayed for 2s, and the third text content is displayed for 6s. Then the total display duration of the first initially screened subtitle region is 1 + 2 + 6 = 9s.

[0114] Step 2034, determine the initially screened subtitle region with the longest total display duration among the initially screened subtitle regions as the subtitle region.

[0115] Exemplarily, of course, some other subtitle region screening strategies can also be adopted to screen the subtitle region.

[0116] For example, when determining the candidate subtitle regions based on the text regions, text regions with the inclination angle of the upper or lower border greater than the angle threshold can be directly removed and not used as candidate subtitle regions. Since subtitles are usually in a regular direction (horizontal or vertical), text data in an irregular direction can be directly removed.

[0117] Again, since subtitles are usually in white or black fonts, after obtaining the text list through recognition, text data corresponding to text content displayed in other colors can be deleted from the text list, and the method provided in this application can be used to identify the subtitle region with the remaining text list.

[0118] Exemplarily, after obtaining the subtitle region of the video, the computer device can recognize the subtitle of the video according to the text content belonging to the subtitle region.

[0119] For example, trim the text content in the text data corresponding to the subtitle region and use it as the subtitle of the video.

[0120] In summary, the method provided in this embodiment first obtains the video frame image of the video, then performs text recognition on the video frame image using the OCR model, removes duplicates from the candidate text content obtained by text recognition to obtain a text list containing text content, thereby extracting the text data in the video, which is convenient for determining the subtitle region according to the text data.

[0121] The method provided in this embodiment first regularizes to obtain candidate subtitle regions according to the text regions, and regularizes multiple text regions obtained through text recognition to obtain several approximate regions of the subtitle regions, which is convenient for subsequent recognition of subtitle regions according to the subtitle region recognition strategy.

[0122] The method provided in this embodiment determines whether a candidate subtitle region is a region for displaying watermarks, video titles, etc. with a long display time and a single display content by calculating the repetition rate of the text content displayed on each candidate subtitle region, and removes these candidate subtitle regions to obtain a preliminarily screened subtitle region.

[0123] The method provided in this embodiment removes regions that only display text content for a short time from the preliminarily screened subtitle regions by calculating the total display duration of each preliminarily screened subtitle region. Since subtitle regions usually display text content for a long time, according to this feature, the preliminarily screened subtitle region with the longest total display duration among the preliminarily screened subtitle regions can be determined as the subtitle region.

[0124] Exemplarily, an exemplary embodiment of obtaining a training sample of a speech-to-text model by using the method provided in this application is given.

[0125] Figure 8 The flowchart of the subtitle region recognition method provided in an exemplary embodiment of this application is shown. This method can be executed by a computer device. For example, it can be executed by a terminal or a server as shown in Figure 1 The method includes the following steps.

[0126] Step 601, the computer device performs data acquisition.

[0127] Exemplarily, first, videos of popular user accounts in a video application are obtained. A popular user account is a user account with a large number of fans, a large number of video clicks, or one of the top few in the leaderboard. Exemplarily, all videos under these popular accounts are obtained as the videos for which subtitle regions are to be recognized.

[0128] Step 602, the computer device performs a subtitle extraction service.

[0129] Exemplarily, the subtitle region recognition method provided in this application is used to recognize the subtitle regions in the video. For example, as shown in Figure 9As shown, first, video OCR frame extraction processing 802 is performed on UGC (User Generated Content) (extracting video frame images, performing text recognition on the video frame images to obtain recognition results, and removing duplicates from the recognition results to obtain a text list) to obtain text content, the display duration 803 of the text content, and the text area 804 of the text content. Then, the text area 804 is normalized to obtain multiple candidate subtitle areas, the repetition rate of each candidate subtitle area is calculated, and duplicate text judgment 805 is performed to select the preliminary screening subtitle areas with a repetition rate lower than the repetition rate threshold. Then, the total display duration of the preliminary screening subtitle areas is calculated, and duration judgment 806 is performed: The preliminary screening subtitle area with the longest total display duration (duration) is selected as the subtitle area 807.

[0130] Step 603, the computer device performs post-processing on the text content in the subtitle area.

[0131] For example, the post-processing includes at least one of short sentence merging, special symbol stripping, text density stripping, text word count stripping, duplicate recognition merging, and single letter and digit removal. Exemplarily, short sentence merging is used to merge ultra-short sentences (e.g., "ah", "okay") in the text content. Special symbol stripping is used to remove non-text data (e.g., emojis) in the text content. Text density stripping is used to remove overly long sentences from the text content. Text word count stripping is used to strip the text content according to the stripping word count. For example, every 2 - 14 words are stripped. Duplicate recognition merging is used to merge data of duplicate text content. Single letter and digit removal is used to remove single letters or digits of other non-target languages (e.g., Chinese) from the text content.

[0132] Step 604, the computer device verifies the delivery quality.

[0133] Exemplarily, the computer device uses the manual annotation results of video subtitles to verify the subtitles automatically recognized. Exemplarily, sampling detection is performed on the obtained subtitle recognition results, a test set is randomly constructed from the recognition results for confidence verification. If the confidence level is within the range of 95 ± 3%, it is determined that the recognition result is accurate, and the recognition result is delivered as data 605. The text content in the recognition result and the audio corresponding to the corresponding time period in the video are used as training samples for the speech-to-text model. Exemplarily, the confidence level is equal to: the ratio of the number of correctly recognized words in the subtitle recognition result to the total number of words in the subtitle recognition result.

[0134] In summary, the method provided in this embodiment can accurately identify the subtitle content in the video by using the subtitle area recognition method provided in this application. Then, based on the recognized subtitle content and the audio corresponding to the corresponding time period in the video, training samples for the speech-to-text model can be obtained. Training the speech-to-text model according to the subtitle content and the audio can save human resources in the sample acquisition process and improve the sample acquisition efficiency.

[0135] The following is an embodiment of the device of this application. For details not described in detail in the device embodiment, reference can be made to the corresponding records in the above method embodiment, which will not be elaborated herein.

[0136] Figure 10 The structural schematic diagram of a subtitle recognition device provided by an exemplary embodiment of this application is shown. This device can be implemented as all or part of a computer device through software, hardware, or a combination of both. This device includes the following devices.

[0137] An identification module 901, configured to identify the text in the video to obtain a text list. The text list includes at least one piece of text data. The text data includes text content, a text area, and a display duration. The text content includes at least one text located on the text area.

[0138] A candidate module 902, configured to normalize the text areas into n candidate subtitle areas. The position deviation between the text areas belonging to the i-th candidate subtitle area and the i-th candidate subtitle area is less than a deviation threshold. n is a positive integer, and i is a positive integer less than or equal to n.

[0139] A screening module 903, configured to screen the subtitle area from the n candidate subtitle areas according to a subtitle area screening strategy. The subtitle area screening strategy is used to determine the candidate subtitle area with the lowest repetition rate of the text content and the longest total display duration among the n candidate subtitle areas as the subtitle area. The total display duration is the sum of the display durations of all the text content belonging to the candidate subtitle area.

[0140] In an optional embodiment, the device further includes:

[0141] A calculation module 904, configured to calculate the repetition rate of the candidate subtitle area. The repetition rate is the ratio of the cumulative duration to the total video duration of the video. The cumulative duration is the sum of the display durations of the same text content.

[0142] The screening module 903 is further configured to determine the candidate subtitle area with the repetition rate of the text content lower than the repetition rate threshold as the preliminarily screened subtitle area.

[0143] The calculation module 904 is further configured to calculate the total display duration of the initially screened subtitle area;

[0144] The screening module 903 is further configured to determine the initially screened subtitle area with the longest total display duration in the initially screened subtitle areas as the subtitle area.

[0145] In an optional embodiment, the calculation module 904 is further configured to obtain the j-th group of text data corresponding to the j-th candidate subtitle area, where the text areas in the j-th group of text data belong to the j-th candidate subtitle area, and j is a positive integer less than or equal to n, and n is a positive integer;

[0146] The calculation module 904 is further configured to group the text data with the same text content in the j-th group of text data into the same text data set, obtaining at least one text data set;

[0147] The calculation module 904 is further configured to calculate the sum of the display durations in each text data set, obtaining at least one of the cumulative durations;

[0148] The calculation module 904 is further configured to calculate the ratio of the largest cumulative duration to the total video duration of the video to obtain the repetition rate;

[0149] The calculation module 904 is further configured to repeat the above four steps to calculate the repetition rate of each candidate subtitle area.

[0150] In an optional embodiment, the calculation module 904 is further configured to calculate the sum of the display durations of the text data corresponding to the initially screened subtitle area, obtaining the total display duration of the initially screened subtitle area.

[0151] In an optional embodiment, the text list includes m pieces of text data, and the text area includes the upper and lower lines of a rectangle, where m is a positive integer;

[0152] The candidate module 902 is further configured to extract one text area from the m text areas as the first text area, determine the first text area as the first candidate subtitle area, and add the first candidate subtitle area to the candidate subtitle area list;

[0153] The candidate module 902 is further configured to repeatedly execute the following steps until the remaining number of the m text areas is 0: extract one text area from the remaining m - k + 1 text areas as the k-th text area, and in response to the first position deviation between the k-th text area and the w-th candidate subtitle area in the candidate subtitle area list being less than the deviation threshold, classify the k-th text area into the w-th candidate subtitle area;

[0154] In response to the second position deviation between the k-th text region and all candidate subtitle regions in the candidate subtitle region list being greater than the deviation threshold, determining the k-th text region as the y-th candidate subtitle region, and adding the y-th candidate subtitle region to the candidate subtitle region list;

[0155] Wherein, the first position deviation includes the difference between the two upper lines and the difference between the two lower lines, the second position deviation includes the difference between the two upper lines or the difference between the two lower lines, y is a positive integer less than or equal to n, k is a positive integer less than or equal to m, w is a positive integer less than or equal to n, and n is a positive integer.

[0156] In an alternative embodiment, the candidate module 902 is further configured to calculate a first height of the k-th text region, where the first height is the difference between the upper line and the lower line of the k-th text region; calculate a second height of the w-th candidate subtitle region, where the second height is the difference between the upper line and the lower line of the w-th candidate subtitle region; and in response to the first height being greater than the second height, determining the k-th text region as the w-th candidate subtitle region;

[0157] Wherein, k is a positive integer less than or equal to m, w is a positive integer less than or equal to n, and n, m are positive integers.

[0158] In an alternative embodiment, the apparatus further includes:

[0159] An acquisition module 905, configured to periodically capture video frame images of the video;

[0160] The recognition module 901 is further configured to recognize text in the video frame image to obtain the text list.

[0161] In an alternative embodiment, the recognition module 901 is further configured to call an optical character recognition (OCR) model to recognize the video frame image, obtain candidate text content in the video frame image and the text region of the candidate text content, and obtain the display time of the candidate text content according to the display time of the video frame image;

[0162] The recognition module 901 is further configured to remove duplicates from the candidate text content to obtain the text content; the duplicate removal includes determining the candidate text content with the earliest display time among multiple candidate text contents with continuous display times, the same text region, and the same candidate text content as the text content, and calculating the display duration of the text content according to the display times of the multiple candidate text contents;

[0163] The recognition module 901 is further configured to generate the text list according to the text content, the text area of the text content, and the display duration.

[0164] In an optional embodiment, the apparatus further includes:

[0165] A subtitle module 906, configured to recognize the subtitles of the video according to the text content belonging to the subtitle area.

[0166] Figure 11 It is a schematic structural diagram of a server provided by an embodiment of the present application. Specifically: The server 1000 includes a central processing unit (English: Central Processing Unit, abbreviated: CPU) 1001, a system memory 1004 including a random access memory (English: Random Access Memory, abbreviated: RAM) 1002 and a read-only memory (English: Read-Only Memory, abbreviated: ROM) 1003, and a system bus 1005 connecting the system memory 1004 and the central processing unit 1001. The server 1000 further includes a basic input / output system (I / O system) 1006 for transmitting information between various devices in the computer, and a mass storage device 1007 for storing an operating system 1013, application programs 1014, and other program modules 1015.

[0167] The basic input / output system 1006 includes a display 1008 for displaying information and an input device 1009 such as a mouse, keyboard, etc. for user input of information. Among them, both the display 1008 and the input device 1009 are connected to the central processing unit 1001 through an input / output controller 1010 connected to the system bus 1005. The basic input / output system 1006 may further include an input / output controller 1010 for receiving and processing inputs from multiple other devices such as a keyboard, mouse, or electronic stylus. Similarly, the input / output controller 1010 also provides output to a display screen, printer, or other types of output devices.

[0168] The mass storage device 1007 is connected to the central processing unit 1001 through a mass storage controller (not shown) connected to the system bus 1005. The mass storage device 1007 and its associated computer-readable medium provide non-volatile storage for the server 1000. That is to say, the mass storage device 1007 may include a computer-readable medium (not shown) such as a hard disk or a compact disc read-only memory (English: Compact Disc Read-Only Memory, abbreviated: CD-ROM) drive.

[0169] Without loss of generality, computer-readable media can include computer storage media and communication media. Computer storage media includes volatile and non-volatile, removable and non-removable media implemented by any method or technology for storing information such as computer-readable instructions, data structures, program modules, or other data. Computer storage media includes RAM, ROM, erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other solid-state storage technologies, CD-ROM, digital versatile disc (DVD) or other optical storage, magnetic tape cartridges, magnetic tape, disk storage or other magnetic storage devices. Of course, those skilled in the art will understand that computer storage media is not limited to the above several types. The above system memory 1004 and mass storage device 1007 can be collectively referred to as memory.

[0170] According to various embodiments of the present application, the server 1000 can also be run by a remote computer on the network connected through a network such as the Internet. That is, the server 1000 can be connected to the network 1012 through the network interface unit 1011 connected to the system bus 1005, or rather, the network interface unit 1011 can also be used to connect to other types of networks or remote computer systems (not shown).

[0171] The present application also provides a terminal, which includes a processor and a memory. At least one instruction is stored in the memory, and the at least one instruction is loaded and executed by the processor to implement the subtitle area recognition method provided by each of the above method embodiments. It should be noted that the terminal can be as follows Figure 12 the provided terminal.

[0172] Figure 12 The structural block diagram of the terminal 1100 provided by an exemplary embodiment of the present application is shown. The terminal 1100 can be: a smart phone, a tablet computer, an MP3 player (Moving Picture Experts Group Audio LayerIII), an MP4 (Moving Picture Experts Group AudioLayer IV) player, a notebook computer or a desktop computer. The terminal 1100 may also be referred to by other names such as user equipment, portable terminal, laptop terminal, desktop terminal, etc.

[0173] Generally, the terminal 1100 includes a processor 1101 and a memory 1102.

[0174] The processor 1101 may include one or more processing cores, such as a quad-core processor, an octa-core processor, etc. The processor 1101 may be implemented in at least one hardware form of DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), or PLA (Programmable Logic Array). The processor 1101 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the wake state, also known as the CPU (Central Processing Unit); the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor 1101 may be integrated with a GPU (Graphics Processing Unit), and the GPU is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 1101 may further include an AI (Artificial Intelligence) processor, and the AI processor is used to process computational operations related to machine learning.

[0175] The memory 1102 may include one or more computer-readable storage media, and the computer-readable storage media may be non-transitory. The memory 1102 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices and flash storage devices. In some embodiments, the non-transitory computer-readable storage medium in the memory 1102 is used to store at least one instruction, and the at least one instruction is used to be executed by the processor 1101 to implement the subtitle area recognition method provided in the method embodiments of the present application.

[0176] In some embodiments, the terminal 1100 may further optionally include a peripheral device interface 1103 and at least one peripheral device. The processor 1101, the memory 1102, and the peripheral device interface 1103 may be connected through a bus or signal lines. Each peripheral device may be connected to the peripheral device interface 1103 through a bus, signal lines, or a circuit board. Specifically, the peripheral device includes at least one of a radio frequency circuit 1104, a display screen 1105, a camera module 1106, an audio circuit 1107, and a power supply 1109.

[0177] The peripheral device interface 1103 can be used to connect at least one I / O (Input / Output) related peripheral device to the processor 1101 and the memory 1102. In some embodiments, the processor 1101, the memory 1102, and the peripheral device interface 1103 are integrated on the same chip or circuit board; in some other embodiments, any one or two of the processor 1101, the memory 1102, and the peripheral device interface 1103 can be implemented on a separate chip or circuit board, and this embodiment does not limit this.

[0178] The radio frequency circuit 1104 is used to receive and transmit RF (Radio Frequency) signals, also known as electromagnetic signals. The radio frequency circuit 1104 communicates with a communication network and other communication devices through electromagnetic signals. The radio frequency circuit 1104 converts an electrical signal into an electromagnetic signal for transmission, or converts a received electromagnetic signal into an electrical signal. Exemplarily, the radio frequency circuit 1104 includes: an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, a codec chipset, a subscriber identity module card, and so on. The radio frequency circuit 1104 can communicate with other terminals through at least one wireless communication protocol. The wireless communication protocol includes but is not limited to: the World Wide Web, a metropolitan area network, an intranet, each generation of mobile communication networks (2G, 3G, 4G, and 5G), a wireless local area network, and / or a WiFi (Wireless Fidelity) network. In some embodiments, the radio frequency circuit 1104 may further include a circuit related to NFC (Near Field Communication), and this application does not limit this.

[0179] The display screen 1105 is used to display the UI (User Interface). The UI may include graphics, text, icons, videos, and any combination thereof. When the display screen 1105 is a touch display screen, the display screen 1105 also has the ability to collect touch signals on or above the surface of the display screen 1105. The touch signals can be input to the processor 1101 as control signals for processing. At this time, the display screen 1105 can also be used to provide virtual buttons and / or a virtual keyboard, also known as soft buttons and / or a soft keyboard. In some embodiments, there may be one display screen 1105, which is provided on the front panel of the terminal 1100; in other embodiments, there may be at least two display screens 1105, which are respectively provided on different surfaces of the terminal 1100 or are in a folding design; in still other embodiments, the display screen 1105 may be a flexible display screen, which is provided on a curved surface or a folding surface of the terminal 1100. Even more, the display screen 1105 can also be set to an irregular non-rectangular shape, that is, a special-shaped screen. The display screen 1105 can be prepared from materials such as LCD (Liquid Crystal Display) and OLED (Organic Light-Emitting Diode).

[0180] The camera module 1106 is used to collect images or videos. Exemplarily, the camera module 1106 includes a front camera and a rear camera. Generally, the front camera is provided on the front panel of the terminal, and the rear camera is provided on the back of the terminal. In some embodiments, there are at least two rear cameras, which are any one of a main camera, a depth-of-field camera, a wide-angle camera, and a telephoto camera respectively, so as to realize the function of background blurring by fusing the main camera and the depth-of-field camera, the function of panoramic shooting by fusing the main camera and the wide-angle camera, and the VR (Virtual Reality) shooting function or other fusion shooting functions. In some embodiments, the camera module 1106 may further include a flash. The flash can be a single-color-temperature flash or a dual-color-temperature flash. A dual-color-temperature flash refers to a combination of a warm-light flash and a cold-light flash, which can be used for light compensation under different color temperatures.

[0181] The audio circuit 1107 may include a microphone and a speaker. The microphone is used to collect sound waves of the user and the environment, and convert the sound waves into electrical signals for input to the processor 1101 for processing, or input to the radio frequency circuit 1104 to enable voice communication. For the purpose of stereo collection or noise reduction, there may be multiple microphones, which are respectively arranged at different parts of the terminal 1100. The microphone may also be an array microphone or an omnidirectional collection microphone. The speaker is used to convert the electrical signal from the processor 1101 or the radio frequency circuit 1104 into sound waves. The speaker may be a traditional thin film speaker or a piezoelectric ceramic speaker. When the speaker is a piezoelectric ceramic speaker, it can not only convert the electrical signal into sound waves audible to humans, but also convert the electrical signal into sound waves inaudible to humans for uses such as ranging. In some embodiments, the audio circuit 1107 may further include a headphone jack.

[0182] The power supply 1109 is used to supply power to each component in the terminal 1100. The power supply 1109 may be alternating current, direct current, a disposable battery or a rechargeable battery. When the power supply 1109 includes a rechargeable battery, the rechargeable battery may be a wired rechargeable battery or a wireless rechargeable battery. A wired rechargeable battery is a battery charged through a wired line, and a wireless rechargeable battery is a battery charged through a wireless coil. The rechargeable battery may also be used to support fast charging technology.

[0183] In some embodiments, the terminal 1100 further includes one or more sensors 1110. The one or more sensors 1110 include but are not limited to: an acceleration sensor 1111, a gyroscope sensor 1112, a pressure sensor 1113, an optical sensor 1115, and a proximity sensor 1116.

[0184] The acceleration sensor 1111 can detect the magnitudes of accelerations on the three coordinate axes of the coordinate system established with the terminal 1100. For example, the acceleration sensor 1111 can be used to detect the components of the gravitational acceleration on the three coordinate axes. The processor 1101 can control the display screen 1105 to display the user interface in a landscape view or a portrait view according to the gravitational acceleration signal collected by the acceleration sensor 1111. The acceleration sensor 1111 can also be used for the collection of game or user's motion data.

[0185] The gyroscope sensor 1112 can detect the body direction and rotation angle of the terminal 1100. The gyroscope sensor 1112 can cooperate with the acceleration sensor 1111 to collect the 3D actions of the user on the terminal 1100. Based on the data collected by the gyroscope sensor 1112, the processor 1101 can implement the following functions: motion sensing (such as changing the UI according to the user's tilting operation), image stabilization during shooting, game control, and inertial navigation.

[0186] The pressure sensor 1113 can be disposed on the side frame of the terminal 1100 and / or the lower layer of the display screen 1105. When the pressure sensor 1113 is disposed on the side frame of the terminal 1100, it can detect the holding signal of the user on the terminal 1100, and the processor 1101 can perform left / right hand recognition or shortcut operations according to the holding signal collected by the pressure sensor 1113. When the pressure sensor 1113 is disposed on the lower layer of the display screen 1105, the processor 1101 can control the operable controls on the UI interface according to the pressure operation of the user on the display screen 1105. The operable controls include at least one of button controls, scroll bar controls, icon controls, and menu controls.

[0187] The optical sensor 1115 is used to collect the ambient light intensity. In one embodiment, the processor 1101 can control the display brightness of the display screen 1105 according to the ambient light intensity collected by the optical sensor 1115. Specifically, when the ambient light intensity is high, the display brightness of the display screen 1105 is increased; when the ambient light intensity is low, the display brightness of the display screen 1105 is decreased. In another embodiment, the processor 1101 can also dynamically adjust the shooting parameters of the camera assembly 1106 according to the ambient light intensity collected by the optical sensor 1115.

[0188] The proximity sensor 1116, also known as the distance sensor, is usually disposed on the front panel of the terminal 1100. The proximity sensor 1116 is used to collect the distance between the user and the front of the terminal 1100. In one embodiment, when the proximity sensor 1116 detects that the distance between the user and the front of the terminal 1100 is gradually decreasing, the processor 1101 controls the display screen 1105 to switch from the lit state to the off state; when the proximity sensor 1116 detects that the distance between the user and the front of the terminal 1100 is gradually increasing, the processor 1101 controls the display screen 1105 to switch from the off state to the lit state.

[0189] Those skilled in the art can understand that Figure 12 the structure shown in does not constitute a limitation on the terminal 1100, and it may include more or fewer components than shown in the figure, or combine some components, or adopt different component arrangements.

[0190] The memory further includes one or more programs, the one or more programs are stored in the memory, and the one or more programs include a method for performing subtitle area recognition provided in the embodiments of the present application.

[0191] The present application also provides a computer device, which includes a processor and a memory. At least one instruction, at least one program, a code set or an instruction set is stored in the storage medium, and the at least one instruction, at least one program, the code set or the instruction set is loaded and executed by the processor to implement the subtitle area recognition method provided by each of the above method embodiments.

[0192] The present application also provides a computer-readable storage medium, in which at least one instruction, at least one program, a code set or an instruction set is stored, and the at least one instruction, at least one program, the code set or the instruction set is loaded and executed by the processor to implement the subtitle area recognition method provided by each of the above method embodiments.

[0193] The present application also provides a computer program product or a computer program, which includes computer instructions stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the subtitle area recognition method provided in the above optional implementation manners.

[0194] It should be understood that "a plurality of" mentioned herein refers to two or more. "And / or" describes the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally represents an "or" relationship between the associated objects before and after.

[0195] Those of ordinary skill in the art can understand that all or part of the steps to implement the above embodiments can be completed by hardware, or can be completed by a program instructing relevant hardware. The program can be stored in a computer-readable storage medium, and the above-mentioned storage medium can be a read-only memory, a disk or an optical disc, etc.

[0196] The above are only optional embodiments of the present application, and are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.

Claims

1. A method for identifying subtitle regions, characterized in that, the method includes: identifying the text in the video to obtain a text list, the text list including at least one piece of text data, the text data including text content, a text region, and a display duration, the text content including at least one character located on the text region, and the text region including the position of a text box for framing the characters; regulating the text regions into n candidate subtitle regions, the deviation between the text regions belonging to the i-th candidate subtitle region and the position of the i-th candidate subtitle region being less than a deviation threshold, n being a positive integer, and i being a positive integer less than or equal to n; screening the subtitle region from the n candidate subtitle regions according to a subtitle region screening strategy; the subtitle region screening strategy is used to determine the candidate subtitle region with the lowest repetition rate of the text content and the longest total display duration among the n candidate subtitle regions as the subtitle region, the repetition rate being the proportion of the cumulative duration of the same text content displayed on the candidate subtitle region to the total video duration of the video, and the total display duration being the sum of the display durations of all the text content belonging to the candidate subtitle region.

2. The method according to claim 1, characterized in that, screening the subtitle region from the n candidate subtitle regions according to the subtitle region screening strategy includes: calculating the repetition rate of the candidate subtitle region, the repetition rate being the ratio of the cumulative duration to the total video duration of the video, and the cumulative duration being the sum of the display durations of the same text content; determining the candidate subtitle regions with the repetition rate of the text content lower than the repetition rate threshold as the preliminary screening subtitle regions; calculating the total display duration of the preliminary screening subtitle regions; determining the preliminary screening subtitle region with the longest total display duration among the preliminary screening subtitle regions as the subtitle region.

3. The method according to claim 2, characterized in that, calculating the repetition rate of the candidate subtitle region includes: obtaining the j-th group of text data corresponding to the j-th candidate subtitle region, the text region in the j-th group of text data belonging to the j-th candidate subtitle region, j being a positive integer less than or equal to n, and n being a positive integer; grouping the text data with the same text content in the j-th group of text data into the same text data set, obtaining at least one text data set; calculating the sum of the display durations in each text data set, obtaining at least one cumulative duration; calculating the ratio of the maximum cumulative duration to the total video duration of the video to obtain the repetition rate; repeating the above four steps to calculate the repetition rate of each candidate subtitle region.

4. The method according to claim 2, characterized in that, calculating the total display duration of the preliminary screening subtitle region includes: calculating the sum of the display durations of the text data corresponding to the preliminary screening subtitle region to obtain the total display duration of the preliminary screening subtitle region.

5. The method according to any one of claims 1 to 4, characterized in that, The text list includes m text data, and the text area includes the upper and lower lines of a rectangle, where m is a positive integer; The step of regularizing the text area into n candidate subtitle areas includes: Selecting one text area from the m text areas as the first text area, determining the first text area as the first candidate subtitle area, and adding the first candidate subtitle area to the candidate subtitle area list; Repeatedly executing the following steps until the remaining number of the m text areas is 0: Selecting one text area from the remaining m - k + 1 text areas as the kth text area. In response to the first position deviation between the kth text area and the wth candidate subtitle area in the candidate subtitle area list being less than the deviation threshold, classifying the kth text area into the wth candidate subtitle area; In response to the second position deviation between the kth text area and all candidate subtitle areas in the candidate subtitle area list being greater than the deviation threshold, determining the kth text area as the yth candidate subtitle area, and adding the yth candidate subtitle area to the candidate subtitle area list; Wherein, the first position deviation includes the difference between the two upper lines and the difference between the two lower lines, the second position deviation includes the difference between the two upper lines or the difference between the two lower lines, y is a positive integer less than or equal to n, k is a positive integer less than or equal to m, w is a positive integer less than or equal to n, and n is a positive integer.

6. The method according to claim 5, wherein, after classifying the kth text area into the wth candidate subtitle area in response to the first position deviation between the kth text area and the wth candidate subtitle area in the candidate subtitle area list being less than the deviation threshold, further includes: Calculating the first height of the kth text area, where the first height is the difference between the upper line and the lower line of the kth text area; Calculating the second height of the wth candidate subtitle area, where the second height is the difference between the upper line and the lower line of the wth candidate subtitle area; In response to the first height being greater than the second height, determining the kth text area as the wth candidate subtitle area; Wherein, k is a positive integer less than or equal to m, w is a positive integer less than or equal to n, and n and m are positive integers.

7. The method according to any one of claims 1 to 4, wherein, The step of recognizing the text in the video to obtain the text list includes: Periodically intercepting the video frame images of the video; Recognizing the text in the video frame images to obtain the text list.

8. The method according to claim 7, wherein, The step of recognizing the text in the video frame images to obtain the text list includes: Invoking an optical character recognition (OCR) model to recognize the video frame images, obtaining the candidate text content in the video frame images and the text area of the candidate text content, and obtaining the display time of the candidate text content according to the display time of the video frame images; Deduplicate the candidate text content to obtain the text content; the deduplication includes determining, among multiple candidate text contents with consecutive display times, the same text area, and the same candidate text content, the candidate text content with the earliest display time as the text content, and calculating the display duration of the text content based on the display times of the multiple candidate text contents. Generate the text list according to the text content, the text area of the text content, and the display duration.

9. According to the method of any one of claims 1 to 4, characterized in that the method further includes: Identifying the subtitles of the video according to the text content belonging to the subtitle area.

10. A subtitle area recognition device, characterized in that the device includes: An identification module for identifying text in a video to obtain a text list, the text list including at least one piece of text data, the text data including text content, a text area, and a display duration, the text content including at least one text located on the text area, and the text area including the position of a text box for framing the text; A candidate module for regularizing the text area into n candidate subtitle areas, the text area belonging to the i-th candidate subtitle area having a position deviation from the position of the i-th candidate subtitle area less than a deviation threshold, n being a positive integer, and i being a positive integer less than or equal to n; A screening module for screening the subtitle area from the n candidate subtitle areas according to a subtitle area screening strategy; the subtitle area screening strategy is used to determine, as the subtitle area, the candidate subtitle area with the lowest repetition rate of the text content and the longest total display duration among the n candidate subtitle areas, the repetition rate being the proportion of the cumulative display duration of the same text content displayed on the candidate subtitle area to the total video duration of the video, and the total display duration being the sum of the display durations of all the text content belonging to the candidate subtitle area.

11. A computer device, the computer device includes: A processor and a memory, where at least one instruction, at least one program, a code set, or an instruction set is stored in the memory, and the at least one instruction, the at least one program, the code set, or the instruction set is loaded and executed by the processor to implement the subtitle area recognition method according to any one of claims 1 to 9.

12. A computer-readable storage medium, characterized in that at least one instruction, at least one program, a code set, or an instruction set is stored in the storage medium, and the at least one instruction, the at least one program, the code set, or the instruction set is loaded and executed by a processor to implement the subtitle area recognition method according to any one of claims 1 to 9.

13. A computer program product, characterized in that The computer program product includes computer instructions which are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions to implement the subtitle area recognition method according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • A method for caption area of positioning video

    CN101102419A

  • Video caption screening method, device and equipment and storage medium

    CN111723790A