Video subtitle processing method and system and storage medium

By combining and cropping the subtitle text boxes in the video frame image, and correcting the subtitle text with speech recognition results, the problems of incomplete subtitle recognition of short drama videos and inaccurate timelines are solved, and efficient and accurate subtitle generation is achieved.

CN120299016APending Publication Date: 2025-07-11SHENZHEN CALF ANIMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510345322.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-24
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

The prior art has problems in the recognition of short drama video subtitles, such as incomplete recognition, missed single words, and inability to accurately give the subtitle timeline, resulting in inaccurate subtitle text and requiring a lot of manual proofreading.

Method used

By detecting the text boxes in the video frame image, combining the subtitle text boxes, cropping the video frame image to obtain the image to be recognized, and correcting the subtitle text with the speech recognition results to generate an accurate subtitle file.

Benefits of technology

It improves the accuracy and efficiency of subtitle recognition, reduces the need for manual proofreading, and ensures the accuracy of subtitle text and the accuracy of the timeline.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120299016A_ABST
    Figure CN120299016A_ABST
Patent Text Reader

Abstract

The invention relates to a video subtitle processing method, and the method comprises the steps: merging subtitle text boxes which can comprise all subtitle texts in a video frame image, thereby avoiding the situation of character missing and character missing when a to-be-recognized image is subjected to character recognition. By calibrating the position of the combined subtitle character box, when the to-be-recognized image is cut, it is guaranteed that the to-be-recognized image only comprises the subtitle text, it is avoided that non-subtitle character content in the video frame image is recognized when the subtitle text is generated according to character recognition of the to-be-recognized image, and the subtitle recognition accuracy is improved. Only the characters in the to-be-recognized image are recognized, so that the subtitle recognition efficiency is improved. And subtitle characters are recognized through voice, so that the subtitle recognition result can be further corrected.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of video analysis technology, and in particular, to a method, system, and storage medium for video subtitle processing. Background Art

[0002] With the rapid development of the short drama industry, a large number of videos need to have their subtitles extracted, recognized, translated, added, and synthesized. Since during the production of a large number of short drama videos, their subtitles have been suppressed into the video image content, and the original subtitles often have problems of poor storage. When obtaining copyrighted content from a third party, often only the video file is obtained, but the text subtitle file is missing.

[0003] In the prior art, methods for recognizing and extracting subtitles from the image content of a video or extracting text from audio have been widely applied to various video subtitle generation requirements. However, for short drama videos that require accurate recognition, translation, and addition of video subtitles, these methods have problems such as incomplete subtitle recognition, missing subtitles in some frames, or excessive recognized subtitle text, with misrecognition and missing single characters, etc., and often require a large amount of manual post-proofreading. At the same time, there is also the problem of being unable to accurately give the subtitle time axis, and there are also problems such as some characters being recognized as homophones. Therefore, there are many deficiencies in the prior art and it is unable to effectively generate accurate subtitle text.

[0004] This application first merges text boxes, where the merged subtitle text boxes can include all the text in the video frame image, avoiding the situation of missing characters when performing text recognition on the subtitle image. Secondly, by screening the merged text boxes, the merged subtitle text boxes that only contain subtitles are determined, and then the positions of the merged subtitle text boxes are calibrated. When intercepting the image to be recognized, it is ensured that the image to be recognized only includes subtitle text, avoiding recognizing non-subtitle text content in the video frame image, which is beneficial to improving the accuracy of subtitle recognition. By only recognizing the text in the image to be recognized and not the entire video frame image, it is beneficial to accelerating the efficiency of subtitle recognition. Summary of the Invention

[0005] This application provides a method, system, and storage medium for video subtitle processing to solve the problem that the prior art cannot effectively generate accurate subtitle text.

[0006] In a first aspect, this application provides a method, including:

[0007] Detect and recognize the text in the video frame image to obtain text boxes, where the text boxes include subtitle text boxes;

[0008] Merge according to the positions of the text boxes, and statistically calculate the calibration positions of the merged subtitle text boxes of the video;

[0009] Based on the merged subtitle text box, calculate and obtain the calibration position of the merged subtitle text box of the video;

[0010] Crop the video frame image according to the position of the merged subtitle text box to obtain an image to be recognized;

[0011] Recognize the text in the image to be recognized and correct it through the speech recognition result to obtain a subtitle file;

[0012] Translate the subtitle file and synthesize it into the video.

[0013] Optionally, the merging according to the position of the text box to obtain a merged text box includes:

[0014] Based on the overlapping area of adjacent text boxes, merge and de-duplicate the adjacent text boxes that meet the preset ratio in the horizontal and vertical directions to obtain the merged text box.

[0015] Optionally, the text box further includes a non-subtitle text box, the merged text box further includes a merged non-subtitle text box, and the calculating and obtaining the calibration position of the merged subtitle text box of the video according to the merged subtitle text box includes:

[0016] Detect and recognize all the merged text boxes;

[0017] Determine the merged text boxes at the four corners of the video frame image and the merged text boxes that are not left-right symmetric with respect to the video frame image as the merged non-subtitle text boxes;

[0018] Filter the merged non-subtitle text boxes to determine the merged subtitle text boxes;

[0019] Statistically calculate the position of the merged subtitle text box frame by frame to obtain the calibration position of the merged subtitle text box of the video.

[0020] Optionally, the cropping the video frame image according to the calibration position of the merged subtitle text box to obtain an image to be recognized includes:

[0021] Detect the first distance between the upper border of the merged subtitle text box and the upper boundary of the video frame image;

[0022] Detect the second distance between the lower border of the merged subtitle text box and the lower boundary of the video frame image;

[0023] According to the first distance and the second distance, retain the area of the calibration position of the merged subtitle text box and crop the video frame image to obtain an image to be recognized.

[0024] Optionally, the recognition of the text in the subtitle image and the correction through the speech recognition result to obtain a subtitle file includes:

[0025] Obtain the timestamp corresponding to the video frame;

[0026] Perform text recognition on the text in the image to be recognized and correct it through the speech recognition result to obtain subtitle text;

[0027] Generate a subtitle file according to the timestamp corresponding to the video frame and the subtitle text.

[0028] Optionally, the performing text recognition on the text in the image to be recognized and correcting it through the speech recognition result to obtain subtitle text includes:

[0029] Perform text recognition on the text in the image to be recognized to obtain a text recognition result;

[0030] Perform speech text recognition on the audio of the video to obtain a speech recognition result;

[0031] Perform text correction on the text recognition result according to the speech recognition result to obtain subtitle text.

[0032] Optionally, the performing text recognition on the text in the image to be recognized to obtain a text recognition result includes:

[0033] Stitch the images to be recognized in order of frames from top to bottom in groups, perform text recognition on the stitched multiple images to be recognized in the horizontal direction, and separate the stitched images to be recognized by black bars or white bars;

[0034] Separate and restore the recognition results of the stitched multiple images to be recognized in order of frames;

[0035] Merge the frames with the same subtitles in order of frames for the separated recognition results to obtain the text recognition result.

[0036] In a second aspect, the present application provides a video subtitle processing system, including:

[0037] A text box acquisition module, configured to detect and recognize text in a video frame image to obtain a text box, where the text box includes a subtitle text box;

[0038] A merged text box module, configured to merge according to the positions of the text boxes to obtain a merged text box, where the merged text box includes a merged subtitle text box;

[0039] A subtitle position calibration module, configured to statistically calculate the calibration position of the merged subtitle text box of the video according to the merged subtitle text box;

[0040] An image to be recognized acquisition module, configured to crop the video frame image according to the calibrated position of the merged subtitle text box to obtain an image to be recognized;

[0041] A subtitle file acquisition module, configured to perform text recognition on the image to be recognized and correct it through the speech recognition result to obtain a subtitle file;

[0042] A subtitle file import module, configured to translate the subtitle file and synthesize it into the video.

[0043] In a third aspect, an embodiment of the present application provides a terminal device, including: a memory, a processor, and a computer program stored in the memory and executable on the processor, where when the processor executes the computer program, the above-mentioned method is implemented.

[0044] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, where the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the above-mentioned method is implemented.

[0045] The above technical solution provided by the embodiment of the present application has the following advantages compared with the prior art: First, by merging text boxes, where the merged subtitle text box can include all text in the video frame image, avoiding the situation of missing characters when performing text recognition on the subtitle image. Secondly, by screening the merged text boxes, determining the merged subtitle text boxes that only contain subtitles, and then calibrating the positions of the merged subtitle text boxes, ensuring that the image to be recognized only includes subtitle text when intercepting the image to be recognized, avoiding recognizing non-subtitle text content in the video frame image, which is beneficial to improving the accuracy of subtitle recognition. By only recognizing the text in the image to be recognized without recognizing the entire video frame image, it is beneficial to speed up the efficiency of subtitle recognition. At the same time, speech recognition is performed on the audio content of the video to further correct the content recognized from the video image. Description of the Drawings

[0046] The drawings here are incorporated into the description and form a part of this description, showing embodiments consistent with the present invention and used together with the description to explain the principles of the present invention.

[0047] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0048] One or more embodiments are exemplarily illustrated by the pictures in the corresponding drawings. These exemplary illustrations do not constitute limitations on the embodiments. Elements with the same reference numerals in the drawings represent similar elements. Unless otherwise stated, the drawings in the figures do not constitute a scale limitation.

[0049] Figure 1 It is a flowchart of a video subtitle processing method provided by an embodiment of the present application;

[0050] Figure 2 It is a schematic diagram of a text box provided by an embodiment of the present application;

[0051] Figure 3 It is a schematic diagram of combining subtitle text boxes and non - subtitle text boxes provided by an embodiment of the present application;

[0052] Figure 4 Provided by an embodiment of the present application Figure 1 The specific process schematic diagram of S300 in

[0053] Figure 5 Provided by an embodiment of the present application Figure 1 The specific process schematic diagram of S400 in

[0054] Figure 6 It is a schematic diagram of an image to be recognized provided by an embodiment of the present application;

[0055] Figure 7 Provided by an embodiment of the present application Figure 1 The specific process schematic diagram of S500 in

[0056] Figure 8 Provided by an embodiment of the present application Figure 7 The specific process schematic diagram of S502 in

[0057] Figure 9 It is a schematic diagram of splicing an image to be recognized provided by an embodiment of the present application;

[0058] Figure 10 It is a schematic diagram of the structure of a video subtitle processing system provided by an embodiment of the present application.

[0059] Figure 11 It is a schematic diagram of the structure of a terminal device provided by an embodiment of the present application. Detailed implementation manners

[0060] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the following will clearly and completely describe the technical solutions in the embodiments of this application with reference to the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are some but not all of the embodiments of this application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of this application without creative efforts shall fall within the scope of protection of this application.

[0061] The following disclosure provides many different embodiments or examples for implementing different structures of the present invention. To simplify the disclosure of the present invention, components and settings of specific examples are described below. Of course, they are only examples and are not intended to limit the present invention. In addition, the present invention may repeat reference numerals and / or letters in different examples. This repetition is for the purpose of simplification and clarity and does not itself indicate the relationship between the various embodiments and / or settings discussed.

[0062] With the rapid development of the short drama industry, a number of short drama video platforms with overseas expansion needs have emerged in China, and there are a large number of videos that require subtitle extraction, recognition, translation, and addition and synthesis.

[0063] Since in the production process of a large number of short drama videos, the subtitles have been suppressed into the video image content, and the original subtitles often have problems of poor storage. When obtaining copyright content from a third party, often only the video file is obtained, but the text subtitle file is missing. Therefore, there is a need to accurately extract video subtitles and perform translation, addition, and synthesis.

[0064] Currently, the publicly available methods for extracting subtitles by recognizing the image content of a video or extracting text by audio recognition have been widely applied to various video subtitle generation requirements. Such as existing video subtitle extraction tools like VideoSubFinder and video - subtitle - extractor, image text recognition tools such as paddleocr of Baidu, Tencent Cloud ocr, and Alibaba Cloud ocr; and speech recognition models on modelscope for generating text by speech recognition.

[0065] However, for a type of short drama videos that require accurate subtitle recognition, translation, and addition, these methods all have certain defects. For software such as VideoSubFinder and video - subtitle - extractor, there are problems such as incomplete subtitle recognition and missing subtitles for some frames in VideoSubFinder, or excessive recognized subtitle text, misrecognition, and missing single characters in video - subtitle - extractor, and a large amount of manual post - proofreading is often required.

[0066] For methods and systems related to speech recognition, there are problems such as the inability to accurately provide the subtitle timeline in the video, and some words being recognized as homophones. As a result, the generated subtitles cannot be used for translation and added back to the original video.

[0067] The present invention proposes a video subtitle processing method, as Figure 1 shown, Figure 1 which is a video subtitle processing method provided by an embodiment of the present application. The method includes:

[0068] S100, detecting and recognizing the text in the video frame image to obtain a text box, where the text box includes a subtitle text box.

[0069] The recognition of video subtitles first requires detecting and recognizing the text in the video frame image. For short drama videos, each one consists of many episodes, and their subtitles generally have a set of parameter settings and are usually located in a fixed area. Therefore, the subtitle position can be detected by detecting a small number of episodes. For a small number of inconsistent episodes, detection processing can be performed on a single episode. Extract a small number of episodes, and first extract one video frame image every n frames. Assuming the video frame rate is 30 FPS, then n can be set to 5, and n can be preset according to the actual situation and will not be specifically limited here. Use TextDetector of yolo or paddleocr to detect the text position in the video frame image. Since the position where the subtitles are synthesized into the short drama video is basically at the same horizontal position, usually at most two lines, it is necessary to expand the detected position in the up and down directions. Among them, the video frame image may not contain text, or may include subtitles and non-subtitle text. The non-subtitle text, for example, appears as the text on a shop billboard in the video image. Generally, the subtitle text box has left-right symmetry with respect to the center line of the video frame image and is located in the lower area of the video frame image. The non-subtitle text box can appear at any position in the video frame image. When the subtitle is recognized, the subtitle text box is obtained, and when the non-subtitle is recognized, the non-subtitle text box is obtained.

[0070] Figure 2 which is a schematic diagram of the text box provided by an embodiment of the present application, as Figure 2 shown. The text box A has a left-right symmetric structure with respect to the center line of the video frame image and is located in the center of the lower part of the video frame. Therefore, the text box A is a subtitle text box. The text box C is in the upper right corner of the video frame image. Therefore, the text box C is a non-subtitle text box.

[0071] S200, merging according to the positions of the text boxes to obtain a merged text box, where the merged text box includes a merged subtitle text box.

[0072] In the embodiments of the present application, the recognized subtitle text boxes and non-subtitle text boxes in S100 are merged respectively to obtain merged subtitle text boxes and merged non-subtitle text boxes. Specifically, based on the overlapping area of adjacent text boxes, the adjacent text boxes that meet the preset ratio are merged and de-duplicated in the horizontal and vertical directions to obtain the merged text boxes.

[0073] In a possible implementation manner, the detected text boxes are de-duplicated and merged according to their positions. If the overlapping ratio of the text boxes meets the preset ratio, such as 80%, multiple text boxes are merged into a larger text box. Among them, merging in the horizontal direction means that the text boxes overlap or are adjacent in the horizontal direction (for example, the distance between the upper boundaries of the text boxes does not exceed 20% of the height of the largest text box), and expanding in the vertical direction means that adjacent text boxes are included as much as possible, and they are merged into a horizontal block. Deleting text boxes by merging text boxes is beneficial to avoiding missing characters when performing text recognition on subtitle images, and is also beneficial to facilitating the subsequent recognition of the text content in the text boxes.

[0074] Figure 3 This is a schematic diagram of the merged subtitle text box and the merged non-subtitle text box provided by the embodiments of the present application. As Figure 3 shown, if the overlapping area ratio of the subtitle text box A and the subtitle text box B in the video frame image detected and recognized meets the preset ratio, the subtitle text box A and the subtitle text box B are merged and combined into a merged subtitle text box in the horizontal and vertical directions. The merged subtitle text box includes the areas of the subtitle text box A and the subtitle text box B. If the overlapping area ratio of the non-subtitle text box C and the non-subtitle text box D in the video frame image detected and recognized meets the preset ratio, the non-subtitle text box C and the non-subtitle text box D are merged and combined into a merged non-subtitle text box in the horizontal and vertical directions. The merged non-subtitle text box includes the areas of the non-subtitle text box C and the non-subtitle text box D.

[0075] S300. According to the merged subtitle text box, the calibration position of the merged subtitle text box of the video is statistically calculated. Figure 4 This is provided by the embodiments of the present application Figure 1 The specific process schematic diagram of S300 in Figure 4 shown. The text box also includes a non-subtitle text box, and the merged text box also includes a merged non-subtitle text box. Obtaining the calibration position of the merged subtitle text box according to the merged subtitle text box includes:

[0076] S301. Detect all the merged text boxes.

[0077] In an embodiment of the present application, in a possible implementation manner, YOLO or PaddleOCR interfaces can be used to detect each video frame image to obtain the dt_box of each frame image slice, where the dt_box is the coordinate of the detected text box, such as [(422, 1559), (655, 1559), (655, 1660), (422, 1660)]. Here, it is not required that these four points must form a rectangle, but it is almost a rectangle. At the same time, the frame number frame_no needs to be recorded. frame_no is the frame number of the video. The video is composed of a series of consecutive images, and each frame image corresponds to a unique increasing consecutive frame number in sequence. Since there are various possibilities for the dt_box detected for a line of subtitles, multiple overlapping boxes may be detected, or adjacent text boxes may be detected right next to each other. Therefore, operations such as deduplication and merging are required.

[0078] S302, Determine the combined non-subtitle text boxes in the combined text boxes at the four corners of the video frame image and the combined text boxes that do not have left-right symmetry for the video frame image as the combined non-subtitle text boxes.

[0079] In an embodiment of the present application, the combined non-subtitle text boxes in the combined text boxes are screened out by detecting the coordinate positions of the combined text boxes. Since the non-subtitle text boxes may appear at any position in the video frame image and do not have left-right symmetry with respect to the lines in the video frame image, while the subtitle text boxes cannot appear at the four corners of the video frame image, the positions of the combined non-subtitle text boxes in the combined text boxes can be determined according to the above conditions.

[0080] S303, Filter the combined non-subtitle text boxes to determine the combined subtitle text boxes.

[0081] S304, Statistically calculate the positions of the combined subtitle text boxes frame by frame to obtain the calibration positions of the combined subtitle text boxes of the video.

[0082] In the embodiments of the present application, after screening out non-subtitle text boxes in S302, only the merged subtitle text boxes remain. The calibration position is determined based on the remaining merged subtitle text boxes. Since the positions where subtitles are synthesized into short video clips are basically at the same horizontal position, usually at most two lines, it is necessary to expand the detected positions upward and downward by a certain distance to prevent missed recognition. In one possible implementation, it can be observed manually, and the subtitle position can be calibrated by selecting the video subtitles on the computer screen and performing upward and downward position expansion; in another possible automatic implementation, the detected merged subtitle text boxes can be counted frame by frame to obtain the average positions of the upper and lower boundaries, and then the average positions of the upper and lower boundaries are expanded upward and downward by 10% to 60% as the calibration position of the merged subtitle text boxes. Uniform calibration of the video subtitle positions can be obtained by both of these two implementation methods, and usually there are no boundary restrictions on the left and right.

[0083] S400, crop the video frame image according to the position of the merged subtitle text box to obtain an image to be recognized. Figure 5 Provided by the embodiments of the present application Figure 1 The specific process schematic diagram of S400 in Figure 5 As shown, the step of cropping the video frame image according to the calibration position of the merged subtitle text box to obtain an image to be recognized includes:

[0084] S401, detect the first distance between the upper border of the merged subtitle text box and the upper boundary of the video frame image;

[0085] S402, detect the second distance between the lower border of the merged subtitle text box and the lower boundary of the video frame image;

[0086] S403, according to the first distance and the second distance, retain the area of the calibration position of the merged subtitle text box, and crop the video frame image to obtain an image to be recognized.

[0087] Figure 6 The schematic diagram of the image to be recognized provided by the embodiments of the present application, as Figure 6 shown, in the embodiments of the present application, the distance between the upper border of the detected merged subtitle text box and the upper boundary of the video frame image is used as the first distance, and the distance between the lower border of the merged subtitle text box and the lower boundary of the video frame image is used as the second distance. The top and bottom regions of the entire image are calculated and cropped according to the first and second distances, and the remaining rectangular image is the image to be recognized. The image to be recognized includes subtitle text. When recognizing subtitles, only the image to be recognized is recognized, without recognizing the entire video frame image, which is beneficial to improving the efficiency of subtitle recognition.

[0088] S500 recognizes the text in the image to be recognized and corrects it based on the speech recognition result to obtain a subtitle file.

[0089] Figure 7 This is provided by an embodiment of the present application Figure 1 The specific process diagram of S500 in Figure 7 As shown, the recognition of the text in the subtitle image to obtain a subtitle file includes:

[0090] S501 obtains the timestamp corresponding to the video frame.

[0091] In the embodiment of the present application, since the generated subtitle files are marked with frame numbers, it is necessary to further calculate the video timestamp corresponding to each frame number. It can be implemented by using ffmpeg or opencv to read the timestamp of the specified frame number. In one implementation, when the timestamp recorded by the video itself is incorrect and the read timestamp < 0, the timestamp needs to be corrected through the frame rate and frameno.

[0092] Taking the processing of opencv as an example:

[0093]

[0094]

[0095] S502 performs text recognition on the text in the image to be recognized and corrects it based on the speech recognition result to obtain subtitle text.

[0096] As Figure 8 shown, Figure 8 This is provided by an embodiment of the present application Figure 7 The specific process diagram of S502 in

[0097] Specifically, the performing text recognition on the text in the image to be recognized and correcting it based on the speech recognition result to obtain subtitle text includes:

[0098] Specifically, the performing text recognition on the text in the image to be recognized to obtain a text recognition result includes:

[0099] S1 splices the images to be recognized in the order of frames from top to bottom in groups, performs text recognition on the spliced multiple images to be recognized in the horizontal direction, and separates the images to be recognized with black bars or white bars;

[0100] S2 separates and restores the recognition results of the spliced multiple images to be recognized in the order of frames.

[0101] S3. Merge the separated recognition results frame by frame to combine frames with the same subtitles, and obtain the text recognition result.

[0102] Here, the number of spliced images in each group is 1 - n, and n is generally limited to within 20. When the number of spliced images is 1, it belongs to the case where no splicing is done.

[0103] In the embodiment of the present application, a batch of images to be recognized are cropped from the video frame images according to step S400. It should be noted that the images to be recognized are spliced in groups from top to bottom in the frame order according to the frame number frame_no, where frame_no is the frame number of the video. The video is composed of a series of consecutive images, and each frame image will have a unique frame number in sequence. As Figure 9 shown, Figure 9 This is a schematic diagram of the spliced image to be recognized provided by the embodiment of the present application. When multiple images to be recognized are spliced together for text recognition, generally, text recognition is performed in the horizontal direction, and the recognition results are: the first line of subtitles, the second line of subtitles, the third line of subtitles, and the fourth line of subtitles. However, there may also be text recognition results in the vertical direction. For example, the recognition results are: the first first first first, one two three four, line line line line, etc. In view of this, the present application provides a solution to insert black bars or white bars in the middle of each image to be recognized to separate them, so as to avoid getting incorrect recognition results in the vertical direction during text recognition, avoid recognition errors, and is conducive to improving the accuracy of subtitle recognition.

[0104] It should be noted that in the embodiment of the present application, there is no need to perform any methods such as black and white binarization processing and edge extraction on the images to be recognized. Directly using the deep learning solution can obtain extremely high accuracy.

[0105] In the embodiment of the present application, the horizontal text is quickly recognized by the ocr method. The paddleocr interface is used to recognize a batch of spliced images to be recognized, and the dt_box and rec_res of a batch of spliced images to be recognized are obtained. Among them, each batch of dt_box is the coordinates of multiple detected text boxes, and rec_res is the recognized multiple texts and confidence probabilities. Then, the dt_box and rec_res are restored and calculated according to the frame number frame_no and position according to the splicing rule, the recognition results are split by frame, and the new frame_no, <dt_box, rec_res> is recorded.

[0106] Splicing the images to be recognized here can effectively utilize the GPU to accelerate the text recognition speed and improve the efficiency of subtitle recognition.

[0107] Meanwhile, according to frame_no, determine the frames without subtitle content in <dt_box, rec_res> and delete them from the result sequence.

[0108] For the case where the number of stitched images is 1, that is, when no stitching is performed, the embodiment is that when the first text-containing image to be recognized is recognized, frame skipping recognition can be performed, and text recognition is performed on the nth image to be recognized after several images to be recognized. Since video subtitles usually have a display time of more than 0.2s, frame skipping recognition will not reduce the accuracy and improves the subtitle recognition efficiency.

[0109] Generally, due to the small probability of misrecognition in the subtitle recognition program, it is necessary to judge whether the images to be recognized in several consecutive frames are the same subtitle according to the overlap ratio of the text recognition results of several consecutive images to be recognized. According to the obtained <dt_box, rec_res> records with frame_no information, judge whether the subtitles of two consecutive frames with each consecutive frame number are the same. For example, if 80% of the area of dt_box overlaps and more than 80% of the content marked by rec_res is the same, it is regarded as the same subtitle. When there are significant text differences between a certain frame and the previous frame, start from this frame and calculate the next subtitle information. After looping this operation, several batches of subtitle information can be obtained. Judging whether several consecutive frames are the same subtitle through the overlap ratio of text recognition results is beneficial to improving the accuracy of subtitle recognition.

[0110] Judge whether it is the same subtitle through the overlap ratio of the text recognition results of several frames in S1. After judgment, several batches of frame sequences with increasing frame numbers and each containing the same content will be generated, and the start frame number and end frame number of each batch of frame sequences will be recorded to calculate the timestamp of the subtitle file. It should be noted that if there is a small number of intervening frames (such as 1-2 frames) between the tail and the head frame numbers of two consecutive frame sequences, but the text recognition content of these two tail and head images to be recognized is the same, this situation occurs because of missed subtitle recognition within the frame. It is necessary to merge these consecutive images to be recognized and re-calibrate the start frame number and end frame number of this merged batch.

[0111] For example, in two consecutive frame sequences, there are 1-2 frames in the middle of the tail frame of the previous sequence and the head frame of the next sequence. Due to the background being too white and the subtitle text also being white, etc., the subtitles of the middle frame images to be recognized are not detected. Then it is necessary to merge the frames to be recognized with the same text recognition content across frames and re-calibrate the start frame number and end frame number of these two image sequences to be recognized for subsequent calculation of the timestamp of the subtitle file. Each image to be recognized corresponding to a frame number has a corresponding time point on the time axis, such as <11, 00:00:00, 366>, <43, 00:00:01, 433>, etc.

[0112] For example:

[0113] [11-43]

[0114] 00:00:00,366-->00:00:01,433

[0115] Hello

[0116] [45-55]

[0117] 00:00:01,500-->00:00:01,833

[0118] Hello

[0119] It can be merged according to the frame numbers as:

[0120] [11-55]

[0121] 00:00:00,366-->00:00:01,833

[0122] Hello

[0123] In the embodiment of the present application, after text recognition of the image to be recognized through S1, frame-by-frame separation and restoration through S2, and merging of frames with the same subtitles through S3, a text recognition result arranged in frame order is obtained.

[0124] S5022. Perform speech text recognition on the audio of the video to obtain a speech recognition result.

[0125] In the embodiment of the present application, for text position detection and localization and text recognition, relatively high-current recognition rate algorithms such as yolo algorithm and paddleocr are used, but are not limited to these software, so no specific limitations are made. For speech recognition, whisper of openai or a speech recognition model of modelscope can be used, etc.

[0126] S5023. Correct the text recognition result according to the speech recognition result to obtain subtitle text.

[0127] In the embodiment of the present application, after obtaining a text recognition result by performing text recognition on the image to be recognized through S5021, it is also necessary to perform speech recognition on the audio part corresponding to the image to be recognized to obtain a speech recognition result.

[0128] Specifically, in very few cases, there will be missing characters, or the situation where multiple frames of text do not exactly correspond, or single characters appear, or non-subtitle text appears, etc. in the text recognition result obtained by recognizing the image to be recognized, resulting in inaccurate text recognition results. At this time, the pyannote_whisper software is used to obtain a speech recognition result by listening to and recording the audio of the video, so as to correct the text recognition result of the video content.

[0129] In the first possible embodiment, since there will be a problem of timestamp deviation between the timestamps of the subtitles recognized from the video content and the timestamps recognized from the audio, and there may be a problem that multiple consecutive texts are merged in the text recognized from the audio, therefore, a method similar to the merge algorithm needs to be used to compare line by line. For example, when performing text recognition on the image to be recognized, the text recognition result is as follows:

[0130] [1-8]

[0131] 00:00:00,033-->00:00:00,266

[0132] Hi

[0133] [11-43]

[0134] 00:00:00,366-->00:00:01,433

[0135] Hello

[0136] And when performing speech text recognition on the audio of the video, the speech recognition result is as follows:

[0137] [00:00.040-->00:01.420]

[0138] Hi, Hello.

[0139] In view of this, since the image text recognition technology only recognizes according to the text on each frame of the image to be recognized, the text recognition result is two text subtitles. And according to the audio to get the speech recognition result, since there is a language logical relationship between "Hi" and "Hello" in the audio and the interval is short, the speech recognition result is one text subtitle. At this time, it can be judged by the program that the content recognized by the speech and the video image is the same, but there are slight differences in the timestamp markings. At this time, the result should be based on the recognition of the video content.

[0140] In the second possible implementation manner, when there are some street signs, book contents or backgrounds with colors close to the text in the calibration position area of the image to be recognized, at this time, subtitle misrecognition is likely to occur in text recognition. For example, when the image content in the video just includes some books, newspapers, etc., at this time, the subtitle content is very likely to also recognize the text here as subtitle text, and the text recognition result is as follows:

[0141] Let's take a picture together #Daily

[0142] Hi, Hello

[0143] By performing text recognition, both "Let's take a photo together" and "#Daily" are recognized as subtitle text. However, the actual subtitle text is only "Hi, hello". In view of this, it is necessary to perform speech text recognition on the audio of the video to obtain the following speech recognition result: Hi, hello. Correct the text recognition result according to the speech recognition result, and only retain the content of "Hi, hello".

[0144] The third possible embodiment is that when the background color of the subtitle area in the image to be recognized is similar to the color of the subtitle text itself, text recognition may not be able to accurately recognize the subtitle text because the text blends with the background. At this time, speech recognition is required to correct the text recognition result. For example, according to the image recognition result, it is "One for the betrothal gift of Yiyi", while the speech recognition result is "One for the betrothal gift of 100 million". At this time, it is necessary to combine natural language processing to correct it to "One for the betrothal gift of 100 million".

[0145] The fourth possible embodiment is that in a certain video frame of text recognition, there may be a situation of missing words or unrecognizable words, and it is necessary to fill in the missing text through adjacent frames or through speech.

[0146] S503. Generate a subtitle file according to the timestamp corresponding to the video frame and the subtitle text.

[0147] According to the timestamp obtained in S501 and the subtitle text obtained in S502, each continuous frame sequence is taken as a group, the start timestamp is calculated, and it is written in the standard format of the srt file. The following video subtitle file can be obtained. Among them, the numbers "00:00:00,366" and "00:00:01,833" represent the start and end timestamps of the subtitle text on this line, that is, the duration of this line of subtitle. Since in the actual operation process, there are usually two processing schemes for the translated subtitles. The first is to directly erase the original subtitles and then paste the translated subtitles. The other is to add and synthesize the translated subtitles at the specified position. To ensure the best experience in both modes, it is necessary to ensure that the subtitles to be synthesized are merged according to the correct timestamp. Therefore, the timestamp of the video image content recognition is required, rather than the timestamp of the speech recognition. Here, through the embodiment, the video image content is recognized to obtain the timestamped subtitles, and a subtitle file in the following format can be obtained: 1

[0149] 00:00:01,040-->00:00:01,366

[0150] Hi 2

[0152] 00:00:01,840-->00:00:03,366

[0153] Hello 3

[0155] 00:00:03,600-->00:00:05,366

[0156] The weather is nice today 4

[0158] 00:00:05,640-->00:00:08,820

[0159] How are you

[0160] For the S600, translate the subtitle file and synthesize it into the video.

[0161] In the embodiment of the present application, the obtained srt subtitle file is batch translated through a large language model or traditional translation tools, and it is ensured that the subtitle content, format, and timestamp correspond. Tools such as ffmpeg are used to synthesize the translated subtitles into the video at the specified position, or an external subtitle form can be used during playback, and the external subtitles need to be encrypted. For example, translating the above srt subtitle file into English format is as follows: 1

[0163] 00:00:01,040-->00:00:01,366

[0164] hi 2

[0166] 00:00:01,840-->00:00:03,366

[0167] hello 3

[0169] 00:00:03,600-->00:00:05,366

[0170] The weather is nice today 4

[0172] 00:00:05,640-->00:00:08,820

[0173] How are you

[0174] As Figure 10 shown, Figure 10 FIG. is a schematic structural diagram of a video subtitle processing system provided by an embodiment of the present application, including:

[0175] A text box acquisition module 710, configured to detect and recognize text in a video frame image to obtain a text box, where the text box includes a subtitle text box;

[0176] The merged text box module 720 is used to merge according to the positions of the text boxes to obtain a merged text box, where the merged text box includes a merged subtitle text box;

[0177] The calibrated subtitle position module 730 statistically calculates the calibrated positions of the merged subtitle text boxes of the video according to the merged subtitle text boxes;

[0178] The image to be recognized acquisition module 740 is used to crop the video frame image according to the calibrated positions of the merged subtitle text boxes to obtain an image to be recognized;

[0179] The subtitle file acquisition module 750 is used to perform text recognition on the image to be recognized and correct it through the speech recognition result to obtain a subtitle file;

[0180] The subtitle file import module 760 is used to translate the subtitle file and synthesize it into the video.

[0181] As Figure 11 shown, Figure 11 is a schematic structural diagram of the terminal device provided by the embodiment of the present application. The computer-readable storage medium 800 of this embodiment includes: a server 810( Figure 11 only one is shown in the figure), a client 820, and a computer program 821 stored in the client 820 and operable on the at least one client 820. The client 820 executes the computer program 821 to send a request to the server 810, and the server 810 feeds back a result to implement the steps in the above method embodiment.

[0182] Those skilled in the art can clearly understand that for the convenience and brevity of description, only the above division of each functional unit and module is used as an example. In actual applications, the above functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiment can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of each functional unit and module are only for the convenience of mutual distinction and do not limit the protection scope of the present application. The specific working processes of the units and modules in the above system can refer to the corresponding processes in the foregoing method embodiments and will not be described in detail here.

[0183] In the above embodiments, the descriptions of the respective embodiments have their own emphases. For parts not detailed or recorded in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0184] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed in this application can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of this application.

[0185] In the embodiments provided in this application, it should be understood that the disclosed device / terminal device and method can be implemented in other ways. For example, the device / terminal device embodiments described above are merely illustrative. For example, the division of the modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of the device or unit can be in an electrical, mechanical or other form.

[0186] In addition, the functional units in each embodiment of this application can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit.

[0187] When the integrated module / unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, to implement all or part of the processes in the above-described embodiment methods of this application, it can also be completed by a computer program instructing relevant hardware. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above-described method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file, or some intermediate form, etc. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium, etc. It should be noted that the content included in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.

[0188] To implement all or part of the processes in the above-described embodiment methods of this application, it can also be completed by a computer program product. When the computer program product runs on a terminal device, the terminal device can implement the steps in the above-described method embodiments when executed.

[0189] The above-described embodiments are only used to illustrate the technical solutions of this application, rather than to limit it. Although this application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included in the protection scope of this application.

Claims

1. A video subtitle processing method, characterized in that, The method includes: Detect and recognize the text in the video frame image to obtain text boxes, where the text boxes include subtitle text boxes; Merge according to the positions of the text boxes to obtain merged text boxes, where the merged text boxes include merged subtitle text boxes; Statistically calculate the calibration positions of the merged subtitle text boxes of the video according to the merged subtitle text boxes; Crop the video frame image according to the positions of the merged subtitle text boxes to obtain an image to be recognized; Recognize the text in the image to be recognized and correct it through the speech recognition result to obtain a subtitle file; Translate the subtitle file and synthesize it into the video.

2. The method according to claim 1, characterized in that The merging according to the positions of the text boxes to obtain merged text boxes includes: Based on the overlapping area of adjacent text boxes, merge and de-duplicate the adjacent text boxes that meet the preset ratio in the horizontal and vertical directions to obtain the merged text boxes.

3. The method according to claim 1, wherein The text boxes further include non-subtitle text boxes, and the merged text boxes further include merged non-subtitle text boxes. The statistically calculating the calibration positions of the merged subtitle text boxes of the video according to the merged subtitle text boxes includes: Detect and recognize all the merged text boxes; Determine the merged text boxes at the four corners of the video frame image and the merged text boxes that are not left-right symmetric with respect to the video frame image as the merged non-subtitle text boxes; Filter the merged non-subtitle text boxes to determine the merged subtitle text boxes; Statistically calculate the positions of the merged subtitle text boxes frame by frame to obtain the calibration positions of the merged subtitle text boxes of the video.

4. The method according to claim 1, wherein The cropping the video frame image according to the calibration positions of the merged subtitle text boxes to obtain an image to be recognized includes: Detect the first distance between the upper border of the merged subtitle text box and the upper boundary of the video frame image; Detect the second distance between the lower border of the merged subtitle text box and the lower boundary of the video frame image; According to the first distance and the second distance, retain the area at the calibration position of the merged subtitle text box and crop the video frame image to obtain an image to be recognized.

5. The method according to claim 1, wherein The recognizing the text in the subtitle image and correcting it through the speech recognition result to obtain a subtitle file includes: Obtain the time stamp corresponding to the video frame; Perform text recognition on the text in the image to be recognized and correct it through the speech recognition result to obtain subtitle text; Generate a subtitle file according to the time stamp corresponding to the video frame and the subtitle text.

6. The method according to claim 5, characterized in that The performing text recognition on the text in the image to be recognized and correcting it through the speech recognition result to obtain subtitle text includes: Perform text recognition on the text in the image to be recognized to obtain a text recognition result; Perform speech text recognition on the audio of the video to obtain a speech recognition result; Correct the text recognition result according to the speech recognition result to obtain subtitle text.

7. The method according to claim 6, wherein The performing text recognition on the text in the image to be recognized to obtain a text recognition result includes: Stitch the to-be-recognized images in groups from top to bottom in frame order, perform text recognition on the stitched multiple to-be-recognized images in the horizontal direction, and separate the stitched to-be-recognized images by black bars or white bars; Separate and restore the recognition results of the stitched multiple to-be-recognized images in frame order; Merge the frames with the same subtitles in frame order for the separated recognition results to obtain the text recognition result.

8. A video subtitle processing system, characterized in that The system includes: A text box acquisition module for detecting and recognizing text in a video frame image to obtain a text box, where the text box includes a subtitle text box; A merged text box module for merging according to the positions of the text boxes to obtain a merged text box, where the merged text box includes a merged subtitle text box; A calibrated subtitle position module for statistically calculating the calibrated positions of the merged subtitle text boxes of the video according to the merged subtitle text boxes; A to-be-recognized image acquisition module for cropping the video frame image according to the calibrated positions of the merged subtitle text boxes to obtain to-be-recognized images; A subtitle file acquisition module for performing text recognition on the to-be-recognized images and correcting them with the speech recognition results to obtain subtitle files; A subtitle file import module for translating the subtitle files and synthesizing them into the video.

9. A terminal device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the method according to any one of claims 1 to 7.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the method according to any one of claims 1 to 7.