Bullet screen-based video subtitle correction method and device

By obtaining barrage data for video subtitle typo detection, and using character recognition and similarity analysis to generate correction information, the existing video subtitle error correction methods are solved, and automated and efficient subtitle error correction and viewing experience are achieved.

CN120343352APending Publication Date: 2025-07-18SHANGHAI KUANYU DIGITAL TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510696226.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-27
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The existing video subtitle error correction methods rely on manual review efficiency and high cost, and the natural language processing-based methods lack effective utilization of video content and audience feedback, resulting in limited error correction accuracy.

Method used

By obtaining barrage data, character recognition technology is used to identify video frames in subtitle intervals, and typos are detected based on barrage vocabulary, including similarity analysis, pronunciation similarity analysis and semantic similarity analysis, and error subtitle correction information is generated for automatic error correction.

Benefits of technology

It improves the efficiency and accuracy of subtitle error correction, improves the viewer's viewing experience, improves the degree of automation of subtitle correction, and ensures the consistency of visual effects after subtitle replacement.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120343352A_ABST
    Figure CN120343352A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a bullet screen-based video subtitle correction method and device. The method comprises the steps that bullet screen data in a to-be-processed video are acquired, and the bullet screen data comprise bullet screen sending time and bullet screen vocabularies; performing character recognition processing on video frames in the subtitle recognition interval of the video to be processed to obtain video subtitles in the video frames; according to the bullet screen vocabularies, carrying out wrongly written character detection on the video subtitles; and if it is detected that wrongly written characters exist in the video subtitles, generating wrong subtitle correction information to correct the video subtitles. According to the scheme, the error points of the subtitles can be judged by utilizing the bullet screen information, so that wrongly written characters in the video subtitles are automatically detected, the corresponding error subtitle correction information is generated, the subtitle correction efficiency and accuracy are improved, and the watching experience of audiences is improved; the character recognition technology is used for automatically recognizing subtitle characters in a video picture, and the automation degree of subtitle correction is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present application relate to the field of Internet technologies, and more particularly, to a method and apparatus for correcting video subtitles based on bullet screens. Background Art

[0002] With the booming development of online video platforms, a large amount of video content has been uploaded and shared. As an important element to assist viewers in understanding video content, the accuracy of video subtitles is crucial. However, in practical applications, due to various reasons such as manual input and translation errors, there are often typos in video subtitles, which not only affect the viewing experience of viewers but may also mislead viewers' understanding of video content.

[0003] Currently, the correction of video subtitles mainly relies on manual review by creators, which is inefficient, costly, and difficult to ensure the timely discovery and correction of all typos. Although there are also some automatic correction methods based on natural language processing technologies, these methods usually only analyze the subtitle text itself and lack the effective utilization of the actual video content and viewers' feedback, resulting in limited correction accuracy. Summary of the Invention

[0004] In view of the above problems, the present application is proposed to provide a method, apparatus, computing device, computer storage medium, and computer program product for correcting video subtitles based on bullet screens that overcome the above problems or at least partially solve the above problems.

[0005] According to one aspect of the embodiments of the present application, a method for correcting video subtitles based on bullet screens is provided, including:

[0006] Obtaining bullet screen data in a video to be processed, where the bullet screen data includes the bullet screen sending time and bullet screen words;

[0007] Performing character recognition processing on video frames within the subtitle recognition interval of the video to be processed to obtain video subtitles in the video frames, where the time span of the subtitle recognition interval is a time window centered on the bullet screen sending time and extending forward by a first preset duration and backward by a second preset duration;

[0008] Detecting typos in the video subtitles according to the bullet screen words;

[0009] If a typo is detected in the video subtitles, generating error subtitle correction information to correct the video subtitles.

[0010] Further, detecting typos in the video subtitles according to the bullet screen words further includes:

[0011] Performing similarity analysis on the bullet screen words and the video subtitles, and determining whether there are typos in the video subtitles according to the similarity analysis results.

[0012] Further, performing a similarity analysis on the bullet screen vocabulary and the video subtitles, and determining whether there are typos in the video subtitles according to the similarity analysis results further includes:

[0013] Performing a pronunciation similarity analysis on the bullet screen vocabulary and the video subtitles, and determining whether there are typos in the video subtitles according to the pronunciation similarity analysis results.

[0014] Further, performing a pronunciation similarity analysis on the bullet screen vocabulary and the video subtitles, and determining whether there are typos in the video subtitles according to the pronunciation similarity analysis results further includes:

[0015] Count the bullet screen word frequencies corresponding to each bullet screen vocabulary within the bullet screen word frequency statistical interval, where the bullet screen word frequency statistical interval spans a time window centered on the bullet screen sending time, extending forward by a third preset duration and backward by a fourth preset duration;

[0016] Calculate the pronunciation similarity between the bullet screen vocabulary with a bullet screen word frequency greater than or equal to the preset word frequency threshold and the subtitle vocabulary in the video subtitles;

[0017] If the pronunciation similarity is greater than or equal to the preset pronunciation similarity threshold, determine that there is a typo in the subtitle vocabulary.

[0018] Further, performing a similarity analysis on the bullet screen vocabulary and the video subtitles, and determining whether there are typos in the video subtitles according to the similarity analysis results further includes:

[0019] If there is a preset target vocabulary in the bullet screen vocabulary, extract the text content to be detected from the bullet screen vocabulary according to the preset target vocabulary;

[0020] Perform a semantic similarity analysis on the text content to be detected and the subtitle vocabulary in the video subtitles, and determine whether there are typos in the video subtitles according to the semantic similarity analysis results.

[0021] Further, the wrong subtitle correction information includes: the subtitle vocabulary with typos, start time, end time, and subtitle vocabulary position information.

[0022] Further, generating the wrong subtitle correction information further includes:

[0023] Taking the video frame where the video subtitle with typos is located as a reference, move forward and backward frame by frame and perform character recognition until the subtitle vocabulary with typos no longer appears in the recognized video subtitles, and record the video timestamps of the video frames where the subtitle vocabulary with typos first appears and last appears as the start time and end time;

[0024] Calculate the subtitle vocabulary position information corresponding to the subtitle vocabulary with typos in the video frame;

[0025] Generate error subtitle correction information based on the misspelled subtitle words, start time, end time, and subtitle word position information.

[0026] Furthermore, the method further includes: replacing the video subtitle with misspelled words according to the error subtitle correction information.

[0027] Furthermore, replacing the video subtitle with misspelled words according to the error subtitle correction information further includes:

[0028] Determine the corrected text after error correction;

[0029] Scale the corrected text according to the first word length of the corrected text and the second word length of the video subtitle with misspelled words;

[0030] According to the error subtitle correction information, replace the video subtitle with misspelled words in the video frame by using the scaled corrected text.

[0031] According to another aspect of the embodiments of the present application, a video subtitle correction device based on bullet screens is provided, including:

[0032] An acquisition module, adapted to acquire bullet screen data in a video to be processed, where the bullet screen data includes bullet screen sending time and bullet screen words;

[0033] A character recognition module, adapted to perform character recognition processing on video frames within a subtitle recognition interval of the video to be processed to obtain video subtitles in the video frames, where the time span of the subtitle recognition interval is a time window centered on the bullet screen sending time and extending forward by a first preset duration and backward by a second preset duration;

[0034] A detection module, adapted to detect misspelled words in the video subtitle according to the bullet screen words;

[0035] A generation module, adapted to generate error subtitle correction information for correcting the video subtitle if misspelled words are detected in the video subtitle.

[0036] According to yet another aspect of the embodiments of the present application, a computing device is provided, including: a processor, a memory, a communication interface, and a communication bus, and the processor, the memory, and the communication interface complete communication with each other through the communication bus;

[0037] The memory is used to store at least one executable instruction, and the executable instruction causes the processor to perform the operations corresponding to the above-mentioned video subtitle correction method based on bullet screens.

[0038] According to another aspect of the embodiments of the present application, a computer storage medium is provided, in which at least one executable instruction is stored, and the executable instruction causes the processor to perform the operations corresponding to the above-mentioned bullet-screen based video subtitle correction method.

[0039] According to still another aspect of the embodiments of the present application, a computer program product is provided, including at least one executable instruction, and the executable instruction causes the processor to perform the operations corresponding to the above-mentioned bullet-screen based video subtitle correction method.

[0040] Based on the bullet-screen based video subtitle correction method and device provided by the embodiments of the present application, the error points of the subtitles can be judged by using the bullet-screen information, so as to automatically detect the typos in the video subtitles and generate corresponding error subtitle correction information, improving the efficiency and accuracy of subtitle error correction and improving the viewing experience of the audience; the character recognition technology is used to automatically recognize the subtitle text in the video picture, improving the automation degree of subtitle correction.

[0041] The above description is only an overview of the technical solutions of the embodiments of the present application. In order to be able to understand the technical means of the embodiments of the present application more clearly, it can be implemented according to the content of the specification. And in order to make the above and other purposes, features and advantages of the embodiments of the present application more obvious and understandable, the following specifically describes the specific embodiments of the present application. Description of the Drawings

[0042] By reading the following detailed description of the preferred embodiments, various other advantages and benefits will become clear to those of ordinary skill in the art. The drawings are only for the purpose of showing the preferred embodiments and are not considered to be a limitation of the embodiments of the present application. And throughout the drawings, the same reference numerals are used to represent the same components. In the drawings:

[0043] Figure 1 A flowchart showing the process of a bullet-screen based video subtitle correction method according to an embodiment of the present application is shown;

[0044] Figure 2A A flowchart showing the process of a bullet-screen based video subtitle correction method according to another embodiment of the present application is shown;

[0045] Figure 2B A schematic diagram of a video frame before video subtitle correction is shown;

[0046] Figure 2C A schematic diagram of a video frame after video subtitle correction is shown;

[0047] Figure 3 A block diagram showing the structure of a bullet-screen based video subtitle correction device according to an embodiment of the present application is shown;

[0048] Figure 4The structural schematic diagram of a computing device according to an embodiment of the present application is shown. Detailed implementation manners

[0049] Exemplary embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although the exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present disclosure can be more thoroughly understood and the scope of the present disclosure can be fully conveyed to those skilled in the art.

[0050] First, the noun terms involved in one or more embodiments of the present application are explained.

[0051] OCR: Optical Character Recognition, optical character recognition.

[0052] The inventors of the present application have found that, as a unique interactive form of online videos, bullet comments reflect the real-time feedback and concerns of viewers when watching videos. In bullet comments, viewers often correct subtitle errors and a large number of recurring words appear, and this information provides new ideas for subtitle correction. Therefore, the inventors of the present application have proposed a correction scheme for video subtitles based on bullet comments through creative labor. The above correction scheme is introduced below in combination with specific embodiments:

[0053] Figure 1 The flowchart of the method for correcting video subtitles based on bullet comments according to an embodiment of the present application is shown, as Figure 1 shown, the method includes the following steps:

[0054] Step S101, obtain the bullet comment data in the video to be processed, where the bullet comment data includes the bullet comment sending time and the bullet comment words.

[0055] The bullet comment data is the data related to bullet comments sent by users. Among them, the bullet comment data includes: the bullet comment sending time and the bullet comment words. The bullet comment sending time refers to the playback time point of the video when the user sends the bullet comment, which is the key identifier for time synchronization between the bullet comment and the video content. For example, the bullet comment sending time is "00:01:30", indicating that this bullet comment was sent when the video was played to 1 minute and 30 seconds. The bullet comment words are the text content input by the user when sending the bullet comment. These words can be comments on the video content, emotional expressions, plot discussions, pointing out typos, etc. For example, "clear sky", "catch bug: clear sky", "typo: clear sky", etc. The bullet comment words can reflect the immediate feedback of the user on the video content.

[0056] Step S102: Perform character recognition processing on the video frames within the subtitle recognition interval of the video to be processed, and obtain the video subtitles in the video frames. The time span of the subtitle recognition interval is a time window centered on the bullet screen sending time, extending forward by a first preset duration and backward by a second preset duration.

[0057] The subtitle recognition interval is a dynamic time window formed by expanding forward and backward on the video timeline based on the bullet screen sending time, and is used to define the operation range of subtitle recognition.

[0058] The first preset duration and the second preset duration are pre-configured time offsets, which determine the front and rear boundaries of the subtitle recognition interval. The values of the first preset duration and the second preset duration can be flexibly set according to actual needs.

[0059] In this step, the first preset duration and the second preset duration can be the same or different. When the values of the first preset duration and the second preset duration are the same, it extends equally in both directions from the bullet screen sending time as the center of the timeline. For example, if both are set to 5 seconds and the bullet screen sending time is the 60th second of the video, the subtitle recognition interval is from the 55th second to the 65th second, and all video frames within this interval will be included in the OCR recognition range.

[0060] If the first preset duration and the second preset duration are different, it extends unequally in both directions from the bullet screen sending time as the center. For example, when the first preset duration is set to 3 seconds and the second preset duration is set to 7 seconds, if the bullet screen is sent at the 60th second, the subtitle recognition interval is from the 57th second to the 67th second. By dynamically adjusting the first preset duration and the second preset duration, the subtitle recognition interval can be adapted to different types of video content and analysis requirements, effectively improving the accuracy and effectiveness of subtitle recognition.

[0061] Extending forward means expanding in the direction of decreasing video timestamp value, that is, obtaining the video frames before the bullet screen sending time for analyzing the video subtitles before the bullet screen is triggered;

[0062] Extending backward means expanding in the direction of increasing video timestamp value, that is, obtaining the video frames after the bullet screen sending time to facilitate tracking the video subtitles after the bullet screen is sent.

[0063] Specifically, after detecting the bullet screen, it indicates that the user may have corrected the typos in the video subtitles through the bullet screen. Therefore, character recognition processing can be performed on the video frames near the time point where the bullet screen appears in the video.

[0064] Since the bullet screen sending time is associated with the playback time point of the video, the subtitle recognition interval can be constructed by extending the first preset duration forward and the second preset duration backward respectively with the bullet screen sending time as the reference point. The subtitle recognition interval is [bullet screen sending time - first preset duration, bullet screen sending time + second preset duration]. Then, the video frames within the subtitle recognition interval are extracted from the video, and an optical character recognition engine (such as an OCR or CRNN model) is used to perform character recognition on the video frames. Through character recognition processing, the video subtitles in each video frame within the subtitle recognition interval can be obtained.

[0065] Step S103, perform typos detection on the video subtitles according to the bullet screen vocabulary.

[0066] After the video subtitles are recognized, using the bullet screen vocabulary as the basis for typos detection, perform typos detection on the video subtitles to determine whether there are typos in the video subtitles. For example, based on a similarity analysis algorithm, calculate the semantic similarity and / or pronunciation similarity between the bullet screen vocabulary and the subtitle vocabulary in the video subtitles, and determine whether there are typos in the video subtitles according to the similarity results; or, based on a context association analysis method, analyze whether there are emotional and semantic conflicts between the video subtitles and the bullet screen vocabulary. If there are obvious emotional and semantic conflicts between the video subtitles and the bullet screen vocabulary, then combine with an error rule library to judge whether there are typos in the video subtitles. mainly combine the emotional tendency and semantic information expressed by the bullet screen vocabulary, analyze the context of the video subtitles. If there are obvious emotional and semantic conflicts between the video subtitles and the bullet screen vocabulary, and there are corresponding error words in the error rule library, then determine that there are typos in the video subtitles. Other typos detection methods can also be used, which will not be elaborated here.

[0067] Step S104, if typos are detected in the video subtitles, generate error subtitle correction information to correct the video subtitles.

[0068] In the case where typos are detected in the video subtitles, error subtitle correction information can be generated. The error subtitle correction information contains the key information required to correct the video subtitles. For example, it can include the subtitle vocabulary with typos, start time, end time, subtitle vocabulary position information, etc. Thus, the video subtitles can be corrected according to the generated error subtitle correction information.

[0069] According to the video subtitle correction method based on bullet screen provided by the embodiments of the present application, using bullet screen information can determine the error points of the subtitles, thereby automatically detecting typos in the video subtitles and generating corresponding error subtitle correction information, improving the efficiency and accuracy of subtitle error correction, and improving the viewing experience of the audience; using character recognition technology to automatically recognize the subtitle text in the video picture improves the automation degree of subtitle correction.

[0070] Figure 2A shows a schematic flowchart of a bullet-screen based video subtitle correction method according to another embodiment of the present application, as Figure 2A shown, the method includes the following steps:

[0071] Step S201, obtain bullet-screen data in the video to be processed, where the bullet-screen data includes the bullet-screen sending time and bullet-screen words.

[0072] Specifically, the video to be processed refers to a video with a need for video subtitle correction. After the video to be processed is successfully uploaded and released for a period of time, the number of bullet screens sent by users for the video to be processed is monitored and counted in real time, and the number of bullet screens is compared with a preset bullet-screen number threshold. If the number of bullet screens is greater than or equal to the preset bullet-screen number threshold, the bullet-screen data in the video to be processed can be obtained from the bullet-screen sentence library. Among them, those skilled in the art can flexibly set the preset bullet-screen number threshold according to actual needs. For example, it can be 50, 200, 500, etc.

[0073] Performing corresponding video subtitle correction after the number of bullet screens reaches the preset bullet-screen number threshold is convenient for having enough bullet screens for video subtitle correction, thereby improving the accuracy of video subtitle correction based on bullet screens.

[0074] Step S202, perform character recognition processing on the video frames within the subtitle recognition interval of the video to be processed to obtain the video subtitles in the video frames, where the time span of the subtitle recognition interval is a time window centered on the bullet-screen sending time, extending forward by a first preset duration and backward by a second preset duration.

[0075] After detecting a bullet screen, it indicates that the user may have corrected the misspelled words in the video subtitles through the bullet screen. Therefore, character recognition processing can be performed on the video frames near the time point where the bullet screen appears in the video.

[0076] Specifically, the way to determine which video frames to perform character recognition processing can be by constructing a subtitle recognition interval. Among them, the subtitle recognition interval is a dynamic time window formed by expanding forward and backward along the video time axis based on the bullet-screen sending time, and is used to limit the operation range of subtitle recognition.

[0077] Since the bullet-screen sending time is associated with the playback time point of the video, the subtitle recognition interval can be constructed by extending forward by a first preset duration and backward by a second preset duration respectively based on the bullet-screen sending time. Among them, the subtitle recognition interval can be expressed as [bullet-screen sending time - first preset duration, bullet-screen sending time + second preset duration].

[0078] After determining the subtitle recognition interval, video frames within the subtitle recognition interval can be extracted from the video to be processed, and an optical character recognition engine (such as an OCR or CRNN model) can be used to perform character recognition on the video frames. Through character recognition processing, video subtitles in each video frame within the subtitle recognition interval can be obtained.

[0079] After recognizing the video subtitles, it is necessary to detect whether there are typos in the video subtitles based on the barrage vocabulary. For example, a similarity analysis can be performed on the barrage vocabulary and the video subtitles, and it can be determined whether there are typos in the video subtitles according to the similarity analysis result. The similarity analysis result reflects the similarity between the barrage vocabulary and the video subtitles. The similarity analysis result can be a similarity degree. Specifically, step S203 can be used to identify whether there are typos in the video subtitles:

[0080] Step S203: Perform a pronunciation similarity analysis on the barrage vocabulary and the video subtitles, and determine whether there are typos in the video subtitles according to the pronunciation similarity analysis result.

[0081] Specifically, the video subtitles can be segmented to obtain subtitle words. Then, the barrage vocabulary and the subtitle words can be respectively converted into pinyin by using speech recognition and / or pinyin conversion technology. The pronunciation similarity analysis algorithm is used to analyze the pronunciation similarity between the pinyin of the barrage vocabulary and the pinyin of the subtitle words, and a pronunciation similarity analysis result is obtained. This pronunciation similarity analysis result reflects the similarity between the pinyin of the barrage vocabulary and the pinyin of the subtitle words. Thus, it can be determined whether there are typos in the video subtitles according to this pronunciation similarity analysis result.

[0082] For example, the edit distance algorithm or the pinyin similarity algorithm is used to calculate the pronunciation similarity between the pinyin of the barrage vocabulary and the pinyin of the subtitle words, and it is judged whether the pronunciation similarity is greater than or equal to a preset pronunciation similarity threshold. If so, it is determined that the two have similar pronunciations and thus it is determined that there is a typo in the subtitle word, and thus it is determined that there is a typo in the video subtitle; if not, it is determined that the two do not have similar pronunciations and thus it is determined that there is no typo in the video subtitle.

[0083] In an alternative embodiment, if there are typos in the video subtitles, many users may correct the typos in the video subtitles. Therefore, the bullet screen words used for correcting typos may appear repeatedly within a short period of time. There may be typos in the video subtitles associated with such bullet screen words. Therefore, it is possible to first perform bullet screen word repetition detection. For each time point when a bullet screen appears, count the word frequency of the bullet screen words that appear in the bullet screens near that time point. Specifically, count the bullet screen word frequencies corresponding to each bullet screen word within the bullet screen word frequency statistics interval, where the bullet screen word frequency statistics interval spans a time window that extends third preset duration forward and fourth preset duration backward centered on the bullet screen sending time; the bullet screen word frequency statistics interval is a time window formed by extending a specific duration forward and backward from the sending time of a single bullet screen, and is used to count the occurrence frequencies of all bullet screen words within this window. This interval includes the context bullet screens of a single bullet screen into the analysis scope in the way of "extending forward and backward" to count the word frequencies of bullet screen words within a short period of time. The third preset duration and the fourth preset duration may be the same or different.

[0084] Then, compare the bullet screen word frequency with a preset word frequency threshold, and screen out the bullet screen words whose bullet screen word frequency is greater than or equal to the preset word frequency threshold. The preset word frequency threshold is a word frequency critical value preset in the word frequency statistics task and is used to screen or filter bullet screen words that meet specific word frequency requirements. The screened bullet screen words are potential error words. Calculate the pronunciation similarity between the bullet screen words whose bullet screen word frequency is greater than or equal to the preset word frequency threshold and the subtitle words in the video subtitles. For example, use the edit distance algorithm or the pinyin similarity algorithm to calculate the pronunciation similarity between the pinyin of the bullet screen words and the pinyin of the subtitle words, and determine whether the pronunciation similarity is greater than or equal to the preset pronunciation similarity threshold. If so, it can be determined that they have similar pronunciations, and thus it can be determined that there are typos in the subtitle words, and further it can be determined that there are typos in the video subtitles; if not, it can be determined that they do not have similar pronunciations, and further it can be determined that there are no typos in the video subtitles.

[0085] In an alternative embodiment, for the case where there are typos in the video subtitles, when the user sends a barrage, they may input a target word indicating the wrong semantics. For example, target words such as "typo", "error", "misspelling", "finding bugs", etc. Therefore, a target word library can be established in advance. The target word library stores target words that can represent wrong semantics. The barrage words can be matched with the target word library to detect whether the barrage words contain preset target words. If there is a match, it can be determined that there are preset target words in the barrage words. Then, the text content to be detected is extracted from the barrage words according to the preset target words. For example, the text content after the preset target word can be extracted. For the case where there are punctuation marks after the preset target word, the interference items (punctuation marks) need to be removed. Then, semantic similarity analysis is performed on the text content to be detected and the subtitle words in the video subtitles. For example, natural language processing technologies such as word vector models (Word2Vec, GloVe, etc.) or pre-trained language models (BERT, etc.) can be used to calculate the semantic similarity between the text content to be detected and the subtitle words. Whether there are typos in the video subtitles is determined according to the semantic similarity analysis result. The semantic similarity analysis result shows the semantic similarity between the text content to be detected and the subtitle words. The higher the semantic similarity, the lower the possibility of having typos. On the contrary, the lower the semantic similarity, the higher the possibility of having typos.

[0086] In an alternative embodiment, for the case where there are typos in the video subtitles, when the user sends a bullet comment, they may input a target word indicating the wrong semantics. For example, target words such as "typo", "error", "misspelling", "finding bugs", etc. Therefore, a target word library can be established in advance. The target word library stores target words that can represent wrong semantics. The bullet comment words can be matched with the target word library to detect whether the bullet comment words contain the preset target words. If there is a match, it can be determined that the bullet comment words contain the preset target words, and then the text content to be detected can be extracted from the bullet comment words according to the preset target words. For example, the text content after the preset target word can be extracted. For the case where there is a punctuation mark after the preset target word, the interference item (punctuation mark) needs to be removed; in order to improve the accuracy of typo recognition, matching can be performed from two dimensions: semantic and phonetic similarity. Specifically, semantic similarity analysis is performed on the text content to be detected and the subtitle words in the video subtitles to obtain the semantic similarity analysis result; phonetic similarity analysis is performed on the bullet comment words and the subtitle words to obtain the phonetic similarity analysis result; according to the semantic similarity analysis result and the phonetic similarity analysis result, it is determined whether there are typos in the subtitle words. For example, it is determined whether the semantic similarity analysis result and the phonetic similarity analysis result meet the preset conditions. For example, the phonetic similarity result is the phonetic similarity degree, and the phonetic similarity degree is greater than or equal to the preset phonetic similarity threshold, and the semantic similarity analysis result is the semantic similarity degree, and the semantic similarity degree is less than or equal to the preset semantic similarity threshold, then it is determined that there are typos in the subtitle words.

[0087] It should be noted that this application can use multiple methods at the same time to determine whether there are typos in the video subtitles. For example, word frequency and target words, semantic matching combination.

[0088] Step S204, if it is detected that there are typos in the video subtitles, then based on the video frame where the video subtitle with typos is located, move forward and backward frame by frame and perform character recognition until the subtitle words with typos no longer appear in the recognized video subtitles, and record the video timestamps of the video frames of the video subtitles with typos that first appear and last appear as the start time and end time.

[0089] In the case where typos are detected in the video subtitles, the video frames corresponding to the video subtitles can be determined. Taking this video frame as a reference point, a frame-by-frame scanning mechanism is started before and after. Specifically, in the order of video playback, character recognition is performed on the corresponding video frames frame by frame forward and backward, the corresponding video subtitles are extracted, and it is checked whether there are subtitle words with typos in the above video subtitles. Once the subtitle words with typos detected previously no longer appear in the video subtitles recognized in a certain video frame, the scanning process ends. At this time, the timestamp corresponding to the video frame where the subtitle word with a typo first appears is recorded as the start time, and the timestamp corresponding to the video frame where the subtitle word with a typo last appears is recorded as the end time.

[0090] Step S205, calculate the subtitle word position information corresponding to the subtitle words with typos in the video frame.

[0091] After determining the video frame interval with typos, using image processing technology, analyze the pixel position information of the subtitle words in the video frame to accurately obtain the specific xy position coordinates of the subtitle.

[0092] Specifically, use image segmentation technology (such as color threshold segmentation, edge detection, or deep learning semantic segmentation model) to separate the subtitle area from the video frame. Since subtitles usually have fixed background colors, font colors, or border features (such as black background with white text, white stroke), therefore, the subtitle area can be quickly locked by setting color channel thresholds (such as RGB value range) or using the Hough transform to detect rectangular contours. For example, if the subtitle background is black and the text is white, after converting the image corresponding to the video frame into a grayscale image, the subtitle can be extracted by setting a brightness threshold (such as pixel value > 200 is recognized as the text area).

[0093] For the separated subtitle area, use connected component analysis or OCR (Optical Character Recognition) post-processing method to further locate the pixel coordinates of the typo words.

[0094] Among them, connected component analysis means regarding the text in the subtitle area as connected pixel blocks, and obtaining the upper left and lower right coordinates by marking the bounding box of each character. For example, if the four characters "artificial intelligence" appear as four connected areas in the image, the (x1, y1, x2, y2) coordinates of each character block can be calculated respectively.

[0095] OCR post-processing means that during the character recognition process, modern OCR engines will output the accurate coordinate information of each recognized character. Directly extract the coordinates of each character in the typo word, and then integrate them into the position information at the word level (such as the upper left coordinate of the first character is the start position of the word, and the lower right coordinate of the last character is the end position).

[0096] Combine the character coordinates with the subtitle text content to determine the character index of the wrong word in the entire subtitle text (such as "人工至能" in "欢迎学习人工至能课程", the starting index is 6 and the ending index is 10), and calculate its two-dimensional pixel coordinates in the video frame. Finally, the character index, the upper left corner coordinates (x_min, y_min) and lower right corner coordinates (x_max, y_max) of the entire word, and the word width and height (width = x_max-x_min, height = y_max-y_min) are encapsulated as structured data as the location information of the subtitle word.

[0097] Through this position information, the display position of the typo in the video screen can be accurately located, providing an accurate reference for subsequent corrections.

[0098] Step S206, generating wrong subtitle correction information according to the subtitle words with wrong characters, the start time, the end time, and the subtitle word position information.

[0099] After determining the start time, end time, subtitle vocabulary position information and subtitle vocabulary with typos, the above information can be integrated to generate wrong subtitle correction information. Therefore, the wrong subtitle correction information contains four core parts: the subtitle vocabulary with typos, the start time and end time of the typos in the video, and the subtitle vocabulary position information. This information is structured and organized together to form a complete correction instruction set. For example, the correction information may be stored in JSON format. In this way, subsequent correction processing can quickly obtain all necessary information and achieve accurate subtitle correction. In addition, the wrong subtitle correction information can also include correction text.

[0100] Step S207: Replace the video subtitles with typos according to the erroneous subtitle correction information.

[0101] Based on the generated wrong subtitle correction information, the video subtitles with wrong characters are replaced. First, the video frame interval containing the wrong characters is located according to the start time and the end time, and then for each video frame, the subtitle words with wrong characters are found in the corresponding video subtitles according to the subtitle word position information, and replaced with correct words, for example, text stickers can be used for replacement, or the wrong subtitle words can be directly replaced.

[0102] In an alternative implementation, a typo correction request containing the above error subtitle correction information can be sent to the processing end. For example, the typo correction request can be preferentially sent to the review interface of the video publisher (creator) to remind the video publisher to confirm the errors in the video subtitles and make fine-tuning modifications to the subtitles. To avoid affecting the viewing experience of users due to the long-term non-response of the video publisher, a timeout can be preset, and the current time is compared with the preset timeout. If the current time exceeds the preset timeout, that is, the video publisher fails to process the typo correction request within the specified time, the next review process is entered, and the typo correction request is sent to the corresponding reviewer of the video platform. The reviewer reviews the typo correction request to determine whether the detected typos are accurate and whether the correction suggestions are reasonable. If the review is passed, the review result and the error subtitle correction information are associated and stored in the review library.

[0103] For the case where the reviewer conducts the review, the video playback client on the user side needs to communicate with the server regularly to obtain the reviewed error subtitle correction information. When the user plays the video, the video playback client replaces the original error subtitle words with text stickers at the subtitle display position according to the obtained error subtitle correction information.

[0104] To ensure the visual effect of video viewing, after determining the corrected text after error correction, the corrected text can be scaled according to the first text length of the corrected text and the second text length of the video subtitle with typos. Then, according to the error subtitle correction information, the scaled corrected text is used to replace the video subtitle with typos in the video frame.

[0105] Specifically, the first text length of the corrected text and the second text length of the video subtitle with typos are calculated in combination with factors such as character type (Chinese, English, numbers, etc.), font style (such as Song typeface, boldface), and font size.

[0106] By comparing the first text length and the second text length, the length difference between the two is analyzed. If the corrected text is longer than the subtitle to be replaced, subsequent reduction processing is required; if the corrected text is shorter, enlargement processing is required to adapt to the display space of the subtitle to be replaced, ensuring that the replaced text occupies the same area as the original error text. In this step, the entire video subtitle can be replaced, or only the subtitle words with typos can be replaced. For the case of only replacing the subtitle words with typos, the second text length here refers to the text length of the subtitle words with typos.

[0107] When the text length of the corrected text is greater than the text length of the subtitle to be replaced, the overall display size of the corrected text can be reduced by adjusting the font size, character spacing, and line spacing. Specifically, the font size is gradually decreased according to a certain scaling factor (determined based on the ratio of the length difference to the original subtitle length), while appropriately compressing the spacing between characters and between lines. During the reduction process, the display effect of the corrected text is monitored in real time to ensure that the text is clear and distinguishable, and will not become blurred or difficult to read due to excessive reduction.

[0108] If the text length of the corrected text is less than the text length of the subtitle to be replaced, the font size, character spacing, and line spacing can be increased accordingly to enable the corrected text to better fill the display area of the subtitle to be replaced. During the enlargement operation, it is also necessary to control the enlargement ratio to prevent the font size from being too large and covering other text, or exceeding the reasonable display range of the video frame, affecting the overall aesthetics and coordination. At the same time, the system will perform edge smoothing processing on the enlarged text to avoid jagged edges due to font size enlargement and ensure the visual effect of the subtitle.

[0109] After completing the scaling process of the corrected text, based on the error subtitle correction information generated previously, accurately locate the specific position of the video subtitle with typos in the video frame. The subtitle word position information recorded in the error subtitle correction information provides an accurate guide for the replacement operation.

[0110] According to these position information, accurately cover the scaled corrected text to the subtitle area with typos in the original video frame. During the replacement process, it is also ensured that the display style (such as font, color, bold, underline, etc.) of the corrected text is consistent with the original subtitle, maintaining the unity of the video subtitle style.

[0111] Figure 2B The schematic diagram of the video frame before video subtitle correction is shown, as Figure 2B shown. In this video frame, two types of bullet screen situations and video subtitles are schematically listed. The bullet screens include: Clear sky, Clear sky, Bug report: Clear sky, Clear sky. Among them, 1 refers to sending the bullet screen "Clear sky" alone, 2 refers to sending the bullet screen "Bug report: Clear sky", 3 refers to the video subtitle "Clear the sky". Using the method provided by this application, it can be determined that there are typos in the video subtitle and the typos are corrected. Replace "Clear" in the video subtitle with "Clear sky". The replaced video frame is as Figure 2C shown. In Figure 2C , the video subtitle is corrected to 4 "Clear sky".

[0112] According to the method for correcting video subtitles based on bullet screens provided by the embodiments of the present application, the bullet screen information can be used to determine the error points of the subtitles, so as to automatically detect the typos in the video subtitles and generate corresponding error subtitle correction information, improving the efficiency and accuracy of subtitle error correction and enhancing the viewing experience of the audience; the character recognition technology is used to automatically recognize the subtitle text in the video picture, improving the automation degree of subtitle correction; the scaling process is carried out in combination with the text length of the corrected text and the text length of the subtitle to be replaced, and the corrected text after the scaling process is used for replacement, ensuring that the replaced text occupies the same area as the original wrong text and maintaining the consistency of subtitle typesetting, thus enhancing the visual effect of video viewing.

[0113] Figure 3 Fig. shows a structural block diagram of a device for correcting video subtitles based on bullet screens according to an embodiment of the present application. As Figure 3 shown, the device includes: an acquisition module 301, a character recognition module 302, a detection module 303, and a generation module 304.

[0114] The acquisition module 301 is adapted to acquire bullet screen data in a video to be processed, where the bullet screen data includes the bullet screen sending time and bullet screen vocabulary.

[0115] The character recognition module 302 is adapted to perform character recognition processing on video frames within a subtitle recognition interval of the video to be processed to obtain video subtitles in the video frames, where the time span of the subtitle recognition interval is a time window centered on the bullet screen sending time and extending forward by a first preset duration and backward by a second preset duration.

[0116] The detection module 303 is adapted to detect typos in the video subtitles according to the bullet screen vocabulary.

[0117] The generation module 304 is adapted to generate error subtitle correction information for correcting the video subtitles if typos are detected in the video subtitles.

[0118] Optionally, the detection module is further adapted to: perform similarity analysis on the bullet screen vocabulary and the video subtitles, and determine whether there are typos in the video subtitles according to the similarity analysis result.

[0119] Optionally, the detection module is further adapted to: perform pronunciation similarity analysis on the bullet screen vocabulary and the video subtitles, and determine whether there are typos in the video subtitles according to the pronunciation similarity analysis result.

[0120] Optionally, the detection module is further adapted to: count the bullet screen word frequencies corresponding to each bullet screen vocabulary within a bullet screen word frequency statistics interval, where the time span of the bullet screen word frequency statistics interval is a time window centered on the bullet screen sending time and extending forward by a third preset duration and backward by a fourth preset duration.

[0121] Calculate the pronunciation similarity between the bullet screen vocabulary with a bullet screen word frequency greater than or equal to a preset word frequency threshold and the subtitle vocabulary in the video subtitle;

[0122] If the pronunciation similarity is greater than or equal to a preset pronunciation similarity threshold, it is determined that there are typos in the subtitle vocabulary.

[0123] Optionally, the detection module is further adapted to: if there is a preset target vocabulary in the bullet screen vocabulary, extract the text content to be detected from the bullet screen vocabulary according to the preset target vocabulary;

[0124] Perform semantic similarity analysis on the text content to be detected and the subtitle vocabulary in the video subtitle, and determine whether there are typos in the video subtitle according to the result of the semantic similarity analysis.

[0125] Optionally, the incorrect subtitle correction information includes: the subtitle vocabulary with typos, start time, end time, and subtitle vocabulary position information.

[0126] Optionally, the generation module is further adapted to: based on the video frame where the video subtitle with typos is located, move forward and backward frame by frame and perform character recognition until the subtitle vocabulary with typos no longer appears in the recognized video subtitle, and record the video timestamps of the video frames where the subtitle vocabulary with typos first appears and last appears as the start time and end time;

[0127] Calculate the subtitle vocabulary position information corresponding to the subtitle vocabulary with typos in the video frame;

[0128] Generate incorrect subtitle correction information according to the subtitle vocabulary with typos, start time, end time, and subtitle vocabulary position information.

[0129] Optionally, the device further includes: a replacement module, adapted to perform replacement processing on the video subtitle with typos according to the incorrect subtitle correction information.

[0130] Optionally, the replacement module is further adapted to: determine the corrected text after error correction;

[0131] Perform scaling processing on the corrected text according to the first text length of the corrected text and the second text length of the video subtitle with typos;

[0132] Use the scaled corrected text to replace the video subtitle with typos in the video frame according to the incorrect subtitle correction information.

[0133] The descriptions of the above modules refer to the corresponding descriptions in the method embodiments and will not be elaborated here.

[0134] According to the video subtitle correction device based on bullet screens provided by the embodiments of the present application, the error points of the subtitles can be determined using bullet screen information, so as to automatically detect typos in the video subtitles and generate corresponding error subtitle correction information, improving the efficiency and accuracy of subtitle error correction and enhancing the viewing experience of the audience; the character recognition technology is used to automatically recognize the subtitle text in the video picture, improving the automation degree of subtitle correction.

[0135] The embodiments of the present application provide a non-volatile computer storage medium, and the computer storage medium stores at least one executable instruction or computer program, and the executable instruction or computer program can enable a processor to execute the operations corresponding to the method for correcting video subtitles based on bullet screens in any of the above method embodiments.

[0136] The embodiments of the present application provide a computer program product, and the computer program product includes at least one executable instruction or computer program, and the executable instruction or computer program can enable a processor to execute the operations corresponding to the method for correcting video subtitles based on bullet screens in any of the above method embodiments.

[0137] Figure 4 The structural schematic diagram of the computing device embodiment of the present application is shown, and the specific implementation of the computing device is not limited in the specific embodiments of the present application.

[0138] As Figure 4 shown, the computing device may include: a processor 402, a communications interface 404, a memory 406, and a communication bus 408.

[0139] Among them: the processor 402, the communications interface 404, and the memory 406 communicate with each other through the communication bus 408. The communications interface 404 is used to communicate with network elements of other devices such as clients or other servers. The processor 402 is used to execute the program 410, and specifically can execute the relevant steps in the method embodiment for correcting video subtitles based on bullet screens for the computing device described above.

[0140] Specifically, the program 410 may include program codes, and the program codes include computer operation instructions.

[0141] The processor 402 may be a central processing unit (CPU), or an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present application. One or more processors included in the computing device may be of the same type of processor, such as one or more CPUs; or may be of different types of processors, such as one or more CPUs and one or more ASICs.

[0142] A memory 406 for storing a program 410. The memory 406 may include a high-speed RAM memory and may also include non-volatile memory, such as at least one disk memory.

[0143] The program 410 may specifically be configured to cause the processor 402 to execute the method for correcting video subtitles based on bullet screens in any of the above method embodiments. For the specific implementation of each step in the program 410, reference may be made to the corresponding steps and descriptions in the corresponding units in the above embodiments of the method for correcting video subtitles based on bullet screens, which will not be elaborated herein. Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the devices and modules described above may refer to the corresponding process descriptions in the foregoing method embodiments, which will not be repeated herein.

[0144] The algorithms and displays provided herein are not inherently related to any particular computer, virtual system, or other device. Various general-purpose systems may also be used in conjunction with the teachings provided herein. The structure required to construct such systems will be apparent from the above description. In addition, the embodiments of the present application are not directed to any specific programming language. It should be understood that the content of the embodiments of the present application described herein can be implemented using various programming languages, and the description of the specific language above is for disclosing the best mode of the embodiments of the present application.

[0145] In the specification provided herein, a large number of specific details are set forth. However, it can be understood that the embodiments of the present application may be practiced without these specific details. In some instances, well-known methods, structures, and technologies have not been shown in detail so as not to obscure the understanding of this specification.

[0146] Similarly, it should be understood that, for the purpose of streamlining the present disclosure and facilitating the understanding of one or more of the various inventive aspects, in the foregoing description of the exemplary embodiments of the embodiments of the present application, the various features of the embodiments of the present application are sometimes grouped together into a single embodiment, figure, or description thereof. However, the disclosed method should not be construed as reflecting an intention that the claimed embodiments of the present application require more features than are expressly recited in each claim. Rather, as reflected in the following claims, the inventive aspects lie in less than all the features of the single foregoing disclosed embodiment. Thus, the claims following the detailed description are hereby expressly incorporated into the detailed description, with each claim standing on its own as a separate embodiment of the embodiments of the present application.

[0147] Those skilled in the art will appreciate that the modules in the devices in the embodiments can be adaptively changed and disposed in one or more devices different from the embodiments. The modules or units or components in the embodiments can be combined into one module or unit or component, and in addition, they can be divided into multiple sub-modules or sub-units or sub-components. Any combination can be adopted for all the features disclosed in this specification (including the accompanying claims, abstract, and drawings) and all the processes or units of any method or device so disclosed, unless at least some of such features and / or processes or units are mutually exclusive. Each feature disclosed in this specification (including the accompanying claims, abstract, and drawings) can be replaced by an alternative feature providing the same, equivalent, or similar purpose, unless otherwise expressly stated.

[0148] In addition, those skilled in the art will be able to understand that, although some of the embodiments described herein include certain features included in other embodiments rather than other features, the combination of the features of different embodiments means that it is within the scope of the embodiments of the present application and forms different embodiments. For example, in the following claims, any one of the claimed embodiments can be used in any combination.

[0149] Each component embodiment of the embodiments of the present application can be implemented in hardware, or in software modules running on one or more processors, or in a combination thereof. Those skilled in the art should understand that a microprocessor or a digital signal processor (DSP) can be used in practice to implement some or all of the functions of some or all of the components according to the embodiments of the present application. The embodiments of the present application can also be implemented as a device or apparatus program (such as a computer program and a computer program product) for executing part or all of the methods described herein. Such a program implementing the embodiments of the present application can be stored on a computer-readable medium, or can be in the form of one or more signals. Such signals can be downloaded from an Internet website, or provided on a carrier signal, or provided in any other form.

[0150] It should be noted that the above embodiments illustrate the embodiments of the present application rather than limit the embodiments of the present application, and those skilled in the art can design alternative embodiments without departing from the scope of the appended claims. In the claims, any reference signs placed between parentheses shall not be construed as limiting the claim. The word "comprising" does not exclude the presence of elements or steps not listed in the claim. The word "a" or "an" preceding an element does not exclude the presence of a plurality of such elements. The embodiments of the present application can be implemented by means of hardware including several different elements and by means of a suitably programmed computer. In the unit claims listing several devices, several of these devices can be embodied by the same item of hardware. The use of the words first, second, and third, etc. does not denote any order. These words can be interpreted as names.

Claims

1. A method for correcting video subtitles based on bullet screens, comprising: Obtaining bullet screen data in a video to be processed, wherein the bullet screen data includes the bullet screen sending time and bullet screen words; Performing character recognition processing on video frames within a subtitle recognition interval of the video to be processed to obtain video subtitles in the video frames, wherein the time span of the subtitle recognition interval is a time window centered on the bullet screen sending time, extending forward by a first preset duration and backward by a second preset duration; Detecting spelling mistakes in the video subtitles according to the bullet screen words; If a spelling mistake is detected in the video subtitles, generating error subtitle correction information for correcting the video subtitles.

2. The method according to claim 1, wherein, The detecting spelling mistakes in the video subtitles according to the bullet screen words further includes: Performing similarity analysis on the bullet screen words and the video subtitles, and determining whether there are spelling mistakes in the video subtitles according to the similarity analysis result.

3. The method according to claim 2, wherein The performing similarity analysis on the bullet screen words and the video subtitles, and determining whether there are spelling mistakes in the video subtitles according to the similarity analysis result further includes: Performing pronunciation similarity analysis on the bullet screen words and the video subtitles, and determining whether there are spelling mistakes in the video subtitles according to the pronunciation similarity analysis result.

4. The method according to claim 3, wherein The performing pronunciation similarity analysis on the bullet screen words and the video subtitles, and determining whether there are spelling mistakes in the video subtitles according to the pronunciation similarity analysis result further includes: Counting the bullet screen word frequencies corresponding to each bullet screen word within a bullet screen word frequency statistics interval, wherein the time span of the bullet screen word frequency statistics interval is a time window centered on the bullet screen sending time, extending forward by a third preset duration and backward by a fourth preset duration; Calculating the pronunciation similarity between bullet screen words with bullet screen word frequencies greater than or equal to a preset word frequency threshold and subtitle words in the video subtitles; If the pronunciation similarity is greater than or equal to a preset pronunciation similarity threshold, determining that there is a spelling mistake in the subtitle words.

5. The method according to any one of claims 2-4, wherein, The performing similarity analysis on the bullet screen words and the video subtitles, and determining whether there are spelling mistakes in the video subtitles according to the similarity analysis result further includes: If there is a preset target word in the bullet screen words, extracting the text content to be detected from the bullet screen words according to the preset target word; Performing semantic similarity analysis on the text content to be detected and subtitle words in the video subtitles, and determining whether there are spelling mistakes in the video subtitles according to the semantic similarity analysis result.

6. The method according to any one of claims 1-5, wherein, The error subtitle correction information includes: subtitle words with spelling mistakes, start time, end time, and subtitle word position information.

7. The method according to claim 6, wherein The generating error subtitle correction information further includes: Taking the video frame where the video subtitle with a spelling mistake is located as a reference, moving forward and backward frame by frame and performing character recognition until the subtitle words with the spelling mistake no longer appear in the recognized video subtitles, and recording the video timestamps of the video frames where the subtitle words with the spelling mistake first appear and last appear as the start time and end time; Calculating the subtitle word position information corresponding to the subtitle words with the spelling mistake in the video frames; Generate error subtitle correction information based on the subtitle vocabulary with typos, start time, end time, and subtitle vocabulary position information.

8. The method according to any one of claims 1-7, wherein, The method further includes: Perform a replacement process on the video subtitle with typos according to the error subtitle correction information.

9. The method according to claim 8, wherein, The performing a replacement process on the video subtitle with typos according to the error subtitle correction information further includes: Determine the corrected text after error correction; Perform a scaling process on the corrected text according to the first character length of the corrected text and the second character length of the video subtitle with typos; According to the error subtitle correction information, use the scaled corrected text to replace the video subtitle with typos in the video frame.

10. A video subtitle correction device based on bullet screens, comprising: An acquisition module, adapted to acquire bullet screen data in a video to be processed, wherein the bullet screen data includes bullet screen sending time and bullet screen vocabulary; A character recognition module, adapted to perform character recognition processing on video frames within a subtitle recognition interval of the video to be processed to obtain video subtitles in the video frames, wherein the time span of the subtitle recognition interval is a time window centered on the bullet screen sending time, extending forward by a first preset duration and backward by a second preset duration; A detection module, adapted to detect typos in the video subtitles according to the bullet screen vocabulary; A generation module, adapted to generate error subtitle correction information for correcting the video subtitles if typos are detected in the video subtitles.

11. A computing device, comprising: A processor, a memory, a communication interface, and a communication bus, through which the processor, the memory, and the communication interface complete communication with each other; The memory is used to store at least one executable instruction, and the executable instruction causes the processor to perform operations corresponding to the method for correcting video subtitles based on bullet screens as described in any one of claims 1-9.

12. A computer storage medium, in which at least one executable instruction is stored, and the executable instruction causes a processor to perform operations corresponding to the method for correcting video subtitles based on bullet screens as described in any one of claims 1-9.

13. A computer program product, including at least one executable instruction, and the executable instruction causes a processor to perform operations corresponding to the method for correcting video subtitles based on bullet screens as described in any one of claims 1-9.