Video subtitle error detection method, device and equipment and storage medium

By combining visual modal features of subtitle text and user lip movements or sign language images, this method detects typos in video subtitles, solving the problem of low detection accuracy in existing technologies and achieving more efficient typo detection.

CN115659957BActive Publication Date: 2026-05-05IFLYTEK CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
IFLYTEK CO LTD
Filing Date
2022-10-28
Publication Date
2026-05-05

AI Technical Summary

Technical Problem

Existing video subtitle error correction systems rely solely on plain text information, resulting in low accuracy in typo detection and wasting a significant amount of manpower and time.

Method used

By acquiring videos containing subtitles and user lip-sync and/or sign language images that match the subtitles, the system identifies the subtitle text and extracts the visual modal features of the lip-sync or sign language image sequences. These features are then fused with the text modal features to determine the real text and compare it to detect typos.

Benefits of technology

It significantly improves the accuracy of misspelling detection by fusing visual and textual modal features to assist in real text prediction, thereby enhancing the accuracy of misspelling detection results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115659957B_ABST
    Figure CN115659957B_ABST
Patent Text Reader

Abstract

This application discloses a method, apparatus, device, and storage medium for detecting typos in video subtitles. For videos containing user lip-sync and / or sign language images, it identifies the subtitle text, extracts lip-sync image sequences and / or sign language image sequences from the video, extracts the textual modal features of the subtitle text, extracts the lip-sync modal features of the lip-sync image sequences, and extracts the sign language modal features of the sign language image sequences. These lip-sync and / or sign language modal features are used as visual modal features, and the visual modal features and textual modal features are fused. Based on the fused features, the true text contained in the video is determined. This application, in addition to considering the textual modal features of the subtitle text, further fuses the visual modal features of lip-sync / sign language in the video, making the prediction results more accurate. Based on this, by comparing the true text and the subtitle text, the typo detection result is determined, greatly improving the accuracy of typo detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of natural language processing technology, and more specifically, to a method, apparatus, device, and storage medium for detecting typos in video subtitles. Background Technology

[0002] With the development of information technology and the continuous emergence of media platforms, an era characterized by diversified information transmission forms and multiple transmission sources has arrived. Current multimedia information includes examples such as self-media personalities delivering professional knowledge or disseminating social hot topics in front of a camera, and online video conferencing provided by various types of video conferencing software. These types of multimedia information generally include video of the user speaking, and corresponding subtitles are also provided to improve communication efficiency. Furthermore, to facilitate understanding of information for the hearing impaired, some multimedia videos, in addition to including subtitles, also feature sign language interpreters providing sign language explanations.

[0003] Due to carelessness by subtitle creators or the immaturity of related subtitle generation technology, numerous video subtitles on video platforms contain typos; typos are also frequently seen in subtitles generated in real-time by video conferencing software. This phenomenon poses a serious threat to the accuracy of information transmission and the breadth of cultural dissemination. Relying solely on manual proofreading and correction of these texts would consume a significant amount of manpower and time.

[0004] In today's era of booming artificial intelligence, especially thanks to advancements in natural language processing technology, various text error detection and correction systems have emerged to help people efficiently check and correct textual errors. Taking video subtitles as an example, existing error correction systems generally identify video subtitles, perform error correction processing based on the context of the subtitle text information, locate potential errors, and return the results to the user.

[0005] Existing error correction methods only utilize plain text information for error correction, resulting in low accuracy in typo detection. Summary of the Invention

[0006] In view of the above problems, this application is proposed to provide a method, apparatus, device, and storage medium for detecting typos in video subtitles, so as to improve the accuracy of typo detection in video subtitles. The specific solution is as follows:

[0007] Firstly, a method for detecting typos in video subtitles is provided, including:

[0008] Acquire a video containing subtitles and images of user lip movements and / or sign language that match the subtitles;

[0009] Identify subtitle text in the video, and extract the user's lip movement process in the video into a lip shape image sequence, and / or extract the user's sign language movement process in the video into a sign language image sequence;

[0010] Extract the text modal features of the subtitle text, and extract the lip modal features of the lip shape image sequence, and / or extract the sign language modal features of the sign language image sequence, using the lip modal features and / or the sign language modal features as visual modal features;

[0011] The visual modal features and the text modal features are fused to obtain the fused features;

[0012] Identify the real text contained in the video based on fusion features;

[0013] By comparing the actual text and the subtitle text, the typo detection results of the video subtitles are obtained.

[0014] Secondly, a video subtitle typo detection device is provided, comprising:

[0015] The video acquisition unit is used to acquire a video containing subtitles and images of user lip movements and / or sign language that match the subtitles;

[0016] A video preprocessing unit is used to identify subtitle text in the video, and to extract the user's lip movement process in the video into a lip shape image sequence, and / or to extract the user's sign language movement process in the video into a sign language image sequence.

[0017] The feature extraction unit is used to extract the text modal features of the subtitle text, and extract the lip modal features of the lip shape image sequence, and / or extract the sign language modal features of the sign language image sequence, wherein the lip modal features and / or the sign language modal features are used as visual modal features;

[0018] The feature fusion unit is used to fuse the visual modal features and the text modal features to obtain fused features;

[0019] The real text determination unit is used to determine the real text contained in the video based on fused features;

[0020] The misspelling detection unit is used to compare the real text and the subtitle text to obtain the misspelling detection results of the video subtitles.

[0021] Thirdly, a video subtitle typo detection device is provided, including: a memory and a processor;

[0022] The memory is used to store programs;

[0023] The processor is used to execute the program to implement the various steps of the video subtitle typo detection method described above.

[0024] Fourthly, a storage medium is provided on which a computer program is stored, which, when executed by a processor, implements the various steps of the video subtitle typo detection method described above.

[0025] Using the above technical solution, this application identifies the subtitle text in a video containing subtitles and user lip movements and / or sign language images matching the subtitles, extracts the user's lip movement process in the video into a lip image sequence, extracts the user's sign language movement process in the video into a sign language image sequence, and then extracts the text modal features of the subtitle text, as well as the lip modal features of the lip image sequence and the sign language modal features of the sign language image sequence. The lip modal features and / or the sign language modal features are used as visual modal features, and the visual modal features and text modal features are fused. Based on the fused features, the real text contained in the video is determined, and the real text and subtitle text are compared to obtain the typo detection result. Therefore, this application, when detecting typos in video subtitles, not only considers the textual modal features of the subtitle text but also further integrates the visual modal features of the video, such as sign language modal features and lip-shape modal features. These visual modal features can better assist in the prediction of real text, making the prediction results more accurate. Based on this, by comparing the real text and the subtitle text, the typo detection results are determined, greatly improving the accuracy of typo detection. Attached Figure Description

[0026] Various other advantages and benefits will become apparent to those skilled in the art upon reading the following detailed description of preferred embodiments. The accompanying drawings are for illustrative purposes only and are not intended to limit the scope of this application. Furthermore, the same reference numerals denote the same parts throughout the drawings. In the drawings:

[0027] Figure 1 This is a flowchart illustrating a video subtitle typo detection method provided in an embodiment of this application.

[0028] Figure 2 This example illustrates the process of marking typos in a video frame image.

[0029] Figure 3 An example is a schematic diagram of the structure of a video text recognition model;

[0030] Figure 4 An example is a schematic diagram of the structure of an image processing module;

[0031] Figure 5 This example illustrates the structure of a text processing module;

[0032] Figure 6 An example is a schematic diagram of the structure of a multimodal fusion module;

[0033] Figure 7 A schematic diagram illustrating the processing flow of a multimodal fusion module is provided.

[0034] Figure 8 This is a schematic diagram of a video subtitle typo detection device provided in an embodiment of this application;

[0035] Figure 9 This is a schematic diagram of the structure of the video subtitle typo detection device provided in the embodiments of this application. Detailed Implementation

[0036] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0037] This application provides a video subtitle typo detection scheme, which can be applied to the task of detecting typos in subtitles of videos containing lip movements or sign language information. Examples include detecting typos in subtitles of video conferences recorded by video conferencing apps, or detecting typos in subtitles of multimedia videos with sign language demonstrations.

[0038] The proposed solution can be implemented based on a terminal with data processing capabilities, such as a mobile phone, computer, server, or cloud platform.

[0039] Next, combined Figure 1 The video subtitle typo detection method of this application may include the following steps:

[0040] Step S100: Obtain a video containing subtitles and user lip movements and / or sign language images that match the subtitles.

[0041] Specifically, the video to be detected contains subtitle information, as well as lip-reading / sign language images of the user's speech process that match the subtitles.

[0042] The video can be user-recorded or downloaded from the internet. The position of the subtitle text in the video is not limited; for example, it can be arranged in rows or columns.

[0043] The subtitle text in the video can include Chinese characters and non-Chinese characters, such as English letters, special symbols, numbers, etc.

[0044] For example Figure 2 which is a frame image in the video and contains subtitle text. It can be known that the character "副" in the subtitle "说明在那副图上面" is a misspelling, and the correct one should be "幅".

[0045] Step S110: Identify the subtitle text in the video, and extract the lip shape / sign language information in the video into a lip shape / sign language image sequence.

[0046] Specifically, the video consists of multiple frame images, and the subtitle text in each frame image of the video can be identified. For example, first determine the text block picture where the text is located in the image. This process can use an image text recognition algorithm (such as an OCR algorithm, etc.) to identify the text block picture where the text is located in the image. Further, identify the subtitle text contained in the text block picture.

[0047] In addition, in order to assist in identifying the real text corresponding to the video, a lip shape / sign language image sequence can be extracted from the video in this step. Specifically, the lip movement process of the user in the video can be extracted into a lip shape image sequence, and / or the sign language movement process of the user in the video can be extracted into a sign language image sequence.

[0048] It should be noted that if the video only contains lip shape images, only the lip shape image sequence can be extracted. If the video only contains sign language images, only the sign language image sequence can be extracted. If the video contains both lip shape images and sign language images, either one or both of the lip shape image sequence and the sign language image sequence can be extracted.

[0049] Step S120: Extract the text modality features of the subtitle text, and extract the visual modality features of the lip shape / sign language image sequence. [[ID=2,1]]

[0050] Specifically, in this step, the text modality features of the subtitle text, that is, the text features, are extracted. When extracting the text modality features, a set text feature extraction algorithm can be used, or a pre-trained natural language processing model can be used.

[0051] The process of extracting the visual modality features of the lip shape / sign language image sequence can specifically be to extract the lip shape modality features of the lip shape image sequence, and / or extract the sign language modality features of the sign language image sequence, and use the lip shape modality features and / or the sign language modality features as the visual modality features.

[0052] In this step, a set image vision algorithm can be used to extract the visual modality features of the lip shape image sequence and the sign language image sequence, or a pre-trained neural network model can be used to extract the visual modality features of the lip shape image sequence and the sign language image sequence.

[0053] Step S130: Fuse the visual modal features and the text modal features to obtain fused features.

[0054] Specifically, visual modal features and textual modal features describe relevant information from both visual (lip-reading visual modality and sign language visual modality) and textual perspectives, respectively. In order to more accurately capture the real text contained in the video, this step fuses the visual modal features and textual modal features, resulting in richer information and stronger expressive power in the fused features.

[0055] Step S140: Determine the real text contained in the video based on the fusion features.

[0056] Specifically, after obtaining the fused features in the above steps, the real text contained in the video can be predicted based on the fused features. This step can use a pre-trained neural network model to predict the real text.

[0057] The actual text predicted through this step is the correct text that should be included in the video as determined by this application.

[0058] Step S150: Compare the real text and the subtitle text to obtain the typo detection result of the video subtitles.

[0059] Specifically, in this step, the real text can be used as a benchmark to compare the subtitle text with the real text, determine whether the subtitle text contains typos, and the specific content of the typos, and obtain the typo detection result of the video subtitles.

[0060] For example, in this step, it is possible to match whether there are characters in the subtitle text that are inconsistent with the real text. If so, the inconsistent characters in the subtitle text are treated as typos.

[0061] The video subtitle typo detection method provided by the embodiments of this application, for a video containing subtitles and user lip shapes and / or sign language images matching the subtitles, identifies the subtitle text therein, extracts the user's lip movement process in the video into a sequence of lip shape images, extracts the user's sign language movement process in the video into a sequence of sign language images, and then extracts the text modality features of the subtitle text, extracts the lip shape modality features of the sequence of lip shape images, and extracts the sign language modality features of the sequence of sign language images. Using the lip shape modality features and / or the sign language modality features as visual modality features, fuses the visual modality features and the text modality features, and determines the real text contained in the video based on the fused features. Compares the real text and the subtitle text to obtain the typo detection result. It can be seen that when this application detects typos in video subtitles, on the basis of considering the text modality features of the subtitle text, it further fuses the visual modality features in the video, such as sign language modality features and lip shape modality features. This visual modality feature can better assist in predicting the real text, making the prediction result more accurate. On this basis, by comparing the real text and the subtitle text, the typo detection result is determined, greatly improving the accuracy of typo detection.

[0062] Optionally, after obtaining the typo detection result in the above step S150, if it is confirmed that the video contains typos, the position of the typos in the video frame picture can be further determined, and then, according to the position, mark the typos in the video frame picture to visually display the typos.

[0063] Reference Figure 2 , for the typo "副" identified in the video frame picture, it is marked in the form of a rectangular box.

[0064] Of course, the marking form of typos is not limited to rectangular box marking, and other various types of marking methods can also be used, such as highlighting, underlining, etc.

[0065] In this embodiment, the process of determining the position of the typos in the video frame picture may specifically include:

[0066] First, determine the first position information of the text block picture where the typo is located, and this first position information is the position information of the text block picture in the video frame picture.

[0067] Further determine the sorting order of the typo in the subtitle text included in the text block picture.

[0068] Based on the first position information, determine the position of the first character in the text block picture in the video frame picture, and in the way of sliding offset according to the estimated width of each character, offset backward from the position of the first character by the width of the number of characters in the sorting order to locate the position of the typo in the video frame picture.

[0069] In some embodiments of this application, the process of fusing the visual modal features and the text modal features in step S130 to obtain fused features is described.

[0070] Optionally, the visual modal features and textual modal features extracted in step S120 can be in vector form. The vector dimensions of the visual modal features and textual modal features can be the same or different. Based on this, when performing feature fusion in this step, the two vector features can be fused to obtain fused features.

[0071] When performing vector fusion, various fusion methods can be used. In this embodiment, a gated fusion method is provided to fuse visual modal and text modal features in vector form to obtain fused features.

[0072] By employing a gated fusion approach, visual modal features are used as the gate to extract some features from text modal features, resulting in fused features. In other words, from the perspective of visual modal features, the most important part of text modal features is extracted as the feature representation of the fusion of visual and character modal features.

[0073] Optionally, this application provides several different gating fusion methods, such as: bitwise multiplication gating fusion, bitwise addition or division gating fusion, etc. For ease of description, the following embodiments only use bitwise multiplication gating fusion as an example.

[0074] Furthermore, to avoid the loss of global features at the text language level, in this embodiment, the above-mentioned fusion features can be added to the text modal features to obtain residual fusion features, which are used as the final fusion features.

[0075] To enhance the richness of visual modal feature representation, before feature fusion in step S130, representation shift and nonlinear transformation processing can be added to the visual modal features to obtain processed visual modal features, which can then be fused with text modal features in step S130.

[0076] In some embodiments of this application, steps S120-S140 described in the foregoing embodiments can be obtained by processing a pre-trained video text recognition model.

[0077] For a video text recognition model, it can be configured as follows: extract lip shape modal features from an input lip shape image sequence, and / or extract sign language modal features from an input sign language image sequence, use the lip shape modal features and / or the sign language modal features as visual modal features, extract text modal features from the input subtitle text, fuse the visual modal features and text modal features, and predict the internal state representation of the real text contained in the video based on the fused features.

[0078] The input to the video text recognition model may include subtitle text extracted from the video, as well as lip-shape image sequences and / or sign language image sequences.

[0079] In this embodiment, by pre-training a video text recognition model, the powerful learning ability of the neural network model can be utilized to extract the visual modal features of lip-shape image sequences and / or sign language image sequences and the text modal features of subtitle text. Based on this, the real text is predicted after fusion.

[0080] Next, combined Figure 3 As shown, this embodiment provides an optional component structure for a video text recognition model.

[0081] A video text recognition model may include an image processing module, a text processing module, a multimodal fusion module, and an output module. Among them:

[0082] An image processing module is used to extract lip shape modal features from an input lip shape image sequence, and / or to extract sign language modal features from an input sign language image sequence, using the lip shape modal features and / or the sign language modal features as visual modal features.

[0083] If both lip-shape modality features and sign language modality features are extracted simultaneously, then the dimensions of the two modality features are the same, and the two modality features are combined to form the visual modality features. Of course, if only lip-shape modality features or sign language modality features can be extracted, then the feature of one of the extracted modalities can be used alone as the visual modality feature.

[0084] The text processing module is used to extract the text modal features of the input subtitle text.

[0085] The multimodal fusion module is used to fuse the visual modal features and the text modal features to obtain fused features.

[0086] The output module is used to determine the real text contained in the video based on the fusion features.

[0087] The output module can be trained using the MLM (Masked Language Model) method. Based on the fusion features output by the multimodal fusion module, it predicts the real text contained in the video.

[0088] Next, each of the above modules will be explained in detail.

[0089] 1. Image processing module

[0090] This embodiment describes an optional component structure of the image processing module, such as... Figure 4 As shown, it may include:

[0091] The image standardization module is used to standardize the input lip shape image sequence and / or sign language image sequence to obtain the processed lip shape image sequence and / or sign language image sequence.

[0092] The input to the image normalization module can be one or both of lip-shape image sequences and sign language image sequences.

[0093] Because videos may contain rich information beyond the speaker's lip and hand gestures, such as surrounding objects, the input lip and sign language image sequences also contain other interfering information. Directly extracting features from these image sequences makes it difficult to learn features specific to the current subtitle's lip / sign language modality. Therefore, to better adapt to subsequent modules for feature extraction and ensure the quality of visual modality feature extraction, this step uses an image standardization module to standardize the lip and sign language image sequences. This can be achieved by processing distorted text block images through algorithms such as image rotation, stretching, and scaling. The processed lip and sign language image sequences contain images of a set size, such as a matrix of size [96, 384].

[0094] The image feature extraction module is used to extract visual modal features from the processed lip shape image sequence and / or sign language image sequence.

[0095] like Figure 4 For example, the image feature extraction module can be composed of several visual feature recognition blocks connected in series. Each visual feature recognition block can include several convolutional layers, batch normalization layers, and nonlinear layers. The size and number of convolutional kernels in the convolutional layers within different visual feature recognition blocks can vary to enrich the perspectives of visual modality feature extraction, thereby resulting in a richer and more accurate final visual modality feature representation.

[0096] The linear transformation module is used to perform a linear transformation on the dimensions of the visual modal features to output visual modal features with the same dimensions as the text modal features.

[0097] Specifically, the number of channels of the visual modal features extracted by the image feature extraction module may not be directly matched with the dimension of the text modal features extracted by the text processing module. Therefore, it is necessary to perform a linear transformation on the dimension of the visual modal features through the linear transformation module to output visual modal features with the same dimension as the text modal features.

[0098] 2. Text Processing Module

[0099] This embodiment describes an optional structural configuration for the text processing module, such as... Figure 5 As shown, it may include:

[0100] The text preprocessing module is used to edit the input subtitle text to a set length by padding with specified characters, and to determine the feature representation of the edited subtitle text.

[0101] Specifically, to standardize the length of different subtitle texts, this embodiment uses a text preprocessing module to edit the subtitle text to a set length using padding. For subtitle text shorter than the set length, a set padding character, such as [PAD], can be added to the end of the subtitle text to supplement it to the set length. For subtitle text longer than the set length, the set length can be truncated from the first character as one edited subtitle text. If the remaining length still exceeds the set length, the truncating operation is repeated. If the remaining length does not exceed the set length, the remaining part is used as another edited subtitle text.

[0102] For each edited subtitle text, a pre-trained tokenizer can be used to encode the subtitle text into a feature representation that the model can recognize. Specifically, the edited subtitle text is segmented into words, and each word is encoded to obtain the token feature representation corresponding to the word.

[0103] Among them, the pre-trained tokenizer can adopt pre-trained model structures such as BERT tokenizer.

[0104] The text modal feature extraction module is used to encode the feature representation of the subtitle text to obtain the text modal features of the subtitle text.

[0105] Specifically, the text modality feature extraction module can use a pre-trained model (such as BERT, Transformer, etc.) to encode the feature representation of the subtitle text after it has been processed by the text preprocessing module, so as to obtain the text modality features of the subtitle text.

[0106] 3. Multimodal fusion module

[0107] This embodiment describes an optional structural composition of the multimodal fusion module, such as... Figure 6 As shown, it may include: a feature editing module, a gating fusion module, and a residual connection module.

[0108] The processing flow of each module is combined Figure 7 Explanation:

[0109] The feature editing module is used to perform representation shifts and nonlinear transformations on visual modal features to obtain processed visual modal features.

[0110] Visual modal features may include one of lip-shape modal features and sign language modal features, or a combination of the two.

[0111] To enhance the representation of visual modal features, representation shift and nonlinear transformation can be applied. Representation shift involves adding a learnable bias parameter to each position of the visual modal feature. Nonlinear transformation uses nonlinear function layers, such as ReLU, sigmoid, or tanh layers, to nonlinearly transform the shifted visual modal features to a relatively small range near 0. For example, the range of the sigmoid transformation is (0, 1), and the range of the tanh transformation is (-1, 1).

[0112] The gated fusion module is used to fuse the processed visual modal features and the text modal features using a gated fusion method to obtain fused features.

[0113] Specifically, in this embodiment, a gating fusion module is designed to perform bitwise multiplication, bitwise addition, or bitwise division to fuse the processed visual modal features and text modal features to obtain fused features.

[0114] Figure 7 Taking the bitwise multiplication gating fusion method as an example, by using the bitwise multiplication gating fusion method, visual modal features are used as the gate to extract some features from the text modal features to obtain the fused features. That is, from the perspective of visual modal features, the most important part of the text modal features is extracted as the feature representation of visual modality and character modality fusion.

[0115] After the visual modal features are processed by the feature editing module, they have additional representation offsets and nonlinear transformations compared to text modal features. This maps the visual modal features to a relatively small range near 0, such as the sigmoid function's range being (0,1). The range and distribution of text modal features remain unchanged. To put it figuratively, each position in the edited visual modal features is like a faucet (fully open corresponds to the upper bound of the nonlinear function's range, fully closed corresponds to the lower bound), controlling the information at the corresponding position in the text modal features. The more open the faucet is at that position in the visual modal features, the more information is retained in the corresponding position in the text modal features, and vice versa. Clearly, this positional multiplication yields the text modal feature portion whose retention is controlled by the visual perspective. In other words, from the perspective of the visual modal features, the most important part of the text modal features is extracted as the feature representation for the fusion of visual and character modal features.

[0116] The residual connection module is used to add the fused feature to the text modal feature to obtain the residual fused feature, which is used as the final fused feature.

[0117] Furthermore, to avoid the loss of global features at the text language level, this embodiment can also add the above-mentioned fused features to the text modal features through the residual connection module to obtain residual fused features, which are used as the final fused features.

[0118] In some embodiments of this application, in order to further improve the accuracy of typo detection, after step S150, comparing the real text and the subtitle text to obtain the typo detection results in the video, a post-processing operation for typo verification can be further added.

[0119] In this embodiment, the post-processing for misspelling verification can be performed from the perspective of sentence semantic fluency, specifically including:

[0120] S1. Delete the typos identified in the subtitle text to obtain the edited text with the typos removed.

[0121] S2. Using a pre-trained language model, calculate the perplexity of the subtitle text and the edited text with typos removed.

[0122] Specifically, perplexity is an indicator that measures the semantic fluency of a sentence; the more fluent a sentence is semantically, the lower its perplexity.

[0123] A language model is a probabilistic model used to calculate the probability that a sentence is a semantically correct sentence. Perplexity is a sentence-length-normalized metric related to the probability that a language model predicts a sentence. For a perfectly correct sentence, the lower the perplexity of the language model, the better the language model. Conversely, if a very good language model has been selected, then for a given sentence, if the perplexity of the language model is very low, it means that the sentence is highly likely to be correct.

[0124] In this step, to verify whether the previously identified typos are indeed typos, the confusion level of the subtitle text and the edited text after deleting the typos are calculated separately.

[0125] S3. If the confusion level of the edited text after deleting the typo is less than that of the subtitle text, and the absolute value of the difference between the two is greater than a set threshold, then the typo is taken as the final typo detection result; otherwise, the typo is removed from the final typo detection result.

[0126] Understandably, if the perplexity of the edited text after deleting the typo is less than that of the subtitle text, and the absolute value of the difference between the two is greater than a set threshold, it indicates that the semantics of the edited text after deleting the typo are more coherent than that of the subtitle text before deletion. In other words, the deleted typo was indeed a typo, and therefore it can be added to the final typo detection result. Conversely, if the perplexity is less than a set threshold, it indicates that the typo identified in the previous steps is a pseudo-typo, and it can be removed from the final typo detection result, meaning it will not be ultimately identified as a typo.

[0127] In this embodiment, the accuracy of misspelling recognition is further improved by adding a post-processing operation that performs secondary verification of misspellings from the perspective of sentence semantic fluency.

[0128] The following describes the video subtitle typo detection device provided in the embodiments of this application. The video subtitle typo detection device described below can be referred to in correspondence with the video subtitle typo detection method described above.

[0129] See Figure 8 , Figure 8 This is a schematic diagram of the structure of a video subtitle typo detection device disclosed in an embodiment of this application.

[0130] like Figure 8 As shown, the device may include:

[0131] Video acquisition unit 11 is used to acquire a video containing subtitles and images of user lip movements and / or sign language that match the subtitles;

[0132] The video preprocessing unit 12 is used to identify subtitle text in the video, and to extract the user's lip movement process in the video into a lip shape image sequence, and / or to extract the user's sign language movement process in the video into a sign language image sequence.

[0133] The feature extraction unit 13 is used to extract the text modal features of the subtitle text, and extract the lip modal features of the lip image sequence, and / or extract the sign language modal features of the sign language image sequence, using the lip modal features and / or the sign language modal features as visual modal features;

[0134] The feature fusion unit 14 is used to fuse the visual modal features and the text modal features to obtain fused features;

[0135] The real text determination unit 15 is used to determine the real text contained in the video based on fusion features;

[0136] The misspelling detection unit 16 is used to compare the real text and the subtitle text to obtain the misspelling detection result of the video subtitles.

[0137] Optionally, if the visual modal features and the text modal features are both in vector form, then the process by which the feature fusion unit fuses the visual modal features and the text modal features to obtain the fused features may include:

[0138] A gated fusion method is used to fuse visual modal features and text modal features in vector form to obtain fused features.

[0139] Optionally, the above gating fusion methods may include gating fusion methods of bitwise multiplication, bitwise addition or division, etc.

[0140] Optionally, after fusing the vector-based visual modal features and text modal features using a gated fusion method, the aforementioned feature fusion unit may further include:

[0141] The fused features are added to the text modal features to obtain residual fused features, which are used as the final fused features.

[0142] Optionally, before fusing the vector-based visual modal features and text modal features using a gated fusion method, the aforementioned feature fusion unit may further include:

[0143] The visual modal features are subjected to representation shift and nonlinear transformation to obtain the processed visual modal features.

[0144] Optionally, the processing of the above-mentioned feature extraction unit 13, feature fusion unit 14, and real text determination unit 15 can be implemented by a pre-trained video text recognition model. The video text recognition model is configured to extract lip shape modal features from the input lip shape image sequence, and / or extract sign language modal features from the input sign language image sequence, use the lip shape modal features and / or the sign language modal features as visual modal features, extract text modal features from the input subtitle text, fuse the visual modal features and text modal features, and predict the internal state representation of the real text contained in the video based on the fused features.

[0145] The video text recognition model may include: an image processing module, a text processing module, a multimodal fusion module, and an output module;

[0146] An image processing module is used to extract lip shape modal features from an input lip shape image sequence, and / or to extract sign language modal features from an input sign language image sequence, using the lip shape modal features and / or the sign language modal features as visual modal features;

[0147] The text processing module is used to extract the text modal features of the input subtitle text;

[0148] A multimodal fusion module is used to fuse the visual modal features and the text modal features to obtain fused features;

[0149] The output module is used to determine the real text contained in the video based on the fusion features.

[0150] Optionally, the above-mentioned multimodal fusion module may further include:

[0151] The feature editing module is used to perform representation shift and nonlinear transformation on the visual modal features to obtain the processed visual modal features;

[0152] The gated fusion module is used to fuse the processed visual modal features and the text modal features using a gated fusion method to obtain fused features;

[0153] The residual connection module is used to add the fused feature to the text modal feature to obtain the residual fused feature, which is used as the final fused feature.

[0154] Optionally, the image processing module described above may further include:

[0155] The image standardization module is used to standardize the input lip shape image sequence and / or sign language image sequence to obtain the processed lip shape image sequence and / or sign language image sequence.

[0156] The image feature extraction module is used to extract visual modal features from the processed lip shape image sequence and / or sign language image sequence;

[0157] The linear transformation module is used to perform a linear transformation on the dimensions of the visual modal features to output visual modal features with the same dimensions as the text modal features.

[0158] Optionally, the above text processing module may further include:

[0159] The text preprocessing module is used to edit the input subtitle text to a set length by filling in set characters, and to determine the feature representation of the edited subtitle text;

[0160] The text modal feature extraction module is used to encode the feature representation of the subtitle text to obtain the text modal features of the subtitle text.

[0161] Optionally, the process by which the above-mentioned misspelling detection unit compares the real text and the subtitle text to obtain the misspelling detection result of the video subtitles may include:

[0162] The system checks if any characters in the subtitle text are inconsistent with the actual text. If such characters are found, they are considered typos in the video subtitles.

[0163] Optionally, the apparatus of this application may further include: a typo verification unit, configured to: after comparing the real text and the subtitle text to obtain a typo detection result for the video subtitles, delete the typos identified in the subtitle text to obtain an edited text with the typos removed; use a pre-trained language model to calculate the perplexity of the subtitle text and the edited text with the typos removed; if the perplexity of the edited text with the typos removed is less than the perplexity of the subtitle text, and the absolute value of the difference between the two is greater than a set threshold, then the typo is taken as the final typo detection result; otherwise, the typo is removed from the final typo detection result.

[0164] Optionally, the apparatus of this application may further include: a misspelling marking unit, configured to: after comparing the real text and the subtitle text to obtain a misspelling detection result for the video subtitles, determine the position of the misspelling in the video frame image; and mark the misspelling in the video frame image according to the position.

[0165] The video subtitle typo detection device provided in this application embodiment can be applied to video subtitle typo detection equipment, such as terminals: mobile phones, computers, etc. Optionally, Figure 9 The hardware structure block diagram of the video subtitle typo detection device is shown below. Figure 9The hardware structure of a video subtitle typo detection device may include: at least one processor 1, at least one communication interface 2, at least one memory 3, and at least one communication bus 4;

[0166] In this embodiment of the application, the number of processor 1, communication interface 2, memory 3, and communication bus 4 is at least one, and processor 1, communication interface 2, and memory 3 communicate with each other through communication bus 4;

[0167] Processor 1 may be a central processing unit (CPU), an application-specific integrated circuit (ASIC), or one or more integrated circuits configured to implement embodiments of the present invention.

[0168] Memory 3 may include high-speed RAM, and may also include non-volatile memory, such as at least one disk storage device;

[0169] The memory stores a program, which the processor can call. The program is used for:

[0170] Acquire a video containing subtitles and images of user lip movements and / or sign language that match the subtitles;

[0171] Identify subtitle text in the video, and extract the user's lip movement process in the video into a lip shape image sequence, and / or extract the user's sign language movement process in the video into a sign language image sequence;

[0172] Extract the text modal features of the subtitle text, and extract the lip modal features of the lip shape image sequence, and / or extract the sign language modal features of the sign language image sequence, using the lip modal features and / or the sign language modal features as visual modal features;

[0173] The visual modal features and the text modal features are fused to obtain the fused features;

[0174] Identify the real text contained in the video based on fusion features;

[0175] By comparing the actual text and the subtitle text, the typo detection results of the video subtitles are obtained.

[0176] Optionally, the refined and extended functions of the program can be found in the description above.

[0177] This application embodiment also provides a storage medium that can store a program suitable for execution by a processor, the program being used for:

[0178] Acquire a video containing subtitles and images of user lip movements and / or sign language that match the subtitles;

[0179] Identify subtitle text in the video, and extract the user's lip movement process in the video into a lip shape image sequence, and / or extract the user's sign language movement process in the video into a sign language image sequence;

[0180] Extract the text modal features of the subtitle text, and extract the lip modal features of the lip shape image sequence, and / or extract the sign language modal features of the sign language image sequence, using the lip modal features and / or the sign language modal features as visual modal features;

[0181] The visual modal features and the text modal features are fused to obtain the fused features;

[0182] Identify the real text contained in the video based on fusion features;

[0183] By comparing the actual text and the subtitle text, the typo detection results of the video subtitles are obtained.

[0184] Optionally, the refined and extended functions of the program can be found in the description above.

[0185] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0186] The various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. The various embodiments can be combined as needed, and the same or similar parts can be referred to each other.

[0187] The above description of the disclosed embodiments enables those skilled in the art to make or use this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for detecting typos in video subtitles, characterized in that, include: Acquire a video containing subtitles and images of user lip movements and / or sign language that match the subtitles; Identify subtitle text in the video, and extract the user's lip movement process in the video into a lip shape image sequence, and / or extract the user's sign language movement process in the video into a sign language image sequence; Extract the text modal features of the subtitle text, and extract the lip modal features of the lip shape image sequence, and / or extract the sign language modal features of the sign language image sequence, using the lip modal features and / or the sign language modal features as visual modal features; The visual modal features and the text modal features are fused to obtain the fused features; Identify the real text contained in the video based on fusion features; By comparing the actual text and the subtitle text, the typo detection results of the video subtitles are obtained.

2. The method according to claim 1, characterized in that, The visual modal features and the text modal features are both in vector form; The process of fusing the visual modal features and the text modal features to obtain the fused features includes: A gated fusion method is used to fuse visual modal features and text modal features in vector form to obtain fused features.

3. The method according to claim 2, characterized in that, After fusing the vector-based visual modal features and text modal features using a gated fusion method to obtain the fused features, the process also includes: The fused features are added to the text modal features to obtain residual fused features, which are used as the final fused features.

4. The method according to claim 2, characterized in that, Before fusing the vector-based visual modal features and text modal features using a gated fusion method, the following steps are also included: The visual modal features are subjected to representation shift and nonlinear transformation to obtain the processed visual modal features.

5. The method according to claim 1, characterized in that, The process of extracting the visual modal features and text modal features and fusing them, and determining the real text contained in the video based on the fused features, is obtained by processing a pre-trained video text recognition model. The video text recognition model is configured to extract lip shape modal features from an input lip shape image sequence, and / or extract sign language modal features from an input sign language image sequence, use the lip shape modal features and / or the sign language modal features as visual modal features, extract text modal features from the input subtitle text, fuse the visual modal features and text modal features, and predict the internal state representation of the real text contained in the video based on the fused features.

6. The method according to claim 5, characterized in that, The video text recognition model includes: an image processing module, a text processing module, a multimodal fusion module, and an output module; The image processing module is used to extract lip shape modal features from the input lip shape image sequence, and / or to extract sign language modal features from the input sign language image sequence, using the lip shape modal features and / or the sign language modal features as visual modal features; The text processing module is used to extract the text modal features of the input subtitle text; A multimodal fusion module is used to fuse the visual modal features and the text modal features to obtain fused features; The output module is used to determine the real text contained in the video based on the fusion features.

7. The method according to claim 6, characterized in that, The multimodal fusion module includes: The feature editing module is used to perform representation shift and nonlinear transformation on the visual modal features to obtain the processed visual modal features; The gated fusion module is used to fuse the processed visual modal features and the text modal features using a gated fusion method to obtain fused features; The residual connection module is used to add the fused feature to the text modal feature to obtain the residual fused feature, which is used as the final fused feature.

8. The method according to claim 6, characterized in that, The image processing module includes: The image standardization module is used to standardize the input lip shape image sequence and / or sign language image sequence to obtain the processed lip shape image sequence and / or sign language image sequence. The image feature extraction module is used to extract visual modal features from the processed lip shape image sequence and / or sign language image sequence; The linear transformation module is used to perform a linear transformation on the dimensions of the visual modal features to output visual modal features with the same dimensions as the text modal features.

9. The method according to claim 6, characterized in that, The text processing module includes: The text preprocessing module is used to edit the input subtitle text to a set length by filling in set characters, and to determine the feature representation of the edited subtitle text; The text modal feature extraction module is used to encode the feature representation of the subtitle text to obtain the text modal features of the subtitle text.

10. The method according to any one of claims 1-9, characterized in that, The process of comparing the real text and the subtitle text to obtain the typo detection results for the video subtitles includes: The system checks if any characters in the subtitle text are inconsistent with the actual text. If such characters are found, they are considered typos in the video subtitles.

11. The method according to any one of claims 1-9, characterized in that, After comparing the real text and the subtitle text to obtain the typo detection results for the video subtitles, the method further includes: The identified typos in the subtitle text are deleted to obtain the edited text with the typos removed; Using a pre-trained language model, the perplexity of the subtitle text and the edited text with typos removed is calculated respectively; If the perplexity of the edited text after deleting the typo is less than that of the subtitle text, and the absolute value of the difference between the two is greater than a set threshold, then the typo is taken as the final typo detection result; otherwise, the typo is removed from the final typo detection result.

12. The method according to any one of claims 1-9, characterized in that, After comparing the real text and the subtitle text to obtain the typo detection results for the video subtitles, the method further includes: Determine the location of the misspelled word in the video frame image; The misspelled words are marked in the video frame image according to the stated position.

13. A video subtitle typo detection device, characterized in that, include: The video acquisition unit is used to acquire a video containing subtitles and images of user lip movements and / or sign language that match the subtitles; A video preprocessing unit is used to identify subtitle text in the video, and to extract the user's lip movement process in the video into a lip shape image sequence, and / or to extract the user's sign language movement process in the video into a sign language image sequence. The feature extraction unit is used to extract the text modal features of the subtitle text, and extract the lip modal features of the lip shape image sequence, and / or extract the sign language modal features of the sign language image sequence, wherein the lip modal features and / or the sign language modal features are used as visual modal features; The feature fusion unit is used to fuse the visual modal features and the text modal features to obtain fused features; The real text determination unit is used to determine the real text contained in the video based on fused features; The misspelling detection unit is used to compare the real text and the subtitle text to obtain the misspelling detection results of the video subtitles.

14. A video subtitle typo detection device, characterized in that, include: Memory and processor; The memory is used to store programs; The processor is used to execute the program to implement each step of the video subtitle typo detection method as described in any one of claims 1 to 12.

15. A storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements each step of the video subtitle typo detection method as described in any one of claims 1 to 12.

Citation Information

Patent Citations

  • Video subtitle identification method and system

    CN106529529A

  • Sign language recognition method and system based on double-flow space-time diagram convolutional neural network

    CN111325099A