Providing subtitles for video content

The system automates subtitle generation using machine-trained models to process audio data and identify breaks, addressing the labor-intensive nature of manual subtitle creation and enhancing audience reach and monetization.

JP2025515560AActive Publication Date: 2025-05-20VOYAGERX INC
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2024557955
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-03-24
Filing Date
2023-04-26
Publication Date
2025-05-20
Estimated Expiration
2043-04-26

Smart Images

  • Figure 2025515560000001_ABST
    Figure 2025515560000001_ABST
Patent Text Reader

Abstract

The present disclosure relates to a system and method for providing subtitles for a video. Audio of a video is transcribed to obtain caption text for the video. A first machine-trained model identifies sentences in the caption text. A second model identifies sentence breaks in the sentences identified using the first machine-trained model. Based on the identified sentences and sentence breaks, one or more words in the caption text are grouped into clip captions that are displayed for a corresponding clip of the video.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] This application relates to a system and method for providing subtitles to video. Summary of the Invention

[0002] One aspect of the disclosure provides a method for providing subtitles for a video, the method including: processing audio data of the video to generate a timed script in a first language including a first sequence of words and a timestamp for each word in the first sequence of words; processing the first sequence of words to calculate a sentence-final probability for each word in the first sequence of words using a first machine-trained model; determining a first word of the first sequence as a first sentence-final word based on the sentence-final probability of the first word to define a first sentence ending with the first word; processing the first sentence to calculate a sentence-final probability for at least one word in the first sentence using a second machine-trained model; The method includes one or more of the following steps: calculating an intra-sentence break probability; determining a second word of the first sentence as a clip end word based on the intra-sentence break probability of the second word and defining a first clip text that ends with the second word, wherein the definition of the first clip text further defines a first clip period that corresponds to the first clip text and ends when the second word is spoken in the video; and generating subtitle data in the first language including the first clip text and information indicating a first clip period in which the first clip text will be displayed as subtitles in the first language.

[0003] In an embodiment, the method further includes determining the third word of the first sentence as another clip end word based on the intra-sentence break probability of the third word, thereby defining a second clip text starting with a word immediately following the second word and ending with the third word, the definition of the second clip text further defining a second clip corresponding to the second clip text and ending at the point when the third word is spoken in the video.

[0004] In an embodiment, the timed script does not include punctuation indicating the end of the first sentence or an intra-sentence break in the first sentence.

[0005] In an embodiment, the first machine-trained model is trained using a plurality of punctuated texts, each of which includes one or more final punctuation marks, and is configured to calculate, for at least one word in the input text, a probability that it is immediately followed by at least one final punctuation mark.

[0006] In an embodiment, the second machine-trained model is trained using a plurality of punctuated texts, each of which includes one or more sentence-breaking punctuation marks, and is configured to calculate a probability that, for at least one word in the input sentence, it is immediately followed by at least one sentence-breaking punctuation mark.

[0007] In an embodiment, the at least one end-of-sentence punctuation mark comprises one of a period, a question mark, an exclamation mark, and an ellipsis, and the at least one intra-sentence delimiting punctuation mark comprises one of a comma, a colon, a semicolon, and an ellipsis.

[0008] In the method, processing audio data of the video to generate a timed script may include performing a speech-to-text translation (STT) process of the audio data, in which audio corresponding to a second word is transcribed into the second word, a time in the video at which the second word is spoken is determined, and the time at which the second word is spoken is specified in the timed script for the second word. The information indicative of the first clip duration may include the time at which the second word was spoken determined by the STT process. Generating the first language subtitle data may include associating the time at which the second word was spoken, determined by the STT process, with the first clip text as an end of the first clip duration in accordance with a predefined subtitle file format.

[0009] In the method, processing the audio data of the video to generate a timed script may include one or more of the following steps: identifying silence and non-silence in the audio data, where the non-silence sounds include corresponding sounds of a second word; transcribing the corresponding sounds of the second word into the second word to obtain a first word sequence; determining, for the second word, an end time at which the corresponding sound of the second word ends in the video; and including the determined end time in the timed script as a timestamp for the second word.

[0010] In the method, processing the audio data of the video to generate a timed script may include one or more of the following steps: obtaining a pre-written script of the video that includes a first word sequence but does not include a timestamp of the first word sequence; for each word of the first word sequence, identifying a corresponding sound in the audio data that identifies a first sound that corresponds to a second word; determining an end time of the first sound when the first sound ends in the video; and combining the first word sequence with the determined end time to generate the timed script such that the determined end time is specified as a timestamp of the second word.

[0011] In the method, the timed script may include a second word timestamp indicating a time the second word is spoken in the video, and generating the first language caption data may include designating the second word timestamp as an end of the first clip period according to a predefined subtitle format. In an embodiment, the first clip text begins with a third word of the first sentence, the timed script includes a third word timestamp indicating a time when speech of the third word begins in the video, and generating the first language caption data includes the third word timestamp as a start of the first clip period according to a predefined subtitle format.

[0012] In an embodiment, the first language subtitle data is configured such that the entire first clip text is displayed as subtitles to the video at the start of the first clip period and remains uninterrupted until the end of the first clip period.

[0013] In an embodiment, the first clip text further includes a fourth word between the third word and the second word, and the first language subtitle data does not include a timestamp of the fourth word, such that the first clip text is displayed as subtitles without reference to the fourth word.

[0014] In the method, the information indicating the first clip duration may include a first timestamp indicating a start time of the first clip in the video and may further include a second timestamp indicating an end time of the first clip in the video, and the first clip text may be displayed together with the video without interruption from the start time of the first clip to the end time of the first clip.

[0015] In this method, the timestamp of each word may define the time when the sound of the word ends in the video. The timestamp of each word defines the time when the sound of the word begins in the video.

[0016] In an embodiment, the method may further include one or more of the following steps: translating the first sentence into a first translation sentence in a second language, the first translation sentence ending with a first translation word; processing the first translation sentence to calculate a sentence break probability of at least one word of the first translated sentence using a third machine-trained model; determining a second translation word of the first translated sentence as a clip end word based on the sentence break probability of the second translated word, defining a first translated clip text ending with the second translated word; and generating second language subtitle data including the first translated clip text and information indicating a second language period in which the first translated clip text is to be displayed as a second language subtitle. In an embodiment, the first clip duration for displaying the first clip text is the same or substantially the same as the second language duration for displaying the first translated clip text, regardless of whether the second word that ends the first clip text semantically corresponds to the second translation word that ends the first translated clip text.

[0017] In embodiments, the first translated clip text in the second language may not correspond in meaning to the first clip text in the first language. The first translated clip text may be a translation of the first clip text.

[0018] In an embodiment, generating the second language subtitle data includes designating the time when the second word is spoken in the first language as the end of the second language period such that the first clip period and the second language period end simultaneously.

[0019] In an embodiment, the timed script includes a timestamp of the second word indicating a time the second word was spoken in the video, and generating the second language subtitle data includes designating, in the second language subtitle data, the timestamp of the second word as an end of the first clip period and designating a second language period according to a predefined subtitle format.

[0020] In an embodiment, the first clip text begins with a third word of the first sentence, the first translated clip text begins with a third translation word of the first translation sentence, the timed script includes a timestamp of the third word indicating a time when the sound of the third word begins in the video, and generating the second language subtitle data further includes designating, in the second language subtitle data, the timestamp of the third word as the start of a first clip period and a second language period, wherein the first clip period and the second language period are identical regardless of whether the third word semantically corresponds to the third translation word.

[0021] In an embodiment, the first translated sentence does not include punctuation marks indicating intra-sentence breaks in the first translated sentence, and the third machine-trained model is trained using a plurality of punctuated sentences in the second language and is configured to calculate a probability that, for at least one word in the input sentence, it is immediately followed by at least one intra-sentence break punctuation mark.

[0022] In an embodiment, where "n" is a natural number greater than "2", the first sentence is divided into "n" clip texts based on at least the first word and the second word, and the first translated sentence is divided into the same "n" number of translated clip texts.

[0023] In an embodiment, the method further includes one of determining a third word of the first sentence as a clip-ending word, where a second clip text is defined beginning with a word immediately following the second word and ending with the third word. In an embodiment, defining the second clip text further defines a second clip corresponding to the second clip text and ending when the third word is spoken in the video, where the third word of the first sentence is identified as a clip-ending word based on at least one of an intra-sentence break probability of the third word and a length of silence following the sound of the third word in the video.

[0024] In an embodiment, the first language subtitle data is configured such that the entire first clip text is displayed as a subtitle on the video at the start of the first clip period and maintained uninterrupted until the end of the first clip period. In an embodiment, the second language subtitle data is configured such that the entire first translated clip text is displayed as a subtitle on the video at the start of the second language period and maintained uninterrupted until the end of the second language period. [Brief description of the drawings]

[0025] [Figure 1] FIG. 2 is a flow chart diagram of one embodiment of providing subtitles to a video. [Diagram 2] 2 presents an embodiment of a method for translating subtitles into different languages. [Diagram 3] It presents a platform user interface where subtitles can be combined with video clips and edited together. [Figure 4] 1 is a diagrammatic representation of the association between a text file and a video when a sentence model is run on the video and its text file(s). [Diagram 5] 1 is a diagrammatic representation of an embodiment showing an association between a text file and a video when an AI model is run on the video and its text file(s) to generate a clip. [Figure 6] 13 is a diagrammatic representation of an embodiment showing associations between text files in different languages ​​and a video when an AI model is run on the video and its text files to translate them into another language. [Figure 7A] An embodiment of a method for creating subtitles in one or more languages ​​for a video clip is presented, detecting sentence ends, intra-sentence breaks, and optional translation into another language. [Figure 7B] An embodiment of a method for creating subtitles in one or more languages ​​for a video clip is presented, detecting sentence ends, intra-sentence breaks, and optional translation into another language. [Figure 8]1 is a diagrammatic representation of an exemplary machine in the form of a computer system that can be used to perform any of the methods disclosed herein. [Figure 9A] 1 shows an example of transcribed text resulting from speech-to-text translation (STT) processing of a video. [Figure 9B] FIG. 9B shows an exemplary timed script with time codes added to the transcribed text of FIG. 9A. [Figure 10] FIG. 9B is an example of calculating the sentence-final probability of words in the transcribed text of FIG. 9A. [Figure 11] 9A shows example sentences identified from the transcribed text. [Figure 12A] We show identifying clip end words in the example sentences of FIG. 11 based on intra-sentence segmentation probabilities. [Figure 12B] This is an example in which the example sentence in FIG. 11 is divided into two clip texts. [Figure 13] 12C is an example of subtitle data generated from the timed script of FIG. 9B using the two clip texts of FIG. 12B. [Figure 14] A sentence-by-sentence translation of the sentences in FIG. [Figure 15A] FIG. 14 illustrates identifying clip-ending words in the translated sentence. [Figure 15B] Indicates splitting the translation into two translated clip texts. [Figure 16] An example of generating translated subtitle data using the translated clip text of FIG. 15B will be described below.

[0026] The accompanying drawings, in which like reference numbers refer to identical or functionally similar elements throughout the separate views, and which, together with the following detailed description, are incorporated in and form a part of this specification, serve to further illustrate embodiments of concepts comprising the claimed disclosure and to explain various principles and advantages of those embodiments.

[0027] The methods and systems disclosed herein have been represented, where appropriate, by conventional symbols in the drawings, showing only those specific details that are relevant to understanding the embodiments of the disclosure, so as not to obscure the disclosure with details that will be readily apparent to those skilled in the art having the benefit of the description herein. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0028] Hereinafter, the embodiments of the present invention will be described with reference to the drawings. These embodiments are provided for a better understanding of the present invention, and the present invention is not limited to these embodiments. Changes and modifications obvious from the embodiments still fall within the scope of the present invention. Meanwhile, the original claims form part of the detailed description of this application.

[0029] The need to provide video subtitles Many creators monetize their videos on platforms such as YouTube. The more views a video gets, the more money a creator can make, so it is important to reach a larger audience. Providing subtitles is a way to attract more viewers. However, creating subtitles without the use of automated technology can be labor-intensive and time-consuming.

[0030] The technology presented This application discloses solutions, systems, and methods for generating, processing, and presenting subtitles for a video (a target video). The solutions, systems, and methods presented in this application are collectively referred to herein as "the present technology" or "the presented technology."

[0031] Non-limiting embodiments Hereinafter, the embodiments of the present technology will be described with reference to the drawings. The present technology is not limited to the described embodiments. Changes and modifications obvious from the described embodiments still fall within the scope of the present technology.

[0032] FIGURES shoW non-limiting examples The drawings are described in detail for understanding non-limiting embodiments of the technology. The drawings are for illustrative purposes and are not intended to limit the technology to the illustrated embodiments.

[0033] Subtitle Format The subtitles may be stored as a single file. A variety of subtitle formats may be used, such as SubRip, SubViewer, Timed Text Markup Language (TTML), SBV (YouTube format), Distribution Format Exchange Profile (DFXP), and Web Video Text Tracks (Web VTT). In embodiments, the subtitles for the target video may be stored using formats other than the examples and may be stored as multiple interrelated files.

[0034] Components of subtitles In embodiments, subtitles for a target video include at least two components: (1) text to display (collectively, "caption text" for the target video), and (2) timing information (timestamp, timecode) for displaying the text. Subtitles may include one or more additional components. For example, markup (bold, italics, underline), font, character size, spacing, and position information may be included in the subtitles. In embodiments, the term "caption text" or "caption data" refers to the entire text that is displayed as a subtitle for the target video.

[0035] FIG. 13 shows exemplary closed caption data 1300, including a sequence number 1312 for the clip, timecodes 1314 indicating the start and end of the clip, and clip text 1240 for the clip.

[0036] Caption text taken from the video's audio Subtitling of a video includes text to visualize speech or sounds in the video. The technology may obtain such text from processing of the audio of the video. The technology may use audio recorded with the video (live recording) and audio recorded separately from the video (dubbing, narration). In embodiments, audio that is part of, associated with, or related to the video may be used to obtain the caption text.

[0037] Speech-to-text translation (STT) to obtain caption text The present technology can use speech-to-text translation (STT) technology on the audio of the video. The STT technology can analyze components of the audio, remove noise from the audio, recognize one or more utterances (words) from the audio, recognize one or more languages ​​of the utterances, and transcribe the recognized utterances into text data (STT text or transcribed text) in the recognized language(s). The STT process can transcribe the audio word-by-word or character-by-character to obtain a sequence of spoken words. At least a portion of the obtained STT text (or a modified version thereof) can be used as caption text to visualize the recognized speech in the video. In an embodiment, a different audio transcription technology from the example can be used. FIG. 9A shows a transcribed text 920 obtained by processing the audio of the video 910.

[0038] Using pre-prepared scripts In an embodiment, a screenplay or script prepared for shooting a video can be used as text data (caption text) for subtitles of the video. The present technology can extract lines of text from a screenplay of a video prepared in advance, determine a portion (clip) of the video corresponding to the lines of text, and display the lines of text as subtitles for the determined portion of the video. In an embodiment, text other than a screenplay or script can be used.

[0039] Missing punctuation in STT text or pre-prepared script In embodiments, the subtitles are generated from processing text having punctuation, where the techniques may perform one or more of removing punctuation, verifying punctuation, and identifying additional punctuation.

[0040] The language of the caption data The caption text may be in one or more languages ​​spoken in the video, which for audio in the video is referred to hereinafter as the "original language" or "first language."

[0041] Use a combination of transcribed text and pre-prepared scripts In an embodiment, STT text (transcribed text) obtained from the audio of the video, a pre-prepared script of the video, and a combination of the two may be used as text data (caption text) of the subtitles of the video. For example, when creating subtitles based on a pre-prepared script of the video, the technology may use one or more words in the STT text to modify (replace), add, or delete one or more words in the script to reflect what is actually spoken in the video. In another example, when creating subtitles based on a pre-prepared script of the video, the technology may use one or more words in the script to replace, add, or delete one or more words in the STT text. For example, slang spoken in the video may be replaced or deleted in the subtitles.

[0042] Determining the timing of caption text components To synchronize a target video with its caption text, the technique determines, calculates, or selects the timing of one or more words (components) of the caption text. In embodiments, the technique determines a start time and an end time for each word of the caption text. In embodiments, timing information may be determined for one or more components other than words (e.g., characters, clauses, phrases, sentences, paragraphs) of the caption text.

[0043] Determining the timing of caption text based on matching sound timing The technology can analyze the audio of the target video to identify silence (and / or noise), identify sounds (or speech) separated by silence or noise, and determine the timing (start / end time) of the identified sounds. In an embodiment, the technology determines the timing of one or more words in a given script (caption text) that does not have timing information. The technology may identify matching sounds in the target video based on a simulated pronunciation of the word and determine the start and / or end time of the sound as the timing of the word(s). In an embodiment, the technology determines the timing of the caption text words when transcribing the audio of the target video. The technology may use the start time of the speech as the start time of the transcribed text of the speech and the end time of the speech as the end time of the transcribed text. In an embodiment, the timing of components (letters, words, phrases) of the caption text may be determined based on the timing of the corresponding sounds of the components in the target video using a process other than the example.

[0044] Timing information format In embodiments, timing information for components of the caption text may be stored using one or more of a time from the start of the target video, a time to the start of the target video, a frame number, and a code capable of indicating a particular time within the target video. In embodiments, any data format capable of indicating a point in time or segment within the target video may be used.

[0045] Timed script The caption text and associated timing information are hereinafter collectively referred to as "timed text data" or "timed script." A timed script may be a single text file that contains the sequence of words in the target video (the caption text) and the timing of each word. In embodiments, a timed script may be stored using a file format other than text, or may be stored using multiple files. FIG. 9A shows a transcribed text 920 obtained by processing the audio of a video 910. FIG. 9B shows an exemplary timed script 940 in which timing information is added to each word of the transcribed text 920. In the timed script 940, the word "dream" 952 is associated with a start time 954 and an end time 956 of its corresponding sound.

[0046] Adjust timecode in a timed script to sync with the video's audio Given a script with time codes, the technique can adjust or verify the time codes so that words in the script are synchronized with corresponding sounds in the video.

[0047] Clip and Clip Cation The technique can process a timed script to determine clips (portions) of a target video in which to present subtitles and to determine corresponding text (clip cations) to display as subtitles for the clips. The term "clip" (or "video clip") refers to a portion (or period) of a video that displays (or maintains) the same subtitle text. The term "clip caption" (or "clip text") refers to text that is displayed as a subtitle for the corresponding clip.

[0048] Same captions maintained throughout the clip In an embodiment, the entire clip caption is displayed at the beginning of the clip, remains for the duration of the clip, and disappears at the end of the clip. The same clip caption (clip text) may be displayed throughout the entire clip without modification or interruption. In an embodiment, visual effects or markups (bold, italics, underline) may be applied to only a portion of a single clip while maintaining the same text characters. In other embodiments, the words in a single clip caption are displayed sequentially according to their individual timing information (timecode) such that the entire clip caption is displayed at the end of the clip. In an embodiment, the clip caption may be displayed in a different manner than the example, so long as the entire clip caption is displayed at least at the point of the clip.

[0049] Defining clips using clip text timing In an embodiment, the technology may first define a clip caption and then define a corresponding clip based on the determined timing information of the clip caption. For example, if a clip caption is defined to have a start word and an end word, the start time (timecode) of the start word is determined as the start time of the clip and the end time (timecode) of the end word is determined as the clip. A predetermined time adjustment may be applied to determine the start time of the clip based on the timing of the start word and to determine the end time of the clip based on the timing of the end word. In an embodiment, the technology may first define a clip and then define its clip caption to include all text of the corresponding period.

[0050] Clip caption with words In an embodiment, a clip caption (single clip) is defined to include one or more words. A single word cannot be separated into two clips. In an embodiment, a single clip includes a fragment of a word if only the fragment is spoken in the video or if there is a long silence between the spoken fragment and other subsequent fragment(s). In an embodiment, a clip caption may be defined using higher grammatical units (phrases, clauses, sentences).

[0051] Grouping of words to define clip / clip-cation In an embodiment, the technology groups two or more consecutive words in a caption text into a single clip clip caption (clip text). Words may be grouped by sentence, such that two words in a single sentence are included in a single clip caption. In an embodiment, two words in a sentence may be separated into two clip captions if there is a long silence between the two words or if the sentence is too long for a single clip. In an embodiment, a single clip may contain words from two different sentences. In an embodiment, words may be grouped using grammatical units other than sentences (phrases, clauses) or segments of caption text other than grammatical units.

[0052] Identifying grammatical units / segments In embodiments, the technology may process the caption text to identify grammatical units (words, phrases, clauses, sentences) or other segments within the caption text. In embodiments, the technology may identify grammatical units or other segments by reference to punctuation marks (periods, question marks, exclamation marks, commas, etc.) within a script provided as caption text. In embodiments, the technology may determine potential locations of punctuation marks in STT text that does not have punctuation marks. Exemplary processes for identifying grammatical units or other segments in caption text are described later in this disclosure.

[0053] Machine-trained models for identifying sentences In embodiments, a machine-trained sentence identification model (hereinafter "sentence model" or "sentence artificial intelligence") is used to identify one or more sentences in the caption text (caption data). The sentence model processes the caption text and identifies the beginning and / or end of one or more sentences in the caption text. In embodiments, techniques other than machine-trained models may be used.

[0054] Sentence model input - word sequence In an embodiment, the sentence model is configured to receive as its input a predetermined number of words (e.g., 200 words). In an embodiment, the caption text (STT text, pre-written script) is split into several smaller word strings to meet the predetermined requirements for the input of the sentence model. If the word string is shorter than the predetermined number, one or more dummy words or null values ​​may be input together with the word string. In an embodiment, the sentence model may be flexible to receive inputs of different sizes. In an embodiment, the input data size may be defined using units other than number of words (e.g., number of characters).

[0055] Word Prescreening In an embodiment, certain words are removed from the input text to the sentence model. For example, articles ("a", "an", and "the") may be excluded from the input to the computed sentence model because articles generally do not terminate sentences. In an embodiment, if a screenplay or script contains words outside the line text that describe scenes in the video (e.g., "laughter", "background music"), such words may be excluded from the input to the sentence model.

[0056] Sentence model output - end / beginning probability In an embodiment, the sentence model is configured to calculate, for one or more words in the input text, the probability that the word is the last word of a sentence (final probability) and / or the probability that the word is the first word of a sentence (initial probability). In an embodiment, the final probability of a word represents the probability that a particular final punctuation mark follows the word. In an embodiment, the initial probability of a word represents the probability that the word follows a particular final punctuation mark. In an embodiment, since the initial probability of a word is the same as the initial probability of the following word, a sentence model that calculates final probabilities may be referred to as a sentence model that calculates initial probabilities. According to FIG. 10, the sentence model 1010 calculates a final probability 1020 (in percentage) for each word in the input text 920. Word 1022 has a final probability of 99%.

[0057] Sentence-ending probability for each punctuation mark In an embodiment, a sentence model is used to calculate multiple sentence ending probabilities for a single word, each corresponding to a final punctuation mark (period, question mark, exclamation mark, ellipsis). The presented technique can add the multiple sentence ending probabilities to calculate a representative sentence ending probability or take the highest value of the multiple sentence ending probabilities. In an embodiment, separate sentence models may be used for different final punctuation marks.

[0058] Adjusting the sentence model output In embodiments, the sentence-ending probabilities (or sentence-beginning probabilities) calculated by the sentence model can be adjusted based on a variety of factors: predefined default probability values ​​for the word itself, certain neighboring words, the presence of well-known or established phrases, or can be used to adjust certain grammar tools or techniques.

[0059] A predefined threshold for determining whether a word is at the end of a sentence In an embodiment, if the word's sentence-ending probability is greater than a predefined threshold, the word is determined as a sentence-ending word. This threshold may be specific to one or more words, may be common to all words, may be set or adjusted by the sentence model, may be set manually by a user, or may be set by a software administrator or programmer. This threshold may be different for each word and its translation, or may be uniform across the language (same for the word and all translations). In an embodiment, a sentence-ending word may be determined using one or more criteria other than the predefined threshold. In FIG. 10, if this threshold is 90 percent, four words 1022, 1024, 1026, 1028 are identified as sentence-ending words.

[0060] Determining a sentence by its final word In an embodiment, a word immediately following a sentence final word in a caption text may be determined as the initial word of the next sentence. The first word of the caption text data is another initial word. One or more words from an initial word to an immediately following sentence final word constitute a sentence. A word sequence can be identified as a sentence, but the identified sentence may not be a grammatically complete sentence. In FIG. 11, the STT text 920 is divided into five segments 1110-1150 based on four sentence final words 1022-1028. Four sentences 1110-1140 are identified. In an embodiment, the last segment 1150 is combined with the start port of another STT text immediately following the STT text 920 to form a complete sentence (or clip) such that the start word of the segment "I" is used as the clip start word.

[0061] Definition of Clip by Text In an embodiment, a clip caption (clip text) and its corresponding clip may be defined to include all words of one or more complete sentences. For example, in FIG. 11, each of the four sentences 1110 to 1140 in FIG. 11 may be defined as the clip text of a single clip.

[0062] When a sentence is defined as the clip text of a single clip, the clip (clip period) can be defined using the timing information of the start word and the end word of the sentence. The start time (timestamp, time code) of the start word can be used as the start time of the clip, and the end time (timestamp, time code) of the end word can be used as the end time of the clip. For example, in the first sentence 1110 of FIG. 11, which is used as the clip text of a single clip, the start time "00:00,175" 984 of the start word "so" 982 is used as the start of the clip, and the end time "00:00,720" 956 of the end word "dream" 952 is used as the end of the clip. In an embodiment, clips corresponding to individual sentences can be combined to form longer clips, and a clip corresponding to a single sentence can be split into two or more clips based on in-sentence cuts.

[0063] Timing Adjustment of Clip In an embodiment, an adjustment may be applied so that the clip starts a predetermined time earlier (or later) than the first word. In an embodiment, an adjustment may be applied so that the clip ends a predetermined time later (or earlier) than the last word. In an embodiment, the start and end of the clip may be defined in a different way than the example, as long as the synchronization between the clip and its corresponding sentence(s) is not impaired.

[0064] In-Sentence Delimiters for Defining Clip Captions In embodiments, the technology can process at least a portion of the caption text to identify one or more sentence breaks and define clip cations and corresponding clips based on the sentence breaks. For example, one or more sentences identified using the sentence model can be further analyzed to identify one or more breaks within the sentence, and the identified sentence breaks can be used to split a clip containing the sentence.

[0065] Intra-sentence model In embodiments, the technology may use a machine-trained intra-sentence break identification model (hereinafter "intra-sentence model") to identify one or more breaks within a sentence of the caption text. The intra-sentence model may be configured to receive a sequence of words and output, for each word in the input, a probability that the word is followed by an intra-sentence break or that the word immediately precedes an intra-sentence break (hereinafter "intra-sentence break probability").

[0066] Intra-sentence model input - sentences identified using the sentence model In an embodiment, the intra-sentence model is configured to receive as its input one or more sentences identified using the sentence model. In an embodiment, the intra-sentence model is configured to receive a portion of the caption text without reference to a sentence identified using the sentence model. The intra-sentence model may have a maximum number of words for its input (e.g., 50 words) or may be shorter than the maximum number of the sentence model (e.g., 300 words).

[0067] Excluding short sentences from the input of intra-sentence models In an embodiment, if a sentence is shorter than a certain length (e.g., number of characters) allowed for a single clip, there may be no need to separate the sentence into two or more clips, and the sentence can be excluded from the input of the intra-sentence model.

[0068] Output of the intra-sentence model - intra-sentence segmentation probability In an embodiment, the sentence break probability of a word represents the probability that the word immediately precedes (or follows) one or more sentence punctuation marks that indicate a sentence break (e.g., a comma, a dash, an ellipsis, a semicolon, etc.) In an embodiment, the sentence break probability of a word represents the probability that the word is the last word (or the first word) of a phrase or clause.

[0069] Various embodiments of intra-sentence model probabilities In an embodiment, the intra-sentence model assigns a different probability value to each word based on the probability that each word immediately precedes various types of intra-sentence breaking punctuation marks. For example, there is a 70% probability that the punctuation mark after the word is a comma, an 80% probability that it is an ellipsis, and a 90% probability that the punctuation mark is a semicolon. The intra-sentence model then selects the most probable punctuation mark for the word, which in this case is a semicolon for the word.

[0070] In an embodiment, the intra-sentence model simply assigns a probability score to each word being an intra-sentence delimiter word based on the probability that each word is adjacent to or immediately precedes a comma, regardless of what punctuation may follow the word, and assigns a single probability score to each word in the examined sentence.

[0071] Depending on the embodiment, the intra-sentence model may select the punctuation mark with the highest probability for each word, or may assign a punctuation probability for each word based on the most probable punctuation mark (i.e., the exclamation mark in this case). In an embodiment, the sentence model compares the different probabilities of each word being immediately preceded by various punctuation marks, and all of these comparisons can be used against the various probability scores of each of the other words in the text.

[0072] 12A, the intra-sentence model 1210 calculates intra-sentence break probabilities 1220 (percentages) for words in the input sentence 1110. The model 1210 did not calculate an intra-sentence break probability for word 952 because word 952 is a sentence-final word.

[0073] Sentence-defined clip splitting In embodiments, a clip defined to include or contain one or more sentences identified using the Sentence AI may be split into two or more clips by one or more intra-sentence breaks identified using the intra-sentence model. In certain embodiments, clips may be defined after identifying intra-sentence breaks using timestamp information of the sentence ends and intra-sentence breaks.

[0074] In an embodiment, the intra-sentence model determines segments or portions within a sentence by determining the location of intra-sentence breaks, preferably by determining the location of words immediately preceding a comma. These sentence segments or portions may be divided by intra-sentence punctuation as described above, or in alternative embodiments, by spaces, pauses, or other determinations made by the intra-sentence model.

[0075] In an embodiment, these intra-sentence breaks defining sentence portions or segments may then be used to mark the location of the intra-sentence breaks in the STT text, the text file, and / or corresponding locations in the video clip and / or audio file, while the clips defined by the sentences may be further timestamped and / or further divided into additional clips.

[0076] Determining intra-sentence breaks - threshold In an embodiment, for a word to be considered in a particular position within a sentence, for example a sentence break word, or a word immediately preceding a sentence break that may be defined by sentence punctuation, its sentence break probability must meet or exceed a predetermined threshold. In Figure 12A, if the threshold is 90%, then word 962, which has a sentence break probability of 98%, is identified as a sentence break word.

[0077] This threshold may be defined for each word, common to all words, set or adjusted by a sentence model, manually set by a user, or by a software administrator or programmer. Also, the threshold may vary by language. For example, in English, an 85% sentence break probability or a score of 85 may be assigned to a specified threshold considered as the last word before a sentence break within a sentence. However, in Korean, it may be set to 80% or a score of 80. This threshold may be specific to one or more words or uniform for all words in the entire language. When the probability of a comma or punctuation mark assigned to a word meets a predefined threshold, that word is considered the last word within a sentence segment by the in-sentence model or, in various other embodiments, is considered to occupy a specific position within the sentence.

[0078] Adjustment of Threshold Comma / Punctuation Probability Values In an embodiment, the in-sentence model can determine or adjust the probability value of punctuation. The threshold may vary depending on various words, positions, or spaces or may be common to all words in the language.

[0079] Using the Sentence Break Word as the Clip End Word In an embodiment, the present technology uses the sentence break word as the clip end word such that the sentence identified in the STT text is split into two or more clip texts. According to FIGS. 12A and 12B, the word "tomorrow" 962 is identified as the sentence break word and as the clip end word, whereby the first sentence 1110 of the text 920 is split into two clip texts 1240, 1260.

[0080] Generation of Subtitle Data from a Timed Script 13 illustrates exemplary subtitle data 1300 generated based on clip end words in a timed script 940. Among the words in the script 940, four sentence end words 1022-1028 identified using the sentence model 1010 are used as clip end words, and the intra-sentence delimiter word 962 identified using the intra-sentence model 1210 is also used as a clip end word. Additionally, the first word "so" 982 of the script 940 is used as the word that starts the clip.

[0081] 13, the first sentence 1110 of the text 920 is split into two clip texts 1240, 1260, while the second sentence 1120 remains a single clip text. The subtitle data 1300 includes three segments 1310, 1320, 1330 corresponding to the clip texts 1240, 1260, 1120, respectively. The first segment 1310 of the clip text 1240 defines the serial number 1312 of the clip text, and a time code 1314 defines the clip period in the video 910 during which the clip text 1240 is displayed as subtitles.

[0082] Storage location of sentence separator The locations of the identified intra-sentence break words can be marked in the text or STT file, which can then be linked (i.e., via timestamp information) to the location of the words in the video and / or associated audio (audio in the video, or dubbed audio). When the intra-sentence model parses the full text file, it identifies the ending words of each sentence segment and marks their respective locations in the text and subsequently in the video / audio. Thus, the marked locations (times / frames in the video) are indicators of the end of the sentence segments, and each new sentence starts from the end of the last sentence.

[0083] Results of intra-sentence identification Once the intra-sentence breaks are identified, the text, data, or STT file may be time-stamped, and the corresponding locations in the video clip and accompanying audio may also be time-stamped. The identified or time-stamped locations in the clip may then be used to further split the clip into smaller clips. The clip initially defined by the sentence model may be further cut, marked, identified, or spliced ​​into new clips at the identified locations of punctuation or specific pauses due to the intra-sentence model's determination of intra-sentence breaks. Thus, the clip created by the sentence model's identification of sentence end words may contain one or more other sentences that may be identified by the intra-sentence model, and the first clip is split into separate clips, each with a clip caption consisting of a sentence segment.

[0084] Storing timestamp information for sentences / sentence separators In embodiments, the technology may store or mark the location of each initial and final sentence word in the caption text (e.g., STT text), the timed script, or separate data connected to the caption text or the timed script, so that the identified sentences can be linked to corresponding portions (clips) of the target video.

[0085] In embodiments, the technique may store or mark the location of each intra-sentence break, which may be the end time of the word immediately preceding the break, or the start time of the word immediately following the break.

[0086] Alternative predefined probabilities In many embodiments, association probability values ​​or scores are provided between each word and various punctuation marks used in the associated language, such as probability values ​​for end-of-sentence punctuation marks, such as periods, or intra-sentence punctuation marks indicating pauses, such as commas. One or more of the AI ​​models already described, or alternative algorithms, can use these pre-provided punctuation probability values ​​for each word to determine whether a comma, period, or any other suitable punctuation mark available needs to be inserted adjacent to the word. The probability of punctuation marks for each word may be different for each side adjacent to the word. However, in embodiments, only one location adjacent to each word on the side where punctuation marks are most likely to be present is considered.

[0087] Post-output adjustment using intra-sentence models In an embodiment, the intra-sentence model may generate or adjust punctuation probability values ​​for words in the text file after its initial output. Adjustments may be made based on a number of factors, including but not limited to the default settings or punctuation probabilities for each word, the presence of punctuation in the input text, the presence of identified sentence-ending words, the probability value of the word being a sentence-ending word, and the assigned probability values ​​and punctuation probability values ​​of other words in the text. In an embodiment, the technology may take silence into account to adjust the intra-sentence break probability. For example, if a word is followed by a pause or silence longer than a predetermined length, the word may have a higher intra-sentence break probability.

[0088] Using sentence and intra-sentence models in sequences In an embodiment, a sentence model is first run on the input text to determine an initial set of sentences to determine an initial set of clips derived from the original video, each clip containing one complete sentence. A second intra-sentence model is then subsequently run to enhance the output of the sentence model and identify intra-sentence breaks within the identified sentences / clips, thereby further identifying and deriving clips that require splitting the already identified clips into additional clips, or possibly combining different clips as needed to complete the sentences.

[0089] Using two separate models In an embodiment, the technology uses two separate models: one model to identify sentences (sentence model) and the other to distinguish (intra-sentence model). The intra-sentence model is to find the right place in each sentence to further break down the sentence, which can be done more accurately than the sentence model. The intra-sentence model may be more accurate in finding commas in sentences because it is an AI model trained primarily for this purpose and provides inputs of sentences that are already defined both when it is trained and when it is used with input data. Since the two AI models may have different inputs, different outputs, require different training datasets, and require different training techniques to meet their objectives, it may be efficient to separate the sentence model and the intra-sentence model.

[0090] Combining sentence-intra-sentence models In embodiments, the present technology may train a single machine-trained model to perform the functions of the sentence model and the intra-sentence model. In embodiments, the present technology may train a sentence model and an intra-sentence model and then combine the two trained models into a single model.

[0091] Probability Table In embodiments, the techniques may use a static table that includes a plurality of words and one or more predetermined probability values ​​for each word. The one or more predetermined probability values ​​for a word may include one or more of an end-of-sentence probability for the word and an intra-sentence break probability for the word.

[0092] Other Factors for Determining Sentence and Subsentence Breaks In an embodiment, if a word is followed by silence (or a pause) longer than a predetermined time, the present technology may determine the word as a sentence end or may increase the sentence end probability of the word. In an embodiment, if a word is followed by silence (or a pause) longer than a predetermined time, the presented technology may increase the probability of a sentence break of the word or determine that a sentence break follows the word. In an embodiment, if the number of sentences in the input text has been determined or is known, the word with the highest probability may be selected as a sentence end word to fill the number. The present technology may consider one or more factors other than punctuation to identify sentences or sentence breaks and may configure the sentence model or intra-sentence model accordingly.

[0093] Clips that are too long or too short The length of the clip duration (or clip text) may indicate that it is too long and needs to be split further, or that it is too short and some clips need to be combined. In an embodiment, the technology uses one or more of the AI ​​models described above to split a clip into multiple clips if the clip length exceeds a predetermined maximum length, or to combine a clip with other clips if the clip is less than a predetermined minimum length.

[0094] For example, an intra-sentence model may be deployed on clips that are deemed too long to be further broken down into sentence parts. Alternatively, a sentence or intra-sentence model may be utilized to combine clips with surrounding clips, whether they are segments of other sentences or other complete sentences. A user may manually split clips and set configurations to split clips that are too long, and a maximum clip length may be manually set by the user.

[0095] Ignoring punctuation In embodiments, one or more identifiable punctuation marks (or sentence breaks) may be ignored when defining sentences and clips. For example, ignoring punctuation marks may result in longer clips that may contain multiple punctuation marks, especially if the punctuation marks do not strongly correspond to pauses or are not strong indicators of the end of a sentence. In embodiments, one or more identifiable punctuation marks may be ignored if the word punctuation rate is not high enough relative to the threshold, even if the threshold is met.

[0096] Clip caption onscreen length display time limit In embodiments, a time limit may be set for how long a clip caption may be displayed on a video clip. For example, a clip caption may be limited to being displayed on a video clip for a maximum defined period of time. In these instances, the caption text may be removed, or the clip may be shortened, or split into several clips. Captions may also have a minimum time limit that they must be displayed.

[0097] Training a machine-trainable model The present techniques may use various known training techniques to obtain a machine-trained model with desirable performance. In embodiments, the present techniques may use machine learning techniques, which may include, but are not limited to, deep neural networks, autoencoders, vibrations or other types of autoencoders, and generative adversarial networks.

[0098] For example, training of a model is complete when, for each of the input data in the training dataset, the output from the model is within a predefined tolerance from the corresponding desired output data (labels) in the training dataset.

[0099] Datasets for training machine-trainable models To prepare a machine-trained model, the techniques may develop or create a dataset for training a machine-trainable model. The training dataset includes a number of data pairs, each pair including input data for training a machine-trainable model and desired output data (labels) from the model in response to the input data.

[0100] Training data for sentence models In an embodiment, to train the sentence model to calculate a sentence-final probability for each word in an input text without punctuation, the training input data may include word sequences without punctuation, and the corresponding training output data (desired output for the input) may be values ​​indicative of each sentence-final symbol (e.g., 100% for sentence-final words and 0% for other words). Word sequences without punctuation may be generated by removing punctuation from properly punctuated text. In an embodiment, the training output data may be an indication of a particular sentence-final punctuation. The training dataset may be in a format different from the example.

[0101] Training data for intra-sentence models In an embodiment, the intra-sentence model may be trained primarily on a set of properly punctuated sentences. In an embodiment, the training input data includes one or more complete sentences without punctuation, and the corresponding training output data is a value indicative of the intra-sentence break of the sentence (e.g., 100% for words immediately followed by intra-sentence punctuation, 0% for other words). In an embodiment, the training dataset may be configured differently from the examples.

[0102] Different models from static tables In embodiments, a machine-trained model differs from a static table of words and their corresponding probabilities in that the model can output different values ​​for the same word. In Figure 10, words 1022 and 1024 have different end-of-sentence probability values ​​but the same text "dream."

[0103] Language Sensitivity of Models In embodiments, the technology may train and configure separate versions of sentence models (and intra-sentence models) for different languages. To provide subtitles for a video recorded in a first language, the technology may need to use a first language version of a machine-trained model. Training of the first language model may rely primarily on a training dataset in the first language and may use additional training datasets set in one or more foreign languages. In embodiments, the technology may train a single model to handle two or more languages.

[0104] Working with video in the Clip View In an embodiment, the target video may be marked, markered, edited, or time-stamped to indicate the time location of the clips. In an embodiment, the target video is divided into multiple portions, each portion corresponding to an identified clip.

[0105] Translation subtitles In embodiments, the present technology can generate one or more translated subtitles for a target video using subtitles generated in the original audio language (spoken language, original language) of the target video. As previously described, the caption text in the original language is either provided as a script or generated by one or more audio processing techniques (STT) or other AI methods. The caption text in the original language may be translated into a desired foreign language (the language into which the caption text is translated is hereinafter referred to as the "translation language" or "second language").

[0106] Sentence-by-sentence translation In an embodiment, the technology translates the caption text into the translation language sentence by sentence. If the caption text in the original language includes a sentence-ending symbol, the sentences separated by the sentence-ending symbol may be translated individually. If the caption text in the original language does not have punctuation information, as in the case of STT text, the technology performs a sentence-by-sentence translation using sentences identified using a sentence model. FIG. 14 shows a translation 1400 of Korean STT text 920. Korean K 1400 includes three translation sentences 1410, 1420, 1450 corresponding to the original language text 1110, 1120, 1150, respectively.

[0107] In certain embodiments, two or more sentences may be translated together. Sentence-by-sentence translation may use sentences (identified sentences) as units of translation, but may allow two or more sentences to be translated together. In embodiments, translation units other than sentences may be used (e.g., by word, by phrase, by clip, or a combination of different translation units).

[0108] Translated subtitles - clips based on caption text in the original language In embodiments, to provide the translated subtitles, the translated caption text in the translation language may use or employ the same clip ("original clip") defined based on the caption text in the original language (defined with timestamps in the original language) and be synchronized with the audio of the target video in the original language. As described above, the technology may determine the clip based on end-of-sentence and intra-sentence breaks identified using sentence and intra-sentence models in the original language. The technology may use determining clips for the translated subtitles as well as the subtitles in the original language.

[0109] However, in some embodiments, the technology may process the translated caption text date to identify end-of-sentence and identified mid-sentences using a translation language version of the sentence and an intra-sentence model, and may determine clips that differ from the defined clip based on the original language subtitle text.

[0110] Assigning translated caption text to the original clip In an embodiment, if the translated subtitle follows the original clip and the clip contains only a sentence, the translation may be assigned in its entirety to the same clip, if the clip contains more than one sentence, the entire translation may be assigned to the same clip, maintaining the same order of the sentences.

[0111] In embodiments, the translated subtitle follows the original clip(s), and when a sentence is split into two or more clips (using sentence and intra-sentence models), the translation may be split into the same number of clips such that the original and translated sentences are synchronized when the original and translated subtitles are displayed together.

[0112] Split the translation into as many clips as there are original language sentences In an embodiment, when an original language sentence is split into two or more clips in the original language subtitles, further processing of the translation may be performed because the translation may not have punctuation or other indications for splitting the original clip into two or more clips.

[0113] The technology can deploy a third AI model to identify intra-sentence breaks in the translated sentence. The third AI model can be a version of the intra-sentence model that is trained and configured in the translation language. ("translation intra-sentence model"). The translation intra-sentence model can be deployed for each translated sentence, and then aims to split the sentence into a number of sentence parts that match the number of clips that the sentence is split into in the original language.

[0114] According to FIG. 15A-FIG. 16, an intra-sentence model 1510 (e.g., a Korean version of the model 1210) calculates intra-sentence break probabilities 1520 (in percentage) for words in a translation 1410 of an original language sentence 1110. The model 1510 did not calculate a probability for the word 1524 because it is a sentence-ending word. In an embodiment, when an original language sentence 1110 is divided into two clip texts 1240, 1260 to form an original language subtitle data 1300, the translation 1410 is divided into an equal number (two) of clip texts 1610, 1620. To divide a sentence into “n” (a natural number greater than 2) clip texts, “n-1” word(s) with maximum intra-sentence break probabilities are determined as intra-sentence clip ending word(s). In Figure 15A, the word 1522 with the highest intra-sentence break probability (84%) is identified as the only clip-ending word in the translated sentence 1410. In Figure 15B, the translated sentence 1410 is split into two clip texts 1542, 1544 to obtain a set of translated clip texts 1530.

[0115] This selection can be made regardless of whether the word's intra-sentence break probability (84%) is greater than a predefined threshold (e.g., 90%) for identifying clip-ending words in the original language sentence. In an embodiment, even if there are two or more words with intra-sentence break probabilities greater than the predefined threshold, only one ("n-1") word with the greatest intra-sentence break probability can be identified as the clip-ending word to split the translation sentence into two ("n") clip texts.

[0116] Generate translated subtitle data with the same time codes as the original language subtitles 16, translated subtitle data (Korean) 1600 can be obtained by replacing text in the original language subtitles 1300 clip by clip, rather than word by word. Each clip text 1240, 1260, 1120 of the original language subtitles 1300 is replaced with its corresponding translated text clip 1542, 1544, 1420, respectively.

[0117] In an embodiment, the sequence numbers and time codes of the original language subtitles 1300 may be maintained in the translated subtitle data 1600. In the translated subtitle data 1600, the first translated clip text 1542 replaces the first original language clip text 1240 while maintaining the same sequence numbers 1312 and time codes 1314. Because the translated clip text 1542 is not spoken in the video, the timing of displaying the translated clip text 1542 is determined such that the translated clip text 1542 is synchronized with the sound of the corresponding clip text 1240.

[0118] Input and output of in-sentence models In an embodiment, the translation intra-sentence model is very similar to the original language intra-sentence model described above, being trained in a particular language and specialized to that language to divide sentences in that language into sentence parts or segments. The construction, training, and operation of the translation intra-sentence model can be understood with reference to the construction, training, and operation of the original language intra-sentence model.

[0119] In an embodiment, the translation intra-sentence model is provided with individual sentences as input, and in an embodiment, each sentence is 30 words or less in the translation language. The translation intra-sentence model can then identify intra-sentence breaks in the translation sentence, in an embodiment, by identifying words that have the highest probability of being the word immediately preceding a comma, i.e., their comma probability, and in other embodiments, by identifying words that have a punctuation probability of a punctuation mark that serves as an intra-sentence break in the translation language. As a result, words that meet or exceed the threshold probability may be identified as intra-sentence break words in the translation language.

[0120] Sentence segments as output of the in-sentence model In embodiments, the technique may then attempt to match clips defined in the first language with sentences and / or sentence portions in the translation language if the clips defined in the first language closely match sentences in the second language. If a complete sentence in one language is equivalent to a complete sentence in the second language, a perfect match clip has been produced. Text from a translation language text file is then provided for the clip, and a subtitle of the entire sentence can become the translation clip caption associated with the video and audio.

[0121] However, in an embodiment, if there are multiple sentence portions, each corresponding to a clip, the matching translation portions output by the translation intra-sentence model may match the clip defined by the original language AI model and be displayed as translation clip captions that may be displayed along with the original language clip caption.

[0122] No sentence ending determination in the translation language In embodiments, when the technique performs sentence-by-sentence translation, there may be no determination of sentence endings or definition of sentences using a translation language sentence model, since individual sentences in the source language are translated into individual sentences in the target language. Although the subtitle generating software or platform may have a translation language sentence model for generating subtitles for a video recorded in the translation language, the translation language sentence model may not be used in the process of generating subtitles by translating the source language subtitles.

[0123] Live or recorded video The techniques may be implemented, performed, or executed on stored or pre-stored video, In embodiments, the techniques may be applied to provide subtitles for live video streams, as well as to generate subtitles in real-time while video is being recorded live.

[0124] Timing of Processing In embodiments, processes or actions for providing subtitles for a video may be performed when the video is being recorded, when the video recording is paused, stopped, or ended, or when the video is saved or loaded to a particular application, computing, storage device, or cloud network.

[0125] Figure 1 FIG. 1 is a flow chart illustrating an example method 100 for providing subtitles to a video. A video is received or loaded from a local or remote data store 105 for further processing. At least one or more text files related to the video or the video's associated audio (e.g., dubbed audio) are received or obtained (110). The text files are input to a sentence model (115). The sentence model identifies one or more sentence-final words based on sentence-final probability values. Words that meet or exceed a predetermined threshold probability value may be identified as sentence-final words. The locations of the sentence-final words may be marked and / or time-stamped within the text file and / or video file. The identified sentences are defined and output (120). The sentences are then input to an intra-sentence model (125). The intra-sentence model is run individually for each sentence to attempt to define one or more sentence segments or portions by identifying intra-sentence breaks. A sentence break may be detected when a word has a sentence break probability that meets or exceeds a certain probability threshold, which may be considered by the sentence model to be immediately preceding a comma or another sentence break punctuation mark that serves as a sentence break. Sentence portions separated by the identified sentence breaks are obtained (130). Based on the identified sentence endings and the identified sentence breaks, clips (clip text and corresponding clip durations) are defined (135), and a subtitle file is generated according to the defined clips. When playing the video, the clip texts are displayed sequentially during the corresponding clip durations as subtitles (140).

[0126] Figure 2 FIG. 2 is a flow chart illustrating an exemplary method 200 for generating and displaying translated subtitles for a video. A video is received or loaded (205) for further processing. A text file (original spoken language) associated with the video or its associated audio is received (210). Steps 215, 220, 225, 230, 235 of defining clips are the same as steps 115-135 of FIG. 1. The identified sentences from the text file are then translated separately (240) to obtain the translated sentences. A translation intra-sentence model is then run (245) on the translated sentences to determine intra-sentence breaks for the translated sentences. The clips for the translated sentences are defined to match the clips already defined in step 235 for the original spoken language text file. For example, if a sentence in the original language is split into two clips, the translation intra-sentence model aims to generate two sentence parts from the translated sentence to define as many clips as there are sentences in the original language. The translation subtitles may be generated based on clips defined for the translation such that each clip text in the translation subtitle and the corresponding clip text in the original language subtitle share the same (or substantially the same) clip duration (250). When playing the video, the clip texts of the translation subtitles are displayed sequentially as subtitles during their respective clip durations (255).

[0127] Figure 3 FIG. 3 presents a platform user interface where subtitles are combined with video clips and the ability to simultaneously edit the video clips and clip captions is provided. The user interface ("UI") 300 may be deployed as part of a video editing software or application and may include a video playback screen 305 where a selected video clip or a complete video may be played along with subtitles linked to the video. The UI 300 may also include a clip editor panel 310 for one or more selected clips, which presents editable manually entered captions in field 315. The clip editor panel 310 may also include video cuts 320, 325 for each selected clip. Each video cut is a segment of the video clip and is identified by a subtitle / caption word, punctuation, sound, or pause connected to that portion of the clip. This allows editing of each position or segment of the clip by applying edits to the displayed word, punctuation, sound, or pause, which then instantly applies the same edit to the corresponding segment in the clip. This makes editing video and removing pauses, silences, or certain undesirable portions much more efficient.

[0128] Thus, each identified word, punctuation mark, or pause that represents a video cut may be individually selected, edited, deleted, or manipulated, which directly affects the portion of the clip that corresponds to the word that also represents the video cut. For example, if a clip consists of a speaker speaking the phrase "I have a dream," when the video cut identified with the word "I" is deleted, its associated linked video and audio are also automatically deleted, leaving only the subtitle / caption text, video, and audio of "have a dream" in the clip. After the deletion has occurred, when the video is played, only the "have a dream" portion is played.

[0129] If there are silences or pauses identified and / or marked by punctuation or pauses in the linked text file, these may also be displayed as individual video cuts 320, allowing their removal and deletion to easily and automatically remove the corresponding video and audio portions (i.e., pauses or silences) in the clip. The UI 300 may also include a button to quickly remove all identified pauses, stops, silences, selected words or punctuation, and / or other undesirable sounds from one or more clips, or from the entire video, with one click. Thus, the UI 300 allows the editing and removal of undesirable portions or video cuts 320 from a video clip to be much easier, due to the link between the subtitle text and the corresponding video / audio portions.

[0130] Figure 4 4 shows a diagrammatic representation of a video 410 and its corresponding script text 415 (STT text) displayed along a timeline. The script text 415 is presented as blocks of words, with each word block representing the duration of the corresponding word in the video. The speech of word 1 occurs at t 10 Starts with t 11 In an embodiment, the script text 415 is obtained from the STT processing of the audio of the video, and word 1 is the audio of the video. 10 From 11 In the embodiment, the timestamp (time code) of word 1 is obtained by transcribing it into t 10 and t 11 and can be included in the subtitle data of the video. 11 From 21are separated by silence until . In an embodiment, the sentence model processes the script text 415 to calculate a sentence-ending probability (or a sentence-beginning probability) of one or more words in the text. Word 3 is determined to be the ending word of sentence 1 based on its sentence-ending probability. In an embodiment, the sentence model identifies or determines a punctuation mark that ends a sentence in the text 415. For example, in FIG. 4, the sentence model determines a period as the ending punctuation mark for sentence 1 and a question mark for sentence 3.

[0131] Figure 5 5 shows a diagrammatic representation of a clip 510 defined for a video 410 and script text 415. In an embodiment, the intra-sentence model processes sentences in the text 415 to detect one or more intra-sentence breaks. For example, in FIG. 4, the intra-sentence model determines that a comma is used as an intra-sentence break in sentence 2. Based on the end-of-sentence punctuation and the intra-sentence breaks, sentence 1 forms the clip text for sling clip 511. Sling clip 511 is a clip that begins at its start time t 10 (start time code of the first word 1) and end time t 32 (the end time code of the last word 3) and the time code t 10 and t 32 is included in the subtitle data of the video to indicate the timing of presenting words 1 to 3. In the embodiment, the time code t 21 is not removed from the subtitle data if it does not affect the presentation of words 1-3 as subtitles. 52 (the end of word 5 just before the identified comma) and time code t 61 (the start of word 6 immediately following the identified comma) is used to split sentence 2 into two clips 512, 513, with words 4-5 forming the clip text of clip 512 and words 6-9 forming the clip text of clip 513.

[0132] Figure 6 FIG. 6 shows a diagrammatic representation of an association between the original spoken language script text 415 and the translated script text 610 to provide the translated language subtitle data. In an embodiment, the translated script text 610 is obtained by translating the original language text 415 sentence by sentence. The translation of sentence 1 itself forms the translated clip text of clip 512. The translation of sentence 2 is split into the same number of clip texts (two) as the original language sentence 2, regardless of the number of intra-sentence breaks identified by running the intra-sentence model (in the translated language) on the translated sentence 2 (e.g., the translated sentence 2 is split into two to match the number of clip texts identified in the original language sentence 2, even if it would be more natural to have no intra-sentence breaks). The clips 511-513 defined based on the original language text 415 are maintained in creating the translated subtitle data, since the translated text 610 does not have its own timecode. In an embodiment, words may be swapped or placed in different clips even if they directly match in the translation text to make sense. For example, word 4 is translated to word tr5 in the translation language, but word tr5 is displayed as a subtitle in clip 513 and word 4 is displayed as a subtitle in clip 512. In an embodiment, clip-ending words in the original and translated languages ​​do not match in meaning, e.g., word 3 and word tr3, which end in the same clip 511, do not match in meaning because the translation of word 3 is word tr2 and not word tr3.

[0133] Figures 7A and 7B 7A and 7B present an embodiment of a method for detecting sentence endings, intra-sentence breaks, and optional translation into another language to create subtitles in one or more languages ​​for a video clip. In this embodiment of a method 700 for creating subtitles from a video, a video is received by a system or platform (705), and may optionally be received along with one or more associated text files (710). These text files may include a transcription or pre-made subtitles associated with the video. One or more text files may additionally or alternatively be generated by the present technology (715) via the AI ​​models or audio extraction or transcription methods described herein. A sentence model may be run on the text file, taking in a certain maximum input size of text to run efficiently. If the word's probability value meets or exceeds a threshold (which may be pre-defined or determined by the sentence model or other methods), the word is marked or identified as a sentence ending word (725), and the location of the identified sentence ending word is marked in the text file and / or the corresponding video and associated audio file (730). The marking may be done in any suitable manner, such as adding a timestamp, manipulating or modifying metadata or any associated text, audio or video files, or any associated video editing files. After the locations of the sentence-final words have been established, the video is divided into several clips (735). Each clip encompasses only one complete sentence. Each clip ends with a sentence-final word. A second AI model, an intra-sentence model, may then be run in a similar manner on the sentences / clips generated by the sentence model to further refine the generated clips by identifying intra-sentence breaks (740).

[0134] Intra-sentence breaks are identified by the comma or punctuation probability of words in the text file, i.e. the probability that a word precedes a punctuation mark that may split a sentence. In an embodiment, a word may have different punctuation probabilities for each space adjacent to it, i.e. there may be different punctuation probabilities for both sides of a word. For example, a particular punctuation mark may have a punctuation probability of 50% immediately to the left of a word and a punctuation probability of 85% to the right. A side of a word may also include a space between the word and the punctuation mark that follows it, either before or after the space. In an embodiment, the punctuation probability of only one side of a word is considered, which may vary depending on the language of the text. For example, in English the space immediately to the right of a word is the space that follows the word and is generally considered. If the punctuation probability of a word meets or exceeds a threshold (this threshold may be predefined or determined by an AI model or any other method), then the word is identified as adjacent to a punctuation mark (745) or the particular space is identified as intra-sentence punctuation or a comma (745).

[0135] Once the sentence breaks are identified (745), their locations may be marked (750) in the text file and / or the corresponding video locations. Marking may be done in any suitable manner, such as adding a timestamp, manipulating or modifying metadata or any associated data in the text, audio or video file, or any associated video editing files. Alternative or additional optional AI refinement models may be deployed (755) on the identified sentences and / or sentence parts in the text file to identify sentence endings and sentence breaks using other factors. These factors may include, but are not limited to, one or more of pauses or silences in the clip or audio, specific phrases, specific words, or clip length, sentence length, or other user configurations and settings. Optionally, the original text file may be translated (760) into another language. Translation (760) into the translation language is performed sentence by sentence, with the sentences identified by the sentence model. The translated sentences may then be input (765) into an intra-sentence translation AI to identify sentence breaks that match the defined clip. The translation intra-sentence model then outputs translated captions that match the defined clip structure (770), and the locations of these sentences / sentence parts can then be marked (775) in the translation text file and / or video file to match them with the corresponding video clips (780). The translated captions are displayed (785) along with subtitles in the original language of each video clip. Some or all of the generated captions, in any or all languages, can then be displayed on a user interface (790), and clip captions can be combined with the corresponding video clips and edited together on the user interface. For an example of a user interface, see FIG. 3 of this application.

[0136] Figure 8 - Example of a User Computing System Architecture 8 illustrates the architecture of an exemplary computing device 800 that can be used to perform one or more features of the present technology. The overall architecture of the computing device 800 includes an arrangement of computer hardware and software modules that can be used to implement one or more aspects of the present disclosure. The computing device 800 may include more (or less) elements than those illustrated in FIG. 8. However, it is not necessary to illustrate all of these elements to provide a useful disclosure.

[0137] The exemplary computing device 800 includes a processor 810, a network interface 820, a computer-readable medium 830, and an input / output device interface 840, all of which may communicate with each other via a communication bus. The network interface 820 may provide connectivity to one or more networks or computing systems. The processor 810 may also communicate with a memory 850 and may further provide output information to one or more output devices, such as a display (e.g., a display 841), a speaker, etc., via the input / output device interface 840. The input / output device interface 840 may also accept input from one or more input devices, such as a camera 842 (e.g., a 3D depth camera), a keyboard, a mouse, a digital pen, a microphone, a touch screen, a gesture recognition system, a voice recognition system, an accelerometer, a gyroscope, etc.

[0138] Memory 850 may contain computer program instructions (which in some implementations may be grouped as modules) that are executed by processor 810 to implement one or more aspects of the present disclosure. Memory 850 may include RAM, ROM, and / or other persistent, secondary, non-transitory computer-readable media.

[0139] Memory 850 may store an operating system 851 that provides computer program instructions for use by processor 810 in the general management and operation of computing device 800. Memory 850 may further include computer program instructions and other information for implementing one or more aspects of the present disclosure.

[0140] In one embodiment, for example, memory 850 includes a user interface module 852 that generates a user interface (and / or instructions therefor) for display via, for example, a browser or an application installed on computing device 800. In addition to and / or in combination with user interface module 852, memory 850 may include a video processing module 853, a text processing module 854, and a machine training model 854, which may be executed by processor 810.

[0141] Although the example of FIG. 8 shows a single processor, a single network interface, a single computer-readable medium, a single input / output device interface, a single memory, a single camera, and a single display, in other implementations computing device 1500 may have a plurality of one or more of these components (e.g., two or more processors and / or two or more memories).

[0142] Processing Using a Remote Computing Device In an embodiment, one or more processes of the present technology may be performed by the exemplary computing device 800, by a remote server, or by a combination of the exemplary computing device 800 and a remote server. For example, if a smartphone does not have a machine-trained model in a local data store, it may communicate with a remote computing server or a cloud computing system to perform one or more processes of the present technology.

[0143] Computer Executable Instructions A logical block, module, or unit described in connection with the implementations disclosed herein may be implemented or executed by a computing device having at least one processor, at least one memory, and at least one communication interface. Elements of a method, process, or algorithm described in connection with the implementations disclosed herein may be embodied directly in hardware, in a software module executed by at least one processor, or in a combination of the two. Computer-executable instructions for implementing a method, process, or algorithm described in connection with the implementations disclosed herein may be stored in a non-transitory computer-readable storage medium.

[0144] Alternative implementations and obvious modifications Although the implementations of the present invention have been disclosed in connection with specific implementations and examples, it will be understood by those skilled in the art that the present invention extends beyond the specifically disclosed implementations to other alternative implementations and / or uses of the invention, as well as obvious modifications and equivalents thereof. In addition, while a number of variations of the present invention have been shown and described in detail, other modifications that are within the scope of the present invention will be readily apparent to those skilled in the art based on this disclosure. It is also contemplated that various combinations or subcombinations of the specific features and aspects of the implementations may be made within the scope of one or more of the inventions. Thus, it should be understood that various features and aspects of the disclosed implementations may be combined with or substituted for one another to form various modes of the disclosed invention. Thus, the scope of the invention disclosed herein should not be limited by the specific disclosed implementations described above, and it is contemplated that various changes in form and detail may be made without departing from the spirit and scope of the present disclosure as set forth in the claims.

Claims

1. 1. A method of providing subtitles to a video, comprising: processing audio data of the video to generate a timed script in a first language including a first sequence of words and a timestamp for each word in the first sequence of words; processing the first sequence of words to calculate a sentence-ending probability for at least one word in the first sequence of words using a first machine-trained model; determining the first word of the first sequence as a first sentence-final word based on the sentence-final probability of a first word and defining a first sentence that ends with the first word; processing the first sentence to calculate an intra-sentence segmentation probability for at least one word of the first sentence using a second machine-trained model; determining the second word of the first sentence as a clip end word based on the intra-sentence break probability of a second word, defining a first clip text that ends with the second word, the definition of the first clip text further defining a first clip period that corresponds to the first clip text and ends when the second word is spoken in the video; generating subtitle data in a first language, the subtitle data including the first clip text and information indicative of the first clip period during which the first clip text is displayed as a subtitle in the first language; The method further comprises: translating the first sentence into a first translation sentence in a second language, the first translation sentence ending with a first translation word; processing the first translated sentence to calculate an intra-sentence break probability for at least one word of the first translated sentence using a third machine-trained model; determining the second translated word of the first translated sentence as a clip end word based on the intra-sentence break probability of the second translated word, thereby defining a first translated clip text that ends with the second translated word; generating second language subtitle data including the first translated clip text and information indicating a second language period during which the first translated clip text is displayed as a subtitle in the second language; The method, wherein the first clip duration for displaying the first clip text is the same or substantially the same as the second language duration for displaying the first translated clip text, regardless of whether the second word that ends the first clip text semantically corresponds to the second translation word that ends the first translated clip text.

2. the timed script includes timestamps for the second words indicating times when the second words were spoken in the video; 2. The method of claim 1 , wherein generating the second language subtitle data includes designating, in the second language subtitle data, a timestamp of the second word as an end of the first clip period and the second language period according to a predefined subtitle format.

3. the first clip text begins with a third word of the first sentence, and the first translated clip text begins with a third translation word of the first translated sentence; the timed script includes a timestamp for the third word indicating a time when the sound of the third word begins in the video; 3. The method of claim 2, wherein generating the second language subtitle data further comprises designating, in the second language subtitle data, a timestamp of the third word as a start of the first clip period and the second language period, wherein the first clip period and the second language period are the same regardless of whether the third word semantically corresponds to the third translation word.

4. the first translation does not include punctuation marks indicating intra-sentence breaks in the first translation; 4. The method of claim 1, wherein the third machine-trained model is trained using a plurality of punctuated sentences in the second language and is configured to calculate a probability that, for at least one word in an input sentence, it is immediately followed by at least one sentence-breaking punctuation mark.

5. The first sentence is divided into "n" clip texts based on at least the first word and the second word, where "n" is a natural number greater than "2"; The method of claim 1 , wherein the first translation is divided into "n" equal translated clip texts.

6. determining a third word of the first sentence as a clip end word, wherein a second clip text is defined that starts with a word immediately following the second word and ends with the third word; the second clip text definition further defines a second clip corresponding to the second clip text and ending when the third word is spoken in the video; 6. The method of claim 5, wherein the third word in the first sentence is identified as a clip end word based on at least one of the intra-sentence break probability of the third word and the length of silence following the sound of the third word in the video.

7. the first language subtitle data is configured such that the entire first clip text is displayed as a subtitle for the video at the start of the first clip period and is maintained uninterrupted until the end of the first clip period; 4. The method of claim 1, wherein the second language subtitle data is configured such that the entire first translated clip text is displayed as subtitles for the video at the start of the second language period and maintained uninterrupted until the end of the second language period.

8. 4. The method of claim 1, wherein generating the second language subtitle data includes designating a time when the second word is spoken in the first language as an end of the second language period such that the first clip period and the second language period end simultaneously.

9. The method of claim 1 , wherein the first translated clip text in the second language does not correspond in meaning to the first clip text in the first language.

10. The method of claim 1 , wherein the first translated clip text is a translation of the first clip text.

11. Processing the audio data of the video to generate the timed script includes performing a speech-to-text translation (STT) process of the audio data, where audio corresponding to the second word is transcribed into the second word, a time in the video at which the second word was spoken is determined, and the time at which the second word was spoken is assigned to the timed script for the second word; the information indicative of the first clip duration includes a time at which the second word was spoken as determined by the STT process; 4. The method of claim 1, wherein generating the first language subtitle data comprises associating with the first clip text the time at which the second word was spoken, as determined by the STT process, as an end of the first clip period, in accordance with a predefined subtitle file format.

12. Processing the audio data of the video to generate the timed script includes: identifying silences and non-silences within the audio data, the non-silences including corresponding sounds of the second word; transcribing corresponding sounds of the second word into the second word to obtain the first word sequence; determining, for the second word, an end time at which a corresponding sound of the second word ends in the video; and including the determined end time in the timed script as a timestamp of the second word.

13. Processing the audio data of the video to generate the timed script includes: obtaining a pre-written script of the video that includes the first word sequence but does not include a timestamp for the first word sequence; identifying, for each word of the first sequence of words, a corresponding sound in the audio data that identifies a first sound that corresponds to the second word; determining an end time of the first sound when the first sound ends within the video; and combining the determined end time with the first sequence of words to generate the timed script such that the determined end time is specified as a timestamp of the second word in the timed script.

14. 1. A method of providing subtitles to a video, comprising: processing audio data of the video to generate a timed script in a first language including a first sequence of words and a timestamp for each word in the first sequence of words; processing the first sequence of words to calculate a sentence-ending probability for at least one word in the first sequence of words using a first machine-trained model; determining the first word of the first sequence as a first sentence-final word based on the sentence-final probability of a first word and defining a first sentence that ends with the first word; processing the first sentence to calculate an intra-sentence segmentation probability for at least one word of the first sentence using a second machine-trained model; determining a second word of the first sentence as a clip end word based on the intra-sentence break probability of the second word, defining a first clip text that ends with the second word, the definition of the first clip text further defining a first clip period that corresponds to the first clip text and ends when the second word is spoken in the video; generating subtitle data in a first language including the first clip text and information indicative of the first clip period during which the first clip text is displayed as subtitles in the first language.

15. the timed script includes timestamps for the second words indicating times when the second words were spoken in the video; 15. The method of claim 14, wherein generating the first language subtitle data includes designating a timestamp of the second word as an end of the first clip period according to a predefined subtitle format.

16. the first clip text begins with the third word of the first sentence; the timed script includes a timestamp for the third word indicating a time when the sound of the third word begins in the video; 16. The method of claim 15, wherein generating the first language subtitle data includes designating a timestamp of the third word as a start of the first clip period in accordance with the predetermined subtitle format.

17. 17. The method of claim 16, wherein the first language subtitle data is configured such that the entire first clip text is displayed as a subtitle for the video at the beginning of the first clip period and remains uninterrupted until the end of the first clip period.

18. 17. The method of claim 16, wherein the first clip text further includes a fourth word between the third word and the second word, and the first language subtitle data does not include a timestamp of the fourth word, such that the first clip text is displayed as subtitles without reference to the fourth word.

19. the timed script does not include a punctuation mark indicating the end of the first sentence or a sentence break in the first sentence; the first machine-trained model is trained using a plurality of punctuated texts, each of the punctuated texts including one or more final punctuation marks, and is configured to calculate, for at least one word in an input text, a probability that it is immediately followed by at least one final punctuation mark; 19. The method of any one of claims 14 to 18, wherein the second machine-trained model is trained using a plurality of punctuated texts, each of which includes one or more sentence-breaking punctuation marks, and is configured to calculate a probability that, for at least one word in an input sentence, it is immediately followed by at least one sentence-breaking punctuation mark.

20. 19. The method of claim 14, further comprising determining a third word of the first sentence as another clip end word based on the intra-sentence break probability of the third word, thereby defining a second clip text starting with a word immediately after the second word and ending with the third word, the definition of the second clip text further defining a second clip corresponding to the second clip text and ending at the point when the third word is spoken in the video.

21. 19. The method of claim 14, further comprising determining a third word of the first sentence as another clip end word based on the intra-sentence break probability of the third word, thereby defining a second clip text starting with a word immediately after the second word and ending with the third word, the definition of the second clip text further defining a second clip corresponding to the second clip text and ending at the point when the third word is spoken in the video.

22. Processing the audio data of the video to generate the timed script includes performing a speech-to-text translation (STT) process of the audio data, where audio corresponding to the second word is transcribed into the second word, a time in the video at which the second word was spoken is determined, and the time at which the second word was spoken is assigned to the timed script for the second word; the information indicative of the first clip duration includes a time at which the second word was spoken as determined by the STT process; 19. The method of claim 14, wherein generating the first language subtitle data comprises associating with the first clip text the time at which the second word was spoken, as determined by the STT process, as an end of the first clip period, in accordance with a predefined subtitle file format.

23. Processing the audio data of the video to generate the timed script includes: identifying silences and non-silences within the audio data, the non-silences including corresponding sounds of the second word; transcribing corresponding sounds of the second word into the second word to obtain the first word sequence; determining, for the second word, an end time at which a corresponding sound of the second word ends in the video; and including the determined end time in the timed script as a timestamp of the second word.

24. Processing the audio data of the video to generate the timed script includes: obtaining a pre-written script of the video that includes the first sequence of words but does not include a timestamp for the first sequence of words; identifying, for each word of the first sequence of words, a corresponding sound in the audio data that identifies a first sound that corresponds to the second word; determining an end time of the first sound when the first sound ends within the video; and combining the determined end time with the first sequence of words to generate the timed script such that the determined end time is specified as a timestamp of the second words.

25. 19. A method according to any one of claims 14 to 18, wherein the information indicating the first clip duration includes a first timestamp indicating a start time of the first clip in the video and further includes a second timestamp indicating an end time of the first clip in the video, and the first clip text is displayed together with the video without interruption from the start time of the first clip to the end time of the first clip.

26. 19. The method of claim 14, wherein the timestamps for each word of the first sequence of words define a time when the sound of the corresponding word ends or starts in the video.

27. 19. The method of claim 14, wherein the timed script further comprises a second sequence of words, the timed script further comprising a timestamp for at least one word of the second sequence of words.

28. A non-transitory computer readable medium storing instructions that, when executed by at least one processor, cause the at least one processor to perform the method of any one of the preceding claims.

29. 1. A system for providing closed captioning to a video, comprising: At least one processor; At least one memory, said memory storing instructions that, when executed by said at least one processor, cause said at least one processor to perform a method according to any one of the preceding claims.

Citation Information

Patent Citations

  • Generation system and retrieving method for video contents file

    JP2002374494A

  • Scene meta information generation apparatus and scene meta information generating method

    JP2019198074A

  • System and method for simultaneous multilingual dubbing of video-audio programs

    US20200211565A1