Subtitle generation method and device, electronic equipment, storage medium and program product
By performing speech recognition and grammatical analysis on video and audio data, and combining audio segment information to segment and merge text segments, subtitles that meet the requirements are generated. This solves the problem of insufficient control over subtitle length and display duration in existing technologies, and improves the subtitle's ability to aid understanding.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- BEIJING ZITIAO NETWORK TECH CO LTD
- Filing Date
- 2022-05-31
- Publication Date
- 2026-08-04
AI Technical Summary
Existing technologies cannot effectively control the length and display duration of individual subtitles when generating subtitles for videos, resulting in a reduced subjective experience of the subtitles and affecting the user's comprehension.
By extracting video audio data for speech recognition, and combining grammatical analysis with the pronunciation information and timestamp information of audio segments, text segments are segmented and merged to generate subtitle data that meets the preset single subtitle length requirements.
It enables effective control over the sentence length and display duration of individual subtitles, improving the subtitles' ability to aid understanding and reducing ambiguity.
Smart Images

Figure CN117201876B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of multimedia technology, and in particular to a method, apparatus, electronic device, storage medium, and program product for generating subtitles. Background Technology
[0002] Subtitles are textual content displayed on video frames, generated based on dialogue, explanatory information, and other data. Because subtitles help users understand video content, generating subtitles for videos is extremely important.
[0003] Currently, video subtitle generation typically involves extracting audio from the video after video production, performing speech recognition on the extracted audio to obtain the corresponding text, then restoring punctuation to obtain text segments, and finally displaying these text segments in the corresponding video frames according to their corresponding time. This method cannot effectively control the length of text segments (i.e., individual subtitles), significantly reducing the subjective experience of subtitle delivery. Summary of the Invention
[0004] To address the aforementioned technical problems, this disclosure provides a subtitle generation method, apparatus, electronic device, storage medium, and program product.
[0005] In a first aspect, embodiments of this disclosure provide a subtitle generation method, including:
[0006] Extract audio data from the video to be processed, perform speech recognition on the audio data, and obtain the text data corresponding to the audio data;
[0007] The text data is obtained by acquiring multiple segmentation positions determined based on grammatical analysis, as well as the pronunciation object information and timestamp information of the audio segments corresponding to each character in the text data;
[0008] Based on the multiple segmentation positions, the pronunciation object information and timestamp information of the audio segments corresponding to each character, the text data is segmented into multiple text segments; the audio segments corresponding to each character in the text segment belong to the same pronunciation object, and the duration of the blank segments in the audio segments corresponding to the text segment is less than a preset duration;
[0009] Based on the semantics of each text segment and the timestamp information of the audio segment corresponding to each character, the multiple text segments are merged to obtain multiple merged segments that are semantically coherent and meet the preset single subtitle sentence length requirements;
[0010] Based on the multiple merged segments, subtitle data corresponding to the video to be processed is generated.
[0011] In some embodiments, the merging based on the semantics of each text segment and the timestamp information of the corresponding audio segment includes:
[0012] Whether adjacent text segments can be merged is determined based on whether the preset single subtitle length requirement is met after merging the adjacent text segments;
[0013] Whether adjacent text segments can be merged is determined based on whether the semantics of the adjacent text segments are grammatically correct after merging.
[0014] If the text segment can be merged with both preceding and following text segments, then the two adjacent text segments with shorter pause durations between audio segments will be merged.
[0015] In some embodiments, the preset single subtitle length requirement includes: characters per second (CPS) requirement and / or the maximum display duration requirement for a single subtitle in the video.
[0016] In some embodiments, segmenting the text data into multiple text segments based on the plurality of segmentation positions, the pronunciation object information of the audio segment corresponding to each character, and the timestamp information includes:
[0017] The text data is input into the text processing module to obtain the multiple text fragments output by the text processing module;
[0018] The text processing module includes: a sub-module for segmentation based on the multiple segmentation positions, a sub-module for text segmentation based on the pronunciation object information of the audio segments corresponding to each character, and a sub-module for text segmentation based on the timestamp information of the audio segments corresponding to each character.
[0019] In some embodiments, the subtitle data is a text-formatted subtitle (SRT) file.
[0020] In some embodiments, the method further includes: fusing the subtitle data with the video to be processed to obtain a target video with subtitles.
[0021] Secondly, embodiments of this disclosure provide a subtitle generation apparatus, including:
[0022] An audio processing module is used to extract audio data from the video to be processed, perform speech recognition on the audio data, and obtain the text data corresponding to the audio data.
[0023] The acquisition module is used to acquire multiple segmentation positions of the text data determined based on grammatical analysis, as well as the pronunciation object information and timestamp information of the audio segments corresponding to each character in the text data;
[0024] The text segmentation module is used to segment the text data into multiple text segments based on the multiple segmentation positions, the pronunciation object information and timestamp information of the audio segments corresponding to each character; the audio segments corresponding to each character in the text segments belong to the same pronunciation object, and the duration of the blank segments in the audio segments corresponding to the text segments is less than a preset duration;
[0025] The merging module is used to merge the multiple text segments according to the semantics of each text segment and the timestamp information of the audio segment corresponding to each character, so as to obtain multiple merged segments that are semantically fluent and meet the preset single subtitle sentence length requirements.
[0026] The generation module is used to generate subtitle data corresponding to the video to be processed based on the multiple merged segments.
[0027] Thirdly, embodiments of this disclosure also provide an electronic device, including: a memory and a processor; the memory is configured to store computer program instructions; the processor is configured to execute the computer program instructions, causing the electronic device to implement the subtitle generation method as described in the first aspect and any one of the first aspects.
[0028] Fourthly, embodiments of this disclosure also provide a readable storage medium, including: computer program instructions; the computer program instructions are executed by an electronic device to cause the electronic device to implement the subtitle generation method as described in the first aspect and any one of the first aspects.
[0029] Fifthly, embodiments of this disclosure also provide a computer program product, including computer program instructions that, when executed by an electronic device, implement the subtitle generation method as described in the first aspect and any one of the first aspects.
[0030] This disclosure provides a subtitle generation method, apparatus, electronic device, storage medium, and program product. The method involves extracting audio from a video to be processed and performing speech recognition on the extracted audio data to obtain text data corresponding to the audio data. Multiple segmentation positions are determined based on grammatical analysis of the text data, along with the pronunciation object information and timestamp information of the audio segments corresponding to each character in the text data. Based on the multiple segmentation positions, the pronunciation object information, and the timestamp information of the audio segments corresponding to each character, the text data is segmented to obtain multiple text segments that meet the requirements. Then, based on the semantics of each text segment and the timestamp information of the audio segments corresponding to each character, the multiple text segments are merged to obtain multiple merged segments that are semantically fluent and meet the preset single subtitle sentence length requirements. Subtitle data corresponding to the video to be processed is generated based on the multiple merged segments. This method, by combining text and audio dimension features for segmentation and merging, can better control the sentence length of a single subtitle and the display duration of a single subtitle in the video, significantly improving the subtitle's comprehension assistance effect. Attached Figure Description
[0031] The accompanying drawings, which are incorporated in and form a part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure.
[0032] To more clearly illustrate the technical solutions in the embodiments of this disclosure or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, those skilled in the art can obtain other drawings based on these drawings without creative effort.
[0033] Figure 1 This is a flowchart of the subtitle generation method described in the embodiments of this disclosure;
[0034] Figure 2 A flowchart illustrating a subtitle generation method provided in an embodiment of this disclosure;
[0035] Figure 3 A flowchart illustrating a subtitle generation method provided in an embodiment of this disclosure;
[0036] Figure 4 A flowchart illustrating a subtitle generation method provided in another embodiment of this disclosure;
[0037] Figure 5 A flowchart illustrating a subtitle generation method provided in another embodiment of this disclosure;
[0038] Figure 6 This is a schematic diagram of the structure of a subtitle generation device provided in an embodiment of the present disclosure;
[0039] Figure 7 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present disclosure. Detailed Implementation
[0040] To better understand the above-mentioned objectives, features, and advantages of this disclosure, the solutions disclosed herein will be further described below. It should be noted that, unless otherwise specified, the embodiments and features described herein can be combined with each other.
[0041] Numerous specific details are set forth in the following description in order to provide a full understanding of this disclosure, but this disclosure may also be implemented in other ways different from those described herein; obviously, the embodiments in the specification are only some, and not all, of the embodiments of this disclosure.
[0042] Currently, generating subtitles for videos typically involves the following process: extracting audio from the video, performing speech recognition on the audio data to obtain the corresponding text data, restoring punctuation to the text data to obtain segmented text fragments, generating subtitle data based on the time of the corresponding video segments, and then fusing the subtitle data with the video to obtain a video with subtitles. This method, however, relies heavily on punctuation restoration during text data fragmentation, making it difficult to control the sentence length of individual subtitles. This affects subtitle layout and display duration in the video, reducing the subjective experience of using subtitles and failing to provide effective comprehension assistance.
[0043] For example, if a single subtitle is long, meaning it contains a large number of characters, and the screen size of an electronic device is limited, the subtitle needs to be displayed in multiple lines. However, when the subtitle occupies many lines, the area it occupies will expand, potentially obscuring more of the video and affecting the user's viewing experience. In addition, a longer subtitle will increase its display time in the video, which will also affect the user's viewing experience.
[0044] For example, some short sentences are spoken at a fast pace, and the sentence length of a single subtitle is short, meaning that the number of characters contained in a single subtitle is small, but the pronunciation duration of each character is short. Therefore, the display time of the subtitle in the video is short, and users may not have enough time to read the subtitle content in detail, thus failing to achieve the purpose of subtitles in assisting understanding.
[0045] For example, the same text may express different meanings depending on the length of the pauses. Subtitles obtained by restoring punctuation may not accurately express the meaning of the same text at different audio positions.
[0046] Based on this, this disclosure provides a subtitle generation method. The method involves extracting audio data from the video to be processed and performing speech recognition on the audio data to obtain corresponding text data. Multiple segmentation positions are determined based on grammatical analysis of the text data, along with the pronunciation object information and timestamp information of the audio segments corresponding to each character in the text data. Based on these segmentation positions, the pronunciation object information, and the timestamp information of the audio segments corresponding to each character, the text data is segmented into multiple text segments that meet the requirements. Then, according to the semantics of each text segment and the timestamp information of the audio segments corresponding to each character, the multiple text segments are merged to obtain multiple merged segments that are semantically fluent and meet the preset single subtitle sentence length requirements. Subtitle data corresponding to the video to be processed is generated based on the multiple merged segments. This method, by combining text and audio dimension features for segmentation and merging, can better control the sentence length of a single subtitle and the display duration of a single subtitle in the video, significantly improving the subtitle's comprehension aid effect. Furthermore, the method fully considers the blank time between audio segments corresponding to characters during the merging and segmentation process, so that the same speech content expressing different meanings is segmented and merged in different ways. Therefore, this method can also effectively reduce the occurrence of ambiguity.
[0047] For example, the subtitle generation method provided in this embodiment can be executed by an electronic device. The electronic device can be a tablet computer, mobile phone (such as a foldable phone, a large-screen phone, etc.), wearable device, in-vehicle device, augmented reality (AR) / virtual reality (VR) device, laptop computer, ultra-mobile personal computer (UMPC), netbook, personal digital assistant (PDA), smart TV, smart screen, high-definition TV, 4K TV, smart speaker, smart projector, and other Internet of Things (IoT) devices. This disclosure does not limit the specific type of electronic device. Furthermore, this disclosure does not limit the type of operating system of the electronic device. For example, Android system, Linux system, Windows system, iOS system, etc.
[0048] Based on the foregoing description, this disclosure will use electronic devices as examples, combined with accompanying drawings and application scenarios, to elaborate in detail on the subtitle generation method provided by this disclosure.
[0049] Figure 1 This is a flowchart illustrating a subtitle generation method according to an embodiment of this disclosure. Please refer to [link / reference]. Figure 1 As shown, the method in this embodiment includes:
[0050] S101. Extract audio data from the video to be processed, and perform speech recognition on the audio data to obtain the text data corresponding to the audio data.
[0051] The video to be processed is the video to which subtitles are to be added. The electronic device can acquire the video to be processed. The video to be processed can be recorded by the user using the electronic device, downloaded from the internet, or created by the user using video processing software. This disclosure does not limit the method of acquiring the video to be processed. Furthermore, this disclosure does not limit the video content, duration, storage format, resolution, or other parameters of the video to be processed.
[0052] Electronic devices can extract audio data from videos and convert it into text data. For example, an electronic device can convert audio data into text data using a speech recognition model. This disclosure does not limit the parameters of the speech recognition model; for example, the speech recognition model can be a deep neural network module, a convolutional neural network model, etc. Alternatively, the electronic device can also utilize other existing speech recognition tools or methods to convert audio data into text data. This disclosure does not limit the implementation method of speech recognition in electronic devices.
[0053] The text data can include continuous sequences of characters. For example, the text data could include "Today I was happy that I went to the amusement park with my parents," without punctuation. It should be noted that since audio data can correspond to one or more languages, the generated text data can also include characters corresponding to one or more languages.
[0054] Of course, during speech recognition, it's also advisable to convert the audio into a language to facilitate subsequent segmentation. For example, the speech recognition result for an audio segment could be "Hello," or it could be "hello." Since the proportion of Chinese characters in the overall text data is relatively high, the former can be chosen if the goal is to improve the consistency of language types in the subtitles, while the latter can be chosen if the goal is to increase the fun of the subtitles.
[0055] S102. Obtain multiple segmentation positions of the text data determined based on grammatical analysis, as well as the pronunciation object information and timestamp information of the audio segments corresponding to each character in the text data.
[0056] Electronic devices can analyze text data using syntactic analysis models to obtain multiple segmentation points. Syntactic analysis can include punctuation position analysis, syntactic feature analysis, etc. Through syntactic analysis, multiple clause positions can be obtained, which are the segmentation points.
[0057] Electronic devices can identify the audio segments corresponding to different audio segments by performing sound object recognition on audio data. Then, by combining the correspondence between the audio segments corresponding to different audio segments and text data, they can obtain the sound object information of the audio segment corresponding to each character.
[0058] Electronic devices can segment audio data to obtain the timestamp information of the audio segment corresponding to each character. The timestamp information can include the start time and the end time.
[0059] S103. Based on multiple segmentation positions, the pronunciation object information and timestamp information of the audio segments corresponding to each character, the text data is segmented to obtain multiple text segments.
[0060] In this process, the audio segments corresponding to each character in the segmented text belong to the same pronunciation object, and the duration of the blank segments in the audio segments corresponding to the text segments is less than the preset duration.
[0061] The text data can be segmented into multiple text fragments by a text processing module. The text processing module can include multiple sub-modules. Each sub-module is used to segment the input text data according to one or more of the aforementioned features. After the text data has been processed by multiple sub-modules, it can be divided into multiple first text fragments.
[0062] Among them, after the text data is segmented by the text processing module, the text is then processed through... Figure 2 as well as Figure 3 The illustrated embodiment is provided as an example.
[0063] S104. Based on the semantics of each text segment and the timestamp information of the audio segment corresponding to each character, merge multiple text segments to obtain multiple merged segments that are semantically coherent and meet the preset single subtitle length requirements.
[0064] The semantics of text fragments can be obtained through semantic analysis. Based on the semantics, it is possible to determine whether the content to be expressed by adjacent text fragments is continuous and coherent. This can then serve as a basis for merging text fragments, avoiding the merging of semantically incoherent text fragments and preventing a poor user experience.
[0065] The pause duration between text segments can be determined by using the timestamp information of the audio segments corresponding to each character. Specifically, the pause duration between adjacent text segments can be determined based on the end time of the audio segment corresponding to the last character of the previous text segment and the start time of the audio segment corresponding to the first character of the next text segment. During the merging process, there is a tendency to merge two adjacent text segments with shorter pause durations. A shorter pause duration indicates a stronger continuity of the content to be expressed in the audio data, resulting in a more complete representation of the content in the audio data after merging, thus making it easier for users to understand.
[0066] In addition, during the merging process, it is also necessary to determine whether the merging of multiple text segments meets the preset single subtitle length requirement, thereby controlling the subtitle length and the display duration of the subtitle on the screen.
[0067] By combining the above three aspects to merge text fragments, we can obtain merged fragments that are semantically coherent and meet the preset subtitle sentence length requirements.
[0068] For example, text fragment 1, text fragment 2, and text fragment 3 are three consecutive first text fragments. Based on semantics, it is determined that text fragment 1 and text fragment 2 can be merged, and text fragment 2 and text fragment 3 can be merged. Furthermore, the pause duration between text fragment 1 and text fragment 2 is t1, and the pause duration between text fragment 2 and text fragment 3 is t2. Since t1 is less than t2, merging text fragment 1 and text fragment 2 is more reasonable. Additionally, merging text fragment 1 and text fragment 2 satisfies the preset single subtitle length requirement; therefore, the merging conditions are met. Thus, text fragment 1 and text fragment 2 can be merged.
[0069] It should be noted that the merged segment obtained by merging text segment 1 and text segment 2 may be the merged segment corresponding to the final single subtitle, or it may be necessary to merge the merged segment with the adjacent text segment 3 to obtain the merged segment corresponding to the single subtitle.
[0070] S105. Generate subtitle data corresponding to the video to be processed based on multiple merged segments.
[0071] Each merged segment corresponds to one subtitle. Multiple merged segments are converted into subtitle files of a preset format in sequence to obtain the subtitle data corresponding to the video to be processed.
[0072] The subtitle data can be, but is not limited to, an SRT file.
[0073] The method provided in this embodiment, by combining text and audio dimension features to segment text data and merge the segmented text fragments, can better control the sentence length of a single subtitle and the display duration of a single subtitle in the video, without affecting semantic understanding, thus greatly improving the subtitle's comprehension-aid effect; in addition, this method can also effectively reduce the occurrence of ambiguity.
[0074] Combination Figure 1 As can be seen from the description of the illustrated embodiment, when an electronic device can segment text data through a text processing module (which can also be understood as a text processing model), the connection order of each sub-module in the text processing module can be flexibly set. Figure 2 and Figure 3 Two different methods are illustrated by example.
[0075] Assuming in Figure 2 as well as Figure 3 In the illustrated embodiment, the text processing module includes: a first segmentation module for segmenting text data based on punctuation analysis, a second segmentation module for segmenting text data based on grammatical features, a third segmentation module for segmenting based on the pronunciation object information corresponding to the audio data, and a fourth segmentation module for segmenting based on the timestamp information of the audio segments corresponding to each character in the text data.
[0076] Figure 2 This is a schematic diagram of the structure of a text processing module provided in one embodiment of this disclosure. Please refer to [link / reference]. Figure 2 As shown, the output of the first segmentation module is connected to the input of the second segmentation module, the output of the second segmentation module is connected to the input of the third segmentation module, and the output of the third module is connected to the input of the fourth segmentation module. (Combined...) Figure 2 The structure of the text processing module in the illustrated embodiment can be understood as the various segmentation modules connected in a serial manner.
[0077] The first segmentation module receives text data as input and performs punctuation analysis (which can also be understood as punctuation restoration) to obtain the sentence positions of multiple punctuation marks. Based on these sentence positions, the text data can be segmented into text fragments. The text fragments output by the first segmentation module are input to the second segmentation module, which performs grammatical feature analysis to determine multiple segmentation positions. Based on these multiple segmentation positions, the text fragments from the first segmentation module can be further segmented or adjusted to obtain multiple text fragments. The text fragments output by the second segmentation module and audio data are input to the third segmentation module, which performs pronunciation object recognition on the audio data. The text processing module first determines the start and end positions of the audio segments corresponding to different pronunciation objects. Then, based on the audio segments corresponding to different pronunciation objects, it determines the segmentation positions in the text data. The text segments are then further segmented based on these determined positions, so that each segment corresponds to a single pronunciation object. Next, the fourth segmentation module determines the pause duration of adjacent characters based on the start and end times of the audio segments corresponding to each character. It then compares the pause durations of adjacent characters with a preset duration, grouping adjacent characters with pause durations shorter than the preset duration into one text segment, and segmenting adjacent characters with pause durations greater than or equal to the preset duration into two different text segments. Finally, the last submodule of the text processing module (the fourth segmentation module) outputs multiple text segments, which represent the final segmentation result of the text data.
[0078] This application does not limit the value of the preset duration. For example, it can be 0.4 seconds, 0.5 seconds, 0.6 seconds, etc. The preset duration can be obtained by statistical analysis of the pause duration between audio segments corresponding to each character in a large amount of audio data.
[0079] As one possible implementation, the text processing module includes segmentation modules that can be implemented using corresponding machine learning models. For example, the first segmentation module can be implemented based on a pre-trained punctuation recovery processing model, the second segmentation module can be implemented based on a pre-trained grammatical feature analysis model, the pronunciation object segmentation module can be implemented based on a pre-trained audio processing model, and the pause duration segmentation module can be implemented based on a pre-trained character processing model. This disclosure does not limit the type of machine learning model or model parameters used in each segmentation module.
[0080] Figure 3 This is a schematic diagram of the structure of a text processing module provided in one embodiment of this disclosure. Please refer to [link / reference]. Figure 3As shown, the text processing module includes segmentation modules connected in parallel. The first and second segmentation modules receive raw text data as input, respectively; the third segmentation module receives audio data and raw text data as input; and the fourth segmentation module receives raw text data as input, with each character in the text data carrying timestamp information. Each segmentation module determines the segmentation position based on its respective input to segment the text data. Then, the segmentation results output by each segmentation module are fused to obtain multiple text fragments.
[0081] The processing methods of each segmentation module in the text processing module can be found in the following reference: Figure 2 The description of the illustrated embodiments will not be repeated here for the sake of brevity.
[0082] It should be noted that the connection methods of the various segmentation modules included in the text processing module are not limited to those described above. Figure 2 as well as Figure 3 For example, other methods can also be used to implement this. For instance, serial and parallel connection methods can be combined; for example, the first and second segmentation modules can be connected serially, the third and fourth segmentation modules can be connected serially, and the first and second segmentation modules can be connected as a whole in parallel with the third and fourth segmentation modules as another whole.
[0083] In addition, it should be noted that the connection order of the various segmentation modules included in the text processing module can be flexibly adjusted according to different scenarios. For example, in scenarios with many pronunciation objects, segmentation processing can be performed first based on the pronunciation objects, and then segmentation processing can be performed based on punctuation analysis, grammatical feature analysis, and the timestamp information of the audio segments corresponding to each character.
[0084] Figure 4 This is a flowchart illustrating a subtitle generation method provided in one embodiment of the present disclosure. Figure 4 The illustrated embodiments are primarily used to demonstrate how an electronic device can merge text fragments. Please refer to [link to relevant documentation]. Figure 4 As shown, electronic devices can merge text fragments by calling the merging module. The merging module includes: an indicator module, a semantic analysis module, a pause duration comparison module, and a text splicing module.
[0085] The metrics module can determine whether merging two input text segments meets the preset subtitle length requirement. The preset subtitle length requirement mainly refers to the retention time of a single subtitle in the video. To easily determine whether the generated single subtitle meets the requirement, the preset subtitle requirement can be either a preset maximum number of characters per second (CPS) or a preset maximum display duration of a single subtitle in the video. These two metrics can effectively reflect the retention time of a single subtitle in the video.
[0086] In addition, the semantic analysis module can determine whether the two input text segments can be merged based on their respective semantics, and output identification information indicating whether the text segments can be merged to the text concatenation module. For example, if the semantic analysis module outputs an identifier of 1, it means that the segments can be merged, and if it outputs an identifier of 0, it means that the segments cannot be merged.
[0087] The pause duration comparison module is used to determine the pause duration comparison results between multiple adjacent text segments based on the timestamp information of the audio segments corresponding to each character in the text segment.
[0088] The text splicing module combines the results or indications output by the aforementioned indicator module, semantic analysis module, and pause duration comparison module to determine the merging scheme. It splices together text segments that meet the preset subtitle sentence length requirements, are semantically fluent, and have short pause durations between text segments, thereby obtaining multiple merged segments.
[0089] During implementation, the indicator module and the semantic analysis module can interact with each other. For example, the indicator module can output the judgment result to the semantic analysis module. The semantic analysis module can judge the combination of text segments that meet the preset subtitle sentence length requirements. For the combination of text segments that do not meet the preset subtitle sentence length requirements, the semantic continuity and fluency are not judged, thereby reducing the workload of the semantic analysis module and improving the efficiency of subtitle generation.
[0090] Suppose that after the text data is segmented, it yields N text segments, namely text segment 1, text segment 2 to text segment N.
[0091] For example, the electronic device can sequentially determine whether merging text fragment 1 and text fragment 2, and text fragment 2 and text fragment 3, meets the preset subtitle length requirements. If, based on semantic features, it is determined that text fragment 1 and text fragment 2 can be merged, and text fragment 2 and text fragment 3 can also be merged, but there is a pause duration between text fragment 1 and text fragment 2, then text fragment 1 and text fragment 2 are merged to obtain merged fragment 1. Afterwards, the electronic device can determine whether merged fragment 1 and text fragment 3 can be merged based on the preset subtitle length requirements and the semantics of the text fragments. If they can be merged, then merged fragment 1 and text fragment 3 are merged to obtain a new merged fragment 1. Alternatively, the electronic device can also determine whether text fragment 3 and text fragment 4 can be merged based on the preset subtitle length requirements and the semantics of the text fragments. If they can be merged, then merged fragment 3 and text fragment 4 are merged to obtain merged fragment 2. The electronic device can compare the subtitle effect of the merged fragment obtained by merging the new merged fragment 1 with text fragment 3 with the subtitle effect of the merged fragment obtained by merging text fragment 3 with text fragment 4, and determine the final merging scheme for text fragment 3.
[0092] By analogy, a merging scheme for each text segment can be obtained.
[0093] It should be noted that the process of determining whether merging two text segments meets the preset subtitle length requirement, determining whether merging is possible based on the semantics of the two text segments, and comparing the pause duration between the audio segments corresponding to adjacent text segments can be performed in parallel. Then, the results of the judgments output by the three processes are combined for merging.
[0094] It should also be noted that the above merging process can go through multiple rounds of processing. For example, if the merged segments obtained in the first round of merging are all short, the merged segments obtained in the first round can be used as input to perform another round of merging processing, so that the length of a single subtitle is infinitely close to the preset subtitle length requirement.
[0095] Another possible implementation is that since text fragments 1 to N contain a small number of characters, multiple rounds of merging may be required. In the first to m1th rounds of merging, merging can be carried out based on the preset subtitle length requirements, the semantics of the text fragments, and the pause duration between the corresponding audio fragments. In the subsequent m1+1th to Mth rounds of merging, merging can be carried out based on the preset subtitle length requirements and the semantic features of the text fragments.
[0096] In some cases, electronic devices can also obtain different merging results based on the aforementioned preset subtitle length requirements, the semantics of the text segment, and the pause duration characteristics between the corresponding audio segments. This results in multiple versions of subtitle data. Then, based on the subtitle effects presented by each version, the best-looking subtitle data is selected. Alternatively, multiple versions of subtitle data can be presented to the user, allowing them to preview the effects of each version and select the final version based on their input.
[0097] The method provided in this disclosure allows for the merging of multiple text segments to obtain single subtitles of appropriate sentence length, ensuring that each subtitle has a suitable display duration in the video and improving its comprehension aid effect. For example, the solution provided in this disclosure can divide a single sentence with a large number of characters into multiple sentences, each presented by a separate single subtitle, avoiding the problems of long single subtitles, disorganized multi-line display, and long display time. For short sentences with a fast speaking speed, the characters corresponding to the short sentence can be combined with the characters of adjacent sentences, increasing the retention time of the subtitle corresponding to the short sentence in the video and ensuring that users have enough time to clearly see the content of the subtitle. Furthermore, the method provided in this disclosure determines which text segment to merge with a text segment that has stronger content continuity by measuring the pause duration between the audio segments corresponding to the text segment, effectively reducing ambiguity and ensuring that the subtitle data accurately expresses the content of the audio data.
[0098] Figure 5 A flowchart illustrating a subtitle generation method provided in another embodiment of this disclosure. Please refer to... Figure 5 As shown, the method in this embodiment is Figure 1 Based on the illustrated embodiment, after step S104, the method further includes:
[0099] S106. Merge the subtitle data with the video to be processed to obtain the target video with subtitles.
[0100] The video data of the video to be processed consists of consecutive video frame images in the video. For each individual subtitle included in the subtitle data, according to the pre-set subtitle display style, each individual subtitle is superimposed on the video frame image of the corresponding display time period, thereby obtaining the target video with subtitles.
[0101] The display time period corresponding to a single subtitle can be determined based on the start time of the audio segment corresponding to the first character and the end time of the audio segment corresponding to the last character of the subtitle. Then, based on the start and end times corresponding to the single subtitle data, the video frame images within the corresponding display time period are determined, and the single subtitle is superimposed on all the video frame images within the corresponding display time period according to a pre-set display style. By performing the above processing on each subtitle in the subtitle data, the target video with subtitles is obtained.
[0102] The subtitle sentence lengths obtained by the method provided in this embodiment are more suitable for user reading, which can greatly improve the user experience.
[0103] By way of example, this disclosure also provides a subtitle generation apparatus.
[0104] Figure 6 This is a schematic diagram of a subtitle generation apparatus provided according to an embodiment of this disclosure. Please refer to [link / reference]. Figure 6 As shown, the device 600 provided in this embodiment includes:
[0105] The audio processing module 601 is used to extract audio data from the video to be processed, perform speech recognition on the audio data, and obtain the text data corresponding to the audio data.
[0106] The acquisition module 602 is used to acquire multiple segmentation positions of the text data determined based on grammatical analysis, as well as the pronunciation object information and timestamp information of the audio segments corresponding to each character in the text data.
[0107] The text segmentation module 603 is used to segment the text data into multiple text segments by taking the multiple segmentation positions, the pronunciation object information and timestamp information of the audio segments corresponding to each character; the audio segments corresponding to each character in the text segments belong to the same pronunciation object, and the duration of the blank segments in the audio segments corresponding to the text segments is less than a preset duration.
[0108] The merging module 604 is used to merge the multiple text segments according to the semantics of each text segment and the timestamp information of the audio segment corresponding to each character, so as to obtain multiple merged segments that are semantically fluent and meet the preset single subtitle sentence length requirements.
[0109] The generation module 605 is used to generate subtitle data corresponding to the video to be processed based on the multiple merged segments.
[0110] As one possible implementation, the merging module 604 is specifically used to determine whether adjacent text segments can be merged based on whether the merged text segments meet the preset single subtitle sentence length requirement; to determine whether adjacent text segments can be merged based on whether the semantics of the adjacent text segments are fluent after merging; and, if the text segment can be merged with the two adjacent text segments before and after it, then the two adjacent text segments with the shorter pause duration between the audio segments are merged.
[0111] As one possible implementation, the preset single subtitle length requirement includes: characters per second (CPS) requirement and / or the maximum display duration of a single subtitle in the video.
[0112] As one possible implementation, the text segmentation module 603 is specifically used to input the text data to the text processing module and obtain the plurality of text segments output by the text processing module; wherein, the text processing module includes: a sub-module for segmentation based on the plurality of segmentation positions, a sub-module for text segmentation based on the pronunciation object information of the audio segment corresponding to each character, and a sub-module for text segmentation based on the timestamp information of the audio segment corresponding to each character.
[0113] As one possible implementation, the subtitle data is a text-formatted subtitle SRT file.
[0114] As one possible implementation, the device 600 further includes a fusion module 606, used to fuse the subtitle data with the video to be processed to obtain a target video with subtitles.
[0115] The subtitle generation device provided in this embodiment can be used to execute the technical solution of any of the foregoing method embodiments. Its implementation principle and technical effect are similar, and can be referred to the detailed description of the foregoing method embodiments. For the sake of brevity, it will not be repeated here.
[0116] By way of example, this disclosure also provides an electronic device.
[0117] Figure 7 This is a schematic diagram of the structure of an electronic device provided according to an embodiment of this disclosure. Please refer to [link / reference]. Figure 7 As shown, the electronic device 700 provided in this embodiment includes a memory 701 and a processor 702.
[0118] The memory 701 can be a separate physical unit, connected to the processor 702 via a bus 703. Alternatively, the memory 701 and processor 702 can be integrated together, implemented in hardware, etc.
[0119] The memory 701 is used to store program instructions, and the processor 702 calls the program instructions to execute the subtitle generation method provided in any of the above method embodiments.
[0120] Optionally, when some or all of the methods in the above embodiments are implemented by software, the electronic device 700 may also include only the processor 702. The memory 701 for storing programs is located outside the electronic device 700, and the processor 702 is connected to the memory via circuits / wires to read and execute the programs stored in the memory.
[0121] The processor 702 can be a central processing unit (CPU), a network processor (NP), or a combination of a CPU and an NP.
[0122] The processor 702 may further include a hardware chip. This hardware chip may be an application-specific integrated circuit (ASIC), a programmable logic device (PLD), or a combination thereof. The PLD may be a complex programmable logic device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof.
[0123] The memory 701 may include volatile memory, such as random-access memory (RAM); the memory may also include non-volatile memory, such as flash memory, hard disk drive (HDD) or solid-state drive (SSD); the memory may also include a combination of the above types of memory.
[0124] This disclosure also provides a readable storage medium, including: computer program instructions, which, when executed by at least one processor of an electronic device, cause the electronic device to implement the subtitle generation method provided in any of the above method embodiments.
[0125] This disclosure also provides a computer program product, including computer program instructions that, when executed by an electronic device, implement the subtitle generation method provided in any of the above method embodiments.
[0126] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0127] The above description is merely a specific embodiment of this disclosure, enabling those skilled in the art to understand or implement it. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this disclosure. Therefore, this disclosure is not to be limited to the embodiments described herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A method for generating subtitles, characterized in that, include: Extract audio data from the video to be processed, perform speech recognition on the audio data, and obtain the text data corresponding to the audio data; The text data is obtained by acquiring multiple segmentation positions determined based on grammatical analysis, as well as the pronunciation object information and timestamp information of the audio segments corresponding to each character in the text data; Based on the multiple segmentation positions, the pronunciation object information and timestamp information of the audio segments corresponding to each character, the text data is segmented into multiple text segments; the audio segments corresponding to each character in the text segment belong to the same pronunciation object, and the duration of the blank segments in the audio segments corresponding to the text segment is less than a preset duration; Based on the semantics of each text segment and the timestamp information of the audio segment corresponding to each character, the multiple text segments are merged to obtain multiple merged segments that are semantically coherent and meet the preset single subtitle sentence length requirements; Based on the multiple merged segments, subtitle data corresponding to the video to be processed is generated.
2. The method according to claim 1, characterized in that, The merging based on the semantics of each text segment and the timestamp information of the corresponding audio segment includes: Whether adjacent text segments can be merged is determined based on whether the preset single subtitle length requirement is met after merging the adjacent text segments; Whether adjacent text segments can be merged is determined based on whether the semantics of the adjacent text segments are grammatically correct after merging. If the text segment can be merged with both preceding and following text segments, then the two adjacent text segments with shorter pause durations between audio segments will be merged.
3. The method according to claim 1 or 2, characterized in that, The preset single subtitle length requirements include: characters per second (CPS) requirement and / or the maximum display duration of a single subtitle in the video.
4. The method according to claim 1, characterized in that, The text data is segmented into multiple text segments based on the multiple segmentation positions, the pronunciation object information of the audio segment corresponding to each character, and the timestamp information, including: The text data is input into the text processing module to obtain the multiple text fragments output by the text processing module; The text processing module includes: a sub-module for segmentation based on the multiple segmentation positions, a sub-module for text segmentation based on the pronunciation object information of the audio segments corresponding to each character, and a sub-module for text segmentation based on the timestamp information of the audio segments corresponding to each character.
5. The method according to claim 1, characterized in that, The subtitle data is a text-formatted subtitle SRT file.
6. The method according to claim 1, characterized in that, The method further includes: The subtitle data is fused with the video to be processed to obtain a target video with subtitles.
7. A subtitle generation device, characterized in that, include: An audio processing module is used to extract audio data from the video to be processed, perform speech recognition on the audio data, and obtain the text data corresponding to the audio data. The acquisition module is used to acquire multiple segmentation positions of the text data determined based on grammatical analysis, as well as the pronunciation object information and timestamp information of the audio segments corresponding to each character in the text data; The text segmentation module is used to segment the text data into multiple text segments based on the multiple segmentation positions, the pronunciation object information and timestamp information of the audio segments corresponding to each character; the audio segments corresponding to each character in the text segments belong to the same pronunciation object, and the duration of the blank segments in the audio segments corresponding to the text segments is less than a preset duration; The merging module is used to merge the multiple text segments according to the semantics of each text segment and the timestamp information of the audio segment corresponding to each character, so as to obtain multiple merged segments that are semantically fluent and meet the preset single subtitle sentence length requirements. The generation module is used to generate subtitle data corresponding to the video to be processed based on the multiple merged segments.
8. An electronic device, characterized in that, include: Memory and processor; The memory is configured to store computer program instructions; The processor is configured to execute the computer program instructions, causing the electronic device to implement the subtitle generation method as described in any one of claims 1 to 6.
9. A readable storage medium, characterized in that, include: Computer program instructions; The computer program instructions are executed by at least one processor of the electronic device, causing the electronic device to implement the subtitle generation method as described in any one of claims 1 to 6.
10. A computer program product comprising computer program instructions, characterized in that, When the computer program instructions are executed by an electronic device, they implement the subtitle generation method as described in any one of claims 1 to 6.