Audio and video synchronization method and device, electronic device and storage medium
By automatically processing audio clips to match the subtitle display time, the problem of audio and video synchronization in film and television dubbing is solved, achieving efficient audio and video synchronization, reducing labor costs and improving synchronization efficiency.
Patent Information
- Application Number
- CN202411794539.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-06
- Publication Date
- 2025-10-03
- Estimated Expiration
- 2044-12-06
AI Technical Summary
In the dubbing process of film and television dramas, the existing technology causes the problem of audio and video asynchrony due to language differences, especially the overlapping sounds and audio and video asynchrony caused by differences in subtitle text length and pronunciation system. Manual adjustment is inefficient and costly.
By obtaining the subtitle file and subtitle timeline of the video file, the initial audio clip is automatically processed to generate a target audio clip that matches the subtitle display time, and audio and video synchronization is achieved by adopting hard synchronization and soft synchronization strategies.
The target audio file synchronized with the video file can be generated efficiently and quickly without human intervention, reducing labor costs, improving the efficiency of audio and video synchronization, and ensuring the deep integration and harmonious resonance of video and audio.
Smart Images

Figure CN119629411B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computers, and in particular to a method and device for synchronizing audio and video, an electronic device, and a storage medium. Background Art
[0002] In recent years, with the continuous improvement in the quality of domestic film and television dramas, many popular domestic film and television dramas have also appeared on foreign television screens and overseas broadcasting platforms. In order to better promote the export of film and television dramas, the subtitles of film and television dramas can be translated into the language of the target market and dubbed. However, due to the differences in the length of subtitle texts and pronunciation systems (such as speaking speed) between different languages, simply dubbing the video directly can easily cause "overlapping" or "asynchronous sound and picture". For example, for clips with dense lines, if the dubbing is based on the subtitle start time of the original drama subtitle timeline, if the current audio clip is too long, it is very likely that the current audio clip has not yet finished and the next audio clip has begun. This overlapping of the previous and next audio clips causes the "overlapping" phenomenon, which is also a phenomenon of "asynchronous sound and picture".
[0003] To solve the problem of audio and video asynchrony in dubbing production between different languages in film and television dramas, the audio duration and playback time are usually manually adjusted. However, this method has high labor costs and low efficiency. Summary of the Invention
[0004] In view of this, the present disclosure proposes a method and device for audio and video synchronization, an electronic device and a storage medium, which can automatically generate a target audio file synchronized with a video file efficiently and quickly, and achieve audio and video synchronization without human intervention, thereby reducing labor costs and improving audio and video synchronization efficiency.
[0005] According to one aspect of the present disclosure, a method for synchronizing audio and video is provided, comprising: obtaining a subtitle file corresponding to a video file, the subtitle file comprising a plurality of translated subtitle texts and a subtitle timeline; the translated subtitle text comprising a translation of the original subtitle text of the video file into a target language; the subtitle timeline being the timeline adopted by the original subtitle text of the video file, the subtitle timeline representing a subtitle display time of the subtitle text and a subtitle character identifier, the subtitle character identifier representing a character speaking the subtitle text, the subtitle display time comprising a start time and an end time of the subtitle text display; for an i-th translated subtitle text in the subtitle file, obtaining a subtitle timer for the i-th translated subtitle text; The method comprises the following steps: obtaining an i-th initial audio segment corresponding to the video file and determining the subtitle display time and subtitle character identification corresponding to the i-th translated subtitle text based on the subtitle timeline; wherein the i-th initial audio segment is an audio segment generated based on the i-th translated subtitle text; processing the i-th initial audio segment into a target audio segment that matches the subtitle display time corresponding to the i-th translated subtitle text according to the segment duration of the i-th initial audio segment, the subtitle display time and subtitle character identification corresponding to the i-th translated subtitle text; and generating a target audio file synchronized with the video file according to the target audio segments that match the subtitle display time of each translated subtitle text in the subtitle file.
[0006] According to another aspect of the present disclosure, another audio-visual synchronization method is provided, comprising: obtaining a subtitle file corresponding to a video file, the subtitle file comprising a plurality of translated subtitle texts and a subtitle timeline; the translated subtitle text comprising a translation of the original subtitle text of the video file into a target language; the subtitle timeline being the timeline adopted by the original subtitle text of the video file, the subtitle timeline representing a subtitle display time and a subtitle character identifier of the subtitle text, the subtitle character identifier representing a character speaking the subtitle text, the subtitle display time comprising a start time and an end time of the subtitle text display; when the subtitle timeline represents a subtitle density less than or equal to a specified density threshold, executing an audio-visual hard synchronization strategy based on the subtitle file to obtain a target audio file synchronized with the video file; the subtitle density represents a density of the subtitle text; and when the subtitle timeline represents a subtitle density greater than the specified density threshold, executing an audio-visual soft synchronization strategy based on the subtitle file to obtain a target audio file synchronized with the video file.
[0007] According to another aspect of the present disclosure, a sound and picture synchronization device is provided, comprising: a first acquisition module, for acquiring a subtitle file corresponding to a video file, the subtitle file comprising a plurality of translated subtitle texts and a subtitle timeline; the translated subtitle text comprising a translation of the original subtitle text of the video file into a target language; the subtitle timeline being the timeline adopted by the original subtitle text of the video file, the subtitle timeline representing the subtitle display time of the subtitle text and a subtitle character identifier, the subtitle character identifier representing the character who speaks the subtitle text, the subtitle display time comprising a start time and an end time of the subtitle text display; a second acquisition module, for acquiring the i-th translated subtitle text in the subtitle file, The invention relates to an i-th initial audio segment corresponding to the video file, and determines the subtitle display time and subtitle character identification corresponding to the i-th translated subtitle text based on the subtitle timeline; wherein the i-th initial audio segment is an audio segment generated based on the i-th translated subtitle text; a processing module is used to process the i-th initial audio segment into a target audio segment that matches the subtitle display time corresponding to the i-th translated subtitle text according to the segment length of the i-th initial audio segment, the subtitle display time and subtitle character identification corresponding to the i-th translated subtitle text; a generation module is used to generate a target audio file synchronized with the video file according to the target audio segments that match the subtitle display time of each translated subtitle text in the subtitle file.
[0008] According to another aspect of the present disclosure, another audio-visual synchronization device is provided, comprising: a file acquisition module, configured to acquire a subtitle file corresponding to a video file, the subtitle file comprising a plurality of translated subtitle texts and a subtitle timeline; the translated subtitle text comprising a translation of the original subtitle text of the video file into a target language; the subtitle timeline being the timeline adopted by the original subtitle text of the video file, the subtitle timeline representing a subtitle display time and a subtitle character identifier of the subtitle text, the subtitle character identifier representing a character speaking the subtitle text, the subtitle display time comprising a start time and an end time of the subtitle text display; a first execution module, configured to, when the subtitle density represented by the subtitle timeline is less than or equal to a specified density threshold, execute a hard audio-visual synchronization strategy based on the subtitle file to obtain a target audio file synchronized with the video file; and a second execution module, configured to, when the subtitle density represented by the subtitle timeline is greater than the specified density threshold, execute a soft audio-visual synchronization strategy based on the subtitle file to obtain a target audio file synchronized with the video file.
[0009] According to another aspect of the present disclosure, an electronic device is provided, comprising: a processor; and a memory for storing processor-executable instructions; wherein the processor is configured to implement the above method when executing the instructions stored in the memory.
[0010] According to another aspect of the present disclosure, a non-volatile computer-readable storage medium is provided, on which computer program instructions are stored, wherein the computer program instructions implement the above method when executed by a processor.
[0011] According to another aspect of the present disclosure, a computer program product is provided, including a computer-readable code, or a non-volatile computer-readable storage medium carrying the computer-readable code. When the computer-readable code runs in a processor of an electronic device, the processor in the electronic device executes the above method.
[0012] According to various aspects of the present disclosure, by adopting the subtitle display time and subtitle character identification indicated by the subtitle timeline adopted by the original subtitle text of the video file, the initial audio segment corresponding to the translated subtitle text into which the original subtitle text is translated can be automatically processed into a target audio segment that matches the subtitle display time, thereby efficiently and quickly generating a target audio file synchronized with the video file, and achieving the audio-visual synchronization effect without manual intervention, thereby greatly reducing labor costs.
[0013] Further features and aspects of the present disclosure will become apparent from the following detailed description of exemplary embodiments with reference to the attached drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] The accompanying drawings, which are incorporated in and constitute a part of the specification, illustrate exemplary embodiments, features, and aspects of the disclosure and, together with the description, serve to explain the principles of the disclosure.
[0015] Figure 1 A flowchart of a method for synchronizing audio and video according to an embodiment of the present disclosure is shown.
[0016] Figure 2 A schematic diagram showing a subtitle timeline and subtitle text according to an embodiment of the present disclosure is shown.
[0017] Figure 3 A schematic diagram of audio and video synchronization logic according to an embodiment of the present disclosure is shown.
[0018] Figure 4 A schematic diagram illustrating another audio-video synchronization logic according to an embodiment of the present disclosure is shown.
[0019] Figure 5 A schematic diagram illustrating another audio-video synchronization logic according to an embodiment of the present disclosure is shown.
[0020] Figure 6 A schematic diagram illustrating another audio-video synchronization logic according to an embodiment of the present disclosure is shown.
[0021] Figure 7 A schematic diagram illustrating another audio-video synchronization logic according to an embodiment of the present disclosure is shown.
[0022] Figure 8 A schematic diagram illustrating another audio-video synchronization logic according to an embodiment of the present disclosure is shown.
[0023] Figure 9 A schematic diagram illustrating another audio-video synchronization logic according to an embodiment of the present disclosure is shown.
[0024] Figure 10 A flowchart of a hard audio-video synchronization strategy according to an embodiment of the present disclosure is shown.
[0025] Figure 11 A schematic diagram illustrating another audio-video synchronization logic according to an embodiment of the present disclosure is shown.
[0026] Figure 12 A schematic diagram illustrating another audio-video synchronization logic according to an embodiment of the present disclosure is shown.
[0027] Figure 13 A schematic diagram illustrating another audio-video synchronization logic according to an embodiment of the present disclosure is shown.
[0028] Figure 14 A schematic diagram illustrating another audio-video synchronization logic according to an embodiment of the present disclosure is shown.
[0029] Figure 15 A schematic diagram illustrating another audio-video synchronization logic according to an embodiment of the present disclosure is shown.
[0030] Figure 16 A schematic diagram illustrating another audio-video synchronization logic according to an embodiment of the present disclosure is shown.
[0031] Figure 17 A flowchart of an audio-video soft synchronization strategy according to an embodiment of the present disclosure is shown.
[0032] Figure 18 A block diagram of an audio and video synchronization device according to an embodiment of the present disclosure is shown.
[0033] Figure 19 A block diagram of an electronic device 1900 according to an embodiment of the present disclosure is shown. DETAILED DESCRIPTION
[0034] Various exemplary embodiments, features, and aspects of the present disclosure will be described in detail below with reference to the accompanying drawings. The same reference numerals in the accompanying drawings represent elements with the same or similar functions. Although various aspects of the embodiments are shown in the accompanying drawings, the drawings are not necessarily drawn to scale unless otherwise indicated.
[0035] The word “exemplary” is used exclusively herein to mean “serving as an example, example, or illustration.” Any embodiment described herein as “exemplary” is not necessarily to be construed as preferred or advantageous over other embodiments.
[0036] The term "and / or" herein simply describes an association relationship between associated objects, indicating that three relationships can exist. For example, "A and / or B" can represent three situations: A exists alone, A and B exist simultaneously, and B exists alone. Furthermore, the term "at least one" herein represents any combination of at least two of any one or more of a plurality of items. For example, "at least one of A, B, and C" can represent any one or more elements selected from the set consisting of A, B, and C. In the description of this disclosure, "plurality" means two or more, unless otherwise specifically defined.
[0037] It should be understood that the terms "first," "second," and the like in the claims, specification, and drawings of the present disclosure are used to distinguish between different objects, rather than to describe a specific order. The terms "include" and "comprising" used in the specification and claims of the present disclosure indicate the presence of the described features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or combinations thereof.
[0038] In addition, numerous specific details are provided in the following detailed description to better illustrate the present disclosure. Those skilled in the art will appreciate that the present disclosure can be practiced without certain specific details. In some instances, methods, means, components, and circuits well known to those skilled in the art are not described in detail in order to highlight the main points of the present disclosure.
[0039] As mentioned above, in order to better carry out the export of domestic film and television dramas, the subtitles of the film and television dramas can be translated into the target language of the target country and dubbed. However, due to the differences in the length of the subtitle texts and the pronunciation system (such as speaking speed) between different languages, simply dubbing directly can easily cause the phenomenon of "overlapping" or "sound and picture are not synchronized". In order to solve the problem of sound and picture synchronization in the dubbing production of film and television dramas, that is, to solve the problem of sound and picture synchronization between the original video screen and the dubbing audio of other languages, the embodiment of the present disclosure proposes a method for synchronizing sound and picture, which can realize the full automation of accelerating, shifting, merging and other processing of the initial audio segment without manual intervention according to the segment length of the initial audio segment corresponding to the translated subtitle text and the subtitle time axis used by the original subtitle text of the video file, to efficiently generate a target audio file synchronized with the video file, and can also combine parameters such as acceleration ratio, offset, and number of merges to achieve a better sound and picture synchronization effect, which is conducive to ensuring the deep integration and harmonious resonance of the sound and picture of the video file and the target audio file, thereby bringing a smooth audio-visual experience to the audience, and significantly reducing the labor cost of the dubbing-related business link.
[0040] The audio and video synchronization method of the embodiment of the present disclosure can be deployed on various terminal devices through software or hardware modification. The terminal device involved in the embodiment of the present disclosure may refer to a device with a wireless connection function and / or a wired connection function. The wireless connection function refers to the ability to connect to other devices through wireless connection methods such as wifi and Bluetooth. The terminal device involved in the embodiment of the present disclosure may also communicate with other devices through a wired connection function. The terminal device involved in the embodiment of the present disclosure may be a touch screen, a non-touch screen, or a device without a screen. The touch screen device can be controlled by clicking, sliding, etc. on the display screen with a finger or a stylus. The non-touch screen device can be connected to an input device such as a mouse, keyboard, touch panel, etc., and the terminal device can be controlled through the input device. For example, a device without a screen can be a Bluetooth speaker without a screen. For example, the terminal device of the present application may include but is not limited to user equipment (UE), mobile device, user terminal, terminal, handheld device, tablet computer, laptop computer, PDA, computing device, etc.
[0041] The audio and video synchronization method of the embodiment of the present disclosure can also be deployed on a server. The server can be located in the cloud or locally. It can be a physical device or a virtual device, such as a virtual machine, a container, etc., and has a wireless communication function, wherein the wireless communication function can be set in the chip (system) or other parts or components of the server. It can refer to a device with a wireless connection function. The wireless connection function means that it can be connected to other servers or terminal devices through wireless connection methods such as Wi-Fi and Bluetooth. The server involved in the embodiment of the present disclosure can also have the function of communicating through a wired connection. For example, the server of the embodiment of the present disclosure can be located in the cloud, communicate with the terminal device, receive the subtitle file of the video file sent by the terminal device, and use the audio and video synchronization method deployed on the server to generate a target audio file synchronized with the video file based on the subtitle file, and return it to the terminal device to generate the target audio file for the user in the terminal device.
[0042] Figure 1 FIG. 1 is a flow chart showing a method for synchronizing audio and video according to an embodiment of the present disclosure. The method can be applied to electronic devices such as the above-mentioned terminal device or server. Figure 1 As shown, the method includes: steps S11 to S14.
[0043] In step S11, the subtitle file corresponding to the video file is obtained;
[0044] Among them, the subtitle file includes multiple translated subtitle texts and a subtitle timeline; the translated subtitle text includes the original subtitle text of the video file translated into the target language; the subtitle timeline is the timeline used by the original subtitle text of the video file, and the subtitle timeline represents the subtitle display time and subtitle character identification of the subtitle text. The subtitle character identification represents the character who speaks the subtitle text (that is, the speaker of the subtitle text, or the role who speaks the subtitle text), and the subtitle display time includes the start time and end time of the subtitle text display.
[0045] Wherein, the original subtitle text and the translated subtitle text of the video file are texts of different languages, for example, the original subtitle text is Chinese text, and the translated subtitle text is Thai text, English text, etc. Those skilled in the art can translate the original subtitle text corresponding to the video file into the translated subtitle text of any target language according to actual needs, and this embodiment of the present disclosure is not limited to this. It should be understood that the subtitle display time and subtitle character identification of the translated subtitle text and the original subtitle text are consistent, that is, the subtitle time axis adopted by the original subtitle text is also the subtitle time axis required to be adopted by the translated subtitle text.
[0046] For example, Figure 2 A schematic diagram showing a subtitle timeline and subtitle text is shown, such as Figure 2As shown, the subtitle time axis includes the subtitle display time (t0, t1, ..., t 13 ) and subtitle character identification (speaker1, speaker2, speaker3). For example, the first translated subtitle text "auspicious day" is displayed at the start time t0 and the end time t1. The corresponding subtitle character identification is "speaker1", and so on.
[0047] In step S12, for the i-th translated subtitle text in the subtitle file, the i-th initial audio segment corresponding to the i-th translated subtitle text is obtained, and the subtitle display time and subtitle character identification corresponding to the i-th translated subtitle text are determined based on the subtitle timeline; wherein the i-th initial audio segment is an audio segment generated based on the i-th translated subtitle text.
[0048] The i-th original audio segment corresponding to the i-th translated subtitle text can be obtained sequentially according to the order of the subtitle texts indicated in the subtitle timeline, and the subtitle display time and subtitle character identifier corresponding to the i-th translated subtitle text can be determined. Here, "one" translated subtitle text refers to a translated subtitle text between a start time and a corresponding end time in the subtitle timeline, that is, a subtitle text displayed simultaneously on the screen during the same period.
[0049] In actual applications, the initial audio segment corresponding to each translated subtitle text can be obtained in advance through manual dubbing, or the audio synthesis technology known in the art can be used to synthesize the translated subtitle text into the corresponding initial audio segment based on the audio features such as the tone and timbre of the character's speech in the original audio segment of the original subtitle text. This embodiment of the present disclosure does not limit this.
[0050] As mentioned above, the subtitle timeline used by the original subtitle text is also the subtitle timeline required to be used by the translated subtitle text. Therefore, based on the subtitle timeline, the subtitle display time and subtitle character identification of the original subtitle text corresponding to the i-th translated subtitle text can be determined, that is, the subtitle display time and subtitle character identification corresponding to the i-th translated subtitle text can be determined.
[0051] In step S13, the i-th initial audio segment is processed into a target audio segment that matches the subtitle display time corresponding to the i-th translated subtitle text based on the segment length of the i-th initial audio segment, the subtitle display time corresponding to the i-th translated subtitle text, and the subtitle character identification.
[0052] It should be understood that the subtitle display time of the subtitle text on the subtitle timeline is equivalent to the playback time of the dubbing audio clip corresponding to the subtitle text, or in other words, the display time of the subtitle text should be aligned with the playback time of the dubbing audio clip corresponding to the subtitle text, so that the audio and video synchronization effect can be achieved when playing the video. However, due to the differences in the length of the subtitle text and the pronunciation system (such as speaking speed) between different languages, the initial audio clip corresponding to the translated subtitle text may not be aligned with the subtitle display time indicated by the original subtitle timeline of the video file. For example, the segment length of the initial audio clip may be less than the subtitle display time indicated on the subtitle timeline, or it may be greater than the subtitle display time indicated on the subtitle timeline. If the playback of the initial audio clip corresponding to the translated subtitle text is directly controlled according to the original subtitle timeline, the phenomenon of audio and video being out of sync may occur. Therefore, the segment length of the initial audio clip corresponding to each translated subtitle text, the subtitle display time corresponding to each translated subtitle text, and the subtitle character identification can be used to process the initial audio clip corresponding to each translated subtitle text into a target audio clip that matches the subtitle display time corresponding to each translated subtitle text to generate a target audio file synchronized with the video file.
[0053] Specifically, in one possible implementation, step S13, processing the i-th initial audio segment into a target audio segment that matches the subtitle display time of the i-th translated subtitle text based on the segment duration of the i-th initial audio segment, the subtitle display time corresponding to the i-th translated subtitle text, and the subtitle character identifier, may include:
[0054] Step S1300: If the duration of the i-th initial audio segment is less than or equal to the subtitle display duration of the i-th translated subtitle text, determine the i-th initial audio segment as a target audio segment that matches the subtitle display duration of the i-th translated subtitle text;
[0055] Step S1301: When the duration of the i-th initial audio segment is greater than the subtitle display duration of the i-th translated subtitle text, the ratio between the duration of the i-th initial audio segment and the subtitle display duration of the i-th translated subtitle text is determined as the speedup ratio corresponding to the i-th initial audio segment.
[0056] Step S1302: When the acceleration ratio corresponding to the i-th initial audio segment is less than or equal to a preset acceleration ratio threshold, the i-th initial audio segment is accelerated according to the acceleration ratio corresponding to the i-th initial audio segment to obtain a target audio segment that matches the subtitle display time of the i-th translated subtitle text.
[0057] The subtitle display duration of the i-th translated subtitle text is determined based on the subtitle display time of the i-th translated subtitle text. It should be understood that if the start time and end time of the translated subtitle text display are known, the subtitle display duration of the translated subtitle text (i.e., the duration between the end time and the start time) can be obtained. Those skilled in the art can customize the acceleration ratio threshold according to actual needs. Setting the acceleration ratio threshold is equivalent to setting a maximum acceleration threshold to prevent the accelerated audio speed from being too fast and affecting the listening experience.
[0058] Taking into account that, in the above step S1300, the i-th initial audio segment is directly determined as the target audio segment that matches the subtitle display time of the i-th translated subtitle text, there may be a blank period after the target audio segment corresponding to the i-th translated subtitle text and before the start time of display of the i+1-th translated subtitle text. Therefore, the target audio segment can be pushed into the audio queue, and the blank audio segment is used to fill the blank after the target audio segment to ensure that the subsequently generated target audio segment is aligned with the subtitle display time indicated on the subtitle timeline; wherein, the audio queue is used to store the target audio segment and the blank audio segment in sequence, and the blank audio segment is an audio segment without sound.
[0059] Therefore, after obtaining the target audio segment that matches the subtitle display time of the i-th translated subtitle text, the method may further include: step S1303, pushing the target audio segment into the audio queue, and based on the first deviation between the segment end time of the target audio segment aligned to the subtitle timeline (that is, after aligning the start time of the target audio segment with the start time of the i-th translated subtitle text) and the start time of the display of the i+1-th translated subtitle text on the subtitle timeline, filling the target audio segment in the audio queue with a blank audio segment corresponding to the first deviation.
[0060] For example, Figure 3 As shown, the first initial audio segment (P1) of the first English subtitle text "auspicious day" (that is, the first translated subtitle text) corresponding to the first original subtitle text "Auspicious Day" is shorter than the subtitle display duration of the first English subtitle text (that is, the duration between t0 and t1). Then, the P1 can be added to the audio queue as the target audio segment M1 of the first English subtitle text, and then the segment end time t can be aligned to the subtitle timeline (that is, the start time of M1 is aligned with the start time of the first English subtitle text) according to M1. M1 The first deviation (i.e., t2-t M1), fill the blank audio segment (K1) after M1 in the audio queue, where the segment length of K1 is t2-t M1 duration.
[0061] For example, Figure 4 As shown, the segment duration of the second initial audio segment (P2) corresponding to the second translated subtitle text "young lady" corresponding to the second original subtitle text "girl" is greater than the subtitle display duration of the second translated subtitle text (i.e., the duration between t2 and t3). At this time, the ratio between the segment duration of P2 and the subtitle display duration (i.e., the duration between t2 and t3) is calculated as the acceleration ratio corresponding to P2 (e.g., the acceleration ratio is 1.1). If the acceleration ratio corresponding to P2 (e.g., 1.1) is less than the preset acceleration ratio threshold (e.g., the acceleration ratio threshold is 1.25), then accelerating P1 according to the acceleration ratio corresponding to P2 (e.g., 1.1) can obtain the target audio segment M2 that matches the subtitle display time (i.e., t2 and t3), that is, accelerating P1 to a target audio segment M2 at 1.1 times the speed. Then, M2 can be pushed to the audio queue and aligned to the segment cutoff time t on the subtitle timeline according to M2. M2 (At this time t M2 = t3) and the start time t4 of the third translated subtitle text “It’s time to get married” on the subtitle timeline (i.e., t4-t M2 ), fill the blank audio segment (K2) after M2 in the audio queue, where the segment length of K2 is t2-t M2 It should be understood that if M2 is aligned to the end time of the segment on the subtitle timeline, M2 The first deviation t4-t from the start time t4 of the third translated subtitle text display on the subtitle timeline M2 If it is 0, that is, there is no blank between the end time of the second translated subtitle text and the start time of the third translated subtitle text, then the blank audio segment may not be filled after M2 in the audio queue.
[0062] In practical applications, after the initial audio clip is accelerated according to the acceleration ratio, the acceleration ratio can be marked on the accelerated target audio clip, so that in some scenarios where the original speed needs to be restored, the marked acceleration ratio can be used to restore the target audio clip to the original speed of the initial audio clip.
[0063] It is also possible that the acceleration ratio corresponding to the initial audio segment is greater than a preset acceleration ratio threshold. In this case, if the acceleration ratio is used, the difference from the original speech rate may be too large, affecting the listening experience. Therefore, in one possible implementation, the above step S13, based on the segment length of the i-th initial audio segment, the subtitle display time corresponding to the i-th translated subtitle text, and the subtitle character identifier, processes the i-th initial audio segment into a target audio segment that matches the subtitle display time of the i-th translated subtitle text, which may include:
[0064] Step S1304: If the acceleration ratio corresponding to the i-th initial audio segment is greater than the acceleration ratio threshold, a determination is made as to whether there are sufficient blank segments before and after the i-th initial audio segment based on the subtitle display time of the i-th translated subtitle text. The existence of sufficient blank segments includes the presence of blank segments before and / or after the i-th initial audio segment, and the duration of the blank segments is greater than or equal to the acceleration offset corresponding to the i-th initial audio segment.
[0065] The acceleration offset includes the difference between the duration of the audio segment obtained by accelerating the i-th initial audio segment using the acceleration ratio threshold and the subtitle display duration corresponding to the i-th translated subtitle text; the duration of the blank segment before the i-th initial audio segment includes the duration of the blank audio segment filled before the i-th initial audio segment; the duration of the blank segment after the i-th initial audio segment includes the duration on the subtitle timeline from the end time of the i-th translated subtitle text display to the start time of the i+1-th translated subtitle text display;
[0066] Step S1305: If there are sufficient blank segments before and after the i-th initial audio segment, the i-th initial audio segment is accelerated according to the acceleration ratio threshold to obtain a target audio segment that matches the subtitle display time of the i-th translated subtitle text;
[0067] Step S1306: Push the target audio segment into the audio queue by occupying a blank segment before the i-th initial audio segment and / or occupying a blank segment after the i-th initial audio segment.
[0068] In actual applications, after the i-th initial audio clip is accelerated according to the acceleration ratio threshold to obtain the target audio clip, the blank audio clip filled before the i-th initial audio clip in the audio queue can be compressed according to the above-mentioned acceleration offset, or the blank space on the subtitle timeline from the end time of the i-th translated subtitle text display to the start time of the i+1-th translated subtitle text display can be directly occupied until the above-mentioned acceleration offset requirement is met. At this time, the target audio clip accelerated according to the acceleration ratio threshold can be pushed into the audio queue.
[0069] In actual applications, any occupation method in this field can be used in the above step S1306 to occupy the blank segments before and / or after the i-th initial audio segment, as long as the target audio segment pushed into the audio queue does not overlap with the target audio segments before and after it. This embodiment of the present disclosure does not limit this.
[0070] Considering that directly compressing the blank audio segment before or occupying the blank audio segment after the target audio segment based on the acceleration offset may cause the target audio segment to be too early or too late relative to the subtitle display time, an offset threshold can be set to limit the amount of compression or occupation of the blank segment. Specifically, step S1306, which pushes the target audio segment to the audio queue by occupying the blank segment before and / or after the i-th initial audio segment, may include:
[0071] Step S13061: If there are sufficient blank segments before and after the i-th initial audio segment, the required compression amount of the blank segment before the i-th initial audio segment and the amount of the blank segment after the i-th initial audio segment occupied by the target audio segment are determined based on a preset offset threshold, an acceleration offset, and the duration of the blank segment before the i-th initial audio segment. The offset threshold is used to limit the maximum compression amount of the blank segment.
[0072] Step S13062: If the occupied amount is less than or equal to the duration of the blank segment after the i-th initial audio segment, based on the compression amount, compress the blank audio segment that is filled after the target audio segment corresponding to the i-1th translated subtitle text in the audio queue, and push the target audio segment that matches the subtitle display time of the i-th translated subtitle text into the audio queue and arrange it after the compressed blank audio segment.
[0073] Step S13063, based on the second deviation between the segment cutoff time when the target audio segment corresponding to the i-th translated subtitle text in the audio queue is aligned to the subtitle timeline (that is, the target audio segment is arranged after the compressed blank audio segment) and the start time when the i+1-th translated subtitle text is displayed on the subtitle timeline, a blank audio segment corresponding to the second deviation is filled after the target audio segment corresponding to the i-th translated subtitle text in the audio queue.
[0074] For example, Figure 5As shown, the segment duration of the third initial audio segment (P3) corresponding to the third translated subtitle text "It's time to get married" corresponding to the third original subtitle text "It's time to get married" is longer than the subtitle display duration of the third translated subtitle text (i.e., the duration between t4 and t5). At this time, the ratio between the segment duration of P3 and the subtitle display duration (i.e., the duration between t4 and t5) is calculated as the acceleration ratio corresponding to P3. If the acceleration ratio corresponding to P3 (e.g., 1.6) is greater than the preset acceleration ratio threshold (e.g., the acceleration ratio threshold is 1.25), and if the duration of the blank segment before P3 (i.e., the blank audio segment K2) (the duration between t4 and t3, T 34 ) plus the duration of the blank segment after P3 (i.e. the duration T between the end time t5 of the third translated subtitle text and the start time t6 of the fourth translated subtitle text). 56 ) is greater than or equal to the difference between the duration of the audio segment obtained by accelerating P3 using the acceleration ratio threshold of 1.25 and the subtitle display duration of the third translated subtitle text (i.e., the duration between t4 and t5), for example, T 34 +T 56 The duration of the audio clip obtained by accelerating P3 using the acceleration ratio threshold of 1.25 is 12.8 seconds. The subtitle display duration of the third translated subtitle text (i.e., the duration between t4 and t5) is 10 seconds. Since the acceleration offset (i.e., the difference between 12.8 seconds and 10 seconds is 2.8 seconds) is less than the above T 34 +T 56 If the speedup threshold of 1.25 is used to speed up P3, the resulting audio clip is 15 seconds long. The difference of 5 seconds between 15 seconds and 10 seconds is greater than the above T 34 +T 56 4 seconds, it is considered that there is no sufficient blank segment before and after P3;
[0075] Then, if there are sufficient blank segments before and after P3, P3 can be accelerated using the acceleration ratio threshold of 1.25 to obtain the target audio segment M3 corresponding to the third translated subtitle text; at this time, since the segment length of the target audio segment M3 obtained by accelerating using the acceleration ratio threshold of 1.25 is longer than the corresponding subtitle display duration (i.e., the duration between t4 and t5), then if M3 is directly pushed into the audio queue, a large amount of blank space between t5 and t6 will be occupied. Therefore, if there is a blank audio segment (such as K2) before P3, the blank audio segment (such as K2) before P3 can be compressed, that is, the blank audio segment can be shortened. The duration of the blank audio segment (such as K2) and the compression amount (i.e., the shortened duration) of the blank audio segment (such as K2) can be determined according to the preset offset threshold, the acceleration offset of the target audio segment M3, and the duration of the blank audio segment (such as K2). Specifically, when the duration of the blank audio segment is greater than or equal to the offset threshold and the acceleration offset is less than or equal to the offset threshold, the compression amount of the blank audio segment (such as K2) is the acceleration offset; when the duration of the blank audio segment is greater than or equal to the offset threshold and the acceleration offset is greater than the offset threshold, the compression amount of the blank audio segment is the offset threshold; when the duration of the blank audio segment is less than the offset threshold And when the acceleration offset is greater than the duration of the blank audio segment, the compression amount of the blank audio segment is the duration of the blank audio segment (that is, the blank audio segment is removed); when the duration of the blank audio segment is less than the offset threshold and the acceleration offset is less than the duration of the blank audio segment, the compression amount of the blank audio segment is the acceleration offset; after determining the compression amount of the blank audio segment (such as K2), the difference between the acceleration offset and the compression amount can be used as the occupancy of the blank segment after M3 to P3, and when the occupancy of the blank segment after P3 by M3 is less than or equal to the duration of the blank segment after P3, the execution of M3 is pushed to the audio queue for processing. For example, if the acceleration offset is 2.8 seconds, the offset threshold is 2 seconds, and the duration of K2 is 3 seconds, then the compression amount of K2 is the maximum compression amount of 2 seconds limited by the offset threshold, which is equivalent to shortening the duration of K2 by 2 seconds. Then the duration of K2 after shortening is 1 second. At this time, the difference between the acceleration offset of 2.8 seconds and the compression amount of 2 is 0.8 seconds, which is the amount of space occupied by M3 in the blank segment after P3. If the occupied amount of 0.8 seconds is less than or equal to the duration between t5 and t6 (equivalent to the fact that there is still enough space between t5 and t6 for the target audio segment to occupy under the limitation of the above offset threshold), then in this case, Figure 5 As shown, M3 can be pushed to the audio queue. The M3 pushed to the audio queue will occupy the blank space of 0.8 seconds after t5, that is, the M3 pushed to the audio queue corresponds to the segment cutoff time t on the subtitle timeline. M3There is a 0.8 second difference between it and t5, and the starting time of M3 is 2 seconds before t4, that is, it occupies the 2 seconds of blank space before t4; and if the acceleration offset is 1.8 seconds, the offset threshold is 2 seconds, and the duration of K2 is 3 seconds, then the compression amount of K2 is the acceleration offset 1.8 seconds (that is, K2 only needs to be compressed for 1.8 seconds), which can meet the duration requirement of M3. At this time, it is equivalent to shortening the duration of K2 by 1.8 seconds. Then the shortened duration of K2 is 1.2 seconds. In this case, the occupancy of the blank segment after P3 is 0. In this case, Figure 6 As shown, M3 can be pushed to the audio queue, and the M3 pushed to the audio queue corresponds to the segment cutoff time t on the subtitle timeline. M3 =t5, the starting time is 1.8 seconds before t4; and if the acceleration offset is 2.8 seconds, the offset threshold is 2 seconds, and the duration of K2 is 1.5 seconds, then the compression amount of K2 is 1.5 seconds of the duration of K2, which is equivalent to shortening the duration of K2 by 1.5 seconds, that is, removing K2. In this case, the duration of the blank segment after P3 is greater than or equal to the 1.3 seconds occupied by the blank segment after P3. In this case, Figure 7 As shown, M3 can be pushed to the audio queue. The M3 pushed to the audio queue will occupy the blank space of 1.3 seconds after t5, that is, the M3 pushed to the audio queue corresponds to the segment cutoff time t on the subtitle timeline. M3 The difference between the target audio segment and t5 is 1.3 seconds, and the starting time is t3. The maximum compression amount of the blank segment is limited by the offset threshold, which can prevent the target audio segment from being too far ahead or behind the corresponding subtitle display time.
[0076] Then, after pushing M3 to the audio queue, Figure 7 As shown, the segment cutoff time t on the subtitle timeline can be aligned according to M3 M3 The second deviation (ie, t6-t M3 ), fill the blank audio segment (K3) after M3 in the audio queue, wherein the segment length of the blank audio segment (K3) is t6-t M3 It should be understood that if M3 is aligned to the end time of the segment on the subtitle timeline, M3 The second deviation t6-t from the start time t6 of the fourth translated subtitle text display on the subtitle timeline M3 If it is 0, then no blank audio segment will be filled after M3 in the audio queue.
[0077] It should be understood that the implementation method of the above-mentioned steps S13061 to S13063 is equivalent to giving priority to moving the target audio segment forward when there are sufficient blank segments before and after the initial audio segment, that is, giving priority to compressing the blank segments before the initial audio segment, and then occupying the blank segments after the initial audio segment; in actual applications, the target audio segment can also be given priority to moving backward based on the offset threshold, that is, giving priority to occupying the blank segments after the initial audio segment, that is, first occupying as many blank segments after the initial audio segment as possible within the allowed range, and then compressing the blank segments before the initial audio segment as needed. The embodiments of the present disclosure do not limit this.
[0078] It should be understood that if there is a sufficient blank segment only after the i-th initial audio segment, that is, there is no blank segment before the i-th initial audio segment, the target audio segment corresponding to the i-th translated subtitle text can be directly pushed into the audio queue. At this time, the target audio segment corresponding to the i-th translated subtitle text is arranged after the target audio segment corresponding to the i-1-th translated subtitle text; of course, it can also be limited that the target audio segment’s occupancy of the blank segment after the initial audio segment cannot exceed the offset threshold. If the occupancy exceeds the offset threshold, you can refer to the processing method when there are no sufficient blank segments before and after the initial audio segment in the following text to obtain a target audio segment that meets the requirements.
[0079] Among them, it is also possible that the target audio segment occupies a larger amount of the blank segment after the i-th initial audio segment than the duration of the blank segment after the i-th initial audio segment. For example, if the acceleration offset is 2.8 seconds, the offset threshold is 2 seconds, and the duration of K2 is 3 seconds, then the compression amount of K2 is the maximum compression amount of 2 seconds limited by the offset threshold. The difference between the acceleration offset of 2.8 seconds and the compression amount of 2 is 0.8 seconds, which is the occupancy of the blank segment after P3. If the duration of the blank segment after P3 (the duration between t5 and t6) is 0.7 seconds, it means that the occupancy of 0.8 seconds is greater than the duration between t5 and t6. That is to say, although the duration of K2 is 3 seconds, The sum of the duration of 0.7 seconds between t5 and t6, 3.7 seconds, is greater than the accelerated offset of 2.8 seconds of the initial audio segment. However, under the limitation of the above-mentioned offset threshold, there is not enough blank space between t5 and t6 for the target audio segment M3 to occupy. In this case, the processing method described later when there are not enough blank segments before and after the i-th initial audio segment can be adopted to obtain a target audio segment that meets the requirements. That is, when the target audio segment occupies the blank segment after the i-th initial audio segment longer than the duration of the blank segment after the i-th initial audio segment, it can also be considered that there are not enough blank segments before and after the i-th initial audio segment.
[0080] In one possible implementation, step S13, processing the i-th initial audio segment into a target audio segment that matches the subtitle display time of the i-th translated subtitle text based on the segment duration of the i-th initial audio segment, the subtitle display time corresponding to the i-th translated subtitle text, and the subtitle character identifier, may include:
[0081] Step S1307: If the acceleration ratio corresponding to the i-th initial audio segment is greater than the acceleration ratio threshold and there are no sufficient blank segments before or after the i-th initial audio segment, determine, based on the subtitle timeline, whether the i-th translated subtitle text and the (i+1)-th translated subtitle text correspond to the same subtitle character identifier. The absence of sufficient blank segments includes the following: there are no blank segments before or after the i-th initial audio segment, or the duration of the existing blank segments is less than the acceleration offset corresponding to the i-th initial audio segment, or the amount of blank segments occupied by the target audio segment after the i-th initial audio segment is greater than the duration of the blank segments after the i-th initial audio segment.
[0082] Step S1308: When the i-th translated subtitle text and the (i+1)-th translated subtitle text correspond to the same subtitle character identifier, the i-th initial audio segment corresponding to the i-th translated subtitle text and the (i+1)-th initial audio segment corresponding to the (i+1)-th translated subtitle text are merged to obtain a merged audio segment.
[0083] Step S1309: Process the merged audio segment into a target audio segment that matches the merged display time based on the segment duration of the merged audio segment, the merged display time corresponding to the i-th translated subtitle text and the i+1-th translated subtitle text, and the subtitle character identifier; wherein the merged display time includes the start time of the i-th translated subtitle text display to the end time of the i+1-th translated subtitle text display.
[0084] In step S1307, it is determined whether the i-th translated subtitle text and the i+1-th translated subtitle text correspond to the same subtitle character identifier, that is, it is determined whether the i-th translated subtitle text and the i+1-th translated subtitle text are spoken by the same character; in step S1308, the two initial audio clips are merged, that is, the two initial audio clips are spliced. The embodiment of the present disclosure does not limit the merging method of the audio clips; in step S1309, the merged audio clip can be used as the new i-th initial audio clip, and with reference to the processing process of the above steps S1300 to S1309, the merged audio clip is processed into a target audio clip that matches the merged display time according to the clip length of the merged audio clip, the merged display time corresponding to the i-th translated subtitle text and the i+1-th translated subtitle text, and the subtitle character identifier. No further details are given here.
[0085] In actual application, after obtaining the target audio segment that matches the merged display time, the method also includes: step S1310, pushing the target audio segment that matches the merged display time into the audio queue, and based on the third deviation between the segment cutoff time of the target audio segment that matches the merged display time and the start time of the i+2th translated subtitle text display on the subtitle timeline (that is, after the target audio segment is aligned with the start time of the i-th translated subtitle text), filling in the blank audio segment corresponding to the third deviation after the target audio segment that matches the merged display time in the audio queue.
[0086] For example, Figure 8 As shown in the figure, if the acceleration ratio of the fourth initial audio segment P4 corresponding to the fourth translated subtitle text "lady, be careful" of the fourth original subtitle text "Girl be careful" is greater than the acceleration ratio threshold and there is no sufficient blank segment before and after the fourth initial audio segment P4, it can be determined based on the subtitle timeline that the fourth translated subtitle text and the fifth translated subtitle text "You should get on the sedan" corresponds to the same subtitle character identifier "speaker1", the 4th initial audio segment corresponding to the 4th translated subtitle text and the 5th initial audio segment P5 corresponding to the 5th translated subtitle text can be merged to obtain a merged audio segment P45; then, the merged audio segment P45 can be used as the new 4th initial audio segment, and with reference to the processing process of steps S1300 to S1309 above, the merged audio segment P45 can be processed into a target audio segment M45 that matches the merged display time (i.e., t6 to t9) according to the segment length of the merged audio segment P45, the merged display time corresponding to the 4th translated subtitle text and the 5th translated subtitle text, and the subtitle character identifier; then, the M45 can be pushed into the audio queue, and based on the segment end time t of the target audio segment M45 aligned to the subtitle timeline (i.e., the start time of M45 is aligned with the start time t6 of the 4th translated subtitle text), the target audio segment M45 is aligned to the subtitle timeline. M45 (such as t M45 is exactly equal to t9) and the start time t of the sixth subtitle text display on the subtitle timeline 10 The third deviation t 10 -t M45 , fill the blank audio segment (K4) corresponding to the third deviation after the target audio segment M45 that matches the merged display time in the audio queue, wherein the segment length of K4 is t 10 -t M45 duration.
[0087] In actual applications, the maximum number of merged audio segments (that is, the maximum number of merged initial audio segments that can be merged) can be set. For example, if the maximum merge number is set to 2, and the number of merged initial audio segments in the merged audio segment has reached the maximum merge number, then when the acceleration ratio of the merged audio segment obtained by merging the i-th initial audio segment with the i+1-th initial audio segment is greater than the acceleration ratio threshold and there are no sufficient blank segments before and after the merged audio segment, it is no longer determined whether the subtitle character identifier of the i+1-th translated subtitle text corresponding to the merged audio segment is the same as the subtitle character identifier of the i+2-th translated subtitle text. Instead, refer to the following description of when the i-th initial audio segment is merged. The merged audio segment is processed into a target audio segment that matches the merged display time, provided that there are no sufficient blank segments before and after the initial audio segment, and that the i-th translated subtitle text and the i+1-th translated subtitle text correspond to different subtitle character identifiers. Specifically, the merged audio segment can be directly accelerated according to the acceleration ratio corresponding to the merged audio segment to obtain the target audio segment that matches the merged display time. Alternatively, the text content of the i-th translated subtitle text and the i+1-th translated subtitle text can be shortened to obtain a shortened merged audio segment, and the aforementioned processing steps S1300 to S1306 can be re-executed using the shortened merged audio segment as the new i-th initial audio segment. The text content of the i-th translated subtitle text and the i+1-th translated subtitle text can be shortened based on models and algorithms in related technologies that automatically reduce the number of words without changing the semantics. If the number of initial audio segments merged in the merged audio segment has not reached the maximum merged number, then when the acceleration ratio of the merged audio segment obtained by merging the i-th initial audio segment with the i+1-th initial audio segment is greater than the acceleration ratio threshold and there are no sufficient blank segments before and after the merged audio segment, the next initial audio segment can be merged until the maximum merged number is reached or the next initial audio segment and the i-th initial audio segment do not have the same subtitle character identifier, and the merged audio segment is used as the new i-th initial audio segment to re-execute the processing process from step S1300 to step S1306.
[0088] In one possible implementation, step S13, processing the i-th initial audio segment into a target audio segment that matches the subtitle display time of the i-th translated subtitle text based on the segment duration of the i-th initial audio segment, the subtitle display time corresponding to the i-th translated subtitle text, and the subtitle character identifier, may include:
[0089] Step S1311, when the acceleration ratio corresponding to the i-th initial audio segment is greater than the acceleration ratio threshold, there are no sufficient blank segments before and after the i-th initial audio segment, and the subtitle character identifiers corresponding to the i-th translated subtitle text and the (i + 1)-th translated subtitle text are different, accelerate the i-th initial audio segment according to the acceleration ratio corresponding to the i-th initial audio segment to obtain a target audio segment that matches the subtitle display time of the i-th translated subtitle text;
[0090] Step S1312, push the target audio segment corresponding to the i-th translated subtitle text into the audio queue, and based on the fourth deviation between the segment cut-off moment after aligning the target audio segment corresponding to the i-th translated subtitle text to the subtitle timeline (i.e., after the start moment of the target audio segment is aligned with the start moment of the i-th translated subtitle text) and the start moment of the (i + 1)-th translated subtitle text displayed on the subtitle timeline, fill in blank audio segments corresponding to the fourth deviation after the target audio segment corresponding to the i-th translated subtitle text in the audio queue. Exemplarily, as Figure 9 shown, for the 6th initial audio segment P6 corresponding to the 6th original subtitle text "Good luck shines brightly at the auspicious time" and the 6th English subtitle text "Lucky stars shine brightly and auspicious times arrive", if the acceleration ratio corresponding to this P6 is greater than the acceleration ratio threshold, there are no sufficient blank segments before and after this P6, and the subtitle character identifier corresponding to the 6th translated subtitle text is "speaker2" which is different from the subtitle character identifier "speaker3" corresponding to the 7th translated subtitle text "Lift the sedan", then directly accelerate the 6th initial audio segment P6 according to the acceleration ratio corresponding to the 6th initial audio segment P6. For example, if the segment duration of P6 is 5 seconds and the subtitle display duration corresponding to the 6th original subtitle text is 2 seconds, then the acceleration ratio is 2.5, and P6 can be accelerated at 2.5 times the speed to obtain a target audio segment M6 that matches the subtitle display time (i.e., t 10 and t 11 ). That is, P6 is accelerated into a target audio segment M6 at 2.5 times the speed, and then M6 can be added to the audio queue, and according to M6 aligned to the subtitle timeline (i.e., aligning the start moment of M6 to the start moment t 10 ) of the 6th original subtitle text, the segment cut-off moment t M6 (at this time t M6 = t 11 ) and the start moment t 12 of the 7th translated subtitle text displayed on the subtitle timeline, the fourth deviation (i.e., t 12 - t M6), fill the blank audio segment (K5) after M6 in the audio queue, where the segment length of K5 is t 12 -t M6 This method can be understood as follows: when the acceleration ratio corresponding to the i-th initial audio segment is greater than the acceleration ratio threshold, there are no sufficient blank segments before and after the i-th initial audio segment, and the i-th translated subtitle text and the (i+1)-th translated subtitle text correspond to different subtitle character identifiers, the acceleration ratio threshold is ignored, and the i-th initial audio segment is directly accelerated according to the acceleration ratio corresponding to the i-th initial audio segment to obtain a target audio segment that matches the subtitle display time of the i-th translated subtitle text.
[0091] Taking into account the situation where the acceleration ratio corresponding to the i-th initial audio segment is greater than the acceleration ratio threshold, there are no sufficient blank segments before and after the i-th initial audio segment, and the i-th translated subtitle text and the (i+1)-th translated subtitle text correspond to different subtitle character identifiers, the acceleration ratio threshold is directly ignored. The target audio segment generated by accelerating the i-th initial audio segment according to the acceleration ratio corresponding to the i-th initial audio segment may be too fast, affecting the user's listening experience. Therefore, in one possible implementation, the above-mentioned step S13, which processes the i-th initial audio segment into a target audio segment that matches the subtitle display time of the i-th translated subtitle text based on the segment length of the i-th initial audio segment, the subtitle display time corresponding to the i-th translated subtitle text, and the subtitle character identifier, may include:
[0092] Step S1313: If the acceleration ratio corresponding to the i-th initial audio segment is greater than the acceleration ratio threshold, there are no sufficient blank segments before and after the i-th initial audio segment, and the i-th translated subtitle text and the (i+1)-th translated subtitle text correspond to different subtitle character identifiers, shorten the text content of the i-th translated subtitle text, and obtain a shortened audio segment corresponding to the shortened i-th translated subtitle text.
[0093] Step S1314: Processing the shortened audio segment into a target audio segment that matches the shortened subtitle display time of the i-th translated subtitle text based on the segment duration of the shortened audio segment, the shortened subtitle display time of the i-th translated subtitle text, and the subtitle character identifier;
[0094] Step S1315: Push the target audio segment corresponding to the shortened i-th translated subtitle text into the audio queue, and based on the fifth deviation between the segment cutoff time of the target audio segment corresponding to the shortened i-th translated subtitle text and the start time of the i+1-th translated subtitle text display on the subtitle timeline (that is, after the start time of the target audio segment is aligned with the start time of the shortened i-th translated subtitle text), fill in the blank audio segment corresponding to the fifth deviation after the target audio segment corresponding to the shortened i-th translated subtitle text in the audio queue.
[0095] In step S1313, shortening the text content of the i-th translated subtitle text can be achieved by, for example, removing uncritical words in the i-th translated subtitle text, or replacing longer words in the i-th translated subtitle text with shorter words, etc. As long as the text content of the i-th translated subtitle text can be shortened without changing the original semantics, the present embodiment does not impose any restrictions on this.
[0096] In step S1313, the shortened audio segment corresponding to the i-th translated subtitle text can be obtained by referring to the above-mentioned method for obtaining the initial audio segment. Specifically, it can be obtained by manual dubbing, or it can also be obtained by using audio synthesis technology known in the art. Based on the audio features such as the tone and timbre of the characters' speech in the original audio segment of the original subtitle text, the shortened translated subtitle text can be synthesized into the corresponding shortened audio segment. This embodiment of the present disclosure does not limit this.
[0097] In step S1314, the shortened audio segment can be used as the new i-th initial audio segment. Referring to the processing steps S1300 to S1312 above, the shortened audio segment is processed into a target audio segment that matches the shortened subtitle display time of the i-th translated subtitle text based on the segment duration of the shortened audio segment, the shortened subtitle display time of the i-th translated subtitle text, and the subtitle character identifier. This is not further described here. The shortened subtitle display time of the i-th translated subtitle text remains the subtitle display time of the i-th original subtitle text as specified on the original subtitle timeline.
[0098] In step S1315, the implementation method of the above-mentioned step S1304 can be referred to to push the target audio segment corresponding to the shortened i-th translated subtitle text into the audio queue, and based on the alignment of the target audio segment corresponding to the shortened i-th translated subtitle text to the fifth deviation between the segment end time on the subtitle timeline and the start time of the display of the i+1-th translated subtitle text on the subtitle timeline, a blank audio segment corresponding to the fifth deviation is filled after the target audio segment corresponding to the shortened i-th translated subtitle text in the audio queue, which will not be elaborated here.
[0099] Optionally, based on the above steps S1300 to S1312, the embodiment of the present disclosure provides Figure 10 The flowchart of a hard synchronization strategy of audio and video is shown in FIG. Figure 10 As shown, the audio and video hard synchronization strategy includes:
[0100] Frequency segment Pi;
[0101] Step S902: Determine whether the duration of the i-th initial audio segment Pi is less than or equal to the subtitle display duration Ti of the i-th translated subtitle text; if so, proceed to step S903; if not, proceed to step S904;
[0102] Step S903: determine the i-th initial audio segment Pi as the target audio segment Mi that matches the subtitle display time of the i-th translated subtitle text, and proceed to step S906;
[0103] Step S904: determine whether the acceleration ratio corresponding to the i-th initial audio segment Pi is less than or equal to a preset acceleration ratio threshold; if so, proceed to step S905; if not, proceed to step S907;
[0104] Step S905: Accelerate the i-th initial audio segment Pi according to the acceleration ratio corresponding to the i-th initial audio segment Pi to obtain a target audio segment Mi that matches the subtitle display time of the i-th translated subtitle text, and then proceed to step S906;
[0105] Step S906: Push the target audio segment Mi into the audio queue, and fill in blank audio segments after Mi;
[0106] Step S907, determine whether there are sufficient blank segments before and after the i-th initial audio segment Pi; if so, proceed to step S908; if not, proceed to step S912;
[0107] Step S908: Determine the required compression amount for the blank segment before the i-th initial audio segment Pi and the amount of the blank segment after the i-th initial audio segment Pi that the target audio segment Mi occupies based on the preset offset threshold, the acceleration offset, and the duration of the blank segment before the i-th initial audio segment.
[0108] Step S909: Determine whether the target audio segment occupies the blank segment after the i-th initial audio segment less than or equal to the duration Ti0 of the blank segment after the i-th initial audio segment; if so, proceed to step S910; if not, proceed to step S912;
[0109] Step S910: compressing the blank audio segment Ki-1 filled before the i-th initial audio segment Pi in the audio queue based on the compression amount required for the blank segment;
[0110] Step S911: Accelerate the i-th initial audio segment Pi according to the acceleration ratio threshold to obtain a target audio segment Mi that matches the subtitle display time of the i-th translated subtitle text, and proceed to step S906;
[0111] Step S912: Determine whether the i-th translated subtitle text Zi and the i+1-th translated subtitle text Zi+1 correspond to the same subtitle character identifier; if so, proceed to step S913; if not, proceed to step S914;
[0112] Step S913: Merge the i-th initial audio segment Pi corresponding to the i-th translated subtitle text with the i+1-th initial audio segment Pi+1 corresponding to the i+1-th translated subtitle text to obtain a merged audio segment. The merged audio segment can then be used as the new i-th initial audio segment, and the process can be returned to step S902 to re-execute the audio-video hard synchronization process.
[0113] Step S914: Accelerate the i-th initial audio segment Pi according to the acceleration ratio corresponding to the i-th initial audio segment Pi to obtain a target audio segment Mi that matches the subtitle display time of the i-th translated subtitle text, and proceed to step S906.
[0114] The above-mentioned audio and video hard synchronization process allows the initial audio segment corresponding to each translated subtitle text to be slightly accelerated, merged, or shifted forward or backward (i.e., compressed or occupied with blank segments) relative to the subtitle display time indicated by the subtitle timeline, thereby ensuring that each translated subtitle text and the target audio segment can basically achieve audio and video hard synchronization, and can adapt to audio and video synchronization tasks with slightly higher subtitle density, and can achieve deep fusion of video audio and video.
[0115] In the above-mentioned audio-visual hard synchronization strategy, when the segment length of the i-th initial audio segment is greater than the subtitle display length of the i-th translated subtitle text, priority is given to acceleration processing, then axis shift processing, and then merging processing; in a possible implementation, when the subtitle density (i.e., the density of the subtitle text) is small or the requirements for audio-visual synchronization are not high (i.e., a small amount of audio-visual offset is allowed), priority can also be given to merging processing, then acceleration processing and axis shift processing. Therefore, in a possible implementation, the above-mentioned step S13, based on the segment length of the i-th initial audio segment, the subtitle display time corresponding to the i-th translated subtitle text, and the subtitle character identification, processes the i-th initial audio segment into a target audio segment that matches the subtitle display time of the i-th translated subtitle text, including:
[0116] Step S1320: If the duration of the i-th initial audio segment is greater than the subtitle display duration of the i-th translated subtitle text, determining, based on the subtitle timeline, whether the i+1-th translated subtitle text and the i-th translated subtitle text correspond to the same subtitle character identifier;
[0117] Step S1321: When the i-th translated subtitle text and the (i+1)-th translated subtitle text correspond to the same subtitle character identifier, merge the i-th initial audio segment corresponding to the i-th translated subtitle text and the (i+1)-th initial audio segment corresponding to the (i+1)-th translated subtitle text to obtain a first merged audio segment.
[0118] Step S1322: Process the first merged audio segment into a target audio segment that matches the first merged display time based on the segment length of the first merged audio segment, the first merged display time corresponding to the i-th translated subtitle text and the i+1-th translated subtitle text, and the subtitle character identifier; wherein the first merged display time includes the start time of the i-th translated subtitle text display and the end time of the i+1-th translated subtitle text display.
[0119] In one possible implementation, in step S1322, based on the duration of the first merged audio segment, the first merged display time corresponding to the i-th translated subtitle text and the (i+1)-th translated subtitle text, and the subtitle character identifier, the first merged audio segment is processed into a target audio segment that matches the first merged display time, including:
[0120] Step S13221: If the duration of the first merged audio segment is less than or equal to the merged display duration corresponding to the first merged display time, determine the first merged audio segment as a target audio segment that matches the first merged display time;
[0121] Step S13222: If the duration of the first merged audio segment is greater than the combined display duration of the first combined display time, and the number of initial audio segments merged into the first merged audio segment is less than a preset merge quantity threshold, determining, based on the subtitle timeline, whether the (i+2)th translated subtitle text and the (i+1)th translated subtitle text correspond to the same subtitle character identifier;
[0122] Step S13223: When the (i+2)th translated subtitle text and the (i+1)th translated subtitle text correspond to the same subtitle character identifier, the ratio between the segment duration of the first merged audio segment and the combined display duration of the first combined display time is determined as the speedup ratio corresponding to the first merged audio segment.
[0123] Step S13224: When the acceleration ratio corresponding to the first merged audio segment is less than or equal to a preset acceleration ratio threshold, the first merged audio segment is accelerated according to the acceleration ratio corresponding to the first merged audio segment to obtain a target audio segment that matches the first merged display time.
[0124] In actual application, after obtaining the target audio segment that matches the first merged display time, the method may further include: step S1323, pushing the target audio segment that matches the first merged display time into the audio queue, and based on the target audio segment that matches the first merged display time being aligned to the sixth deviation between the segment cutoff moment on the subtitle timeline and the start moment of the display of the i+2th translated subtitle text on the subtitle timeline, filling the blank audio segment corresponding to the sixth deviation after the target audio segment that matches the first merged display time in the audio queue.
[0125] For example, Figure 11 As shown, the segment duration of the second initial audio segment P2 corresponding to the second translated subtitle text "young lady" is greater than the subtitle display duration of the second translated subtitle text (i.e., the duration between t2 and t3). At this time, based on the subtitle timeline, it is determined that the third translated subtitle text "It's time to get married" corresponds to the same subtitle character identifier "speaker1" as the second translated subtitle text, and the second initial audio segment P2 corresponding to the second translated subtitle text and the third initial audio segment P3 corresponding to the third translated subtitle text are merged to obtain a first merged audio segment P23. If the segment duration of the first merged audio segment P23 is less than or equal to the display duration corresponding to the first merged display time (i.e., the duration between t2 and t5), P23 can be determined as the target audio segment M23 that matches the first merged display time, and M23 is pushed into the audio queue, and the segment end time t is aligned to the subtitle timeline (i.e., the start time of M23 is aligned with the start time t2 of the second translated subtitle text) according to M23. M23 The sixth deviation (ie, t6-t M23 ), fill the blank audio segment (K2) after M23 in the audio queue, where the segment length of K2 is t6-t M23 It should be understood that if M23 is aligned to the end time of the segment on the subtitle timeline, M23 If the sixth deviation from the start time t6 of the display of the fourth translated subtitle text on the subtitle timeline is 0, then a blank audio segment may not be filled after M23 in the audio queue.
[0126] For example, Figure 12As shown, if the duration of the first merged audio segment P23 is greater than the display duration corresponding to the first merged display time (i.e., the duration between t2 and t5), it can be determined whether the number of initial audio segments merged in the first merged audio segment is less than or equal to the preset merge quantity threshold. Among them, those skilled in the art can set the specific value of the merge quantity threshold according to actual needs. For example, it can be set to 4, that is, a maximum of 4 initial audio segments are merged; if the first merged audio segment P23 merges two initial audio segments, which is less than the merge quantity threshold (e.g., 4), it can be determined that the fourth translated subtitle text "lady, be Whether "careful" corresponds to the same subtitle character identifier as the third translated subtitle text; if the subtitle character identifier "speaker2" corresponding to the fourth translated subtitle text is different from the subtitle character identifier "speaker3" corresponding to the third translated subtitle text, then the ratio between the segment length of the first merged audio segment P23 and the display length of the merged display time (i.e., the duration between t2 and t5) is calculated, and determined as the acceleration ratio corresponding to the first merged audio segment P23; then, it is determined whether the acceleration ratio corresponding to the first merged audio segment P23 is less than or equal to a preset acceleration ratio threshold; if it is less than or equal to the preset acceleration ratio threshold, the first merged audio segment P23 can be accelerated according to the acceleration ratio corresponding to the first merged audio segment P23 (i.e., P23 is doubled in speed according to the acceleration ratio) to obtain a target audio segment M23 that matches the first merged display time; then, M23 can be pushed to the audio queue, and aligned to the segment cutoff time t on the subtitle timeline according to M23. M23 (At this time t M23 = t5) and the start time t6 of the fourth translated subtitle text display on the subtitle time axis (ie, t6-t M23 =t6-t5), fill the blank audio segment (K2) after M23 in the audio queue, and the segment length of K2 is t6-t M23 duration.
[0127] In one possible implementation, in step S1322, based on the duration of the first merged audio segment, the first merged display time corresponding to the i-th translated subtitle text and the (i+1)-th translated subtitle text, and the subtitle character identifier, processing the first merged audio segment into a target audio segment that matches the first merged display time may further include:
[0128] Step S13225: When the segment duration of the first merged audio segment is greater than the display duration of the first merged display time and the number of initial audio segments merged in the first merged audio segment is equal to the merge quantity threshold, determine the ratio between the segment duration of the first merged audio segment and the display duration of the first merged display time as the speedup ratio corresponding to the first merged audio segment;
[0129] Step S13226: When the acceleration ratio corresponding to the first merged audio segment is less than or equal to a preset acceleration ratio threshold, the first merged audio segment is accelerated according to the acceleration ratio corresponding to the first merged audio segment to obtain a target audio segment that matches the first merged display time.
[0130] For example, if the segment duration of the first merged audio segment P23 is greater than the display duration corresponding to the first merged display time (i.e., the duration between t2 and t5), and the number of initial audio segments merged in the first merged audio segment P23 is equal to the merge quantity threshold (e.g., 2), that is, the number of initial audio segments merged in the first merged audio segment reaches the upper limit, then there is no need to determine whether the fourth translated subtitle text and the third translated subtitle text correspond to the same subtitle character identifier, and the difference between the segment duration of the first merged audio segment P23 and the display duration of the first merged display time (i.e., the duration between t2 and t5) can be directly calculated. The ratio between them is determined as the acceleration ratio corresponding to the first merged audio segment P23, and then it is determined whether the acceleration ratio corresponding to the first merged audio segment P23 is less than or equal to the preset acceleration ratio threshold. If the acceleration ratio corresponding to P23 is less than or equal to the preset acceleration ratio threshold, the first merged audio segment P23 can be accelerated according to the acceleration ratio corresponding to the first merged audio segment P23 (that is, P23 is doubled in speed according to the acceleration ratio) to obtain the target audio segment M23 that matches the first merged display time; then M23 can be pushed to the audio queue and aligned to the segment end time t on the subtitle timeline according to M23. M23 (At this time t M23 = t5) and the start time t6 of the fourth translated subtitle text display on the subtitle time axis (ie, t6-t M23 =t6-t5), fill the blank audio segment (K2) after M23 in the audio queue, and the segment length of K2 is t6-t M23 duration.
[0131] In one possible implementation, in step S1322, based on the duration of the first merged audio segment, the first merged display time corresponding to the i-th translated subtitle text and the (i+1)-th translated subtitle text, and the subtitle character identifier, processing the first merged audio segment into a target audio segment that matches the first merged display time may further include:
[0132] Step S13227: If the acceleration ratio corresponding to the first merged audio segment is greater than the acceleration ratio threshold, the first merged audio segment is accelerated according to the acceleration ratio threshold to obtain a target audio segment corresponding to the first merged display time;
[0133] Step S13228, based on the length offset of the target audio segment corresponding to the first merged display time relative to the merged display time of the first merged display time, compress the blank audio segment before the first merged audio segment in the audio queue, and push the target audio segment corresponding to the first merged display time into the audio queue, wherein the target audio segment pushed into the audio queue corresponds to the segment cutoff time on the subtitle timeline and is aligned with the cutoff time in the first merged display time.
[0134] For example, Figure 13 As shown, if the acceleration ratio corresponding to the first merged audio segment P23 is greater than the preset acceleration ratio threshold, the first merged audio segment P23 is accelerated according to the acceleration ratio threshold to obtain the corresponding target audio segment M23; since the acceleration ratio corresponding to the first merged audio segment P23 is greater than the acceleration ratio threshold, the segment length of the target audio segment M23 after acceleration according to the acceleration ratio threshold is still greater than the merged display length of the first merged display time (that is, the length between t2 and t5), at this time, the difference between the segment length of the target audio segment M23 and the first merged display length (that is, the length between t2 and t5) can be calculated to obtain the duration deviation For example, if M23 is 10.5 seconds and the duration between t2 and t5 is 10 seconds, the duration offset is 0.5 seconds; then the blank audio segment before the first merged audio segment in the audio queue can be compressed according to the duration offset, that is, the blank audio segment before the first merged audio segment can be shortened by the duration offset. For example, if the duration offset is 0.5 seconds, the blank audio segment (K1) can be shortened by 0.5 seconds, and then the target audio segment M23 can be pushed into the audio queue. At this time, the target audio segment M23 pushed into the audio queue corresponds to the segment cutoff time t on the subtitle timeline. M23 Aligned with the end time t5 in the combined display time (ie t M23 =t5).
[0135] In actual applications, the blank audio segment before the first merged audio segment may include any blank audio segment before the first merged audio segment in the audio queue, and is not limited to the previous adjacent blank audio segment; in this method, there is no need to set a maximum compression limit for the compression amount of the blank audio segment, as long as the duration offset requirement of the target audio segment relative to the merged display duration can be met, which is equivalent to moving the target audio segment forward relative to the merged display time.
[0136] It may also happen that the number of initial audio segments merged in the first merged audio segment is less than a merging number threshold. Therefore, in a possible implementation, the method may further include:
[0137] Step S1325: If the segment duration of the first merged audio segment is greater than the combined display duration of the first combined display time, the number of initial audio segments merged in the first merged audio segment is less than a preset merge quantity threshold, and the (i+2)th translated subtitle text and the (i+1)th translated subtitle text correspond to the same subtitle character identifier, the first merged audio segment is merged with the (i+2)th initial audio segment corresponding to the (i+2)th translated subtitle text to obtain a second merged audio segment.
[0138] Step S1326: Process the second merged audio segment into a target audio segment that matches the second merged display time based on the segment length of the second merged audio segment, the second merged display time corresponding to the i-th translated subtitle text to the i+2-th translated subtitle text, and the subtitle character identifier.
[0139] For example, Figure 14 As shown, if the duration of the first merged audio segment P23 is greater than the display duration corresponding to the first merged display time (i.e., the duration between t2 and t5), it can be determined whether the number of initial audio segments merged in the first merged audio segment is less than or equal to the preset merge quantity threshold; if the first merged audio segment P23 merges two initial audio segments less than the preset merge quantity threshold (such as 4), it can be determined whether the 4th translated subtitle text and the 3rd translated subtitle text correspond to the same subtitle character identifier. If the 4th translated subtitle text and the 3rd translated subtitle text correspond to the same subtitle character identifier, then the first merged audio segment P23 can continue to be merged with the 4th initial audio segment corresponding to the 4th translated subtitle text. P4 is merged to obtain a second merged audio segment P234; then the second merged audio segment P234 can be used as a new first merged audio segment, and referring to the processing flow of the first merged audio segment in the above step S1322, the second merged audio segment is processed into a target audio segment M234 that matches the second merged display time. Specifically, if the segment length of the second merged audio segment P234 is less than or equal to the display length corresponding to the second merged display time (that is, the length between t2 and t7), then P234 can be determined as the target audio segment M234 that matches the second merged display time, and M234 can be pushed into the audio queue, and aligned to the segment cutoff time t on the subtitle timeline according to M234. M234 The sixth deviation (ie, t8-t M234 ), fill the blank audio segment (K2) after M234 in the audio queue, where the segment length of K2 is t8-t M234duration; if the segment duration of the second merged audio segment P234 is greater than the display duration corresponding to the second merged display time (i.e. the duration between t2 and t7), then determine whether the number of initial audio segments merged in the second merged audio segment is less than or equal to the preset merge quantity threshold; if less than, then determine whether the 5th translated subtitle text and the 4th translated subtitle text correspond to the same subtitle character identifier; if so, then the second merged audio segment and the 5th initial audio segment corresponding to the 5th translated subtitle text can continue to be merged until the preset merge quantity threshold is reached, or they correspond to different subtitle character identifiers; then, an acceleration ratio calculation can be performed on the merged audio segment formed by merging multiple initial audio segments, and the merged audio segment can be processed into the corresponding target audio segment based on the acceleration ratio, and pushed into the audio queue to fill in blank audio segments and other processing.
[0140] In one possible implementation, the duration of the i-th initial audio segment may be greater than the subtitle display duration of the i-th translated subtitle text, and the i-th translated subtitle text and the (i+1)-th translated subtitle text may correspond to different subtitle character identifiers. Therefore, step S13, which processes the i-th initial audio segment into a target audio segment that matches the subtitle display duration of the i-th translated subtitle text based on the duration of the i-th initial audio segment, the subtitle display duration corresponding to the i-th translated subtitle text, and the subtitle character identifier, may include:
[0141] Step S1327: If the duration of the i-th initial audio segment is greater than the subtitle display duration of the i-th translated subtitle text, and the i-th translated subtitle text and the (i+1)-th translated subtitle text correspond to different subtitle character identifiers, the ratio between the duration of the i-th initial audio segment and the subtitle display duration corresponding to the i-th translated subtitle text is determined as the speedup ratio corresponding to the i-th initial audio segment.
[0142] Step S1328: If the acceleration ratio corresponding to the i-th initial audio segment is less than or equal to a preset acceleration ratio threshold, the i-th initial audio segment is accelerated according to the acceleration ratio corresponding to the i-th initial audio segment to obtain a target audio segment that matches the subtitle display time of the i-th translated subtitle text;
[0143] Step S1329: Push the target audio segment corresponding to the i-th translated subtitle text into the audio queue, and based on the alignment of the target audio segment corresponding to the i-th translated subtitle text to the seventh deviation between the segment cutoff moment on the subtitle timeline and the start moment of display of the i+1-th translated subtitle text on the subtitle timeline, fill in the blank audio segment corresponding to the seventh deviation after the target audio segment corresponding to the i-th translated subtitle text in the audio queue.
[0144] For example, Figure 15 As shown in the figure, if the fifth initial audio segment P5 corresponding to the fifth translated subtitle text "You should get on the sedan" is longer than the subtitle display duration of the fifth translated subtitle text (i.e., the duration between t8 and t9) and the subtitle character identifier "spraker2" corresponding to the fifth translated subtitle text is the same as the sixth translated subtitle text "Luckystars shine brightly and auspicious times arrive" is different from the subtitle character identifier "spraker3". At this time, the ratio of the segment duration of the fifth initial audio segment P5 to the subtitle display duration of the fifth translated subtitle text (that is, the duration between t8 and t9) can be calculated as the acceleration ratio corresponding to the fifth initial audio segment P5; if the acceleration ratio corresponding to P5 is less than or equal to the preset acceleration ratio threshold, P5 is accelerated according to the acceleration ratio corresponding to P5 to obtain a target audio segment M5 that matches the subtitle display time of the fifth translated subtitle text; then, M5 can be pushed to the audio queue and aligned to the subtitle timeline based on the segment end time t M5 (At this time t M5 = t9) and the start time t of the sixth subtitle text display on the subtitle timeline 10 The seventh deviation between t 10 -t M5 , fill in the blank audio segment (K3) corresponding to the seventh deviation after M5 in the audio queue, and the segment length of K3 is t 10 -t M5 duration.
[0145] In one possible implementation, processing the i-th original audio segment into a target audio segment that matches the subtitle display time of the i-th translated subtitle text based on the segment duration of the i-th original audio segment, the subtitle display time corresponding to the i-th translated subtitle text, and the subtitle character identifier includes:
[0146] Step S1330: If the acceleration ratio corresponding to the i-th initial audio segment is greater than the acceleration ratio threshold, the i-th initial audio segment is accelerated according to the acceleration ratio threshold to obtain a target audio segment corresponding to the i-th translated subtitle text;
[0147] Step S1331: Based on the duration offset of the target audio segment corresponding to the i-th translated subtitle text relative to the subtitle display duration of the i-th translated subtitle text, compress the blank audio segment before the i-th initial audio segment in the audio queue, and push the target audio segment corresponding to the i-th translated subtitle text into the audio queue, wherein the target audio segment pushed into the audio queue is aligned with the segment cutoff time on the subtitle timeline and the cutoff time in the subtitle display time of the i-th translated subtitle text.
[0148] For example, Figure 16 As shown in the figure, if the segment length of the initial audio segment P6 corresponding to the sixth translated subtitle text "Lucky stars shine brightly and auspicious times arrive" is longer than the subtitle display length of the sixth translated subtitle text (ie, t 10 to t 11 The duration between the 6th translated subtitle text and the subtitle character identifier "spraker3" corresponding to the 7th translated subtitle text is different from the subtitle character identifier "spraker4" corresponding to the 7th translated subtitle text, and the acceleration ratio corresponding to the 6th initial audio segment P6 is greater than the acceleration ratio threshold, then the 6th initial audio segment can be accelerated according to the acceleration ratio threshold to obtain the target audio segment M6 corresponding to the 6th translated subtitle text, and then the segment duration of M6 and the subtitle display duration of the 6th translated subtitle text (i.e., t 10 to t 11 The difference between the duration of the segment M6 and the duration of the subtitle display duration of the sixth translated subtitle text is used as the duration offset of the segment M6 relative to the subtitle display duration of the sixth translated subtitle text. Then, it can be determined whether the duration offset is less than or equal to the blank audio segment before the sixth initial audio segment in the audio queue (such as Figure 15 K3 in the 6th initial audio segment), if the duration offset is less than or equal to the duration of the blank audio segment before the 6th initial audio segment, the blank audio segment before the 6th initial audio segment in the audio queue can be compressed according to the duration offset, that is, the blank audio segment before the 6th initial audio segment is shortened by the duration offset. For example, if the duration offset is 2.5 seconds, and the duration offset of 2.5 seconds is less than or equal to the duration of the blank audio segment K3, the blank audio segment K3 can be compressed by 2.5 seconds, and then the target audio segment M6 can be pushed into the audio queue. If the duration of the blank audio segment K3 is equal to the duration offset of 2.5 seconds, it is equivalent to removing the blank audio segment K3 from the audio queue. Then, after pushing the target audio segment M6 to the audio queue, it can be displayed Figure 16 The arrangement order shown is shown. At this time, the target audio segment M6 pushed into the audio queue corresponds to the segment cutoff time t on the subtitle timeline. M6The end time t in the subtitle display time 11 Alignment (ie t M6 =t 11 ).
[0149] In actual applications, the above steps S1330 to S1331 can be performed when the duration offset of the target audio segment corresponding to the i-th translated subtitle text relative to the subtitle display duration of the i-th translated subtitle text is less than or equal to the duration of the blank audio segment adjacent to the i-th initial audio segment in the audio queue; for the case where the duration offset of the target audio segment corresponding to the i-th translated subtitle text relative to the subtitle display duration of the i-th translated subtitle text is greater than the duration of the blank audio segment adjacent to the i-th initial audio segment in the audio queue, for example, the above Figure 16 The duration offset corresponding to the target audio segment M6 is greater than the duration of the blank audio segment K3. In this case, the i-th initial audio segment can be processed into a target audio segment that matches the subtitle display time of the i-th translated subtitle text with reference to the implementation method of the above steps S1311 to S1312, or the above steps S1313 to S1315. Specifically, the i-th initial audio segment can be directly accelerated according to the acceleration ratio corresponding to the i-th initial audio segment to obtain the target audio segment that matches the subtitle display time of the i-th translated subtitle text and push it into the audio queue to fill in the blank audio segment; or, the text content of the i-th translated subtitle text can be shortened, and the shortened audio segment corresponding to the shortened i-th translated subtitle text can be obtained. Then, based on the segment duration of the shortened audio segment, the shortened subtitle display time of the i-th translated subtitle text, and the subtitle character identifier, the shortened audio segment can be processed into a target audio segment that matches the shortened subtitle display time of the i-th translated subtitle text and push it into the audio queue to fill in the blank audio segment. The embodiments of the present disclosure are not limited to this.
[0150] In actual applications, the blank audio segment before the initial audio segment may also include any blank audio segment before the initial audio segment in the audio queue, and is not limited to the previous adjacent blank audio segment; in this method, there is no need to set a maximum compression limit for the compression amount of the blank audio segment, as long as the requirement of the duration offset of the target audio segment relative to the subtitle display duration can be met, which is equivalent to moving the target audio segment forward relative to the subtitle display time.
[0151] Optionally, based on the above steps S1320 to S1331, the embodiment of the present disclosure provides Figure 17 The flowchart of a soft synchronization strategy of audio and video is shown in FIG. Figure 17 As shown, the audio and video soft synchronization strategy includes:
[0152] Step S1701: for the i-th translated subtitle text in the subtitle file, obtain the i-th initial audio segment Pi corresponding to the i-th translated subtitle text Zi;
[0153] Step S1702: Determine whether the duration of the i-th initial audio segment Pi is less than or equal to the subtitle display duration Ti of the i-th translated subtitle text; if so, proceed to step S1703; if not, proceed to step S1704;
[0154] Step S1703: determine the i-th initial audio segment Pi as the target audio segment Mi that matches the subtitle display time of the i-th translated subtitle text, and proceed to step S1709;
[0155] Step S1704: determine whether the i-th translated subtitle text Zi and the i+1-th translated subtitle text Zi+1 correspond to the same subtitle character identifier; if so, proceed to step S1710; if not, proceed to step S1705;
[0156] Step S1705: determine whether the acceleration ratio corresponding to the i-th initial audio segment Pi is less than or equal to a preset acceleration ratio threshold; if so, proceed to step S1706; if not, proceed to step S1707;
[0157] Step S1706: Accelerate the i-th initial audio segment Pi according to the acceleration ratio corresponding to the i-th initial audio segment Pi to obtain a target audio segment Mi that matches the subtitle display time of the i-th translated subtitle text, and then proceed to step S1709;
[0158] Step S1707: Accelerate the i-th initial audio segment Pi according to the acceleration ratio threshold to obtain a target audio segment Mi that matches the subtitle display time of the i-th translated subtitle text;
[0159] Step S1708: Compress the blank audio segment before the i-th initial audio segment in the audio queue based on the duration offset of the target audio segment Mi relative to the subtitle display duration of the i-th translated subtitle text;
[0160] Step S1709: Push the target audio segment Mi corresponding to the i-th translated subtitle text into the audio queue, and fill in a blank audio segment after the target audio segment Mi;
[0161] Step S1710: Merge the i-th initial audio segment Pi corresponding to the i-th translated subtitle text and the i+1-th initial audio segment Pi+1 corresponding to the i+1-th translated subtitle text to obtain a first merged audio segment Pii+1;
[0162] Step S1711, determining whether the duration of the first merged audio segment Pii+1 is less than the combined display duration Tii+1 of the first combined display time; if so, proceeding to step S1712; if not, proceeding to step S1713;
[0163] Step S1712: determine the first merged audio segment Pii+1 as the target audio segment Mii+1 that matches the first merged display time, and proceed to step S1720;
[0164] Step S1713: determine whether the number Nii+1 of the initial audio segments merged into the first merged audio segment Pii+1 is less than a preset merging number threshold; if so, proceed to step S1716; if not, proceed to step S1714;
[0165] Step S1714: Determine whether the i+1th translated subtitle text Zi+1 and the i+2th translated subtitle text Zi+2 correspond to the same subtitle character identifier; if so, proceed to step S1715; if not, proceed to step S1716;
[0166] Step S1715: Merge the first merged audio segment Pii+1 with the i+2th initial audio segment Pi+2 corresponding to the i+2th translated subtitle text to obtain a second merged audio segment Pii+1+2. This second merged audio segment can then be used as the new first merged audio segment and the process can be returned to step S1711 to perform the audio-visual soft synchronization process.
[0167] Step S1716: Determine whether the acceleration ratio corresponding to the first merged audio segment Pii+1 is less than or equal to a preset acceleration ratio threshold; if so, proceed to step S1717; if not, proceed to step S1718;
[0168] Step S1717: Accelerate the first merged audio segment Pii+1 according to the acceleration ratio corresponding to the first merged audio segment Pii+1 to obtain the target audio segment Mii+1 corresponding to the first merged display time, and then proceed to step S1720;
[0169] Step S1718: Accelerate the first merged audio segment Pii+1 according to the acceleration ratio threshold to obtain a first merged target audio segment Mii+1 with matching display time.
[0170] Step S1719: compress the blank audio segment before the first merged audio segment Pii+1 in the audio queue based on the duration offset of the target audio segment Mii+1 relative to the merged display duration of the first merged display time;
[0171] Step S1720: Push the target audio segment Mii+1 corresponding to the first merged display time into the audio queue, and fill in a blank audio segment after the target audio segment Mii+1.
[0172] The above-mentioned soft audio and video synchronization strategy can basically ensure soft audio and video synchronization between audio and video under the same subtitle character identifier "speaker", and has slightly lower requirements for subtitle density. Under the premise of audio and video synchronization, a small amount of subtle audio and video offset can be allowed when the same person speaks. The number of initial audio clips that can be merged by the soft audio and video synchronization strategy can be greater than the number of initial audio clips that can be merged by the above-mentioned hard audio and video synchronization strategy. For example, it can be twice the number of initial audio clips that can be merged by the hard audio and video synchronization strategy. This is not limited in the embodiments of the present disclosure.
[0173] In step S14, a target audio file synchronized with the video file is generated based on the target audio segments in the subtitle file that match the subtitle display time of each translated subtitle text.
[0174] As described above, an audio queue can be used to sequentially store target audio clips and blank audio clips. Thus, generating a target audio file synchronized with a video file based on target audio clips that match the subtitle display times of respective subtitle texts in the subtitle file includes exporting the target audio clips and blank audio clips sequentially arranged in the audio queue as the target audio file. The target audio file can be combined with the video file and the subtitle file to synthesize video data with subtitles and dubbing, so that the video data can be played back to achieve an audiovisual effect of synchronized sound and picture.
[0175] According to the audio-visual synchronization method of the embodiment of the present disclosure, by adopting the subtitle display time and subtitle character identification indicated by the subtitle timeline adopted by the original subtitle text of the video file, the initial audio segment corresponding to the translated subtitle text translated from the original subtitle text can be automatically processed into a target audio segment that matches the subtitle display time, thereby efficiently and quickly generating a target audio file synchronized with the video file, and achieving the audio-visual synchronization effect without manual intervention, greatly reducing labor costs.
[0176] As described above, the two audio and video synchronization strategies provided by the above embodiment of the present disclosure, namely the above steps S1300 to S1315 or Figure 10 The audio and video hard synchronization strategy shown, and the above steps S1320 to S1331 or Figure 17 The audio and video soft synchronization strategies shown are shown. In actual applications, different application scenarios can be set for the above two audio and video synchronization strategies. Therefore, the embodiment of the present disclosure also provides an audio and video synchronization method, including:
[0177] Step S21: Acquire a subtitle file corresponding to a video file, the subtitle file including a plurality of translated subtitle texts and a subtitle timeline; the translated subtitle text including a translation of the original subtitle text of the video file into a target language; the subtitle timeline being the timeline used by the original subtitle text of the video file, the subtitle timeline representing a subtitle display time and a subtitle character identifier, the subtitle character identifier representing a character speaking the subtitle text, and the subtitle display time including a start time and an end time of subtitle text display;
[0178] Step S22: When the subtitle density represented by the subtitle time axis is less than or equal to the specified density threshold, execute the above steps S1300 to S1315 or Figure 10 The audio-video hard synchronization strategy and step S14 shown are used to obtain a target audio file synchronized with the video file;
[0179] Step S23: When the subtitle density represented by the subtitle time axis is greater than the specified density threshold, execute the above steps S1320 to S1331 or Figure 17 The audio-video soft synchronization strategy and step S14 are shown to obtain a target audio file synchronized with the video file.
[0180] The subtitle density represents the density of the subtitle text; the denser the subtitle text, the greater the number of subtitles and the shorter the time interval between subtitles. It should be understood that those skilled in the art can use subtitle density detection algorithms known in the art to detect the subtitle density represented on the subtitle timeline, and this embodiment of the present disclosure is not limited to this. The specified density threshold can also be customized, and this embodiment of the present disclosure is not limited to this.
[0181] In actual applications, the above-mentioned audio and video hard synchronization strategy and the above-mentioned audio and video soft synchronization strategy can be encapsulated into two interfaces respectively. Users can call the required synchronization strategy through the interface according to the actual audio and video synchronization task requirements. For example, when the audio and video synchronization task requires strong synchronization of audio and video, the interface of the audio and video hard synchronization strategy can be called to implement the audio and video hard synchronization logic. When the audio and video synchronization task allows the overall audio and video to have no obvious perceptual offset, the interface of the audio and video soft synchronization strategy can be called to implement the audio and video hard synchronization logic. The embodiments of the present disclosure do not limit this.
[0182] According to the audio-visual synchronization method of the embodiment of the present disclosure, the corresponding audio-visual synchronization strategy can be automatically selected for different subtitle densities, and the thresholds required for acceleration, shifting, merging, etc. can be adaptively set according to the selected audio-visual synchronization strategy; and whether it is the audio-visual synchronization strategy selected according to the task adaptation, or the acceleration, shifting, merging, etc. parameter thresholds adaptively set according to the audio-visual synchronization strategy, they can be appropriately fine-tuned to meet the audio-visual requirements. During the entire audio-visual synchronization process, no human intervention is required, and audio-visual synchronization can be achieved with one click, which greatly reduces labor costs.
[0183] Figure 18 A block diagram of an audio-video synchronization device according to an embodiment of the present disclosure is shown. Figure 18 As shown, the device includes:
[0184] A first acquisition module 181 is configured to acquire a subtitle file corresponding to a video file, wherein the subtitle file includes a plurality of translated subtitle texts and a subtitle timeline; the translated subtitle text includes a translation of the original subtitle text of the video file into a target language; the subtitle timeline is the timeline used by the original subtitle text of the video file, and the subtitle timeline represents the subtitle display time and the subtitle character identifier of the subtitle text, wherein the subtitle character identifier represents the character speaking the subtitle text, and the subtitle display time includes the start and end times of the subtitle text display;
[0185] A second acquisition module 182 is configured to acquire, for an i-th translated subtitle text in the subtitle file, an i-th initial audio segment corresponding to the i-th translated subtitle text, and determine, based on the subtitle timeline, a subtitle display time and a subtitle character identifier corresponding to the i-th translated subtitle text; wherein the i-th initial audio segment is an audio segment generated based on the i-th translated subtitle text;
[0186] a processing module 183 configured to process the i-th initial audio segment into a target audio segment having a subtitle display time that matches the i-th translated subtitle text based on the segment duration of the i-th initial audio segment, the subtitle display time corresponding to the i-th translated subtitle text, and the subtitle character identifier;
[0187] The generating module 184 is configured to generate a target audio file synchronized with the video file based on target audio segments in the subtitle file that match the subtitle display time of each translated subtitle text.
[0188] In one possible implementation, the processing of the i-th initial audio segment into a target audio segment that matches the subtitle display time of the i-th translated subtitle text based on the segment duration of the i-th initial audio segment, the subtitle display time corresponding to the i-th translated subtitle text, and the subtitle character identification includes: when the segment duration of the i-th initial audio segment is less than or equal to the subtitle display time of the i-th translated subtitle text, determining the i-th initial audio segment as a target audio segment that matches the subtitle display time of the i-th translated subtitle text; wherein the subtitle display time of the i-th translated subtitle text is determined based on the subtitle display time of the i-th translated subtitle text.
[0189] In one possible implementation, the processing of the i-th initial audio segment into a target audio segment that matches the subtitle display time of the i-th translated subtitle text based on the segment duration of the i-th initial audio segment, the subtitle display time corresponding to the i-th translated subtitle text, and the subtitle character identifier includes: when the segment duration of the i-th initial audio segment is greater than the subtitle display time of the i-th translated subtitle text, determining the ratio between the segment duration of the i-th initial audio segment and the subtitle display time of the i-th translated subtitle text as the acceleration ratio corresponding to the i-th initial audio segment; and when the acceleration ratio corresponding to the i-th initial audio segment is less than or equal to a preset acceleration ratio threshold, accelerating the i-th initial audio segment according to the acceleration ratio corresponding to the i-th initial audio segment to obtain the target audio segment that matches the subtitle display time of the i-th translated subtitle text.
[0190] In one possible implementation, after obtaining the target audio segment that matches the subtitle display time of the i-th translated subtitle text, the device also includes: a first push filling module, used to: push the target audio segment into an audio queue, and based on the first deviation between the segment cutoff moment when the target audio segment is aligned to the subtitle timeline and the start moment of the display of the i+1-th translated subtitle text on the subtitle timeline, fill the target audio segment in the audio queue with a blank audio segment corresponding to the first deviation; wherein, the audio queue is used to store the target audio segment and the blank audio segment in sequence.
[0191] In a possible implementation, the processing of the i-th initial audio segment into a target audio segment that matches the subtitle display time of the i-th translated subtitle text according to the segment duration of the i-th initial audio segment, the subtitle display time corresponding to the i-th translated subtitle text, and the subtitle character identifier includes: when the acceleration ratio corresponding to the i-th initial audio segment is greater than the acceleration ratio threshold, judging whether there are sufficient blank segments before and after the i-th initial audio segment based on the subtitle display time of the i-th translated subtitle text; wherein the existence of sufficient blank segments includes that there are blank segments before and after the i-th initial audio segment and the duration of the blank segments is greater than or equal to the acceleration offset corresponding to the i-th initial audio segment; wherein the acceleration offset includes the difference between the audio segment duration obtained by accelerating the i-th initial audio segment using the acceleration ratio threshold and the duration of the i-th initial audio segment. The difference between the subtitle display durations corresponding to i translated subtitle texts; the duration of the blank segment before the i-th initial audio segment includes the duration of the blank audio segment filled before the i-th initial audio segment; the duration of the blank segment after the i-th initial audio segment includes the duration on the subtitle timeline from the end time of the display of the i-th translated subtitle text to the start time of the display of the i+1-th translated subtitle text; when there are sufficient blank segments before and after the i-th initial audio segment, the i-th initial audio segment is accelerated according to the acceleration ratio threshold to obtain a target audio segment that matches the subtitle display time of the i-th translated subtitle text; the target audio segment is pushed into the audio queue by occupying the blank segment before the i-th initial audio segment and / or occupying the blank segment after the i-th initial audio segment.
[0192] In a possible implementation, the method of pushing the target audio segment into the audio queue by occupying the blank segment before the i-th initial audio segment and / or occupying the blank segment after the i-th initial audio segment includes: when there are sufficient blank segments before and after the i-th initial audio segment, determining the required compression amount of the blank segment before the i-th initial audio segment and the occupation amount of the blank segment after the i-th initial audio segment by the target audio segment based on a preset offset threshold, the acceleration offset, and the duration of the blank segment before the i-th initial audio segment; wherein the offset threshold is used to limit the maximum compression amount of the blank segment; and when the occupation amount is less than or equal to the duration of the blank segment after the i-th initial audio segment, compressing the target audio segment after the i-1-th translation word in the audio queue based on the compression amount. The method comprises the following steps: compressing the blank audio segment filled after the target audio segment corresponding to the subtitle text; accelerating the i-th initial audio segment according to the acceleration ratio threshold to obtain a target audio segment matching the subtitle display time of the i-th translated subtitle text, and pushing the target audio segment matching the subtitle display time of the i-th translated subtitle text into the audio queue and arranging it after the compressed blank audio segment; based on the target audio segment corresponding to the i-th translated subtitle text in the audio queue being aligned to the second deviation between the segment end time on the subtitle timeline and the start time of the display of the i+1-th translated subtitle text on the subtitle timeline, filling the blank audio segment corresponding to the second deviation after the target audio segment corresponding to the i-th translated subtitle text in the audio queue; wherein, the audio queue is used to store target audio segments and blank audio segments in sequence.
[0193] In a possible implementation, the processing of the i-th initial audio segment into a target audio segment that matches the subtitle display time of the i-th translated subtitle text according to the segment duration of the i-th initial audio segment, the subtitle display time corresponding to the i-th translated subtitle text, and the subtitle character identifier includes: when the acceleration ratio corresponding to the i-th initial audio segment is greater than the acceleration ratio threshold and there are no sufficient blank segments before and after the i-th initial audio segment, based on the subtitle timeline, determining whether the i-th translated subtitle text and the (i+1)th translated subtitle text correspond to the same subtitle character identifier; wherein, the absence of sufficient blank segments includes that there are no blank segments before and after the i-th initial audio segment, or the duration of the existing blank segments is less than the acceleration offset corresponding to the i-th initial audio segment, or the target audio segment corresponds to the i+1th translated subtitle text. The amount of blank segments after the i-th initial audio segment is greater than the duration of the blank segment after the i-th initial audio segment; when the i-th translated subtitle text and the i+1-th translated subtitle text correspond to the same subtitle character identifier, the i-th initial audio segment corresponding to the i-th translated subtitle text and the i+1-th initial audio segment corresponding to the i+1-th translated subtitle text are merged to obtain a merged audio segment; based on the segment duration of the merged audio segment, the merged display time corresponding to the i-th translated subtitle text and the i+1-th translated subtitle text, and the subtitle character identifier, the merged audio segment is processed into a target audio segment that matches the merged display time; wherein the merged display time includes the start time of the display of the i-th translated subtitle text to the end time of the display of the i+1-th translated subtitle text.
[0194] In one possible implementation, after obtaining the target audio segment that matches the merged display time, the device also includes: a second push filling module, which is used to: push the target audio segment that matches the merged display time into the audio queue, and based on the target audio segment that matches the merged display time being aligned to the third deviation between the segment cutoff moment on the subtitle timeline and the start moment of display of the i+2th translated subtitle text on the subtitle timeline, fill the audio queue with a blank audio segment corresponding to the third deviation after the target audio segment that matches the merged display time; wherein, the audio queue is used to store the target audio segment and the blank audio segment in sequence.
[0195] In a possible implementation, the processing of the i-th initial audio segment into a target audio segment that matches the subtitle display time of the i-th translated subtitle text according to the segment duration of the i-th initial audio segment, the subtitle display time corresponding to the i-th translated subtitle text, and the subtitle character identifier includes: when the acceleration ratio corresponding to the i-th initial audio segment is greater than the acceleration ratio threshold, there are no sufficient blank segments before and after the i-th initial audio segment, and the i-th translated subtitle text and the i+1-th translated subtitle text have different subtitle character identifiers, processing the i-th initial audio segment according to the acceleration ratio corresponding to the i-th initial audio segment. The target audio segment of the first translated subtitle text is accelerated to obtain a target audio segment that matches the subtitle display time of the i-th translated subtitle text; the target audio segment corresponding to the i-th translated subtitle text is pushed into the audio queue, and based on the target audio segment corresponding to the i-th translated subtitle text being aligned to the fourth deviation between the segment cutoff moment on the subtitle timeline and the start moment of display of the i+1-th translated subtitle text on the subtitle timeline, a blank audio segment corresponding to the fourth deviation is filled after the target audio segment corresponding to the i-th translated subtitle text in the audio queue; wherein, the audio queue is used to store the target audio segment and the blank audio segment in sequence.
[0196] In a possible implementation, the processing of the i-th initial audio segment into a target audio segment that matches the subtitle display time of the i-th translated subtitle text according to the segment duration of the i-th initial audio segment, the subtitle display time corresponding to the i-th translated subtitle text, and the subtitle character identifier includes: shortening the text content of the i-th translated subtitle text and obtaining a shortened audio segment corresponding to the shortened i-th translated subtitle text when the acceleration ratio corresponding to the i-th initial audio segment is greater than the acceleration ratio threshold, there are no sufficient blank segments before and after the i-th initial audio segment, and the i-th translated subtitle text and the i+1-th translated subtitle text correspond to different subtitle character identifiers; and obtaining a shortened audio segment corresponding to the shortened i-th translated subtitle text according to the segment duration of the shortened audio segment, the shortened The method comprises the following steps: according to the present invention, wherein the target audio segment corresponding to the shortened i-th translated subtitle text is processed into a target audio segment that matches the subtitle display time of the shortened i-th translated subtitle text according to the subtitle display time of the i-th translated subtitle text and the subtitle character identification of the i-th translated subtitle text; the target audio segment corresponding to the shortened i-th translated subtitle text is pushed into an audio queue; and based on the target audio segment corresponding to the shortened i-th translated subtitle text, the target audio segment corresponding to the shortened i-th translated subtitle text is aligned to the fifth deviation between the segment end time on the subtitle timeline and the start time of the display of the i+1-th translated subtitle text on the subtitle timeline; and a blank audio segment corresponding to the fifth deviation is filled after the target audio segment corresponding to the shortened i-th translated subtitle text in the audio queue; wherein, the audio queue is used to store target audio segments and blank audio segments in sequence.
[0197] In a possible implementation, the processing of the i-th initial audio segment into a target audio segment that matches the subtitle display time of the i-th translated subtitle text according to the segment duration of the i-th initial audio segment, the subtitle display time corresponding to the i-th translated subtitle text, and the subtitle character identifier includes: when the segment duration of the i-th initial audio segment is greater than the subtitle display time of the i-th translated subtitle text, based on the subtitle timeline, determining whether the i+1-th translated subtitle text and the i-th translated subtitle text correspond to the same subtitle character identifier; when the i-th translated subtitle text and the i+1-th translated subtitle text correspond to the same subtitle character identifier, In the case of a subtitle character identification, the i-th initial audio segment corresponding to the i-th translated subtitle text and the i+1-th initial audio segment corresponding to the i+1-th translated subtitle text are merged to obtain a first merged audio segment; according to the segment length of the first merged audio segment, the first merged display time corresponding to the i-th translated subtitle text and the i+1-th translated subtitle text, and the subtitle character identification, the first merged audio segment is processed into a target audio segment that matches the first merged display time; wherein the first merged display time includes the start time of the display of the i-th translated subtitle text and the end time of the display of the i+1-th translated subtitle text.
[0198] In one possible implementation, the first merged audio segment is processed into a target audio segment that matches the first merged display time based on the segment duration of the first merged audio segment, the first merged display time corresponding to the i-th translated subtitle text and the i+1-th translated subtitle text, and the subtitle character identification, including: when the segment duration of the first merged audio segment is less than or equal to the merged display time corresponding to the first merged display time, the first merged audio segment is determined as the target audio segment that matches the first merged display time.
[0199] In one possible implementation, processing the first merged audio segment into a target audio segment matching the first merged display time based on the segment duration of the first merged audio segment, the first merged display time corresponding to the i-th translated subtitle text and the i+1-th translated subtitle text, and the subtitle character identifier includes: when the segment duration of the first merged audio segment is greater than the combined display time of the first combined display time, and the number of initial audio segments merged in the first merged audio segment is less than a preset combined number threshold, determining, based on the subtitle timeline, whether the i+2-th translated subtitle text and the i+1-th translated subtitle text correspond to the same subtitle character identifier; when the i+2-th translated subtitle text and the i+1-th translated subtitle text correspond to different subtitle character identifiers, determining the ratio between the segment duration of the first merged audio segment and the combined display time of the first combined display time as the acceleration ratio corresponding to the first merged audio segment; and when the acceleration ratio corresponding to the first merged audio segment is less than or equal to a preset acceleration ratio threshold, accelerating the first merged audio segment according to the acceleration ratio corresponding to the first merged audio segment to obtain the target audio segment matching the first combined display time.
[0200] In a possible implementation, the processing of the first merged audio segment into a target audio segment that matches the first merged display time based on the segment duration of the first merged audio segment, the first merged display time corresponding to the i-th translated subtitle text and the i+1-th translated subtitle text, and the subtitle character identifier includes: when the segment duration of the first merged audio segment is greater than the merged display time of the first merged display time and the number of initial audio segments merged in the first merged audio segment is equal to the merge quantity threshold, determining the ratio between the segment duration of the first merged audio segment and the merged display time of the first merged display time as the acceleration ratio corresponding to the first merged audio segment; when the acceleration ratio corresponding to the first merged audio segment is less than or equal to a preset acceleration ratio threshold, accelerating the first merged audio segment according to the acceleration ratio corresponding to the first merged audio segment to obtain the target audio segment that matches the first merged display time.
[0201] In a possible implementation, the first merged audio segment is processed into a target audio segment that matches the first merged display time based on the segment duration of the first merged audio segment, the first merged display time corresponding to the i-th translated subtitle text and the i+1-th translated subtitle text, and the subtitle character identification, including: when the acceleration ratio corresponding to the first merged audio segment is greater than the acceleration ratio threshold, the first merged audio segment is accelerated according to the acceleration ratio threshold to obtain the target audio segment corresponding to the first merged display time; based on the duration offset of the segment duration of the target audio segment corresponding to the first merged display time relative to the merged display duration of the first merged display time, the blank audio segment in the audio queue that is located before the first merged audio segment is compressed, and the target audio segment corresponding to the first merged display time is pushed into the audio queue, wherein the segment cutoff time corresponding to the target audio segment pushed into the audio queue on the subtitle timeline is aligned with the cutoff time in the first merged display time.
[0202] In one possible implementation, after obtaining the target audio segment that matches the first merged display time, the device also includes: a third push filling module, which is used to: push the target audio segment that matches the first merged display time into the audio queue, and based on the target audio segment that matches the first merged display time being aligned to the sixth deviation between the segment cutoff moment on the subtitle timeline and the start moment of display of the i+2th translated subtitle text on the subtitle timeline, fill the audio queue with a blank audio segment corresponding to the sixth deviation after the target audio segment that matches the first merged display time; wherein, the audio queue is used to store the target audio segment and the blank audio segment in sequence.
[0203] In one possible implementation, the method further includes: a merging module for merging the first merged audio segment with the i+2th initial audio segment corresponding to the i+2th translated subtitle text to obtain a second merged audio segment when the segment duration of the first merged audio segment is greater than the merged display duration of the first merged display time, the number of initial audio segments merged in the first merged audio segment is less than a preset merge quantity threshold, and the i+2th translated subtitle text and the i+1th translated subtitle text correspond to the same subtitle character identifier; the processing module is further used to process the second merged audio segment into a target audio segment that matches the second merged display time based on the segment duration of the second merged audio segment, the second merged display time corresponding to the i-th translated subtitle text to the i+2th translated subtitle text, and the subtitle character identifier.
[0204] In a possible implementation, the processing of the i-th initial audio segment into a target audio segment that matches the subtitle display time of the i-th translated subtitle text based on the segment duration of the i-th initial audio segment, the subtitle display time corresponding to the i-th translated subtitle text, and the subtitle character identifier includes: when the segment duration of the i-th initial audio segment is greater than the subtitle display time of the i-th translated subtitle text and the i+1-th translated subtitle text has different subtitle character identifiers, determining the ratio between the segment duration of the i-th initial audio segment and the subtitle display time corresponding to the i-th translated subtitle text as the acceleration ratio corresponding to the i-th initial audio segment; when the acceleration ratio corresponding to the i-th initial audio segment is less than When the acceleration ratio is greater than or equal to a preset acceleration ratio threshold, the i-th initial audio segment is accelerated according to the acceleration ratio corresponding to the i-th initial audio segment to obtain a target audio segment that matches the subtitle display time of the i-th translated subtitle text; the target audio segment corresponding to the i-th translated subtitle text is pushed into the audio queue, and based on the target audio segment corresponding to the i-th translated subtitle text being aligned to the seventh deviation between the segment end time on the subtitle timeline and the start time of the display of the i+1-th translated subtitle text on the subtitle timeline, a blank audio segment corresponding to the seventh deviation is filled after the target audio segment corresponding to the i-th translated subtitle text in the audio queue; wherein, the audio queue is used to store target audio segments and blank audio segments in sequence.
[0205] In one possible implementation, the processing of the i-th initial audio segment into a target audio segment that matches the subtitle display time of the i-th translated subtitle text based on the segment duration of the i-th initial audio segment, the subtitle display time corresponding to the i-th translated subtitle text, and the subtitle character identifier includes: when the acceleration ratio corresponding to the i-th initial audio segment is greater than the acceleration ratio threshold, accelerating the i-th initial audio segment according to the acceleration ratio threshold to obtain the target audio segment corresponding to the i-th translated subtitle text; compressing a blank audio segment in an audio queue that precedes the i-th initial audio segment based on a duration offset of the segment duration of the target audio segment corresponding to the i-th translated subtitle text relative to the subtitle display time of the i-th translated subtitle text, and pushing the target audio segment corresponding to the i-th translated subtitle text into the audio queue, wherein the target audio segment pushed into the audio queue is aligned with a segment cutoff on the subtitle timeline and a cutoff in the subtitle display time of the i-th translated subtitle text.
[0206] In a possible implementation, a target audio file synchronized with the video file is generated based on target audio segments that match the subtitle display time of each subtitle text in the subtitle file, including: exporting the target audio segments and blank audio segments arranged in sequence in the audio queue as the target audio file.
[0207] According to the device of the embodiment of the present disclosure, by adopting the subtitle display time and subtitle character identification indicated by the subtitle timeline adopted by the original subtitle text of the video file, the initial audio segment corresponding to the translated subtitle text translated from the original subtitle text can be automatically processed into a target audio segment that matches the subtitle display time, thereby efficiently and quickly generating a target audio file synchronized with the video file, and achieving the audio-visual synchronization effect without manual intervention, greatly reducing labor costs.
[0208] Based on the method of the above embodiment of the present disclosure, the embodiment of the present disclosure also provides another audio and video synchronization device, including:
[0209] A file acquisition module is configured to acquire a subtitle file corresponding to a video file, wherein the subtitle file includes a plurality of translated subtitle texts and a subtitle timeline; the translated subtitle text includes a translation of the original subtitle text of the video file into a target language; the subtitle timeline is the timeline used by the original subtitle text of the video file, and the subtitle timeline represents the subtitle display time and the subtitle character identifier of the subtitle text, wherein the subtitle character identifier represents the character speaking the subtitle text, and the subtitle display time includes the start time and the end time of the subtitle text display;
[0210] The first execution module is configured to execute the above steps S1300 to S1315 based on the subtitle file when the subtitle density represented by the subtitle timeline is less than or equal to the specified density threshold. Figure 10 The audio-video hard synchronization strategy and step S14 shown are used to obtain a target audio file synchronized with the video file;
[0211] The second execution module is configured to execute the above steps S1320 to S1331 or S1341 based on the subtitle file when the subtitle density represented by the subtitle timeline is greater than the specified density threshold. Figure 17 The audio-video soft synchronization strategy and step S14 are shown to obtain a target audio file synchronized with the video file.
[0212] According to the device of the embodiment of the present disclosure, it is possible to automatically select the corresponding audio and video synchronization strategy for different subtitle densities, and can adaptively set the thresholds required for acceleration, shifting, merging, and other processing based on the selected audio and video synchronization strategy; and whether it is the audio and video synchronization strategy selected based on task adaptation, or the acceleration, shifting, merging, and other parameter thresholds adaptively set based on the audio and video synchronization strategy, appropriate fine-tuning can be performed to meet audio and video requirements. During the entire audio and video synchronization process, no human intervention is required, and audio and video synchronization can be achieved with one click, greatly reducing labor costs.
[0213] In some embodiments, the functions or modules included in the device provided by the embodiments of the present disclosure can be used to execute the method described in the above method embodiments. The specific implementation can refer to the description of the above method embodiments. For the sake of brevity, it will not be repeated here.
[0214] The present disclosure also provides a computer-readable storage medium having computer program instructions stored thereon, wherein the computer program instructions implement the above method when executed by a processor. The computer-readable storage medium may be a volatile or non-volatile computer-readable storage medium.
[0215] An embodiment of the present disclosure further proposes an electronic device, comprising: a processor; and a memory for storing instructions executable by the processor; wherein the processor is configured to implement the above method when executing the instructions stored in the memory.
[0216] An embodiment of the present disclosure also provides a computer program product, including computer-readable code, or a non-volatile computer-readable storage medium carrying computer-readable code. When the computer-readable code runs in a processor of an electronic device, the processor in the electronic device executes the above method.
[0217] Figure 19 FIG1 shows a block diagram of an electronic device 1900 according to an embodiment of the present disclosure. For example, the electronic device 1900 can be provided as a server or a terminal device. Figure 19 The electronic device 1900 includes a processing component 1922, which further includes one or more processors, and a memory resource represented by a memory 1932 for storing instructions executable by the processing component 1922, such as an application. The application stored in the memory 1932 may include one or more modules, each corresponding to a set of instructions. In addition, the processing component 1922 is configured to execute the instructions to perform the above-described method.
[0218] The electronic device 1900 may further include a power supply component 1926 configured to perform power management of the electronic device 1900, a wired or wireless network interface 1950 configured to connect the electronic device 1900 to a network, and an input / output interface 1958 (I / O interface). The electronic device 1900 may operate based on an operating system stored in the memory 1932, such as Windows Server 2003. TM , Mac OS X TM , Unix TM ,Linux TM , FreeBSD TM or similar.
[0219] In an exemplary embodiment, a non-volatile computer-readable storage medium is also provided, such as a memory 1932 including computer program instructions that can be executed by the processing component 1922 of the electronic device 1900 to perform the above method.
[0220] The present disclosure may be a system, method and / or computer program product. The computer program product may include a computer-readable storage medium carrying computer-readable program instructions for causing a processor to implement various aspects of the present disclosure.
[0221] A computer-readable storage medium can be a tangible device that can hold and store instructions for use by an instruction execution device. A computer-readable storage medium can be, for example, but not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination thereof. More specific examples (a non-exhaustive list) of computer-readable storage media include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanical encoding device, such as a punch card or a raised structure in a groove on which instructions are stored, and any suitable combination thereof. As used herein, a computer-readable storage medium is not to be construed as a transient signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., a light pulse through a fiber optic cable), or an electrical signal transmitted through an electrical wire.
[0222] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to each computing / processing device, or downloaded to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network can include copper transmission cables, fiber optic transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. The network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions to be stored in the computer-readable storage medium in each computing / processing device.
[0223] The computer program instructions for performing the operations of the present disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages such as Smalltalk, C++, and conventional procedural programming languages such as "C" language or similar programming languages. Computer-readable program instructions may be executed entirely on a user's computer, partially on a user's computer, as an independent software package, partially on a user's computer, partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., utilizing an Internet service provider to connect via the Internet). In some embodiments, an electronic circuit, such as a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA), may be personalized by utilizing the state information of the computer-readable program instructions. The electronic circuit may execute the computer-readable program instructions, thereby realizing various aspects of the present disclosure.
[0224] Various aspects of the present disclosure are described herein with reference to flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present disclosure. It should be understood that each block of the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer-readable program instructions.
[0225] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, thereby producing a machine, so that when these instructions are executed by the processor of the computer or other programmable data processing device, a device is generated that implements the functions / actions specified in one or more blocks in the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium, where these instructions cause the computer, programmable data processing device, and / or other device to operate in a specific manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing various aspects of the functions / actions specified in one or more blocks in the flowchart and / or block diagram.
[0226] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device so that a series of operational steps are performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to implement the functions / actions specified in one or more blocks in the flowchart and / or block diagram.
[0227] The flow charts and block diagrams in the accompanying drawings show the possible architecture, functions and operations of the systems, methods and computer program products according to multiple embodiments of the present disclosure. In this regard, each box in the flow chart or block diagram can represent a part of a module, program segment or instruction, and the part of the module, program segment or instruction contains one or more executable instructions for realizing the prescribed logical function. In some alternative implementations, the functions marked in the box can also occur in a sequence different from that marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be implemented by a dedicated hardware-based system that performs the prescribed function or action, or can be implemented by a combination of dedicated hardware and computer instructions.
[0228] While various embodiments of the present disclosure have been described above, the foregoing description is intended to be illustrative, non-exhaustive, and not limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is selected to best explain the principles of the embodiments, their practical applications, or technological improvements in the marketplace, or to enable others skilled in the art to understand the embodiments disclosed herein.
Claims
1. A method for synchronizing audio and video, characterized in that: include: Obtaining a subtitle file corresponding to a video file, wherein the subtitle file includes a plurality of translated subtitle texts and a subtitle timeline; The translated subtitle text includes a translation of the original subtitle text of the video file into a target language; the subtitle timeline is the timeline used by the original subtitle text of the video file, the subtitle timeline represents the subtitle display time and the subtitle character identifier of the subtitle text, the subtitle character identifier represents the character speaking the subtitle text, and the subtitle display time includes the start time and the end time of the subtitle text display; For an i-th translated subtitle text in the subtitle file, obtaining an i-th initial audio segment corresponding to the i-th translated subtitle text, and determining a subtitle display time and a subtitle character identifier corresponding to the i-th translated subtitle text based on the subtitle timeline; wherein the i-th initial audio segment is an audio segment generated based on the i-th translated subtitle text; According to the segment duration of the i-th initial audio segment, the subtitle display time corresponding to the i-th translated subtitle text, and the subtitle character identifier, the i-th initial audio segment is processed into a target audio segment that matches the subtitle display time corresponding to the i-th translated subtitle text; wherein, if the segment duration of the i-th initial audio segment is greater than the subtitle display time of the i-th translated subtitle text, the acceleration ratio corresponding to the i-th initial audio segment is greater than the acceleration ratio threshold, and there are no sufficient blank segments before and after the i-th initial audio segment, if the i-th translated subtitle text is consistent with the i+1-th translated subtitle text, If the subtitle texts correspond to different subtitle character identifiers, the i-th initial audio segment is accelerated. If the i-th translated subtitle text and the i+1-th translated subtitle text correspond to the same subtitle character identifier, the i-th initial audio segment corresponding to the i-th translated subtitle text and the i+1-th initial audio segment corresponding to the i+1-th translated subtitle text are merged to obtain a target audio segment that matches the subtitle display time of the i-th translated subtitle text. The acceleration ratio is the ratio of the segment length of the i-th initial audio segment to the subtitle display time of the i-th translated subtitle text. A target audio file synchronized with the video file is generated according to target audio segments in the subtitle file that match the subtitle display time of each translated subtitle text.
2. The method according to claim 1, characterized in that The processing of the i-th initial audio segment into a target audio segment matching the subtitle display time of the i-th translated subtitle text according to the segment duration of the i-th initial audio segment, the subtitle display time corresponding to the i-th translated subtitle text, and the subtitle character identifier includes: When the segment duration of the i-th initial audio segment is less than or equal to the subtitle display duration of the i-th translated subtitle text, the i-th initial audio segment is determined as a target audio segment that matches the subtitle display time of the i-th translated subtitle text; wherein the subtitle display duration of the i-th translated subtitle text is determined based on the subtitle display time of the i-th translated subtitle text.
3. The method according to claim 1, characterized in that The processing of the i-th initial audio segment into a target audio segment matching the subtitle display time of the i-th translated subtitle text according to the segment duration of the i-th initial audio segment, the subtitle display time corresponding to the i-th translated subtitle text, and the subtitle character identifier includes: When the duration of the i-th initial audio segment is greater than the subtitle display duration of the i-th translated subtitle text, determining a ratio between the duration of the i-th initial audio segment and the subtitle display duration of the i-th translated subtitle text as the speedup ratio corresponding to the i-th initial audio segment; When the acceleration ratio corresponding to the i-th initial audio segment is less than or equal to a preset acceleration ratio threshold, the i-th initial audio segment is accelerated according to the acceleration ratio corresponding to the i-th initial audio segment to obtain a target audio segment that matches the subtitle display time of the i-th translated subtitle text.
4. The method according to claim 2, characterized in that After obtaining the target audio segment that matches the subtitle display time of the i-th translated subtitle text, the method further includes: Pushing the target audio segment into an audio queue, and based on a first offset between a segment cutoff time on the subtitle timeline and a start time for displaying the (i+1)th translated subtitle text on the subtitle timeline, filling the audio queue with a blank audio segment corresponding to the first offset after the target audio segment; The audio queue is used to store target audio segments and blank audio segments in sequence.
5. The method according to claim 3, characterized in that The processing of the i-th initial audio segment into a target audio segment matching the subtitle display time of the i-th translated subtitle text according to the segment duration of the i-th initial audio segment, the subtitle display time corresponding to the i-th translated subtitle text, and the subtitle character identifier includes: If the acceleration ratio corresponding to the i-th initial audio segment is greater than the acceleration ratio threshold, determining whether there are sufficient blank segments before and after the i-th initial audio segment based on the subtitle display time of the i-th translated subtitle text; wherein the presence of sufficient blank segments includes the presence of blank segments before and after the i-th initial audio segment, and the duration of the blank segments is greater than or equal to the acceleration offset corresponding to the i-th initial audio segment; The acceleration offset comprises the difference between the duration of the audio segment obtained by accelerating the i-th initial audio segment using the acceleration ratio threshold and the subtitle display duration corresponding to the i-th translated subtitle text; the duration of the blank segment before the i-th initial audio segment comprises the duration of the blank audio segment filled before the i-th initial audio segment; and the duration of the blank segment after the i-th initial audio segment comprises the duration on the subtitle timeline from the end time of display of the i-th translated subtitle text to the start time of display of the i+1-th translated subtitle text. When there are sufficient blank segments before and after the i-th initial audio segment, speeding up the i-th initial audio segment according to the speedup ratio threshold to obtain a target audio segment that matches the subtitle display time of the i-th translated subtitle text; The target audio segment is pushed into the audio queue by occupying a blank segment before the i-th initial audio segment and / or occupying a blank segment after the i-th initial audio segment.
6. The method according to claim 5, characterized in that Pushing the target audio segment into the audio queue by occupying a blank segment before the i-th initial audio segment and / or occupying a blank segment after the i-th initial audio segment includes: In a case where there are sufficient blank segments before and after the i-th initial audio segment, determining, based on a preset offset threshold, the acceleration offset, and the duration of the blank segment before the i-th initial audio segment, a required amount of compression for the blank segment before the i-th initial audio segment and an amount of the blank segment after the i-th initial audio segment occupied by the target audio segment; wherein the offset threshold is used to limit the maximum amount of compression that can be applied to the blank segment; If the occupied amount is less than or equal to the duration of the blank segment after the i-th initial audio segment, compressing the blank audio segment after the target audio segment corresponding to the i-1th translated subtitle text in the audio queue based on the compression amount; Accelerating the i-th initial audio segment according to the acceleration ratio threshold to obtain a target audio segment that matches the subtitle display time of the i-th translated subtitle text, and pushing the target audio segment that matches the subtitle display time of the i-th translated subtitle text into the audio queue and arranging it after the compressed blank audio segment; Based on the alignment of the target audio segment corresponding to the i-th translated subtitle text in the audio queue to a second offset between the segment cutoff time on the subtitle timeline and the start time of display of the (i+1)-th translated subtitle text on the subtitle timeline, filling the target audio segment corresponding to the i-th translated subtitle text in the audio queue with a blank audio segment corresponding to the second offset; The audio queue is used to store target audio segments and blank audio segments in sequence.
7. The method according to claim 5, characterized in that The processing of the i-th initial audio segment into a target audio segment matching the subtitle display time of the i-th translated subtitle text according to the segment duration of the i-th initial audio segment, the subtitle display time corresponding to the i-th translated subtitle text, and the subtitle character identifier includes: If the acceleration ratio corresponding to the i-th initial audio segment is greater than the acceleration ratio threshold and there are no sufficient blank segments before and after the i-th initial audio segment, determine, based on the subtitle timeline, whether the i-th translated subtitle text and the (i+1)-th translated subtitle text correspond to the same subtitle character identifier; wherein the absence of sufficient blank segments includes that there are no blank segments before and after the i-th initial audio segment, or the duration of the existing blank segments is less than the acceleration offset corresponding to the i-th initial audio segment, or the amount of blank segments occupied by the target audio segment after the i-th initial audio segment is greater than the duration of the blank segments after the i-th initial audio segment; When the i-th translated subtitle text and the (i+1)-th translated subtitle text correspond to the same subtitle character identifier, merging the i-th initial audio segment corresponding to the i-th translated subtitle text and the (i+1)-th initial audio segment corresponding to the (i+1)-th translated subtitle text to obtain a merged audio segment; Based on the segment length of the merged audio segment, the merged display time corresponding to the i-th translated subtitle text and the i+1-th translated subtitle text, and the subtitle character identification, the merged audio segment is processed into a target audio segment that matches the merged display time; wherein the merged display time includes the start time of the display of the i-th translated subtitle text to the end time of the display of the i+1-th translated subtitle text.
8. The method according to claim 7, characterized in that After obtaining the target audio segment that matches the combined display time, the method further includes: Pushing the target audio segment that matches the combined presentation time into the audio queue, and based on aligning the target audio segment that matches the combined presentation time to a third offset between the segment cutoff time on the subtitle timeline and the start time of display of the (i+2)th translated subtitle text on the subtitle timeline, filling the audio queue with a blank audio segment corresponding to the third offset after the target audio segment that matches the combined presentation time; The audio queue is used to store target audio segments and blank audio segments in sequence.
9. The method according to claim 7, characterized in that The processing of the i-th initial audio segment into a target audio segment matching the subtitle display time of the i-th translated subtitle text according to the segment duration of the i-th initial audio segment, the subtitle display time corresponding to the i-th translated subtitle text, and the subtitle character identifier includes: If the acceleration ratio corresponding to the i-th initial audio segment is greater than the acceleration ratio threshold, there are no sufficient blank segments before and after the i-th initial audio segment, and the i-th translated subtitle text and the (i+1)-th translated subtitle text correspond to different subtitle character identifiers, the i-th initial audio segment is accelerated according to the acceleration ratio corresponding to the i-th initial audio segment to obtain a target audio segment that matches the subtitle display time of the i-th translated subtitle text; Pushing the target audio segment corresponding to the i-th translated subtitle text into an audio queue, and based on aligning the target audio segment corresponding to the i-th translated subtitle text to a fourth offset between a segment cutoff time on the subtitle timeline and a start time for display of the (i+1)th translated subtitle text on the subtitle timeline, filling the target audio segment corresponding to the i-th translated subtitle text in the audio queue with a blank audio segment corresponding to the fourth offset; The audio queue is used to store target audio segments and blank audio segments in sequence.
10. The method according to claim 7, characterized in that The processing of the i-th initial audio segment into a target audio segment matching the subtitle display time of the i-th translated subtitle text according to the segment duration of the i-th initial audio segment, the subtitle display time corresponding to the i-th translated subtitle text, and the subtitle character identifier includes: shortening the text content of the i-th translated subtitle text and obtaining a shortened audio segment corresponding to the shortened i-th translated subtitle text, if the acceleration ratio corresponding to the i-th initial audio segment is greater than the acceleration ratio threshold, there are no sufficient blank segments before and after the i-th initial audio segment, and the i-th translated subtitle text and the (i+1)-th translated subtitle text correspond to different subtitle character identifiers; processing the shortened audio segment into a target audio segment that matches the shortened subtitle display time of the i-th translated subtitle text based on the segment duration of the shortened audio segment, the shortened subtitle display time of the i-th translated subtitle text, and the subtitle character identifier; Pushing the target audio segment corresponding to the shortened i-th translated subtitle text into an audio queue, and based on aligning the target audio segment corresponding to the shortened i-th translated subtitle text to a fifth offset between a segment cutoff time on the subtitle timeline and a start time for display of the (i+1)th translated subtitle text on the subtitle timeline, filling the audio segment corresponding to the shortened i-th translated subtitle text in the audio queue with a blank audio segment corresponding to the fifth offset; The audio queue is used to store target audio segments and blank audio segments in sequence.
11. The method according to any one of claims 4, 6, 8-10, characterized in that: The step of generating a target audio file synchronized with the video file according to target audio segments in the subtitle file that match the subtitle display time of each subtitle text comprises: Export the target audio segments and blank audio segments arranged in sequence in the audio queue as a target audio file.
12. A method for synchronizing audio and video, characterized in that: include: Obtaining a subtitle file corresponding to a video file, wherein the subtitle file includes a plurality of translated subtitle texts and a subtitle timeline; The translated subtitle text includes a translation of the original subtitle text of the video file into a target language; the subtitle timeline is the timeline used by the original subtitle text of the video file, the subtitle timeline represents the subtitle display time and the subtitle character identifier of the subtitle text, the subtitle character identifier represents the character speaking the subtitle text, and the subtitle display time includes the start time and the end time of the subtitle text display; For an i-th translated subtitle text in the subtitle file, obtaining an i-th initial audio segment corresponding to the i-th translated subtitle text, and determining a subtitle display time and a subtitle character identifier corresponding to the i-th translated subtitle text based on the subtitle timeline; wherein the i-th initial audio segment is an audio segment generated based on the i-th translated subtitle text; According to the segment duration of the i-th initial audio segment, the subtitle display time corresponding to the i-th translated subtitle text, and the subtitle character identifier, the i-th initial audio segment is processed into a target audio segment that matches the subtitle display time corresponding to the i-th translated subtitle text; wherein, when the segment duration of the i-th initial audio segment is greater than the subtitle display time of the i-th translated subtitle text, if the i-th translated subtitle text and the i+1-th translated subtitle text correspond to the same subtitle character identifier, the i-th initial audio segment corresponding to the i-th translated subtitle text is processed. Merging the i+1th initial audio segment corresponding to the i+1th translated subtitle text; if the i-th translated subtitle text and the i+1th translated subtitle text correspond to different subtitle character identifiers, determining whether the acceleration ratio corresponding to the i-th initial audio segment is less than or equal to a preset acceleration ratio threshold, thereby accelerating the i-th initial audio segment to obtain a target audio segment that matches the subtitle display time of the i-th translated subtitle text, wherein the acceleration ratio is a ratio of the segment duration of the i-th initial audio segment to the subtitle display time of the i-th translated subtitle text; A target audio file synchronized with the video file is generated according to target audio segments in the subtitle file that match the subtitle display time of each translated subtitle text.
13. The method according to claim 12, characterized in that The processing of the i-th initial audio segment into a target audio segment matching the subtitle display time of the i-th translated subtitle text according to the segment duration of the i-th initial audio segment, the subtitle display time corresponding to the i-th translated subtitle text, and the subtitle character identifier includes: When the segment duration of the i-th initial audio segment is less than or equal to the subtitle display duration of the i-th translated subtitle text, the i-th initial audio segment is determined as a target audio segment that matches the subtitle display time of the i-th translated subtitle text; wherein the subtitle display duration of the i-th translated subtitle text is determined based on the subtitle display time of the i-th translated subtitle text.
14. The method according to claim 12, characterized in that The processing of the i-th initial audio segment into a target audio segment matching the subtitle display time of the i-th translated subtitle text according to the segment duration of the i-th initial audio segment, the subtitle display time corresponding to the i-th translated subtitle text, and the subtitle character identifier includes: If the duration of the i-th initial audio segment is greater than the subtitle display duration of the i-th translated subtitle text, determining, based on the subtitle timeline, whether the i+1-th translated subtitle text and the i-th translated subtitle text correspond to the same subtitle character identifier; When the i-th translated subtitle text and the (i+1)-th translated subtitle text correspond to the same subtitle character identifier, merging the i-th initial audio segment corresponding to the i-th translated subtitle text and the (i+1)-th initial audio segment corresponding to the (i+1)-th translated subtitle text to obtain a first merged audio segment; Based on the segment length of the first merged audio segment, the first merged display time corresponding to the i-th translated subtitle text and the i+1-th translated subtitle text, and the subtitle character identification, the first merged audio segment is processed into a target audio segment that matches the first merged display time; wherein the first merged display time includes the start time of the display of the i-th translated subtitle text and the end time of the display of the i+1-th translated subtitle text.
15. The method according to claim 14, characterized in that The processing of the first merged audio segment into a target audio segment matching the first merged display time according to the segment duration of the first merged audio segment, the first merged display time corresponding to the i-th translated subtitle text and the (i+1)-th translated subtitle text, and the subtitle character identifier includes: When the segment length of the first merged audio segment is less than or equal to the merged display duration corresponding to the first merged display time, the first merged audio segment is determined as a target audio segment matching the first merged display time.
16. The method according to claim 14, characterized in that The processing of the first merged audio segment into a target audio segment matching the first merged display time according to the segment duration of the first merged audio segment, the first merged display time corresponding to the i-th translated subtitle text and the (i+1)-th translated subtitle text, and the subtitle character identifier includes: If the segment duration of the first merged audio segment is greater than the combined display duration of the first combined display time, and the number of initial audio segments merged into the first merged audio segment is less than a preset merge quantity threshold, determining, based on the subtitle timeline, whether the (i+2)th translated subtitle text and the (i+1)th translated subtitle text correspond to the same subtitle character identifier; If the (i+2)th translated subtitle text and the (i+1)th translated subtitle text correspond to different subtitle character identifiers, determining a ratio between a segment duration of the first merged audio segment and a combined display duration of the first combined display time as a speedup ratio corresponding to the first merged audio segment; When the acceleration ratio corresponding to the first merged audio segment is less than or equal to a preset acceleration ratio threshold, the first merged audio segment is accelerated according to the acceleration ratio corresponding to the first merged audio segment to obtain a target audio segment that matches the first merged display time.
17. The method according to claim 16, characterized in that The processing of the first merged audio segment into a target audio segment matching the first merged display time according to the segment duration of the first merged audio segment, the first merged display time corresponding to the i-th translated subtitle text and the (i+1)-th translated subtitle text, and the subtitle character identifier includes: When the segment duration of the first merged audio segment is greater than the combined display duration of the first combined display time, and the number of initial audio segments merged into the first merged audio segment is equal to the merge quantity threshold, determining a ratio between the segment duration of the first merged audio segment and the combined display duration of the first merged display time as the speedup ratio corresponding to the first merged audio segment; When the acceleration ratio corresponding to the first merged audio segment is less than or equal to a preset acceleration ratio threshold, the first merged audio segment is accelerated according to the acceleration ratio corresponding to the first merged audio segment to obtain a target audio segment that matches the first merged display time.
18. The method according to claim 16 or 17, characterized in that The processing of the first merged audio segment into a target audio segment matching the first merged display time according to the segment duration of the first merged audio segment, the first merged display time corresponding to the i-th translated subtitle text and the (i+1)-th translated subtitle text, and the subtitle character identifier includes: When the acceleration ratio corresponding to the first merged audio segment is greater than the acceleration ratio threshold, accelerating the first merged audio segment according to the acceleration ratio threshold to obtain a target audio segment corresponding to the first merged display time; Based on the duration offset of the target audio segment corresponding to the first merged display time relative to the merged display duration of the first merged display time, the blank audio segment in the audio queue that is located before the first merged audio segment is compressed, and the target audio segment corresponding to the first merged display time is pushed into the audio queue, wherein the target audio segment pushed into the audio queue corresponds to the segment cutoff time on the subtitle timeline and is aligned with the cutoff time in the first merged display time.
19. The method according to claim 14, wherein After obtaining the target audio segment matching the first combined display time, the method further includes: Pushing the target audio segment that matches the first combined presentation time into an audio queue, and aligning the target audio segment that matches the first combined presentation time to a sixth offset between a segment cutoff time on the subtitle timeline and a start time for presentation of the (i+2)th translated subtitle text on the subtitle timeline, filling the audio queue with a blank audio segment corresponding to the sixth offset after the target audio segment that matches the first combined presentation time; The audio queue is used to store target audio segments and blank audio segments in sequence.
20. The method according to claim 16, wherein The method further comprises: If the segment duration of the first merged audio segment is greater than the combined display duration of the first combined display time, the number of initial audio segments merged in the first merged audio segment is less than a preset merge quantity threshold, and the (i+2)th translated subtitle text and the (i+1)th translated subtitle text correspond to the same subtitle character identifier, merging the first merged audio segment with the (i+2)th initial audio segment corresponding to the (i+2)th translated subtitle text to obtain a second merged audio segment; Based on the segment length of the second merged audio segment, the second merged display time corresponding to the i-th translated subtitle text to the i+2-th translated subtitle text, and the subtitle character identification, the second merged audio segment is processed into a target audio segment that matches the second merged display time.
21. The method according to claim 12, wherein The processing of the i-th initial audio segment into a target audio segment matching the subtitle display time of the i-th translated subtitle text according to the segment duration of the i-th initial audio segment, the subtitle display time corresponding to the i-th translated subtitle text, and the subtitle character identifier includes: If the duration of the i-th initial audio segment is greater than the subtitle display duration of the i-th translated subtitle text, and the i-th translated subtitle text and the (i+1)-th translated subtitle text correspond to different subtitle character identifiers, determine the ratio between the duration of the i-th initial audio segment and the subtitle display duration corresponding to the i-th translated subtitle text as the speedup ratio corresponding to the i-th initial audio segment; If the acceleration ratio corresponding to the i-th initial audio segment is less than or equal to a preset acceleration ratio threshold, accelerating the i-th initial audio segment according to the acceleration ratio corresponding to the i-th initial audio segment to obtain a target audio segment that matches the subtitle display time of the i-th translated subtitle text; Pushing the target audio segment corresponding to the i-th translated subtitle text into the audio queue, and based on aligning the target audio segment corresponding to the i-th translated subtitle text to the seventh offset between the segment end time on the subtitle timeline and the start time of display of the (i+1)th translated subtitle text on the subtitle timeline, filling the target audio segment corresponding to the i-th translated subtitle text in the audio queue with a blank audio segment corresponding to the seventh offset; The audio queue is used to store target audio segments and blank audio segments in sequence.
22. The method according to claim 21, characterized in that The processing of the i-th initial audio segment into a target audio segment matching the subtitle display time of the i-th translated subtitle text according to the segment duration of the i-th initial audio segment, the subtitle display time corresponding to the i-th translated subtitle text, and the subtitle character identifier includes: When the acceleration ratio corresponding to the i-th initial audio segment is greater than the acceleration ratio threshold, accelerating the i-th initial audio segment according to the acceleration ratio threshold to obtain a target audio segment corresponding to the i-th translated subtitle text; Based on the duration offset of the target audio segment corresponding to the i-th translated subtitle text relative to the subtitle display duration of the i-th translated subtitle text, the blank audio segment in the audio queue that is located before the i-th initial audio segment is compressed, and the target audio segment corresponding to the i-th translated subtitle text is pushed into the audio queue, wherein the target audio segment pushed into the audio queue is aligned with the segment cutoff time on the subtitle timeline and the cutoff time in the subtitle display time of the i-th translated subtitle text.
23. The method according to any one of claims 19, 21-22, characterized in that: Generating a target audio file synchronized with the video file according to a target audio segment in the subtitle file that matches the subtitle display time of each subtitle text includes: Export the target audio segments and blank audio segments arranged in sequence in the audio queue as a target audio file.
24. A method for synchronizing audio and video, characterized in that: include: Obtaining a subtitle file corresponding to a video file, wherein the subtitle file includes a plurality of translated subtitle texts and a subtitle timeline; The translated subtitle text includes a translation of the original subtitle text of the video file into a target language; the subtitle timeline is the timeline used by the original subtitle text of the video file, the subtitle timeline represents the subtitle display time and the subtitle character identifier of the subtitle text, the subtitle character identifier represents the character speaking the subtitle text, and the subtitle display time includes the start time and the end time of the subtitle text display; When the subtitle density represented by the subtitle timeline is less than or equal to a specified density threshold, performing the method according to any one of claims 2 to 11 based on the subtitle file to obtain a target audio file synchronized with the video file; the subtitle density represents the density of the subtitle text; When the subtitle density represented by the subtitle timeline is greater than the specified density threshold, the method according to any one of claims 13 to 23 is executed based on the subtitle file to obtain a target audio file synchronized with the video file.
25. An audio and video synchronization device, characterized in that: include: A first acquisition module is used to acquire a subtitle file corresponding to a video file, wherein the subtitle file includes a plurality of translated subtitle texts and a subtitle timeline; The translated subtitle text includes a translation of the original subtitle text of the video file into a target language; the subtitle timeline is the timeline used by the original subtitle text of the video file, the subtitle timeline represents the subtitle display time and the subtitle character identifier of the subtitle text, the subtitle character identifier represents the character speaking the subtitle text, and the subtitle display time includes the start time and the end time of the subtitle text display; A second acquisition module is configured to acquire, for an i-th translated subtitle text in the subtitle file, an i-th initial audio segment corresponding to the i-th translated subtitle text, and determine, based on the subtitle timeline, a subtitle display time and a subtitle character identifier corresponding to the i-th translated subtitle text; wherein the i-th initial audio segment is an audio segment generated based on the i-th translated subtitle text; The processing module is configured to process the i-th initial audio segment into a target audio segment that matches the subtitle display time corresponding to the i-th translated subtitle text according to the segment duration of the i-th initial audio segment, the subtitle display time corresponding to the i-th translated subtitle text, and the subtitle character identifier; wherein, if the segment duration of the i-th initial audio segment is greater than the subtitle display time of the i-th translated subtitle text, the acceleration ratio corresponding to the i-th initial audio segment is greater than the acceleration ratio threshold, and there is no sufficient blank segment before and after the i-th initial audio segment, if the i-th translated subtitle text is consistent with the i+1-th subtitle text, the target audio segment is processed. If the i-th translated subtitle text corresponds to a different subtitle character identifier, the i-th initial audio segment is accelerated; if the i-th translated subtitle text and the i+1-th translated subtitle text correspond to the same subtitle character identifier, the i-th initial audio segment corresponding to the i-th translated subtitle text and the i+1-th initial audio segment corresponding to the i+1-th translated subtitle text are merged to obtain a target audio segment that matches the subtitle display time of the i-th translated subtitle text, and the acceleration ratio is the ratio of the segment length of the i-th initial audio segment to the subtitle display time of the i-th translated subtitle text; The generating module is configured to generate a target audio file synchronized with the video file according to target audio segments in the subtitle file that match the subtitle display time of each translated subtitle text.
26. An electronic device, characterized in that: include: processor; a memory for storing processor-executable instructions; The processor is configured to implement the method of any one of claims 1 to 11 or claims 12 to 23 when executing the instructions stored in the memory.
27. A non-volatile computer-readable storage medium having computer program instructions stored thereon, characterized in that: When the computer program instructions are executed by a processor, the method of any one of claims 1 to 11 or claims 12 to 23 is implemented.
Citation Information
Patent Citations
Video processing method and device, electronic equipment and storage medium
CN113207044A
Video dubbing method and related device, electronic equipment and storage medium
CN117177024A
Automatic speech translation dubbing of pre-recorded video
CN117201889A