Audio subtitle alignment method and device, medium and electronic equipment

By segmenting long audio files into short audio files and extracting and splicing their feature information, the problem of excessive resource consumption in long audio subtitle alignment is solved, achieving efficient and accurate subtitle alignment.

CN116527979BActive Publication Date: 2026-03-31BEIJING YOUZHUJU NETWORK TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-05-11
Publication Date
2026-03-31

AI Technical Summary

Technical Problem

Existing technologies consume excessive machine resources during the automatic timing of long audio clips, resulting in low efficiency and poor accuracy in subtitle alignment.

Method used

The long audio is divided into multiple short audio segments, and feature information is extracted from each segment. The segments are then spliced ​​together when the duration of each short audio segment is less than or equal to a second preset duration, and then matched with the subtitle text.

Benefits of technology

It improves the efficiency and accuracy of subtitle feature extraction, enhances the accuracy of subtitle alignment, and reduces machine resource consumption.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116527979B_ABST
    Figure CN116527979B_ABST
Patent Text Reader

Abstract

The present disclosure relates to the technical field of audio recognition, in particular, to a subtitle alignment method and device for audio, a medium and an electronic device. The method comprises: obtaining a target audio and a target subtitle text of the target audio; if a time length of the target audio is greater than a first preset time length, performing slicing processing on the target audio according to a slice time length to obtain a plurality of first target audios; determining first audio feature information of each first target audio; if the time length of the target audio is less than or equal to a second preset time length, splicing all the first audio feature information to obtain target audio feature information of the target audio, wherein the second preset time length is greater than the first preset time length; and generating subtitle information corresponding to the target audio according to the target subtitle text and the target audio feature information. In this way, the occupation of excessive machine resources can be avoided, the matching of the target subtitle text and the target audio feature information is realized through one alignment, and the accuracy of the alignment result is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates to the technical field of audio recognition, in particular, to a subtitle alignment method and device for audio, a medium and an electronic device. BACKGROUND

[0002] In a video subtitle application scenario, a user has a function requirement of automatic axis marking. The automatic axis marking, also called automatic subtitle alignment, is to automatically match a prepared subtitle text to an audio and generate a time axis. The function is applicable to a case of having an audio file and a subtitle text at the same time. In a matching process, the longer the audio is, the more machine resources are required. Therefore, the implementation of automatic axis marking of a long audio becomes a difficulty. SUMMARY

[0003] This section is provided to introduce the general concepts of the present disclosure in a simplified form, which will be described in detail in the following detailed description section. This section does not intend to identify key or essential features of the claimed technology nor is it intended to be used to limit the scope of the claimed technology.

[0004] In a first aspect, the present disclosure provides a subtitle alignment method for audio, comprising:

[0005] obtaining a target audio and a target subtitle text of the target audio;

[0006] if a time length of the target audio is greater than a first preset time length, performing slice processing on the target audio according to a slice time length to obtain a plurality of first target audios;

[0007] determining first audio feature information of each of the first target audios;

[0008] if the time length of the target audio is less than or equal to a second preset time length, splicing all the first audio feature information to obtain target audio feature information of the target audio, wherein the second preset time length is greater than the first preset time length;

[0009] generating subtitle information corresponding to the target audio according to the target subtitle text and the target audio feature information.

[0010] In a second aspect, the present disclosure provides a subtitle alignment device for audio, comprising:

[0011] an obtaining module configured to obtain a target audio and a target subtitle text of the target audio;

[0012] a first processing module configured to, if a time length of the target audio is greater than a first preset time length, perform slice processing on the target audio according to a slice time length to obtain a plurality of first target audios;

[0013] The first determining module is used to determine the first segment audio feature information of each of the first target audio segments;

[0014] The second processing module is used to concatenate all the first audio feature information to obtain the target audio feature information of the target audio if the duration of the target audio is less than or equal to the second preset duration, wherein the second preset duration is greater than the first preset duration.

[0015] The first generation module is used to generate subtitle information corresponding to the target audio based on the target subtitle text and the target audio feature information.

[0016] Thirdly, this disclosure provides a computer-readable medium having a computer program stored thereon, which, when executed by a processing device, implements the steps of the above-described audio subtitle alignment method.

[0017] Fourthly, this disclosure provides an electronic device, comprising:

[0018] A storage device on which computer programs are stored;

[0019] A processing device for executing the computer program in the storage device to implement the steps of the above-described audio subtitle alignment method.

[0020] The above technical solution involves slicing target audio longer than a first preset duration to obtain multiple first target audio segments, thereby determining the first audio feature information of each first target audio segment. If the duration of the target audio is less than or equal to a second preset duration, all the first audio feature information is concatenated to obtain the target audio feature information. Based on the target subtitle text and the target audio feature information, subtitle information corresponding to the target audio is generated. This approach divides long audio segments into multiple short audio segments for feature extraction, avoiding excessive machine resource consumption. After extracting the corresponding audio feature information, if the duration of the target audio is less than or equal to the second preset duration, the various audio feature information segments can be combined into a comprehensive target audio feature information segment during subtitle alignment. This single alignment achieves the matching of the target subtitle text and the target audio feature information. Therefore, the efficiency and accuracy of feature extraction from the subtitle text can be effectively improved, and the efficiency of subtitle alignment can also be improved to a certain extent. Combined with the target subtitle text, highly accurate subtitle information corresponding to the target audio can be generated, achieving timeline matching between the target audio and the target subtitle text and improving the accuracy of the alignment results.

[0021] Other features and advantages of this disclosure will be described in detail in the following detailed description section. Attached Figure Description

[0022] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and elements are not necessarily drawn to scale. In the drawings:

[0023] Figure 1 This is a flowchart of an audio subtitle alignment method provided according to one embodiment of the present disclosure.

[0024] Figure 2 This is a flowchart of an audio subtitle alignment method provided according to one embodiment of the present disclosure.

[0025] Figure 3 This is a block diagram of an audio subtitle alignment device provided according to one embodiment of the present disclosure.

[0026] Figure 4 This is a schematic diagram of the structure of an electronic device provided according to one embodiment of the present disclosure. Detailed Implementation

[0027] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.

[0028] It should be understood that the steps described in the method embodiments of this disclosure may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of this disclosure is not limited in this respect.

[0029] The term "comprising" and its variations as used herein are open-ended inclusions, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Definitions of other terms will be given in the description below.

[0030] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are used only to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependencies.

[0031] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".

[0032] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.

[0033] All actions involving the acquisition of signals, information, or data in this disclosure are carried out in accordance with the relevant data protection laws and policies of the country where the location is situated, and with the authorization granted by the owner of the relevant device.

[0034] It is understood that before using the technical solutions disclosed in the various embodiments of this disclosure, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this disclosure in an appropriate manner in accordance with relevant laws and regulations, and user authorization should be obtained.

[0035] For example, upon receiving a user's active request, a prompt message is sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the software or hardware, such as the electronic device, application, server, or storage medium performing the operations of this disclosed technical solution, based on the prompt message.

[0036] As an optional but non-limiting implementation, in response to a user's active request, sending a prompt message to the user can be done via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose "agree" or "disagree" to provide personal information to the electronic device.

[0037] It is understood that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of this disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of this disclosure.

[0038] Meanwhile, it is understood that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws, regulations and related provisions.

[0039] Figure 1 This is a flowchart illustrating an audio subtitle alignment method according to one embodiment of the present disclosure. This method can be applied to terminals such as smartphones, tablets, personal computers (PCs), laptops, and other devices, and can also be applied to servers. Figure 1 As shown, the method may include S101 to S105.

[0040] S101, obtain the target audio and the target subtitle text of the target audio.

[0041] For example, the target audio and its target subtitle text can be uploaded by the user or retrieved from a database. If the retrieved data is a video file and its target subtitle text, audio data can be extracted from the target video and converted into a standard format target audio. The target audio and target subtitle text are in a one-to-one correspondence; that is, the content expressed by the target audio and target subtitle text is identical.

[0042] S102, if the duration of the target audio is greater than the first preset duration, then the target audio is sliced ​​according to the slice duration to obtain multiple first target audios.

[0043] S103, determine the first audio feature information for each first target audio.

[0044] Typically, audio feature information can be extracted from the audio input to the feature extraction model based on a feature extraction model (such as the Transformer model). Each initial audio feature can contain multi-dimensional audio features, such as pre-defined dimensions like pitch, loudness, and timbre. However, due to the attention mechanism in the Transformer model, its ability to extract audio feature information from long audio files may be poor, potentially affecting the final alignment result. Conversely, if the audio input to the feature extraction model is short, the extraction of audio feature information based on this model is more accurate and does not consume excessive machine resources during the extraction process. Therefore, the target audio can be classified based on a first preset duration to determine the subtitle alignment method corresponding to the target audio duration.

[0045] The first preset duration can be set in advance, for example, it can be set to 5 minutes.

[0046] If the duration of the target audio is not greater than the first preset duration, it can be determined that the target audio is relatively short. The audio feature information can be extracted directly through the feature extraction model. This can improve the accuracy of the extracted audio feature information while avoiding excessive use of machine resources. Then, based on the target subtitle text and the extracted audio feature information, the subtitle information corresponding to the target audio can be generated, that is, the target audio and the target subtitle text can be matched on the timeline.

[0047] It is worth noting that subtitle information can include not only the subtitle text itself, but also the time information corresponding to each character in the subtitle text.

[0048] If the duration of the target audio exceeds a first preset duration, it can be determined that the target audio is relatively long, and it can be segmented. For example, the target audio can be segmented according to a preset segment duration. If the target audio duration is 30 minutes and the segment duration is preset to 10 minutes, the target audio can be divided into three consecutive first target audio segments, corresponding to the content from 0 minutes to 10 minutes, the content from 10 minutes to 20 minutes, and the content from 20 minutes to 30 minutes, respectively. The first audio feature information of each of the above three consecutive first target audio segments can be determined by a feature extraction model. In this way, the accuracy of the determined first audio feature information of the first target audio segments can be improved, thereby improving the accuracy of the target audio feature information obtained. At the same time, in the process of determining the first segment audio feature information of each first target audio segment, excessive machine resources can be avoided, the feature extraction model can be prevented from being overloaded, and the application range of the feature extraction model can be expanded.

[0049] S104, if the duration of the target audio is less than or equal to the second preset duration, then all the first audio feature information is spliced ​​together to obtain the target audio feature information of the target audio.

[0050] The second preset duration is longer than the first preset duration. The second preset duration can be predetermined based on the upper limit of the optimal duration corresponding to the alignment model capable of forcibly aligning the target subtitle text and target audio feature information. For example, during the alignment model training process, the optimal duration range of audio that the alignment model can handle can be determined based on the confidence level of the alignment information corresponding to training data (sample audio and text corresponding to the sample audio) of different durations. For instance, the upper limit of the optimal duration can be determined based on the duration range of audio corresponding to training data where the confidence level of the alignment information is higher than the first threshold, and this upper limit can be set as the second preset duration. For example, the second preset duration can be set to 20 minutes.

[0051] If the duration of the target audio is less than or equal to the second preset duration, it can be determined that the alignment model can achieve the matching of target subtitle text and target audio feature information through a single alignment. Compared to concatenating multiple alignment information, this effectively improves the accuracy of the alignment result, resulting in more accurate subtitle information. All first audio feature information can be concatenated to obtain the target audio feature information. Specifically, when slicing the target audio, each first target audio segment can have corresponding time information, and the first audio feature information can be concatenated based on the order of the time information corresponding to each first target audio segment.

[0052] As mentioned above, the target audio is sliced ​​to obtain three consecutive first target audio segments. If the first audio feature information obtained by the feature extraction model for the first target audio segments from 0 min to 10 min is ctc_logits1, the first audio feature information obtained by the feature extraction model for the first target audio segments from 10 min to 20 min is ctc_logits2, and the first audio feature information obtained by the feature extraction model for the first target audio segments from 20 min to 30 min is ctc_logits3, then the target audio feature information obtained by concatenating these segments can be [ctc_logits1, ctc_logits2, ctc_logits3]. For example, the audio feature information extracted by the feature extraction model can be frame-level audio feature information, and the corresponding alignment model can output frame-level alignment information, with 40ms as one frame. Thus, the audio representation corresponding to the first audio feature information can be [slice duration * 60 * 1000 / 40, number of dimensions], and the audio representation corresponding to the target audio feature information can be [target audio duration * 60 * 1000 / 40, number of dimensions]. The alignment information includes characters in the audio frame and the target subtitle text that matches the audio frame.

[0053] S105, Generate subtitle information corresponding to the target audio based on the target subtitle text and target audio feature information.

[0054] For example, the output information of the alignment model described above can be used to generate subtitle information corresponding to the target audio, thus achieving timeline matching between the target audio and the target subtitle text. The alignment model can be modeled based on a forced alignment algorithm from related technologies, using sample audio provided by an open-source database and the corresponding text as training samples, trained through machine learning. Specifically, the target subtitle text and target audio feature information can be input into the pre-trained alignment model, and the model output is the alignment information between the target subtitle text and the target audio feature information. This alignment model can be stored locally for local access each time it is used, or it can be stored on a third-party platform for access each time it is used; no specific limitation is made here. The alignment information output by the alignment model can be frame-level alignment information. Based on the frame-level alignment information, time-level alignment information can be determined to generate the subtitle information corresponding to the target audio.

[0055] The above technical solution involves slicing target audio longer than a first preset duration to obtain multiple first target audio segments, thereby determining the first audio feature information of each first target audio segment. If the duration of the target audio is less than or equal to a second preset duration, all the first audio feature information is concatenated to obtain the target audio feature information. Based on the target subtitle text and the target audio feature information, subtitle information corresponding to the target audio is generated. This approach divides long audio segments into multiple short audio segments for feature extraction, avoiding excessive machine resource consumption. After extracting the corresponding audio feature information, if the duration of the target audio is less than or equal to the second preset duration, the various audio feature information segments can be combined into a comprehensive target audio feature information segment during subtitle alignment. This single alignment achieves the matching of the target subtitle text and the target audio feature information. Therefore, the efficiency and accuracy of feature extraction from the subtitle text can be effectively improved, and the efficiency of subtitle alignment can also be improved to a certain extent. Combined with the target subtitle text, highly accurate subtitle information corresponding to the target audio can be generated, achieving timeline matching between the target audio and the target subtitle text and improving the accuracy of the alignment results.

[0056] Optionally, in S103, determining the first segment audio feature information of each first target audio may include:

[0057] The first target audio is input into a pre-trained feature extraction model to obtain the audio feature information of the first target audio.

[0058] The feature extraction model can be an encoder in a model trained based on sample audio and corresponding text. It's worth noting that this feature extraction model is trained using machine learning methods, enabling the extraction of audio feature information. After training, the encoder can be used as the feature extraction model. This model can be stored locally for local access or stored on a third-party platform for access; no specific limitation is made here. The audio feature information extracted by the feature extraction model can be frame-level audio feature information.

[0059] Optionally, in S105, generating subtitle information corresponding to the target audio based on the target subtitle text and target audio feature information may include:

[0060] Based on the target subtitle text and target audio feature information, the alignment information of each frame of audio in the target audio is determined, wherein the alignment information includes the characters in the target subtitle text that match the target subtitle text of that frame.

[0061] Based on the alignment information and frame length of each audio frame, subtitle information corresponding to the target audio is generated.

[0062] For example, based on the alignment model described above, if the second frame of the target audio and the first character of the target subtitle text are determined to be a set of alignment information, the third frame of the target audio and the first character of the target subtitle text are determined to be a set of alignment information, and the fourth frame of the target audio and the second character of the target subtitle text are determined to be a set of alignment information, and the frame length of the audio is 40ms, then the duration corresponding to the first character of the target subtitle text on the timeline can be determined to be 80ms. Its start time is the start time of the second frame of the audio, and its end time is the end time of the third frame of the audio. That is, the first character corresponds to the 40ms to 120ms of the target audio. In this way, the time-level alignment information can be determined based on the frame-level alignment information, achieving the matching of the target audio and the target subtitle text on the timeline.

[0063] Because users have diverse usage scenarios, the duration of audio may exceed the second preset duration. However, the alignment model can only process a limited amount of data at a time, making it difficult to achieve subtitle alignment for ultra-long audio in a single alignment. Even if the alignment model can run under heavy load, the accuracy of its output is not ideal. Therefore, this disclosure also provides the following embodiments to solve this problem.

[0064] In an optional embodiment, the audio subtitle alignment method provided in this disclosure may further include:

[0065] Based on the first audio feature information and the target subtitle text, determine the subtitle text segment corresponding to the first target audio; based on the first audio feature information and the subtitle text segment, determine the alignment information of each frame of audio in each first target audio; splice the alignment information of each first target audio to generate the subtitle information corresponding to the target audio.

[0066] For example, for each first target audio, the subtitle text segment corresponding to the first target audio can be determined from the target subtitle text based on its first audio feature information and the target subtitle text, using the minimum edit distance algorithm in related technologies.

[0067] The first audio feature information and subtitle text fragment of the first target audio can be input into the alignment model described above to obtain the alignment information of each frame of the first target audio. Thus, after determining the alignment information of each frame of the first target audio, the alignment information of each frame of the first target audio can be spliced ​​temporally according to the time information of each first target audio to generate the subtitle information corresponding to the target audio.

[0068] In this way, the amount of data corresponding to each first audio feature can be adapted to the alignment model, so that the amount of data processed by the alignment model each time is within the allowable range, so as to output more accurate alignment information.

[0069] In another optional embodiment, the audio subtitle alignment method provided in this disclosure can be as follows: Figure 2 As shown, Figure 2 This is a flowchart of an audio subtitle alignment method according to one embodiment of the present disclosure. Figure 2 As shown, the method may include S201 to S205.

[0070] S201, if the duration of the target audio is greater than the second preset duration, then multiple consecutive first target audios are merged to obtain multiple second target audios.

[0071] The duration of each second target audio segment shall not exceed the second preset duration.

[0072] S202, for each second target audio, the first audio feature information in the second target audio is spliced ​​together to obtain the second audio feature information.

[0073] For example, the target audio duration is 60 minutes, and the slice duration is preset to 10 minutes, resulting in 6 consecutive first target audio segments, corresponding to the content of the target audio from 0 minutes to 10 minutes, the content of the target audio from 10 minutes to 20 minutes, the content of the target audio from 20 minutes to 30 minutes, the content of the target audio from 30 minutes to 40 minutes, the content of the target audio from 40 minutes to 50 minutes, and the content of the target audio from 50 minutes to 60 minutes. The first audio feature information of the 6 consecutive first target audio segments is ctc_logit1, ctc_logit2, ctc_logit3, ctc_logit4, ctc_logit5, and ctc_logit6, respectively.

[0074] By merging multiple consecutive first target audio segments, multiple second target audio segments can be obtained. Taking a second preset duration of 20 minutes as an example, resulting in three consecutive second target audio segments, these three segments correspond to the content of the target audio segments from 0 minutes to 20 minutes, 20 minutes to 40 minutes, and 40 minutes to 60 minutes, respectively. The second audio feature information of the three consecutive second target audio segments is ctc_logit1+ctc_logit2, ctc_logit3+ctc_logit4, and ctc_logit5+ctc_logit6, respectively.

[0075] Thus, by splicing together the first audio feature information from the second target audio, the amount of data corresponding to each second audio feature information can be adapted to the alignment model. This reduces the number of alignment steps while keeping the amount of data processed by the alignment model within an acceptable range, thereby outputting more accurate alignment information.

[0076] S203, Based on the second audio feature information and the target subtitle text, determine the subtitle text segment corresponding to the second target audio.

[0077] For example, based on the minimum edit distance algorithm, the subtitle text segment corresponding to the second target audio can be determined from the target subtitle text according to the second audio feature information and the target subtitle text. If the duration of the target audio is longer than a second preset duration, it indicates that the target audio is relatively long, and the corresponding content in the target subtitle text is also relatively long. The second target audio is a part of the target audio, and the subtitle text segment corresponding to the second target audio is also a part of the target subtitle text. For example, the target subtitle text may contain 3000 characters. Taking the content of the second target audio corresponding to 20 to 40 minutes of the target audio as an example, if the minimum edit distance algorithm determines characters 800 to 1700 in the target subtitle text corresponding to the second target audio, then characters 800 to 1700 in the target subtitle text can be identified as the subtitle text segment corresponding to the second target audio.

[0078] S204, based on the second audio feature information and the subtitle text fragment, determine the alignment information of each frame of audio in each second target audio.

[0079] For example, the alignment information for each frame of audio in each second target audio can be determined sequentially using the alignment model described above. The alignment information may include the characters in the target subtitle text that match the audio frame.

[0080] S205, the alignment information of each second target audio is spliced ​​together to generate subtitle information corresponding to the target audio.

[0081] For example, for the second target audio corresponding to the target audio from 0 min to 20 min, corresponding to the 1st to 799th characters in the target subtitle text, the alignment information of the second target audio can be obtained based on the alignment model; for the second target audio corresponding to the target audio from 20 min to 40 min, corresponding to the 800th to 1700th characters in the target subtitle text, the alignment information of the second target audio can be obtained based on the alignment model; for the second target audio corresponding to the target audio from 40 min to 60 min, corresponding to the 1701st to 300th characters in the target subtitle text, the alignment information of the second target audio can be obtained based on the alignment model. Based on the time sequence of the above three second target audios and their corresponding target audio, the alignment information of the three second target audios can be concatenated. In this way, the concatenated audio is the same as the target audio, and the concatenated subtitle may be the same as the target subtitle text. Based on the concatenation of alignment information in time, a complete match between the target audio and the target subtitle text on the time axis can be achieved, that is, the subtitle information corresponding to the target audio can be generated.

[0082] In the process of determining the subtitle text segment corresponding to the second target audio based on the minimum edit distance algorithm, the algorithm's insufficient precision may lead to inaccurate subtitle text segments. Consequently, when these segments are concatenated to obtain the concatenated subtitle text, differences may be found when comparing it with the target subtitle text. This could result in duplicate or missing text segments in the concatenated subtitle text. Therefore, merging multiple consecutive first target audio segments can keep the amount of data processed by the alignment model within acceptable limits, outputting more accurate alignment information while reducing the number of second target audio segments to be concatenated during the alignment process. This reduces the likelihood of the aforementioned problems and improves the accuracy of the alignment results.

[0083] To further address the aforementioned issues, after concatenating the alignment information of each second target audio segment to generate subtitle information corresponding to the target audio, the audio subtitle alignment method provided in this disclosure may further include:

[0084] The spliced ​​subtitle text in the subtitle information is compared with the target subtitle text to determine if there is any missing text;

[0085] If missing text is determined, then based on the target subtitle text, determine the first subtitle text segment and the second subtitle text segment adjacent to the missing text, wherein the time information corresponding to the first subtitle text segment is earlier than the time information corresponding to the second subtitle text segment.

[0086] The text insertion time is determined based on the last moment in the alignment information of the first subtitle text segment and the earliest moment in the alignment information of the second subtitle text segment;

[0087] Based on the text insertion time, missing text is inserted into the spliced ​​subtitle text to obtain updated subtitle information.

[0088] For example, the second target audio corresponding to the target audio from 0 min to 20 min corresponds to the first to 790th characters in the target subtitle text; the second target audio corresponding to the target audio from 20 min to 40 min corresponds to the 800th to 1700th characters in the target subtitle text. Thus, comparing the concatenated subtitle text in the determined subtitle information with the target subtitle text, characters 791 to 799 are missing; that is, characters 791 to 799 are missing text. Based on the time information of the subtitle text segments, characters 1 to 790 can be identified as the first subtitle text segment; characters 800 to 1700 are identified as the second subtitle text segment. If the last time in the alignment information of the first subtitle text segment is 19min58s, and the earliest time in the alignment information of the second subtitle text segment is 20min07s, with a difference of 9s, then characters 791 to 799 can be inserted within these 9s. For example, each character in the missing text can occupy the same duration in the text insertion time; that is, the 791st character corresponds to the 1st second, the 792nd character to the 2nd second, and so on. In this way, directly inserting the missing text within the text insertion time to obtain the updated subtitle information ensures that the final subtitle text matches the target subtitle text and avoids affecting the alignment information of other characters.

[0089] To further address the aforementioned issues, after concatenating the alignment information of each second target audio segment to generate subtitle information corresponding to the target audio, the audio subtitle alignment method provided in this disclosure may further include:

[0090] If, based on the target subtitle text, it is determined that there is duplicate text in the alignment information corresponding to the adjacent second target audio, then the confidence level of the duplicate text in the adjacent second target audio is determined respectively;

[0091] In adjacent second target audios, duplicate text in the alignment information of the second target audio with low confidence is removed to obtain updated subtitle information.

[0092] Determining the confidence level of the repeated text in adjacent second target audio segments may include:

[0093] Determine the matching confidence of each character of the repeated text in two adjacent second target audios.

[0094] The confidence level of the repeated text in the second target audio is determined based on the matching confidence level of each character in the second target audio.

[0095] For example, if two adjacent subtitle text segments are spliced ​​together, and a repeated character segment appears at the end of the first subtitle text segment and the beginning of the second subtitle text segment, and this character segment appears once in the corresponding position of the target subtitle text, then the repeated character segment can be identified as duplicate text. For the second target audio corresponding to 0min to 20min, the characters corresponding to the first character to the 810th character in the target subtitle text; for the second target audio corresponding to 20min to 40min, the characters corresponding to the 800th character to the 1700th character in the target subtitle text. Thus, comparing the spliced ​​subtitle text in the determined subtitle information with the target subtitle text, characters 800 to 810 appear simultaneously in the subtitle text segments of two adjacent second target audio segments, meaning characters 800 to 810 are duplicate text. This ensures that the subtitle text in the final subtitle information is consistent with the target subtitle text.

[0096] As an example, the duplicate text can be deleted from the subtitle text segment of the second target audio corresponding to the target audio from 0 min to 20 min, or from the subtitle text segment of the second target audio corresponding to the target audio from 20 min to 40 min.

[0097] As another example, the alignment model can output frame-level alignment information along with the confidence level of that alignment information. Based on the confidence level of the alignment information for each frame of audio, the matching confidence of each character of the repeated text in two adjacent second target audio segments can be determined. For each second target audio segment, the average matching confidence of each character of the repeated text in that segment can be determined as the confidence level of the repeated text in that segment. In adjacent second target audio segments, repeated text in the alignment information of the segment with lower confidence can be deleted to obtain updated subtitle information. For example, if the confidence level of the repeated text in the second target audio segments corresponding to target audio segments from 0 min to 20 min is low, then the text content from the 800th to the 810th character in the alignment information of those segments can be deleted. If the time corresponding to characters 800 to 810 in the second target audio corresponding to target audio from 0 min to 20 min is from 19 min 20 s to 20 min, and the time corresponding to characters 800 to 810 in the second target audio corresponding to target audio from 20 min to 40 min is from 21 min to 21 min 50 s, then after removing duplicate text, there will be no corresponding subtitles in the audio from 19 min 20 s to 20 min, and the audio from 21 min to 21 min 50 s will correspond to characters 800 to 810. This ensures that the subtitle text in the final subtitle information is consistent with the target subtitle text, and maximizes the accuracy of the subtitle information while avoiding affecting the alignment information of other characters.

[0098] Based on the same inventive concept, this disclosure also provides an audio subtitle alignment device. Figure 3 This is a block diagram of an audio subtitle alignment device provided according to one embodiment of the present disclosure. Figure 3 As shown, the audio subtitle alignment device 300 may include:

[0099] The acquisition module 301 is used to acquire the target audio and the target subtitle text of the target audio;

[0100] The first processing module 302 is used to slice the target audio according to the slice duration if the duration of the target audio is greater than the first preset duration, so as to obtain multiple first target audios.

[0101] The first determining module 303 is used to determine the first audio feature information of each of the first target audio;

[0102] The second processing module 304 is used to splice all the first audio feature information to obtain the target audio feature information of the target audio if the duration of the target audio is less than or equal to the second preset duration, wherein the second preset duration is greater than the first preset duration.

[0103] The first generation module 305 is used to generate subtitle information corresponding to the target audio based on the target subtitle text and the target audio feature information.

[0104] The above technical solution involves slicing target audio longer than a first preset duration to obtain multiple first target audio segments, thereby determining the first audio feature information of each first target audio segment. If the duration of the target audio is less than or equal to a second preset duration, all the first audio feature information is concatenated to obtain the target audio feature information. Based on the target subtitle text and the target audio feature information, subtitle information corresponding to the target audio is generated. This approach divides long audio segments into multiple short audio segments for feature extraction, avoiding excessive machine resource consumption. After extracting the corresponding audio feature information, if the duration of the target audio is less than or equal to the second preset duration, the various audio feature information segments can be combined into a comprehensive target audio feature information segment during subtitle alignment. This single alignment achieves the matching of the target subtitle text and the target audio feature information. Therefore, the efficiency and accuracy of feature extraction from the subtitle text can be effectively improved, and the efficiency of subtitle alignment can also be improved to a certain extent. Combined with the target subtitle text, highly accurate subtitle information corresponding to the target audio can be generated, achieving timeline matching between the target audio and the target subtitle text and improving the accuracy of the alignment results.

[0105] Optionally, the device 300 further includes:

[0106] The third processing module is used to merge multiple consecutive first target audios to obtain multiple second target audios if the duration of the target audio is greater than the second preset duration, wherein the duration of each second target audio does not exceed the second preset duration.

[0107] The fourth processing module is used to concatenate the first audio feature information in each second target audio for each second target audio to obtain second audio feature information.

[0108] The second determining module is used to determine the subtitle text segment corresponding to the second target audio based on the second audio feature information and the target subtitle text;

[0109] The third determining module is used to determine the alignment information of each frame of audio in each second target audio according to the second audio feature information and the subtitle text fragment, wherein the alignment information includes the frame of audio and the characters in the target subtitle text that match the frame of audio;

[0110] The second generation module is used to splice the alignment information of each second target audio to generate subtitle information corresponding to the target audio.

[0111] Optionally, the device 300 further includes:

[0112] The comparison module is used to compare the spliced ​​subtitle text in the subtitle information with the target subtitle text to determine whether there is any missing text;

[0113] The fourth determining module is used to determine, if it is determined that the missing text exists, a first subtitle text segment and a second subtitle text segment adjacent to the missing text based on the target subtitle text, wherein the time information corresponding to the first subtitle text segment is earlier than the time information corresponding to the second subtitle text segment;

[0114] The fifth determining module is used to determine the text insertion time based on the last time in the alignment information corresponding to the first subtitle text segment and the earliest time in the alignment information corresponding to the second subtitle text segment;

[0115] The first update module is used to insert the missing text into the spliced ​​subtitle text based on the text insertion time, so as to obtain the updated subtitle information.

[0116] Optionally, the device 300 further includes:

[0117] The sixth determining module is used to determine the confidence level of the repeated text in the adjacent second target audio if the target subtitle text is determined to have repeated text in the alignment information corresponding to the adjacent second target audio.

[0118] The second update module is used to delete the duplicate text in the alignment information of the second target audio with low confidence in adjacent second target audios, so as to obtain updated subtitle information.

[0119] Optionally, the sixth determining module is used to determine the confidence level of the repeated text in adjacent second target audio in the following ways:

[0120] Determine the matching confidence of each character of the repeated text in two adjacent second target audios;

[0121] The confidence level of the repeated text in the second target audio is determined based on the matching confidence level of each character in the second target audio.

[0122] Optionally, the first determining module 303 is used to determine the first segment audio feature information of each of the first target audio segments in the following manner:

[0123] The first target audio is input into a pre-trained feature extraction model to obtain the audio feature information of the first target audio. The feature extraction model is an encoder in a model trained based on sample audio and the text corresponding to the sample audio.

[0124] Optionally, the second processing module 304 includes:

[0125] The first determining submodule is used to determine the alignment information of each frame of audio in the target audio based on the target subtitle text and the target audio feature information, wherein the alignment information includes the frame of audio and the characters in the target subtitle text that match the frame of audio;

[0126] The generation submodule is used to generate subtitle information corresponding to the target audio based on the alignment information and frame length of each audio frame.

[0127] The following is for reference. Figure 4 This document illustrates a structural schematic diagram of an electronic device 600 suitable for implementing embodiments of the present disclosure. The terminal devices in these embodiments may include, but are not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 4 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.

[0128] like Figure 4 As shown, electronic device 600 may include a processing device (e.g., a central processing unit, a graphics processor, etc.) 601, which can perform various appropriate actions and processes according to a program stored in read-only memory (ROM) 602 or a program loaded from storage device 608 into random access memory (RAM) 603. RAM 603 also stores various programs and data required for the operation of electronic device 600. Processing device 601, ROM 602, and RAM 603 are interconnected via bus 604. Input / output (I / O) interface 605 is also connected to bus 604.

[0129] Typically, the following devices can be connected to I / O interface 605: input devices 606 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 607 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 608 including, for example, magnetic tapes, hard disks, etc.; and communication devices 609. Communication device 609 allows electronic device 600 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 4 An electronic device 600 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.

[0130] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 609, or installed from a storage device 608, or installed from a ROM 602. When the computer program is executed by the processing device 601, it performs the functions defined in the methods of embodiments of this disclosure.

[0131] It should be noted that the computer-readable medium described in this disclosure can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in connection with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.

[0132] In some implementations, clients and servers can communicate using any currently known or future-developed network protocol such as HTTP (Hypertext Transfer Protocol) and can interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet of Things), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any currently known or future-developed networks.

[0133] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.

[0134] The aforementioned computer-readable medium carries one or more programs that, when executed by the electronic device, cause the electronic device to:

[0135] Obtain the target audio and the target subtitle text of the target audio;

[0136] If the duration of the target audio is greater than the first preset duration, the target audio is sliced ​​according to the slice duration to obtain multiple first target audios;

[0137] Determine the first audio feature information for each of the first target audio segments;

[0138] If the duration of the target audio is less than or equal to the second preset duration, then all the first audio feature information is concatenated to obtain the target audio feature information of the target audio, wherein the second preset duration is greater than the first preset duration;

[0139] Based on the target subtitle text and the target audio feature information, subtitle information corresponding to the target audio is generated.

[0140] Computer program code for performing the operations of this disclosure can be written in one or more programming languages ​​or a combination thereof, including but not limited to object-oriented programming languages ​​such as Java, Smalltalk, and C++, as well as conventional procedural programming languages ​​such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0141] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0142] The modules described in the embodiments of this disclosure can be implemented in software or in hardware. The names of the modules are not necessarily limiting in certain circumstances; for example, the first processing module can also be described as a "first slicing module".

[0143] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: Field Programmable Gate Arrays (FPGAs), Application-Specific Integrated Circuits (ASICs), Application Standard Products (ASSPs), System-on-Chip (SoCs), Complex Programmable Logic Devices (CPLDs), and so on.

[0144] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0145] According to one or more embodiments of this disclosure, Example 1 provides an audio subtitle alignment method, including:

[0146] Obtain the target audio and the target subtitle text of the target audio;

[0147] If the duration of the target audio is greater than the first preset duration, the target audio is sliced ​​according to the slice duration to obtain multiple first target audios;

[0148] Determine the first audio feature information for each of the first target audio segments;

[0149] If the duration of the target audio is less than or equal to the second preset duration, then all the first audio feature information is concatenated to obtain the target audio feature information of the target audio, wherein the second preset duration is greater than the first preset duration;

[0150] Based on the target subtitle text and the target audio feature information, subtitle information corresponding to the target audio is generated.

[0151] According to one or more embodiments of this disclosure, Example 2 provides the method of Example 1, the method further comprising:

[0152] If the duration of the target audio is greater than the second preset duration, then multiple consecutive first target audios are merged to obtain multiple second target audios, wherein the duration of each second target audio does not exceed the second preset duration;

[0153] For each second target audio, the first audio feature information in the second target audio is concatenated to obtain the second audio feature information;

[0154] Based on the second audio feature information and the target subtitle text, determine the subtitle text segment corresponding to the second target audio;

[0155] Based on the second audio feature information and the subtitle text fragment, the alignment information of each frame of audio in each second target audio is determined, wherein the alignment information includes the frame of audio and the characters in the target subtitle text that match the frame of audio;

[0156] The alignment information of each second target audio is concatenated to generate subtitle information corresponding to the target audio.

[0157] According to one or more embodiments of this disclosure, Example 3 provides the method of Example 2, the method further comprising:

[0158] The spliced ​​subtitle text in the subtitle information is compared with the target subtitle text to determine whether there is any missing text;

[0159] If the missing text is determined to exist, then based on the target subtitle text, a first subtitle text segment and a second subtitle text segment adjacent to the missing text are determined, wherein the time information corresponding to the first subtitle text segment is earlier than the time information corresponding to the second subtitle text segment;

[0160] The text insertion time is determined based on the last moment in the alignment information of the first subtitle text segment and the earliest moment in the alignment information of the second subtitle text segment;

[0161] The missing text is inserted into the spliced ​​subtitle text based on the text insertion time to obtain updated subtitle information.

[0162] According to one or more embodiments of this disclosure, Example 4 provides the method of Example 2, wherein after the alignment information of each second target audio is spliced ​​to generate subtitle information corresponding to the target audio, the method further includes:

[0163] If the target subtitle text is determined to have duplicate text in the alignment information corresponding to the adjacent second target audio, then the confidence level of the duplicate text in the adjacent second target audio is determined respectively.

[0164] In adjacent second target audios, the repeated text in the alignment information of the second target audio with low confidence is deleted to obtain updated subtitle information.

[0165] According to one or more embodiments of this disclosure, Example 5 provides the method of Example 4, wherein determining the confidence level of the repeated text in adjacent second target audio segments includes:

[0166] Determine the matching confidence of each character of the repeated text in two adjacent second target audios;

[0167] The confidence level of the repeated text in the second target audio is determined based on the matching confidence level of each character in the second target audio.

[0168] According to one or more embodiments of this disclosure, Example 6 provides the method of Example 1, wherein determining the first segment audio feature information for each of the first target audio segments includes:

[0169] The first target audio is input into a pre-trained feature extraction model to obtain the audio feature information of the first target audio. The feature extraction model is an encoder in a model trained based on sample audio and the text corresponding to the sample audio.

[0170] According to one or more embodiments of this disclosure, Example 7 provides the method of Example 1, wherein generating subtitle information corresponding to the target audio based on the target subtitle text and the target audio feature information includes:

[0171] Based on the target subtitle text and the target audio feature information, the alignment information of each frame of audio in the target audio is determined, wherein the alignment information includes the frame of audio and the characters in the target subtitle text that match the frame of audio;

[0172] Based on the alignment information and frame length of each audio frame, subtitle information corresponding to the target audio is generated.

[0173] According to one or more embodiments of this disclosure, Example 8 provides an audio subtitle alignment device, comprising:

[0174] The acquisition module is used to acquire the target audio and the target subtitle text of the target audio;

[0175] The first processing module is used to slice the target audio according to the slice duration if the duration of the target audio is greater than the first preset duration, so as to obtain multiple first target audios.

[0176] The first determining module is used to determine the first segment audio feature information of each of the first target audio segments;

[0177] The second processing module is used to concatenate all the first audio feature information to obtain the target audio feature information of the target audio if the duration of the target audio is less than or equal to the second preset duration, wherein the second preset duration is greater than the first preset duration.

[0178] The first generation module is used to generate subtitle information corresponding to the target audio based on the target subtitle text and the target audio feature information.

[0179] According to one or more embodiments of this disclosure, Example 9 provides an apparatus of Example 8, the apparatus further comprising:

[0180] The third processing module is used to merge multiple consecutive first target audios to obtain multiple second target audios if the duration of the target audio is greater than the second preset duration, wherein the duration of each second target audio does not exceed the second preset duration.

[0181] The fourth processing module is used to concatenate the first audio feature information in each second target audio for each second target audio to obtain second audio feature information.

[0182] The second determining module is used to determine the subtitle text segment corresponding to the second target audio based on the second audio feature information and the target subtitle text;

[0183] The third determining module is used to determine the alignment information of each frame of audio in each second target audio according to the second audio feature information and the subtitle text fragment, wherein the alignment information includes the frame of audio and the characters in the target subtitle text that match the frame of audio;

[0184] The second generation module is used to splice the alignment information of each second target audio to generate subtitle information corresponding to the target audio.

[0185] According to one or more embodiments of this disclosure, Example 10 provides an apparatus of Example 9, the apparatus further comprising:

[0186] The comparison module is used to compare the spliced ​​subtitle text in the subtitle information with the target subtitle text to determine whether there is any missing text;

[0187] The fourth determining module is used to determine, if it is determined that the missing text exists, a first subtitle text segment and a second subtitle text segment adjacent to the missing text based on the target subtitle text, wherein the time information corresponding to the first subtitle text segment is earlier than the time information corresponding to the second subtitle text segment;

[0188] The fifth determining module is used to determine the text insertion time based on the last time in the alignment information corresponding to the first subtitle text segment and the earliest time in the alignment information corresponding to the second subtitle text segment;

[0189] The first update module is used to insert the missing text into the spliced ​​subtitle text based on the text insertion time, so as to obtain the updated subtitle information.

[0190] According to one or more embodiments of this disclosure, Example 11 provides the apparatus of Example 9, the apparatus further comprising:

[0191] The sixth determining module is used to determine the confidence level of the repeated text in the adjacent second target audio if the target subtitle text is determined to have repeated text in the alignment information corresponding to the adjacent second target audio.

[0192] The second update module is used to delete the duplicate text in the alignment information of the second target audio with low confidence in adjacent second target audios, so as to obtain updated subtitle information.

[0193] According to one or more embodiments of this disclosure, Example 12 provides an apparatus of Example 11, wherein a sixth determining module is configured to determine the confidence level of the repeated text in adjacent second target audio in the following manner:

[0194] Determine the matching confidence of each character of the repeated text in two adjacent second target audios;

[0195] The confidence level of the repeated text in the second target audio is determined based on the matching confidence level of each character in the second target audio.

[0196] According to one or more embodiments of this disclosure, Example 13 provides the apparatus of Example 8, wherein a first determining module is configured to determine first segment audio feature information for each of the first target audio segments in the following manner:

[0197] The first target audio is input into a pre-trained feature extraction model to obtain the audio feature information of the first target audio. The feature extraction model is an encoder in a model trained based on sample audio and the text corresponding to the sample audio.

[0198] According to one or more embodiments of this disclosure, Example 14 provides the apparatus of Example 8, wherein the second processing module includes:

[0199] The first determining submodule is used to determine the alignment information of each frame of audio in the target audio based on the target subtitle text and the target audio feature information, wherein the alignment information includes the frame of audio and the characters in the target subtitle text that match the frame of audio;

[0200] The generation submodule is used to generate subtitle information corresponding to the target audio based on the alignment information and frame length of each audio frame.

[0201] According to one or more embodiments of the present disclosure, Example 15 provides a computer-readable medium having a computer program stored thereon that, when executed by a processing device, implements the steps of the method described in any of Examples 1 to 7.

[0202] According to one or more embodiments of this disclosure, Example 16 provides an electronic device including: a storage device having a computer program stored thereon; and a processing device for executing the computer program in the storage device to implement the steps of the method described in any of Examples 1 to 7.

[0203] The above description is merely a preferred embodiment of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features disclosed in this disclosure that have similar functions.

[0204] Furthermore, while the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in a sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, while several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of this disclosure. Certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments.

[0205] Although the subject matter has been described using language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely illustrative forms of implementing the claims. Regarding the apparatus in the above embodiments, the specific manner in which the various modules perform their operations has been described in detail in the embodiments relating to the method, and will not be elaborated upon here.

Claims

1. A method of subtitle alignment for audio, characterized by, The method comprises: obtaining target audio and target subtitle text of the target audio; if a time length of the target audio is greater than a first preset time length, performing slicing processing on the target audio according to a slice time length to obtain a plurality of first target audios; determining first audio feature information of each of the first target audios; if the time length of the target audio is less than or equal to a second preset time length, splicing all the first audio feature information to obtain target audio feature information of the target audio, wherein the second preset time length is greater than the first preset time length; determining alignment information of each frame of audio in the target audio according to the target subtitle text and the target audio feature information, wherein the alignment information comprises the frame of audio and a character in the target subtitle text matched with the frame of audio; generating subtitle information corresponding to the target audio according to the alignment information of each frame of audio and a frame length.

2. The method of claim 1, wherein, The method further comprises: if the time length of the target audio is greater than the second preset time length, merging a plurality of continuous first target audios to obtain a plurality of second target audios, wherein a time length of each of the second target audios does not exceed the second preset time length; for each of the second target audios, splicing each of the first audio feature information in the second target audio to obtain second audio feature information; determining a subtitle text segment corresponding to the second target audio according to the second audio feature information and the target subtitle text; determining alignment information of each frame of audio in each of the second target audios according to the second audio feature information and the subtitle text segment, wherein the alignment information comprises the frame of audio and a character in the target subtitle text matched with the frame of audio; splicing the alignment information of each of the second target audios to generate subtitle information corresponding to the target audio.

3. The method of claim 2, wherein, After the splicing of the alignment information of each of the second target audios to generate the subtitle information corresponding to the target audio, the method further comprises: comparing spliced subtitle text in the subtitle information with the target subtitle text to determine whether there is missing text; if it is determined that there is missing text, determining a first subtitle text segment and a second subtitle text segment adjacent to the missing text according to the target subtitle text, wherein time information corresponding to the first subtitle text segment is earlier than time information corresponding to the second subtitle text segment; determining a text insertion time based on a last time in corresponding alignment information of the first subtitle text segment and an earliest time in corresponding alignment information of the second subtitle text segment; inserting the missing text in the spliced subtitle text based on the text insertion time to obtain updated subtitle information.

4. The method of claim 2, wherein, After the splicing of the alignment information of each of the second target audios to generate the subtitle information corresponding to the target audio, the method further comprises: If it is determined, based on the target subtitle text, that there is repeated text in alignment information corresponding to adjacent second target audios, confidence degrees of the repeated text in the adjacent second target audios are respectively determined; In the adjacent second target audios, the repeated text in the alignment information of the second target audio with the smaller confidence degree is deleted to obtain updated subtitle information.

5. The method of claim 4, wherein, The confidence degrees of the repeated text in the adjacent second target audios are respectively determined, including: The matching confidence degrees of each character of the repeated text in two adjacent second target audios are respectively determined; The confidence degree of the repeated text in the second target audio is determined according to the matching confidence degree of each character in the second target audio.

6. The method of claim 1, wherein, The first audio feature information of each first target audio is determined, including: The first target audio is input into a pre-trained feature extraction model to obtain audio feature information of the first target audio, wherein the feature extraction model is an encoder in a model obtained by training based on sample audio and text corresponding to the sample audio.

7. An apparatus for aligning subtitles with audio, characterized by Including: An acquisition module is configured to acquire target audio and target subtitle text of the target audio; A first processing module is configured to perform slicing processing on the target audio according to a slicing time length to obtain a plurality of first target audios if a time length of the target audio is greater than a first preset time length. A first determination module is configured to determine first audio feature information of each first target audio. A second processing module is configured to splice all the first audio feature information to obtain target audio feature information of the target audio if the time length of the target audio is less than or equal to a second preset time length, wherein the second preset time length is greater than the first preset time length. A first generation module is configured to determine alignment information of each frame of audio in the target audio according to the target subtitle text and the target audio feature information, wherein the alignment information includes the frame of audio and a character in the target subtitle text matched with the frame of audio, and generate subtitle information corresponding to the target audio according to the alignment information of each frame of audio and a frame length.

8. A computer readable medium having stored thereon a computer program, characterized in that, The program is executed by a processing device to implement steps of the method in any one of claims 1-6.

9. An electronic device, comprising: Including: A storage device having at least one computer program stored thereon; At least one processing device is configured to execute the at least one computer program in the storage device to implement steps of the method in any one of claims 1-6.

Citation Information

Patent Citations

  • Sound signal subtitle matching method and device

    CN106792097A

  • Method and device for automatically adding subtitle fragments and computer equipment

    CN112738563A