Video translation methods, devices, computer equipment, and storage media

CN119360854BActive Publication Date: 2026-08-14SHENZHEN COOCAA NETWORK TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-16
Publication Date
2026-08-14

AI Technical Summary

Technical Problem

[0003]本发明实施例提供一种视频翻译方法、装置、计算机设备及存储介质,以解决视频翻译问题

Benefits of technology

[0036]上述视频翻译方法、装置、计算机设备及存储介质,通过提取视频中的原始音频以及视频中的关键帧;根据所述原始音频以及所述关键帧,以判断是否在云端已提前存储翻译后的目标翻译音频;若判断没有提前存储翻译后的目标翻译音频,则将所述原始音频转换为目标翻译文本;提取所述原始音频特征,以根据所述目标翻译文本与提取到的所述原始音频特征合成目标翻译音频,其中,所述提取到的所述原始音频特征,包括对所述原始音频的语气特征进行提取;将所述原始音频替换为目标翻译音频。本申请通过判断云端存储目标翻译音频是否存在和分段翻译,从而最大化缩短视频翻译后加载时长,以提升用户视频观看的流畅度。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119360854B_ABST
    Figure CN119360854B_ABST
Patent Text Reader

Abstract

This invention discloses a video translation method, apparatus, computer device, and storage medium. The method involves extracting original audio and keyframes from a video; determining whether the target translated audio has been pre-stored in the cloud based on the original audio and keyframes; if no target translated audio has been pre-stored, converting the original audio into target translated text; extracting features from the original audio; and synthesizing the target translated audio based on the target translated text and the extracted original audio features, wherein the extracted original audio features include extracting tone features from the original audio; and finally, replacing the original audio with the target translated audio. This application maximizes the reduction of video loading time after translation by determining whether the target translated audio exists in the cloud and by translating in segments, thereby improving the smoothness of video viewing for users.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of video translation, and more particularly to a video translation method, apparatus, computer device, and storage medium. Background Technology

[0002] As people's living standards improve, they have more and more opportunities to watch foreign language videos. However, most existing foreign language videos are in the foreign language with Chinese subtitles, which affects viewers' enjoyment. Therefore, there is an urgent need for a method to solve the language conversion problem when watching foreign language videos. Summary of the Invention

[0003] This invention provides a video translation method, apparatus, computer device, and storage medium to solve the video translation problem.

[0004] The first aspect of this application provides a video translation method, including:

[0005] Extract the original audio and keyframes from the video;

[0006] Based on the original audio and the keyframes, determine whether the translated target audio has been pre-stored in the cloud;

[0007] If it is determined that the target translated audio has not been stored in advance, the original audio segments are converted into target translated text.

[0008] The original audio features are extracted to synthesize target translated audio based on the target translated text and the extracted original audio features, wherein the extracted original audio features include the extraction of tone features from the original audio.

[0009] Optionally, converting the original audio segments into the target translation text includes:

[0010] Based on the content of the original audio and the target translation language, the content of the original audio is segmented and converted into target translation text, and semantic calibration is performed on the target translation text.

[0011] Optionally, converting the original audio segments into the target translation text includes:

[0012] The original audio is divided into several original audio segments;

[0013] Based on the playback time of the original audio segments, the estimated playback time of each audio segment is obtained sequentially.

[0014] Based on the estimated playback time of each audio segment, the original audio segments are sequentially converted into target translation text.

[0015] Optionally, features are extracted from the original audio to synthesize target translated audio based on the target translated text and the extracted audio features, including:

[0016] Extract the audio features of the original audio and analyze the audio features of the original audio;

[0017] Based on the audio feature analysis results and the target translation text, the target translation audio with the same features as the original audio is generated.

[0018] Optionally, based on the estimated playback time of each audio segment, the original audio segments are sequentially converted into target translated text, including the following:

[0019] Check the translation progress of the original audio to determine if there is any untranslated original audio.

[0020] If there is untranslated audio, calculate the target translated audio generation time based on the estimated playback time of the audio segment;

[0021] Based on the target audio generation time, the original audio segments are sequentially converted into target translated text.

[0022] Optionally, the target translated audio is synthesized, followed by:

[0023] Replace the original audio with the target translated audio;

[0024] The target audio and video animation are merged to form the translated target video.

[0025] A second aspect of this application provides a video translation device, comprising:

[0026] The audio extraction module is used to extract the original audio and keyframes from the video.

[0027] The judgment module determines, based on the original audio and the keyframes, whether the translated target audio has been pre-stored in the cloud.

[0028] The first audio processing module is used to convert the original audio into target translated text if it is determined that the target translated audio has not been stored in advance.

[0029] The second audio processing module extracts the original audio features to synthesize the target translated audio based on the target translated text and the extracted original audio features, wherein the extracted original audio features include the extraction of tone features from the original audio;

[0030] The audio output module replaces the original audio with the target translated audio.

[0031] Optionally, the second audio processing module includes:

[0032] Extract the audio features of the original audio and analyze the audio features of the original audio;

[0033] Based on the audio feature analysis results and the target translation text, the target translation audio with the same features as the original audio is generated.

[0034] A third aspect of this application provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the video translation method described above.

[0035] A fourth aspect of this application provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the aforementioned video translation method.

[0036] The aforementioned video translation method, apparatus, computer equipment, and storage medium extract the original audio and keyframes from the video; determine whether the target translated audio has been pre-stored in the cloud based on the original audio and keyframes; if it is determined that the target translated audio has not been pre-stored, the original audio is converted into target translated text; features of the original audio are extracted to synthesize target translated audio based on the target translated text and the extracted original audio features, wherein the extracted original audio features include the extraction of tone features of the original audio; and the original audio is replaced with target translated audio. This application maximizes the reduction of video loading time after translation by determining whether the target translated audio exists in the cloud and translating in segments, thereby improving the smoothness of video viewing for users. Attached Figure Description

[0037] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the description of the embodiments of the present invention will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0038] Figure 1 This is a schematic diagram of an application environment for a video translation method according to an embodiment of the present invention;

[0039] Figure 2 This is a flowchart of a video translation method according to an embodiment of the present invention;

[0040] Figure 3 This is a schematic diagram of a video translation device according to an embodiment of the present invention;

[0041] Figure 4 This is a schematic diagram of a computer device according to an embodiment of the present invention. Detailed Implementation

[0042] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0043] The video translation method provided in this embodiment of the invention can be applied to, for example... Figure 1 The application environment is shown. Specifically, this video translation method is applied in a video translation system, which includes, for example, […]. Figure 1 The diagram illustrates a client and a cloud platform. The client communicates with the cloud via a network to replace the original audio with the target translated audio. The client, also known as the user terminal, refers to the program that provides local services to the client, corresponding to the cloud platform. The client can be installed on, but is not limited to, various personal computers, laptops, smartphones, tablets, and portable wearable devices. The cloud platform can be implemented using a standalone cloud or a cloud cluster consisting of multiple cloud platforms.

[0044] It should be understood that, when used in this specification and the appended claims, the term "comprising" indicates the presence of the described features, integrals, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or collections thereof.

[0045] It should also be understood that the term “and / or” as used in this specification and the appended claims refers to any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.

[0046] As used in this specification and the appended claims, the term "if" may be interpreted, depending on the context, as "when," "once," "in response to determination," or "in response to detection." Similarly, the phrase "if determined" or "if [described condition or event] is detected" may be interpreted, depending on the context, as meaning "once determined," "in response to determination," "once [described condition or event] is detected," or "in response to detection of [described condition or event]."

[0047] Furthermore, in the description of this invention and the appended claims, the terms "first," "second," "third," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.

[0048] References to "one embodiment" or "some embodiments" as described in this specification mean that one or more embodiments of the invention include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "comprising," "including," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.

[0049] It should be understood that the sequence number of each step in the following embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0050] To illustrate the technical solution of the present invention, specific embodiments are described below.

[0051] In one embodiment, such as Figure 2 As shown, a video translation method is provided, which is applied to... Figure 1 Taking the cloud as an example, the explanation includes the following steps:

[0052] S10: Extract the raw audio and keyframes from the video.

[0053] S20: Based on the original audio and the keyframes, determine whether the translated target audio has been pre-stored in the cloud.

[0054] S30: If it is determined that the target translated audio has not been stored in advance, the original audio segments are converted into target translated text.

[0055] Specifically, this application provides a video player that extracts the original audio and some keyframe content from the video when the user is playing the video, determines whether the original audio has been translated in advance, and if it is determined that the target translated audio has not been stored in advance, then the original audio is segmented and converted into target translated text.

[0056] For example, when a user uses the video player in this video to play an English version of a video, the video player will extract the original audio and some keyframe content from the video in advance, determine whether the original audio has been translated into the target translation text in advance, and store the original audio in the video player's cloud. If, based on the video player's pre-extraction of the original audio and some keyframe content, it is determined that the original audio has not been translated in advance, then the original audio will be segmented and converted into the target translation text. The target translation text is determined according to the target translation language, which is the language the user needs, such as Chinese, Japanese, Korean, French, etc.

[0057] In one embodiment, converting the original audio segments into target translation text includes: converting the content of the original audio into target translation text segments according to the content of the original audio and the target translation language, and performing semantic calibration on the target translation text.

[0058] Specifically, translating audio from a video takes time. If playback is delayed until the translation is complete, the video loading time will be too long, affecting the user's viewing experience. Therefore, in this application, the original audio from the video and the target language to be translated are first obtained. Then, the content of the original audio is segmented and converted into the target translation text, and the target translation text is semantically calibrated to ensure the accuracy of the translation.

[0059] In this embodiment, by extracting the original audio and keyframes from the video, it is determined whether the target translated audio has been pre-stored in the cloud based on the original audio and keyframes. If it is determined that the target translated audio has not been pre-stored, the original audio is segmented and converted into target translated text, and semantic calibration is performed on the translated text. This ensures translation accuracy while reducing video loading time and improving the user's viewing experience.

[0060] S40: Extract the original audio features to synthesize the target translated audio based on the target translated text and the extracted original audio features.

[0061] Specifically, after obtaining the target translation text translated from the original audio, the audio features of the original audio are extracted. Based on the content of the target translation text and the extracted audio features, audio synthesis technology is used to generate the target translation audio, and the original audio is replaced with the generated target translation audio. The extracted original audio features include the extraction of tone features from the original audio.

[0062] In this embodiment, by extracting the original audio features, a target translated audio is synthesized based on the target translated text and the audio features, and the original audio is replaced with the target translated audio. This allows the audio in the target language to be played directly when the video is played, eliminating the need for subtitles to understand the dialogue.

[0063] In this embodiment, the original audio and keyframes of the video are extracted. Based on the original audio and keyframes, it is determined whether the target translated audio has been pre-stored in the cloud. If it is determined that the target translated audio has not been pre-stored, the original audio is converted into target translated text. Features of the original audio are extracted, and the target translated audio is synthesized based on the target translated text and the extracted original audio features. The extracted original audio features include the extraction of tone features from the original audio. The original audio is then replaced with the target translated audio. This embodiment maximizes the reduction of video loading time after translation by determining whether the target translated audio exists in the cloud and translating in segments, thereby improving the smoothness of video viewing for users.

[0064] In one embodiment, a video translation method is provided. Step S30, which involves converting the original audio segments into target translation text, specifically includes the following steps:

[0065] S31: The original audio is segmented to form several original audio segments;

[0066] S32: According to the playback time of the original audio segment, obtain the estimated playback time of each audio segment in sequence;

[0067] S33: Based on the estimated playback time of each audio segment, the original audio segments are sequentially converted into target translation text.

[0068] Specifically, when a user plays a video, translating the original audio in the video into target audio takes time. If the original audio is translated completely, it will increase the user's waiting time. Therefore, this application segments the original audio into several original audio segments. According to the playback time of the original audio segments, the estimated playback time of each audio segment is obtained in sequence. Based on the estimated playback time of each audio segment, the original audio segments are converted into target translation text in sequence.

[0069] For example, when playing an English movie and needing to translate it into Chinese audio output, all the original audio in the video is obtained, the audio to be translated is segmented into several original audio segments to be translated, the estimated playback time of each audio segment is obtained, and before this estimated playback time, the audio segment needs to be translated into the target translation text. The target translation audio is generated through audio synthesis technology and stored in the cloud of the video player. Then, the translation is performed in advance according to the playback time of the audio segments.

[0070] In this embodiment, the original audio is segmented into several original audio segments. The estimated playback time of each audio segment is obtained sequentially according to the playback time of the original audio segments. Based on the estimated playback time of each audio segment, the original audio segments are converted into target translation text in sequence, ensuring that when the audio segment is played, it can be played directly without waiting.

[0071] In one embodiment, a video translation method is provided, wherein in step S40, features of the original audio are extracted to synthesize target translated audio based on the target translated text and the extracted audio features, specifically including the following steps.

[0072] S41: Extract the audio features of the original audio and analyze the audio features of the original audio.

[0073] S42: Based on the audio feature analysis results and the target translation text, generate the target translation audio with the same features as the original audio.

[0074] Specifically, in order to ensure the consistency between the target translation audio and the original audio, it is necessary to extract and analyze the sound features of the original audio when generating the target translation audio. This includes analyzing the sound timbre, sound frequency, and other features of the original audio. Based on the analysis results and the target translation text, a target translation audio with the same features as the original audio is generated.

[0075] In this embodiment, by extracting the audio features of the original audio and analyzing them, a target translated audio with the same features as the original audio is generated based on the audio feature analysis results and the target translated text. This ensures the consistency between the target translated audio and the original audio, providing users with a better video viewing experience.

[0076] In one embodiment, in step S33, the original audio segments are sequentially converted into target translation text based on the estimated playback time of each audio segment. Prior to this, the following steps are included:

[0077] S331: Check the translation progress of the original audio and determine whether there is any original audio that has not been translated.

[0078] S332: If there is untranslated audio, calculate the target translated audio generation time based on the estimated playback time of the audio segment.

[0079] S333: Based on the target translation audio generation time, the original audio segments are sequentially converted into target translation text.

[0080] Specifically, during the translation of the original audio in the video, the video player checks the translation progress of the original audio to determine if there are any untranslated original audio segments. If there is untranslated audio, the target translation audio generation time is calculated based on the estimated playback time of the audio segment. Based on the target translation audio generation time, the original audio segments are sequentially converted into target translation text.

[0081] For example, the video is divided into several original audio segments, and the audio is translated according to the playback time of the audio. The target audio generation time is the time required to translate the next original audio segment. That is, the video player will divide the original audio into several original audio segments according to the time required to translate each audio segment, and the playback time of the original audio segment is the time required to translate the next original audio segment. Each original audio segment is translated sequentially according to the playback time of the target audio.

[0082] In this embodiment, by checking the translation progress of the original audio, it is determined whether there are any untranslated original audio segments. If there are untranslated audio segments, the target translation audio generation time is calculated based on the estimated playback time of the audio segments. Based on the target translation audio generation time, the original audio segments are sequentially converted into target translation text, ensuring that each audio segment has sufficient time for translation and guaranteeing the smoothness of video playback.

[0083] In one embodiment, after step S50, i.e., synthesizing the target translation audio, the original audio is replaced with the target translation audio of the video translation, and then the target translation audio and video animation are merged to form the translated target video, ensuring that the user can watch it directly without having to translate it again the next time.

[0084] The aforementioned video translation method, apparatus, computer equipment, and storage medium extract the original audio and keyframes from the video; determine whether the target translated audio has been pre-stored in the cloud based on the original audio and keyframes; if it is determined that the target translated audio has not been pre-stored, the original audio is converted into target translated text; features of the original audio are extracted to synthesize target translated audio based on the target translated text and the extracted original audio features, wherein the extracted original audio features include the extraction of tone features of the original audio; and the original audio is replaced with target translated audio. This application maximizes the reduction of video loading time after translation by determining whether the target translated audio exists in the cloud and translating in segments, thereby improving the smoothness of video viewing for users.

[0085] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0086] In one embodiment, a video translation device is provided, which corresponds one-to-one with the video translation methods described in the above embodiments. For example... Figure 3 As shown, the video translation device includes an audio extraction module, a judgment module, a first audio processing module, a second audio processing module, and an audio output module. Detailed descriptions of each functional module are as follows:

[0087] The audio extraction module is used to extract the original audio and keyframes from the video.

[0088] The judgment module determines, based on the original audio and the keyframes, whether the translated target audio has been pre-stored in the cloud.

[0089] The first audio processing module is used to convert the original audio into target translated text if it is determined that the target translated audio has not been stored in advance.

[0090] Specifically, this application provides a video player that extracts the original audio and some keyframe content from the video when the user is playing the video, determines whether the original audio has been translated in advance, and if it is determined that the target translated audio has not been stored in advance, then the original audio is segmented and converted into target translated text.

[0091] For example, when a user uses the video player in this video to play an English version of a video, the video player will extract the original audio and some keyframe content from the video in advance, determine whether the original audio has been translated into the target translation text in advance, and store the original audio in the video player's cloud. If, based on the video player's pre-extraction of the original audio and some keyframe content, it is determined that the original audio has not been translated in advance, then the original audio will be segmented and converted into the target translation text. The target translation text is determined according to the target translation language, which is the language the user needs, such as Chinese, Japanese, Korean, French, etc.

[0092] In one embodiment, converting the original audio segments into target translation text includes: converting the content of the original audio into target translation text segments according to the content of the original audio and the target translation language, and performing semantic calibration on the target translation text.

[0093] Specifically, translating audio from a video takes time. If playback is delayed until the translation is complete, the video loading time will be too long, affecting the user's viewing experience. Therefore, in this application, the original audio from the video and the target language to be translated are first obtained. Then, the content of the original audio is segmented and converted into the target translation text, and the target translation text is semantically calibrated to ensure the accuracy of the translation.

[0094] In this embodiment, by extracting the original audio and keyframes from the video, it is determined whether the target translated audio has been pre-stored in the cloud based on the original audio and keyframes. If it is determined that the target translated audio has not been pre-stored, the original audio is segmented and converted into target translated text, and semantic calibration is performed on the translated text. This ensures translation accuracy while reducing video loading time and improving the user's viewing experience.

[0095] The second audio processing module extracts the original audio features to synthesize the target translated audio based on the target translated text and the extracted original audio features, wherein the extracted original audio features include the extraction of tone features from the original audio;

[0096] The audio output module replaces the original audio with the target translated audio.

[0097] Specifically, after obtaining the target translation text translated from the original audio, the audio features of the original audio are extracted. Based on the content of the target translation text and the extracted audio features, audio synthesis technology is used to generate the target translation audio, and the generated target translation audio replaces the original audio.

[0098] In this embodiment, by extracting the original audio features, a target translated audio is synthesized based on the target translated text and the audio features, and the original audio is replaced with the target translated audio. This allows the audio in the target language to be played directly when the video is played, eliminating the need for subtitles to understand the dialogue.

[0099] Optionally, the second audio processing module includes:

[0100] Extract the audio features of the original audio and analyze the audio features of the original audio;

[0101] Based on the audio feature analysis results and the target translation text, the target translation audio with the same features as the original audio is generated.

[0102] Specifically, in order to ensure the consistency between the target translation audio and the original audio, it is necessary to extract and analyze the sound features of the original audio when generating the target translation audio. This includes analyzing the sound timbre, sound frequency, and other features of the original audio. Based on the analysis results and the target translation text, a target translation audio with the same features as the original audio is generated.

[0103] Specific limitations regarding the video translation device can be found in the limitations of the video translation method above, and will not be repeated here. Each module in the aforementioned video translation device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device in hardware form, or stored in the memory of a computer device in software form, so that the processor can call and execute the corresponding operations of each module.

[0104] In one embodiment, a computer device is provided, which may be cloud-based, and its internal structure diagram may be as follows: Figure 4 As shown, the computer device includes a processor, memory, network interface, and database connected via a system bus. The processor provides computing and control capabilities. The memory includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores the operating system, computer programs, and the database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The database stores the replacement of the original audio with the target translated audio. The network interface communicates with external terminals via a network connection. When executed by the processor, the computer program implements a video translation method.

[0105] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the video translation method described in the above embodiment, for example... Figure 1 The diagram shown illustrates an application environment for the video translation method. Figures 2 to 3 As shown, to avoid repetition, it will not be described again here. Alternatively, when the processor executes the computer program, it implements the functions of each module / unit in this embodiment of the video translation device, for example... Figure 3 The video translation function shown will not be described again here to avoid repetition.

[0106] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When executed by a processor, the computer program implements the video translation method described in the above embodiment, for example... Figure 1 The diagram shown illustrates an application environment for the video translation method. Figures 2 to 3 As shown, to avoid repetition, it will not be described again here. Alternatively, when the computer program is executed by the processor, it implements the functions of each module / unit in this embodiment of the video translation device, for example... Figure 3 The video translation function shown will not be described again here to avoid repetition.

[0107] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. This computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in a variety of forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0108] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.

[0109] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.

Claims

1. A video translation method, characterized in that, include: Extract the original audio and keyframes from the video; Based on the original audio and the keyframes, determine whether the translated target audio has been pre-stored in the cloud; If it is determined that the target translated audio has not been stored in advance, the original audio segments are converted into target translated text. The original audio features are extracted to synthesize target translated audio based on the target translated text and the extracted original audio features. The extracted original audio features include the extraction of tone features, timbre and frequency of the original audio. Converting the original audio segments into target translation text includes: obtaining the target translation language; Based on the content of the original audio and the target translation language, the content of the original audio is segmented and converted into target translation text, and semantic calibration is performed on the target translation text; Extracting features from the original audio to synthesize target translated audio based on the target translated text and the extracted audio features includes: extracting audio features from the original audio; analyzing the audio features of the original audio; and generating target translated audio with the same features as the original audio based on the audio feature analysis results and the target translated text. The method further includes: dividing the original audio into several original audio segments according to the time required to translate each audio segment, and the playback time of the original audio segment is the translation time of the next original audio segment; and translating each original audio segment sequentially according to the playback time of the target audio to be translated. Based on the estimated playback time of each audio segment, the original audio segments are sequentially converted into target translated text. Before this, the process includes: checking the translation progress of the original audio to determine if there is any untranslated original audio; if there is untranslated audio, calculating the target translated audio generation time based on the estimated playback time of the audio segments; and sequentially converting the original audio segments into target translated text based on the target translated audio generation time.

2. The video translation method according to claim 1, characterized in that, Converting the original audio segments into target translated text includes: The original audio is divided into several original audio segments; Based on the playback time of the original audio segments, the estimated playback time of each audio segment is obtained sequentially. Based on the estimated playback time of each audio segment, the original audio segments are sequentially converted into target translation text.

3. The video translation method according to claim 1, characterized in that, Synthesize the target translated audio, which then includes: Replace the original audio with the target translated audio; The target audio and video animation are merged to form the translated target video.

4. A video translation device, characterized in that, For implementing the method of claim 1, the video translation device includes: The audio extraction module is used to extract the original audio and keyframes from the video. The judgment module determines, based on the original audio and the keyframes, whether the translated target audio has been pre-stored in the cloud. The first audio processing module is used to convert the original audio into target translated text if it is determined that the target translated audio has not been stored in advance. The second audio processing module extracts the original audio features to synthesize the target translated audio based on the target translated text and the extracted original audio features. The extracted original audio features include the extraction of tone features, timbre and frequency of the original audio. The audio output module replaces the original audio with the target translated audio.

5. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the video translation method according to any one of claims 1 to 3.

6. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the video translation method according to any one of claims 1 to 3.

Citation Information

Patent Citations

  • Method and system for automatically translating accompanying sounds

    CN112423106A

  • Video translation method, system and device and storage medium

    CN112562721A