Video live broadcast content translation method and device, electronic equipment and storage medium

This modular video live streaming translation method solves the problem of poor real-time performance in video live streaming translation, achieving audio and video synchronization and real-time multilingual broadcasting, and is suitable for live streaming platforms and various types of users.

CN121122280APending Publication Date: 2025-12-12TERMINUSBEIJING TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202511027378.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-24
Publication Date
2025-12-12

AI Technical Summary

Technical Problem

Existing technologies for live video translation suffer from poor real-time performance, failing to achieve millisecond-level response times. This results in a lack of synchronization between video and audio, impacting the effectiveness of the broadcast and audience reach.

Method used

The video live streaming translation method adopts a modular design, including speech recognition, translation, and text-to-speech modules, to realize the separation, translation, and synthesis of audio and video. By deploying functional modules locally, audio transmission delays are avoided, and end-to-end automated translation and broadcasting are achieved.

Benefits of technology

It achieves fully automated and real-time processing, improves processing efficiency, ensures audio and video synchronization, supports multi-language expansion and cross-platform streaming, and is suitable for stable operation in live streaming environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121122280A_ABST
    Figure CN121122280A_ABST
Patent Text Reader

Abstract

The invention provides a video live broadcast content translation method and device, electronic equipment and a storage medium, and relates to the technical field of video processing. The method is applied to a terminal, the terminal comprises a voice recognition module, a translation module and a text-to-voice module, and the method comprises the following steps: acquiring a live video stream, and splitting an audio track and a video track of the live video stream to obtain an audio and a video; a voice recognition module is adopted to recognize the audio, so that the audio is transferred into a text, and the text is the text of the first language; translating the text of the first language into a text of a second language by adopting a translation module; a text-to-language module is adopted to convert the text of the second language into voice; and performing video synthesis on the voice and the video to obtain a target video. According to the method and the device, the function module is deployed locally, so that the delay generated by transmitting the audio to the server is avoided, the processing efficiency is improved, and the real-time rebroadcasting of the live broadcast content in other languages is realized.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates to the technical field of video processing, and in particular, to a video live content translation method and device, electronic equipment and storage medium. BACKGROUND

[0002] With the development of the Internet, mobile communication and cloud computing, real-time video live has become an important way for people to obtain information, disseminate content and interact. Whether it is online meetings, online teaching, or short video live, e-commerce live, the consumption of content in video form is growing exponentially. At the same time, the demand for global communication is increasingly evident, especially when exporting live content (such as international exhibitions, brand promotion, e-commerce sales, and corporate promotion), language barriers often arise, especially the lack of real-time translation capabilities, making it difficult for non-native speakers to understand live content in a timely manner, greatly affecting the dissemination effect and audience coverage.

[0003] In related technologies, the real-time translation of live video is poor, with a delay that cannot achieve a second-level response, causing the picture and voice to be out of sync. SUMMARY

[0004] Therefore, the purpose of the present disclosure is to provide a video live content translation method, device, electronic equipment and storage medium, which can solve the existing problems.

[0005] To achieve the above purpose, in a first aspect, the present disclosure provides a video live content translation method, comprising: applying to a terminal, the terminal comprising a speech recognition module, a translation module and a text-to-speech module, the method comprising: obtaining a video live stream, splitting the video live stream into an audio track and a video track to obtain audio and video; using the speech recognition module to recognize the audio to transcribe the audio into text, the text being text in a first language; using the translation module to translate the text in the first language into text in a second language; using the text-to-speech module to convert the text in the second language into speech; and video synthesizing the speech and the video to obtain a target video.

[0006] Secondly, a translation device for live video content is also provided, applied to a terminal. The terminal includes a speech recognition module, a translation module, and a text-to-speech module. The device includes: an acquisition unit configured to acquire a live video stream and split the live video stream into audio and video tracks to obtain audio and video; a first adoption unit configured to use the speech recognition module to recognize the audio and transcribe the audio into text, wherein the text is in a first language; a second adoption unit configured to use the translation module to translate the text in the first language into text in a second language; a third adoption unit configured to use the text-to-speech module to convert the text in the second language into speech; and a synthesis unit configured to synthesize the speech and the video to obtain a target video.

[0007] Thirdly, an electronic device is also provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, the processor running the computer program to implement the method of the first aspect.

[0008] Fourthly, a computer-readable storage medium is also provided, on which a computer program is stored, the computer program being executed by a processor to implement the method described in any one of the first aspects.

[0009] Fifthly, a computer program product is also provided, comprising a computer program that is executed by a processor to implement the method described in any one of the first aspects.

[0010] In summary, this disclosure offers at least the following advantages: Through modularization, it achieves fully automated and real-time processing, overcoming the limitations of existing solutions that rely heavily on manual operation and offline processing. It enables end-to-end automated translation and broadcasting, allowing for continuous and stable operation in live streaming environments. Furthermore, by deploying functional modules locally, this disclosure avoids the latency associated with transmitting audio to the server, improving processing efficiency and enabling real-time broadcasting of live content in other languages. Attached Figure Description

[0011] In the accompanying drawings, unless otherwise specified, the same reference numerals throughout the various drawings denote the same or similar parts or elements. These drawings are not necessarily drawn to scale. It should be understood that these drawings depict only some embodiments disclosed in this disclosure and should not be construed as limiting the scope of this disclosure.

[0012] Figure 1 A flowchart illustrating a method for translating live video content according to an embodiment of this disclosure is shown;

[0013] Figure 2Another flowchart of a method for translating live video content according to an embodiment of this disclosure is shown;

[0014] Figure 3 A schematic diagram of a translation device for live video content according to an embodiment of the present disclosure is shown;

[0015] Figure 4 A schematic diagram of the structure of an electronic device provided in an embodiment of the present disclosure is shown;

[0016] Figure 5 A schematic diagram of a storage medium provided according to an embodiment of the present disclosure is shown. Detailed Implementation

[0017] The present disclosure will now be described in further detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are for illustrative purposes only and are not intended to limit the invention. Furthermore, it should be noted that, for ease of description, only the parts relevant to the invention are shown in the accompanying drawings.

[0018] It should be noted that, unless otherwise specified, the embodiments and features described in this disclosure can be combined with each other. This disclosure will now be described in detail with reference to the accompanying drawings and embodiments.

[0019] Figure 1 This disclosure illustrates a method for translating live video content. In embodiments of this disclosure, the method is applied to a terminal, which includes a speech recognition module, a translation module, and a text-to-speech module. The method includes:

[0020] Step S101: Obtain the live video stream, and split the live video stream into audio and video tracks to obtain audio and video.

[0021] Step S102: The audio is recognized by a speech recognition module to transcribe the audio into text, which is text in a first language.

[0022] Step S103: Using the translation module, the text in the first language is translated into text in the second language.

[0023] Step S104: Use the text-to-language module to convert the text in the second language into speech.

[0024] Step S105: Combine the audio and video to obtain the target video.

[0025] All of the above modules can achieve their functions by calling the corresponding models.

[0026] This disclosure achieves fully automated and real-time processing through modularization, overcoming the problems of existing solutions that rely heavily on manual operation and offline processing (such as translating and uploading after recording). It realizes end-to-end automated translation and broadcasting, which can run continuously and stably in a live streaming environment. Furthermore, by deploying functional modules locally, this disclosure avoids the latency caused by transmitting audio to the server, improves processing efficiency, and enables real-time broadcasting of live content in other languages.

[0027] In some optional implementations of any embodiment of this disclosure, the step of using the text-to-speech module to convert the text in the second language into speech includes: replicating the voice of the target live streamer to obtain the sound features of the voice, the sound features including sound quality information, timbre information and rhythm information; and converting the text in the second language into speech having the sound features.

[0028] In some optional implementations of any embodiment of this disclosure, obtaining the live video stream includes: pulling the live video stream from a first live streaming platform; the step of synthesizing the audio and the video to obtain a target video includes: mixing and packaging the audio track of the audio and the video track of the video according to timestamps to obtain the target video to be pushed, wherein the target video to be pushed is used to push to a second live streaming platform so that the live streaming terminal corresponding to the second live streaming platform plays the target video of the second language audio.

[0029] Optionally, the terminal further includes a streaming module; the method further includes: using the streaming module to push the target video to be pushed to a second live streaming platform.

[0030] In some optional implementations of any embodiment of this disclosure, the step of using the translation module to translate the text in the first language into the text in the second language includes: calling the translation model in the translation module to perform real-time text translation of the text in the first language to obtain the text in the second language.

[0031] Optionally, there are multiple translation models, and the method further includes: obtaining real-time live streaming demand information, and selecting the translation model to be invoked from among the multiple translation models based on the real-time live streaming demand information.

[0032] Live streaming demand information can reflect the processing power required by the translation model. Different translation models have different processing powers.

[0033] In some optional implementations of any embodiment of this disclosure, the method further includes: acquiring hot words corresponding to various live streaming scene information; inputting the speech and text of the hot words corresponding to the various live streaming scene information into the speech recognition model in the speech recognition module; the step of using the speech recognition module to recognize the audio to transcribe the audio into text includes: acquiring live streaming scene information of the live video stream; inputting the live streaming scene information and the audio into the speech recognition model to obtain the text output by the speech recognition model.

[0034] Live streaming scene information can reflect the current activity scene of the people in the live stream, such as a sales scene or a performance scene.

[0035] Figure 2 This illustrates another method for translating live video content according to embodiments of the present disclosure. For example... Figure 2 As shown, the translation methods for this live video content include:

[0036] 1. Input live stream: The system pulls the live stream (RTMP / HLS format) from the original live streaming platform;

[0037] 2. Audio and video separation: The streaming media decoding module uses FFmpeg to split the stream into audio and video tracks; caching: The split audio and video are cached in memory to ensure a stable input for subsequent modules;

[0038] 3. ASR Recognition: Using deployed speech recognition models (such as SenseVoice and Paraformer), audio is transcribed into Chinese text;

[0039] 4. Large-scale model translation: Use large language models (such as Qwen3, DeepSeek, etc.) or NMT models to translate Chinese into English in real time;

[0040] 5. TTS speech synthesis: Use a speech synthesis system (such as GPT-SoVITS, CosyVoice) to replicate the voice of the broadcaster, replicating the broadcaster's voice quality, timbre, speaking rhythm and other characteristics, and then convert the English text into speech;

[0041] 6. Audio and video recombination: Repackage the original video track and the newly synthesized English audio track into a new RTMP stream;

[0042] 7. External Streaming: The synthesis results are pushed to overseas live streaming platforms through the streaming module.

[0043] The audio and video quality and synchronization are significantly improved. Existing solutions often suffer from audio-visual asynchrony and frame loss during the audio-video synthesis stage. This disclosure ensures high synchronization between English speech and the original video footage through unified timestamp management and a dynamic audio alignment strategy, while also guaranteeing the integrity of audio quality and frame rate after synthesis.

[0044] Supporting multilingual expansion and cross-platform streaming, compared to one-way translation from Chinese to English, this invention can be flexibly extended to multilingual scenarios such as Chinese-Japanese and Chinese-Korean by supporting multi-model switching and hot word injection mechanisms, and supports simultaneous streaming to multiple overseas platforms.

[0045] Easy for users to deploy and reuse, traditional solutions require the integration of multiple complex toolchains, while this disclosure adopts a modular architecture (containerable and supports API calls), with low deployment costs, strong maintainability, and good scalability, making it suitable for various users such as live streaming platforms, broadcast control centers, and conference organizers.

[0046] This disclosure provides a device for translating live video content. This device is used to execute the live video content translation method described in the above embodiments, such as... Figure 3 As shown, the device is applied to a terminal, which includes a speech recognition module, a translation module, and a text-to-speech module. The device includes: an acquisition unit 301, configured to acquire a live video stream and split the live video stream into audio and video tracks to obtain audio and video; a first adoption unit 302, configured to use the speech recognition module to recognize the audio and transcribe the audio into text in a first language; a second adoption unit 303, configured to use the translation module to translate the text in the first language into text in a second language; a third adoption unit 304, configured to use the text-to-speech module to convert the text in the second language into speech; and a synthesis unit 305, configured to synthesize the speech and the video to obtain a target video.

[0047] The video live streaming content translation device and the video live streaming content translation method provided in the above embodiments of this disclosure are based on the same inventive concept and have the same beneficial effects as the methods adopted, run or implemented by the applications stored therein.

[0048] This disclosure also provides an electronic device corresponding to the video live streaming content translation method provided in the foregoing embodiments, for executing the aforementioned video live streaming content translation method. This disclosure is not limiting.

[0049] Please refer to Figure 4 This illustrates a schematic diagram of an electronic device provided by some embodiments of the present disclosure. For example... Figure 4As shown, the electronic device 40 includes: a processor 400, a memory 401, a bus 402, and a communication interface 403. The processor 400, the communication interface 403, and the memory 401 are connected via the bus 402. The memory 401 stores a computer program that can run on the processor 400. When the processor 400 runs the computer program, it executes the method provided in any of the foregoing embodiments of this disclosure.

[0050] The memory 401 may include high-speed random access memory (RAM) or non-volatile memory, such as at least one disk storage device. Communication between this system network element and at least one other network element is achieved through at least one communication interface 403 (which can be wired or wireless), such as the Internet, wide area network, local area network, or metropolitan area network.

[0051] Bus 402 can be an ISA bus, PCI bus, or EISA bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. Memory 401 is used to store programs. After receiving an execution instruction, processor 400 executes the program. The video live streaming content translation method disclosed in any of the foregoing embodiments of this disclosure can be applied to processor 400, or implemented by processor 400.

[0052] The processor 400 may be an integrated circuit chip with signal processing capabilities. In implementation, each step of the above method can be completed by the integrated logic circuitry in the hardware of the processor 400 or by instructions in software form. The processor 400 may be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it may also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), an off-the-shelf programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this disclosure. The general-purpose processor may be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of this disclosure can be directly embodied in the execution of a hardware decoding processor, or executed by a combination of hardware and software modules in the decoding processor. The software modules may reside in random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other mature storage media in the art. The storage medium is located in memory 401. The processor 400 reads the information in memory 401 and, in conjunction with its hardware, completes the steps of the above method.

[0053] The electronic device provided in this disclosure and the video live streaming content translation method provided in this disclosure are based on the same inventive concept and have the same beneficial effects as the methods they employ, operate, or implement.

[0054] This disclosure also provides a computer-readable storage medium corresponding to the video live streaming content translation method provided in the foregoing embodiments. Please refer to [link / reference]. Figure 5 The computer-readable storage medium shown is an optical disc 50, on which a computer program (i.e., a program product) is stored. When the computer program is run by a processor, it executes the method for translating live video content provided in any of the foregoing embodiments.

[0055] It should be noted that examples of the computer-readable storage medium may also include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other optical and magnetic storage media, which will not be elaborated here.

[0056] The computer-readable storage medium provided in the above embodiments of this disclosure and the video live streaming content translation method provided in the embodiments of this disclosure are based on the same inventive concept and have the same beneficial effects as the methods adopted, run or implemented by the applications stored therein.

[0057] It should be noted that:

[0058] In the foregoing text, the terms "comprising," "including," or any other variations thereof are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element. Furthermore, it should be noted that the scope of the methods and apparatuses in this disclosure is not limited to performing functions in the order shown or discussed, but may also include performing functions substantially simultaneously or in the reverse order, depending on the functions involved. For example, the described methods may be performed in a different order than described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.

[0059] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this disclosure, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of this disclosure.

[0060] The embodiments of this disclosure have been described above with reference to the accompanying drawings. These are merely specific implementations of this disclosure, but this disclosure is not limited to the specific implementations described above. The specific implementations described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this disclosure without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this disclosure.

Claims

1. A method for translating live video content, characterized in that, Applied to a terminal, the terminal including a speech recognition module, a translation module, and a text-to-speech module, the method includes: Acquire the live video stream, and split the live video stream into audio and video tracks to obtain audio and video; The audio is recognized by a speech recognition module to transcribe the audio into text, wherein the text is in a first language. The translation module is used to translate text in the first language into text in the second language. The text-to-speech module is used to convert the text in the second language into speech. The audio and video are combined to obtain the target video.

2. The method according to claim 1, characterized in that, The step of using the text-to-speech module to convert text in the second language into speech includes: The voice of the target live streamer is replicated to obtain the voice features, which include sound quality information, timbre information, and rhythm information. The text in the second language is converted into speech with the aforementioned sound features.

3. The method according to claim 1, characterized in that, The acquisition of the live video stream includes: pulling the live video stream from the first live streaming platform; The step of synthesizing the speech and the video to obtain the target video includes: The audio track of the speech and the video track of the video are mixed and packaged according to timestamps to obtain the target video to be pushed. The target video to be pushed is used to push to the second live streaming platform so that the live streaming terminal corresponding to the second live streaming platform can play the target video of the speech in the second language.

4. The method according to claim 3, characterized in that, The terminal also includes a streaming module; The method further includes: The push module is used to push the target video to the second live streaming platform.

5. The method according to claim 1, characterized in that, The step of using the translation module to translate text in the first language into text in the second language includes: The translation model in the translation module is invoked to perform real-time text translation of the text in the first language, thereby obtaining the text in the second language.

6. The method according to claim 5, characterized in that, The translation model is multiple, and the method further includes: Obtain real-time live streaming demand information, and select the translation model to be invoked from multiple translation models based on the real-time live streaming demand information.

7. The method according to claim 1, characterized in that, The method further includes: Get the hot keywords corresponding to various live streaming scenarios; The speech and text of the hot words corresponding to the various live streaming scene information are input into the speech recognition model in the speech recognition module; The step of using a speech recognition module to recognize the audio and transcribe it into text includes: Obtain the live streaming scene information of the live video stream; The live stream scene information and the audio are input into the speech recognition model to obtain the text output by the speech recognition model.

8. A translation device for live video content, characterized in that, Applied to a terminal, the terminal includes a speech recognition module, a translation module, and a text-to-speech module, and the device includes: The acquisition unit is configured to acquire a live video stream, and to split the live video stream into audio and video tracks to obtain audio and video. The first employing unit is configured to use a speech recognition module to recognize the audio in order to transcribe the audio into text, wherein the text is text in a first language; The second employing unit is configured to employ the translation module to translate text in the first language into text in the second language; The third unit is configured to use the text-to-language module to convert the text in the second language into speech; The synthesis unit is configured to synthesize the speech and the video to obtain a target video.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, The processor executes the computer program to implement the method as described in any one of claims 1-7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, The program is executed by a processor to implement the method as described in any one of claims 1-7.

Citation Information

Cited By

  • Live broadcast multi-language interpretation method, device and system and storage medium

    CN121567890A