Multi-mode AI-driven short video automatic translation and speech synthesis system
Through the multimodal AI-driven short video automatic translation and speech synthesis system, the problems of violation detection, format processing, voice separation and speech synthesis in the existing technology are solved, and efficient and natural cross-language short video translation and speech synthesis are achieved, improving the user experience.
Patent Information
- Application Number
- CN202511107005.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-08
- Publication Date
- 2025-10-28
AI Technical Summary
Existing short video processing technologies have problems such as low efficiency in detecting illegal content, complex and time-consuming audio and video format processing, difficulty in separating human voices from backgrounds, unnatural speech translation, and poor accuracy in lip image and audio alignment. These problems lead to a fragmented overall process and make it difficult to meet the needs of real-time, high-quality translation and synthesis.
A multimodal AI-driven short video automatic translation and speech synthesis system uses AI modules such as convolutional neural networks, FFmpeg, non-negative matrix factorization, neural machine translation models, and adversarial networks to achieve violation detection, format transcoding, audio and video separation, speech synthesis, and lip synchronization, building a full-process automated processing.
It achieves high efficiency, authenticity and automation in short video translation and speech synthesis, improves processing efficiency and naturalness, ensures synchronization between speech and picture, and enhances users' viewing experience and immersion.
Smart Images

Figure CN120856930A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of audio and video technology, specifically relating to a multimodal AI-driven short video automatic translation and speech synthesis system. Background Technology
[0002] With the rapid development of economic globalization and internet technology, short videos have become a key medium for cross-cultural communication. Currently, AI technology is widely used in the field of audio and video processing, such as convolutional neural networks for content detection, format compatibility engines for audio and video transcoding, and non-negative matrix factorization for audio separation.
[0003] Existing technologies have many limitations in the field of short video processing; the detection of illegal content often relies on manual screening or simple rule filtering, which is inefficient and prone to omissions; in terms of audio and video format processing, the transcoding process is complex and time-consuming, easily destroying the original quality, and it is difficult to accurately separate human voice from the background; traditional translation is mostly text document processing, which cannot be directly applied to human voice signals, and does not consider the matching of speech rhythm and intonation, resulting in stiff and unnatural synthesized audio; when aligning lip images with audio, it mostly relies on manual annotation or simple algorithms, which has poor accuracy and makes it difficult to achieve natural synchronization between lip movements and sound; the various stages of the overall process are fragmented, failing to form an efficient and collaborative multimodal processing mode, making it difficult to meet the needs of real-time, high-quality translation synthesis.
[0004] To address the aforementioned issues, this invention proposes a multimodal AI-driven automatic translation and speech synthesis system for short videos. Summary of the Invention
[0005] To address the shortcomings of existing technologies, this invention provides a power plant peak-shaving thermal storage system based on a high-pressure electrode boiler, solving the problems of poor integration and low efficiency in existing technologies for automatic translation and speech synthesis of short videos.
[0006] The objective of this invention can be achieved through the following technical solutions:
[0007] A multimodal AI-driven short video automatic translation and speech synthesis system, comprising the following components:
[0008] The video acquisition end receives audio and video data and uses the first AI module to detect illegal content.
[0009] If any content violates regulations, it will be removed; otherwise, it will be retained.
[0010] The initial processing unit extracts and retains the audio and video data, and the second AI module identifies the container format and encoding parameters. It then performs real-time transcoding on non-standard audio and video data and outputs standard audio and video data.
[0011] Perform a separation operation on standard audio and video data to obtain standard audio streams and standard video streams;
[0012] A third AI module is used to separate the standard audio stream, resulting in a clean human voice signal and independent background audio.
[0013] The video translation module extracts pure human voice signals and uses the fourth AI module to generate target language text associated with the pure human voice signals.
[0014] The prosodic features of the pure human voice signal are extracted in time sequence, and a preliminary audio stream is generated by combining it with the human voice module. The duration of the continuous audio segments is compared with the pure human voice signal and the speed is adjusted to obtain a secondary audio stream. The independent background audio and the secondary audio stream are then fitted in time sequence to generate the final audio stream.
[0015] The video synthesis end divides the standard video stream by frame number to obtain a standard video frame sequence, and extracts the lip image frame sequence associated with the standard video frame sequence;
[0016] The secondary audio stream is then divided according to the frame order of the standard video frame sequence to obtain the secondary audio frame sequence. The fifth AI module generates a lip vertex displacement coordinate sequence that matches the secondary audio frame sequence frame by frame, and renders the corresponding frames of the lip image frame sequence.
[0017] The rendered lip image frame sequence is fitted to a standard video frame sequence, and the final audio frame sequence is time-aligned to synthesize an audio-video synchronized sequence.
[0018] At the video output end, the determined audio and video synchronization sequence is compressed and encapsulated using predefined encoding rules to output the final audio and video data.
[0019] As a further aspect of the present invention, the first AI module in this system includes a pre-trained convolutional neural network and TensorFlow;
[0020] The second AI module includes FFmpeg and a format compatibility engine;
[0021] The third AI module includes Non-negative Matrix Factorization (NMF) and Spleeter;
[0022] The fourth AI module includes a neural machine translation model with an encoder-decoder architecture and Sipmpleon-listening;
[0023] The fifth AI module includes an adversarial network-driven 3D facial deformation model and LipSync.
[0024] As a further aspect of the present invention, in the video acquisition terminal, the audio and video data submitted by the user to the system includes local files, network URLs, and real-time streaming media;
[0025] The audio and video data submitted by users to this system are cached in the cache pool and included in the violation detection queue in the order of caching.
[0026] When the first AI module is idle, it extracts the first audio / video data from the violation detection queue for violation detection.
[0027] If the violation detection passes, the first audio / video data will be added to the storage pool;
[0028] If the violation detection fails, a clearing operation is performed, removing the first audio / video data from the violation detection queue and cache pool.
[0029] As a further aspect of the present invention, the specific method for outputting standard audio and video data in the preliminary processing terminal is as follows:
[0030] Extract any audio or video data from the storage pool, denoted as YP;
[0031] The second AI module is used to identify the container format CF associated with the audio and video data YP. YP and encoding parameter EP YP It is then compared with the container format set CF and the encoding parameter set EP preset by the operator;
[0032] in, And CF i Corresponding EP i j is the value preset by the operator, i is the counting index, 1≤i≤j;
[0033] If the container format is CF YP Belongs to the container format set CF and has encoding parameter EP YP It belongs to the encoding parameter set EP, and CF YP Corresponding EP YP If the comparison is successful, the audio / video data YP will be marked as the standard audio / video data SYP.
[0034] Conversely, if the comparison fails, the container format CF associated with the audio / video data YP will be changed. YP and encoding parameter EP YP Real-time transcoding to any pair of CFs from the container format set CF and the encoding parameter set EP. i With EP i The real-time transcoded audio and video data YP is recorded as the standard audio and video data SYP.
[0035] As a further aspect of the present invention, the specific method for performing the separation operation on the standard audio and video data in the preliminary processing terminal is as follows:
[0036] The standard audio and video data SYP is demultiplexed to separate the standard audio stream and standard video stream in the standard audio and video data SYP, which are denoted as SAS and SVS respectively. The duration of the standard audio stream SAS is equal to the duration of the standard video stream SVS.
[0037] Based on the established standard audio stream SAS, a separation operation is performed by the third AI module to determine the human voice track and the background audio track to obtain a clean human voice signal and an independent background audio, denoted as PVS and IBA, respectively.
[0038] As a further aspect of the present invention, the specific method for generating the final audio stream in the video translation terminal is as follows:
[0039] Extract the pure human voice signal PVS, and use the fourth AI module to perform text recognition on the pure human voice signal PVS to generate first language text;
[0040] The first language text is translated into the user-defined target language to obtain the target language text, denoted as TLT. The target language is derived from the operator's preset target language library.
[0041] The duration and total number of moments of the pure human voice signal PVS are determined. Based on the fourth AI module, the prosodic features corresponding to the moments in the pure human voice signal PVS are extracted to obtain the prosodic feature sequence.
[0042] Then, the prosodic feature sequence and the target language text TLT are input into the human voice module pre-built by the operator to generate a preliminary audio stream associated with the target language text TLT, denoted as FAS;
[0043] Determine the duration of all consecutive audio segments in the initial audio stream FAS, and record them in chronological order as the audio segment duration sequence CAS. FAS-1 CAS FAS-2 ,...,CAS FAS-m , where m is the total number of consecutive audio segments in the initial audio stream PAS;
[0044] Similarly, the duration sequence of consecutive audio segments associated with the pure human voice signal PVS is determined by CAS. PVS-1 CAS PVS-2 ,...,CAS PVS-m ;
[0045] Extracting the duration sequence of consecutive audio segments using CAS FAS-1 CAS FAS-2 ,...,CAS FAS-m CAS duration of any audio segment FAS-n CAS with audio continuous segment duration sequence PVS-1 CASPVS-2 ,...,CAS PVS-m CAS duration of any audio segment PVS-n Perform a comparison, where n is the counting index, with a value ranging from 1 to m;
[0046] If the duration of a continuous audio segment is CAS FAS-n CAS of audio continuous segment duration PVS-n If the error is greater than the error percentage threshold of 5%, then the duration of the audio segment CAS is... FAS-n Perform speed adjustment, and adjust the duration of continuous audio segments using CAS. FAS-n CAS of audio continuous segment duration PVS-n The error is controlled within 5% of the error percentage threshold;
[0047] Similarly, for the duration sequence of consecutive audio segments, CAS FAS-1 CAS FAS-2 ,...,CAS FAS-m The duration of all consecutive audio segments is adjusted and summed so that the summed audio duration is equal to the duration of the pure human voice signal, and this is denoted as the secondary audio stream EAS.
[0048] The determined secondary audio stream EAS is then fitted with the independent background audio IBA to generate the final audio stream, denoted as LAS.
[0049] As a further aspect of the present invention, the specific method for extracting the lip image frame sequence associated with the standard video frame sequence in the video synthesis end is as follows:
[0050] Extract standard video stream SVS;
[0051] Determine the total number of frames in the standard video stream SVS, denoted as o;
[0052] Divide the standard video stream into frames based on the total number of frames o, resulting in the standard video frame sequence SVS1, SVS2, ..., SVS o ;
[0053] Extract the standard video frame sequence SVS1, SVS2, ..., SVS o Any standard video frame in the video is denoted as SVS. u Where u is the counting index, 1≤u≤o;
[0054] Extracting Standard Video Frames (SVS) using facial recognition technology u The lip image is used to obtain the lip image frame, denoted as LIF. u ;
[0055] And so on, to obtain the standard video frame sequence SVS1, SVS2, ..., SVS o The lip image frames associated with all standard video frames are summarized to form the lip image frame sequence LIF1, LIF2, ..., LIF o .
[0056] As a further aspect of the present invention, the specific method for rendering the corresponding frames of the lip image frame sequence in the video synthesis terminal is as follows:
[0057] S81. Extract the determined secondary audio stream EAS and the lip image frame sequence LIF1, LIF2, ..., LIF o ;
[0058] According to the lip image frame sequence LIF1, LIF2, ..., LIF o The secondary audio stream EAS is divided into frames according to the frame order, resulting in the secondary audio frame sequence EAS1, EAS2, ..., EAS. o ;
[0059] S82. Determine the lip image frame sequence LIF1, LIF2, ..., LIF o Any lip image frame in LIF u and the corresponding secondary audio frame sequences EAS1, EAS2, ..., EAS o Secondary audio frames EAS u ;
[0060] S83, Transfer the lip image frame to LIF u and secondary audio frames EAS u Input the fifth AI module to generate EAS audio frames. u Matching lip vertex displacement coordinates LVD u ;
[0061] S84, using LVD (Lip Vertex Displacement Coordinates) u LIF image frames of the lips u Rendering is performed to make the lip image frame LIF u Conforms to secondary audio frames EAS u The shape of the mouth;
[0062] S85. Repeat steps S82 to S84 to render the lip image frame sequence LIF1, LIF2, ..., LIF o From all the lip image frames, obtain the rendered sequence EAS1, EAS2, ..., EAS2, ... o The matched sequence of lip image frames is denoted as LIF′1, LIF′2, ..., LIF′. o .
[0063] As a further aspect of the present invention, the specific method for synthesizing the audio-video synchronization sequence in the video synthesis terminal is as follows:
[0064] Extract the lip image frame sequence LIF'1, LIF'2, ..., LIF' o And fit it to the standard video frame sequence SVS1, SVS2, ..., SVS o In this process, the fitted standard video frame sequence SVS'1,SVS'2,...,SVS' is obtained. o ;
[0065] Extract the final audio stream LAS, according to the standard video frame sequence SVS'1, SVS'2, ..., SVS' o The final audio stream LAS is divided into frames according to the frame order, resulting in the final audio frame sequence LAS1, LAS2, ..., LAS o ;
[0066] The final audio frame sequence LAS1, LAS2, ..., LAS o With the standard video frame sequence SVS'1,SVS'2,...,SVS' o Alignment and fitting are performed frame by frame to obtain the audio and video synchronization sequences LAS1-SVS'1, LAS2-SVS'2, ..., LAS associated with the audio and video data SYP. o -SVS′ o .
[0067] As a further aspect of the present invention, the specific method for outputting the final audio and video data at the video output terminal is as follows:
[0068] Extract the audio and video synchronization sequences LAS1-SVS'1, LAS2-SVS'2, ..., LAS associated with the audio and video data SYP. o -SVS' o The system uses predefined encoding rules to compress and encapsulate the data, outputting the final audio and video data.
[0069] The beneficial effects of this invention are:
[0070] (1) This invention achieves high efficiency, realism and automation of short video translation and speech synthesis through the full-process integration of multimodal AI technology. Its advantages are first reflected in the full-link automated processing: from violation detection, format transcoding to audio and video separation, no manual intervention is required, which greatly improves the processing efficiency. The key breakthrough is the deep optimization of speech processing - by separating pure human voice and extracting prosodic features, combined with the target language audio generated by the human voice module, after time alignment and speed adjustment, it is fitted with the background sound, which not only preserves the original emotional rhythm, but also ensures the naturalness of the speech. The most innovative technology is lip-sync technology: based on 3D lip vertex displacement rendering driven by audio frames, the translated lip shape is accurately matched with the synthesized speech, which significantly improves the visual realism.
[0071] (2) This invention effectively improves processing efficiency and compliance through intelligent process management: The system supports diverse audio and video input sources and uses a cache pool and sequential queue mechanism for task scheduling. Combined with the automatic violation detection triggered when the AI module is idle, it ensures the timely interception and removal of illegal content, reduces the burden of manual review and ensures platform security. In the format processing stage, the system automatically compares the format and encoding parameters through a preset standard set and realizes real-time transcoding. This not only supports a wide range of media formats but also ensures the uniformity and compatibility of the input data in the subsequent processing flow, significantly improving the robustness and smoothness of the system.
[0072] (3) This invention significantly improves the naturalness and audio-visual synchronization of cross-language speech synthesis through refined prosodic transplantation and duration calibration technology. The system uses the fourth AI module to accurately extract the prosodic feature sequence of the original human voice and combines it with the translated target language text to generate a preliminary audio stream, ensuring that the emotional expression and rhythm of the translated speech are close to the original voice. More importantly, by comparing the duration of the new and old audio streams segment by segment and performing intelligent speed adjustment, the system strictly controls the error within the threshold, ultimately ensuring that the total duration of the secondary audio stream is completely consistent with the original voice. This design effectively solves the common problem of lip-syncing and voice mismatch in cross-language dubbing. Finally, the precise duration of the translated human voice (secondary audio stream) is fitted with the separated original background audio track to generate the final audio stream. While achieving language conversion, the auditory atmosphere and background sound effects of the original video are perfectly preserved, significantly improving the viewing smoothness and immersion of multilingual short videos.
[0073] (4) This invention significantly improves the visual realism and overall immersiveness of cross-language short videos through highly intelligent frame-level rendering and dual synchronization mechanisms. Its core advantage lies in achieving pixel-level lip-sync and precise audio-video frame alignment: The system uses the fifth AI module to divide the secondary audio stream into a frame sequence and generates corresponding lip vertex displacement coordinates frame by frame. Through dynamic rendering, the lip shape of each frame is perfectly matched with the mouth shape of the translated speech, which completely solves the common problem of "voice-image asynchrony" in traditional translation videos. Furthermore, the system seamlessly fits the rendered lip-shape frame sequence to the original video frame, and at the same time divides the final audio stream according to the video frame rate and performs frame-level alignment to construct a strictly synchronized audio-video sequence. This dual synchronization guarantee (lip-shape-speech synchronization, audio-video frame synchronization) not only makes the lip movement after translation highly natural, but also ensures the fit between background sound effects, translated voice and screen movements, greatly improving the viewing experience for multilingual users. Attached Figure Description
[0074] The invention will now be further described with reference to the accompanying drawings.
[0075] Figure 1 This is a schematic diagram of the system described in this invention;
[0076] Figure 2 This is a flowchart illustrating the method described in Embodiment 2 of the present invention;
[0077] Figure 3 This is a flowchart illustrating the method described in Embodiment 3 of the present invention;
[0078] Figure 4 This is a flowchart illustrating the method described in Embodiment 4 of the present invention. Detailed Implementation
[0079] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0080] Example 1
[0081] Multimodal AI-driven short video automatic translation and speech synthesis system, such as Figure 1 As shown, the specific steps include the following:
[0082] This system is primarily designed to serve users who require video and audio translation. It analyzes and processes the audio and video data uploaded by users to the system and outputs the desired audio and video data (in other languages).
[0083] Because the system needs to adjust the lip movements of people in real time during operation, it is necessary to ensure that the audio and video data uploaded by users contain a recognizable human face.
[0084] This system is interconnected with the user's operating terminal to receive audio and video data transmitted from the user's operating terminal in real time. The operating terminal includes mobile phones, computers, and multimedia devices used by the user for communication.
[0085] This system mainly includes the following modules:
[0086] The video acquisition end is interconnected with the user's operating terminal and receives audio and video data (including local files, network URLs and real-time streaming media) submitted by the user to the system in real time. The first AI module detects whether there is any illegal content in the audio and video data. Specifically, the video acquisition end, as the front end of the system, establishes a real-time interconnection channel with the user's operating terminal through a standardized network interface with the data receiving end.
[0087] It needs to be explained that the first AI module includes a pre-trained convolutional neural network and TensorFlow (based on convolutional neural networks and TensorFlow, a violation detection model for audio and video data is built to realize violation detection of audio and video data content).
[0088] After receiving audio and video data submitted by the user through the terminal, the local end first stores the audio and video data in the cache pool of the local system and checks whether there is an empty space in the violation detection queue. If there is an empty space, the audio and video data is added to the end of the violation detection queue to wait for violation detection. If there is no empty space, it waits for an empty space.
[0089] Then, the first AI module pre-built by the operator is invoked, and it is checked whether the first AI module is in an idle state. If the first AI module is not in an idle state, then continue to wait.
[0090] If the first AI module is idle, it will take the first audio and video data from the violation detection queue and perform violation detection. If violation content is detected in the audio and video data submitted by the user, it will perform a clearing operation on the audio and video data (audio and video data containing violation content cannot be subjected to subsequent operations).
[0091] The clearing operation involves retrieving the audio / video data from the violation detection queue and the cache pool, and simultaneously clearing it.
[0092] If the violation detection passes, the corresponding audio and video data will be included in the storage pool, awaiting processing by the initial processing unit.
[0093] The initial processing end primarily retrieves the audio and video data that has passed violation detection from the storage pool and performs initial processing. The initial processing mainly includes using the second AI module to identify the container format and encoding parameters of the audio and video data, and calibrating the container format and encoding parameters of the audio and video data to make them conform to the standard container format and standard encoding parameters specified by this system (the standard container format and standard encoding parameters are determined by the operators of this system and belong to the container format set and encoding parameter set, respectively).
[0094] It is necessary to explain here that the second AI module includes FFmpeg and a format compatibility engine. FFmpeg can parse the container format and encoding parameters of audio and video data. The format compatibility engine compares the parsed container format and encoding parameters with the standard container format and standard encoding parameters preset by the operator. If they match, no processing is done. If they do not match, standardization is performed (through transcoding, the audio and video data is transformed into standard audio and video data).
[0095] Next, standard audio and video data are extracted, and the video and audio parts are separated to obtain the standard audio stream and standard video stream associated with the standard audio and video data.
[0096] Next, the standard audio stream is separated again. Because a standard audio stream may contain human voice information and background music at the same time, if the separation operation is not performed, the background music will interfere with the human voice information. Therefore, the third AI module is used to separate the background audio track and the human voice track to obtain a pure human voice signal and independent background audio.
[0097] It needs to be explained that the third AI module includes Non-negative Matrix Factorization (NMF) and Spleeter. The NMF method can effectively matrix-divide the standard audio stream, decompose the audio signal into different audio track components, and then input them into Spleeter for separation, outputting two separated audio tracks: a clean vocal signal and independent background audio.
[0098] The video translation terminal primarily processes the pure human voice signal, utilizes the fourth AI module to perform language recognition on the pure human voice signal, and converts it into text to obtain the target language text;
[0099] The fourth AI module includes a neural machine translation model with an encoder-decoder architecture and Sipmpleon-listening, used for language recognition and text generation of clean human voice signals.
[0100] This process also includes determining the prosodic features of the pure human voice signal according to the time sequence. The prosodic features include fundamental frequency, pitch, phoneme duration, and pause distribution, which are used for subsequent speech synthesis operations.
[0101] Then, the prosodic features—fundamental frequency, pitch, phoneme duration, and pause distribution—are input into the human voice module pre-built by the operator. Combined with the target language text and the user-preset language to be translated, an audio segment containing the prosodic features of the person and in the user-preset language to be translated is generated and recorded as the initial audio stream.
[0102] Next, the duration of each continuous audio segment in the initial audio stream is compared with the duration of each continuous audio segment in the clean human voice signal. The duration of each continuous audio segment in the initial audio stream is adjusted by speed adjustment to reduce the segment error between each continuous audio segment in the initial audio stream and each continuous audio segment in the clean human voice signal (the reason is that due to language differences, the same content is expressed but the duration of the audio segments is different, so in order to avoid a sense of disjointedness after final video synthesis, the duration of the continuous audio segments needs to be calibrated). The initial audio stream after reducing the segment error is denoted as: the secondary audio stream.
[0103] Then, the secondary audio stream is fitted with the independent background audio in chronological order (human voice information and background music are merged again) to generate the final audio stream.
[0104] The video synthesis end primarily processes standard video streams. First, it divides the standard video stream by frame count to obtain a standard video frame sequence. Each frame represents a moment in time, meaning one moment corresponds to one standard video frame.
[0105] Face recognition technology is used to process each standard video frame in the standard video frame sequence, extract the face portion, and then further separate the lip image frames (to facilitate subsequent mouth shape / lip shape adjustment, with the goal of adjusting the mouth shape to correspond to the lip shape / lip shape in the translated secondary audio stream). All lip image frames are summarized in chronological order to obtain: lip image frame sequence.
[0106] Next, the secondary audio stream is divided into secondary audio stream frame sequences according to the order of the standard video frame sequence. The secondary audio stream frame sequences correspond one-to-one with the standard video frame sequences. The secondary audio stream frame sequences and the lip image frame sequences are input into the fifth AI module, and the lip vertex displacement coordinate sequence is matched frame by frame with the secondary audio frame sequence (lip shape correction processing is performed on each lip image frame to make the lip shape in the lip image frame correspond to the mouth shape / lip shape in the translated secondary audio stream).
[0107] The fifth AI module includes an adversarial network-driven 3D facial deformation model and LipSync. The LipSync component extracts key features related to pronunciation from a sequence of secondary audio frames, and generates a set of target lip shape parameters for each secondary audio frame (the target lip shape parameters represent the state of the mouth shape / mouth shape under the current audio).
[0108] Next, all target lip shape parameters are input into the adversarial network-driven 3D facial deformation model. Based on the corresponding lip image frame, the lip vertex displacement coordinates are generated, and the corresponding lip image frame is rendered (so that the translated speech visually perfectly matches the speaker's lip movements, thereby enhancing the realism and immersion of the viewing experience).
[0109] Extract the determined sequence of lip vertex displacement coordinates corresponding to the lip image frame sequence, and render the corresponding lip image frames in the lip image frame sequence so that the mouth shape of the lip image frames in the lip image frame sequence corresponds to each secondary audio frame in the secondary audio frame sequence.
[0110] Next, the rendered lip image frame sequence is fitted to the standard video frame sequence (using the method of corresponding lip image frames with standard video frames). Then, the final audio frame sequence is further fitted to the result of fitting the lip image frame sequence to the standard video frame sequence through the corresponding frame method, and the audio and video synchronization sequence is output (sound and video are synchronized).
[0111] The purpose of this embodiment is to achieve automatic conversion of audio and video data between different languages, providing users with the following services: After a user uploads audio and video data, the system first receives it through the video acquisition end and the first AI module detects illegal content, clearing the illegal audio and video data and storing the compliant audio and video data in the storage pool; the preliminary processing end uses the second AI module to calibrate the audio and video format parameters, separating the standard audio stream and the standard video stream; the third AI module separates the human voice and background music in the audio, obtaining a clean human voice signal and independent background audio; the video translation end uses the fourth AI module to perform speech recognition, text generation, and speech synthesis, generating a preliminary audio stream and calibrating the segment errors to obtain a secondary audio stream, which is then fitted with the independent background audio; the video synthesis end processes the video stream, extracts the lip image frame sequence and calibrates it with the audio, uses the fifth AI module to generate target lip shape parameters, renders the lip image frames and synchronously fits them to the video frame sequence, and finally the video output end compresses and encapsulates the audio and video synchronized sequence to meet the user's needs for automatic translation and synthesis of short videos in multiple scenarios.
[0112] Example 2
[0113] This embodiment discloses a method for separating audio and video data based on Embodiment 1, such as... Figure 2 As shown, the specific steps include the following:
[0114] As shown in Example 1, before performing the separation operation on audio and video data, it is necessary to verify the container format and encoding parameters of the audio and video data uploaded by the user. This involves the operator's preset container format set CF and encoding parameter set EP. Encoding parameter set And CF i Corresponding EP i j is a value preset by the operator, and i is a counting index, ranging from 1 to j;
[0115] The specific steps for separating audio and video data are as follows:
[0116] First, extract any audio or video data from the storage pool. To facilitate subsequent differentiation, label this audio or video data as: YP.
[0117] Next, the second AI module identifies the container format and encoding parameters associated with the audio and video data YP, and labels them as: CF YP and EP YP ;
[0118] Extract the operator's preset container format set CF and encoding parameter set EP, and associate the audio / video data YP with the container format CF. YP and encoding parameter EP YP Perform verification;
[0119] If the container format is CF YP Belongs to the container format set CF and has encoding parameter EP YP If the data belongs to the encoding parameter set EP and they correspond to each other, it means that the audio and video data conforms to the container format and encoding parameters preset by the operator, and is regarded as: standard audio and video data, marked as: SYP;
[0120] If the container format set CF YP and the encoding parameter set EP YP If any one or more of the container formats or encoding parameters do not conform to the operator's preset format, or if they do not correspond to each other, then the audio / video data YP needs to be transcoded in real time to change the container format CF associated with the audio / video data YP. YP and encoding parameter EP YP Real-time transcoding to any pair of container formats CF from the container format set CF and the encoding parameter set EP. i and encoding parameter EP i The real-time transcoded audio and video data YP is then denoted as standard audio and video data SYP.
[0121] At this point, the standard audio / video data SYP is obtained. Subsequent separation operations will then be performed, linking the standard audio / video data SYP to the container format CF. i and encoding parameter EP i Demultiplexing is performed (based on the inconsistency in container formats and encoding parameters between audio and video), further separating the independently encoded standard audio stream and standard video stream from the standard audio and video data SYP, denoted as SAS and SVS respectively;
[0122] After obtaining the standard audio stream SAS, the standard audio stream SAS is then input into the third AI module pre-built by the operator for another separation operation. The purpose of this step is to identify and separate the human voice track and the background audio track to obtain a pure human voice signal and an independent background audio, which are respectively denoted as PVS and IBA.
[0123] This embodiment aims to lay a data foundation for subsequent targeted audio processing and synthesis, improve the accuracy and professionalism of audio and video processing, ensure efficient connection between each stage, meet the needs of fine separation and optimization of audio and video data, thereby improving the overall quality and effect of audio and video data processing, and enhancing the expressiveness and user experience of subsequent audio and video data.
[0124] Example 3
[0125] This embodiment further discloses a method for generating the final audio stream based on embodiment 2, such as... Figure 3 As shown, the specific steps include the following:
[0126] According to the content in Example 2, a pure human voice signal PVS can be obtained. The pure human voice signal PVS is an audio segment containing only human voice content. First, the pure human voice signal PVS is subjected to text recognition (speech recognition, recognizing the text) by the fourth AI module, and the recognized text is recorded and marked as: first language text (at this time, it is the original language text of the pure human voice signal).
[0127] Next, the first language text is translated into the target language set by the user (as required by the user) (this step can be achieved using existing translation software, where the target language also comes from the target language library of these translation software). The translated first language text is marked as: target language text, and denoted as TLT.
[0128] Next, the duration and total number of moments of the pure human voice signal PVS are obtained (the total number of moments corresponds to the number of frames in the standard audio and video data SYP, with one frame being one moment). The prosodic features corresponding to the moments are extracted from the pure human voice signal PVS by the fourth AI module to obtain the prosodic feature sequence.
[0129] The prosodic feature sequence and the target language text TLT are then input into the operator's pre-built voice module (the part that can be achieved by existing technology will not be elaborated in this solution). The voice module is used to simulate an audio segment that is similar to the voice (prosodic features) in the original standard audio and video data SYP and whose reading content is the target language text TLT. This segment is recorded as the preliminary audio stream and labeled as: FAS.
[0130] Next, the duration of all continuous audio segments in the initial audio stream FAS is determined. A segment of audio in the initial audio stream FAS where the human voice is interrupted due to breathing or other reasons is recorded as an audio continuous segment (for example, the interval between two breaths is considered an audio continuous segment, as two breaths define two interruptions in sound). The duration of this audio continuous segment is recorded as the audio continuous segment duration. All audio continuous segment durations are summarized in chronological order to obtain the audio continuous segment duration sequence, represented as: CAS FAS-1 CAS FAS-2 ,...,CAS FAS-m , where m is the total number of consecutive audio segments in the initial audio stream PAS.
[0131] Repeat the above steps to process the clean vocal signal PVS in the same way, and determine the duration sequence of the continuous audio segments associated with the clean vocal signal PVS, denoted as: CAS PVS-1 CAS PVS-2 ,...,CAS PVS-m Since the rhythmic features are the same, the number of interruptions in the human voice is also the same. Therefore, the total number of continuous audio segments in the pure human voice signal PVS is the same as the total number of continuous audio segments in the initial audio stream FAS, both being m.
[0132] CAS from the duration sequence of consecutive audio segments FAS-1 CAS FAS-2 ,...,CAS FAS-m Extract the duration of any consecutive audio segment and label it as: CAS FAS-n Where n is the counting index, and its value ranges from 1 to m;
[0133] Then from the duration sequence of consecutive audio segments CAS PVS-1 CAS PVS-2 ,...,CAS PVS-m Extract the duration of an audio segment that corresponds to the duration of the audio segment, and label it as: CAS PVS-n and CAS FAS-n With CAS PVS-n Perform a comparison operation;
[0134] Using |1-CASFAS-n / CAS PVS-n |*100%=α calculates the duration of the continuous audio segment (CAS) FAS-n CAS of audio continuous segment duration PVS-n The error α between them, if α = 0%, represents the duration of the continuous audio segment CAS. FAS-n CAS of audio continuous segment duration PVS-n There is no error between them, and no adjustment is needed;
[0135] If α < 95%, it indicates CAS FAS-n Compared to CAS PVS-n The error percentage exceeded the operator's preset threshold of 5%, requiring adjustment of the CAS. FAS-n Adjust the speed to enable CAS. FAS-n The duration compared to CAS PVS-n The duration is controlled within 5% of the error percentage threshold;
[0136] Where α < 95% and CAS FAS-n The duration is less than CAS PVS-n Then CAS needs to be modified. FAS-n Adjust the speed by reducing the magnification; conversely, if α < 95% and CAS... FAS-n The duration is longer than CAS PVS-n Then CAS needs to be modified. FAS-n Adjustments were made to increase the speed, so that CAS... FAS-n The duration compared to CAS PVS-n The duration is controlled within 5% of the error percentage threshold.
[0137] Repeat the above operation to perform CAS on the duration sequence of consecutive audio segments. FAS-1 CAS FAS-2 ,...,CAS FAS-m The duration of all continuous audio segments is sped up, and the sped-up sequence of the sped-up audio segments is summarized. The summed audio duration is then compared with the duration of the pure human voice signal. If they are the same, no further processing is needed. If they are different, the sped-up speed needs to be finely adjusted so that the summed audio duration of the sped-up audio segments equals the duration of the pure human voice signal. The summed audio duration of the sped-up audio segments is recorded as a secondary audio stream, denoted as EAS.
[0138] At this point, we have obtained the translated audio: the secondary audio stream EAS. Then, we fit the secondary audio stream EAS with the independent background audio IBA to obtain the complete audio, which is denoted as the final audio stream, and represented as LAS.
[0139] The core of this embodiment lies in performing speech recognition on a pure human voice signal to obtain a first language text, which is then translated into the user-defined target language text. Subsequently, combining the duration and prosodic features of the pure human voice signal, a preliminary audio stream is generated using a human voice module. Next, by comparing the duration of continuous audio segments in the preliminary audio stream with that in the pure human voice signal, the speed of the preliminary audio stream is adjusted to ensure that its duration is consistent with that of the pure human voice signal, resulting in a secondary audio stream. Finally, the secondary audio stream is fitted with independent background audio to generate a complete final audio stream.
[0140] The purpose of this embodiment is to achieve accurate translation and synthesis of audio and video data, ensure that the audio duration is consistent with the original pure human voice signal, and integrate background audio to improve audio and video synchronization and integrity, enhance the user's viewing experience, and meet the user's needs for multilingual audio and video content.
[0141] Example 4
[0142] This embodiment, based on Embodiments 1, 2, and 3, further discloses a method for extracting lip image frame sequences from standard video streams and combining them with secondary audio frame sequences to perform lip shape correction and rendering on the lip image frame sequences, such as... Figure 4 As shown, the specific steps include the following:
[0143] The standard video stream SVS, obtained by separating the standard audio and video data SYP, can be obtained from the content described in Example 2.
[0144] The standard video stream (SVS) is divided into frame numbers, and the total number of frames in the SVS is counted, denoted as o. The resulting o-frame standard video is recorded as a standard video frame sequence, represented as: SVS1, SVS2, ..., SVS o ;
[0145] From the determined standard video frame sequence SVS1, SVS2, ..., SVS o Extract any standard video frame and label it as SVS. u , where u is the counting index, with a value ranging from 1 to 0.
[0146] The following steps are for standard video frame SVS u Example processing is performed using the standard video frame sequence SVS1, SVS2, ..., SVS o The remaining standard video frames are processed according to the standard video frame SVS. u The same process is applied.
[0147] First, facial recognition technology is used to analyze standard video frames SVS. uPerform face recognition, identify the facial features, and extract the lip image portion to obtain a lip image frame, denoted as LIF. u ;
[0148] Similarly, repeat this step to obtain the standard video frame sequence SVS1, SVS2, ..., SVS o The associated sequence of lip image frames is represented as: LIF1, LIF2, ..., LIF o .
[0149] Then, extract the secondary audio stream EAS containing only human voice signals after translation processing. Then, sort the secondary audio stream EAS according to the lip image frame sequence LIF1, LIF2, ..., LIF o The temporal sequence is divided into frames (i.e., time-based division, with one frame representing one time point). The result of the division is denoted as a secondary audio frame sequence, represented as: EAS1, EAS2, ..., EAS o .
[0150] Then, from the determined secondary audio frame sequence EAS1, EAS2, ..., EAS o Extract a LIF image frame that matches the lip image frame u The secondary audio frame corresponding to the frame (time) is denoted as: EAS u .
[0151] LIF image frame of lips u and secondary audio frames EAS u Simultaneously input into the fifth AI module, and generate EAS with the secondary audio frame. u The matched lip vertex displacement coordinates are denoted as LVD. u ;
[0152] Then use the lip vertex displacement coordinate LVD u LIF image frames of the lips u The lips are adjusted and rendered to produce a higher LIF image frame. u Conforms to secondary audio frames EAS u The lip movements of a human voice.
[0153] Repeat the above steps for the lip image frame sequence LIF1, LIF2, ..., LIF o All lip image frames in the sequence are adjusted and rendered to obtain the rendered result that is consistent with the secondary audio frame sequence EAS1, EAS2, ..., EAS. o The sequence of lip image frames matched for the corresponding frame (time) is represented as: LIF'1, LIF'2, ..., LIF' o .
[0154] This embodiment first divides the standard video stream into frames to obtain a standard video frame sequence. Then, it uses face recognition technology to extract lip image frames from each standard video frame, forming a lip image frame sequence. Next, it divides the secondary audio stream into frames according to the temporal order of the lip image frame sequence to obtain a secondary audio frame sequence. The lip image frames and corresponding secondary audio frames are input into the fifth AI module to generate matching lip vertex displacement coordinates. Based on these coordinates, the lips in the lip image frames are adjusted and rendered so that the lip movements match the audio lip shapes. This process is repeated to process all lip image frames, resulting in a rendered lip image frame sequence that precisely matches the secondary audio frame sequence.
[0155] The purpose of this embodiment is to achieve accurate synchronization between the lip movements of the characters in the video and the translated audio content, thereby improving the quality of audio-visual synchronization and enhancing the realism and immersion of the viewing experience.
[0156] All data in the formulas described above are numerical calculations performed with dimensions removed. Furthermore, any content not described in detail in this specification is existing technology known to those skilled in the art.
[0157] The above description is merely an example and illustration of the present invention. Those skilled in the art can make various modifications or additions to the specific embodiments described, or use similar methods to replace them, as long as they do not deviate from the invention or exceed the scope defined in the claims, all of which should fall within the protection scope of the present invention.
[0158] It should be stated that all user data collected in this application was collected with the user's consent and authorization. Furthermore, the uses of user data are legal and compliant, and the use and processing of user data comply with the relevant laws, regulations, and standards of the relevant regions.
Claims
1. A multimodal AI-driven short video automatic translation and speech synthesis system, characterized in that, This system includes the following: The video acquisition end receives audio and video data and uses the first AI module to detect illegal content. If any content violates regulations, it will be removed; otherwise, it will be retained. The initial processing unit extracts and retains the audio and video data, and the second AI module identifies the container format and encoding parameters. It then performs real-time transcoding on non-standard audio and video data and outputs standard audio and video data. Perform a separation operation on standard audio and video data to obtain standard audio streams and standard video streams; A third AI module is used to separate the standard audio stream, resulting in a clean human voice signal and independent background audio. The video translation module extracts pure human voice signals and uses the fourth AI module to generate target language text associated with the pure human voice signals. The prosodic features of the pure human voice signal are extracted in time sequence, and a preliminary audio stream is generated by combining it with the human voice module. The duration of the continuous audio segments is compared with the pure human voice signal and the speed is adjusted to obtain a secondary audio stream. The independent background audio and the secondary audio stream are then fitted in time sequence to generate the final audio stream. The video synthesis end divides the standard video stream by frame number to obtain a standard video frame sequence, and extracts the lip image frame sequence associated with the standard video frame sequence; The secondary audio stream is then divided according to the frame order of the standard video frame sequence to obtain the secondary audio frame sequence. The fifth AI module generates a lip vertex displacement coordinate sequence that matches the secondary audio frame sequence frame by frame, and renders the corresponding frames of the lip image frame sequence. The rendered lip image frame sequence is fitted to a standard video frame sequence, and the final audio frame sequence is time-aligned to synthesize an audio-video synchronized sequence. At the video output end, the determined audio and video synchronization sequence is compressed and encapsulated using predefined encoding rules to output the final audio and video data.
2. The multimodal AI-driven short video automatic translation and speech synthesis system according to claim 1, characterized in that, The first AI module described in this system includes a pre-trained convolutional neural network and TensorFlow; The second AI module includes FFmpeg and a format compatibility engine; The third AI module includes Non-negative Matrix Factorization (NMF) and Spleeter; The fourth AI module includes a neural machine translation model with an encoder-decoder architecture and Sipmpleon-listening; The fifth AI module includes an adversarial network-driven 3D facial deformation model and LipSync.
3. The multimodal AI-driven short video automatic translation and speech synthesis system according to claim 1, characterized in that, In the video acquisition terminal, the audio and video data submitted by the user to the system includes local files, network URLs, and real-time streaming media; The audio and video data submitted by users to this system are cached in the cache pool and included in the violation detection queue in the order of caching. When the first AI module is idle, it extracts the first audio / video data from the violation detection queue for violation detection. If the violation detection passes, the first audio / video data will be added to the storage pool; If the violation detection fails, a clearing operation is performed, removing the first audio / video data from the violation detection queue and cache pool.
4. The multimodal AI-driven short video automatic translation and speech synthesis system according to claim 3, characterized in that, The specific method for outputting standard audio and video data in the preliminary processing terminal is as follows: Extract any audio or video data from the storage pool, denoted as YP; The second AI module is used to identify the container format CF associated with the audio and video data YP. YP and encoding parameter EP YP It is then compared with the container format set CF and the encoding parameter set EP preset by the operator; in, And CF i Corresponding EP i j is the value preset by the operator, i is the counting index, 1≤i≤j; If the container format is CF YP Belongs to the container format set CF and has encoding parameter EP YP It belongs to the encoding parameter set EP, and CF YP Corresponding EP YP If the comparison is successful, the audio / video data YP will be marked as the standard audio / video data SYP. Conversely, if the comparison fails, the container format CF associated with the audio / video data YP will be changed. YP and encoding parameter EP YP Real-time transcoding to any pair of CFs from the container format set CF and the encoding parameter set EP. i With EP i The real-time transcoded audio and video data YP is recorded as the standard audio and video data SYP.
5. The multimodal AI-driven short video automatic translation and speech synthesis system according to claim 4, characterized in that, In the preliminary processing unit, the specific method for performing the separation operation on the standard audio and video data is as follows: The standard audio and video data SYP is demultiplexed to separate the standard audio stream and standard video stream in the standard audio and video data SYP, which are denoted as SAS and SVS respectively. The duration of the standard audio stream SAS is equal to the duration of the standard video stream SVS. Based on the established standard audio stream SAS, a separation operation is performed by the third AI module to determine the human voice track and the background audio track to obtain a clean human voice signal and an independent background audio, denoted as PVS and IBA, respectively.
6. The multimodal AI-driven short video automatic translation and speech synthesis system according to claim 5, characterized in that, The specific method for generating the final audio stream in the video translation terminal is as follows: Extract the pure human voice signal PVS, and use the fourth AI module to perform text recognition on the pure human voice signal PVS to generate first language text; The first language text is translated into the user-defined target language to obtain the target language text, denoted as TLT. The target language is derived from the operator's preset target language library. The duration and total number of moments of the pure human voice signal PVS are determined. Based on the fourth AI module, the prosodic features corresponding to the moments in the pure human voice signal PVS are extracted to obtain the prosodic feature sequence. Then, the prosodic feature sequence and the target language text TLT are input into the human voice module pre-built by the operator to generate a preliminary audio stream associated with the target language text TLT, denoted as FAS; Determine the duration of all consecutive audio segments in the initial audio stream FAS, and record them in chronological order as the audio segment duration sequence CAS. FAS-1 CAS FAS-2 ,...,CAS FAS-m , where m is the total number of consecutive audio segments in the initial audio stream PAS; Similarly, the duration sequence of consecutive audio segments associated with the pure human voice signal PVS is determined by CAS. PVS-1 CAS PVS-2 ,...,CAS PVS-m ; Extracting the duration sequence of consecutive audio segments using CAS FAS-1 CAS FAS-2 ,...,CAS FAS-m CAS duration of any audio segment FAS-n CAS with audio continuous segment duration sequence PVS-1 CAS PVS-2 ,...,CAS PVS-m CAS duration of any audio segment PVS-n Perform a comparison, where n is the counting index, with a value ranging from 1 to m; If the duration of a continuous audio segment is CAS FAS-n CAS of audio continuous segment duration PVS-n If the error is greater than the error percentage threshold of 5%, then the duration of the audio segment CAS is... FAS-n Perform speed adjustment, and adjust the duration of continuous audio segments using CAS. FAS-n CAS of audio continuous segment duration PVS-n The error is controlled within 5% of the error percentage threshold; Similarly, for the duration sequence of consecutive audio segments, CAS FAS-1 CAS FAS-2 ,...,CAS FAS-m The duration of all consecutive audio segments is adjusted and summed so that the summed audio duration is equal to the duration of the pure human voice signal, and this is denoted as the secondary audio stream EAS. The determined secondary audio stream EAS is then fitted with the independent background audio IBA to generate the final audio stream, denoted as LAS.
7. The multimodal AI-driven short video automatic translation and speech synthesis system according to claim 6, characterized in that, In the video synthesis terminal, the specific method for extracting the lip image frame sequence associated with the standard video frame sequence is as follows: Extract standard video stream SVS; Determine the total number of frames in the standard video stream SVS, denoted as o; Divide the standard video stream into frames based on the total number of frames o, resulting in the standard video frame sequence SVS1, SVS2, ..., SVS o ; Extract the standard video frame sequence SVS1, SVS2, ..., SVS o Any standard video frame in the video is denoted as SVS. u Where u is the counting index, 1≤u≤o; Extracting Standard Video Frames (SVS) using facial recognition technology u The lip image is used to obtain the lip image frame, denoted as LIF. u ; And so on, to obtain the standard video frame sequence SVS1, SVS2, ..., SVS o The lip image frames associated with all standard video frames are summarized to form the lip image frame sequence LIF1, LIF2, ..., LIF o .
8. The multimodal AI-driven short video automatic translation and speech synthesis system according to claim 7, characterized in that, In the video synthesis terminal, the specific method for rendering the corresponding frames of the lip image frame sequence is as follows: S81. Extract the determined secondary audio stream EAS and the lip image frame sequence LIF1, LIF2, ..., LIF o ; According to the lip image frame sequence LIF1, LIF2, ..., LIF o The secondary audio stream EAS is divided into frames according to the frame order, resulting in the secondary audio frame sequence EAS1, EAS2, ..., EAS. o ; S82. Determine the lip image frame sequence LIF1, LIF2, ..., LIF o Any lip image frame in LIF u and the corresponding secondary audio frame sequences EAS1, EAS2, ..., EAS o Secondary audio frames EAS u ; S83, Transfer the lip image frame to LIF u and secondary audio frames EAS u Input the fifth AI module to generate EAS audio frames. u Matching lip vertex displacement coordinates LVD u ; S84, using LVD (Lip Vertex Displacement Coordinates) u LIF image frames of the lips u Rendering is performed to make the lip image frame LIF u Conforms to secondary audio frames EAS u The shape of the mouth; S85. Repeat steps S82 to S84 to render the lip image frame sequence LIF1, LIF2, ..., LIF o From all the lip image frames, obtain the rendered sequence EAS1, EAS2, ..., EAS2, ... o The matched sequence of lip image frames is denoted as LIF′1, LIF′2, ..., LIF′. o .
9. The multimodal AI-driven short video automatic translation and speech synthesis system according to claim 8, characterized in that, The specific method for synthesizing the audio-video synchronization sequence in the video synthesis terminal is as follows: Extract the lip image frame sequence LIF'1, LIF'2, ..., LIF' o And fit it to the standard video frame sequence SVS1, SVS2, ..., SVS o In this process, the fitted standard video frame sequence SVS'1,SVS'2,...,SVS' is obtained. o ; Extract the final audio stream LAS, according to the standard video frame sequence SVS'1, SVS'2, ..., SVS' o The final audio stream LAS is divided into frames according to the frame order, resulting in the final audio frame sequence LAS1, LAS2, ..., LAS o ; The final audio frame sequence LAS1, LAS2, ..., LAS o With the standard video frame sequence SVS'1,SVS'2,...,SVS' o Alignment and fitting are performed frame by frame to obtain the audio and video synchronization sequences LAS1-SVS'1, LAS2-SVS'2, ..., LAS' associated with the audio and video data SYP. o -SVS′ o .
10. The multimodal AI-driven short video automatic translation and speech synthesis system according to claim 9, characterized in that, The specific method for outputting the final audio and video data at the video output terminal is as follows: Extract the audio and video synchronization sequences LAS1-SVS'1, LAS2-SVS'2, ..., LAS associated with the audio and video data SYP. o -SVS' o The system uses predefined encoding rules to compress and encapsulate the data, outputting the final audio and video data.
Citation Information
Patent Citations
Video translation method, system and device and storage medium
CN112562721A
Live stream review intervention method and device, storage medium and equipment
CN114339292A
Video dubbing system and method, electronic equipment and storage medium
CN119152854A
End-to-end video translation method and device based on artificial intelligence and medium
CN119211653A
Automatic video translation system based on FFMPEG, TTS and Wav2Lip
CN119830924A
Cited By
Cross-platform video adaptation response method and system
CN121603733A