Video processing method, device and equipment for translation and dubbing, and medium
Patent Information
- Application Number
- CN202610946262.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-29
- Publication Date
- 2026-09-25
AI Technical Summary
[0009]鉴于以上内容,有必要提供一种面向翻译及配音的视频处理方法、装置、设备及介质,旨在解决面向翻译及配音的视频处理效果不佳的问题
[0014]由以上技术方案可以看出,本发明能够对初始视频执行自动语音识别,得到带有时间戳、角色及语音活动边界的初始转写片段,以避免出现角色语气不符合场景、上下文指代断裂问题;对初始转写片段执行字幕级片段重构,避免由于字幕切分不合理影响字幕可读性及配音可用性;生成多模态语境及术语表,并根据视频处理模式对多个待处理片段进行校准,能够避免术语误识别,阻断错误传递;对多个片段级译文执行时长与可说性修正,能够解决译文时长不匹配原视频时间窗口问题;基于主客观协同质量评估的双引擎生成质量评分,并根据评分进行优化,能够量化配音质量,从而生成面向翻译及配音的高质量视频。
Smart Images

Figure CN122824935A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence technology, and in particular to a video processing method, apparatus, device, and medium for translation and dubbing. Background Technology
[0002] With the rapid growth of short videos, film clips, online courses, corporate promotional videos, cross-border e-commerce videos, and social media content, the demand for cross-language dissemination of video content has increased significantly. Traditional video translation and dubbing typically rely on manual processes such as dictation, translation, subtitle creation, character-specific dubbing, audio track mixing, and video export. This process is time-consuming, costly, and difficult to adapt to the large-scale, multilingual, and rapidly iterating demands of video localization.
[0003] Existing automated video translation and dubbing technologies typically include modules such as speech recognition, machine translation, text-to-speech, and audio-visual synthesis. The system generally first extracts speech from the video and transcribes it into text, then translates the transcribed text into the target language, subsequently uses TTS (Text-to-Speech) to synthesize the target speech, and finally combines the target speech with the original video for output. While this method can reduce manual labor costs to some extent, it still faces many challenges in practical applications.
[0004] First, ASR (Automatic Speech Recognition) results are easily affected by factors such as background music, ambient noise, multiple speakers overlapping, accents, proper nouns, and unreasonable sentence breaks, leading to problems such as misspelled words, missing words, repeated text, and unreasonable segment boundaries. If ASR results are directly entered into the translation stage without calibration, errors will continue to propagate to the translated text, subtitles, and target speech, affecting the final dubbing quality.
[0005] Secondly, existing machine translation typically focuses on plain text or subtitles, lacking full utilization of video footage, character relationships, plot background, speaking styles, terminology, and scene tone. For content such as names of people, places, brands, cultural expressions, and character titles, the lack of pre-translation context modeling and terminology constraints can easily lead to mistranslations, inconsistencies, and tone that doesn't match the character.
[0006] Furthermore, video dubbing not only requires semantic accuracy in the translation but also demands that the target speech be delivered naturally within the original speaking time window. Different languages exhibit significant differences in expression length and speech rate, meaning that directly translated target text may be too long or too short, leading to a mismatch between the TTS-synthesized audio and the original video's lip movements, gestures, and subtitle rhythm. Existing methods often address the duration issue by simply speeding up or stretching the audio, which easily results in mechanical, incomprehensible, or unnatural-sounding speech.
[0007] Furthermore, existing automated tools have weak capabilities in handling multi-role scenarios, and some even lack the concept of roles altogether. For videos with multiple speakers, systems often can only perform simple segmentation based on silence or audio clips, making it difficult to identify the relationships between different roles or assign a stable and characteristic voice to each role. As a result, the same character may exhibit voice drift in different segments, and different characters may be assigned similar or mismatched voices, leading to unstable automatic dubbing effects for multi-role videos. In actual production, manual recording for each role is still necessary.
[0008] Furthermore, existing technologies suffer from low processing efficiency and the tendency for errors to accumulate when dealing with long videos and complex scenes. Long videos typically contain more characters, scene transitions, terminology repetition, cross-segment references, and contextual dependencies. If the system lacks a global character table, terminology table, and quality status management, problems such as inconsistent translation styles, inconsistent character voices, inconsistent terminology, and timeline misalignment can easily arise after segmented processing. Summary of the Invention
[0009] In view of the above, it is necessary to provide a video processing method, apparatus, device and medium for translation and dubbing, aiming to solve the problem of poor video processing effect for translation and dubbing.
[0010] A video processing method for translation and dubbing, the method comprising: In response to a video processing instruction triggered based on an initial video, automatic speech recognition is performed on the initial video to obtain an initial transcription segment with timestamps, roles, and boundaries of speech activities; The initial transcribed segment is reconstructed at the subtitle level to obtain multiple segments to be processed; A multimodal context and terminology list are generated based on the initial video, the initial transcribed segment, and the multiple segments to be processed; Obtain the video processing mode of the initial video, and calibrate the multiple segments to be processed according to the video processing mode to obtain multiple calibrated segments; Based on the multimodal context and the terminology, translation operations are performed on the multiple calibration segments to obtain multiple segment-level translations, and the multiple segment-level translations are modified in terms of duration and describability to obtain multiple modified translations; Based on the timestamp, the role, the speech activity boundary, and the multimodal context, the target speech is generated according to the multiple corrected translations, and the initial video and the target speech are synthesized and rendered to obtain an initial dubbing video; The initial quality score of the initial dubbing video is generated by a dual-engine system based on a combined subjective and objective quality assessment. The initial dubbing video is then optimized based on the initial quality score and the video processing mode to obtain the target video after translation and dubbing.
[0011] A video processing apparatus for translation and dubbing, the video processing apparatus for translation and dubbing comprising: The recognition unit is used to perform automatic speech recognition on the initial video in response to a video processing instruction triggered based on the initial video, and to obtain an initial transcription segment with timestamps, roles and speech activity boundaries; The reconstruction unit is used to perform subtitle-level segment reconstruction on the initial transcription segment to obtain multiple segments to be processed. The generation unit is configured to generate a multimodal context and terminology list based on the initial video, the initial transcription segment, and the multiple segments to be processed; A calibration unit is used to acquire the video processing mode of the initial video, and calibrate the multiple segments to be processed according to the video processing mode to obtain multiple calibration segments; The correction unit is configured to perform translation operations on the multiple calibration segments according to the multimodal context and the terminology list to obtain multiple segment-level translations, and to perform duration and speakability corrections on the multiple segment-level translations to obtain multiple corrected translations. The rendering unit is used to generate target speech based on the timestamp, the character, the speech activity boundary and the multimodal context, according to the multiple corrected translations, and to synthesize and render the initial video and the target speech to obtain an initial dubbing video; The optimization unit is used to generate an initial quality score for the initial dubbing video based on a dual-engine system of subjective and objective collaborative quality assessment, and to optimize the initial dubbing video according to the initial quality score and the video processing mode to obtain the target video after translation and dubbing.
[0012] A computer device, the computer device comprising: A memory that stores at least one instruction; and a processor that executes the instructions stored in the memory to implement the video processing method for translation and dubbing.
[0013] A computer-readable storage medium storing at least one instruction, which is executed by a processor in a computer device to implement the video processing method for translation and dubbing.
[0014] As can be seen from the above technical solutions, this invention can perform automatic speech recognition on the initial video to obtain an initial transcribed segment with timestamps, characters, and boundaries of voice activities, thus avoiding problems such as character tone not matching the scene and broken contextual references; it performs subtitle-level segment reconstruction on the initial transcribed segment to avoid affecting the readability of subtitles and the usability of dubbing due to unreasonable subtitle segmentation; it generates a multimodal context and terminology list, and calibrates multiple segments to be processed according to the video processing mode, which can avoid terminology misidentification and block error transmission; it performs duration and speakability correction on multiple segment-level translations, which can solve the problem of translation duration not matching the original video time window; it generates a quality score based on a dual-engine system of subjective and objective collaborative quality assessment, and optimizes according to the score, which can quantify dubbing quality, thereby generating high-quality videos for translation and dubbing. Attached Figure Description
[0015] Figure 1 This is a flowchart of a preferred embodiment of the video processing method for translation and dubbing of the present invention.
[0016] Figure 2 This is a functional block diagram of a preferred embodiment of the video processing apparatus for translation and dubbing of the present invention.
[0017] Figure 3 This is a schematic diagram of the structure of a computer device that implements a preferred embodiment of the video processing method for translation and dubbing according to the present invention. Detailed Implementation
[0018] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be described in detail below with reference to the accompanying drawings and specific embodiments.
[0019] like Figure 1 The diagram shown is a flowchart of a preferred embodiment of the video processing method for translation and dubbing according to the present invention. The order of the steps in this flowchart can be changed, and some steps can be omitted, depending on different requirements.
[0020] The video processing method for translation and dubbing is applied to one or more computer devices. The computer device is a device that can automatically perform numerical calculations and / or information processing according to pre-set or stored instructions. Its hardware includes, but is not limited to, microprocessors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), embedded devices, etc.
[0021] The computer device can be any electronic product that can interact with the user, such as a personal computer, tablet computer, smartphone, personal digital assistant (PDA), interactive network television (IPTV), smart wearable device, etc.
[0022] The computer equipment may also include network equipment and / or user equipment. The network equipment includes, but is not limited to, a single network server, a server group consisting of multiple network servers, or a cloud based on cloud computing consisting of a large number of hosts or network servers.
[0023] The server can be a standalone server or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDN), and big data and artificial intelligence platforms.
[0024] Artificial intelligence (AI) is the theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results.
[0025] Foundational artificial intelligence technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interactive systems, and mechatronics. AI software technologies mainly encompass computer vision, robotics, biometrics, speech processing, natural language processing, and machine learning / deep learning.
[0026] The network in which the computer device is located includes, but is not limited to, the Internet, wide area network, metropolitan area network, local area network, and virtual private network (VPN).
[0027] S10, in response to a video processing instruction triggered based on the initial video, performs Automatic Speech Recognition (ASR) on the initial video to obtain an initial transcription segment with timestamps, roles, and speech activity boundaries.
[0028] In this embodiment, the initial video can be a film clip, online course, corporate promotional video, cross-border e-commerce video, interview program, or multi-role dialogue video, etc.
[0029] In this embodiment, the initial video may be uploaded by the triggerer of the video processing instruction.
[0030] In this embodiment, the video processing command can be automatically triggered when a video is detected being uploaded to a designated platform.
[0031] In this embodiment, the metadata of the initial video may include, but is not limited to: video duration, resolution, frame rate, audio track information, and file format.
[0032] In this embodiment, an audio and video processing program can be invoked to extract the metadata from the initial video and use it as input for subsequent ASR, background separation, and TTS (Text To Speech) alignment processing.
[0033] In this embodiment, the initial transcription segment includes at least: segment number, start time, end time, and transcribed text, and may also optionally include speaker information, speaker identifier, word-level timestamp, confidence level, and other information.
[0034] S11, perform subtitle-level segment reconstruction on the initial transcription segment to obtain multiple segments to be processed.
[0035] In this embodiment, the step of performing subtitle-level segment reconstruction on the initial transcribed segment to obtain multiple segments to be processed includes: Obtain the multidimensional reconstruction constraints constructed based on the timestamp, the role, and the voice activity boundary; The initial transcribed segment is reconstructed at the subtitle level according to the multidimensional reconstruction constraints to obtain the multiple segments to be processed. Specifically, the initial transcription fragments with a length greater than the maximum length threshold are split, and the initial transcription fragments with a length less than the minimum length threshold are merged. Each segment to be processed corresponds to a character.
[0036] The multidimensional reconstruction constraints may also include, but are not limited to: semantic integrity, pause position, time window, text length, subtitle readability, speaker boundary, etc.
[0037] This involves splitting excessively long segments, merging excessively short or semantically fragmented segments, and preventing content from different speakers from being mixed into the same segment, thereby obtaining a segment structure suitable for subtitle display, translation, and TTS synthesis.
[0038] In the above embodiments, by reconstructing subtitle-level segments, the problems of existing ASR output segments being too long or too short, semantically fragmented, mixed segments across speakers, or unsuitable for subtitle display can be solved, enabling the transcribed text to form a more reasonable segment structure according to semantic integrity, time window, speaker boundaries, and subtitle readability.
[0039] S12, generate a multimodal context and terminology list based on the initial video, the initial transcription segment, and the multiple segments to be processed.
[0040] In this embodiment, after the ASR is completed and before the formal translation, multimodal analysis can be performed on the initial video footage, audio, initial transcription fragments, speaker information, and target language to generate video themes, plot backgrounds, scene tone, character relationships, character traits, speaking styles, and translation considerations as a multimodal context.
[0041] In this embodiment, entities or terms such as personal names, place names, organization names, brand names, professional terms, role titles, cultural expressions, and high-frequency phrases can be extracted from video footage, audio, and transcribed text, and combined with the target language to generate a project-level translation terminology list.
[0042] The glossary may include information such as source language terms, target language translations, contextual descriptions, categories, and confidence levels.
[0043] S13, obtain the video processing mode of the initial video, and calibrate the multiple segments to be processed according to the video processing mode to obtain multiple calibrated segments.
[0044] In this embodiment, the video processing mode may include a pipeline mode and an intelligent agent mode. Both modes support both manual retouching and process control in high-quality content creation, as well as automated production of batch videos.
[0045] In this embodiment, the pipeline mode is mainly suitable for video translation and dubbing scenarios that are sensitive to terminology, have complex character relationships, require high dubbing quality, require manual step-by-step review, or require partial modification of intermediate results; the intelligent agent mode is mainly suitable for video localization scenarios that require batch short video, multilingual content generation, or automated processing.
[0046] In the intelligent agent mode, configuration parameters such as target language, maximum number of optimization iterations, quality threshold, cost threshold, whether to automatically skip manual confirmation, and whether to enable closed-loop optimization can be pre-configured. Based on the above configuration parameters, dubbing tasks can be created, and task status, cost budget, quality goals, and optimization strategies can be initialized.
[0047] In the intelligent agent mode, upon receiving the initial video, the system can automatically extract video metadata and the original audio, separate human voice from background noise, and perform quality checks on the separation results to determine whether the background audio track meets the requirements for subsequent mixing and whether there is any risk of residual human voice from the original language. Subsequently, it performs ASR speech recognition, generates initial transcription segments, and records the timestamp, speaker, confidence level, and speech activity boundaries of each segment.
[0048] In this embodiment, calibrating the plurality of segments to be processed according to the video processing mode to obtain a plurality of calibrated segments includes: When the video processing mode is pipeline mode, high-risk segments among the multiple segments to be processed are identified according to preset rules, the triggerer of the video processing instruction receives the processing instruction for the high-risk segments, and the high-risk segments are calibrated according to the processing instruction to obtain the multiple calibrated segments; or When the video mode is in agent mode, the agent is invoked to identify the high-risk and low-risk segments among the multiple segments to be processed; the agent is used to automatically correct the low-risk segments and record the high-risk segments as candidate problems.
[0049] Among them, high-risk segments can be identified from the multiple segments to be processed based on obvious misspellings, omissions, duplicate text, abnormal sentence breaks, punctuation errors, unreasonable segment boundaries, and suspected misidentification of terms.
[0050] The processing instructions may include confirmation, modification, or skip instructions.
[0051] In addition to the high-risk segments, the triggerer can also issue processing instructions for multimodal contexts, role information, and terminology.
[0052] In pipeline mode, the content confirmed by the triggerer can be used as a constraint input for subsequent translation and dubbing synthesis.
[0053] In the agent mode, high-risk segments can be recorded as candidate problems for subsequent quality scoring and optimization.
[0054] The intelligent agent can be a large language model, etc., which has the ability to automatically identify the risk level of each segment to be processed and to automatically correct it.
[0055] This embodiment solves the problems of existing technologies that directly enter translation after ASR, lack of high-risk transcription exposure, lack of terminology entity extraction and pre-translation context constraints by sequentially performing multimodal context understanding and transcription calibration after ASR. It reduces the risk of errors such as misspellings, omissions, abnormal sentence breaks, and misidentification of terms being transmitted to the translation, subtitling and TTS stages.
[0056] S14, perform translation operations on the multiple calibration segments according to the multimodal context and the terminology to obtain multiple segment-level translations, and perform duration and speakability corrections on the multiple segment-level translations to obtain multiple corrected translations.
[0057] In this embodiment, translation can be performed on each calibration segment based on multimodal context, terminology, speaker information, and target language, generating multiple segment-level translations.
[0058] In the translation process, priority is given to ensuring consistency in terminology, tone of voice, and contextual coherence, and making the translation suitable for spoken delivery.
[0059] In this embodiment, the process of performing time and describability corrections on the multiple fragment-level translations to obtain multiple corrected translations includes: Obtain the translation length of each segment-level translation, and predict the corresponding text-to-speech (TTS) duration based on the translation length of each segment-level translation; For each segment-level translation, calculate the time difference between the TTS duration and the original time window length of the corresponding calibration segment; When the time difference exceeds a first threshold, the segment-level translation is compressed, rewritten in a colloquial style, or rewritten semantically, or a modification prompt is sent to the triggerer; or When the time difference is less than or equal to the first threshold and greater than or equal to the second threshold, the speech rate of the segment-level translation is adjusted, the audio is stretched or compressed, the segment time is fine-tuned, or the audio-visual offset is corrected.
[0060] Among them, the start time, end time, duration, target language, and length of each segment-level translation can be used to comprehensively determine whether the translation can be naturally completed within the original time window.
[0061] Using the first threshold and the second threshold as the upper and lower boundaries, an adjustable range can be obtained. Segment-level translations with time differences within this range do not need to be rewritten and can be finely adjusted directly.
[0062] In particular, by combining speech rate adjustment or timeline fine-tuning, the target speech can be more easily aligned with the original video.
[0063] In the above embodiments, the translation duration alignment mechanism of dubbing speakability can solve the problem that ordinary machine translation does not consider whether the target speech can be spoken naturally within the original time window. Through translation length budgeting, translation compression and rewriting, TTS duration detection, speech rate adjustment and time axis fine-tuning, the matching degree between the target speech and the original video actions, lip movements and subtitle rhythm can also be improved.
[0064] S15, according to the timestamp, the role, the speech activity boundary and the multimodal context, the target speech is generated based on the multiple corrected translations, and the initial video and the target speech are synthesized and rendered to obtain the initial dubbing video.
[0065] In this embodiment, target timbre can be assigned to different speakers based on speaker information and target language, and TTS audio can be generated as the target speech based on modified translation, timbre, emotion tags, pauses, stress, speech rate and prosody parameters.
[0066] For segments that require emphasizing emotions, expression parameters such as calm, excitement, sadness, anger, whisper, and seriousness can be configured to make the target speech more consistent with the multimodal context and character state of the video.
[0067] The above embodiments can solve the problems of mismatched emotions, unnatural pauses, abrupt speech rate, and discontinuous expression between adjacent segments in existing TTS dubbing, enabling the generated target speech to produce a more natural and scene-appropriate expression effect based on the multimodal context of the video, the character's state, the content of the segment, and the duration of the segment. In this embodiment, depending on actual needs, the initial video can be separated into human voice and background audio first, and it can be determined whether the background audio track contains residual original speaker audio, lost background audio, pure dialogue, or separation failure. If the background audio track is of reliable quality, the target speech and background audio are subjected to loudness normalization, volume control, background ducking (audio ducking / sidechain compression), and mixing. If there is a risk of crosstalk, the original audio backfill path that may lead to leakage of the original language voice or crosstalk is skipped.
[0068] The above-mentioned audio and video mixed rendering protection mechanism resolves the contradiction between preserving background sound and leaking the original language voice. By separating the voice and background for quality judgment, controlling the background sound intensity, protecting the mixing, and avoiding crosstalk risks, the clarity of the target voice and the quality of the final product are improved while preserving the original video atmosphere.
[0069] Furthermore, the target speech can be synthesized with the initial video according to the timeline of the segments. Specifically, the target speech, background sound, subtitles, and the original video footage of the initial video can be synthesized and rendered to generate the initial dubbed video.
[0070] In the above embodiments, by using multi-role recognition, role modeling, and timbre binding mechanisms, the problems of existing automation tools lacking role concepts, being unable to stably assign timbres to different roles, timbre drift of the same role across segments, and timbre confusion between different roles can be solved, enabling stable, distinguishable, and role-character-compliant automatic dubbing for multi-speaker videos. Through a global state maintenance mechanism for long videos and complex scenes, the problems of inconsistent roles, inconsistent terminology, context breaks, timeline misalignment, and accumulation of quality issues during long video segmentation can be solved, improving the stability and processing efficiency of automatic dubbing for long videos.
[0071] S16, the initial quality score of the initial dubbing video is generated by the dual engine based on the subjective and objective collaborative quality assessment, and the initial dubbing video is optimized according to the initial quality score and the video processing mode to obtain the target video after translation and dubbing.
[0072] In this embodiment, the initial quality score generated by the dual-engine system based on subjective and objective collaborative quality assessment for the initial dubbed video includes: Obtain an objective indicator evaluation engine based on objective evaluation indicators, and obtain a subjective perception evaluation engine based on subjective perception evaluation indicators. The objective evaluation engine is used to score the quality of each segment in the initial dubbing video, resulting in an objective score for each segment. The initial dubbing video is scored using the objective indicator evaluation engine to obtain an overall objective score for the initial dubbing video. The subjective perception evaluation engine is used to score the quality of each segment in the initial dubbing video, resulting in a subjective score for each segment. The subjective perception evaluation engine is used to score the quality of the initial dubbing video, resulting in an overall subjective score for the initial dubbing video. The objective score for each segment is combined with the corresponding subjective score to obtain a segment-level quality score for each segment. The overall objective score and the overall subjective score of the initial dubbing video are combined to obtain the global quality score of the initial dubbing video.
[0073] The dual engines can be used to obtain comprehensive quality scores at the segment, character, and video levels.
[0074] The evaluation dimensions of the dual engines may include audio and video synchronization, TTS duration deviation, speech rate and rhythm, volume and loudness, subtitle timing matching, translation naturalness, emotional fit, character consistency, timbre similarity, background voice retention, and overall viewing experience.
[0075] The objective indicator evaluation engine is used to calculate quantifiable indicators. The objective indicator evaluation engine includes at least one or more of the following indicators: audio and video synchronization error, subtitle timeline deviation, speech duration deviation, speech rate outliers, volume loudness, background noise persistence, timbre similarity, speaker consistency, MOS (Mean Opinion Score) prediction score, lip-sync accuracy, and audio intelligibility.
[0076] The subjective perception evaluation engine is used to simulate the comprehensive judgment of human beings on the quality of dubbing. The subjective perception evaluation engine includes at least one or more of the following indicators: translation accuracy, naturalness of expression, emotional fit, fluency of hearing, consistency of character, context matching, and overall viewing experience.
[0077] By integrating the quality scores from both engines, it is possible to identify low-scoring segments and their corresponding problem types.
[0078] In this embodiment, after quality scoring, a structured repair operation generation task, an operation verification and execution task, and a re-scoring, rollback, and best result selection task can be executed sequentially to optimize the initial dubbed video.
[0079] Specifically, optimizing the initial dubbing video based on the initial quality score and the video processing mode to obtain the target video after translation and dubbing includes: When the video processing mode is the pipeline mode, the initial dubbing video, the segment-level quality score of each segment, and the global quality score are displayed on the designated display interface; for segments to be optimized whose segment-level quality scores are less than the score threshold, the segments to be optimized are locally modified to regenerate replacement segments; the replacement segments are then used to replace the initial dubbing video to obtain the target video; or When the video mode is the agent mode, the agent is invoked to perform problem localization to obtain the problem localization result. Based on the segment-level quality score of the segment to be optimized, the global quality score, and the problem localization result, a structured repair operation is generated. The structured repair operation is filtered to obtain an executable operation. The executable operation is executed on the segment to be optimized to obtain the optimized segment, and the optimized segment is replaced in the initial dubbing video to obtain the target video.
[0080] The structured repair operations may include, but are not limited to: source text correction, translation correction, translation rewriting, segment splitting, segment merging, timestamp adjustment, speech rate adjustment, audio and video offset correction, emotion adjustment, TTS style adjustment, TTS recombination, volume adjustment, speaker reassignment, subtitle correction, background sound processing, and re-rendering.
[0081] Furthermore, priorities can be assigned to each structured repair operation based on the severity of the problem and the estimated quality benefits.
[0082] The filtering of the structured repair operation to obtain executable operations includes: The structured repair operation is subjected to parameter validation, severity screening, execution order sorting, duplicate operation deduplication, cost assessment, and operation disabling filtering to retain the executable operations.
[0083] After executing the executable operation, it can be determined whether TTS re-synthesis or video re-rendering is required based on the operation type. For example, operations such as translation correction, mood adjustment, timbre replacement, or speech rate parameter adjustment can trigger TTS re-synthesis of the corresponding segment; operations such as timeline adjustment, volume adjustment, subtitle correction, background sound modification, or lip-sync correction can trigger video re-rendering.
[0084] Furthermore, after replacing the optimized segment with the initial dubbing video to obtain the target video, the method further includes: The dual engines are invoked to perform iterations; wherein, during each iteration, a quality score for each generated video is generated sequentially, and the generated video is optimized based on the generated quality score; When the stopping condition is met, the video with the highest quality score is determined as the optimal video, and the optimal video is output. The process includes recording an iteration log, which is used to incrementally optimize the agent.
[0085] For example, in each iteration, the dual-engine quality scoring can be performed again on the repaired target video, and the new score can be compared with the historical version. If the score improves, the current result is retained; if the score decreases, a key dimension deteriorates significantly, or there is no significant benefit after multiple rounds of optimization, the system can revert to the historical best version or stop optimization.
[0086] Among these features, task snapshots can be saved before each round of optimization to support version rollback and selection of the best result.
[0087] The process, from quality scoring to generating structured repair operations, to operation verification and execution, and then to re-scoring, rollback, and best result selection, is a complete iterative process.
[0088] Among them, when the quality threshold is reached, the maximum number of iterations is reached, the score enters a plateau period, there are no executable problems, or the cost threshold is exceeded, the stopping condition can be determined, and the dubbing video with the best quality or that meets the threshold can be output.
[0089] It can also simultaneously generate scoring records, repair operation records, optimization trajectories, and quality reports.
[0090] Furthermore, the high-scoring samples ultimately adopted, the low-scoring samples after manual correction, and the records of repair operations and quality score changes can be used as feedback data for incremental training of subsequent translation models, TTS parameter selection models, quality assessment models, or optimization strategy models. Thus, the quality score can not only serve as the basis for final selection of finished products but also as a feedback signal for agents such as the generation and optimization models, enabling the system to form an end-to-end closed loop of "generation-evaluation-feedback-optimization".
[0091] The above embodiments address the shortcomings of existing video dubbing systems, such as the lack of multi-dimensional quality assessment, structured repair, automatic recompositing, re-rendering, and rollback optimization. This enables the system to continuously improve dubbing quality across dimensions including translation accuracy, audio-visual synchronization, speech rate and rhythm, emotional matching, volume balance, and speaker consistency. The dual-engine quality assessment and score feedback optimization mechanism, which combines objective and subjective evaluation, solves the problems of existing technologies' inability to quantitatively assess dubbing quality, reliance on manual quality control, and the system's inability to learn from errors. By combining objective indicator evaluation with subjective perception evaluation, a comprehensive quality score can be generated for each segment, character, and complete video. This quality score serves as a feedback signal to drive retranslation, rewriting, recompositing, re-rendering, and generation strategy optimization, giving the video dubbing system sustainable improvement capabilities.
[0092] In this embodiment, the pipeline mode allows for the introduction of manual confirmation and feedback at each key node. Combined with quality scoring, the final quality of the finished product is quantitatively evaluated and problems are identified. This reduces the risks of ASR error transmission, terminology mistranslation, inconsistent character tone, voice drift, TTS duration mismatch, and mixing crosstalk, thereby improving the controllability, traceability, and final quality of the video dubbing process.
[0093] In this embodiment, through the intelligent agent model, video dubbing generation, quality diagnosis, structured repair, local recompositing, re-rendering, version rollback, and optimal result selection can be completed without human intervention. Simultaneously, through a dual-engine quality evaluation system that combines subjective and objective assessments, the dubbing quality, which originally relied on human perception, can be transformed into a calculable, comparable, and optimizable score, thereby improving the automation, quality stability, and continuous optimization capabilities of batch video dubbing.
[0094] In this embodiment, the target video, subtitle file, translation, glossary, etc., after translation and dubbing can be output for users to preview, download, or further modify.
[0095] This embodiment provides a role status maintenance mechanism for multi-speaker, multi-role, and long-video scenarios. It not only identifies speakers in audio segments but also establishes role-level identity representations by combining video footage, character appearance, role location, dialogue context, address relationships, and voiceprint features. For videos with multiple speakers, differentiated target timbres matching the character's characteristics can be assigned to different roles, and the binding relationship between roles and timbres can be maintained. This binding relationship remains effective across segments, scenes, sub-segments, and optimization rounds to prevent timbreries of the same character in different scenarios or similar timbres assigning to different characters, making it difficult for viewers to distinguish between them. For long videos or complex narrative videos, the video can be divided into multiple processing segments, while maintaining a global role table, terminology table, timeline status, and historical quality records. This ensures consistency in translation style, character timbres, address expressions, and semantic context after segment processing, thereby reducing role confusion, contextual breaks, and timbre instability in automatic dubbing of long videos.
[0096] As can be seen from the above technical solutions, this invention can perform automatic speech recognition on the initial video to obtain an initial transcribed segment with timestamps, characters, and boundaries of voice activities, thus avoiding problems such as character tone not matching the scene and broken contextual references; it performs subtitle-level segment reconstruction on the initial transcribed segment to avoid affecting the readability of subtitles and the usability of dubbing due to unreasonable subtitle segmentation; it generates a multimodal context and terminology list, and calibrates multiple segments to be processed according to the video processing mode, which can avoid terminology misidentification and block error transmission; it performs duration and speakability correction on multiple segment-level translations, which can solve the problem of translation duration not matching the original video time window; it generates a quality score based on a dual-engine system of subjective and objective collaborative quality assessment, and optimizes according to the score, which can quantify dubbing quality, thereby generating high-quality videos for translation and dubbing.
[0097] like Figure 2 The diagram shown is a functional block diagram of a preferred embodiment of the video processing apparatus 11 for translation and dubbing according to the present invention. The video processing apparatus 11 for translation and dubbing includes a recognition unit 110, a reconstruction unit 111, a generation unit 112, a calibration unit 113, a correction unit 114, a rendering unit 115, and an optimization unit 116. The module / unit referred to in this invention is a series of computer program segments that can be executed by a processor and perform a fixed function, and which are stored in memory. In this embodiment, the functions of each module / unit will be described in detail in subsequent embodiments.
[0098] The recognition unit 110 is used to perform automatic speech recognition on the initial video in response to a video processing instruction triggered based on the initial video, and obtain an initial transcription segment with timestamps, roles and speech activity boundaries. The reconstruction unit 111 is used to perform subtitle-level segment reconstruction on the initial transcription segment to obtain multiple segments to be processed; The generation unit 112 is used to generate a multimodal context and terminology list based on the initial video, the initial transcription segment, and the multiple segments to be processed; The calibration unit 113 is used to acquire the video processing mode of the initial video, and calibrate the multiple segments to be processed according to the video processing mode to obtain multiple calibration segments. The correction unit 114 is used to perform translation operations on the multiple calibration segments according to the multimodal context and the terminology to obtain multiple segment-level translations, and to perform duration and speakability corrections on the multiple segment-level translations to obtain multiple corrected translations. The rendering unit 115 is used to generate target speech based on the timestamp, the character, the speech activity boundary and the multimodal context, according to the multiple corrected translations, and to synthesize and render the initial video and the target speech to obtain an initial dubbing video. The optimization unit 116 is used to generate an initial quality score for the initial dubbing video based on a dual-engine system of subjective and objective collaborative quality assessment, and to optimize the initial dubbing video according to the initial quality score and the video processing mode to obtain the target video after translation and dubbing.
[0099] As can be seen from the above technical solutions, this invention can perform automatic speech recognition on the initial video to obtain an initial transcribed segment with timestamps, characters, and boundaries of voice activities, thus avoiding problems such as character tone not matching the scene and broken contextual references; it performs subtitle-level segment reconstruction on the initial transcribed segment to avoid affecting the readability of subtitles and the usability of dubbing due to unreasonable subtitle segmentation; it generates a multimodal context and terminology list, and calibrates multiple segments to be processed according to the video processing mode, which can avoid terminology misidentification and block error transmission; it performs duration and speakability correction on multiple segment-level translations, which can solve the problem of translation duration not matching the original video time window; it generates a quality score based on a dual-engine system of subjective and objective collaborative quality assessment, and optimizes according to the score, which can quantify dubbing quality, thereby generating high-quality videos for translation and dubbing.
[0100] like Figure 3 The diagram shown is a structural schematic of a computer device that implements a preferred embodiment of the video processing method for translation and dubbing according to the present invention.
[0101] The computer device 1 may include a memory 12, a processor 13, and a bus (the arrow in the figure represents the bus), and may also include a computer program stored in the memory 12 and executable on the processor 13, such as a video processing program for translation and dubbing.
[0102] Those skilled in the art will understand that the schematic diagram is merely an example of computer device 1 and does not constitute a limitation on computer device 1. Computer device 1 can be either a bus topology or a star topology. Computer device 1 may also include more or fewer other hardware or software than shown in the diagram, or different component arrangements. For example, computer device 1 may also include input / output devices, network access devices, etc.
[0103] It should be noted that the computer device 1 described is merely an example. Other existing or future electronic products that are adaptable to this invention should also be included within the scope of protection of this invention and are incorporated herein by reference.
[0104] The memory 12 includes at least one type of readable storage medium, such as flash memory, portable hard drive, multimedia card, card-type memory (e.g., SD or DX memory), magnetic memory, magnetic disk, optical disk, etc. In some embodiments, the memory 12 can be an internal storage unit of the computer device 1, such as a portable hard drive of the computer device 1. In other embodiments, the memory 12 can be an external storage device of the computer device 1, such as a plug-in portable hard drive, smart media card (SMC), secure digital card (SD), flash card, etc., equipped on the computer device 1. Furthermore, the memory 12 can include both internal and external storage units of the computer device 1. The memory 12 can be used not only to store application software and various types of data installed on the computer device 1, such as code for video processing programs for translation and dubbing, but also to temporarily store data that has been output or will be output.
[0105] In some embodiments, the processor 13 may be composed of integrated circuits, such as a single packaged integrated circuit or multiple integrated circuits packaged with the same or different functions, including combinations of one or more central processing units (CPUs), microprocessors, digital processing chips, graphics processors, and various control chips. The processor 13 is the control unit of the computer device 1, connecting various components of the computer device 1 via various interfaces and lines. It executes programs or modules stored in the memory 12 (e.g., executing video processing programs for translation and dubbing) and calls data stored in the memory 12 to perform various functions of the computer device 1 and process data.
[0106] The processor 13 executes the operating system of the computer device 1 and various installed applications. The processor 13 executes these applications to implement the steps in the various video processing method embodiments for translation and dubbing described above, for example... Figure 1 The steps are shown.
[0107] For example, the computer program may be divided into one or more modules / units, which are stored in the memory 12 and executed by the processor 13 to complete the present invention. The one or more modules / units may be a series of computer-readable instruction segments capable of performing specific functions, which describe the execution process of the computer program in the computer device 1. For example, the computer program may be divided into an identification unit 110, a reconstruction unit 111, a generation unit 112, a calibration unit 113, a correction unit 114, a rendering unit 115, and an optimization unit 116.
[0108] The integrated unit implemented as a software functional module described above can be stored in a computer-readable storage medium. This software functional module, stored in a storage medium, includes several instructions to cause a computer device (which may be a personal computer, computer equipment, or network device, etc.) or processor to execute portions of the video processing methods for translation and dubbing described in the various embodiments of this invention.
[0109] If the modules / units integrated in the computer device 1 are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments of the present invention can also be implemented by a computer program instructing related hardware devices. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above.
[0110] The computer program includes computer program code, which may be in the form of source code, object code, executable file, or some intermediate form. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording media, USB flash drive, portable hard drive, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory, etc.
[0111] Furthermore, the computer-readable storage medium may primarily include a stored program area and a stored data area, wherein the stored program area may store the operating system, an application program required for at least one function, etc.; and the stored data area may store data created based on the use of blockchain nodes, etc.
[0112] The blockchain referred to in this invention is a novel application model of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanisms, and encryption algorithms. Essentially, a blockchain is a decentralized database, a chain of data blocks linked together using cryptographic methods. Each data block contains information about a batch of network transactions, used to verify the validity of the information (anti-counterfeiting) and generate the next block. A blockchain can include an underlying blockchain platform, a platform product service layer, and an application service layer.
[0113] The bus can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This bus can be divided into address bus, data bus, control bus, etc. For ease of representation, in... Figure 3 The bus is represented by only one straight line, but this does not mean that there is only one bus or one type of bus. The bus is configured to enable communication between the memory 12 and at least one processor 13, etc.
[0114] Although not shown, the computer device 1 may also include a power supply (such as a battery) to power various components. Preferably, the power supply can be logically connected to the at least one processor 13 through a power management device, thereby enabling functions such as charging management, discharging management, and power consumption management. The power supply may also include one or more DC or AC power supplies, recharging devices, power fault detection circuits, power converters or inverters, power status indicators, and other arbitrary components. The computer device 1 may also include various sensors, Bluetooth modules, Wi-Fi modules, etc., which will not be described in detail here.
[0115] Furthermore, the computer device 1 may also include a network interface. Optionally, the network interface may include a wired interface and / or a wireless interface (such as a Wi-Fi interface, a Bluetooth interface, etc.), which is typically used to establish communication connections between the computer device 1 and other computer devices.
[0116] Optionally, the computer device 1 may further include a user interface, which may be a display, an input unit (such as a keyboard), and optionally, a standard wired interface or a wireless interface. Optionally, in some embodiments, the display may be an LED display, a liquid crystal display, a touch-sensitive liquid crystal display, or an OLED (Organic Light-Emitting Diode) touchscreen, etc. The display may also be appropriately referred to as a screen or display unit, used to display information processed in the computer device 1 and to display a visual user interface.
[0117] It should be understood that the embodiments described are for illustrative purposes only and are not limited to this structure in the scope of the patent application.
[0118] It will be understood by those skilled in the art that Figure 3 The structure shown does not constitute a limitation on the computer device 1, and may include fewer or more components than shown, or combine certain components, or have different component arrangements.
[0119] Combination Figure 1 The memory 12 in the computer device 1 stores multiple instructions to implement a video processing method for translation and dubbing, and the processor 13 can execute the multiple instructions to achieve the following: In response to a video processing instruction triggered based on an initial video, automatic speech recognition is performed on the initial video to obtain an initial transcription segment with timestamps, roles, and boundaries of speech activities; The initial transcribed segment is reconstructed at the subtitle level to obtain multiple segments to be processed; A multimodal context and terminology list are generated based on the initial video, the initial transcribed segment, and the multiple segments to be processed; Obtain the video processing mode of the initial video, and calibrate the multiple segments to be processed according to the video processing mode to obtain multiple calibrated segments; Based on the multimodal context and the terminology, translation operations are performed on the multiple calibration segments to obtain multiple segment-level translations, and the multiple segment-level translations are modified in terms of duration and describability to obtain multiple modified translations; Based on the timestamp, the role, the speech activity boundary, and the multimodal context, the target speech is generated according to the multiple corrected translations, and the initial video and the target speech are synthesized and rendered to obtain an initial dubbing video; The initial quality score of the initial dubbing video is generated by a dual-engine system based on a combined subjective and objective quality assessment. The initial dubbing video is then optimized based on the initial quality score and the video processing mode to obtain the target video after translation and dubbing.
[0120] Specifically, the processor 13's implementation method for the above instructions can be found in [reference needed]. Figure 1 The descriptions of the relevant steps in the corresponding embodiments are not repeated here.
[0121] It should be noted that all the data involved in this case was legally obtained.
[0122] If any AI models, software tools, or components not belonging to this company appear in the embodiments of this invention, they are merely illustrative examples and do not represent actual use. All user personal information involved in the embodiments of this invention has been obtained by an entity authorized (with the knowledge and consent) or fully authorized by all parties through various legal and compliant means. The collection, storage, use, processing, transmission, provision, and disclosure of the information, data, and signals involved all comply with relevant laws and regulations and do not violate public order and good morals.
[0123] In the several embodiments provided by this invention, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of modules is only a logical functional division, and other division methods may be used in actual implementation.
[0124] This invention can be used in a wide variety of general-purpose or special-purpose computer system environments or configurations. Examples include: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, and distributed computing environments including any of the above systems or devices. This invention can be described in the general context of computer-executable instructions, such as program modules, that are executed by a computer. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform specific tasks or implement specific abstract data types. This invention can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.
[0125] The modules described as separate components may or may not be physically separate. The components shown as modules may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.
[0126] Furthermore, the functional modules in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or in the form of hardware plus software functional modules.
[0127] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the present invention can be implemented in other specific forms without departing from the spirit or essential characteristics of the present invention.
[0128] Therefore, the embodiments should be considered exemplary and non-limiting in all respects, and the scope of the invention is defined by the appended claims rather than the foregoing description. Thus, all variations falling within the meaning and scope of equivalents of the claims are intended to be embraced within the invention. No appended diagram markings in the claims should be construed as limiting the scope of the claims.
[0129] Furthermore, it is clear that the word "comprising" does not exclude other units or steps, and the singular does not exclude the plural. Multiple units or devices described in this invention can also be implemented by a single unit or device through software or hardware. Terms such as "first," "second," etc., are used to indicate names and do not indicate any specific order.
[0130] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention.
Claims
1. A video processing method for translation and dubbing, characterized in that, The video processing method for translation and dubbing includes: In response to a video processing instruction triggered based on an initial video, automatic speech recognition is performed on the initial video to obtain an initial transcription segment with timestamps, roles, and boundaries of speech activities; The initial transcribed segment is reconstructed at the subtitle level to obtain multiple segments to be processed; A multimodal context and terminology list are generated based on the initial video, the initial transcribed segment, and the multiple segments to be processed; Obtain the video processing mode of the initial video, and calibrate the multiple segments to be processed according to the video processing mode to obtain multiple calibrated segments; Translation operations are performed on the multiple calibration segments based on the multimodal context and the terminology to obtain multiple segment-level translations, and duration and speakability corrections are performed on the multiple segment-level translations to obtain multiple corrected translations; Based on the timestamp, the role, the speech activity boundary, and the multimodal context, the target speech is generated according to the multiple corrected translations, and the initial video and the target speech are synthesized and rendered to obtain an initial dubbing video; The initial quality score of the initial dubbing video is generated by a dual-engine system based on a combined subjective and objective quality assessment. The initial dubbing video is then optimized based on the initial quality score and the video processing mode to obtain the target video after translation and dubbing.
2. The video processing method for translation and dubbing as described in claim 1, characterized in that, The process of performing subtitle-level segment reconstruction on the initial transcribed segment yields multiple segments to be processed, including: Obtain the multidimensional reconstruction constraints constructed based on the timestamp, the role, and the voice activity boundary; The initial transcribed segment is reconstructed at the subtitle level according to the multidimensional reconstruction constraints to obtain the multiple segments to be processed. Specifically, the initial transcription fragments with a length greater than the maximum length threshold are split, and the initial transcription fragments with a length less than the minimum length threshold are merged. Each segment to be processed corresponds to a character.
3. The video processing method for translation and dubbing as described in claim 1, characterized in that, The step of calibrating the plurality of segments to be processed according to the video processing mode to obtain a plurality of calibrated segments includes: When the video processing mode is pipeline mode, high-risk segments among the multiple segments to be processed are identified according to preset rules, the triggerer of the video processing instruction receives the processing instruction for the high-risk segments, and the high-risk segments are calibrated according to the processing instruction to obtain the multiple calibrated segments; or When the video mode is in agent mode, the agent is invoked to identify the high-risk and low-risk segments among the multiple segments to be processed; the agent is used to automatically correct the low-risk segments and record the high-risk segments as candidate problems.
4. The video processing method for translation and dubbing as described in claim 3, characterized in that, The process of performing time and describability corrections on the multiple fragment-level translations yields multiple corrected translations, including: Obtain the translation length of each segment-level translation, and predict the corresponding text-to-speech (TTS) duration based on the translation length of each segment-level translation; For each segment-level translation, calculate the time difference between the TTS duration and the original time window length of the corresponding calibration segment; When the time difference exceeds a first threshold, the segment-level translation is compressed, rewritten in a colloquial style, or rewritten semantically, or a modification prompt is sent to the triggerer; or When the time difference is less than or equal to the first threshold and greater than or equal to the second threshold, the speech rate of the segment-level translation is adjusted, the audio is stretched or compressed, the segment time is fine-tuned, or the audio-visual offset is corrected.
5. The video processing method for translation and dubbing as described in claim 3, characterized in that, The initial quality score generated by the dual-engine system based on a combined subjective and objective quality assessment includes: Obtain an objective indicator evaluation engine based on objective evaluation indicators, and obtain a subjective perception evaluation engine based on subjective perception evaluation indicators. The objective evaluation engine is used to score the quality of each segment in the initial dubbing video, resulting in an objective score for each segment. The initial dubbing video is scored using the objective indicator evaluation engine to obtain an overall objective score for the initial dubbing video. The subjective perception evaluation engine is used to score the quality of each segment in the initial dubbing video, resulting in a subjective score for each segment. The subjective perception evaluation engine is used to score the quality of the initial dubbing video, resulting in an overall subjective score for the initial dubbing video. The objective score for each segment is combined with the corresponding subjective score to obtain a segment-level quality score for each segment. The overall objective score and the overall subjective score of the initial dubbing video are combined to obtain the global quality score of the initial dubbing video.
6. The video processing method for translation and dubbing as described in claim 5, characterized in that, The optimization of the initial dubbing video based on the initial quality score and the video processing mode to obtain the target video after translation and dubbing includes: When the video processing mode is the pipeline mode, the initial dubbing video, the segment-level quality score of each segment, and the global quality score are displayed on the designated display interface; for segments to be optimized whose segment-level quality scores are less than the score threshold, the segments to be optimized are locally modified to regenerate replacement segments; the replacement segments are then used to replace the initial dubbing video to obtain the target video; or When the video mode is the agent mode, the agent is invoked to perform problem localization to obtain the problem localization result. Based on the segment-level quality score of the segment to be optimized, the global quality score, and the problem localization result, a structured repair operation is generated. The structured repair operation is filtered to obtain an executable operation. The executable operation is executed on the segment to be optimized to obtain the optimized segment, and the optimized segment is replaced in the initial dubbing video to obtain the target video.
7. The video processing method for translation and dubbing as described in claim 6, characterized in that, After replacing the optimized segment with the initial dubbing video to obtain the target video, the method further includes: The dual engines are invoked to perform iterations; wherein, during each iteration, a quality score for each generated video is generated sequentially, and the generated video is optimized based on the generated quality score; When the stopping condition is met, the video with the highest quality score is determined as the optimal video, and the optimal video is output. The process includes recording an iteration log, which is used to incrementally optimize the agent.
8. A video processing device for translation and dubbing, characterized in that, The video processing device for translation and dubbing includes: The recognition unit is used to perform automatic speech recognition on the initial video in response to a video processing instruction triggered based on the initial video, and to obtain an initial transcription segment with timestamps, roles and speech activity boundaries; The reconstruction unit is used to perform subtitle-level segment reconstruction on the initial transcription segment to obtain multiple segments to be processed. The generation unit is configured to generate a multimodal context and terminology list based on the initial video, the initial transcription segment, and the multiple segments to be processed; A calibration unit is used to acquire the video processing mode of the initial video, and calibrate the multiple segments to be processed according to the video processing mode to obtain multiple calibration segments; The correction unit is used to perform translation operations on the multiple calibration segments according to the multimodal context and the terminology list to obtain multiple segment-level translations, and to perform duration and speakability corrections on the multiple segment-level translations to obtain multiple corrected translations. The rendering unit is used to generate target speech based on the timestamp, the character, the speech activity boundary and the multimodal context, according to the multiple corrected translations, and to synthesize and render the initial video and the target speech to obtain an initial dubbing video; The optimization unit is used to generate an initial quality score for the initial dubbing video based on a dual-engine system of subjective and objective collaborative quality assessment, and to optimize the initial dubbing video according to the initial quality score and the video processing mode to obtain the target video after translation and dubbing.
9. A computer device, characterized in that, The computer device includes: A memory that stores at least one instruction; and a processor that executes the instructions stored in the memory to implement the video processing method for translation and dubbing as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores at least one instruction, which is executed by a processor in a computer device to implement the video processing method for translation and dubbing as described in any one of claims 1 to 7.