Video processing method and device and electronic equipment

By combining multimodal large language model translation technology with subtitle text and audio frame sequences, the problem of translation accuracy in foreign language videos with low audio quality is solved, achieving higher translation accuracy and video processing effect.

CN121750897APending Publication Date: 2026-03-27VIVO MOBILE COMM CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-05
Publication Date
2026-03-27

AI Technical Summary

Technical Problem

When processing foreign language videos, existing technologies suffer from reduced audio quality, leading to poor translation accuracy and impacting video processing results.

Method used

By using a multimodal large language model to translate subtitle text and audio frame sequences during video playback, translated text in different languages ​​corresponding to the subtitle text is generated.

Benefits of technology

It improves the translation accuracy of electronic devices under low audio quality conditions and enhances video processing performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121750897A_ABST
    Figure CN121750897A_ABST
Patent Text Reader

Abstract

The invention discloses a video processing method and device and electronic equipment, and belongs to the technical field of artificial intelligence, and the method comprises the steps: translating a first audio frame sequence in a first video based on a first subtitle text in a first screen image corresponding to the first audio frame sequence in a process of playing the first video, and obtaining a first subtitle text in a first screen image corresponding to the first audio frame sequence; obtaining a first translated text of the first language; the first language is a language different from a language corresponding to the first caption text; and generating and displaying a second subtitle text based on the first translated text.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of artificial intelligence technology, specifically relating to a video processing method, apparatus, and electronic device. Background Technology

[0002] With the widespread adoption of mobile internet, more and more users are using electronic devices to watch videos in their daily lives. For some foreign language videos on domestic and international video websites, electronic devices can generate translated subtitles for these videos, either offline or online, to facilitate user viewing.

[0003] However, the methods described above typically require high-quality audio to be translated. Lower audio quality directly impacts audio recognition accuracy. For example, noisy environments can lead to inaccurate voice recognition. This results in poor translation accuracy from electronic devices, consequently leading to suboptimal video processing performance. Summary of the Invention

[0004] The purpose of this application is to provide a video processing method, apparatus, and electronic device that can improve the accuracy of translation by electronic devices, thereby improving the video processing effect of electronic devices.

[0005] In a first aspect, embodiments of this application provide a video processing method, which includes: during the playback of a first video, translating the first audio frame sequence in the first video based on the first subtitle text in a first screen image corresponding to the first audio frame sequence to obtain a first translated text in a first language; the first language being a language different from the language corresponding to the first subtitle text; and generating and displaying a second subtitle text based on the first translated text.

[0006] Secondly, embodiments of this application provide a video processing apparatus, which includes: a translation module and a display module; the translation module is used to translate the first audio frame sequence in the first video based on the first subtitle text in the first screen image corresponding to the first audio frame sequence during the playback of the first video, to obtain a first translated text in a first language; the first language is a language different from the language corresponding to the first subtitle text; the display module is used to generate and display a second subtitle text based on the first translated text translated by the translation module.

[0007] Thirdly, embodiments of this application provide an electronic device including a processor and a memory, wherein the memory stores programs or instructions executable on the processor, and the programs or instructions, when executed by the processor, implement the steps of the method described in the first aspect.

[0008] Fourthly, embodiments of this application provide a readable storage medium on which a program or instructions are stored, which, when executed by a processor, implement the steps of the method described in the first aspect.

[0009] Fifthly, embodiments of this application provide a chip, the chip including a processor and a communication interface, the communication interface being coupled to the processor, the processor being used to run programs or instructions to implement the method as described in the first aspect.

[0010] In a sixth aspect, embodiments of this application provide a computer program / program product stored in a storage medium, which is executed by at least one processor to implement the method described in the first aspect.

[0011] In this embodiment, during the playback of the first video, the first audio frame sequence in the first video is translated based on the first subtitle text in the first screen image corresponding to the first audio frame sequence, resulting in a first translated text in a first language; the first language is a language different from the language corresponding to the first subtitle text; based on the first translated text, a second subtitle text is generated and displayed. In this solution, the electronic device can acquire the subtitle text displayed on the screen during real-time playback of the first video, and then supplement the translation of the audio frame sequence corresponding to that subtitle text based on that subtitle text. That is, during the translation of the first audio frame sequence, the electronic device can more accurately determine the content to be translated through the first subtitle text and the first audio frame sequence, i.e., two different types of information. Therefore, even if the audio quality of the audio frame sequence acquired by the electronic device is poor, the electronic device can still achieve accurate translation, thus improving the accuracy of the electronic device's translation and further enhancing the video processing effect of the electronic device. Attached Figure Description

[0012] Figure 1 This is one of the flowcharts of a video processing method provided in the embodiments of this application;

[0013] Figure 2 This is one of the schematic diagrams illustrating an example of acquiring video data provided in this application embodiment;

[0014] Figure 3 This is a second schematic diagram illustrating an example of acquiring video data provided in this application embodiment;

[0015] Figure 4 This is a second flowchart of a video processing method provided in an embodiment of this application;

[0016] Figure 5 This is a flowchart of a model training method provided in an embodiment of this application;

[0017] Figure 6A This is one of the schematic diagrams illustrating a video playback format provided in the embodiments of this application;

[0018] Figure 6B This is a second example of a video playback format provided in the embodiments of this application;

[0019] Figure 6C This is the third example of a video playback format provided in the embodiments of this application;

[0020] Figure 6D This is the fourth example of a video playback format provided in the embodiments of this application;

[0021] Figure 6E This is the fifth example of a video playback format provided in the embodiments of this application;

[0022] Figure 7A This is one of the example diagrams of a video image provided in the embodiments of this application;

[0023] Figure 7B This is a second example of a video image provided in an embodiment of this application;

[0024] Figure 8 This is a flowchart of acquiring video images provided in an embodiment of this application;

[0025] Figure 9A This is the third example of a video image provided in the embodiments of this application;

[0026] Figure 9B This is the fourth example of a video image provided in the embodiments of this application;

[0027] Figure 10 This is a flowchart of extracting subtitle text provided in an embodiment of this application;

[0028] Figure 11 This is a flowchart of a method for detecting subtitle text provided in an embodiment of this application;

[0029] Figure 12 This is a schematic diagram of the structure of a video processing device provided in an embodiment of this application;

[0030] Figure 13 This is one of the hardware structure diagrams of an electronic device provided in the embodiments of this application;

[0031] Figure 14 This is a second schematic diagram of the hardware structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0032] The technical solutions of the embodiments of this application will be clearly described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application are within the scope of protection of this application.

[0033] The terms "first," "second," etc., used in this application's specification are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such terms can be used interchangeably where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first," "second," etc., are generally of the same class, without limiting the number of objects. For example, a first object can be one or more, where "more" means at least two. Furthermore, in the specification, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects are in an "or" relationship.

[0034] The terms "at least one," "at least one," etc., used in this application's specification refer to any one, any two, or a combination of two or more of the included objects. For example, "at least one of a, b, and c" can mean: "a," "b," "c," "a and b," "a and c," "b and c," and "a, b, and c," where a, b, and c can be single or multiple, and multiple means at least two. Similarly, "at least two" refers to two or more, and its meaning is similar to that of "at least one."

[0035] The terminology used in the implementation section of this application is only for explaining specific embodiments of this application and is not intended to limit this application. The terminology involved in the embodiments of this application is explained below.

[0036] Subtitle Translation (ST): Converting dialogue in the source language of a video or audio file into text in the target language and presenting it as subtitles.

[0037] Automatic Speech Recognition (ASR): A technology that uses artificial intelligence to convert human speech signals into digital text, also known as speech-to-text (STT).

[0038] Machine translation (MT) uses computer technology to automatically convert video or audio from source language text into target language text as required by the user. In other words, it automatically converts one natural language text into another, aiming to achieve fast and accurate cross-language information transmission.

[0039] Multimodal Large Language Models (MLLM): Artificial intelligence models that can simultaneously process and understand multiple modalities of data such as text, images, audio, and video, and achieve semantic alignment and collaborative reasoning between different modalities through cross-modal fusion technology.

[0040] Incremental translation is a phased and progressive translation strategy. Its core is to break down the original text into multiple manageable parts, such as paragraphs, sentences, or phrases, and translate them one by one in a certain order. At the same time, the translated content is dynamically adjusted in combination with the context, and finally integrated into a complete translation.

[0041] Automatic speech recognition (ASR) is a technology that converts human speech into text, belonging to the interdisciplinary field of artificial intelligence and speech signal processing. Its core goal is to achieve efficient and accurate speech-to-text conversion by analyzing the acoustic features, language patterns, and semantic information in speech through algorithmic models.

[0042] In ASR (Automatic Recognition Scale), the confidence score is a metric that measures the reliability of the system's recognition results. It is typically expressed numerically, such as a probability value between 0 and 1. It helps users or downstream systems assess the accuracy of the recognition results.

[0043] Object detection (OD) utilizes computer vision techniques to automatically identify and locate multiple target objects in images or videos, while outputting their category labels and bounding box coordinates. The core of object detection is solving two sub-problems: classification, determining whether a target object exists in an image and identifying its category; and localization, precisely labeling the location of the target object using bounding boxes.

[0044] Optical Character Recognition (OCR): This technology uses image algorithms to analyze characters in an image and convert them into editable computer text content.

[0045] Perplexity (PPL) is a core metric used in Natural Language Processing (NLP) and Automatic Speech Recognition (ASR) to measure the performance of a Language Model (LM). It reflects the "degree of confusion" or "predictive uncertainty" of the model regarding the test data. The lower the value, the more accurate and confident the model's predictions are.

[0046] With the widespread adoption of mobile internet, more and more users are using electronic devices to watch videos in their daily lives. Many foreign language videos on domestic and international video websites lack Chinese subtitles, making it difficult for users to accurately understand the video content and severely impacting their viewing experience.

[0047] Therefore, to address these issues, some app developers have introduced subtitle translation features, which translate dialogue and narration in audio and video from one language to another and display them as subtitles on the playback interface. Typically, video websites can generate translated subtitles for their videos offline; however, this method suffers from latency and incomplete coverage.

[0048] Furthermore, addressing the shortcomings of offline subtitle translation solutions, an online subtitle translation solution has been introduced to the market. This solution can acquire the audio playback content from the device in real time and request the translated text from a cloud service. Online subtitle translation typically uses an ASR (Automatic Sound Responsive) approach combined with machine translation. Specifically, online subtitle translation first inputs real-time audio into an online translation model to obtain the corresponding text through ASR, and then uses machine translation to convert the source language text into the target language text.

[0049] Meanwhile, multimodal large language model technology can provide an end-to-end implementation for online subtitle translation. Multimodal large language models can process multimodal data, including audio. For example, when audio data is input into a multimodal large model, it can directly obtain and output the translated text in the target language. Compared to ASR cascaded machine translation, the end-to-end multimodal large model solution for online subtitle translation has the following advantages: lower processing latency; avoidance of error accumulation problems inherent in cascaded solutions; understanding of multimodal inputs; support for session context preservation; and higher accuracy.

[0050] However, for online subtitle translation, whether using ASR cascaded machine translation or an end-to-end multimodal large model approach, high-quality input audio is crucial. Degraded audio quality directly impacts audio recognition accuracy, potentially reducing translation accuracy. For example, in noisy environments, the model may be affected by noise, leading to inaccurate voice recognition; background sounds similar to human voices, such as background music, may be misidentified as the target voice; and in scenarios with multiple speakers simultaneously, interference between different voices can prevent the model from distinguishing their content. Furthermore, audio translation suffers from accuracy issues with certain words, such as names, place names, and technical terms. These factors contribute to the poor accuracy of translation on electronic devices.

[0051] The video processing method provided in this application will be described in detail below with reference to the accompanying drawings, through specific embodiments and application scenarios.

[0052] The video processing method provided in this application can be applied to scenarios involving online video or audio translation. The following examples illustrate the video processing method provided in this application.

[0053] Scenario 1: When a user wants to watch an English video on an electronic device, but the video lacks a Chinese subtitle, the device can retrieve the subtitle text from the screen image corresponding to the currently playing audio frame. Based on this subtitle text, it translates the audio within the same audio frame sequence as the currently playing audio frame to obtain the translated text in the user's desired language, such as Chinese. The subtitle text is identical for each audio frame within the same audio frame sequence. The electronic device can then display this Chinese translated text so the user can understand the content of the English video.

[0054] Scenario 2: When a user is watching an English video on an electronic device, but the translation provided by the video application is incorrect, the electronic device can obtain the subtitle text from the screen image corresponding to the currently playing audio frame. Based on this subtitle text, it can translate audio within the same audio frame sequence as the currently playing audio frame to obtain a more accurate translation. For example, it can accurately distinguish between the main character's speech and the speech of passersby in the background based on the subtitles. The subtitle text corresponding to each audio frame within the same audio frame sequence is identical. The electronic device can then display this more accurate translation, allowing the user to understand the actual content of the main character's speech and avoiding interference from the speech of passersby in the background.

[0055] It should be noted that the above scenarios 1 and 2 are merely exemplary examples of some scenarios that may be applied to the embodiments of this application. In actual implementation, the embodiments of this application can also be applied to more and more possible scenarios involving online translation, and the embodiments of this application are not limited here.

[0056] Based on the above-mentioned scenario applied in the embodiments of this application, the video processing method provided in this application, during the playback of a first video, translates the first audio frame sequence in the first video based on the first subtitle text in the first screen image corresponding to the first audio frame sequence, to obtain a first translated text in a first language; the first language is a language different from the language corresponding to the first subtitle text; and a second subtitle text is generated and displayed based on the first translated text. In this solution, the electronic device can obtain the subtitle text in the video frame displayed on the screen during the real-time playback of the first video, and supplement the translation of the audio frame sequence corresponding to the subtitle text based on the subtitle text. That is, during the translation of the first audio frame sequence, the electronic device can more accurately know the content to be translated through the first subtitle text and the first audio frame sequence, i.e., two different types of information. Therefore, even if the audio quality of the audio frame sequence obtained by the electronic device is poor, the electronic device can still achieve accurate translation, thus improving the accuracy of the electronic device's translation and further improving the video processing effect of the electronic device.

[0057] The video processing method provided in this application is executed by a video processing device, which can be an electronic device, or a functional module or entity within an electronic device. This application does not limit the specific implementation of this method. The following will use an electronic device as an example to illustrate the video processing method provided in this application.

[0058] This application provides a video processing method. Figure 1 A flowchart illustrating a video processing method provided in an embodiment of this application is shown. Figure 1 As shown, the video processing method provided in this application embodiment may include the following steps 201 and 202.

[0059] Step 201: During the playback of the first video, the electronic device translates the first audio frame sequence in the first video based on the first subtitle text in the first screen image corresponding to the first audio frame sequence, and obtains the first translated text in the first language.

[0060] In some embodiments of this application, the first video described above may include subtitle text.

[0061] In some embodiments of this application, the aforementioned subtitle text may be subtitles embedded in the video or external subtitles.

[0062] In some embodiments of this application, if the online subtitle translation function of the electronic device is enabled during the playback of the first video, the electronic device can obtain the aforementioned first subtitle text.

[0063] In some embodiments of this application, the online subtitle translation function can be manually activated by the user, or it can be automatically activated by the electronic device when it detects that the video being viewed by the user is a foreign language video.

[0064] In some embodiments of this application, when the electronic device has the online subtitle translation function enabled, the electronic device can obtain the audio source language of the video and the first language mentioned above.

[0065] In some embodiments of this application, the aforementioned audio source language can be understood as the language used by the audio used in the first video playback.

[0066] In some embodiments of this application, the first language mentioned above can be a language that the user can understand and use proficiently, such as the user's native language.

[0067] In some embodiments of this application, the language of the audio source or the first language may include, but is not limited to, any of the following: Chinese, English, French, Russian, Korean, Japanese, German, Spanish, Arabic, and Italian.

[0068] It should be noted that the language of the audio source mentioned above and the first language mentioned above can be different languages.

[0069] Understandably, since users cannot directly understand the audio source language of the first video, they cannot know the video content, such as the dialogue of the characters. Therefore, users can trigger the online subtitle translation function of their electronic devices, allowing the devices to provide a translated text in the user's preferred language. This translated text then enables the user to accurately understand the video content.

[0070] In some embodiments of this application, during the playback of a first video by an electronic device, the electronic device can automatically determine the audio source language corresponding to the first video based on the first video.

[0071] It should be noted that the language of the audio source mentioned above is different from the first language mentioned above.

[0072] In some embodiments of this application, the electronic device can acquire screen images displayed on its screen interface and audio frames of played audio content at a fixed frame rate. For example, the fixed frame rate can be 30 frames / s, so the electronic device can acquire screen images and audio frames once every approximately 33ms, or 1000 / 30ms.

[0073] In some embodiments of this application, the first screen image corresponding to the first audio frame sequence can be the screen image acquired by the electronic device at the start time of the first audio frame sequence, that is, the screen image acquired by the electronic device at the start time of the first audio frame in the first audio frame sequence.

[0074] For example, such as Figure 2 As shown, the electronic device can acquire a screen image at the beginning of each audio frame, and the audio frame can be a sampled segment of audio output within the current frame time range.

[0075] In some embodiments of this application, the audio frames in the first audio frame sequence can be matched with the first subtitle text. That is, it can be understood that the audio content contained in the audio frames in the first audio frame sequence can be a portion of the first subtitle text.

[0076] In some embodiments of this application, the audio frames in the first audio frame sequence described above may also include background audio, i.e., content that may not be displayed in the subtitle text.

[0077] In some embodiments of this application, the background audio may include, but is not limited to, at least one of the following: ambient background sound, background music, and voice audio of people in the background.

[0078] In some embodiments of this application, when the first audio frame sequence contains content not displayed in the aforementioned subtitle text, during the subsequent translation of the first audio frame sequence by the electronic device based on the first subtitle text, the electronic device can automatically obtain the audio to be translated from the first audio frame sequence based on the first subtitle text, that is, automatically remove the background audio contained in the first audio frame sequence before translating the audio to be translated. This reduces or avoids the interference of background audio on the translation process, thereby improving the accuracy of the translation.

[0079] In some embodiments of this application, before the electronic device translates the first audio frame sequence, the electronic device may preprocess the audio frames contained in the first audio frame sequence to improve the accuracy of subsequent translation by enhancing the audio instructions of the audio to be translated.

[0080] In some embodiments of this application, the above preprocessing may include, but is not limited to, at least one of the following: noise reduction, speech enhancement, and echo cancellation.

[0081] In some embodiments of this application, the aforementioned first screen image can be understood as the image displayed on the screen of an electronic device, that is, it includes all image content displayed on the screen.

[0082] In some embodiments of this application, the first screen image described above includes a video playback area.

[0083] In some embodiments of this application, the subtitle text obtained by the electronic device based on any one of the audio frames in the first audio frame sequence is the first subtitle text. That is, the first subtitle text can be displayed at the beginning of the first audio frame of the first audio sequence and remains unchanged until the end of the audio in the first audio sequence.

[0084] For example, such as Figure 3 As shown, the audio frame sequence a can include four audio frames, such as audio frame i, audio frame i+1, audio frame i+2, and audio frame i+3. Each audio frame can correspond to a screen image, such as screen image i corresponding to audio frame i, screen image i+1 corresponding to audio frame i+1, screen image i+2 corresponding to audio frame i+2, and screen image i+3 corresponding to audio frame i+3. The subtitle text contained in screen images i, i+1, i+2, and i+3 is all subtitle text a.

[0085] Similarly, the audio frame sequence a+1 can include three audio frames, such as audio frame i+4, audio frame i+5, and audio frame i+6. Each audio frame can correspond to a screen image, such as screen image i+4 corresponding to audio frame i+4, screen image i+5 corresponding to audio frame i+5, and screen image i+6 corresponding to audio frame i+6. The subtitle text contained in screen images i+4, i+5, and i+6 is all subtitle text a+1.

[0086] In this embodiment of the application, the first language is a language different from the language corresponding to the first subtitle text.

[0087] In some embodiments of this application, the language of the first subtitle text may be the same as or different from the language of the audio source.

[0088] For example, the first video can be a video in which both the audio and subtitles are in English, that is, the language of the first subtitle text is the same as the language of the audio source; or, the first video can be a video in which the audio is in Korean and the subtitles are in English, that is, the language of the first subtitle text is different from the language of the audio source.

[0089] In some embodiments of this application, before "translating the first audio frame sequence in the first video based on the first subtitle text in the first screen image corresponding to the first audio frame sequence to obtain the first translated text in the first language" in step 201 above, the video processing method provided in this application may also include the following steps A1 and A2.

[0090] Step A1: The electronic device acquires the screen image corresponding to the first audio frame in the first audio frame sequence.

[0091] In some embodiments of this application, the first audio frame may be a video frame in the first video that meets certain conditions.

[0092] In some embodiments of this application, the audio frames that meet the above conditions may include any of the following:

[0093] The screen image corresponding to the audio frame preceding the audio frame does not contain subtitle text;

[0094] The subtitle text in the screen image corresponding to the audio frame is different from the subtitle text in the screen image corresponding to the previous audio frame.

[0095] It is understandable that the first audio frame mentioned above does not belong to the same audio sequence as the previous audio frame; that is, the first audio frame mentioned above is the first audio frame of the first audio sequence.

[0096] In some embodiments of this application, the first audio frame sequence is an audio frame sequence in the first video that is the first audio frame and matches the first subtitle text.

[0097] Step A2: The electronic device obtains the first subtitle text based on the screen image corresponding to the first audio frame.

[0098] In some embodiments of this application, the first subtitle text may be subtitle text acquired in real time by an electronic device based on a first audio frame.

[0099] It should be noted that the methods for how an electronic device obtains the first subtitle text based on the screen image corresponding to the first audio frame, and the methods for how an electronic device determines which audio frames are included in the first audio frame sequence, can be specifically referred to in steps 301 and 302 below, and will not be repeated here in the embodiments of this application.

[0100] In some embodiments of this application, combined with Figure 1 ,like Figure 4 As shown, step 201 can be implemented through steps 201a and 201b.

[0101] Step 201a: During the playback of the first video, the electronic device acquires each audio frame and inputs the audio frame into the translation model corresponding to the first language. If an audio frame is an audio frame that meets the conditions, the subtitle text corresponding to that audio frame is input into the translation model.

[0102] In some embodiments of this application, during the playback of the first video by the electronic device, each time an audio frame is acquired, the electronic device can input that audio frame into the translation model. If the audio frame meets the condition that it is the first audio frame in an audio frame sequence, the electronic device can input that audio frame and its corresponding subtitle text into the translation model. If the audio frame does not meet the condition that it is not the first audio frame in an audio frame sequence, since the translation model has already received the subtitle text corresponding to that audio frame sequence, the electronic device can input the audio frame into the translation model, thus reducing the repeated input of the same subtitle text. Simultaneously, the electronic device expects the translation model to refer to the previously provided subtitle text during the translation of the audio frame sequence to improve the accuracy of the audio translation task.

[0103] In some embodiments of this application, during the playback of the first video by the electronic device, if an audio frame (e.g., audio frame a) satisfies a certain condition, and the electronic device then acquires another audio frame (e.g., audio frame a+n) that satisfies the same condition, the electronic device can determine at least one audio frame between the two audio frames as an audio frame belonging to the same audio frame sequence as audio frame a. That is, the electronic device can use audio frames a to a+(n-1) as the audio frame sequence corresponding to audio frame a.

[0104] In some embodiments of this application, the above translation model can be a multimodal large model, that is, the translation model can process and understand subtitle text and audio frames at the same time to realize the translation of audio frames.

[0105] In some embodiments of this application, the aforementioned multimodal large model can be a model obtained by fine-tuning an existing multimodal large model, and the model has the ability to translate with reference text when audio and reference text such as subtitles are mixed input.

[0106] In some embodiments of this application, the above translation model can use a pre-trained model that covers the range of selectable languages ​​and has good audio translation capabilities within that range as the base model, and construct a model trained with corresponding fine-tuned data for that language range.

[0107] It should be noted that, due to limitations in the richness and difficulty of acquiring the corpus, in order to maintain a consistent level across all languages, it is necessary to limit the range of selectable languages ​​and set the source language and target language at the beginning of the subtitle translation.

[0108] In some embodiments of this application, during the training process of the above translation model, each training sample used for fine-tuning training can correspond to an audio segment with complete semantics, which includes audio data information, task instructions, and a set of audio frame data arranged in time sequence. The set of audio frame data can include the start and end time of each audio frame, the expected output translated text, and the reference text at that moment.

[0109] For example, in a training sample where the corresponding audio segment is "I would like to book the conference room for 3 PM tomorrow afternoon," the source language is English (en), and the translation language is Chinese (zh), the electronic device can set task instructions for the translation model, such as "Please translate the input audio in English into text in Chinese, with reference to the most recently provided reference text during the translation process." This allows the translation model to understand the task to be performed. The relevant code is shown below:

[0110] {

[0111] "id": "...", # Sample ID

[0112] "audio": {...}, # Audio data information

[0113] "source_language": "en", #audio source language

[0114] "target_language": "zh", #translation language

[0115] "instruction": "Please translate the input audio in English into textin Chinese, with reference to the most recently provided reference textduring the translation process." #TASK INSTRUCTION

[0116] "segments": [

[0117] {"start": t0, "end": t1, "text": "I", "ref_text": "I would like to book the conference room for 3 PM tomorrow afternoon."}, #Audio frame time range, translated text, reference text

[0118] {"start": t2, "end": t3, "text": "want", "ref_text": ""},

[0119] {"start": t4, "end": t5, "text": "Reservation", "ref_text": ""},

[0120] {"start": t6, "end": t7, "text": "Meeting Room", "ref_text": ""},

[0121] {"start": t8, "end": t9, "text": ",Tomorrow at 3 PM.", "ref_text": ""} ]

[0123] }

[0124] The translation model acquires the first audio frame from time t0 to t1, and the translated text corresponding to the first audio frame is "I"; the translation model acquires the second audio frame from time t2 to t3, and the translated text corresponding to the second audio frame is "want"; the translation model acquires the third audio frame from time t4 to t5, and the translated text corresponding to the third audio frame is "book"; the translation model acquires the fourth audio frame from time t6 to t7, and the translated text corresponding to the fourth audio frame is "meeting room"; the translation model acquires the fifth audio frame from time t8 to t9, and the translated text corresponding to the fifth audio frame is "tomorrow afternoon at 3 pm".

[0125] It should be noted that, to improve the translation accuracy of the aforementioned translation model, during the training process, rich, high-quality triplet fine-tuning data can be constructed for the source audio language, target translation language, and reference text language. Simultaneously, a certain proportion of noisy data can be added to the training data to enhance the model's generalization ability and robustness to different scenarios. For example, scenarios with poor audio quality, missing reference text, or character errors. The specific model training process can be as follows: Figure 5 As shown.

[0126] Step 201b: The electronic device translates the first audio frame sequence based on the first subtitle text using a translation model to obtain the first translated text.

[0127] In some embodiments of this application, the electronic device can sample real-time audio frames, i.e. audio data, from the played audio at a fixed frequency and input them, along with the extracted subtitle text, into the translation model to obtain real-time subtitle translation text output.

[0128] In some embodiments of this application, step 201b above can be specifically implemented by step 201b1 below.

[0129] Step 201b1: The electronic device uses a translation model to incrementally translate the audio frames in the first audio frame sequence based on the first subtitle text, thereby obtaining the first translated text.

[0130] In some embodiments of this application, when the translation model acquires audio frames in the first audio frame sequence, the electronic device can perform an audio-to-text operation on the audio frames to obtain the audio text of the audio frames. Then, the translation model can determine whether the text composed of at least one audio text from the acquired first audio frame sequence is semantically complete. If semantically complete, it translates the audio frames corresponding to at least one audio text based on the first subtitle text and outputs the corresponding translated text. Further, if the translation model continues to receive audio frames, it can translate based on the newly received audio frames and update the translated text for display. Thus, after the translation model has translated all the audio frames in the first audio frame sequence, the aforementioned first translated text is obtained.

[0131] In some embodiments of this application, the method by which an electronic device determines whether text is semantically complete may include, but is not limited to, at least one of the following: grammatical structure, speech silence pauses, and ASR confidence.

[0132] For example, if the translation model obtains the subject, predicate, and object, it can determine that the text is semantically complete; or, if the translation model determines that the audio text is not finished, but there is a pause in the middle of the audio, such as a 0.5s pause, the translation model can determine that the text is semantically complete; or, the translation model can calculate a confidence score for the text corresponding to each audio frame, and determine whether the text is semantically complete based on the confidence score.

[0133] Exemplarily, assume that the first subtitle text is "I would like to book the conference room for 3 PM tomorrow afternoon", and the first audio frame sequence includes 5 audio frames, such as audio frame 1 "I would like", audio frame 2 "to book", audio frame 3 "the conference room", audio frame 4 "for 3 PM", and audio frame 5 "tomorrow afternoon". The translation model can first obtain the subtitle text and audio frame 1 to determine whether the semantics of audio frame 1 "I would like" is complete; when the semantics of audio frame 1 is incomplete and the translation model receives audio frame 2, the translation model can determine the semantics of audio frame 1 and audio frame 2, that is, whether the semantics of "I would like to book" is complete; if the translation model determines that the semantics is complete, the translation model can translate the "I would like to book" based on the subtitle text, such as obtaining translation text 1 "I want to book" and displaying it on the screen in the form of subtitles; then, when the translation model receives audio frame 3 "the conference room", incremental translation can be performed based on translation text 1 to obtain translation text 2 "I want to book the conference room" and update the display on the screen in the form of subtitles; then, when the translation model receives audio frame 4 "for 3 PM", incremental translation can be performed based on translation text 2 to obtain translation text 3 "I want to book the conference room at 3 PM" and update the display on the screen in the form of subtitles; finally, when the translation model receives audio frame 5 "tomorrow afternoon", incremental translation can be performed based on translation text 3 to obtain translation text 4 "I want to book the conference room at 3 PM tomorrow afternoon" and update the display on the screen in the form of subtitles.

[0134] In this way, the electronic device can use the translation model to perform incremental translation on the audio frames in the first audio frame sequence based on the first subtitle text, and obtain the first translation text. That is, the electronic device can directly display the corresponding translation on the screen when receiving an audio frame with complete semantics, thereby reducing the latency caused by performing translation again when the translation model receives a complete sentence. In this way, the convenience of audio translation is improved.

[0135] In this embodiment, the electronic device can input subtitle text and audio frames into the translation model to more accurately determine the content to be translated through two different types of information. Thus, even if the audio quality is poor, accurate translation can still be achieved, thereby improving the accuracy of the electronic device's translation and further enhancing the video processing effect of the electronic device.

[0136] Step 202: The electronic device generates and displays the second subtitle text based on the first translated text.

[0137] In some embodiments of this application, the display format of the aforementioned second subtitle text can be actively set by the user or set by default by the electronic device.

[0138] In some embodiments of this application, the above-mentioned display format may include, but is not limited to, at least one of the following: display font size, display font, display color, and display position.

[0139] In some embodiments of this application, while the electronic device displays the aforementioned second subtitle text, the electronic device can also generate audio outputs with different privacy settings based on the second subtitle text to achieve an effect similar to simultaneous interpretation.

[0140] In the video processing method provided in this application embodiment, the electronic device can acquire subtitle text from the video frame displayed on the screen during real-time playback of the first video, and then supplement the translation of the audio frame sequence corresponding to the subtitle text based on the subtitle text. That is, during the translation of the first audio frame sequence, the electronic device can more accurately determine the content to be translated through the first subtitle text and the first audio frame sequence, i.e., two different types of information. Therefore, even if the audio quality of the audio frame sequence acquired by the electronic device is poor, the electronic device can still achieve accurate translation, thus improving the accuracy of the electronic device's translation and further enhancing the video processing effect of the electronic device.

[0141] In some embodiments of this application, the video processing method provided in this application may further include the following steps 301 and 302.

[0142] Step 301: The electronic device acquires the video image within the video playback area of ​​the first screen image.

[0143] In some embodiments of this application, the aforementioned first screen image can be understood as the image displayed on the screen of an electronic device, that is, it includes all image content displayed on the screen.

[0144] In some embodiments of this application, the display form of the video playback area in the first screen image may include, but is not limited to, any of the following: portrait full-screen playback, portrait partial playback, portrait small window playback, and landscape full-screen playback.

[0145] For example, the above-mentioned vertical full-screen playback format can be as follows: Figure 6A The above-mentioned vertical screen partial playback format can be as shown in the figure; Figure 6B The above-mentioned vertical screen small window playback format can be as shown in the figure; Figure 6C or Figure 6D The above-mentioned landscape full-screen playback format can be as shown in the image; Figure 6E As shown in the figure.

[0146] In some embodiments of this application, in order to obtain the video image within the video playback area, i.e., to exclude interference from other elements in the screen image, the electronic device needs to detect the actual video playback area within the interface.

[0147] In some embodiments of this application, the electronic device can detect the actual video playback area within the interface using a detection model in computer vision. For example, the detection model can be the YOLO model.

[0148] For example, taking a mobile phone as an example, in order to collect data on different forms of video playback areas displayed on the mobile phone, the actual video playback areas within the screen image can be manually labeled, and a YOLO model trained based on a large amount of manually labeled data from real interfaces can achieve high accuracy.

[0149] It should be noted that the actual video playback area mentioned above refers to the area occupied by the actual video content, excluding non-video content such as black borders.

[0150] In some embodiments of this application, the coordinate range of the video playback area can be represented by a rectangle, such as [x1, y1, x2, y2]. Here, (x1, y1) can be the coordinates of the upper left corner of the rectangular area, and (x2, y2) can be the coordinates of the lower right corner of the rectangular area. The coordinates can be represented in pixels.

[0151] In some embodiments of this application, when the electronic device detects the aforementioned video playback area, the electronic device can crop the video playback area in the first screen image to obtain a video image within the video playback area. Thus, the electronic device can determine the video layout type based on the size of the cropped area.

[0152] For example, in the rectangle corresponding to the cropping area, if the width is greater than the height, the video playback interface displayed by the electronic device is in landscape mode; if the height is greater than the width, the video playback interface displayed by the electronic device is in portrait mode.

[0153] In some embodiments of this application, when an electronic device acquires a cropped video image, the electronic device can adjust the cropped video image to a uniform aspect ratio and image resolution. For example, the aspect ratio of a horizontal video image can be adjusted to 16:9; the aspect ratio of a vertical video image can be adjusted to 9:16.

[0154] For example, when the phone displays something like... Figure 6A In the case of interface 11 shown, the mobile phone can detect whether a video playback interface exists in interface 11. Then, if the interface 11 contains a video playback area, it can crop the interface 11 to obtain the actual video playback area, and further adjust the aspect ratio of the cropped video image to obtain the desired result. Figure 7A The video image shown has an aspect ratio of 9:16 31. Similarly, on a mobile phone display, as shown... Figure 6E In the case of interface 15 shown, the phone can obtain the following: Figure 7B The video image shown has an aspect ratio of 16:9. The steps performed by the mobile phone can be as follows: Figure 8 As shown.

[0155] Step 302: The electronic device extracts the text from the effective subtitle region in the video image, and obtains the first subtitle text based on the text.

[0156] In some embodiments of this application, the electronic device may directly use the text in the valid subtitle area as the first subtitle text. Alternatively, the electronic device may detect the category labels corresponding to all text in the video image, and, if the category labels are determined, combine them with the valid subtitle area to determine the first subtitle text.

[0157] In some embodiments of this application, electronic devices may use OCR to extract text from valid caption areas.

[0158] For example, when the video image is in landscape mode, the effective subtitle area can be the bottom 1 / 3 of the height; when the video image is in portrait mode, the effective subtitle area can be the bottom 1 / 2 of the height.

[0159] It should be noted that the position of the above-mentioned effective subtitle area can be determined according to actual needs, and this embodiment of the application does not limit it here.

[0160] In some embodiments of this application, electronic devices can use corresponding detection models for landscape and portrait videos respectively, to take into account differences in layout and displayed content.

[0161] In some embodiments of this application, after the electronic device-based detection model detects the video image, the electronic device can use different category labels to distinguish between subtitles and non-subtitle content. These different category labels may include, but are not limited to, at least one of the following: subtitle text category label, bullet screen text category label, and other text category labels.

[0162] For example, electronic devices can detect, such as, using a detection model. Figure 7A The video image 31 shown is used to determine the category labels corresponding to all the text contained in the video image 31, such as Figure 9A As shown, the category labels corresponding to the text contained in video image 33 are: the subtitle text category label "Good morning, students"; the bullet screen text category labels "4 minutes ago!", "Class is starting", "!!!!!!", "Six minutes of enthusiasm"; and other text category labels "10 people are watching", "User A", "Following", "536,000 followers", "Required Course 1, First Section", "57,900 views", "43,000", "548", "7,697", and "687". Then, the electronic device can combine the effective subtitle area of ​​the vertical version, that is, the bottom 1 / 2 height, to determine "Good morning, students" as the above-mentioned first subtitle text.

[0163] Similarly, electronic devices can detect, for example, using detection models. Figure 7B The video image 32 shown is used to determine the category labels corresponding to all the text contained in the video image 32, such as Figure 9B As shown, the category labels corresponding to the text contained in video image 34 are: the subtitle text category label "Good morning, teacher"; the bullet screen text category labels "4 minutes ago!", "Sixteen minutes ago", "!!!!"; and other text category labels "10 people are watching". Then, the electronic device can combine the effective subtitle area of ​​the horizontal version, that is, the bottom 1 / 3 height, to determine "Good morning, teacher" as the aforementioned first subtitle text.

[0164] In some embodiments of this application, electronic devices can distinguish different category labels by using different marking formats.

[0165] It should be noted that the above marking format can be determined according to actual needs, and this application embodiment does not limit it. For example, marking can be done using different colors, or marking can be done using lines of different thicknesses.

[0166] For example, in Figure 9A and Figure 9BIn the text, thicker solid lines are used to mark subtitle text, thinner solid lines are used to mark bullet screen text, and dashed lines are used to mark other text.

[0167] It should be noted that the steps performed by the electronic device in the above example can be as follows: Figure 10 As shown.

[0168] In some embodiments of this application, the "obtaining the first subtitle text based on the text" in step 302 above can be specifically accomplished through the following steps 302a and 302b.

[0169] Step 302a: The electronic device obtains the perplexity value corresponding to the text through the character recognition model.

[0170] In this embodiment of the application, the perplexity value mentioned above can be used to represent the reasonableness of the text.

[0171] Step 302b: When the confusion value is within a preset range, the electronic device edits the text to obtain the first subtitle text.

[0172] In this embodiment of the application, the above editing process may include at least one of the following: error correction processing and completion processing.

[0173] In some embodiments of this application, the electronic device may normalize the text before calculating the perplexity value of the text, so as to calculate the perplexity value based on the normalized text.

[0174] In some embodiments of this application, the above-mentioned normalization process may include, but is not limited to, at least one of the following: removing special symbols and garbled characters, removing invisible characters at the beginning and end, and compressing consecutive spaces or newlines in the text into a single character. The invisible characters at the beginning and end may be spaces, newlines, tabs, etc.

[0175] In some embodiments of this application, the electronic device may perform language detection on the text before calculating the perplexity value of the text, so as to improve the accuracy of the perplexity value calculated subsequently.

[0176] In some embodiments of this application, the electronic device can use a statistical probability language model to assess whether the text conforms to the distribution characteristics of the training corpus by calculating the generation probability of the text sequence. This probability value is inversely proportional to the perplexity.

[0177] It should be noted that the above language model can be selected according to the language of the subtitle text, such as a small-parameter language model trained on the language data corresponding to the subtitle text.

[0178] In some embodiments of this application, the aforementioned preset range can be used to indicate that the text contains errors but are not serious, i.e., the text itself. Therefore, when the perplexity value is within the preset range, the electronic device can edit the text to use the corrected text as the aforementioned first subtitle text.

[0179] It should be noted that the lower threshold of the aforementioned preset range can be used to filter correct results, while the higher threshold can be used to filter results with serious errors. That is, if the perplexity value of a text is below the lower threshold, the electronic device can confirm that the text is error-free and directly identify it as subtitle text; if the perplexity value of a text is above the higher threshold, the electronic device can discard the text, skipping it in the current frame and continuing detection in the next frame. In other words, as... Figure 11 The process is shown below.

[0180] In some embodiments of this application, electronic devices can use the character recognition model described above to edit text with a perplexity value within a preset range.

[0181] Understandably, when character recognition errors occur in the text, the model can correct them; when characters are missed in the text, the model can complete them.

[0182] For example, if the text extracted by the electronic device based on the video image is "I ate a flat apple today", the above character recognition model can identify, in combination with the context, that replacing "flat" with "apple" will maximize the overall probability of the text sequence. Thus, the character recognition model can correct the text to "I ate an apple today", and the electronic device can use the corrected text as the first subtitle text.

[0183] For example, if the text extracted by the electronic device based on the video image is "I ate an apple today", the character recognition model can predict that the word with the highest probability after "today" is "day" and thus complete the text. Therefore, the character recognition model can complete the text as "I ate an apple today", and the electronic device can use the completed text as the first subtitle text.

[0184] In this embodiment, after the electronic device extracts text based on video images, the electronic device can detect the extracted text to determine whether there are any errors in the extraction process. If errors are found, the extracted text can be adjusted, and the adjusted text can be used as subtitle text. This improves the accuracy of subsequent audio translation based on the subtitle text.

[0185] In some embodiments of this application, during the playback of the first video, for each audio frame and the screen image corresponding to that audio frame acquired, the electronic device needs to extract subtitles from the screen image corresponding to each audio frame in order to obtain the subtitle text corresponding to each screen image.

[0186] In some embodiments of this application, when an electronic device extracts subtitle text based on the latest received screen image, the electronic device can compare the newly extracted subtitle text with the subtitle text extracted from the previous screen image to determine whether the audio frame of the latest received screen image and the audio frame of the previous screen image are the same audio frame sequence.

[0187] For example, when an electronic device acquires audio frame a and screen image a, it can compare the subtitle text extracted from screen image a with the subtitle text extracted from screen image a-1. If the electronic device determines that the two subtitle texts are the same, it can determine that audio frame a and audio frame a-1 belong to the same audio frame sequence. Therefore, the electronic device can input audio frame a into the subsequent translation model to perform the corresponding translation steps. If the electronic device determines that the two subtitle texts are different, it can determine that audio frame a and audio frame a-1 do not belong to the same audio frame sequence, that is, audio frame a is the first audio frame of a new audio frame sequence. Therefore, the electronic device can input audio frame a and the subtitle text extracted from screen image a into the subsequent translation model to perform the corresponding translation steps.

[0188] In some embodiments of this application, since some characters may be incorrectly or missed in the OCR recognition results, the electronic device can use text similarity to determine the identity of the two subtitle texts during the comparison process. This means that if the similarity between two texts is greater than a preset similarity threshold, the two subtitle texts are considered identical. For example, the preset similarity threshold can be any one of 0.6, 0.7, or 0.8.

[0189] It should be noted that the above-mentioned preset similarity threshold can be determined according to actual needs, and this application embodiment does not limit it here.

[0190] For example, suppose subtitle text 1 is "Machine learning requires a lot of data." and subtitle text 2 is "Machine learning needs a lot of data." The electronic device can first perform preprocessing such as de-symbolization and word segmentation on the two subtitle texts respectively to obtain word segmentation set 1 {"machine", "learning", "need", "a lot", "data"} corresponding to subtitle text 1, and word segmentation set 2 {"machine", "learning", "need", "a lot", "data"} corresponding to subtitle text 2. Then, the electronic device can calculate the similarity between subtitle text 1 and subtitle text 2 based on the intersection {"machine", "learning", "a lot", "data"} and the union {"machine", "learning", "need", "need", "a lot", "data"} of word segmentation set 1 and word segmentation set 2, such as 4 / 6. Since this similarity is greater than the preset similarity threshold of 0.6, subtitle text 1 and subtitle text 2 are determined to be the same.

[0191] In this embodiment, the electronic device can extract subtitle text based on real-time acquired screen images and input the extracted subtitle text into the subsequent translation model. In other words, during the translation of the first audio frame sequence, the electronic device can obtain the content to be translated more accurately through two different types of information. Thus, even if the audio quality of the audio frame sequence acquired by the electronic device is poor, the electronic device can still achieve accurate translation. This improves the accuracy of the electronic device's translation and further enhances the video processing effect of the electronic device.

[0192] In view of the various scenarios in which the embodiments of this application can be applied, and in conjunction with the various implementation schemes of the embodiments of this application described above, specific examples are given below to illustrate the implementation process of the embodiments of this application in various scenarios.

[0193] If a user wants to watch an English video on their electronic device, but the video does not have a Chinese translation, the user can manually activate the online translation function on their electronic device and set the audio source language, such as "English," and the language of the subtitles they want to view, such as "Chinese."

[0194] The electronic device can obtain the subtitle text 'a' from the screen image 'a' corresponding to the currently playing audio frame 'a'. Then, the electronic device can compare the subtitle text 'a' extracted from screen image 'a' with the subtitle text 'a-1' extracted from screen image 'a-1'. If the electronic device determines that the two subtitle texts are the same, it can determine that audio frame 'a' and audio frame 'a-1' belong to the same audio frame sequence. Therefore, the electronic device can input audio frame 'a' into the subsequent translation model to perform the corresponding translation steps. If the electronic device determines that the two subtitle texts are different, it can determine that audio frame 'a' and audio frame 'a-1' do not belong to the same audio frame sequence, that is, audio frame 'a' is the first audio frame of a new audio frame sequence. Therefore, the electronic device can input audio frame 'a' and the subtitle text 'a' extracted from screen image 'a' into the subsequent translation model to perform the corresponding translation steps.

[0195] Specifically, assume that the subtitle text is "I would like to book the conference room for 3PM tomorrow afternoon", and the audio frame sequence corresponding to the audio frame a includes 5 audio frames, such as audio frame 1 "I would like", audio frame 2 "to book", audio frame 3 "the conference room", audio frame 4 "for 3 PM", and audio frame 5 "tomorrow afternoon". The translation model can first obtain the subtitle text and audio frame 1 to determine whether the semantics of audio frame 1 "I would like" are complete; when the semantics of audio frame 1 are incomplete and the translation model receives audio frame 2, the translation model can determine the semantics of audio frames 1 and 2, that is, whether the semantics of "I would like to book" are complete; if the translation model determines that the semantics are complete, the translation model can translate the "I would like to book" based on the subtitle text, such as obtaining translation text 1 "I want to book" and displaying it on the screen in the form of subtitles; then, when the translation model receives audio frame 3 "the conference room", it can perform incremental translation based on translation text 1 to obtain translation text 2 "I want to book the conference room" and update and display it on the screen in the form of subtitles; then, when the translation model receives audio frame 4 "for 3 PM", it can perform incremental translation based on translation text 2 to obtain translation text 3 "I want to book the conference room at 3 PM" and update and display it on the screen in the form of subtitles; finally, when the translation model receives audio frame 5 "tomorrow afternoon", it can perform incremental translation based on translation text 3 to obtain translation text 4 "I want to book the conference room at 3 PM tomorrow afternoon" and update and display it on the screen in the form of subtitles.

[0196] It should be noted that each of the above method embodiments, or various possible implementation manners in each method embodiment, can be executed independently, or, on the premise of no contradiction, can also be executed in combination with each other, which can be specifically determined according to actual usage requirements, and the embodiments of the present application do not limit this.

[0197] It should be noted that for the video processing method provided by the embodiments of the present application, the execution subject can be a video processing device. In the embodiments of the present application, taking the video processing device executing the video processing method as an example, the video processing device provided by the embodiments of the present application is described.

[0198] Figure 12 Shows a possible structural schematic diagram of the video processing device involved in the embodiments of the present application. As Figure 12 As shown, the video processing device 70 may include: a translation module 71 and a display module 72;

[0199] The translation module 71 is used to translate the first audio frame sequence in the first video based on the first subtitle text in the first screen image corresponding to the first audio frame sequence during the playback of the first video, to obtain the first translated text in a first language; the first language is a language different from the language corresponding to the first subtitle text.

[0200] Display module 72 is used to generate and display second subtitle text based on the first translated text translated by translation module 71.

[0201] In one possible implementation, the video processing apparatus 70 provided in this application embodiment may further include: an acquisition module; the acquisition module is configured to acquire a screen image corresponding to the first audio frame in the first audio frame sequence before translating the first audio frame sequence in the first video based on the first subtitle text in the first screen image corresponding to the first audio frame sequence to obtain the first translated text in the first language; and acquire the first subtitle text based on the screen image corresponding to the first audio frame; wherein the first audio frame is a video frame in the first video that meets certain conditions; the audio frame that meets certain conditions includes any one of the following: the screen image corresponding to the previous audio frame of the audio frame does not contain subtitle text, the subtitle text in the screen image corresponding to the audio frame is different from the subtitle text in the screen image corresponding to the previous audio frame of the audio frame; the first audio frame sequence is an audio frame sequence in the first video that starts with the first audio frame and matches the first subtitle text.

[0202] In one possible implementation, the video processing apparatus 70 provided in this application embodiment may further include: an acquisition module and an extraction module; the acquisition module is used to acquire a video image within a video playback area in a first screen image; the extraction module is used to extract text from a valid subtitle area in the video image acquired by the acquisition module, and obtain a first subtitle text based on the text.

[0203] In one possible implementation, the translation module 71 is specifically used to: during the playback of the first video, for each audio frame acquired, input the audio frame into the translation model corresponding to the first language, and if an audio frame is an audio frame that meets the conditions, input the subtitle text corresponding to the audio frame into the translation model; and through the translation model, translate the first audio frame sequence based on the first subtitle text to obtain the first translated text.

[0204] In one possible implementation, the translation module 71 is specifically used to incrementally translate the audio frames in the first audio frame sequence based on the first subtitle text using a translation model to obtain the first translated text.

[0205] In the video processing apparatus provided in this application embodiment, the video processing apparatus can acquire subtitle text from the video frame displayed on the screen during real-time playback of a first video, and then supplement the translation of the audio frame sequence corresponding to the subtitle text based on the subtitle text. That is, during the translation of the first audio frame sequence, the video processing apparatus can more accurately determine the content to be translated using both the first subtitle text and the first audio frame sequence—two different types of information. Therefore, even if the audio quality of the audio frame sequence acquired by the video processing apparatus is poor, the video processing apparatus can still achieve accurate translation. This improves the accuracy of the video processing apparatus's translation and further enhances the video processing effect.

[0206] The video processing device in this application embodiment can be an electronic device or a component within an electronic device, such as an integrated circuit or a chip. The electronic device can be a terminal or other devices besides a terminal. For example, the electronic device can be a mobile phone, tablet computer, laptop computer, PDA, in-vehicle electronic device, mobile internet device (MID), augmented reality (AR) / virtual reality (VR) device, robot, wearable device, ultra-mobile personal computer (UMPC), netbook, or personal digital assistant (PDA), etc. It can also be a server, network attached storage (NAS), personal computer (PC), television set (TV), ATM, or self-service machine, etc. This application embodiment does not specifically limit the device.

[0207] The video processing device in this application embodiment can be a device with an operating system. This operating system can be Android, iOS, or other possible operating systems; this application embodiment does not specifically limit the specific operating system used.

[0208] The video processing apparatus provided in this application embodiment can implement the various processes implemented in the above method embodiments, and will not be described again here to avoid repetition.

[0209] Optionally, such as Figure 13 As shown, this application embodiment also provides an electronic device 90, including a processor 91 and a memory 92. The memory 92 stores a program or instructions that can run on the processor 91. When the program or instructions are executed by the processor 91, they implement the various steps of the above-described video processing method embodiment and can achieve the same technical effect. To avoid repetition, they will not be described again here.

[0210] It should be noted that the electronic devices in the embodiments of this application include the mobile electronic devices and non-mobile electronic devices described above.

[0211] Figure 14 A schematic diagram of the hardware structure of an electronic device to implement an embodiment of this application.

[0212] The electronic device 100 includes, but is not limited to, components such as: radio frequency unit 101, network module 102, audio output unit 103, input unit 104, sensor 105, display unit 106, user input unit 107, interface unit 108, memory 109, and processor 110.

[0213] Those skilled in the art will understand that the electronic device 100 may also include a power supply (such as a battery) for supplying power to various components. The power supply may be logically connected to the processor 110 through a power management system, thereby enabling functions such as managing charging, discharging, and power consumption through the power management system. Figure 14 The electronic device structure shown does not constitute a limitation on the electronic device. The electronic device may include more or fewer components than shown, or combine certain components, or have different component arrangements, which will not be elaborated here.

[0214] The processor 110 is used to translate the first audio frame sequence in the first video based on the first subtitle text in the first screen image corresponding to the first audio frame sequence during the playback of the first video, to obtain the first translated text in a first language; the first language is a language different from the language corresponding to the first subtitle text.

[0215] Display unit 106 is used to generate and display second subtitle text based on the first translated text.

[0216] Optionally, the processor 110 is further configured to, before translating the first audio frame sequence in the first video based on the first subtitle text in the first screen image corresponding to the first audio frame sequence to obtain the first translated text in the first language, acquire the screen image corresponding to the first audio frame in the first audio frame sequence; and acquire the first subtitle text based on the screen image corresponding to the first audio frame; wherein the first audio frame is a video frame in the first video that meets certain conditions; the audio frame that meets certain conditions includes any of the following: the screen image corresponding to the previous audio frame of the audio frame does not contain subtitle text, the subtitle text in the screen image corresponding to the audio frame is different from the subtitle text in the screen image corresponding to the previous audio frame of the audio frame; the first audio frame sequence is an audio frame sequence in the first video that starts with the first audio frame and matches the first subtitle text.

[0217] Optionally, the processor 110 is further configured to acquire a video image within the video playback area of ​​the first screen image; and extract text from the valid subtitle area in the video image, and obtain the first subtitle text based on the text.

[0218] Optionally, the processor 110 is specifically configured to, during the playback of the first video, input each audio frame acquired into the translation model corresponding to the first language, and, if an audio frame is an audio frame that meets certain conditions, input the subtitle text corresponding to that audio frame into the translation model; and, through the translation model, translate the first audio frame sequence based on the first subtitle text to obtain the first translated text.

[0219] Optionally, the processor 110 is specifically used to perform incremental translation of audio frames in the first audio frame sequence based on the first subtitle text using a translation model to obtain the first translated text.

[0220] In the electronic device provided in this application embodiment, the electronic device can acquire subtitle text from the video frame displayed on the screen during real-time playback of the first video, and then supplement the translation of the audio frame sequence corresponding to the subtitle text based on the subtitle text. That is, during the translation of the first audio frame sequence, the electronic device can more accurately determine the content to be translated through the first subtitle text and the first audio frame sequence, i.e., two different types of information. Therefore, even if the audio quality of the acquired audio frame sequence is poor, the electronic device can still achieve accurate translation, thus improving the accuracy of the electronic device's translation and further enhancing the video processing effect of the electronic device.

[0221] The electronic device provided in this application embodiment can implement the various processes implemented in the above method embodiments and achieve the same technical effect. To avoid repetition, it will not be described again here.

[0222] For details on the beneficial effects of the various implementation methods in this embodiment, please refer to the beneficial effects of the corresponding implementation methods in the above method embodiments. To avoid repetition, these will not be repeated here.

[0223] It should be understood that, in this embodiment, the input unit 104 may include a graphics processing unit (GPU) 1041 and a microphone 1042. The GPU 1041 processes image data of still images or videos obtained by an image capture device (such as a camera) in video capture mode or image capture mode. The display unit 106 may include a display panel 1061, which may be configured in the form of a liquid crystal display, an organic light-emitting diode, or the like. The user input unit 107 includes at least one of a touch panel 1071 and other input devices 1072. The touch panel 1071 is also called a touch screen. The touch panel 1071 may include a touch detection device and a touch controller. Other input devices 1072 may include, but are not limited to, physical keyboards, function keys (such as volume control buttons, power buttons, etc.), trackballs, mice, and joysticks, which will not be described in detail here.

[0224] The memory 109 can be used to store software programs and various data. The memory 109 may primarily include a first storage area for storing programs or instructions and a second storage area for storing data. The first storage area may store the operating system, application programs or instructions required for at least one function (such as sound playback, image playback, etc.). Furthermore, the memory 109 may include volatile memory or non-volatile memory, or both. The non-volatile memory may be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory can be random access memory (RAM), static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct memory bus RAM (DRRAM). The memory 109 in the embodiments of this application includes, but is not limited to, these and any other suitable types of memory.

[0225] Processor 110 may include one or more processing units; optionally, processor 110 integrates an application processor and a modem processor, wherein the application processor mainly handles operations involving the operating system, user interface, and applications, and the modem processor mainly handles wireless communication signals, such as a baseband processor. It is understood that the aforementioned modem processor may also not be integrated into processor 110.

[0226] This application also provides a readable storage medium storing a program or instructions. When the program or instructions are executed by a processor, they implement the various processes of the above method embodiments and achieve the same technical effect. To avoid repetition, they will not be described again here.

[0227] The processor is the processor in the electronic device described in the above embodiments. The readable storage medium includes computer-readable storage media, such as computer read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk.

[0228] This application embodiment also provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor. The processor is used to run programs or instructions to implement the various processes of the above method embodiments and achieve the same technical effect. To avoid repetition, it will not be described again here.

[0229] It should be understood that the chip mentioned in the embodiments of this application may also be referred to as a system-on-a-chip, system chip, chip system, or system-on-a-chip, etc.

[0230] This application provides a computer program product, which is stored in a storage medium and executed by at least one processor to implement the various processes of the above method embodiments and achieve the same technical effects. To avoid repetition, it will not be described again here.

[0231] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element. Furthermore, it should be noted that the scope of the methods and apparatuses in the embodiments of this application is not limited to performing functions in the order shown or discussed, but may also include performing functions substantially simultaneously or in the reverse order, depending on the functions involved. For example, the described methods may be performed in a different order than described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.

[0232] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a computer software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of this application.

[0233] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.

Claims

1. A video processing method, characterized in that, The method includes: During the playback of the first video, the first audio frame sequence in the first video is translated based on the first subtitle text in the first screen image corresponding to the first audio frame sequence to obtain the first translated text in a first language; the first language is a language different from the language corresponding to the first subtitle text. Based on the first translated text, the second subtitle text is generated and displayed.

2. The method according to claim 1, characterized in that, Before translating the first audio frame sequence in the first video based on the first subtitle text in the first screen image corresponding to the first audio frame sequence to obtain the first translated text in the first language, the method further includes: Obtain the screen image corresponding to the first audio frame in the first audio frame sequence; Based on the screen image corresponding to the first audio frame, obtain the first subtitle text; Wherein, the first audio frame is a video frame in the first video that meets the conditions; The audio frames that meet the conditions include any of the following: The screen image corresponding to the audio frame preceding the audio frame does not contain subtitle text. The subtitle text in the screen image corresponding to the audio frame is different from the subtitle text in the screen image corresponding to the previous audio frame. The first audio frame sequence is the audio frame sequence in the first video that starts with the first audio frame and matches the first subtitle text.

3. The method according to claim 1 or 2, characterized in that, The method further includes: Obtain the video image within the video playback area of ​​the first screen image; The text in the effective subtitle region of the video image is extracted, and the first subtitle text is obtained based on the text.

4. The method according to claim 2, characterized in that, During the playback of the first video, based on the first subtitle text in the first screen image corresponding to the first audio frame sequence, the first audio frame sequence in the first video is translated to obtain the first translated text in the first language, including: During the playback of the first video, each time an audio frame is acquired, the audio frame is input into the translation model corresponding to the first language, and if an audio frame is an audio frame that meets the conditions, the subtitle text corresponding to the audio frame is input into the translation model. Using the translation model, the first audio frame sequence is translated based on the first subtitle text to obtain the first translated text.

5. The method according to claim 4, characterized in that, The step of translating audio frames in the first audio frame sequence based on the first subtitle text using the translation model to obtain the first translated text includes: Using the translation model, based on the first subtitle text, incremental translation is performed on the audio frames in the first audio frame sequence to obtain the first translated text.

6. A video processing apparatus, characterized in that, The video processing device includes: a translation module and a display module; The translation module is used to translate the first audio frame sequence in the first video based on the first subtitle text in the first screen image corresponding to the first audio frame sequence during the playback of the first video, to obtain a first translated text in a first language; the first language is a language different from the language corresponding to the first subtitle text. The display module is used to generate and display the second subtitle text based on the first translated text translated by the translation module.

7. The apparatus according to claim 6, characterized in that, The device further includes: an acquisition module; The acquisition module is used to acquire the screen image corresponding to the first audio frame in the first audio frame sequence before translating the first audio frame sequence in the first video based on the first subtitle text in the first screen image corresponding to the first audio frame sequence to obtain the first translated text in the first language; and to acquire the first subtitle text based on the screen image corresponding to the first audio frame. Wherein, the first audio frame is a video frame in the first video that meets the conditions; The audio frames that meet the conditions include any of the following: The screen image corresponding to the audio frame preceding the audio frame does not contain subtitle text. The subtitle text in the screen image corresponding to the audio frame is different from the subtitle text in the screen image corresponding to the previous audio frame. The first audio frame sequence is the audio frame sequence in the first video that starts with the first audio frame and matches the first subtitle text.

8. The apparatus according to claim 6 or 7, characterized in that, The device further includes: an acquisition module and an extraction module; The acquisition module is used to acquire video images within the video playback area of ​​the first screen image; The extraction module is used to extract the text from the effective subtitle region in the video image acquired by the acquisition module, and obtain the first subtitle text based on the text.

9. The apparatus according to claim 7, characterized in that, The translation module is specifically used for: During the playback of the first video, each time an audio frame is acquired, the audio frame is input into the translation model corresponding to the first language, and if an audio frame is an audio frame that meets the conditions, the subtitle text corresponding to the audio frame is input into the translation model. Using the translation model, the first audio frame sequence is translated based on the first subtitle text to obtain the first translated text.

10. The apparatus according to claim 9, characterized in that, The translation module is specifically used to perform incremental translation of audio frames in the first audio frame sequence based on the first subtitle text using the translation model, so as to obtain the first translated text.

11. An electronic device, characterized in that, It includes a processor and a memory, the memory storing a program or instructions that can run on the processor, the program or instructions being executed by the processor to implement the steps of the video processing method as described in any one of claims 1 to 5.