Video machine translation method and device fusing multi-modal fine-grained information
By extracting and fusing the pictures and audio in the video into the source text in fine-grained information, the existing video machine translation methods are solved in terms of translation accuracy, and a more efficient and accurate video translation effect is achieved.
Patent Information
- Application Number
- CN202510043829.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-10
- Publication Date
- 2025-05-13
AI Technical Summary
The existing video machine translation methods are insufficient in terms of translation accuracy and are easily interfered with video information that is not related to the translation task.
By extracting information from the pictures and audio in the video, fine-grained visual information and fine-grained audio information are obtained, and this information is fused into the source text and input into the machine translation model for translation.
It significantly improves the accuracy of video translation, reduces interference from redundant information, and improves the quality of translation.
Smart Images

Figure CN119996762A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a video machine translation method and device integrating multimodal fine-grained information. Background Art
[0002] Video is an image content that contains audio and multiple frames, and is commonly found in daily life and communication channels such as the Internet. Subtitles are text that displays dialogues and other non-visual information in the video. The goal of video machine translation is to automatically translate the subtitles in the video from the source language to the target language. It is one of the important technologies for realizing the automated processing of video content.
[0003] Current video machine translation methods can combine video images to help translate subtitles, but this may be interfered by video information that is irrelevant to the translation task, making it impossible to guarantee the accuracy of the translation. Summary of the invention
[0004] The present invention provides a video machine translation method and device integrating multimodal fine-grained information, so as to solve the technical problem that the accuracy of video machine translation cannot be guaranteed in the prior art.
[0005] In a first aspect, the present invention provides a video machine translation method integrating multimodal fine-grained information, comprising the following steps.
[0006] Extracting information from a picture in a video to obtain fine-grained visual information in the picture, and extracting information from an audio in the video to obtain fine-grained audio information in the audio; The fine-grained visual information and the fine-grained audio information are merged into a source text to obtain a merged text; the source text is a subtitle to be translated in the video; The fused text is input into a machine translation model to obtain a target translation text.
[0007] In some embodiments, extracting information from a picture in a video to obtain fine-grained visual information in the picture includes: Acquire multiple frames corresponding to target subtitles; the target subtitles are any subtitles to be translated in the video; Fine-grained visual information is obtained based on a multimodal large model and multiple frames corresponding to the target subtitles; the fine-grained visual information includes entity tags, location tags, expression tags, action tags and video description tags.
[0008] In some embodiments, the multiple frames corresponding to the target subtitles include a frame corresponding to the midpoint time of the target subtitles, a frame corresponding to 0.5 seconds before and after the midpoint time of the target subtitles, a frame corresponding to the start time point of the target subtitles, and / or a frame corresponding to the end time point of the target subtitles; the acquiring of fine-grained visual information based on the multimodal large model and the multiple frames corresponding to the target subtitles includes one or more of the following: Processing the picture corresponding to the midpoint time of the target subtitle using the multimodal large model to obtain entity labels, location labels, and expression labels; Using the multimodal large model to process and extract action information from the picture corresponding to the midpoint time of the target subtitle and the picture corresponding to 0.5 seconds before and after the midpoint time of the target subtitle, and obtaining an action label using the multimodal large model according to the action information; The multimodal large model is used to process and extract the description information of the picture corresponding to the start time point of the target subtitle, the picture corresponding to the end time point of the target subtitle, and the picture corresponding to the midpoint time of the target subtitle, and the video description label is obtained according to the description information using the multimodal large model.
[0009] In some embodiments, extracting information from the audio in the video to obtain fine-grained audio information in the audio includes one or more of the following: Extracting speech emotion information of the video using a speech emotion model to obtain a speech emotion label; The speech in the video is intercepted, and a stress label is obtained based on a digital signal corresponding to the speech.
[0010] In some embodiments, fusing the fine-grained visual information and the fine-grained audio information into a source text to obtain a fused text includes: splicing the fine-grained visual information, the speech emotion label in the fine-grained audio information and the source text to obtain a spliced text; Performing normal distribution scaling on the stress labels in the fine-grained audio information, and multiplying the value obtained after the normal distribution scaling by the word embedding of the source text; The obtained product and the concatenated text are used as the fused text.
[0011] In some embodiments, inputting the fused text into a machine translation model to obtain a target translated text includes: The fused text is input into a machine translation model of the Transformer architecture, attention calculation is performed, and a target translation text is generated autoregressively.
[0012] In a second aspect, the present invention provides a video machine translation device integrating multimodal fine-grained information, comprising the following modules.
[0013] A first acquisition module is used to extract information from a picture in a video to obtain fine-grained visual information in the picture, and to extract information from an audio in the video to obtain fine-grained audio information in the audio; A second acquisition module is used to fuse the fine-grained visual information and the fine-grained audio information into a source text to obtain a fused text; the source text is the subtitles to be translated in the video; The third acquisition module is used to input the fused text into a machine translation model to obtain a target translation text.
[0014] In a third aspect, the present invention provides an electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the method for video machine translation integrating multimodal fine-grained information as described in any one of the above is implemented.
[0015] In a fourth aspect, a non-transitory computer-readable storage medium stores a computer program, which, when executed by a processor, implements any of the above-described video machine translation methods for fusing multimodal fine-grained information.
[0016] In a fifth aspect, the present invention further provides a computer program product, comprising a computer program, which, when executed by a processor, implements the video machine translation method for fusing multimodal fine-grained information as described in any one of the above.
[0017] The video machine translation method and device for integrating multimodal fine-grained information provided by the present invention extract information from the picture in the video to obtain fine-grained visual information in the picture, and extract information from the audio in the video to obtain fine-grained audio information in the audio; the fine-grained visual information and the fine-grained audio information are integrated into the source text to obtain integrated text; the source text is the subtitles to be translated in the video; the integrated text is input into the machine translation model to obtain the target translation text. Considering that multimodal information includes fine-grained visual information and fine-grained audio information of the video, these multimodal fine-grained information are integrated into the input to assist translation, which significantly improves the auxiliary role of the video in translation and improves the accuracy of translation. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0019] Figure 1 It is a flow chart of the video machine translation method integrating multi-modal fine-grained information provided by the present invention.
[0020] Figure 2 It is a schematic diagram of the extraction process of fine-grained visual information provided by the present invention.
[0021] Figure 3 It is a schematic diagram of the extraction process of fine-grained audio information provided by the present invention.
[0022] Figure 4 It is a model framework diagram of the video machine translation method that integrates multimodal fine-grained information provided by the present invention.
[0023] Figure 5 It is a structural schematic diagram of a video machine translation device integrating multimodal fine-grained information provided by the present invention.
[0024] Figure 6 It is a structural schematic diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION
[0025] The subtitles in a video are highly relevant to the video itself, and are usually a text transcription of the video's speech. Translating subtitles into other languages can help viewers from different language backgrounds better understand the video content.
[0026] Unlike traditional text machine translation, subtitles in video translation are closely linked to the video screen, usually a textual presentation of the video audio, and are highly relevant to the screen content. Therefore, when translating subtitles, combining video information can help improve translation accuracy, especially when encountering problems such as ambiguous words or reference resolution, the additional information provided by the video screen is particularly important.
[0027] At present, the main method of incorporating video content into the translation process is to extract video frames at regular intervals, that is, to select video frames at a certain frequency, and input these frames into the visual model to extract visual features, and then combine these features with the text and input them into the translation model. However, this method has two main problems: first, it requires selecting multiple frames from the video, which not only reduces the processing speed, but also may introduce redundant information that is irrelevant to the translation task; second, these methods focus too much on the visual content in the video, while ignoring the potential benefit of audio content on video translation.
[0028] Based on the above technical problems, the present invention proposes a video machine translation method that integrates multimodal fine-grained information, extracts information from the picture in the video to obtain fine-grained visual information in the picture, and extracts information from the audio in the video to obtain fine-grained audio information in the audio; the fine-grained visual information and the fine-grained audio information are integrated into the source text to obtain a fused text; the source text is the subtitles to be translated in the video; the fused text is input into the machine translation model to obtain the target translation text. Considering that multimodal information includes fine-grained visual information and fine-grained audio information of the video, these information are integrated into the source text, and machine translation is performed based on the fused text, thereby improving the accuracy of translation.
[0029] In order to make the purpose, technical solution and advantages of the present invention clearer, the technical solution of the present invention will be clearly and completely described below in conjunction with the drawings of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0030] Figure 1 is a flow chart of a video machine translation method integrating multimodal fine-grained information provided by the present invention, such as Figure 1 As shown, the present invention provides a video machine translation method integrating multimodal fine-grained information. The method comprises: Step 101: extract information from the picture in the video to obtain fine-grained visual information in the picture, and extract information from the audio in the video to obtain fine-grained audio information in the audio.
[0031] Specifically, for a video that needs to be subtitled, information is first extracted from the picture and audio of the video to obtain fine-grained visual information in the picture and fine-grained audio information in the audio.
[0032] Fine-grained visual information may include entity tags, location tags, expression tags, action tags, and / or video description tags, etc.
[0033] The fine-grained audio information may include a tone emotion label and / or a stress label in the speech information.
[0034] Step 102: Fusing the fine-grained visual information and the fine-grained audio information into a source text to obtain a fused text; the source text is the subtitles to be translated in the video.
[0035] Specifically, the fine-grained visual information and the fine-grained audio information are fused into the source text, for example, by concatenation or mathematical calculation, to obtain a fused text.
[0036] Step 103: input the fused text into a machine translation model to obtain a target translation text.
[0037] Specifically, the machine translation model is used to perform machine translation, and the obtained fusion text is input into the machine translation model to obtain a target translation text, that is, a target language translation text corresponding to the source text.
[0038] The video machine translation method that integrates multimodal fine-grained information provided in the embodiment of the present application obtains fine-grained visual information and fine-grained audio information in the video, integrates this multimodal fine-grained information into the source text, and then uses the translation model for translation, thereby improving the accuracy of video translation.
[0039] In some embodiments, extracting information from a picture in a video to obtain fine-grained visual information in the picture includes: Acquire multiple frames corresponding to target subtitles; the target subtitles are any subtitles to be translated in the video; Fine-grained visual information is obtained based on a multimodal large model and multiple frames corresponding to the target subtitles; the fine-grained visual information includes entity tags, location tags, expression tags, action tags and video description tags.
[0040] Specifically, for a certain subtitle in the video (ie, the target subtitle), multiple frames corresponding to the target subtitle are obtained, that is, multiple frames within the time period from the start time point of the target subtitle to the end time point of the target subtitle, including the frames corresponding to the start time point and the end time point.
[0041] These images are then processed using multimodal large language models (MLLMs) to extract a variety of information from the images, such as entities, locations, expressions, actions, and image descriptions, and to create a variety of labels to obtain fine-grained visual information.
[0042] Figure 2 is a schematic diagram of the extraction process of fine-grained visual information provided by the present invention, such as Figure 2As shown in the figure, the machine translation task is to translate the English subtitles in the video into Chinese. The start time of the first subtitle in the video (i.e., the original data) is 0:03, the end time is 0:13, and the start time of the last subtitle is 2:24, and the end time is 2:34. For these subtitles, the following operations are performed: obtain multiple frames corresponding to the subtitles, use the multimodal large model to obtain a variety of information in these frames, and make them into fine-grained labels of visual modalities, including location labels, entity labels, facial expression labels, action labels, and video description labels, to form multimodal fine-grained labels (i.e., multimodal fine-grained information), and output them in text form.
[0043] The video machine translation method that integrates multimodal fine-grained information provided in the embodiment of the present application uses a multimodal large language model to accurately obtain a variety of information from multiple frames corresponding to subtitles, and produces a variety of labels, enriches the label types, and ensures the validity of the labels, that is, obtains accurate and effective multimodal fine-grained information, thereby ensuring the accuracy of subsequent translation tasks.
[0044] In some embodiments, the multiple frames corresponding to the target subtitles include a frame corresponding to the midpoint time of the target subtitles, a frame corresponding to 0.5 seconds before and after the midpoint time of the target subtitles, a frame corresponding to the start time point of the target subtitles, and / or a frame corresponding to the end time point of the target subtitles; the acquiring of fine-grained visual information based on the multimodal large model and the multiple frames corresponding to the target subtitles includes one or more of the following: Processing the picture corresponding to the midpoint time of the target subtitle using the multimodal large model to obtain entity labels, location labels, and expression labels; Using the multimodal large model to process and extract action information from the picture corresponding to the midpoint time of the target subtitle and the picture corresponding to 0.5 seconds before and after the midpoint time of the target subtitle, and obtaining an action label using the multimodal large model according to the action information; The multimodal large model is used to process and extract the description information of the picture corresponding to the start time point of the target subtitle, the picture corresponding to the end time point of the target subtitle, and the picture corresponding to the midpoint time of the target subtitle, and the video description label is obtained according to the description information using the multimodal large model.
[0045] Specifically, one or more frames corresponding to different time points are selected to extract different fine-grained visual information and obtain different fine-grained labels.
[0046] For entities, places, and facial expressions, the multimodal large model is used to process the picture corresponding to the midpoint time of the target subtitle to obtain entity labels, place labels, and expression labels.
[0047] For action information, the multimodal large model is used to extract action information from the three pictures corresponding to the midpoint time of the target subtitle and the pictures corresponding to 0.5 seconds before and after the midpoint time of the target subtitle. Then the multimodal large model is used to obtain action labels based on the three extracted action information.
[0048] For the picture description information, the multimodal large model is used to extract the picture description information of the picture corresponding to the start time point, the picture corresponding to the end time point, and the picture corresponding to the midpoint time of the target subtitle. Then the multimodal large model is used to obtain the video description label based on the three extracted picture description information.
[0049] Finally, the obtained entity labels, location labels, expression labels, action labels, and video description labels are used as fine-grained visual information.
[0050] The video machine translation method that integrates multimodal fine-grained information provided in the embodiment of the present application selects pictures corresponding to different time points for extraction according to different fine-grained visual information, which can extract the corresponding information more specifically and accurately and reduce the interference of redundant information.
[0051] In some embodiments, extracting information from the audio in the video to obtain fine-grained audio information in the audio includes one or more of the following: Extracting speech emotion information of the video using a speech emotion model to obtain a speech emotion label; The speech in the video is intercepted, and a stress label is obtained based on a digital signal corresponding to the speech.
[0052] Specifically, fine-grained audio information includes speech emotion labels and stress labels.
[0053] Figure 3 FIG. 1 is a schematic diagram of the extraction process of fine-grained audio information provided by the present invention. Figure 3 As shown, a speech emotion model can be used to extract speech emotion information of a video and automatically annotate to obtain a speech emotion label. The speech emotion model can be a pre-trained emotion recognition model, such as an emotion recognition model trained based on a convolutional neural network, or a speech emotion base model such as emotion2vec.
[0054] The stress labels of speech are obtained by mathematical calculation. First, the speech in the video is intercepted, and the digital signal corresponding to the speech is obtained. Then, the root mean square is calculated based on the digital signal to obtain the stress labels in the multimodal fine-grained speech information. The calculation formula for calculating the root mean square based on the digital signal is as follows: in, is the stress label of the i-th word in the source language, is The corresponding voice digital signal, N is the total number of speech sampling points.
[0055] The video machine translation method that integrates multimodal fine-grained information provided in the embodiment of the present application takes into account the impact of audio content on video translation, obtains speech emotion labels and stress labels of video speech as fine-grained audio information to assist in video translation, thereby improving the accuracy of translation.
[0056] In some embodiments, fusing the fine-grained visual information and the fine-grained audio information into a source text to obtain a fused text includes: splicing the fine-grained visual information, the speech emotion label in the fine-grained audio information and the source text to obtain a spliced text; Performing normal distribution scaling on the stress labels in the fine-grained audio information, and multiplying the value obtained after the normal distribution scaling by the word embedding of the source text; The obtained product and the concatenated text are used as the fused text.
[0057] Specifically, Figure 4 is a model framework diagram of the video machine translation method integrating multimodal fine-grained information provided by the present invention, such as Figure 4 As shown, after obtaining fine-grained visual information and fine-grained audio information, the entity labels, location labels, expression labels, action labels, video description labels in the fine-grained visual information and the speech emotion labels in the fine-grained audio information are spliced with the source text to obtain a spliced text.
[0058] The stress labels in the fine-grained audio information are scaled by the normal distribution, and then the value obtained after the normal distribution scaling is multiplied by the word embedding of the source text. The calculation formula is as follows: in, () represents the embedding layer, Represents the source text, word embeddings representing the source text, represents the stress labels after scaling by normal distribution, Represents the product of the scaled normal distribution of stress labels and the word embedding of the source text.
[0059] Finally, the obtained product and the concatenated text are used together as the fused text.
[0060] In some embodiments, inputting the fused text into a translation model to obtain a target translated text includes: The fused text is input into a machine translation model of the Transformer architecture, attention calculation is performed, and a target translation text is generated autoregressively.
[0061] Specifically, the fused text that integrates multimodal fine-grained information (including fine-grained visual information and fine-grained audio information) is input into the machine translation model of the Transformer architecture, and attention calculation is performed to autoregressively generate the translated text in the target language (i.e., the target translated text). That is, the cross entropy loss function between the target translated text and the source text is calculated, and the target translated text is autoregressively generated based on the cross entropy loss function. The expression of the cross entropy loss function is as follows: in, represents the cross entropy loss function between the generated target sentence (i.e., target translation text) and the source text; n represents the total length of the target sentence to be generated; Represents the tth word in the currently generated target sentence (t is a positive integer); Indicates the words that have been generated, that is, the first word to the t-1th word of the target sentence; Represents the input source text; Represents the model parameters of the machine translation model; represents the predicted probability.
[0062] In the embodiment of the present application, the video translation effect of the proposed method can be verified on the video translation dataset TriFine. The TriFine dataset is a large-scale video machine translation dataset in Chinese and English, covering various video types in mainstream video platforms. In addition, in order to verify the advantages of the method proposed in the present invention in multimodal fine-grained scenarios, experimental tests were also carried out on the ambiguous test set of the TriFine dataset. The pure text method and the existing method based on coarse-grained visual information were selected to compare with the method based on fine-grained information fusion proposed in the present invention. The evaluation indicators used are the three most commonly used evaluation indicators BLEU, METEOR and COMET for translation tasks. The experimental results show that the method proposed in the present invention has achieved good results on the general test set, and is superior to the comparison method in both Chinese to English and English to Chinese translation tasks. This proves the advantages of the method proposed in the present invention in terms of versatility and robustness. Moreover, the effect of the method proposed in the present invention on the ambiguous test set is also significantly better than the comparison method, proving that the method proposed in the present invention can better understand the video by utilizing multimodal fine-grained information, and use multimodal information for disambiguation to improve translation quality.
[0063] The video machine translation method that integrates multimodal fine-grained information provided in the embodiment of the present application solves the problems of low efficiency and susceptibility to interference from redundant information in the current video translation methods. The embodiment of the present application can extract a variety of multimodal fine-grained information from the video screen according to the set information extraction process, and then integrate these multimodal fine-grained information into the input to assist translation, which significantly improves the auxiliary role of video in translation.
[0064] Figure 5 is a schematic diagram of the structure of a video machine translation device integrating multimodal fine-grained information provided by the present invention, such as Figure 5 As shown, the present invention provides a video machine translation device integrating multimodal fine-grained information, including a first acquisition module 501, a second acquisition module 502 and a third acquisition module 503.
[0065] The first acquisition module 501 is used to extract information from the pictures in the video to obtain fine-grained visual information in the pictures, and to extract information from the audio in the video to obtain fine-grained audio information in the audio.
[0066] The second acquisition module 502 is used to merge the fine-grained visual information and the fine-grained audio information into a source text to obtain a merged text; the source text is the subtitles to be translated in the video.
[0067] The third acquisition module 503 is used to input the fused text into a machine translation model to obtain a target translation text.
[0068] Specifically, the above-mentioned video machine translation device integrating multimodal fine-grained information provided by the present invention can implement all the method steps implemented by the above-mentioned video machine translation method embodiment integrating multimodal fine-grained information, and can achieve the same technical effect. The parts and beneficial effects of this embodiment that are the same as those of the method embodiment will not be described in detail here.
[0069] It should be noted that the division of units / modules in the above-mentioned embodiments of the present invention is schematic and is only a logical function division. There may be other division methods in actual implementation. In addition, the functional units in the various embodiments of the present application may be integrated into a processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated units may be implemented in the form of hardware or in the form of software functional units.
[0070] Figure 6 is a schematic diagram of the structure of the electronic device provided by the present invention, such as Figure 6As shown, the electronic device may include: a processor 601, a communications interface 602, a memory 603 and a communication bus 604, wherein the processor 601, the communications interface 602 and the memory 603 communicate with each other through the communication bus 604. The processor 601 may call the logic instructions in the memory 603 to execute a video machine translation method integrating multimodal fine-grained information, the method comprising: Extracting information from a picture in a video to obtain fine-grained visual information in the picture, and extracting information from an audio in the video to obtain fine-grained audio information in the audio; The fine-grained visual information and the fine-grained audio information are merged into a source text to obtain a merged text; the source text is a subtitle to be translated in the video; The fused text is input into a machine translation model to obtain a target translation text.
[0071] In some embodiments, extracting information from a picture in a video to obtain fine-grained visual information in the picture includes: Acquire multiple frames corresponding to target subtitles; the target subtitles are any subtitles to be translated in the video; Fine-grained visual information is obtained based on a multimodal large model and multiple frames corresponding to the target subtitles; the fine-grained visual information includes entity tags, location tags, expression tags, action tags and video description tags.
[0072] In some embodiments, the multiple frames corresponding to the target subtitles include a frame corresponding to the midpoint time of the target subtitles, a frame corresponding to 0.5 seconds before and after the midpoint time of the target subtitles, a frame corresponding to the start time point of the target subtitles, and / or a frame corresponding to the end time point of the target subtitles; the acquiring of fine-grained visual information based on the multimodal large model and the multiple frames corresponding to the target subtitles includes one or more of the following: Processing the picture corresponding to the midpoint time of the target subtitle using the multimodal large model to obtain entity labels, location labels, and expression labels; Using the multimodal large model to process and extract action information from the picture corresponding to the midpoint time of the target subtitle and the picture corresponding to 0.5 seconds before and after the midpoint time of the target subtitle, and obtaining an action label using the multimodal large model according to the action information; The multimodal large model is used to process and extract the description information of the picture corresponding to the start time point of the target subtitle, the picture corresponding to the end time point of the target subtitle, and the picture corresponding to the midpoint time of the target subtitle, and the video description label is obtained according to the description information using the multimodal large model.
[0073] In some embodiments, extracting information from the audio in the video to obtain fine-grained audio information in the audio includes one or more of the following: Extracting speech emotion information of the video using a speech emotion model to obtain a speech emotion label; The speech in the video is intercepted, and a stress label is obtained based on a digital signal corresponding to the speech.
[0074] In some embodiments, fusing the fine-grained visual information and the fine-grained audio information into a source text to obtain a fused text includes: splicing the fine-grained visual information, the speech emotion label in the fine-grained audio information and the source text to obtain a spliced text; Performing normal distribution scaling on the stress labels in the fine-grained audio information, and multiplying the value obtained after the normal distribution scaling by the word embedding of the source text; The obtained product and the concatenated text are used as the fused text.
[0075] In some embodiments, inputting the fused text into a machine translation model to obtain a target translated text includes: The fused text is input into a machine translation model of the Transformer architecture, attention calculation is performed, and a target translation text is generated autoregressively.
[0076] Specifically, the processor 601 may be a central processing unit (CPU), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or a complex programmable logic device (CPLD), and the processor may also adopt a multi-core architecture.
[0077] The logic instructions in the memory 603 can be implemented in the form of software functional units and can be stored in a processor-readable storage medium when sold or used as an independent product. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) or a processor to perform all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), disk or optical disk, etc., various media that can store program codes.
[0078] In some embodiments, a computer program product is further provided. The computer program product includes a computer program. The computer program can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the video machine translation method for fusing multimodal fine-grained information provided by the above-mentioned method embodiments. The method includes: Extracting information from a picture in a video to obtain fine-grained visual information in the picture, and extracting information from an audio in the video to obtain fine-grained audio information in the audio; The fine-grained visual information and the fine-grained audio information are merged into a source text to obtain a merged text; the source text is a subtitle to be translated in the video; The fused text is input into a machine translation model to obtain a target translation text.
[0079] Specifically, the above-mentioned computer program product provided in the embodiment of the present application can implement all the method steps implemented by the above-mentioned method embodiments, and can achieve the same technical effect. The parts and beneficial effects of this embodiment that are the same as the method embodiment will not be described in detail here.
[0080] In some embodiments, a computer-readable storage medium is further provided, wherein the computer-readable storage medium stores a computer program, wherein the computer program is used to enable a computer to execute the video machine translation method for fusing multimodal fine-grained information provided by the above-mentioned method embodiments, the method comprising: Extracting information from a picture in a video to obtain fine-grained visual information in the picture, and extracting information from an audio in the video to obtain fine-grained audio information in the audio; The fine-grained visual information and the fine-grained audio information are merged into a source text to obtain a merged text; the source text is a subtitle to be translated in the video; The fused text is input into a machine translation model to obtain a target translation text.
[0081] Specifically, the above-mentioned computer-readable storage medium provided by the present invention can implement all the method steps implemented by the above-mentioned method embodiments, and can achieve the same technical effects. The parts and beneficial effects that are the same as the method embodiments in this embodiment will not be described in detail here.
[0082] It should be noted that the computer-readable storage medium can be any available medium or data storage device that can be accessed by the processor, including but not limited to magnetic storage (such as floppy disks, hard disks, magnetic tapes, magneto-optical disks (MO), etc.), optical storage (such as CD, DVD, BD, HVD, etc.), and semiconductor storage (such as ROM, EPROM, EEPROM, non-volatile memory (NAND FLASH), solid-state drive (SSD)), etc.
[0083] It should also be noted that the terms "first", "second", etc. in the present invention are used to distinguish similar objects, and are not used to describe a specific order or sequence. It should be understood that the terms used in this way can be interchangeable under appropriate circumstances, so that the embodiments of the present invention can be implemented in an order other than those illustrated or described herein, and the objects distinguished by "first" and "second" are usually of the same type, and the number of objects is not limited. For example, the first object can be one or more.
[0084] In the present invention, the term "and / or" describes the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist at the same time, and B exists alone. The character " / " generally indicates that the associated objects before and after are in an "or" relationship.
[0085] In the present invention, "determine B based on A" means that the factor A should be considered when determining B. It is not limited to "B can be determined based on A alone", but should also include: "determine B based on A and C", "determine B based on A, C and E", "determine C based on A, and further determine B based on C", etc. In addition, it can also include taking A as a condition for determining B, for example, "when A meets the first condition, use the first method to determine B"; for another example, "when A meets the second condition, determine B", etc.; for another example, "when A meets the third condition, determine B based on the first parameter", etc. Of course, it can also be a condition that takes A as a factor for determining B, for example, "when A meets the first condition, use the first method to determine C, and further determine B based on C", etc.
[0086] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the present application may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage and optical storage, etc.) containing computer-usable program codes.
[0087] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer executable instructions. These computer executable instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0088] These processor executable instructions may also be stored in a processor readable memory that can direct a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the processor readable memory produce an article of manufacture including an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0089] These processor-executable instructions may also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for executing on the computer or other programmable device to implement the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0090] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalents, the present invention is also intended to include these modifications and variations.
Claims
1. A video machine translation method integrating multimodal fine-grained information, characterized in that: include: Extracting information from a picture in a video to obtain fine-grained visual information in the picture, and extracting information from an audio in the video to obtain fine-grained audio information in the audio; The fine-grained visual information and the fine-grained audio information are merged into a source text to obtain a merged text; the source text is a subtitle to be translated in the video; The fused text is input into a machine translation model to obtain a target translation text.
2. The video machine translation method integrating multimodal fine-grained information according to claim 1, characterized in that: The extracting information from the picture in the video to obtain fine-grained visual information in the picture includes: Acquire multiple frames corresponding to target subtitles; the target subtitles are any subtitles to be translated in the video; Fine-grained visual information is obtained based on a multimodal large model and multiple frames corresponding to the target subtitles; the fine-grained visual information includes entity tags, location tags, expression tags, action tags and video description tags.
3. The video machine translation method integrating multimodal fine-grained information according to claim 2, characterized in that: The multiple frames corresponding to the target subtitles include a frame corresponding to the midpoint time of the target subtitles, a frame corresponding to 0.5 seconds before and after the midpoint time of the target subtitles, a frame corresponding to the start time point of the target subtitles, and / or a frame corresponding to the end time point of the target subtitles; the acquiring of fine-grained visual information based on the multimodal large model and the multiple frames corresponding to the target subtitles includes one or more of the following: Processing the picture corresponding to the midpoint time of the target subtitle using the multimodal large model to obtain entity labels, location labels, and expression labels; Using the multimodal large model to process and extract action information from the picture corresponding to the midpoint time of the target subtitle and the picture corresponding to 0.5 seconds before and after the midpoint time of the target subtitle, and obtaining an action label using the multimodal large model according to the action information; The multimodal large model is used to process and extract the description information of the picture corresponding to the start time point of the target subtitle, the picture corresponding to the end time point of the target subtitle, and the picture corresponding to the midpoint time of the target subtitle, and the video description label is obtained according to the description information using the multimodal large model.
4. The video machine translation method integrating multimodal fine-grained information according to claim 1, characterized in that: The extracting information from the audio in the video to obtain fine-grained audio information in the audio includes one or more of the following: Extracting speech emotion information of the video using a speech emotion model to obtain a speech emotion label; The speech in the video is intercepted, and a stress label is obtained based on a digital signal corresponding to the speech.
5. The video machine translation method integrating multimodal fine-grained information according to claim 1, characterized in that: The step of fusing the fine-grained visual information and the fine-grained audio information into a source text to obtain a fused text includes: splicing the fine-grained visual information, the speech emotion label in the fine-grained audio information and the source text to obtain a spliced text; Performing normal distribution scaling on the stress labels in the fine-grained audio information, and multiplying the value obtained after the normal distribution scaling by the word embedding of the source text; The obtained product and the concatenated text are used as the fused text.
6. The video machine translation method integrating multimodal fine-grained information according to claim 1, characterized in that: The step of inputting the fused text into a machine translation model to obtain a target translated text comprises: The fused text is input into a machine translation model of the Transformer architecture, attention calculation is performed, and a target translation text is generated autoregressively.
7. A video machine translation device integrating multimodal fine-grained information, characterized in that: include: A first acquisition module is used to extract information from a picture in a video to obtain fine-grained visual information in the picture, and to extract information from an audio in the video to obtain fine-grained audio information in the audio; A second acquisition module is used to fuse the fine-grained visual information and the fine-grained audio information into a source text to obtain a fused text; the source text is the subtitles to be translated in the video; The third acquisition module is used to input the fused text into a machine translation model to obtain a target translation text.
8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the video machine translation method integrating multimodal fine-grained information as described in any one of claims 1 to 6 is implemented.
9. A non-transitory computer-readable storage medium, characterized in that: The non-transitory computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the video machine translation method for fusing multimodal fine-grained information as claimed in any one of claims 1 to 6 is implemented.
10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the video machine translation method integrating multimodal fine-grained information as claimed in any one of claims 1 to 6 is implemented.
Citation Information
Cited By
Automatic video editing method and device based on text and shot similarity, and terminal
CN120499445A
Voice generation method and device
CN121096318A