Video editing device and program
The video editing system addresses the challenge of conveying emotional nuances in dubbed audio by recognizing speaker attributes and modifying translations accordingly, resulting in more engaging and relatable video content.
Patent Information
- Application Number
- JP2023181239
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-10-20
- Publication Date
- 2025-05-02
- Estimated Expiration
- 2043-10-20
AI Technical Summary
Existing video content dubbing techniques struggle to convey the nuances of a speaker's emotions and words, as the dialogue and discourse in dubbed audio often appear uniform and lack the emotional depth of the original.
A video editing system that recognizes sounds in a video file to generate text data indicating the sound and attribute data of the speaker, translates this data into another language, and modifies the translation based on the speaker's attributes, such as age, gender, and tone, to create dubbed audio that better reflects the original emotional nuances.
This approach enables the generation of video content with dubbed audio that effectively conveys the emotional nuances and speaker attributes, enhancing the viewer's experience by making the dialogue more relatable and engaging.
Smart Images

Figure 2025070727000001_ABST
Abstract
Description
[Technical field]
[0001] The present invention relates to a technique for generating moving image files, and more particularly to a technique for generating a moving image file including subtitles and dubbed audio from a moving image file including audio in a foreign language. [Background technology]
[0002] Patent Documents 1, 2, and 3 are documents that disclose this type of technology. The receiving device described in Patent Document 1 buffers a multimedia presentation, extracts audio and video components of the multimedia presentation, performs speech recognition analysis on the audio components to generate subtitle text data, and integrates the subtitle text data with the original video components for output. The transcription server described in Patent Document 2 translates text generated by a speech recognition unit into a user's target language according to a user's profile, and transmits the translation text, which is the translation result, to the user's mobile terminal. The machine translator described in Patent Document 3 processes a timed text track in an online video service, translates the timed text of the timed text track into a predetermined foreign language, and transmits the content of the online video service with the timed text rewritten to the user's viewer client. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] U.S. Patent No. 8,781,824 [Patent Document 2] European Patent No. 2725816 [Patent Document 3] US Patent Application Publication No. 2010 / 0138209 Summary of the Invention [Problem to be solved by the invention]
[0004] However, the video content dubbing technologies proposed to date have had the problem that the dialogue and speech in the dubbed audio is uniform, making it difficult for the viewer to understand the speaker's emotions and the nuances of the words.
[0005] The present invention has been made in consideration of such problems, and aims to provide a technical means that can adjust the language used in a video content depending on the attributes of the person who uttered the lines or comments. [Means for solving the problem]
[0006] In order to solve the above problem, a video editing device which is a preferred embodiment of the present invention is characterized in that it comprises a recognition means for recognizing each sound in a video file to be processed as first text data indicating the sound and attribute data of the person who uttered the sound, a translation means for translating the first text data into second text data in another language, and a modification means for generating third text data by modifying the wording of a part of the translation result indicated by the second text data based on the attribute data, and outputting the generated third text data together with data indicating an image in the video file.
[0007] In this aspect, the recognition means may have an attribute recognition processing means that performs FFT processing on audio data of the audio in the video file and identifies the attributes of the person who uttered the audio based on spectrum data obtained by the FFT processing.
[0008] The recognition means may further include an audio extraction means for extracting only the audio data of the portion in which human vocalizations are recorded from the audio data between the start and end of the video content in the video file, and a division means for dividing the audio data extracted by the audio extraction means into a plurality of audio data each indicating a group of vocalizations, and the attribute recognition processing means may identify the attributes of the person who uttered the audio for each of the plurality of audio data divided by the division means.
[0009] Another preferred aspect of the present invention is a program that causes a computer to realize a recognition function that recognizes each sound in a video file to be processed as first text data indicating the sound and attribute data of the person who uttered the sound, a translation function that translates the first text data into second text data in another language, and a modification function that generates third text data by modifying the wording of some of the translation result indicated by the second text data based on the attribute data, and outputs the generated third text data together with data indicating images in the video file. Effect of the Invention
[0010] According to the present invention, video content including dubbed audio information that easily conveys the speaker's emotions and nuances of the words to the viewer can be generated from video content including audio in a foreign language. [Brief description of the drawings]
[0011] [Figure 1] 1 is a diagram showing an overall configuration of a video editing system according to an embodiment of the present invention; [Diagram 2] 2 is a flowchart showing the operation of the video editing system of FIG. 1. [Diagram 3] FIG. 3 is a diagram showing the process contents of steps S201 to S203 in FIG. [Figure 4] FIG. 3 is a diagram showing the process contents of steps S204 to S210 in FIG. [Diagram 5] FIG. 3 is a diagram showing the process contents of S211 to S214 in FIG. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0012] An embodiment of the present invention will be described below with reference to the drawings. Fig. 1 is a diagram showing the overall configuration of a video editing system 1 according to an embodiment of the present invention. This video editing system 1 generates a second video file from a first video file including audio in a first language (e.g., English) by converting the audio in the video content into subtitles or dubbed audio in a second language (e.g., Japanese).
[0013] The video editing system 1 includes a terminal 10 and a video editing device 20. The terminal 10 and the video editing device 20 are connected via a network 90. The terminal 10 is a smartphone or a personal computer used by a user. The video editing device 20 is a server that operates under the management of a service provider. The video editing device 20 includes a CPU 21, a RAM 22, a ROM 23, a hard disk 24, a communication interface 25, a display 26, a keyboard 27, and a pointing device 28. The CPU 21 uses the RAM 22 as a work area and executes programs stored in the ROM 23 and the hard disk 24. A video editing program 29 is stored in the hard disk 24.
[0014] The video editing program 29 has the following functions. a1.Recognition function This is a function that receives a first moving image file to be processed from terminal 10 via network 90, and recognizes each sound in a first language in the first moving image file as first text data indicating the sound and attribute data of the person who uttered the sound. b1. Translation function This is a function for translating first text data into second text data in a second language. c1. Modification functions This is a function that generates third text data by modifying some of the wording of the translation result indicated by the second text data based on attribute data, and outputs the generated third text data together with data indicating the images in the video file.
[0015] Next, the operation of this embodiment will be described. Fig. 2 is a flowchart showing the operation of this embodiment. Fig. 3 shows the processing contents of steps S201 to S203 in Fig. 2. Fig. 4 shows the processing contents of steps S204 to S210 in Fig. 2. Fig. 5 shows the processing contents of steps S211 to S214 in Fig. 2. Steps S201 to S214 and S231 in these figures are executed by the video editing program 29.
[0016] 2, the terminal 10 performs an upload process. In the upload process, the user performs an operation to select a desired video file from within the memory of the terminal 10. When this operation is performed, the terminal 10 transmits the selected video file to the video editing device 20 as a first video file to be processed. When the CPU 21 of the video editing device 20 receives the first video file from the terminal 10, it stores the first video file in the RAM 22.
[0017] 2 and 3, the CPU 21 of the video editing device 20 performs a noise removal process. In the noise removal process, the CPU 21 decodes the first video file in the RAM 22, stores a series of frame image data from the start to the end of the video content, which is the decoded result, and audio data indicating the sound waveform of the sound from the start to the end of the video content in the RAM 22, and removes noise from the audio data in the RAM 22. This noise removal may be performed by a known noise removal program.
[0018] 2 and 3, the CPU 21 of the video editing device 20 performs an audio extraction process. In the audio extraction process, the CPU 21 extracts only audio data of a portion in which human vocalizations (such as lines or statements) are recorded in the video content from the noise-removed audio data in the RAM 22, and stores this audio data in the RAM 22. More specifically, the CPU 21 searches for a track in which human vocalizations are recorded from a plurality of sound tracks in the noise-removed audio data, extracts only the audio data of the sound waveform recorded on this track, and stores it in the RAM 22 as audio data to be processed later.
[0019] 2 and 3, the CPU 21 of the video editing device 20 performs an audio division process. In the audio division process, the CPU 21 divides the audio data of a person's vocalization in the RAM 22 into multiple pieces of audio data, each of which represents a group of vocalizations. More specifically, in the audio division process, the CPU 21 scans all sound waveforms from the beginning to the end of the video content represented by the original audio data, and whenever a silent point appears, cuts out the sound waveform before that point to create one piece of audio data. This process is repeated. The multiple pieces of audio data obtained by this audio division process represent one line or statement uttered by one person.
[0020] In step S204 in FIG. 2 and FIG. 4, the CPU 21 of the video editing device 20 performs attribute recognition processing. In the attribute recognition processing, the CPU 21 recognizes the attribute of the person who uttered the voice indicated by each of the multiple voice data in the RAM 22, and stores attribute data indicating the recognized attribute in the RAM 22 in association with the voice data. The attributes include the rough age of the person who uttered the voice (e.g., infant, under 10, in their 20s, in their 30s, in their 40s, in their 50s, in their 60s or older), gender (e.g., female or male), and tone of voice (e.g., high, normal, low). More specifically, in the attribute recognition processing, each voice data is subjected to FFT (Fast Fourier transform) processing, and the attribute of the person who uttered the voice is identified based on the spectrum data obtained by the FFT. The identification of the attribute based on the spectrum data may be performed by inputting the voice data into a trained model for attribute recognition.
[0021] 2 and 4, the CPU 21 of the video editing device 20 performs a transcription process. In the transcription process, the CPU 21 performs a text transfer process on each piece of audio data divided by the audio division process in the RAM 22, and sets the result of the text transfer process as first text data indicating the audio of the audio data, and stores the first text data in the RAM 22 in association with a timestamp indicating the appearance timing of the audio in the audio content.
[0022] 2 and 4, the CPU 21 of the video editing device 20 performs a dictionary adjustment process. In the dictionary adjustment process, the words obtained by the transcription process in step S205 are registered in a dictionary library.
[0023] 2 and 4, the CPU 21 of the video editing device 20 performs a reconstruction process. In the reconstruction process, the CPU 21 reconstructs the first text data in the RAM 22 into a longer block of first text data. For example, if the first text data TX(30), TX(31), TX(32), TX(33), TX(34), and TX(35) of six consecutive sounds in the RAM 22 (the numbers in parentheses indicate the order of appearance of the sounds counted from the beginning of the video content) form one meaningful paragraph, the first text data TX(30), TX(31), TX(32), TX(33), TX(34), and TX(35) of the six sounds are consolidated into one block of first text data TX'(30).
[0024] 2 and 4, the CPU 21 of the video editing device 20 performs a translation process. In the translation process, the CPU 21 translates the reconstructed first text data in the RAM 22 into a second language, and stores the translation result in the RAM 22 as second text data indicating translated subtitles for each paragraph from the start to the end of the video content indicated by the original audio data.
[0025] 2 and 4, the CPU 21 of the video editing device 20 performs a paragraph division process. In the paragraph division process, the CPU 21 divides the translated second text data in the RAM 22 into second text data corresponding to subtitles to be displayed in the same image in the video content. For example, if the second text data TX' (30) of one paragraph in the RAM 22 corresponds to subtitles displayed in two separate screens, the second text data TX' (30) is divided into second text data TX'' (30-33) corresponding to the subtitles of the first screen and second text data TX'' (34-35) corresponding to the subtitles of the second screen.
[0026] 2 and 4, the CPU 21 of the video editing device 20 performs a linguistic quality assurance (LQA) process. In the linguistic quality assurance process, the CPU 21 generates third text data by modifying a part of the wording of the translation result indicated by the second text data in the RAM 22 based on the attribute data in the RAM 22, and stores the generated third text data in the RAM 22. For example, if the second text data TX'' (30-33) corresponding to the subtitles of one screen in the RAM 22 is composed of four voices, and the attribute data corresponding to these four voices are approximate age = 20s, gender = female, and tone of voice = normal, in the linguistic quality assurance process, the second text data TX'' (30-33) is modified to use the wording in the normal tone of a woman in her 20s, and the third text data TX'' (30-33) is generated. Furthermore, if the second text data TX'' (34-35) corresponding to the subtitles on the next screen consists of two voices and the attribute data corresponding to these two voices are approximate age = 20s, gender = female, and voice tone = high, in the language quality assurance process, the second text data TX'' (34-35) is converted to the high-pitched speech of a woman in her twenties and becomes the third text data TX''' (34-35). This third text data may be generated by inputting the second text data into a trained model for language quality assurance processing.
[0027] 2 and 5, the CPU 21 of the video editing device 20 performs a time stamp adjustment process. In the time stamp adjustment process, the CPU 21 displays a time stamp adjustment screen on the display 26 of the video editing device 20. The time stamp adjustment screen shows subtitles indicated by a series of third text data in the RAM 22 in association with the appearance timing of the subtitles indicated by the third text data in the video content. The operator can perform an operation to correct the appearance timing of the subtitles on the time stamp adjustment screen. When an operation to correct the appearance timing of the subtitles is performed, the CPU 21 changes the time stamp corresponding to the third text data in the RAM 22 in response to this operation.
[0028] 2 and 5, the CPU 21 of the video editing device 20 performs a post-editing process. In the post-editing process, the CPU 21 displays a post-editing screen on the display 26 of the video editing device 20. The post-editing screen shows subtitles indicated by a series of third text data in the RAM 22 and a correction field for the subtitles. The operator can perform an operation to correct the subtitles on the post-editing screen. When an operation to correct the subtitles is performed, the CPU 21 changes the corresponding third text data in the RAM 22 in response to this operation.
[0029] 2 and 5, the CPU 21 of the video editing device 20 performs a text-to-speech conversion process. In the text-to-speech conversion process, the CPU 21 converts a series of third text data in the RAM 22 into dubbed voice data corresponding to the data, and stores the dubbed voice data in the RAM 22. The conversion from text to dubbed voice may be performed by inputting the third text data into a trained model for conversion.
[0030] 2 and 5, the CPU 21 of the video editing device 20 performs an audio adjustment process. In the audio adjustment process, the CPU 21 displays an audio adjustment screen on the display 26 of the video editing device 20. The audio adjustment screen shows the sound waveform of the dubbed audio data in the RAM 22 and adjustment controls thereof. The operator can perform an operation to adjust the dubbed audio on the audio adjustment screen. When an operation to adjust the dubbed audio is performed, the CPU 21 changes the dubbed audio data in the RAM 22 in response to this operation.
[0031] In step S130 of FIG. 2, the terminal 10 performs a review and adjust process. In the review and adjust process, the terminal 10 receives image data, third text data, and dubbed audio data constituting one video content in the RAM 22 of the video editing device 20, and displays a review and adjust screen in which these data are arranged in chronological order. On the review and adjust screen, the user can perform an operation to correct the subtitles indicated by the third text data or the dubbed audio indicated by the dubbed audio data. When an operation to correct the subtitles or the dubbed audio is performed, the terminal 10 transmits data indicating the correction contents to the video editing device 20. When the CPU 21 of the video editing device 20 receives the data indicating the correction contents, the CPU 21 changes the third text data and the dubbed audio data in the RAM 22 according to the data.
[0032] In step S231 in FIG. 2 and FIG. 5, the CPU 21 of the video editing device 20 performs a video publishing process. In the video publishing process, the CPU 21 uploads the image data, the third text data, and the dubbed audio data constituting one video content in the RAM 22 to a specified video posting site. This upload may be performed as a video file of a subtitled version of the video content including the image data and the third text data, or as a video file of a dubbed version of the video content including the image data and the dubbed audio data. In addition, the image data, the third text data, and the dubbed audio data constituting the video content may be uploaded as a non-compressed file, or may be uploaded as a compressed and encoded file. There are two modes for uploading the subtitled version of the video content. In the first mode, the image data constituting the video content is a video file, the third text data is a subtitle file, and the image and the subtitle are uploaded as separate files. In the second mode, the third text data is written as a subtitle into the video file of the image data of the video content, and the image and the subtitle are uploaded as one file. In the first aspect, if the application playing the video file does not support the subtitle file, the subtitles cannot be displayed. However, in the second aspect, even if the application playing the video file does not support the subtitle file, the subtitles can be displayed.
[0033] The above is the details of this embodiment. The video editing device 20 of this embodiment includes a recognition means for recognizing each sound in a video file to be processed as first text data indicating the sound and attribute data of the person who uttered the sound, a translation means for translating the first text data into second text data in another language, and a modification means for generating third text data by modifying a part of the wording of the translation result indicated by the second text data based on the attribute data, and outputting the generated third text data together with data indicating an image in the video file. Thus, video content including dubbed audio information that easily conveys the speaker's feelings and nuances of words to the viewer can be generated from video content including audio in a foreign language.
[0034] Although the embodiment of the present invention has been described above, the following modifications may be made to this embodiment.
[0035] (1) In the above embodiment, the types of attributes recognized in the attribute recognition process are not limited to the approximate age, sex, and tone of voice of the person who uttered the voice. Attributes other than these may be specified. For example, the historical background of the person who uttered the voice (whether the person lived in the Middle Ages or early modern times, or whether the person lives in the present day) may be recognized as an attribute of the person.
[0036] (2) In the above embodiment, both the time stamp adjustment process and the post-editing process, or either one of them, may be omitted. In an embodiment in which neither the time stamp adjustment process nor the post-editing process is executed, the process immediately proceeds to the text-to-speech conversion process after the execution of the language quality assurance process, and the third text data that has been modified by the language quality assurance process is converted into dubbed voice data corresponding to that data.
[0037] (3) In the above embodiment, both or either of the audio adjustment process and the review & adjust process may not be executed. In an embodiment in which neither the audio adjustment process nor the review & adjust process is executed, immediately after the execution of the text-to-speech conversion process, the process proceeds to the video publishing process, and the image data, the third text data, and the dubbed audio data constituting the video content are uploaded to a designated video posting site. [Explanation of symbols]
[0038] 1. Video editing system 10 Terminal 20 Video editing equipment 21 CPU 22 RAM 23 ROM 24 Hard Disk 25 Communication Interface 27 Display 28 Keyboard 29 Pointing Device 29 Video Editing Programs 90 Network
Claims
1. a recognition means for recognizing each sound in a moving image file to be processed as first text data indicating the sound and attribute data of a person who uttered the sound; a translation means for translating the first text data into second text data in another language; a modification means for generating third text data by modifying a part of the wording of the translation result indicated by the second text data based on the attribute data, and outputting the generated third text data together with data indicating an image in the video file; A video editing device comprising:
2. The recognition means includes: and an attribute recognition processing means for performing an FFT process on the audio data of the audio in the video file and identifying the attributes of the person who emitted the audio based on the spectrum data obtained by the FFT process.
2. The video editing device according to claim 1.
3. The recognition means includes: a voice extraction means for extracting only voice data of a portion in which human vocalization is recorded from the voice data between the start and end of the video content in the video file; a division means for dividing the voice data extracted by the voice extraction means into a plurality of voice data each representing a group of vocal sounds; Equipped with The attribute recognition processing means identifies an attribute of a person who uttered the voice for each of the plurality of voice data divided by the dividing means.
3. The video editing apparatus according to claim 2.
4. On the computer, a recognition function for recognizing each sound in a moving image file to be processed as first text data indicating the sound and attribute data of a person who uttered the sound; a translation function for translating the first text data into second text data in another language; a modification function of generating third text data by modifying a part of the wording of the translation result indicated by the second text data based on the attribute data, and outputting the generated third text data together with data indicating an image in the video file; A program to achieve this.
Citation Information
Patent Citations
Voice translation device, method and program
JP2012073941A
Translation device, translation method, and translation program
JP2020134719A
Audio and video translator
US20220358905A1
Video translation platform
US20230325611A1
Speech processing device and speech processing method
WO2019202804A1