Video editing device and program

JP7917890B2Active Publication Date: 2026-09-09WOVN TECH INC
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
JP2023181239
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2023-10-20
Publication Date
2026-09-09
Estimated Expiration
2043-10-20

Smart Images

  • Figure 0007917890000001
    Figure 0007917890000001
  • Figure 0007917890000002
    Figure 0007917890000002
  • Figure 0007917890000003
    Figure 0007917890000003
Patent Text Reader

Abstract

To provide technical means for adjusting wording according to attributes of a person who uttered a line or discourse in video content.SOLUTION: A video editing device 20 includes: recognition means for recognizing each voice in a video file to be processed as first text data indicating the voice and attribute data of a person who uttered the voice; translation means for translating first text data into second text data in a different language; and modification means that generates third text data by modifying wording of a part of a translation result shown in the second text data based on the attribute data, and outputs the generated third text data together with data showing images in the video file.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a technology for generating video files, and particularly to a technology for generating a video file including subtitles and dubbed audio from a video file including foreign-language speech. [Background Art]

[0002] Patent Documents 1, 2 and 3 are documents disclosing this type of technology. The receiving device described in Patent Document 1 buffers a multimedia presentation, extracts an audio component and a video component of the multimedia presentation, performs speech recognition analysis on the audio component to generate subtitle text data, and integrates the subtitle text data with the original video component for output. The transcription server described in Patent Document 2 translates text generated by a speech recognition unit into a user's target language in accordance with a user profile, and transmits translated text as a translation result to the user's mobile terminal. The machine translator described in Patent Document 3 targets timed text tracks in an online video service, translates timed text in the timed text tracks into a predetermined foreign language, and transmits content of the online video service with the rewritten timed text to a user's viewer client. [Prior Art Documents] [Patent Documents]

[0003] [Patent Document 1] U.S. Patent No. 8781824 Specification [Patent Document 2] European Patent No. 2725816 Specification [Patent Document 3] U.S. Patent Application Publication No. 2010 / 0138209 Specification [Summary of the Invention] [Problem to be Solved by the Invention]

[0004] However, the video content dubbing technologies proposed so far have had the problem that the wording of the dialogue and discourse in the dubbed audio is uniform, making it difficult for viewers to understand the speaker's emotions and nuances of their words.

[0005] This invention has been made in view of these problems, and aims to provide a technical means that can adjust the wording according to the attributes of the person who speaks the lines or statements in the video content. [Means for solving the problem]

[0006] To solve the above problems, a video editing apparatus in a preferred embodiment of the present invention is characterized by comprising: recognition means for recognizing each sound in a video file to be processed as first text data representing the sound and attribute data of the person who uttered the sound; translation means for translating the first text data into second text data in another language; and modification means for generating third text data by modifying some of the wording of the translation result shown in the second text data based on the attribute data, and outputting the generated third text data together with data representing images in the video file.

[0007] In this embodiment, the recognition means may include attribute recognition processing means that applies FFT processing to the audio data of the audio in the video file and identifies the attributes of the person who spoke the audio based on the spectrum data obtained by the FFT processing.

[0008] Furthermore, the recognition means may include an audio extraction means that extracts only the portion of the audio data containing human speech from the audio data in the video file from the start to the end of the video content, and a splitting means that divides the audio data extracted by the audio extraction means into a plurality of audio data, each representing a single set of speech sounds, and the attribute recognition processing means may identify the attribute of the person who uttered the sound for each of the plurality of audio data divided by the splitting means.

[0009] Another preferred embodiment of the present invention provides a computer with a recognition function that recognizes each sound in a video file to be processed as first text data representing the sound and attribute data of the person who made the sound; a translation function that translates the first text data into second text data in another language; and a modification function that generates third text data by modifying some of the wording of the translation result shown in the second text data based on the attribute data, and outputs the generated third text data together with data representing images in the video file. [Effects of the Invention]

[0010] According to the present invention, it is possible to generate video content from video content containing foreign language audio that includes dubbed audio information that makes it easier for viewers to understand the speaker's emotions and the nuances of their words. [Brief explanation of the drawing]

[0011] [Figure 1] This diagram shows the overall configuration of a video editing system, which is one embodiment of the present invention. [Figure 2] Figure 1 is a flowchart showing the operation of the video editing system. [Figure 3] This figure shows the processing details for steps S201 to S203 in Figure 2. [Figure 4] This figure shows the processing details for steps S204 to S210 in Figure 2. [Figure 5] This figure shows the processing details for steps S211 to S214 in Figure 2. [Modes for carrying out the invention]

[0012] One embodiment of the present invention will be described below with reference to the drawings. Figure 1 is a diagram showing the overall configuration of a video editing system 1, which is one embodiment of the present invention. This video editing system 1 generates a second video file from a first video file containing audio in a first language (e.g., English), in which the audio in the video content is replaced with subtitles or dubbed audio in a second language (e.g., Japanese).

[0013] The video editing system 1 comprises a terminal 10 and a video editing device 20. The terminal 10 and the video editing device 20 are connected via a network 90. ​​The terminal 10 is a smartphone or personal computer used by the user. The video editing device 20 is a server operating under the management of the service provider. The video editing device 20 includes a CPU 21, RAM 22, ROM 23, hard disk 24, communication interface 25, display 26, keyboard 27, and pointing device 28. The CPU 21 executes programs stored in the ROM 23 and hard disk 24 while using the RAM 22 as a work area. The hard disk 24 stores a video editing program 29.

[0014] Video editing program 29 has the following functions: a1. Recognition function This function receives a first video file to be processed from terminal 10 via network 90, and recognizes each voice in the first language within the first video file as first text data representing the voice, and attribute data of the person who spoke the voice. b1. Translation function This is a function that translates first text data into second text data in a second language. c1. Modification function This function generates a third text data by modifying some of the wording in the translation result shown in the second text data, based on attribute data, and outputs the generated third text data along with data indicating the images in the video file.

[0015] Next, the operation of this embodiment will be described. Figure 2 is a flowchart of the operation of this embodiment. Figure 3 shows the processing content of steps S201 to S203 in Figure 2. Figure 4 shows the processing content of steps S204 to S210 in Figure 2. Figure 5 shows the processing content of S211 to S214 in Figure 2. Steps S201 to S214 and S231 in these figures are executed by the video editing program 29.

[0016] In step S100 of FIG. 2, the terminal 10 performs an upload process. In the upload process, a user performs an operation of selecting a desired video file from within the memory of the terminal 10. When this operation is performed, the terminal 10 transmits the selected video file as a first video file to be processed to the video editing apparatus 20. When the CPU 21 of the video editing apparatus 20 receives the first video file from the terminal 10, the CPU 21 stores the first video file in the RAM 22.

[0017] In step S201 of FIG. 2 and FIG. 3, the CPU 21 of the video editing apparatus 20 performs a noise removal process. In the noise removal process, the CPU 21 decodes the first video file in the RAM 22, stores, in the RAM 22, a series of frame image data from the start to the end of the decoded video content and audio data indicating a sound waveform of sound from the start to the end of the video content, and removes noise from the audio data in the RAM 22. This noise removal may preferably be performed by a known noise removal program.

[0018] In step S202 of FIG. 2 and FIG. 3, the CPU 21 of the video editing apparatus 20 performs an audio extraction process. In the audio extraction process, the CPU 21 extracts only audio data of a portion where human utterances (such as lines and speeches) are recorded in the video content from the noise-removed audio data in the RAM 22, and stores this audio data in the RAM 22. More specifically, the CPU 21 searches for a track on which human utterances are recorded from a plurality of audio tracks in the noise-removed audio data, extracts only the audio data of the sound waveform recorded on this track, and stores the extracted audio data in the RAM 22 as audio data to be subjected to subsequent processing.

[0019] In step S203 of Figures 2 and 3, the CPU 21 of the video editing device 20 performs audio splitting. In audio splitting, the CPU 21 divides the audio data of human speech in RAM 22 into multiple audio data, each representing a single set of speech. More specifically, in audio splitting, the CPU 21 scans the entire sound waveform from the beginning to the end of the video content represented by the original audio data, and each time a point of silence appears, it cuts out the sound waveform before that point and combines it into a single audio data. This process is repeated. The multiple audio data obtained through this audio splitting process represent a single line of dialogue or statement spoken by a single person.

[0020] In step S204 of Figures 2 and 4, the CPU 21 of the video editing device 20 performs attribute recognition processing. In attribute recognition processing, the CPU 21 recognizes the attributes of the person who made the sound indicated by each of the multiple audio data in RAM 22, and stores attribute data indicating the recognized attributes in RAM 22 in association with the audio data. These attributes include the approximate age of the person who made the sound (e.g., infant, under 10, 20s, 30s, 40s, 50s, 60s or older), gender (e.g., female or male), and voice tone (e.g., high, normal, low). More specifically, in attribute recognition processing, an FFT (Fast Fourier transform) is applied to each audio data, and the attributes of the person who made the sound are identified based on the spectrum data obtained by the FFT. Identification of attributes based on spectrum data is preferably performed by inputting the audio data into a trained model for attribute recognition.

[0021] In step S205 of Figures 2 and 4, the CPU 21 of the video editing device 20 performs transcription processing. In transcription processing, the CPU 21 applies character transcription processing to each of the audio data divided by the audio division processing in RAM 22, and uses the result of the character transcription processing as first text data representing the audio of the audio data. This first text data is stored in RAM 22 in association with a timestamp indicating the timing of the appearance of the audio in the audio content.

[0022] In step S206 of Figures 2 and 4, the CPU 21 of the video editing device 20 performs dictionary adjustment processing. In the dictionary adjustment processing, the words obtained by the transcription processing in step S205 are registered in the dictionary library.

[0023] In step S207 of Figures 2 and 4, the CPU 21 of the video editing device 20 performs a reconstruction process. In the reconstruction process, the CPU 21 reconstructs the first text data in RAM 22 into longer chunks of first text data. For example, if six consecutive audio first text data TX(30), TX(31), TX(32), TX(33), TX(34), TX(35) (the numbers in parentheses indicate the order in which the audio appears from the beginning of the video content) in RAM 22 form a meaningful paragraph, the CPU 21 combines these six audio first text data TX(30), TX(31), TX(32), TX(33), TX(34), TX(35) into a single first text data TX'(30).

[0024] In step S208 of Figures 2 and 4, the CPU 21 of the video editing device 20 performs translation processing. In the translation processing, the CPU 21 translates the reconstructed first text data in RAM 22 into a second language and stores the translation result in RAM 22 as second text data that shows the translated subtitles for each paragraph from the start to the end of the video content indicated by the original audio data.

[0025] In step S209 of Figures 2 and 4, the CPU 21 of the video editing device 20 performs paragraph splitting. In paragraph splitting, the CPU 21 splits the second text data, which has been translated in RAM 22, into second text data corresponding to subtitles to be displayed on the same image in the video content. For example, if the second text data TX'(30) of one paragraph in RAM 22 corresponds to subtitles that are displayed on two separate screens, the CPU 21 splits this second text data TX'(30) into second text data TX''(30-33) corresponding to the subtitle on the first screen and second text data TX''(34-35) corresponding to the subtitle on the second screen.

[0026] In step S210 of Figures 2 and 4, the CPU 21 of the video editing device 20 performs linguistic quality assurance (LQA) processing. In linguistic quality assurance processing, the CPU 21 generates third text data based on attribute data in RAM 22, modifying some of the wording of the translation result shown by the second text data in RAM 22, and stores the generated third text data in RAM 22. For example, if the second text data TX''(30-33) in RAM 22, which corresponds to the subtitles of one screen, consists of four voices, and the attribute data corresponding to these four voices is roughly age = 20s, gender = female, voice tone = normal, then in linguistic quality assurance processing, the second text data TX''(30-33) is modified to use the normal tone of a woman in her 20s, and this modified version becomes the third text data TX'''(30-33). Furthermore, if the second text data TX''(34-35) corresponding to the subtitles on the next screen consists of two audio clips, and the attribute data corresponding to these two audio clips is roughly age = 20s, gender = female, voice tone = high, then in the language quality assurance processing, the second text data TX''(34-35) is modified to use the high-pitched speech of a woman in her 20s, and this becomes the third text data TX'''(34-35). This third text data can be generated by inputting the second text data into a trained model for language quality assurance processing.

[0027] In step S211 of Figures 2 and 5, the CPU 21 of the video editing device 20 performs a timestamp adjustment process. During the timestamp adjustment process, the CPU 21 displays a timestamp adjustment screen on the display 26 of the video editing device 20. The timestamp adjustment screen shows the correspondence between the subtitles indicated by a series of third text data in RAM 22 and the timing of the appearance of those subtitles in the video content. The operator can perform an operation to correct the timing of the subtitle appearance on the timestamp adjustment screen. When an operation to correct the timing of the subtitle appearance is performed, the CPU 21 changes the timestamp corresponding to the relevant third text data in RAM 22 in response to this operation.

[0028] In step S212 of Figures 2 and 5, the CPU 21 of the video editing device 20 performs post-editing. During post-editing, the CPU 21 displays a post-editing screen on the display 26 of the video editing device 20. The post-editing screen shows subtitles indicated by a series of third text data in RAM 22 and a field for editing them. The operator can edit the subtitles on the post-editing screen. When an operation to edit the subtitles is performed, the CPU 21 changes the corresponding third text data in RAM 22 in response to this operation.

[0029] In step S213 of Figures 2 and 5, the CPU 21 of the video editing device 20 performs text-to-speech conversion processing. In text-to-speech conversion processing, the CPU 21 converts a series of third text data in RAM 22 into corresponding dubbed audio data and stores this dubbed audio data in RAM 22. The conversion from text to dubbed audio is preferably performed by inputting the third text data into a trained model for conversion.

[0030] In step S214 of Figures 2 and 5, the CPU 21 of the video editing device 20 performs audio adjustment processing. During audio adjustment processing, the CPU 21 displays an audio adjustment screen on the display 26 of the video editing device 20. The audio adjustment screen shows the sound waveform of the dubbed audio data in RAM 22 and its adjustment controls. The operator can adjust the dubbed audio on the audio adjustment screen. When an operation to adjust the dubbed audio is performed, the CPU 21 modifies the dubbed audio data in RAM 22 in response to this operation.

[0031] In step S130 of Figure 2, terminal 10 performs a review and adjust process. During the review and adjust process, terminal 10 receives image data, third text data, and dubbed audio data that constitute a single video content from the RAM 22 of the video editing device 20, and displays a review and adjust screen in which this data is arranged in chronological order. On the review and adjust screen, the user can modify the subtitles indicated by the third text data or the dubbed audio indicated by the dubbed audio data. When the user modifies the subtitles or dubbed audio, terminal 10 sends data indicating the modifications to the video editing device 20. When the CPU 21 of the video editing device 20 receives the data indicating the modifications, it modifies the third text data and dubbed audio data in the RAM 22 according to this data.

[0032] In step S231 of Figures 2 and 5, the CPU 21 of the video editing device 20 performs video publishing processing. In video publishing processing, the CPU 21 uploads the image data, third text data, and dubbed audio data that constitute a single video content from RAM 22 to a designated video posting site. This upload may be performed as a video file of the subtitled video content, including the image data and third text data, or as a video file of the dubbed video content, including the image data and dubbed audio data. Furthermore, the image data, third text data, and dubbed audio data that constitute the video content may be uploaded as uncompressed files or as compressed encoded files. There are two modes of uploading subtitled video content. The first mode is to upload the image data that constitutes the video content as a video file and the third text data as a subtitle file, and upload the image and subtitle as separate files. The second mode is to write the third text data as subtitles to the video file of the video content image data, and upload the image and subtitle as a single file. In the first embodiment, if the application playing the video file does not support subtitle files, subtitles cannot be displayed. However, in the second embodiment, subtitles can be displayed even if the application playing the video file does not support subtitle files.

[0033] The above describes the details of this embodiment. The video editing device 20 of this embodiment includes recognition means that recognize each audio in the video file to be processed as first text data representing the audio and attribute data of the person who spoke the audio; translation means that translate the first text data into second text data in another language; and modification means that generate third text data by modifying some of the wording of the translation result shown in the second text data based on the attribute data, and output the generated third text data together with data representing images in the video file.Therefore, it is possible to generate video content that includes dubbed audio information that makes it easier for viewers to understand the emotions and nuances of the speaker's words from video content that includes audio in a foreign language.

[0034] The embodiments of the present invention have been described above, but the following modifications may be made to these embodiments.

[0035] (1) In the above embodiment, the types of attributes recognized by the attribute recognition process are not limited to the approximate age, gender, and tone of voice of the person who made the sound. Other attributes may be identified. For example, the historical context of the person who made the sound (whether they are a living person from the Middle Ages or early modern period, or a person living in the present day) may be recognized as an attribute of that person.

[0036] (2) In the above embodiment, it is not necessary to perform both the timestamp adjustment process and the post-editing process, or to perform either one of them. In an embodiment in which neither the timestamp adjustment process nor the post-editing process is performed, immediately after the language quality assurance process is performed, the process proceeds to the text-to-speech conversion process, and the third text data, which has undergone wording modification by the language quality assurance process, is converted into the corresponding dubbed audio data.

[0037] (3) In the above embodiment, it is not necessary to perform both the audio adjustment process and the review and adjustment process, or to perform either one of them. In an embodiment in which neither the audio adjustment process nor the review and adjustment process is performed, immediately after the text-to-speech conversion process, the process proceeds to the video publication process, and the image data, third text data, and dubbed audio data constituting the video content are uploaded to the designated video posting site. [Explanation of Symbols]

[0038] 1. Video editing system 10 devices 20 Video editing equipment 21 CPU 22 RAM 23 ROM 24 hard disks 25 Communication Interfaces 27 displays 28 keyboards 29 Pointing devices 29 Video Editing Programs 90 Networks

Claims

1. A recognition means that extracts only the audio data containing human speech from the audio data of the video content within the video file to be processed, divides the extracted audio data into multiple audio data, each representing a single speech sound, recognizes each of the divided audio data as first text data representing the speech, applies FFT processing to the audio data, and identifies attribute data indicating the approximate age, gender, and tone of voice of the person who uttered the speech based on the spectrum data obtained by the FFT processing. A translation means that reconstructs the first text data into longer chunks of first text data, and then translates them into second text data in another language. Modification means that divides the second text data into second text data corresponding to subtitles to be displayed on the same image in the video file, generates third text data with some of the wording of the translation result shown in the second text data modified based on the attribute data corresponding to each audio contained in the divided second text data, and outputs the generated third text data together with data showing the image in the video file. A video editing device characterized by comprising the following:

2. The aforementioned recognition means is The system includes attribute recognition processing means that identifies the attribute data based on the spectrum data by inputting the audio data into a trained model for attribute recognition. The video editing device according to feature 1.

3. The aforementioned recognition means is The video editing apparatus according to claim 1 or 2, characterized in that each of the divided audio data is subjected to transcription processing, and the first text data, which is the result of the transcription processing, is associated with a timestamp indicating the timing of the appearance of the audio in the video content.

4. On the computer, A recognition function that extracts only the audio data containing human speech from the audio data of the video content from the beginning to the end of the video file to be processed, divides the extracted audio data into multiple audio data, each representing a single speech sound, recognizes each of the divided audio data as first text data representing the speech, applies FFT processing to the audio data, and identifies attribute data indicating the approximate age, gender, and tone of voice of the person who uttered the speech based on the spectrum data obtained by the FFT processing. A translation function that reconstructs the aforementioned first text data into longer chunks of first text data, and then translates them into second text data in another language. The modification function divides the second text data into second text data corresponding to subtitles to be displayed on the same image in the video file, generates third text data with some of the wording of the translation result shown in the second text data modified based on the attribute data corresponding to each audio contained in the divided second text data, and outputs the generated third text data together with data indicating the image in the video file. A program that makes this possible.

Citation Information

Patent Citations

  • A method and system for generating subtitles

    EP2725816A1

  • Voice translation device, method and program

    JP2012073941A

  • Translation device, translation method, and translation program

    JP2020134719A

  • System and Method for Translating Timed Text in Web Video

    US20100138209A1

  • Audio and video translator

    US20220358905A1