Video editing device and program

By identifying the voices in video files and determining the attributes of the characters, and using FFT to process the spectral data to adjust the wording of the translation results, the problem of monotonous wording in existing technologies is solved. This generates dubbing voice information that includes subtle differences in the speaker's emotions or language, thus enhancing the audience's experience.

CN122074151APending Publication Date: 2026-05-22WOVN TECH INC
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
WOVN TECH INC
Filing Date
2024-10-18
Publication Date
2026-05-22

AI Technical Summary

Technical Problem

Existing video dubbing technology struggles to effectively convey the speaker's emotions or subtle nuances in language, resulting in monotonous and repetitive wording in dialogue or speech.

Method used

By recognizing the sounds in the video files and determining the attributes of the characters, the FFT is used to process the spectral data, the wording of the translation results is adjusted to reflect the characteristics of the characters, the modified text data is generated, and it is output together with the video image data.

Benefits of technology

It enables the generation of dubbing audio information from foreign language video content, which includes the speaker's emotions or subtle differences in language, thus enhancing the audience's experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122074151A_ABST
    Figure CN122074151A_ABST
Patent Text Reader

Abstract

The invention provides a technical means capable of adjusting words according to attributes of a person who speaks lines or speech in video content. This video editing device is provided with: a recognition unit that recognizes each voice in a video file to be processed as first text data indicating the voice and attribute data of a person who makes the voice; the translation unit translates the first text data into second text data of another language; and a modification means for generating, on the basis of the attribute data, third text data obtained by modifying a part of the words of the translation result indicated by the second text data, and outputting the generated third text data together with data indicating an image in the video file.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to techniques for generating video files, and more particularly to techniques for generating video files containing subtitles or dubbing from video files containing foreign language audio. Background Technology

[0002] Patent documents 1, 2, and 3 disclose this technology. Patent document 1 describes a receiving device that buffers a multimedia presentation, extracts the audio and video components of the presentation, performs sound recognition analysis on the audio components to generate subtitle text data, and integrates this subtitle text data into the original video component for output. Patent document 2 describes a text transcription server that, based on the user's profile, translates the text generated by the sound recognition unit into the user's target language and sends the translated text as the translation result to the user's handheld terminal. Patent document 3 describes a machine translator that processes timed text tracks in an online video service, translates the timed text in the timed text track into a specified foreign language, and sends the content of the online video service with the rewritten timed text to the user's viewer client.

[0003] Patent Document 1: US Patent No. 8781824 Patent Document 2: Description of European Patent No. 2725816 Patent Document 3: U.S. Patent Application Publication No. 2010 / 0138209 However, the video content dubbing technology proposed to date has the following problems: the wording of the lines or remarks in the dubbed voices is all the same, making it difficult to convey the speaker's emotions or subtle differences in language to the audience. Summary of the Invention

[0004] This invention was made in view of this problem, and its purpose is to provide a technical means to adjust the wording according to the attributes of the person speaking the lines or remarks in the video content.

[0005] To address the aforementioned issues, a preferred embodiment of the present invention, a video editing apparatus, is characterized by comprising: a recognition unit that identifies each sound within a video file containing the object to be processed as first text data representing the sound and attribute data of the person emitting the sound; a translation unit that translates the first text data into second text data in another language; and a modification unit that, based on the attribute data, generates third text data that modifies a portion of the wording of the translation result shown in the second text data, and outputs the generated third text data together with data representing images within the video file.

[0006] In this method, the recognition unit may have an attribute recognition processing unit that performs FFT processing on the sound data in the video file and determines the attribute of the person making the sound based on the spectral data obtained through the FFT processing.

[0007] In addition, the recognition unit may include: a sound extraction unit, which extracts only the sound data recording a person's voice from the sound data between the beginning and end of the video content in the video file; and a segmentation unit, which segments the sound data extracted by the sound extraction unit into multiple sound data representing a group of voices, and the attribute recognition processing unit determines the attribute of the person making the voice for each of the multiple sound data segmented by the segmentation unit.

[0008] Another preferred embodiment of the present invention provides a program that enables a computer to perform the following functions: a recognition function, which identifies each sound in a video file of the object being processed as first text data representing the sound and attribute data of the person emitting the sound; a translation function, which translates the first text data into second text data in another language; and a modification function, which, based on the attribute data, generates third text data that modifies a portion of the wording of the translation result shown in the second text data, and outputs the generated third text data together with data representing images in the video file.

[0009] Invention Effects According to the present invention, it is possible to generate video content containing dubbing audio information that easily conveys the speaker's emotions or subtle nuances of language from video content containing foreign language audio. Attached Figure Description

[0010] Figure 1 This is a diagram illustrating the overall configuration of a video editing system according to an embodiment of the present invention.

[0011] Figure 2 It is shown Figure 1 The flowchart of the actions of the video editing system.

[0012] Figure 3 It is shown Figure 2 The diagram shows the processing steps S201 to S203.

[0013] Figure 4 It is shown Figure 2 A diagram showing the processing steps S204 to S210.

[0014] Figure 5 It is shown Figure 2 A diagram showing the processing content of S211~S214 in the diagram. Detailed Implementation

[0015] Hereinafter, one embodiment of the present invention will be described with reference to the accompanying drawings. Figure 1 This diagram illustrates the overall configuration of a video editing system 1 according to an embodiment of the present invention. The video editing system 1 generates a second video file from a first video file containing audio in a first language (e.g., English), converting the audio within the video content into subtitles or dubbing in a second language (e.g., Japanese).

[0016] The video editing system 1 includes a terminal 10 and a video editing device 20. The terminal 10 and the video editing device 20 are connected via a network 90. ​​The terminal 10 is a smartphone or personal computer used by a user. The video editing device 20 is a server operating under the management of a service provider. The video editing device 20 includes a CPU 21, RAM 22, ROM 23, hard disk 24, communication interface 25, display 26, keyboard 27, and pointing device 28. The CPU 21 uses RAM 22 as its working area and simultaneously executes programs stored in ROM 23 or hard disk 24. The hard disk 24 stores video editing programs 29.

[0017] The video editing program 29 has the following functions.

[0018] a1. Recognition function This is a function that receives the first video file of the processing object from the terminal 10 via the network 90, and identifies each sound of the first language in the first video file as first text data representing the sound and attribute data of the person making the sound.

[0019] b1. Translation function This is a function that translates first text data into second text data in a second language.

[0020] c1. Modification Functionality This is a function that generates a third text data based on attribute data, modifying part of the wording of the translation result shown in the second text data, and outputs the generated third text data together with data representing images within the video file.

[0021] Next, the operation of this embodiment will be explained. Figure 2 This is a flowchart illustrating the operation of this embodiment. Figure 3 It shows Figure 2 The processing content of steps S201 to S203 in the process. Figure 4 It shows Figure 2 The processing content of steps S204 to S210 in the process. Figure 5 It shows Figure 2The processing steps S211 to S214 in these diagrams are executed by the video editing program 29.

[0022] exist Figure 2 In step S100, terminal 10 performs upload processing. During upload processing, the user selects a desired video file from the memory of terminal 10. When this operation is performed, terminal 10 sends the selected video file as the first video file to be processed to video editing device 20. Upon receiving the first video file from terminal 10, the CPU 21 of video editing device 20 stores the first video file in RAM 22.

[0023] exist Figure 2 as well as Figure 3 In step S201, the CPU 21 of the video editing device 20 performs noise removal processing. In this noise removal processing, the CPU 21 decodes the first video file stored in RAM 22 and stores in RAM 22 a series of frame image data representing the video content from start to end, as well as sound data representing the sound waveforms of the video content from start to end. Noise is then removed from the sound data in RAM 22. This noise removal can be performed using a known noise removal procedure.

[0024] exist Figure 2 as well as Figure 3 In step S202, the CPU 21 of the video editing device 20 performs sound extraction processing. In this process, the CPU 21 extracts only the portion of the video content containing human speech (dialogue or remarks) from the noise-removed sound data stored in RAM 22, and stores this sound data in RAM 22. More specifically, it searches for audio tracks containing human speech from multiple audio tracks in the noise-removed sound data, extracts only the sound wave waveform recorded in those tracks, and stores this sound data in RAM 22 as the target for subsequent processing.

[0025] exist Figure 2 as well as Figure 3In step S203, the CPU 21 of the video editing device 20 performs sound segmentation processing. In this process, the CPU 21 divides the human voice data stored in the RAM 22 into multiple sound data representing a single set of voices. More specifically, in the sound segmentation process, the CPU 21 scans the entire sound waveform of the video content shown in the original sound data from the beginning to the end. Whenever a silence point appears, the CPU 21 repeatedly extracts the sound waveform before that point as a single sound data point. The multiple sound data obtained through this sound segmentation process represent a line of dialogue or a statement spoken by a character.

[0026] exist Figure 2 as well as Figure 4 In step S204, the CPU 21 of the video editing device 20 performs attribute recognition processing. In this processing, the CPU 21 identifies the attribute of the person emitting the sound for each of the multiple sound data in the RAM 22, and stores the attribute data representing the identified attribute in the RAM 22, corresponding to the sound data. This attribute includes the approximate age of the person emitting the sound (e.g., toddler, under 10 years old, 20s, 30s, 40s, 50s, over 60 years old), gender (e.g., female or male), and tone of voice (e.g., high, medium, low). More specifically, in the attribute recognition processing, an FFT (Fast Fourier transform) is applied to each sound data, and the attribute of the person emitting the sound is determined based on the spectral data obtained through the FFT. The determination of the attribute based on the spectral data can be performed by inputting the sound data into a learned model used for attribute recognition.

[0027] exist Figure 2 as well as Figure 4 In step S205, the CPU 21 of the video editing device 20 performs transcription processing. In the transcription processing, the CPU 21 performs cross-text processing on each sound data segmented by the sound segmentation processing in the RAM 22, and uses the cross-text processing result as the first text data representing the sound data. The first text data is then stored in the RAM 22 along with a timestamp indicating the time when the sound appears in the sound content.

[0028] exist Figure 2 as well as Figure 4 In step S206, the CPU 21 of the video editing device 20 performs dictionary adjustment processing. During this process, the words obtained from the transcription process in step S205 are registered in the dictionary database.

[0029] exist Figure 2 as well as Figure 4In step S207, the CPU 21 of the video editing device 20 performs reconstruction processing. In the reconstruction processing, the CPU 21 reconstructs the first text data in RAM 22 into a longer set of first text data. For example, if the first text data TX(30), TX(31), TX(32), TX(33), TX(34), TX(35) of six consecutive sounds in RAM 22 (the numbers in parentheses indicate the order of appearance of the sounds from the beginning of the video content) constitute a meaningful segment, the first text data TX(30), TX(31), TX(32), TX(33), TX(34), TX(35) of these six sounds are integrated into a single first text data TX'(30).

[0030] exist Figure 2 as well as Figure 4 In step S208, the CPU 21 of the video editing device 20 performs translation processing. In the translation processing, the CPU 21 translates the reconstructed first text data in RAM 22 into a second language, and stores the translation result in RAM 22 as second text data representing the translated subtitles of each segment of the video content shown in the original audio data from the beginning to the end.

[0031] exist Figure 2 as well as Figure 4 In step S209, the CPU 21 of the video editing device 20 performs segmentation processing. In the segmentation processing, the CPU 21 divides the translated second text data in RAM 22 into second text data corresponding to the subtitles displayed on the same image in the video content. For example, if the second text data TX'(30) of one segment in RAM 22 corresponds to a subtitle displayed in two frames, the second text data TX'(30) is divided into second text data TX''(30-33) corresponding to the subtitle of the previous frame and second text data TX''(34-35) corresponding to the subtitle of the next frame.

[0032] exist Figure 2 as well as Figure 4In step S210, the CPU 21 of the video editing device 20 performs Linguistic Quality Assurance (LQA) processing. In LQA processing, the CPU 21 generates third text data based on attribute data in RAM 22, modifying a portion of the wording of the translation result shown by the second text data in RAM 22, and stores the generated third text data in RAM 22. For example, if the second text data TX''(30-33) corresponding to the subtitle of one frame in RAM 22 consists of four voices, and the attribute data corresponding to these four voices is approximately age = 20+, gender = female, and tone = normal, then in LQA processing, the second text data TX''(30-33) is adjusted to the normal tone of a woman in her 20s, and this is used as the third text data TX'''(30-33). Furthermore, in the case that the second text data TX''(34-35) corresponding to the subtitle of the next screen consists of two voices, and the attribute data corresponding to these two voices are approximately age = 20+, gender = female, and pitch = high, in the language quality assurance processing, the second text data TX''(34-35) is adjusted to the high-pitched wording of a woman in her 20s, and used as the third text data TX'''(34-35). The generation of this third text data can be performed by inputting the second text data into the learned model used for language quality assurance processing.

[0033] exist Figure 2 as well as Figure 5 In step S211, the CPU 21 of the video editing device 20 performs timestamp adjustment processing. During timestamp adjustment processing, the CPU 21 causes the display 26 of the video editing device 20 to display a timestamp adjustment screen. The timestamp adjustment screen correspondingly displays subtitles represented by a series of third text data stored in RAM 22, and the appearance time of the subtitles represented by the third text data within the video content. The operator can correct the appearance time of the subtitles on the timestamp adjustment screen. When correcting the appearance time of the subtitles, the CPU 21 changes the timestamp in RAM 22 corresponding to the corresponding third text data according to the operation.

[0034] exist Figure 2 as well as Figure 5 In step S212, the CPU 21 of the video editing device 20 performs post-editing processing. During post-editing, the CPU 21 causes the display 26 of the video editing device 20 to show a post-editing screen. The post-editing screen displays subtitles and their correction bars based on a series of third text data stored in RAM 22. The operator can correct the subtitles on the post-editing screen. When correcting subtitles, the CPU 21 changes the corresponding third text data in RAM 22 according to the operation.

[0035] exist Figure 2 as well as Figure 5 In step S213, the CPU 21 of the video editing device 20 performs text-to-speech conversion processing. In this process, the CPU 21 converts a series of third text data in RAM 22 into corresponding voice-over audio data and stores the voice-over audio data in RAM 22. The conversion from text to voice-over audio can be performed by inputting the third text data into a learned model used for the conversion.

[0036] exist Figure 2 as well as Figure 5 In step S214, the CPU 21 of the video editing device 20 performs sound adjustment processing. During this processing, the CPU 21 causes the display 26 of the video editing device 20 to show a sound adjustment screen. The sound adjustment screen displays the sound waveform of the dubbing sound data stored in the RAM 22 and its adjustment control. The operator can adjust the dubbing sound on the sound adjustment screen. When adjusting the dubbing sound, the CPU 21 changes the dubbing sound data in the RAM 22 according to the operation.

[0037] exist Figure 2 In step S130, terminal 10 performs review and adjustment processing. During this process, terminal 10 receives image data, third text data, and dubbing audio data constituting a video content from the RAM 22 of video editing device 20, and displays a review and adjustment screen that arranges these data in chronological order. Users can modify the subtitles shown in the third text data or the dubbing audio shown in the dubbing audio data on the review and adjustment screen. When modifying subtitles or dubbing audio, terminal 10 sends data indicating the modification content to video editing device 20. Upon receiving the data indicating the modification content, CPU 21 of video editing device 20 modifies the third text data and dubbing audio data in RAM 22 according to that data.

[0038] exist Figure 2 as well as Figure 5In step S231, the CPU 21 of the video editing device 20 performs video publishing processing. During video publishing, the CPU 21 uploads the image data, third text data, and voice-over data constituting a video content stored in RAM 22 to a designated video submission website. This upload can be performed as a video file containing subtitled video content including image data and third text data, or as a video file containing dubbed video content including image data and voice-over data. Furthermore, the image data, third text data, and voice-over data constituting the video content can be uploaded as uncompressed files or as compressed files. There are two ways to upload subtitled video content. The first way is to upload the image data constituting the video content as a video file and the third text data as a subtitle file, uploading the image and subtitle as separate files. The second way is to write the third text data as subtitles into the video file containing the image data of the video content, uploading the image and subtitle as a single file. In the first method, subtitles cannot be displayed if the application playing the video file is incompatible with the subtitle file. However, in the second method, subtitles can be displayed even if the application playing the video file is incompatible with the subtitle file.

[0039] The above is a detailed description of this embodiment. The video editing apparatus 20 of this embodiment includes: a recognition unit that identifies each sound within a video file as first text data representing the sound and attribute data of the person uttering the sound; a translation unit that translates the first text data into second text data in another language; and a modification unit that, based on the attribute data, generates third text data that modifies a portion of the wording of the translation result shown in the second text data, and outputs the generated third text data along with data representing images within the video file. Therefore, it is possible to generate video content containing dubbed audio information that easily conveys the speaker's emotions or subtle linguistic nuances to the audience from video content containing foreign language sounds.

[0040] The embodiments of the present invention have been described above, but the following modifications may also be made to the embodiments.

[0041] (1) In the above embodiments, the types of attributes identified in the attribute recognition process are not limited to the approximate age, gender, and tone of voice of the person making the voice. Attributes other than these can also be identified. For example, the historical background of the person making the voice (whether they lived in the Middle Ages or the modern era, or a contemporary person) can also be used as an attribute for identification.

[0042] (2) In the above embodiments, it is also possible to omit either or both of the timestamp adjustment processing and post-editing processing. In the method of omitting both timestamp adjustment processing and post-editing processing, after the language quality assurance processing is performed, the text-to-speech conversion processing is immediately initiated, and the third text data with wording modifications after language quality assurance processing is converted into dubbing audio data corresponding to that data.

[0043] (3) In the above embodiments, it is also possible to omit either or both of the audio adjustment processing and the review and adjustment processing. In the method of not performing both audio adjustment processing and the review and adjustment processing, after the text-to-audio conversion processing is performed, the video publishing process is immediately initiated, and the image data, third-party text data, and dubbing audio data constituting the video content are uploaded to the designated video submission website.

[0044] Explanation of reference numerals in the attached figures 1…Video editing system; 10…Terminal; 20…Video editing device; 21…CPU; 22…RAM; 23…ROM; 24…Hard disk; 25…Communication interface; 26…Display; 27…Keyboard; 28…Pointing device; 29…Video editing program; 90…Network.

Claims

1. A video editing device, characterized in that, have: The recognition unit identifies each sound in the video file of the object being processed as first text data representing the sound and attribute data of the person emitting the sound; The translation unit translates the first text data into second text data in another language; as well as The modification unit generates third text data based on the attribute data, which modifies a portion of the wording of the translation result shown in the second text data, and outputs the generated third text data together with data representing images within the video file.

2. The video editing device according to claim 1, characterized in that, The recognition unit has an attribute recognition processing unit, which performs FFT processing on the sound data in the video file and determines the attribute of the person making the sound based on the spectrum data obtained through the FFT processing.

3. The video editing device according to claim 2, characterized in that, The identification unit has the following features: The sound extraction unit extracts only the sound data of the human voice from the sound data of the video content from the beginning to the end of the video file. as well as The segmentation unit divides the sound data extracted by the sound extraction unit into multiple sound data representing a group of sounds. The attribute recognition processing unit determines the attribute of the person who made the sound for each of the multiple sound data segments obtained by the segmentation unit.

4. A program that enables a computer to perform the following functions: The recognition function identifies each sound in the video file of the object being processed as first text data representing the sound and attribute data of the person emitting the sound; The translation function translates the first text data into a second text data in another language; as well as The modification function generates third text data based on the attribute data, which modifies a portion of the wording of the translation result shown in the second text data, and outputs the generated third text data together with data representing the images in the video file.