Information processing apparatus, information processing method, and program
By extracting speech segments, detecting correspondences, and adjusting positions in the information processing device, the problem of misalignment between speech and video after dubbing was solved, achieving higher alignment accuracy and lower inconsistency.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SONY GROUP CORP
- Filing Date
- 2024-10-09
- Publication Date
- 2026-05-22
AI Technical Summary
In existing technologies, the misalignment between audio and video after dubbing causes a sense of disharmony for viewers, and it is necessary to reduce the misalignment between audio and video after dubbing.
The speech segment extraction unit, correspondence detection unit, and position adjustment unit in the information processing device detect and adjust the position of the dubbed speech segments to align them with the original speech segments.
It improves the alignment accuracy during voice dubbing, reduces misalignment between voice and video after dubbing, and reduces the sense of disharmony for viewers.
Smart Images

Figure CN122074153A_ABST
Abstract
Description
Technical Field
[0001] This technology relates to information processing apparatus, information processing methods and programs, and particularly to technology for dubbing video. Background Technology
[0002] Traditionally, in video content such as movies and plays, the original audio added to the video is dubbed into another language. Additionally, for example, Patent Document 1 proposes a technique for modifying video so that the mouth movements of the dubbed actor match the mouth movements of the voice actor speaking the dubbed language.
[0003] Reference List
[0004] Patent documents
[0005] Patent Document 1: JP-H08-6182-A Summary of the Invention
[0006] Technical issues
[0007] In the technology disclosed in Patent Document 1, the processing becomes complex because the video is modified to match the dubbed audio, and there is a possibility of causing a sense of incongruity for the viewer when watching the video. Therefore, there is a need for a method to further reduce the misalignment between the dubbed audio and the video.
[0008] This technology was developed in light of this situation, and its purpose is to reduce the misalignment between voice and video after dubbing.
[0009] Solution to the problem
[0010] The information processing apparatus according to the present technology includes: a speech segment extraction unit that extracts speech segments from each of speech data in a first language added to video data and speech data in a second language after dubbing; a correspondence detection unit that detects a correspondence between the speech segments in the first language and the speech segments in the second language; and a position adjustment unit that adjusts the position of the speech segments in the second language that have been detected to correspond to the speech segments in the first language relative to the speech segments in the first language.
[0011] According to the information processing device, by adjusting the position of the second language speech segment, which has been detected to correspond with the first language speech segment, relative to the first language speech segment, the position of the second language speech segment can be more accurately aligned with the first language speech segment. Attached Figure Description
[0012] Figure 1 This is a diagram showing an overview of an information processing device.
[0013] Figure 2 This is a diagram used to illustrate speech extraction and processing.
[0014] Figure 3 This is a diagram used to illustrate the correspondence detection process.
[0015] Figure 4 This is a diagram illustrating a specific example of correspondence detection processing.
[0016] Figure 5 This is a diagram illustrating an example of correspondence detection.
[0017] Figure 6 This is a diagram used to illustrate an example of position adjustment processing.
[0018] Figure 7 This is a diagram illustrating an example of how the intersection-over-union ratio is calculated.
[0019] Figure 8 This is a flowchart illustrating the process of voice position adjustment.
[0020] Figure 9 This is a diagram used to illustrate the correspondence detection process in the first modified example.
[0021] Figure 10 This is a diagram used to illustrate the correspondence detection process in the first modified example.
[0022] Figure 11 This is a diagram used to illustrate the correspondence detection process in the second modified example.
[0023] Figure 12 This is a flowchart illustrating the process of voice position adjustment in the third modified example.
[0024] Figure 13 This is a flowchart illustrating the process of voice position adjustment in the fourth modified example.
[0025] Figure 14 This is a flowchart illustrating the process of voice position adjustment in the fourth modified example. Detailed Implementation
[0026] Hereinafter, embodiments of the sensor device according to the present technology will be described in the following order with reference to the accompanying drawings.
[0027] <1. Configuration of Information Processing Devices>
[0028] <2. Voice Position Adjustment Processing>
[0029] <3. Modified Example>
[0030] <4. Summary of Examples>
[0031] <5. This technology>
[0032] <1. Configuration of Information Processing Devices>
[0033] Figure 1 This is a diagram showing an outline of the information processing device 1. For example, the information processing device 1 is a computer device such as a personal computer, workstation, mobile terminal device (such as a smartphone or tablet computer), or a server device configured in cloud computing.
[0034] The information processing device 1 includes a CPU (central processing unit) 11, a ROM (read-only memory) 12, a RAM (random access memory) 13, and a non-volatile memory section 14 such as an EEP-ROM (electrically erasable programmable read-only memory).
[0035] The CPU 11 performs various types of processing based on programs stored in the ROM 12 or the non-volatile memory section 14, or programs loaded from the storage section 19 onto the RAM 13.
[0036] RAM 13 also appropriately stores data required by CPU 11 to perform various types of processing.
[0037] As functional units related to the processing mentioned later, the CPU 11 includes a speech segment extraction unit 31, a correspondence detection unit 32, and a position adjustment unit 33. These functional units operate based on a software program that starts in the CPU 11. These functional units will be described in detail later.
[0038] Note that in addition to or together with CPU 11, you can also set up a GPU (Graphics Processing Unit), GPGPU (General-purpose computing on graphics processing units), AI (Artificial Intelligence) processor, etc.
[0039] The CPU 11, ROM 12, RAM 13, and non-volatile memory section 14 are interconnected via bus 23. Bus 23 is also connected to input / output interface 15.
[0040] The input / output interface 15 is connected to the input section 16, which includes control elements, operating devices, etc. For example, various types of control elements or operating devices such as keyboards, mice, buttons, dial pads, touch panels, touchpads, or remote controls can be used as the input section 16.
[0041] The input unit 16 detects user operations and outputs signals corresponding to the input operations to the CPU 11.
[0042] A microphone can also be used as the input unit 16. Voice input by the user can also be used as operation information.
[0043] In addition, the input / output interface 15 is integrated or separately connected to the display unit 17 and the sound output unit 18.
[0044] Display unit 17 performs various types of displays, including LCD (Liquid Crystal Display) and organic EL (electro-luminescence) panels. Display unit 17 displays various types of images on the display screen based on instructions from CPU 11. In addition, display unit 17 also performs the display of operation menus, icons, messages, etc., that is, as a GUI (Graphical User Interface) display.
[0045] The sound output unit 18 outputs sound and includes a speaker, etc.
[0046] The input / output interface 15 is connected in some cases to a storage unit 19, including an HDD (Hard Disk Drive), an SSD (Solid State Drive), etc., and a communication unit 20 for communicating with other devices.
[0047] Storage unit 19 can store various types of data and programs.
[0048] The communication unit 20 performs communication processing via a transmission path such as the Internet, and performs communication through near-field communication, wired communication with peripheral devices, bus communication, etc.
[0049] Additionally, the input / output interface 15 is connected to the drive 21 as needed, and a removable recording medium 22, such as a disk, optical disk, magneto-optical disk, memory card, or USB memory, is appropriately attached to the input / output interface 15.
[0050] The information processing device 1 can acquire video content via network communication through the communication unit 20. Furthermore, the information processing device 1 can use the driver 21 to read video content recorded on the removable recording medium 22.
[0051] The acquired video content is stored in the storage unit 19, and the video included in the video content is displayed on the display unit 17, or the sound included in the video content is output from the sound output unit 18.
[0052] Alternatively, for example, the processing program for this embodiment can be installed on the information processing device 1 via network communication using the communication unit 20 or the removable recording medium 22. Alternatively, the program can be pre-stored in the ROM 12, storage unit 19, etc.
[0053] <2. Voice Position Adjustment Processing>
[0054] Next, the voice position adjustment processing performed by the CPU 11 of the information processing device 1 will be explained.
[0055] Video content such as movies and plays includes not only video data but also audio data such as actors' lines. Among the audio data included in video content are spoken words in a specific language (such as English).
[0056] When voice data is dubbed into a language different from the given language (e.g., Japanese), the process begins by having voice actors dub the lines in the dubbed language while watching the video on location (e.g., in a studio). Then, the dubbed voice data is added to the video content, resulting in video content with the added voice data.
[0057] Note that video content can include both audio data before dubbing and audio data after dubbing.
[0058] In voice dubbing performed in this manner, slight misalignment may occur between the recorded speech and video because the voice-over is collected while the video is being watched, and the characteristics of the language before and after dubbing differ. This misalignment between video and speech unintentionally creates a sense of dissonance for viewers of the video content.
[0059] Therefore, in the information processing device 1, the alignment accuracy during voice dubbing is improved by adjusting the position of the dubbed speech. As a result, in the information processing device 1, the misalignment between the recorded dubbed speech and the video can be reduced, and the sense of disharmony caused to viewers of the video content can be reduced.
[0060] Note that the location of speech refers to the time at which the speech is reproduced, i.e., the time location.
[0061] The following example illustrates a scenario where the pre-dubbing audio (language) added to the video content is in English and the post-dubbing audio (language) is in Japanese. The English audio is referred to as English audio, and the Japanese audio as Japanese audio.
[0062] First, the CPU 11 acquires video content including video data and English voice data, as well as Japanese voice data with dubbing. As a method of acquisition, data can be acquired from other devices via the communication unit 20, or data can be acquired from the removable recording medium 22 via the driver 21.
[0063] Note that because the voice recording of the voice actor is performed once or multiple times for Japanese voice data, in some cases the Japanese voice data is a single data point, or in other cases the Japanese voice data includes multiple separate data points.
[0064] [2.1. Speech Extraction and Processing]
[0065] Figure 2 This is a diagram used to illustrate speech extraction processing. The speech segment extraction unit 31 performs speech extraction processing to extract speech segments corresponding to the segments being spoken by actors, etc., from English and Japanese speech data. In speech extraction processing, such as... Figure 2 As shown, a known speech extraction method for the speech waveform 41 of English and Japanese speech data is used to extract the segment of speech that is being spoken as the speech segment 42.
[0066] In addition, from the English speech data and the Japanese speech data, the speech segment extraction unit 31 extracts the following segments as silence segments 43, which are segments sandwiched between adjacent speech segments 42 and in which no speech is taking place.
[0067] Note that in the following text, the speech waveform 41 of the English speech data is referred to as English speech waveform 51, and the speech waveform 41 of the Japanese speech data is referred to as Japanese speech waveform 61. Without distinguishing between English speech waveform 51 and Japanese speech waveform 61, they are referred to as speech waveform 41.
[0068] Furthermore, the speech segment 42 of the English speech data is referred to as English speech segment 52, and the speech segment 42 of the Japanese speech data is referred to as Japanese speech segment 62. Without distinguishing between English speech segment 52 and Japanese speech segment 62, they are simply referred to as speech segment 42.
[0069] Furthermore, the silence segment 43 of the English speech data is referred to as the English silence segment 53, and the silence segment 43 of the Japanese speech data is referred to as the Japanese silence segment 63. Without distinguishing between the English silence segment 53 and the silence segment 43, they are simply referred to as silence segment 43.
[0070] [2.2. Correspondence Detection and Processing]
[0071] Figure 3 This is a diagram used to illustrate the correspondence detection process. Note that... Figure 3 This illustrates the case where two Japanese speech data sets exist (referred to as Japanese Speech 1 and Japanese Speech 2). Additionally, Figure 3 The volume waveform 44 and voice waveform 41 for each voice data are shown.
[0072] Then, the volume waveform 44 of the English speech data is referred to as the English volume waveform 54, and the volume waveform 44 of the Japanese speech data is referred to as the Japanese volume waveform 64. Without distinguishing between the English volume waveform 54 and the Japanese volume waveform 64, they are simply referred to as volume waveform 44.
[0073] For the English speech segment 52 and the Japanese speech segment 62 extracted by the speech segment extraction unit 31, the correspondence detection unit 32 performs correspondence detection processing to detect the correspondence between the English speech segment 52 and the Japanese speech segment 62.
[0074] Here, the correspondence refers to the relationship between the English speech segment 52, which includes English speaking parts, and the Japanese speech segment 62, which includes Japanese speech that has been dubbed for these speaking parts. In other words, the correspondence refers to the relationship between the English speech segment 52 and the Japanese speech segment 62, which includes Japanese speech that has been dubbed for the English speech in the English speech segment 52.
[0075] exist Figure 3 In the example shown, English speech segment 52a was detected as corresponding to Japanese speech segments 62a, 62b, and 62c in Japanese speech 1. Additionally, English speech segment 52b was detected as corresponding to Japanese speech segment 62d in Japanese speech 1.
[0076] Additionally, English speech segment 52c was detected to correspond to Japanese speech segment 62f in Japanese speech segment 2. Furthermore, English speech segment 52d was detected to correspond to Japanese speech segment 62g in Japanese speech segment 2.
[0077] In addition, English speech segment 52e was detected to correspond to Japanese speech segment 62e in Japanese speech 1 and Japanese speech segments 62h and 62i in Japanese speech 2.
[0078] In this way, the correspondence between English phonetic segment 52 and Japanese phonetic segment 62 is not limited to a one-to-one relationship, but includes an N-to-M relationship (N and M are integers equal to or greater than 1).
[0079] Figure 4 This is a diagram illustrating a specific example of the correspondence detection process. Note that... Figure 4 The English volume waveform 54 and the Japanese voice volume waveform 64 are shown for illustration.
[0080] Here, since neither English nor Japanese is spoken in the sections without dialogue, they become silent sections 43 at almost the same timing and for almost the same duration. Therefore, in a specific example of the correspondence detection processing, such as... Figure 4 As shown, the correspondence detection unit 32 detects the correspondence between the English speech segment 52 and the Japanese speech segment 62 relative to the silence segment 43.
[0081] In view of this, firstly, based on the Intersection over Union (IoU) of the Japanese silence segment 63 and the English silence segment 53, the correspondence detection unit 32 detects the Japanese silence segment 63 that has a correspondence with the English silence segment 53.
[0082] Specifically, the correspondence detection unit 32 sets an English silence segment 53 as the detection target and selects a Japanese silence segment 63 located within a predetermined range before and after the English silence segment 53. Then, the correspondence detection unit 32 calculates the intersection-union ratio (IUGR) between the English silence segment 53 of the detection target and the selected Japanese silence segment 63.
[0083] Note that the correspondence refers to the relationship between the English silence segment 53, which includes the English silence part, and the Japanese silence segment 63 corresponding to that silence part.
[0084] The Cross-Union Ratio (CUI) represents the degree of matching in position and width between the detected English silence segment 53 and the selected Japanese silence segment 63. The CUI is 1 when the English silence segment 53 and the Japanese silence segment 63 match in position and width and completely overlap. The CUI is 0 when the English silence segment 53 and the Japanese silence segment 63 do not overlap at all. Therefore, a higher CUI indicates a stronger correspondence between the detected English silence segment 53 and the selected Japanese silence segment 63.
[0085] Note that the width of the silent section 43 refers to the duration of the silent section 43.
[0086] The correspondence detection unit 32 detects Japanese silence segments 63 whose intersection-over-union ratio with the English silence segment 53 of the detection target is equal to or greater than a predetermined value (e.g., 0.5) as Japanese silence segments 63 that correspond to the English silence segment 53 of the detection target.
[0087] The correspondence detection unit 32 detects Japanese silence segments 63 that correspond to each of all English silence segments 53.
[0088] exist Figure 4 In the example, a Japanese silence segment 63a, which corresponds to the English silence segment 53a, was detected. Additionally, a Japanese silence segment 63c, which corresponds to the English silence segment 53b, was detected. Furthermore, a Japanese silence segment 63d, which corresponds to the English silence segment 53c, was detected.
[0089] Then, the Japanese silence segment 63b between Japanese silence segment 63a and Japanese silence segment 63c was determined to have no correspondence with any of the English silence segments 53.
[0090] Figure 5 This is a diagram illustrating an example of correspondence detection.
[0091] After detecting a Japanese silence segment 63 that corresponds to the English silence segment 53, the correspondence detection unit 32 detects the English speech segment 52 located between adjacent English silence segments 53 and the Japanese speech segment 62 located between the Japanese silence segments 63 that correspond to the English silence segments 53 as a combination that has a correspondence.
[0092] exist Figure 5 In the example, the English speech segment 52f located between the English silence segments 53a and 53b, and the Japanese speech segments 62j and 62k located between the Japanese silence segments 63a and 63c, were detected as corresponding combinations.
[0093] In addition, the English speech segment 52g, located between the English silence segment 53b and the English silence segment 53c, and the Japanese speech segment 62l, located between the Japanese silence segment 63c and the Japanese silence segment 63d, were detected as a corresponding combination.
[0094] In this way, during the correspondence detection process, the correspondence between English speech segment 52 and Japanese speech segment 62 is detected relative to the silence segment 43 of English and Japanese speech. Therefore, even when the speech segments 42 of English and Japanese speech have an N-to-M relationship, the correspondence can be accurately detected, and the computational load can be reduced.
[0095] [2.3. Position Adjustment Processing]
[0096] Figure 6 This is a diagram used to illustrate an example of position adjustment processing.
[0097] When the correspondence detection unit 32 has detected the correspondence between the English speech segment 52 and the Japanese speech segment 62, the position adjustment unit 33 performs a position adjustment process, which adjusts the position of the Japanese speech segment 62 relative to the English speech segment 52 in the front-back direction. Note that the adjustment of the position of the Japanese speech segment 62 strictly refers to the adjustment of the position of the Japanese speech within the Japanese speech segment 62.
[0098] For example, such as Figure 6 As shown, when the English speech segment 52f and the Japanese speech segments 62j and 62k are detected as a corresponding combination, the position adjustment unit 33 adjusts the position of the Japanese speech segments 62j and 62k relative to the English speech segment 52f.
[0099] Specifically, the range between the center position of the English silence segment 53a (located before the English speech segment 52f) and the center position of the English silence segment 53b (located after the English speech segment 53f) is set as the position adjustment range within which the positions of the Japanese speech segments 62j and 62k can be adjusted. Therefore, when adjusting the position of the Japanese speech segment 62, overlap with adjacent Japanese speech segments 62 can be avoided.
[0100] It should be noted that the range of position adjustment is not limited to this. For example, the range of position adjustment may be between the beginning position of English silence segment 53a and the end position of English silence segment 53b, or it may be within English speech segment 52f.
[0101] Next, the position adjustment unit 33 shifts the Japanese speech segments 62j and 62k in the time axis direction at predetermined intervals within the position adjustment range, and calculates the intersection-over-union ratio (IoU) between the English volume waveform 54 of the English speech segment 52a and the Japanese speech volume waveform 64 of the Japanese speech segments 62j and 62k at each position.
[0102] Figure 7 This is a diagram illustrating an example of the cross-union ratio (CUNR) calculation method. As an example of the CUNR calculation method, such as... Figure 7 As shown in A, the position adjustment unit 33 calculates the sum of local maximum values by integrating the local maximum values of the English volume waveform 54 and the daily speech volume waveform 64 in the time axis direction. Additionally, as... Figure 7 As shown in B, the position adjustment unit 33 calculates the sum of local minimum values by integrating the local minimum values of the English volume waveform 54 and the daily speech volume waveform 64 in the time axis direction.
[0103] Then, the position adjustment unit 33 calculates the crossover ratio as the value obtained by dividing the sum of local minimum values by the sum of local maximum values.
[0104] The crossover-union ratio (CUNR) indicates the degree of matching between the English volume waveform 54 and the Japanese volume waveform 64. When the English volume waveform 54 and the Japanese volume waveform 64 are perfectly matched, the CUNR becomes 1, and this value decreases as the difference between the English volume waveform 54 and the Japanese volume waveform 64 increases. Therefore, a larger CUNR indicates a higher probability of positional matching.
[0105] The position adjustment unit 33 calculates the intersection-over-union ratio (IoU) of each of the Japanese speech segments 62j and 62k after shifting at predetermined intervals within the position adjustment range. Then, the position adjustment unit 33 determines the position of the Japanese speech segments 62j and 62k with the largest IoU. The position adjustment unit 33 then moves the Japanese speech segments 62j and 62k to the determined position.
[0106] The position adjustment unit 33 adjusts the position of the Japanese speech segment 62 in each of all combinations with corresponding relationships detected by the correspondence detection unit 32.
[0107] Japanese audio data, whose position was adjusted in this way, was added to the video data to generate video content.
[0108] [2.4. Voice Position Adjustment Processing]
[0109] Figure 8 This is a flowchart illustrating the speech position adjustment process. When the speech position adjustment process begins, in step S1, the speech segment extraction unit 31 performs speech segment extraction processing to extract English speech segment 52 and Japanese speech segment 62 from the English speech data and Japanese speech data, respectively. Additionally, in step S1, the speech segment extraction unit 31 also extracts English silence segment 53 and Japanese silence segment 63.
[0110] Next, in step S2, the correspondence detection unit 32 detects the Japanese silence segment 63 that has a correspondence with the English silence segment 53. Then, the correspondence detection unit 32 performs a correspondence detection process, which detects the correspondence between the English speech segment 52 and the Japanese speech segment 62 relative to the corresponding English silence segment 53 and Japanese silence segment 63.
[0111] In step S3, the position adjustment unit 33 performs a position adjustment process, which adjusts the position of the Japanese speech segment 62 in the front-back direction relative to the English speech segment 52.
[0112] <3. Modified Example>
[0113] Note that the embodiments are not limited to the specific examples described above, and various configurations can be used as examples of modifications.
[0114] For example, although the embodiments illustrate the case where the first language before dubbing is English and the second language after dubbing is Japanese, the first and second languages are not limited to this and can be other languages.
[0115] [3.1. First Modification Example]
[0116] Figure 9 and Figure 10 This diagram illustrates the correspondence detection process in the first modified example. In the correspondence detection process of the above embodiment, the correspondence between the English speech segment 52 and the Japanese speech segment 62, which are sandwiched between the English silence segment 53 and the Japanese silence segment 63, is detected relative to the English silence segment 53 and the Japanese silence segment 63, which have a correspondence. However, the method for detecting the correspondence between the English speech segment 52 and the Japanese speech segment 62 is not limited to this.
[0117] For example, such as Figure 9 As shown, for the extracted English speech segment 52, the correspondence detection unit 32 uses hierarchical clustering technology to generate clusters 71 hierarchically, so as to group the closest English speech segments 52 or clusters into one cluster in sequence. Assuming in Figure 9 In the example, three clusters, 71a, 71b, and 71c, are generated from English speech data.
[0118] Similarly, for the extracted Japanese speech segments 62, the correspondence detection unit 32 uses hierarchical clustering technology to generate clusters 72 hierarchically, so as to group the closest Japanese speech segments 62 or clusters into one cluster in sequence. Assuming in Figure 9 In the example, five clusters, 72a, 72b, 72c, 72d, and 72e, were generated from Japanese speech data.
[0119] After generating clusters in this way, such as Figure 10 As shown, the correspondence detection unit 32 calculates the intersection-union ratio (IoU) between the cluster 71 of English speech data and the cluster 72 of Japanese speech data.
[0120] Then, the correspondence detection unit 32 detects the cluster with the highest crossover-union ratio (in Figure 9 The example shown illustrates the correspondence between English speech segment 52 and Japanese speech segment 62 included in clusters 71b and 72d.
[0121] By doing so, the correspondence detection unit 32 can accurately detect correspondences by sequentially clustering speech segments 42 with different durations.
[0122] [3.2. Second Modification Example]
[0123] Figure 11 This is a diagram used to illustrate the correspondence detection process in the second modified example. As another example of a method for detecting the correspondence between English speech segment 52 and Japanese speech segment 62, such as... Figure 11 As shown, the correspondence detection unit 32 calculates the cross-union ratio between each English speech segment 52 and each of all Japanese speech segments 62.
[0124] Then, the correspondence detection unit 32 detects the Japanese speech segment 62 that has the highest intersection-union ratio with the English speech segment 52 as a Japanese speech segment 62 that has a correspondence with each English speech segment 52.
[0125] [3.3. Third Modification Example]
[0126] Figure 12 This is a flowchart illustrating the speech position adjustment process in the third modified example. In the above embodiment, the speech segment extraction unit 31 directly extracts the English speech segment 52 from the English speech data. However, in some cases, the English speech data includes background noise in addition to the actor's speech (speaking part). Therefore, in the speech position adjustment process in the third modified example, as... Figure 12 As shown, in step S11, the speech segment extraction unit 31 performs sound source separation processing to extract the human speech portion from the English speech data. Note that since the Japanese speech data is collected in a studio or similar environment, it essentially does not include background noise. Therefore, the speech segment extraction unit 31 does not need to perform speech separation processing on the Japanese speech data.
[0127] Next, in step S1, the speech segment extraction unit 31 extracts English speech segment 52 from the English speech data, from which background noise has been removed through sound source separation processing, and extracts Japanese speech segment 62 from the Japanese speech data. Thereafter, similar to the embodiment described above, the CPU 11 performs correspondence detection processing in step S2 and position adjustment processing in step S3.
[0128] [3.4. Fourth Modification Example]
[0129] Figure 13 This is a flowchart illustrating the speech position adjustment process in the fourth modified example. In the above embodiment, the position adjustment unit 33 adjusts the position of the Japanese speech segment 62 relative to the English speech segment 52. However, if the actor's mouth movements can be detected in the video, the position of the Japanese speech segment 62 can be adjusted to match the actor's mouth movements.
[0130] In the voice position adjustment process in the fourth modified example, such as Figure 13As shown, in step S21, the CPU 11, which serves as the motion detection unit, detects the actor's mouth movements from the video data. Next, similar to the embodiment, the CPU 11 performs speech segment extraction processing in step S1 and correspondence detection processing in step S2. Then, in step S22, the position adjustment unit 33 adjusts the position of the Japanese speech segment 62 relative to the actor's mouth movements or the English speech segment 52.
[0131] For example, when an actor's mouth movement is detected, the position adjustment unit 33 adjusts the position of the Japanese speech segment 62 (one or more Japanese speech segments 62) corresponding to the mouth movement, so that, for example, the timing of the start of the mouth movement matches the start of the Japanese speech segment 62. Furthermore, the position adjustment unit 33 can adjust the position of the Japanese speech segment 62 corresponding to the mouth movement, so that, for example, the timing of the end of the mouth movement matches the end of the Japanese speech segment 62.
[0132] Note that, similar to the embodiment, the position of the Japanese speech segment 62, excluding the Japanese speech segment 62 corresponding to mouth movements, is adjusted relative to the English speech segment 52 which corresponds to the Japanese speech segment 62.
[0133] This reduces the misalignment between the video and the dubbed audio.
[0134] [3.5. Fifth Revision Example]
[0135] Figure 14 This is a flowchart illustrating the speech position adjustment process in the fifth modified example. In the above embodiment, the position adjustment unit 33 adjusts the position of the Japanese speech segment 62 relative to the English speech segment 52. However, if the speech of each actor can be detected in the video, the position of the Japanese speech segment 62 can be adjusted to match the speech of each actor.
[0136] In the voice position adjustment process in the fifth modified example, such as Figure 14 As shown, in step S31, the CPU 11, which serves as the speaker diarization processing unit, performs speaker diarization processing, which uses speaker diarization technology to detect the speech and timing of each actor (speaker) from the English speech data. Next, similar to the embodiment, the CPU 11 performs speech segment extraction processing in step S1 and correspondence detection processing in step S2. Then, in step S32, the position adjustment unit 33 adjusts the position of the Japanese speech segment 62 relative to the speech and timing of each actor detected in the speaker diarization processing.
[0137] Therefore, since the position of the Japanese voice segment 62 can be adjusted to match each actor's speech and its timing, the misalignment between the video and the dubbed audio can be reduced.
[0138] <4. Summary of Examples>
[0139] As described above, the information processing apparatus 1, as an embodiment, includes: a speech segment extraction unit 31, which extracts speech segments 42 (English speech segment 52 and Japanese speech segment 62) from each of the first language speech data (English speech data) added to the video data and the dubbed second language speech data (Japanese speech data); a correspondence detection unit 32, which detects the correspondence between the first language speech segments and the second language speech segments; and a position adjustment unit 33, which adjusts the position of the second language speech segments whose correspondence with the first language speech segments has been detected relative to the first language speech segments.
[0140] After detecting the correspondence between English speech segment 52 and Japanese speech segment 62, the information processing device 1 adjusts the position of the Japanese speech segment 62, which corresponds to the English speech segment 52. Thus, the information processing device 1 can more accurately align the position of the Japanese speech segment 62 with the English speech segment 52.
[0141] Therefore, the information processing device 1 can reduce the misalignment between the dubbed audio and video.
[0142] The correspondence detection unit 32 can detect the correspondence between the speech segments of the first language (English speech segment 52) and the speech segments of the second language (Japanese speech segment 62) as an N-to-M relationship (N is an integer equal to or greater than 1, and M is an integer equal to or greater than 1).
[0143] Due to differences in the characteristics of different languages, in some cases the number of speech segments 42 differs before and after dubbing.
[0144] Therefore, the correspondence detection unit 32 can reduce the false detection of correspondences caused by the differences in characteristics between different languages by detecting the N-to-M correspondence between the English speech segment 52 and the Japanese speech segment 62.
[0145] The correspondence detection unit 32 detects the correspondence between the first language speech segment (English speech segment 52) and the second language speech segment (Japanese speech segment 62) relative to the silence segment 43.
[0146] Since neither English nor Japanese is spoken in the segments without dialogue, they become silent segments 43 at almost the same timing and for almost the same duration. Therefore, by detecting the correspondence between English speech segment 52 and Japanese speech segment 62 relative to silent segment 43 (English silent segment 53, Japanese silent segment 63), false detections of correspondence can be reduced.
[0147] The correspondence detection unit 32 detects the silence segments of the first language (English silence segment 53) and the silence segments of the second language (Japanese silence segment 63) that have a correspondence relationship based on the cross-union ratio between the silence segments of the first language and the silence segments of the second language.
[0148] As described above, since silence segment 43 occurs at the same timing regardless of language, the cross-union ratio (CUNR) between English silence segment 53 and Japanese silence segment 63 is calculated, and combinations with high CUNR are detected as corresponding English silence segment 53 and Japanese silence segment 63.
[0149] Therefore, the correspondence between English silence segment 53 and Japanese silence segment 63 can be accurately detected.
[0150] The correspondence detection unit 32 detects the first language speech segment (English speech segment 52) located between adjacent silence segments of the first language (English silence segment 53) and the second language speech segment (Japanese speech segment 62) located between silence segments of the second language (Japanese silence segment 63) that have a correspondence with the adjacent silence segments of the first language as a combination with a correspondence.
[0151] As described above, the correspondence between English silence segment 53 and Japanese silence segment 63 is accurately detected. Therefore, the correspondence detection component 32 can accurately detect the correspondence between English speech segment 52 sandwiched between English silence segments 53 and Japanese speech segment 62 sandwiched between Japanese silence segments 63 that have a correspondence with English silence segments 53.
[0152] Based on the volume waveforms 44 (English volume waveform 54, Japanese volume waveform 64) of the first and second language speech segments 42 (English speech segment 52, Japanese speech segment 62) with detected corresponding relationships, the position adjustment unit 33 adjusts the position of the second language speech segment (Japanese speech segment 62).
[0153] In different languages, the speech waveforms 41 can vary significantly in some cases due to differences in pitch, subject and predicate positions, etc. On the other hand, even if the language differs before and after dubbing, the volume is often kept the same in order to match the emotions, etc.
[0154] Therefore, the position adjustment unit 33 can absorb the differences between the speech waveforms 41 of different languages by adjusting the position of the Japanese speech segment 62 based on the volume waveform 44, and accurately adjust the position.
[0155] The position adjustment unit 33 adjusts the position based on the intersection-over-union ratio between the volume waveforms 44 of the first and second language speech segments (English volume waveform 54, Japanese speech waveform 64).
[0156] Therefore, the position adjustment unit 33 can align the position of the Japanese speech with the position of the English volume waveform 54 and the Japanese speech volume waveform 64 that are most similar to each other.
[0157] The position adjustment unit 33 moves the volume waveform of the second language speech segment (daily speech volume waveform 64) relative to the volume waveform of the first language speech segment (English volume waveform 54) at predetermined intervals, and moves the second language speech segment to a position where the crossover ratio between the volume waveforms of the first and second language speech segments is maximized.
[0158] Therefore, the position adjustment unit 33 can move the Japanese voice to what is considered the optimal position.
[0159] The position adjustment unit 33 moves the volume waveform (daily speech volume waveform 64) of the second language speech segment between positions of a silence segment (English silence segment 53) adjacent to the speech segment of the first language that corresponds to the speech segment of the second language.
[0160] Therefore, when the position of the Japanese speech is moved, the position adjustment unit 33 can reduce the unwanted overlap between Japanese speech sounds.
[0161] The correspondence detection unit 32 performs hierarchical clustering of the speech segments of the first language and the second language, and detects the correspondence between the speech segments of the first language and the speech segments of the second language based on the intersection-union ratio between the clustered speech segments.
[0162] Therefore, the correspondence detection unit 32 can accurately detect correspondences by sequentially clustering speech segments with different durations.
[0163] The speech segment extraction unit 31 performs speech separation processing to extract the human language portion of the first language speech data (English speech data).
[0164] Therefore, even if the speech data includes background noise, the English speech segment 52, which is the part of the speech, can be accurately extracted.
[0165] Furthermore, an information processing method includes: extracting speech segments from each of speech data in a first language added to video data and speech data in a second language after dubbing; detecting a correspondence between the speech segments in the first language and the speech segments in the second language; and adjusting the position of the speech segments in the second language that have been detected to correspond to the speech segments in the first language relative to the speech segments in the first language.
[0166] Furthermore, a program causes the information processing device 1 to perform: speech segment extraction processing, extracting speech segments from each of the speech data of a first language added to the video data and the speech data of a dubbed second language; correspondence detection processing, detecting the correspondence between the speech segments of the first language and the speech segments of the second language; and position adjustment processing, adjusting the position of the speech segments of the second language that have been detected to correspond to the speech segments of the first language relative to the speech segments of the first language.
[0167] This program is a program that implements the functions of the aforementioned information processing device in a computer.
[0168] Furthermore, the recording medium according to the invention, on which such a program according to the invention is recorded, facilitates the provision of the program and makes the invention widely available to the public.
[0169] The program can be pre-recorded in HDDs, ROMs in CPUs, or other recording media built into devices such as computers.
[0170] Alternatively, the program can be temporarily or permanently stored (recorded) on removable recording media such as floppy disks, CD-ROMs (Compact Disc Read Only Memory), MO (Magneto-optical) discs, DVDs (Digital Versatile Discs), magnetic disks, or semiconductor memory. Such removable recording media can be provided as so-called software packages.
[0171] In addition to being installed on a computer or other removable recording medium, the program can also be downloaded from a download site via a network such as a LAN (Local Area Network) or the Internet.
[0172] Note that the beneficial effects described in this specification are merely exemplary and not limiting, and other beneficial effects may exist.
[0173] <5. This technology>
[0174] This technology can also be configured as follows.
[0175] (1) An information processing device, comprising:
[0176] The speech segment extraction unit extracts speech segments from each of the speech data in the first language added to the video data and the speech data in the second language after dubbing.
[0177] A correspondence detection unit detects the correspondence between speech segments of a first language and speech segments of a second language; and
[0178] The position adjustment unit adjusts the position of the second language speech segment, which has been detected to correspond to the speech segment of the first language, relative to the speech segment of the first language.
[0179] (2) The information processing apparatus according to (1), wherein,
[0180] The correspondence detection unit can detect the correspondence between the speech segments of the first language and the speech segments of the second language as N pairs of M relationships, where N is an integer equal to or greater than 1, and M is an integer equal to or greater than 1.
[0181] (3) The information processing apparatus according to (1) or (2), wherein,
[0182] The correspondence detection unit detects the correspondence between speech segments of the first language and speech segments of the second language relative to the silence segment.
[0183] (4) The information processing apparatus according to (3), wherein,
[0184] The correspondence detection unit detects the correspondence between the silence segments of the first language and the silence segments of the second language based on the cross-union ratio between the silence segments of the first language and the silence segments of the second language.
[0185] (5) The information processing apparatus according to (3) or (4), wherein,
[0186] The correspondence detection unit detects the speech segments of the first language located between adjacent silence segments of the first language and the speech segments of the second language located between silence segments of the second language that have corresponding relationships with adjacent silence segments of the first language as a combination with a corresponding relationship.
[0187] (6) The information processing apparatus according to any one of (1) to (5), wherein,
[0188] Based on the volume waveforms of the first language speech segments and the second language speech segments that have been detected to have a corresponding relationship, the position adjustment unit adjusts the position of the second language speech segment.
[0189] (7) The information processing apparatus according to (6), wherein,
[0190] Based on the cross-combination ratio between the volume waveforms of the first language speech segment and the second language speech segment, the position adjustment unit adjusts the position of the second language speech segment.
[0191] (8) The information processing apparatus according to (7), wherein,
[0192] The position adjustment unit moves the volume waveform of the second language speech segment relative to the volume waveform of the first language speech segment at predetermined intervals, and moves the second language speech segment to a position where the cross-interference ratio between the volume waveforms of the first language speech segment and the second language speech segment is maximized.
[0193] (9) The information processing apparatus according to (8), wherein,
[0194] The position adjustment unit moves the volume waveform of the second language speech segment within a range of silent segments adjacent to the speech segment of the first language that corresponds to the speech segment of the second language.
[0195] (10) The information processing apparatus according to any one of (1) to (9), wherein,
[0196] The correspondence detection unit performs hierarchical clustering of speech segments of the first language and speech segments of the second language, and detects the correspondence between speech segments of the first language and speech segments of the second language based on the intersection-union ratio between the clustered speech segments.
[0197] (11) The information processing apparatus according to any one of (1) to (10), wherein,
[0198] The speech segment extraction unit performs speech separation processing on the speech data of the first language to extract the human speech portion.
[0199] (12) The information processing apparatus according to (5) includes:
[0200] The motion detection unit detects mouth movement from video data, wherein...
[0201] The position adjustment unit adjusts the position of the speech segment of the second language to match the mouth movement.
[0202] (13) An information processing method, comprising:
[0203] Extract speech segments from each of the first language speech data added to the video data and the second language speech data after dubbing;
[0204] Detect the correspondence between speech segments of the first language and speech segments of the second language; and
[0205] The position of the second language speech segment, which has been detected to correspond to the first language speech segment, is adjusted relative to the first language speech segment.
[0206] (14) A program that causes an information processing device to perform:
[0207] Speech segment extraction processing: Extract speech segments from each of the first language speech data added to the video data and the dubbed second language speech data;
[0208] Correspondence detection processing detects the correspondence between speech segments of the first language and speech segments of the second language; and
[0209] Position adjustment processing adjusts the position of the second language speech segment, which has been detected to correspond to the first language speech segment, relative to the first language speech segment.
[0210] Reference tag list
[0211] 1. Information processing device
[0212] 11 CPU
[0213] 31. Speech Segment Extraction Unit
[0214] 32 Correspondence Detection Department
[0215] 33 Position Adjustment Section
Claims
1. An information processing apparatus, comprising: The speech segment extraction unit extracts speech segments from each of the speech data in the first language added to the video data and the speech data in the second language after dubbing. A correspondence detection unit detects the correspondence between speech segments of the first language and speech segments of the second language. as well as The position adjustment unit adjusts the position of the second language speech segment, which has been detected to correspond to the speech segment of the first language, relative to the speech segment of the first language.
2. The information processing apparatus according to claim 1, wherein, The correspondence detection unit can detect the correspondence between the speech segments of the first language and the speech segments of the second language as N pairs of M relationships, where N is an integer equal to or greater than 1, and M is an integer equal to or greater than 1.
3. The information processing apparatus according to claim 1, wherein, The correspondence detection unit detects the correspondence between speech segments of the first language and speech segments of the second language relative to the silence segment.
4. The information processing apparatus according to claim 3, wherein, The correspondence detection unit detects the correspondence between the silence segments of the first language and the silence segments of the second language based on the cross-union ratio between the silence segments of the first language and the silence segments of the second language.
5. The information processing apparatus according to claim 3, wherein, The correspondence detection unit detects the speech segments of the first language located between adjacent silence segments of the first language and the speech segments of the second language located between silence segments of the second language that have corresponding relationships with adjacent silence segments of the first language as a combination with a corresponding relationship.
6. The information processing apparatus according to claim 1, wherein, Based on the volume waveforms of the first language speech segments and the second language speech segments that have been detected to have a corresponding relationship, the position adjustment unit adjusts the position of the second language speech segment.
7. The information processing apparatus according to claim 6, wherein, Based on the cross-combination ratio between the volume waveforms of the first language speech segment and the second language speech segment, the position adjustment unit adjusts the position of the second language speech segment.
8. The information processing apparatus according to claim 7, wherein, The position adjustment unit moves the volume waveform of the second language speech segment relative to the volume waveform of the first language speech segment at predetermined intervals, and moves the second language speech segment to a position where the cross-interference ratio between the volume waveforms of the first language speech segment and the second language speech segment is maximized.
9. The information processing apparatus according to claim 8, wherein, The position adjustment unit moves the volume waveform of the second language speech segment within a range of silent segments adjacent to the speech segment of the first language that corresponds to the speech segment of the second language.
10. The information processing apparatus according to claim 1, wherein, The correspondence detection unit performs hierarchical clustering of speech segments of the first language and speech segments of the second language, and detects the correspondence between speech segments of the first language and speech segments of the second language based on the intersection-union ratio between the clustered speech segments.
11. The information processing apparatus according to claim 1, wherein, The speech segment extraction unit performs speech separation processing on the speech data of the first language to extract the human speech portion.
12. The information processing apparatus according to claim 1, comprising: The motion detection unit detects mouth movement from video data, wherein... The position adjustment unit adjusts the position of the speech segment of the second language to match the mouth movement.
13. An information processing method, comprising: Extract speech segments from each of the first language speech data added to the video data and the second language speech data after dubbing; Detect the correspondence between speech segments of the first language and speech segments of the second language; as well as The position of the second language speech segment, which has been detected to correspond to the first language speech segment, is adjusted relative to the first language speech segment.
14. A program that causes an information processing device to perform: Speech segment extraction processing: Extract speech segments from each of the first language speech data added to the video data and the dubbed second language speech data; Correspondence detection processing detects the correspondence between speech segments of the first language and speech segments of the second language; as well as Position adjustment processing adjusts the position of the second language speech segment, which has been detected to correspond to the first language speech segment, relative to the first language speech segment.
Citation Information
Patent Citations
Dubbing system and video image display system
JP1996006182A