Sound processing device, sound processing system, and program
The sound processing device optimally adjusts secondary audio dialogue gain by measuring objective indices at varying times and calculating weighted gains based on speech timing similarity, addressing timing discrepancies and ensuring stable, responsive audio level adjustments.
Patent Information
- Application Number
- JP2021132442
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2021-08-16
- Publication Date
- 2025-08-27
- Estimated Expiration
- 2041-08-16
AI Technical Summary
Existing automated audio mixing technologies struggle to optimally adjust the gain of secondary audio dialogue when the speech timing between main and secondary audio differs, leading to unstable or delayed level adjustments, particularly in real-time applications like live broadcasts.
A sound processing device that automatically adjusts the gain of secondary audio dialogue by measuring objective indices at different effective measurement times, determining speech timing similarity, and calculating weighted gains to align the target audio with a reference audio, using a combination of momentary and average loudness values to ensure stability and responsiveness.
The device enables optimal gain adjustment of secondary audio dialogue even when speech timings differ, reducing delays and ensuring stable, responsive audio level adjustments in real-time scenarios, such as live broadcasts.
Smart Images

Figure 0007730276000002 
Figure 0007730276000003 
Figure 0007730276000004
Abstract
Description
[Technical Field]
[0001] The present invention relates to a sound processing device, a sound processing system, and a program. [Background technology]
[0002] In recent years, with the globalization of society and the spread of the Internet, viewing content produced overseas has become commonplace, and there has been an increasing demand for audio-visual content such as television programs and movies, including multilingual support and audio commentary to support people with disabilities. To meet this demand, there is a movement in Europe, the United States, and Asian countries to use object-based audio services in next-generation broadcasting, which can easily and efficiently realize various audio descriptions. In particular, the United States and South Korea, which have adopted ATSC (Advanced Television Systems Committee) 3.0, have already begun introducing AC-4 or MPEG-H 3DA audio encoding standards into their broadcasting systems.
[0003] As the need for various audio content patterns increases, technologies have been announced that automate and support some of the mixing work performed by engineers. Currently, automated mixing technologies can be broadly divided into those that perform offline processing and those that perform online processing. While numerous products for the former have been released thanks to the recent rise of machine learning technology, the latter have not yet been fully adopted in the market, as they are often used for professional equipment such as audio consoles. As a result, the number of products is still limited and the situations in which they can be applied are limited. The mainstream online processing technology, which is designed for real-time mixing of multi-person conversations, involves crossfading the gain of faders assigned to the speakers while turning them on and off (see, for example, Non-Patent Document 1).
[0004] Furthermore, to prevent viewers from being disturbed by extreme changes in loudness between channels or programs, the International Telecommunication Union-Radiocommunication Sector (ITU-R) has established a method for measuring loudness values, an objective index that takes into account human auditory characteristics, in order to estimate the loudness of program audio and keep it within a certain range (see, for example, Non-Patent Documents 2 and 3). In Japan, ARIB has also formulated loudness operational regulations based on the ITU-R recommendations, and program audio is actually managed based on loudness values in broadcasting sites (see, for example, Non-Patent Document 4). There are three types of standardized loudness values, as shown in Table 1 below, but the basic measurement method is the same. The value for each audio block, which is the smallest unit of measurement, is defined as the momentary loudness value, and the short-term loudness value and average loudness value are basically calculated as the average of the measured values for each block over the respective time widths, although there are differences in gating processing, etc.
[0005] [Table 1]
[0006] When creating secondary audio for television programs and other media, assigning individual audio engineers to each audio segment increases production costs, so the number of secondary audio segments is limited. Object-based audio is being considered for personal adaptation services tailored to viewer preferences, but production costs necessitate automating the production of secondary audio. Automatically adjusting the target dialogue (dialogue in secondary audio) requires objective indicators for the audio levels of both the reference dialogue (adjusted dialogue in the main audio) and the target dialogue. Existing objective indicators include loudness values described in Non-Patent Documents 2 to 4. However, if the three loudness values listed in Table 1 above are used directly for automatic level adjustment, loudness values with short effective measurement times, such as momentary loudness values, tend to change rapidly, making stable level adjustment difficult. Conversely, loudness values with long effective measurement times, such as average loudness values, are calculated from the average value of the values for multiple audio blocks included in that effective measurement time, resulting in a delay in reflecting changes in audio level in the loudness value. Depending on the effective measurement time, the delay can be measured in seconds, making it unsuitable as an indicator in situations where real-time adjustments are required, such as live broadcasts.
[0007] Furthermore, while a medium-length effective measurement time, such as short-term loudness values, is relatively excellent in terms of the stability and delay of the calculated values, it is not clear whether the time width (3 seconds) is an appropriate length as an objective indicator for automatic adjustment. Furthermore, when the start and end of speech are involved, i.e., when silent periods are included within the effective measurement time, the difference in level between the reference dialogue and the target dialogue can suddenly become large, and inappropriate adjustments such as excessive or insufficient adjustment values may occur.
[0008] Therefore, Non-Patent Document 5 discloses a technology for automatically adjusting the gain of audio objects such as dialogue using an objective index that achieves both stability and responsiveness in the production of secondary audio for audiovisual content such as television programs. [Prior art documents] [Non-patent literature]
[0009] [Non-Patent Document 1] https: / / www.solid-state-logic.co.jp / broadcastsound / dialogue-automix / [Non-patent document 2] ITU-R, Rec. ITU-R BS.1770-4, “Algorithms to measure audio program loudness and true-peak audio level”, 2015 [Non-patent document 3] ITU-R, Rec. ITU-R BS.1771-1, “Requirements for loudness and true-peak indicating meters”, 2012 [Non-patent document 4] ARIB, ARIB TR-B32, "Loudness Operational Guidelines for Digital Television Broadcasting Programs," 2016 [Non-patent document 5] Kubo and Oide, "Study on Objective Indicators for Automatic Dialogue Level Adjustment in Audio Description Production," IEICE Technical Report, February 24, 2020, Vol. 119, No. 441 Summary of the Invention [Problem to be solved by the invention]
[0010] Using the technology disclosed in Non-Patent Document 5, it is possible to automatically adjust the audio level of the secondary audio based on the audio level of the main audio. However, even if the adjustment is optimal when the dialogue between the main audio and the secondary audio is similar and spoken at almost the same timing (for example, dubbing between Japanese and English), when the content and timing of speech between the main audio and the secondary audio are completely different (for example, when the secondary audio has almost no relation to the progress of the program and is telling jokes by a comedian), adjustment made with reference to the main audio may not be optimal.
[0011] The object of the present invention, made in consideration of the above circumstances, is to provide an audio processing device and program that automatically and optimally adjusts the gain of secondary audio dialogue when producing secondary audio for audio content such as television programs, even when the timing of speech between the secondary audio dialogue and the main audio dialogue differs. [Means for solving the problem]
[0012] In order to solve the above problem, a sound processing device according to the present invention is a sound processing device that automatically adjusts a target sound object based on a reference sound object that is a sound object of a reference sound that serves as a reference, and a target sound object that is a sound object of a target sound that is to be adjusted, and includes an objective index measurement unit that measures a reference objective index that is an objective index based on a loudness value of the reference sound object, and a target objective index that is an objective index based on a loudness value of the target sound object, an utterance timing determination unit that determines a similarity of utterance timing for the reference sound object and the target sound object, a gain calculation unit that determines a gain of the target sound object based on the similarity so as to bring the target objective index closer to the reference objective index, and a level adjustment unit that adjusts the sound level of the target sound object based on the gain. The objective index measurement unit measures the reference objective index and the target objective index for different effective measurement times, and the gain calculation unit calculates the gain by calculating a difference between the reference objective index and the target objective index for each effective measurement time and performing a weighted addition of the difference, and the higher the similarity, the greater the weight given to the difference when the effective measurement time is short. do.
[0014] To solve the above problems , a sound processing device according to the present invention teeth , An audio processing device that automatically adjusts a target audio object based on a reference audio object, which is an audio object of a reference audio that serves as a reference, and a target audio object, which is an audio object of a target audio that is to be adjusted, comprising: an objective index measurement unit that measures a reference objective index that is an objective index based on a loudness value of the reference audio object, and a target objective index that is an objective index based on a loudness value of the target audio object; an utterance timing determination unit that determines a similarity of utterance timing for the reference audio object and the target audio object; a gain calculation unit that determines a gain of the target audio object based on the similarity so as to bring the target objective index closer to the reference objective index; and a level adjustment unit that adjusts the audio level of the target audio object based on the gain. The objective index measurement unit measures the reference objective index at different effective measurement times, and measures the target objective index at the shortest measurement time among the effective measurement times. The gain calculation unit calculates a difference between a value obtained by weighting and adding the reference objective index and the target objective index. The gain is calculated as The higher the similarity, the greater the weighting of the reference objective index when the effective measurement time is short. do .
[0015] Furthermore, in the sound processing device according to the present invention, the effective measurement time may be set to 60 to 70 seconds, 20 to 30 seconds, and 1.6 to 2 seconds.
[0016] Furthermore, in the sound processing device according to the present invention, the speech timing determination unit may calculate a correlation coefficient between the reference objective index and the target objective index as the degree of similarity.
[0017] Furthermore, in the sound processing device according to the present invention, the objective index measurement unit may include a momentary loudness value measurement unit that measures a momentary loudness value, a skip gate unit that regards the time during which the momentary loudness value exceeds a skip gate threshold as speech time and outputs a momentary loudness value of a fixed-length effective measurement time within the speech time immediately prior to the current time, and a time percentage upper average value calculation unit that calculates, as the objective index, an average value of momentary loudness values of a predetermined time percentage within the effective measurement time.
[0018] In order to solve the above problem, the present invention provides a program that causes a computer to function as the sound processing device. [Effects of the Invention]
[0019] According to the present invention, even when the speech timings of the main audio and the secondary audio are different, it is possible to automatically adjust the gain of the secondary audio dialogue optimally. [Brief explanation of the drawings]
[0020] [Figure 1] 1 is a block diagram illustrating an example of the configuration of a sound processing device according to an embodiment. [Figure 2] 10 is a flowchart illustrating an example of a processing procedure of a sound processing device according to an embodiment. [Figure 3] 2 is a block diagram showing an example of the configuration of an objective index measurement unit in the sound processing device according to the first embodiment. FIG. [Figure 4] 10A and 10B are diagrams illustrating processing by a skip gate unit in the sound processing device according to an embodiment. [Figure 5]10A and 10B are diagrams illustrating processing by a time rate upper average value calculation unit in the sound processing device according to an embodiment. [Figure 6] 10 is a flowchart illustrating an example of a processing procedure of an objective index measurement unit in the sound processing device according to an embodiment. [Figure 7] 1 is a block diagram showing an example of the configuration of a sound processing device according to a first embodiment. [Figure 8] FIG. 1 is a diagram for explaining a role of an objective index measurement unit in the sound processing device according to the first embodiment. [Figure 9] 1 is a table showing examples of the speech content of a reference voice object and a target voice object; [Figure 10] 10 is a graph showing examples of objective indices measured over various effective measurement times for a dialogue of expert commentary, one of the target audio objects shown in FIG. 9. [Figure 11] 10 is a graph showing an example of the variance of objective indices measured while changing the effective measurement time. [Figure 12] 10 is a graph showing an example of a correlation coefficient between objective indices of a reference audio object and a target audio object. [Figure 13] 10 is a graph showing the relationship between the similarity of utterance timing and weighting. [Figure 14] 10 is a flowchart illustrating an example of a processing procedure of a gain calculation unit in the sound processing device according to the first embodiment. [Figure 15] FIG. 10 is a block diagram showing an example of the configuration of a sound processing device according to a second embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0021] Hereinafter, embodiments of the present invention will be described in detail with reference to the drawings.
[0022] 1 is a block diagram showing an example of the configuration of a sound processing device 1 according to an embodiment. The sound processing device 1 shown in FIG. 1 includes an objective index measurement unit group 11s, an utterance timing determination unit 12, a gain calculation unit 13, and a level adjustment unit 14.
[0023] In the automatic production of secondary audio for audio content such as television programs, the sound processing device 1 inputs the main audio dialogue (hereinafter referred to as the "reference audio object") that serves as a comparison reference (standard) for adjusting the volume of the sound and the secondary audio dialogue (hereinafter referred to as the "target audio object") that is to be adjusted, calculates the gain of the target audio object based on multiple objective indicators measured at different effective measurement times and the similarity in the speech timing of the reference audio object and the target audio object, and automatically adjusts the gain of the target audio object.
[0024] The objective index measurement unit group 11s calculates a plurality of objective indexes measured at different effective measurement times. The objective index measurement unit group 11s is made up of a plurality of objective index measurement units 11 (11-1 to 11-N). Each objective index measurement unit 11 divides a reference audio object and a target audio object input in real time from outside the sound processing device 1 into audio blocks of a certain length, and measures a reference objective index, which is an objective index based on the loudness value of the reference audio object, and a target objective index, which is an objective index of the target audio object, for each audio block. Then, the objective index measurement unit 11 outputs the measured reference objective index and target objective index to the gain calculation unit 13.
[0025] The speech timing determination unit 12 determines the similarity of speech timing between the reference audio object and the target audio object. Any known method can be used to determine the similarity of speech timing. For example, the speech timing determination unit 12 may determine the similarity based on whether the reference objective index and the target objective index in a certain-length audio block used for measurement both exceed a predetermined threshold, or the proportion exceeding the threshold may be used as the similarity. Furthermore, the speech timing determination unit 12 may determine the similarity based on the amplitude and zero-crossing count used in audio segment detection for speech recognition, or the degree of overlap of audio segments detected using a Gaussian mixture distribution model, or the correlation coefficient between short-term or long-term loudness values or objective indexes of the reference audio object and the target audio object. When the speech timing determination unit 12 determines the similarity using objective indexes, the speech timing determination unit 12 inputs the reference objective index and the target objective index from the objective index measurement unit group 11s.
[0026] The gain calculation unit 13 calculates a gain required to adjust the target voice object so that the target objective index approaches the reference objective index, based on the objective index input from the objective index measurement unit group 11s and the similarity input from the speech timing determination unit 12. Then, the gain calculation unit 13 outputs the calculated gain to the level adjustment unit 14.
[0027] The level adjustment unit 14 adjusts the audio level by applying the gain input from the gain calculation unit 13 to a target audio object input from outside the sound processing device 1, and outputs the adjusted target audio object to the outside as an adjusted audio object. Here, by applying the gain to the target audio object that is the block next to the measured audio block, the theoretical delay can be reduced to pure calculation time only, so that even in real-time adjustments such as live broadcasting, it is possible to significantly reduce the impact of audio signal delays when using the device of the present application.
[0028] FIG. 2 is a flowchart showing an example procedure for adjusting the audio level of a target audio object in the sound processing device 1. In step S11, the objective index measurement unit group 11s measures objective indexes of the reference audio object and the target audio object. In step S12, the speech timing determination unit 12 determines the similarity of the speech timing between the reference audio object and the target audio object. In step S13, the gain calculation unit 13 calculates a gain from the objective index. In step S14, the level adjustment unit 14 reflects the gain calculated for the previous block in the audio block of the target audio object. In step S15, the sound processing device 1 checks whether or not an adjustment end instruction has been received, and repeats the processes from step S11 to step S14 for each audio block until an adjustment end instruction is received.
[0029] Fig. 3 is a block diagram showing an example configuration of an objective index measurement unit 11 that measures an objective index over a certain effective measurement time. The objective index measurement unit 11 shown in Fig. 3 includes a momentary loudness value measurement unit 110, a skip gate unit 116, and a time rate upper average calculation unit 117. The momentary loudness value measurement unit 110 includes a pre-filter 111, a root mean square unit 112, a weighting unit 113, a summation unit (Σ) 114, and a decibel scale conversion unit (Log) 115. The objective index measurement units 11 have the same configuration, and only the effective measurement time input to the skip gate unit 116 differs.
[0030] In this embodiment, the momentary loudness value measurement unit 110 measures the momentary loudness value using a standardized algorithm. For details, see Non-Patent Documents 2 to 4, which are specifications established by standardization organizations.
[0031] The prefilter 111 applies a two-stage prefilter, such as a K-characteristic filter, to each channel of the input reference audio object or target audio object for each audio block, performs prefiltering, and outputs the result to the square mean unit 112.
[0032] The mean square unit 112 performs mean square processing on the signal input from the pre-filter 111 and outputs the result to the weighting unit 113 .
[0033] The weighting unit 113 multiplies the signal input from the root mean square unit 112 by a weighting coefficient according to the direction of the audio signal for each channel, and outputs the result to the summing unit 114 .
[0034] The summing unit 114 sums the weighted root mean square values of the channels excluding the LFE, and outputs the sum to the decibel scale conversion unit 115 .
[0035] The decibel scale conversion unit 115 converts the signal input from the summing unit 114 into a decibel scale, finds a loudness value (momentary loudness value) for each audio block, and outputs it to the skip gate unit .
[0036] The skip gate unit 116 receives the momentary loudness value for each sound block from the momentary loudness value measurement unit 110 and regards the time during which the momentary loudness value exceeds the skip gate threshold as speech time. On the other hand, the skip gate unit 116 regards the time during which the momentary loudness value is equal to or less than the skip gate threshold as non-speech time and skips it. Then, the skip gate unit 116 outputs the momentary loudness value of the effective measurement time within the speech time immediately before the current time (the relevant time) to the time rate upper average value calculation unit 117. Here, the effective measurement time may be stored in advance in the sound processing device 1 or may be input and set from outside.
[0037] 4A and 4B are diagrams illustrating the processing of the skip gate unit 116. The horizontal axis of the graph shown in FIG. 4 represents time, and the vertical axis represents momentary loudness value. As shown in FIG. 4A, the skip gate unit 116 processes the momentary loudness value t immediately before the current time (the relevant time) t. pv The speaking time in seconds is the effective measurement time t pv The effective measurement time t pv is different for each objective index measurement unit 11. The skip gate unit 116 calculates the value t immediately before the current time t. pvIf there is a non-speech period in t seconds (i.e., a period during which the momentary loudness value is equal to or less than the skip gate threshold), the non-speech period is skipped and the preceding speech period is included in the effective measurement time instead, as shown in Figure 4(b). pv In other words, in Figure 4(b), t pv =t pv1 +t pv2 is.
[0038] Furthermore, the skip gate unit 116 may use a skip gate threshold value that reflects the average loudness value up to that time, rather than a fixed value, used to determine a silent section. This can prevent erroneous determination of a silent section due to the presence of "overlap," in which unintended sounds other than the speaker's voice, such as ambient sounds, are input to the microphone. In addition, it can reduce the impact of differences in the input voice levels depending on the speaker.
[0039] The time rate upper average value calculation unit 117 calculates the effective measurement time t pv The momentary loudness values included in pv The momentary loudness calculation unit 13 calculates the average value of momentary loudness values (that is, momentary loudness values at a predetermined time rate) corresponding to the momentary loudness values (i.e., momentary loudness values at a predetermined time rate [%]) as an objective index and outputs the calculated value to the gain calculation unit 13.
[0040] FIG. 5 is a diagram for explaining the processing of the time rate upper average value calculation unit 117. The horizontal axis of the graph shown in FIG. 5 is time, and the vertical axis is momentary loudness value. The time rate refers to the proportion of a time length that exceeds a certain audio level in a certain time length, and in the present invention, refers to the proportion of a time period in which the momentary loudness value exceeds the gating threshold in the effective measurement time. In the example shown in FIG. 5(a), the time rate [%]=(t o1 +t o2 +t o3 )×100 / t pvAs shown in Fig. 5(b), the time rate changes by changing the gating threshold. Then, the time rate upper average value calculation unit 117 outputs the objective index of the reference voice object to the gain calculation unit 13 as a reference objective index, and outputs the objective index of the target voice object to the gain calculation unit 13 as a target objective index.
[0041] 6 is a flowchart showing an example procedure for measuring an objective index by the objective index measurement unit 11. In step S111, the prefilter 111 performs prefiltering processing on the reference audio object and the target audio object. In step S112, the root mean square unit 112 performs root mean square processing. In step S113, the weighting unit 113 performs weighting processing. In step S114, the summation unit 114 performs summation processing. In step S115, the decibel scale conversion unit 115 performs decibel scale conversion processing. The objective index measurement unit 11 calculates a momentary loudness value through the processing of steps S111 to S115.
[0042] Next, in step S116, skip gate unit 116 performs skip gate processing on the momentary loudness values. In step S117, upper time rate average calculation unit 117 calculates the average value of the momentary loudness values with higher time rates as the objective index. This allows objective index measurement unit 11 to measure an objective index that is based on the loudness value measurement algorithm and has both stability and responsiveness.
[0043] Example 1 Next, a sound processing device 1a according to a first embodiment will be described. Fig. 7 is a block diagram showing an example of the configuration of the sound processing device 1a according to the first embodiment. The sound processing device 1a includes an objective index measurement unit group 11s, an utterance timing determination unit 12, a gain calculation unit 13a, and a level adjustment unit 14. In this embodiment, the objective index measurement unit group 11s includes three objective index measurement units 11 (11-1, 11-2, 11-3).
[0044] The objective index measurement units 11-1, 11-2, and 11-3 measure the reference objective index and the target objective index, respectively, for different effective measurement times. In this embodiment, the effective measurement times of the objective indexes in the objective index measurement units 11-1, 11-2, and 11-3 are long (60 seconds), medium (20 seconds), and short (2 seconds), respectively. The speech timing determination unit 12 then determines the similarity of the speech timing from the correlation coefficient between the reference objective index and the target objective index measured for the short effective measurement time (2 seconds).
[0045] Furthermore, the gain calculation unit 13a calculates the difference between the reference objective index and the target objective index, and performs weighted addition of the difference. At this time, the gain calculation unit 13a increases the weight (weighting coefficient) for the difference when the effective measurement time is short, as the similarity of the utterance timing increases. This will be specifically described below.
[0046] FIG. 8 is a diagram illustrating the role of each objective indicator measurement unit 11. FIG. 8(a) shows an example in which the speech timing of the reference audio object and the target audio object closely match. For example, for content in which the speech content is the same and the speech timing closely matches, such as Japanese and English dubbing, adjustments are made based on objective indicators measured over a short effective measurement time in order to sequentially align the target audio object (English dialogue) with the reference audio object (Japanese dialogue). On the other hand, as shown in FIG. 8(c), when the speech content and speech timing of the reference audio object and the target audio object are completely different (for example, when the secondary audio has almost no relation to the progress of the program and a comedian is telling jokes), it is not possible to align the audio level based on objective indicators measured over a short effective measurement time for the reference audio object, and the overall audio level is adjusted on average using objective indicators measured over a longer effective measurement time. Furthermore, as in Figure 8(b), when the audio level of the target audio object follows the reference audio object at a certain speed, even if not as quickly as in Figure 8(a) (for example, when different speakers provide commentary and analysis for the main audio and secondary audio in a sports broadcast, but the speakers are different but basically speaking in conjunction with the same content shown in the video), adjustments are made based on the objective index measured over an effective measurement time that is intermediate in length between Figures 8(a) and 8(c).
[0047] Next, the criteria for selecting the long, medium, and short effective measurement times are shown below: Fig. 9 is a table showing examples of the speech content of the reference audio object and the target audio object in a professional baseball broadcast.
[0048] Figure 10 is a graph showing the target objective indices measured over various effective measurement times for the expert commentary dialogue of the target audio object shown in Figure 9. The effective measurement time is shown on the left side of Figure 10. "Integrated" refers to the entire time from the start of the content to the current time. Figure 10 shows that as the effective measurement time is extended, the waveform becomes smoother and the amplitude of fluctuations decreases. For long and medium effective measurement times (hereafter referred to as "medium times"), the objective is to averagely match the overall audio level rather than sequentially adjust using objective indices measured over short periods, so it is undesirable for the objective indices to fluctuate significantly over time. Furthermore, even if the value fluctuations are relatively gradual, forcing the target audio object to match the reference audio object when the reference and target objective indices have completely different fluctuation trends is likely to result in unnatural adjustments.
[0049] Therefore, as shown in Figure 11, we calculated the variance of the objective indicators measured while changing the effective measurement time. We found that the variance decreased as the effective measurement time increased, reducing the fluctuation in the objective indicator values. Furthermore, of the seven types of dialogue measured this time, the shortest effective measurement time was around 20 to 30 seconds, and the longest was around 60 to 70 seconds, at which point the integrated value approached and the fluctuations were sufficiently small. Therefore, we selected 20 and 60 seconds as candidates for the medium and long effective measurement times. Furthermore, because the short period is close to the adjustment speed of a mixing engineer, we set it at 1.6 to 2 seconds. In other words, we could select three effective measurement times: 60 to 70 seconds, 20 to 30 seconds, and 1.6 to 2 seconds.
[0050] Figure 12 also shows the correlation coefficient between the reference objective index and the target audio objective index. The correlation coefficient is an index that indicates the similarity in the fluctuation trends of the reference objective index and the target objective index. Although there are some exceptions, the correlation coefficient generally increases as the effective measurement time is extended, and the value converges once a certain length is reached; in other words, the fluctuation trends of the reference objective index and the target objective index become most similar. For three of the six target audio objects measured this time, the effective measurement time was approximately 15 to 30 seconds, and for the remaining three, the change in the correlation coefficient converged when the effective measurement time was approximately 60 to 70 seconds, and the value was close to the correlation coefficient calculated using the integrated (full time) objective index.
[0051] Based on the results of Figures 11 and 12, in this example, three effective measurement times were determined, with the long time being 60 seconds and the medium time being 20 seconds. Note that there may be two or four or more effective measurement times. Furthermore, the effective measurement time may be set using a different standard, or if the candidate effective measurement times differ depending on the standard, these effective measurement times may be used in combination.
[0052] The speech timing determination unit 12 calculates the correlation coefficient between the reference objective index measured during the shortest valid measurement time and the target objective index as the similarity of the speech timing. The objective index is an algorithm based on loudness values, and when the valid measurement time is short, the value increases at the timing of speech and decreases at the end of the speech. Therefore, if the fluctuations in the values are similar between the reference audio object and the target audio object, the speech timing is considered to be similar. The short valid measurement time was set to 2 seconds, which is close to the adjustment speed of a mixing engineer.
[0053] Referring again to FIG. 7, the gain calculation unit 13a includes three difference calculation units 131 (131-1, 131-2, 131-3), three weighting units 132 (132-1, 132-2, 132-3), and a summing unit 133.
[0054] The difference calculation unit 131 calculates an objective index difference, which is the difference between the reference objective index and the target objective index, and outputs the calculated objective index difference to the weighting unit 132.
[0055] The weighting unit 132 calculates a weighted difference by weighting the objective index difference input from the difference calculation unit 131. Then, the weighting unit 132 outputs the calculated weighted difference to the summation unit 133.
[0056] FIG. 13 shows the relationship between the determination result by the speech timing determination unit 12, i.e., the correlation coefficient between the reference objective index (short time) and the target objective index (short time), and the weight multiplied by the weighting unit 132 based on the determination result. As shown in FIG. 13(a), the weights are set based on simple trigonometric functions. When the correlation coefficient is 1, i.e., when the reference voice object and the target voice object are perfectly correlated, this means that the speech timings are perfectly matched, and it is considered that sequential adjustment over a short period of time is sufficient, so the weight is set to 1.0 when the effective measurement time is short. On the other hand, when the correlation coefficient is 0, and the reference voice object and the target voice object are completely uncorrelated, the speech timings of the reference voice object and the target voice object are completely different, and sequential adjustment over a short period of time is not appropriate, so the weight is set to 1.0 when the effective measurement time is long. In other words, the weighting unit 132 increases the weight applied to the objective index difference when the effective measurement time is short, as the similarity increases.
[0057] The correlation can also be negative, indicating that the timing of the reference and target audio objects is exactly opposite (e.g., when they alternate). However, content typically contains common audio, such as background sounds, in addition to dialogue. Therefore, it is unlikely that the timing of the reference and target audio objects will have a meaningful negative correlation. In fact, Figure 8 shows that only one of the six target audio objects exhibits a negative correlation coefficient, and its absolute value is not significant, at most -0.2. Therefore, negative correlation coefficients are treated as completely uncorrelated (0), and a weight of 1.0 is used when the effective measurement time is long. To ensure that the sum of the weights set above is always 1, normalization is performed as shown in Figure 13(b). The vertical lines in Figure 13(b) indicate the weights applied to each target audio object based on the short-term correlation coefficient with the reference audio object.
[0058] The summing unit 133 calculates a gain by summing the weighted differences input from each weighting unit 132. Then, the summing unit 133 outputs the calculated gain to the level adjusting unit .
[0059] 14 is a flowchart showing an example of the processing procedure of the gain calculation unit 13a. In step S131, the difference calculation unit 131 calculates the difference between the objective indices. In step S132, the weighting unit 132 performs weighting based on the similarity of the utterance timing. In step S133, the summation unit 133 performs weighted addition of the differences between the objective indices to calculate the gain.
[0060] Example 2 Next, a sound processing device 1b according to Example 2 will be described. Fig. 15 is a block diagram showing an example of the configuration of the sound processing device 1b according to Example 2. The sound processing device 1b includes an objective index measurement unit group 11s, an utterance timing determination unit 12, a gain calculation unit 13b, and a level adjustment unit 14. In Example 2, the sound processing device 1b includes the gain calculation unit 13b instead of the gain calculation unit 13a of Example 1.
[0061] The objective index measurement units 11-1, 11-2, and 11-3 measure the reference objective index at different effective measurement times and measure the target objective index at the shortest effective measurement time. In this embodiment, the effective measurement times of the objective indexes in the objective index measurement units 11-1, 11-2, and 11-3 are long (60 seconds), medium (20 seconds), and short (2 seconds), respectively. The speech timing determination unit 12 then determines the similarity of the speech timing from the correlation coefficient between the reference objective index and the target objective index measured at the short effective measurement time (2 seconds) of the reference voice object and the target voice object.
[0062] The gain calculation unit 13b calculates the difference between the weighted sum of the reference objective index and the target objective index. At this time, the gain calculation unit 13b increases the weight of the reference objective index when the effective measurement time is short as the similarity of the utterance timing is higher.
[0063] The number of objective index measurement units 11, the effective measurement time, and other concepts in the second embodiment are all the same as those in the first embodiment. However, due to the algorithm of taking the average of audio levels over a long time width for an objective index measured over a long effective measurement time, there is a delay before the actual loudness of the sound is reflected in the value of the objective index, and in the first embodiment, this delay also affects the adjustment value when the weighting of the objective index over a long time period becomes large. Therefore, in the second embodiment, in order to reduce the effect of the delay, the reference audio object that serves as the correct value to be fitted is a value calculated by also taking into account long-term and medium-term reference objective indexes, as in the first embodiment, whereas the target audio object is always fitted to the reference audio object sequentially based on a short-term target objective index.
[0064] The gain calculation unit 13b includes three weighting units 132 (132-1, 132-2, 132-3), a summing unit 133, and a difference calculation unit 131.
[0065] The weighting unit 132 calculates a weighted reference objective index by weighting the reference objective index input from the objective index measurement unit 11. Then, the weighting unit 132 outputs the calculated weighted reference objective index to the summing unit 133.
[0066] The summing unit 133 calculates a sum by summing the weighted reference objective indices input from each weighting unit 132. Then, the summing unit 133 outputs the calculated sum to the difference calculation unit 131.
[0067] The difference calculation unit 131 calculates, as a gain, the difference between the summed value input from the summing unit 133 and the short-term reference objective index input from the objective index measurement unit 11-3. Then, the difference calculation unit 131 outputs the calculated gain to the level adjustment unit 14.
[0068] A computer can be suitably used to function as the above-mentioned sound processing devices 1, 1a, and 1b, and such a computer can be realized by storing a program describing the processing contents for realizing each function of the sound processing devices 1, 1a, and 1b in a storage unit of the computer, and reading and executing this program by the CPU of the computer. Note that this program can be recorded on a computer-readable recording medium.
[0069] The program may also be recorded on a computer-readable medium. The computer-readable medium allows the program to be installed on a computer. Here, the computer-readable medium on which the program is recorded may be a non-transitory recording medium. The non-transitory recording medium is not particularly limited, and may be, for example, a recording medium such as a CD-ROM or a DVD-ROM.
[0070] As described above, the sound processing devices 1, 1a, and 1b according to the present invention and the programs therefor determine the similarity of speech timing between a reference audio object and a target audio object, and determine the gain of the target audio object based on the similarity so as to bring the target objective index closer to the reference objective index. Therefore, according to the present invention, even if the speech timings of the reference dialogue (adjusted main audio dialogue) and the target dialogue (secondary audio dialogue) differ, it is possible to automatically adjust the gain of the target dialogue to an optimum (natural volume).
[0071] In Example 1, the effective measurement times for both the reference objective index and the target objective index are of three types: long, medium, and short. In Example 2, the effective measurement time for the reference objective index is of three types: long, medium, and short, and the effective measurement time for the target objective index is fixed to short. In a dialogue having a high similarity to the reference audio object, including a short time in the effective measurement time for the target objective index is considered to contribute to gain adjustment, so optimal gain adjustment is possible in both Example 1 and Example 2. In a dialogue having a low similarity to the reference audio object and a large variance of the target objective index, it is considered that excessive gain adjustment may occur if a long time is not included in the effective measurement time for the target objective index, so Example 1 is more preferable. In a dialogue having a low similarity to the reference audio object and a small variance of the target objective index, it is considered that a long effective measurement time for the reference objective index is sufficient, so optimal gain adjustment is possible in both Example 1 and Example 2.
[0072] Furthermore, in the present invention, the time during which the momentary loudness value exceeds the skip gate threshold is regarded as speech time, and momentary loudness values for a fixed measurement time are extracted from the speech time immediately prior to the current time, and the average value of the momentary loudness values for a predetermined time percentage is calculated as an objective index. In other words, in the present invention, silent sections (non-speech time) are not included in the measurement time. Therefore, according to the present invention, in the production of secondary audio for video and audio content such as television programs, it is possible to automatically adjust the gain of audio objects such as dialogue using an objective index that is both stable and responsive.
[0073] Furthermore, according to the present invention, when producing multiple patterns of audio, such as secondary audio, for audiovisual content, it is possible to generate dialogue signals of different patterns adjusted to the same volume as the dialogue adjusted by an engineer for the main audio, etc., without increasing the number of mixing engineers required for production, even in live broadcasts. Furthermore, even if the number of variations in audio to be produced simultaneously increases in the future, similar effects can be expected by applying the present invention according to the number of added patterns.
[0074] Furthermore, in audio content, if the difference is within a range of about a few dB, it is generally believed that a temporarily too-low audio object will have a greater impact on the viewer than a temporarily too-high audio object, since it may be masked by background sounds, etc., resulting in less information received by the viewer. Therefore, by changing the weighting value when adjusting to lower the fader, it is possible to reduce the risk that the target audio object will be masked by background sounds, etc., and become difficult to hear when it is automatically adjusted to be lower.
[0075] Furthermore, when producing audio content of a certain level of quality or higher, it is common for a dedicated microphone to be installed for each main speaker. Therefore, it will be clear to those skilled in the art that the present invention can be applied not only to object audio systems and channel-based audio systems, but also to any audio system that will be newly established in the future, as long as it is possible to handle the speaker's audio signal individually.
[0076] Although the above-described embodiments have been described as typical examples, it will be apparent to those skilled in the art that many modifications and substitutions can be made within the spirit and scope of the present invention. Therefore, the present invention should not be construed as being limited by the above-described embodiments, and various modifications and alterations are possible without departing from the scope of the claims. For example, multiple building blocks shown in the block diagrams of the embodiments can be combined into one, or one building block can be divided. [Explanation of symbols]
[0077] 1, 1a, 1b Sound processing device 11s Objective Indicator Measurement Group 11 Objective Indicator Measurement Department 12 Speech timing determination unit 13, 13a, 13b Gain calculation section 14 Level adjustment section 110 Momentary loudness value measurement section 111 Pre-filter 112 Mean Square Section 113 Weighting section 114 Addition Section 115 Decibel scale converter 116 Skip gate section 117 Time rate upper average calculation unit 118 Removal gate section 131 Difference calculation part 132 Weighting section 133 Addition Section
Claims
1. A sound processing device that automatically adjusts a target audio object based on a reference audio object that is an audio object of a reference audio serving as a reference and a target audio object that is an audio object of a target audio to be adjusted, an objective index measurement unit that measures a reference objective index, which is an objective index based on the loudness value of the reference audio object, and a target objective index, which is an objective index based on the loudness value of the target audio object; an utterance timing determination unit that determines a similarity in utterance timing between the reference voice object and the target voice object; a gain calculation unit that determines a gain of the target voice object based on the similarity so as to bring the target objective index closer to the reference objective index; a level adjustment unit that adjusts the audio level of a target audio object based on the gain; Equipped with the objective index measurement unit measures the reference objective index and the target objective index at different effective measurement times, respectively; The gain calculation unit calculates the difference between the reference objective index and the target objective index for each effective measurement time, weights and adds the difference to calculate the gain, and increases the weighting for the difference when the effective measurement time is short as the similarity is higher.
2. A sound processing device that automatically adjusts a target audio object based on a reference audio object that is an audio object of a reference audio serving as a reference, and a target audio object that is an audio object of a target audio to be adjusted, comprising: an objective index measurement unit that measures a reference objective index, which is an objective index based on the loudness value of the reference audio object, and a target objective index, which is an objective index based on the loudness value of the target audio object; an utterance timing determination unit that determines a similarity in utterance timing between the reference voice object and the target voice object; a gain calculation unit that determines a gain of the target voice object based on the similarity so as to bring the target objective index closer to the reference objective index; a level adjustment unit that adjusts the audio level of a target audio object based on the gain; Equipped with the objective index measurement unit measures the reference objective index at different effective measurement times, and measures the target objective index at the shortest measurement time among the effective measurement times; The gain calculation unit calculates the gain as the difference between the weighted sum of the reference objective index and the target objective index, and the higher the similarity, the greater the weight given to the reference objective index when the effective measurement time is short.
3. 3. The sound processing device according to claim 1, wherein the effective measurement time is set to 60 to 70 seconds, 20 to 30 seconds, and 1.6 to 2 seconds.
4. The sound processing device according to claim 1 , wherein the speech timing determination unit calculates a correlation coefficient between the reference objective index and the target objective index as the degree of similarity.
5. The objective index measurement unit a momentary loudness value measurement unit that measures a momentary loudness value; a skip gate unit that regards a time during which the momentary loudness value exceeds a skip gate threshold as a speech time and outputs a momentary loudness value of a fixed length of effective measurement time within the speech time immediately before the current time; a time rate upper average value calculation unit that calculates an average value of momentary loudness values for a predetermined time rate among the momentary loudness values for the effective measurement time as the objective index; The sound processing device according to claim 1 , comprising:
6. A program for causing a computer to function as the sound processing device according to any one of claims 1 to 5.
Citation Information
Patent Citations
Signal processor
JP2016148818A
Loudness regulator and program
JP2017092818A
Acoustic processing device, acoustic processing system, and program
JP2022042892A