Media data playing loudness processing method and device, electronic equipment and medium
By analyzing the distribution of voice loudness in media data and adjusting the playback loudness to achieve balance, the problem of uneven voice loudness is solved, thus improving the user experience.
Patent Information
- Application Number
- CN202411087752.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-08
- Publication Date
- 2026-02-10
AI Technical Summary
Existing technologies suffer from poor user experience due to uneven voice volume during media data playback.
By acquiring voice data from media data and analyzing its loudness distribution, the playback loudness of the media data is adjusted based on the voice loudness metadata to achieve a balance in voice loudness.
It improved the playback effect of media data and enhanced the user's auditory experience.
Smart Images

Figure CN121509736A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of computer processing technology, specifically to methods, apparatus, electronic devices, and media for processing the playback loudness of media data. Background Technology
[0002] When media data is played on a terminal, the loudness varies depending on the media data being created. Therefore, to ensure balanced loudness during playback on the terminal, a balance adjustment is performed based on the overall loudness of the media data. However, this method of loudness balancing can easily lead to significant differences in voice loudness among different media data, thus affecting the user experience. Summary of the Invention
[0003] In view of this, the present disclosure provides a method, apparatus, electronic device and medium for processing the playback loudness of media data to solve the problem of playback loudness equalization.
[0004] Firstly, this disclosure provides a method for processing the playback loudness of media data, the method comprising:
[0005] Acquire media data;
[0006] In response to the inclusion of speech data in the media data, the loudness distribution results corresponding to the speech data are used.
[0007] Based on voice loudness metadata, the playback loudness of media data is adjusted to obtain the target media data for playback.
[0008] Secondly, this disclosure provides a media data playback loudness processing apparatus, the apparatus comprising:
[0009] The first acquisition module is used to acquire media data;
[0010] The first processing module is used to respond to the media data including speech data and to process the speech loudness distribution results corresponding to the speech data.
[0011] The second processing module is used to adjust the playback loudness of media data based on voice loudness metadata in order to obtain the target media data for playback.
[0012] Thirdly, this disclosure provides an electronic device, including: a memory and a processor, which are communicatively connected to each other. The memory stores computer instructions, and the processor executes the computer instructions to perform the media data playback loudness processing method of the first aspect or any corresponding embodiment described above.
[0013] Fourthly, this disclosure provides a computer-readable storage medium storing computer instructions for causing a computer to execute the media data playback loudness processing method of the first aspect or any corresponding embodiment described above.
[0014] Fifthly, the present invention provides a computer program product, including computer instructions for causing a computer to execute the media data playback loudness processing method of the first aspect or any corresponding embodiment described above.
[0015] The media data playback loudness processing method provided by the present invention, based on the loudness distribution results corresponding to the speech data in the media data, clarifies the speech loudness metadata of the speech data, and then adjusts the playback loudness of the media data based on the speech loudness metadata, which can ensure that the loudness of the speech data in the adjusted target media data is relatively balanced, thereby helping to improve the playback effect of the media data. Attached Figure Description
[0016] To more clearly illustrate the technical solutions in the specific embodiments of this disclosure or the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this disclosure. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0017] Figure 1 This is a schematic flowchart of a media data playback loudness processing method according to an embodiment of the present disclosure;
[0018] Figure 2 This is a schematic diagram of the speech loudness distribution results according to an embodiment of the present disclosure;
[0019] Figure 3 This is a flowchart illustrating another method for processing the playback loudness of media data according to an embodiment of the present disclosure;
[0020] Figure 4 This is a flowchart illustrating another method for processing the playback loudness of media data according to an embodiment of the present disclosure;
[0021] Figure 5 This is a flowchart illustrating a method for processing the playback loudness of media data according to an embodiment of the present disclosure;
[0022] Figure 6 This is a schematic flowchart of a media data playback loudness processing method according to an embodiment of the present disclosure;
[0023] Figure 7 This is a structural block diagram of a media data playback loudness processing device according to an embodiment of the present disclosure;
[0024] Figure 8 This is a schematic diagram of the hardware structure of an electronic device according to an embodiment of this disclosure. Detailed Implementation
[0025] To make the objectives, technical solutions, and advantages of the embodiments of this disclosure clearer, the technical solutions of the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this disclosure, and not all embodiments. Based on the embodiments of this disclosure, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this disclosure.
[0026] In related technologies, the loudness of the playback device is adjusted based on the overall loudness of the audio. However, when the loudness of the speech in the audio is uneven, the loudness changes are quite obvious. That is, there are some speech sounds with high loudness and some speech sounds with low loudness, which will affect the user experience.
[0027] In view of this, embodiments of the present disclosure provide a method for processing the loudness of media data playback to solve the problem that users experience different loudness levels of adjacent voice audio when playing media data.
[0028] According to an embodiment of this disclosure, a method for processing the playback loudness of media data is provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.
[0029] This embodiment provides a method for processing the playback loudness of media data, which can be used in the aforementioned mobile terminals, such as mobile phones and tablet computers. Figure 1 This is a flowchart of a media data playback loudness processing method according to an embodiment of the present disclosure, such as... Figure 1 As shown, the process includes the following steps:
[0030] Step S101: Obtain media data.
[0031] Media data refers to the media data to be played received by the playback client, including but not limited to video or audio. There are no restrictions on the format of the media data; it is set according to actual needs. For example, if the playback client has a short video playback application installed, when a user plays a short video using that application, the corresponding video, i.e., the media data, is retrieved from the short video playback application's server.
[0032] Step S102: In response to the inclusion of speech data in the media data, determine the speech loudness metadata of the speech data based on the speech loudness distribution results corresponding to the speech data.
[0033] When it is determined that the media data includes voice data, the loudness distribution of that voice data is analyzed to obtain, for example... Figure 2 The results of the speech loudness distribution are shown.
[0034] The loudness distribution results reveal the loudness mean, loudness range, and loudness variation of the speech data, thus providing speech loudness metadata that characterizes the audio features of the speech data.
[0035] In some examples, the content of speech loudness metadata includes, but is not limited to, any one or more of the following information: the proportion of speech data in the media data (speech_ratio), actual speech loudness (Loudness Range, LRA), average loudness (integrated_loudness), the starting loudness (LRA_start) and ending loudness (LRA_end) of the loudness range, the maximum instantaneous loudness (max_mom_loud), and the maximum short-term loudness (max_short_term_loud), etc.
[0036] Step S103: Based on the voice loudness metadata, adjust the playback loudness of the media data to obtain the target media data for playback.
[0037] By using voice loudness metadata, the loudness distribution of voice data in media data can be clearly identified. This allows for targeted equalization of the overall loudness of the voice data when adjusting the playback loudness of the media data, thereby effectively improving the playback effect of the subsequent target media data.
[0038] The media data playback loudness processing method provided by the present invention, based on the loudness distribution results corresponding to the speech data in the media data, clarifies the speech loudness metadata of the speech data, and then adjusts the playback loudness of the media data based on the speech loudness metadata, which can ensure that the loudness of the speech data in the adjusted target media data is relatively balanced, thereby helping to improve the playback effect of the media data.
[0039] In some optional implementations, the process of determining the speech loudness distribution corresponding to the speech data includes:
[0040] Step a1: Perform speech detection on the media data to identify the media data segments that correspond to the speech data.
[0041] Step a2: Determine the loudness distribution of the media data segment and obtain the loudness distribution result;
[0042] Step a3: Based on the loudness distribution results, determine the loudness distribution results corresponding to the speech data.
[0043] Specifically, to identify speech data in media data, speech detection is performed on the media data to separate the speech data from the background noise data, thereby obtaining media data segments corresponding to the speech data. The content of media data segments includes, but is not limited to, a single word, a dialogue, a sentence, or continuous speech. In some optional implementation scenarios, speech detection can be performed on media data based on audio features (such as root mean square (RMS)) by creating an Audio Event Detection (AED) task. For example, the media data is processed to calculate the RMS value for each time point or time period. A suitable RMS threshold is determined based on the audio characteristics of the speech data and application requirements. The selection of the RMS threshold may need to be adjusted based on experience or experimentation. The RMS value calculated for the current time point or time period is compared with the threshold. If the calculated RMS value exceeds or equals the threshold, speech data is considered to exist; if the RMS value is lower than the threshold, speech data is considered not to exist.
[0044] After identifying the media data segment, analyzing its loudness distribution yields the corresponding loudness distribution result. Since this media data segment corresponds to the speech data within the media data, its loudness distribution result can be directly used as the speech loudness distribution result for the speech data.
[0045] In some examples, if there are multiple media data segments, it indicates that there are multiple discrete speech data segments in the media data. Therefore, in order to ensure the loudness equalization effect of the speech data, the process of determining the speech loudness distribution result corresponding to the speech data includes: according to the order of multiple media data segments, the current loudness distribution result is fused with the previous loudness distribution result in turn, and the fused result is used as the previous loudness distribution result to be fused for the next loudness distribution result, so as to obtain the speech loudness distribution result corresponding to the speech data.
[0046] To facilitate understanding, the following example illustrates the process: During speech detection of media data, if media data segment A is the first detected speech data, its loudness distribution result is used as the initial loudness distribution result for the corresponding speech data. Continuing with speech detection, if media data segment B is also detected as speech data, its loudness distribution result is fused with that of media data segment A, and the fused result is used as the previous loudness distribution result to be fused in the next loudness distribution result. If, during continued detection, media data segment C is also detected as speech data, its loudness distribution result is fused with the fusion result of media data segment B, and the fused result is used as the previous loudness distribution result to be fused in the next loudness distribution result, and so on, until the speech detection is complete. The final fusion result is then used as the loudness distribution result for the corresponding speech data.
[0047] If no new speech data is detected after media data segment B is detected until the speech detection is completed, the fusion result of the loudness distribution of media data segment B and the loudness distribution of media data segment A will be used as the intermediate speech loudness distribution result corresponding to the speech data.
[0048] Determining the speech loudness distribution results corresponding to the speech data through the above method can effectively reduce the interference of redundant media data such as background noise data or silence data on the analysis of speech loudness distribution, thereby helping to improve the reliability and accuracy of the speech loudness distribution results, and providing favorable data support for subsequent speech loudness equalization processing.
[0049] This embodiment provides a method for processing the playback loudness of media data, which can be used in the aforementioned mobile terminals, such as mobile phones and tablet computers. Figure 3 This is a flowchart of a media data playback loudness processing method according to an embodiment of the present disclosure, such as... Figure 3 As shown, the process includes the following steps:
[0050] Step S301: Obtain media data. For details, please refer to [link / reference]. Figure 1 Step S101 of the illustrated embodiment will not be described again here.
[0051] Step S302: In response to the inclusion of speech data in the media data, based on the speech loudness distribution result corresponding to the speech data, determine the speech loudness metadata of the speech data. For details, please refer to [link to details]. Figure 1 Step S102 of the illustrated embodiment will not be described again here.
[0052] Step S303: Based on the voice loudness metadata, adjust the playback loudness of the media data to obtain the target media data for playback.
[0053] Specifically, step S303 includes:
[0054] Step S3031: Based on the overall loudness distribution of the media data, obtain the overall loudness distribution result.
[0055] To ensure the effectiveness of speech loudness balance, the overall loudness distribution of the media data is analyzed to clarify the overall loudness distribution of the media data and thus obtain the overall loudness distribution result.
[0056] Step S3032: Determine the overall loudness metadata based on the overall loudness distribution results to obtain the overall loudness of the media data.
[0057] By analyzing the overall loudness distribution of media data, we can clearly identify the mean loudness, overall loudness range, and loudness variations, thus obtaining overall loudness metadata that characterizes the overall audio features of the media data. This metadata is then used to calculate the overall loudness (Program Loudness, PL) of the media data. The overall loudness can be obtained by analyzing the overall loudness metadata using a pre-defined algorithm or standard. For example, the overall loudness metadata can be mapped to a specific loudness metric or analyzed using a loudness assessment model to arrive at the final loudness value.
[0058] In some examples, the content of the overall loudness metadata includes, but is not limited to, any one or more of the following information: the overall loudness range (LRA) of the media data, the average loudness (integrated_loudness), the starting loudness (LRA_start) and ending loudness (LRA_end) of the loudness range, the maximum instantaneous loudness (max_mom_loud), and the maximum short-term loudness (max_short_term_loud).
[0059] Step S3033: Determine the actual loudness of the speech data using speech loudness metadata.
[0060] The actual speech loudness (DL) of the speech data can be obtained by analyzing the speech loudness metadata. For example, the actual speech loudness can be obtained by analyzing and processing the speech loudness metadata using a preset algorithm or standard.
[0061] Step S3034: Determine the target loudness-to-dialogue ratio of the media data based on the difference between the overall loudness and the actual speech loudness.
[0062] Based on audio playback standards, it is clear that the Loudness-to-Dialogue Ratio (LDR) is determined by the difference between the overall loudness (PL) and the dialogue loudness (DL). Therefore, given the overall loudness of the media data and the actual speech loudness of the speech data, the difference between the two is used as the target loudness-to-dialogue ratio for the media data.
[0063] The formula for determining the target loudness-to-dialogue ratio is as follows: LDR = PL - DL.
[0064] Step S3035: Based on the target loudness-to-dialogue ratio, adjust the playback loudness of the media data to obtain the target media data for playback.
[0065] The target loudness-to-dialogue ratio clarifies the balance between the actual loudness and the overall loudness when performing loudness equalization on voice data. Consequently, when adjusting the playback loudness of media data based on this target loudness-to-dialogue ratio, it ensures that the loudness of the target media data is relatively balanced during the playback of voice data, avoiding it from being too loud or too weak. This effectively improves the playback effect of the target media data and enhances the user experience.
[0066] The media data playback loudness processing method provided in this embodiment determines speech loudness metadata based on speech loudness distribution results and overall loudness metadata based on overall loudness distribution results. It can accurately describe the loudness of speech and overall media data. Furthermore, by comparing the actual speech loudness and the overall loudness, it determines the target loudness-to-dialogue ratio, which enables the adjusted media data to maintain a suitable playback loudness during playback, thereby providing a better auditory experience.
[0067] In some optional implementations, step S3035 above includes:
[0068] Step b1: Obtain the target playback loudness of the voice data;
[0069] Step b2: Determine the dynamic range control parameters of the media data based on the target playback loudness, the actual speech loudness, and the target loudness-to-dialogue ratio.
[0070] Step b3: Adjust the playback loudness of the media data based on the dynamic range control parameters to obtain the target media data for playback.
[0071] Specifically, the target playback loudness of voice data can be determined based on loudness requirement information. Loudness requirement information includes the current playback environment; the noisier the environment, the greater the target loudness. In determining the target loudness, factors such as the noise level of the current environment or the playback capabilities of the playback device can be considered.
[0072] Since the actual loudness of the speech data varies at different times, in order to balance the loudness of the speech data, the dynamic range control parameters of the media data are determined based on the target playback loudness, the actual speech loudness, and the target loudness-to-dialogue ratio. Then, the playback loudness of the media data is adjusted based on these dynamic range control parameters, which can achieve dynamic adjustment and obtain target media data that can improve the auditory effect.
[0073] In some examples, the dynamic range control parameters include the compression ratio of the dynamic range; therefore, step b3 above includes:
[0074] Step b31: Determine the first loudness compression ratio based on the ratio between the actual speech loudness and the target playback loudness;
[0075] Step b32: Determine the second loudness compression ratio based on the ratio between the target loudness-to-dialogue ratio and the specified loudness-to-dialogue ratio;
[0076] Step b33: Based on the comparison between the first loudness compression ratio and the second loudness compression ratio, determine the dynamic range compression ratio.
[0077] Specifically, actual speech loudness refers to the true loudness level of the speech data, while target playback loudness is the desired playback loudness. By calculating the ratio between them, a first loudness compression ratio representing the difference in actual loudness can be obtained. The first loudness compression ratio determines the extent to which the actual speech loudness of the speech data needs to be compressed to make it closer to the target playback loudness. That is, the first loudness compression ratio ratio1 = actual speech loudness anchor_lra / target playback loudness target_lra.
[0078] The specified loudness-to-dialogue ratio is a pre-defined standard or reference ratio. This specified loudness-to-dialogue ratio can be determined based on a loudness-to-dialogue ratio range standard. For example, if the loudness-to-dialogue ratio range standard is 4–8 LU, then the specified loudness-to-dialogue ratio can be any value within 4–8 LU, such as 5 LU. The specified loudness-to-dialogue ratio can be set according to actual needs.
[0079] The second loudness compression ratio, obtained by calculating the ratio between the target loudness-to-dialogue ratio and the specified loudness-to-dialogue ratio, clarifies how to adjust the dynamic range compression to achieve the target loudness-to-dialogue ratio. That is, the second loudness compression ratio ratio2 = target loudness-to-dialogue ratio LDR / specified loudness-to-dialogue ratio.
[0080] The dynamic range compression ratio determines the degree of dynamic compression of the actual loudness range. Therefore, determining the dynamic range compression ratio based on a comparison of the first and second loudness compression ratios allows for greater flexibility in the compression ratio determination process and helps improve the adjusted auditory effect. For example, one loudness compression ratio can be chosen or a weighted average can be used to determine the final dynamic range compression ratio, thereby adjusting the overall loudness balance.
[0081] Preferably, the larger of the first loudness compression ratio and the second loudness compression ratio can be used as the dynamic range compression ratio, which helps to improve the efficiency of determining the compression ratio. Moreover, the determined compression ratio can not only achieve the goal of achieving the target loudness-to-dialogue ratio, but also make the actual speech loudness close to the target playback loudness.
[0082] In other examples, the dynamic range control parameters also include static characteristic thresholds. Therefore, the starting loudness of the actual speech loudness can be determined through speech loudness metadata, and the starting loudness can be used as the static characteristic threshold to ensure that the dynamic range of adjustment is determined based on the starting loudness of the actual speech. This ensures that the adjusted actual speech loudness is always adjusted based on the same starting loudness, thus better maintaining the consistency of speech loudness and making the loudness performance of speech data more stable and reliable.
[0083] In some alternative embodiments, step S3035 further includes:
[0084] Step c1: Obtain the reference loudness of the speech data;
[0085] Step c2: Determine the speech loudness gain of the speech data based on the difference between the reference speech loudness and the actual speech loudness.
[0086] Step c3: Based on the speech loudness gain and dynamic range control parameters, adjust the playback loudness of the media data to obtain the target media data for playback.
[0087] Specifically, the reference speech loudness can be a fixed value determined according to a certain standard or setting, or it can be a dynamic value determined based on the current playback environment or user needs, used as a reference value for adjusting speech loudness.
[0088] By determining the difference between the reference and actual speech loudness, the numerical relationship between the two loudnesses and the direction of loudness adjustment can be clearly identified, thus obtaining the speech loudness gain used to adjust the actual speech loudness. For example, if the actual speech loudness is less than the reference loudness, the difference is positive, indicating that the actual speech loudness needs to be increased based on this difference, resulting in the speech loudness gain for processing the speech data. If the actual speech loudness is greater than the reference loudness, the difference is negative, indicating that the actual speech loudness needs to be decreased based on this difference, resulting in the speech loudness gain for processing the speech data. Determining the speech loudness gain in this way allows for appropriate adjustments to the actual speech loudness based on actual conditions. Furthermore, by adjusting the playback loudness of media data based on the speech loudness gain and dynamic range control parameters, greater precision can be achieved. This effectively avoids over-amplification or over-compression of the actual speech loudness, resulting in target media data for playback that enhances the user's auditory experience.
[0089] In some optional scenarios, if the reference speech loudness is 50dB and the actual speech loudness is 40dB, then the speech loudness gain = 50dB - 40dB = +10dB. When adjusting the actual speech loudness of the speech data, it is necessary to increase the actual speech loudness by amplifying or reducing the loudness amplitude of the speech data so that the adjusted actual speech loudness meets expectations.
[0090] This embodiment provides a method for processing the playback loudness of media data, which can be used in the aforementioned mobile terminals, such as mobile phones and tablet computers. Figure 4 This is a flowchart of a media data playback loudness processing method according to an embodiment of the present disclosure, such as... Figure 4 As shown, the process includes the following steps:
[0091] Step S401: Obtain media data.
[0092] Step S402: Obtain the historical playback configuration information of the playback device.
[0093] The playback device is the device used to play the target media data. Historical playback configuration information may include, but is not limited to, the playback device's external loudness configuration parameters and playback mode during historical media data playback. External loudness configuration parameters indicate the user's volume settings on the device during historical playback. For example, the user may have adjusted the volume up or down at different times or in different situations. Playback modes may include speaker mode (using built-in or external speakers), headphone mode, etc. The selection of these modes will also affect the audio playback quality.
[0094] Step S403: Based on the analysis results of the historical playback configuration information, determine the target loudness equalization mode.
[0095] By analyzing the historical playback configuration information, we can clarify the user's volume preferences during the historical use of the playback device, and then select an appropriate equalization mode as the target loudness equalization mode. This helps ensure that the actual voice loudness after subsequent adjustment better meets the user's expectations.
[0096] The target loudness equalization mode can include, but is not limited to, any of the following equalization modes: speech equalization mode and default equalization mode. Speech equalization mode can be understood as an equalization mode that prioritizes clear playback of speech data and requires loudness equalization of the actual speech loudness. Default equalization mode can be understood as a general equalization mode that performs overall loudness equalization on the media to be played.
[0097] Step S404: In response to the target loudness equalization mode being speech equalization mode, determine whether the media data includes speech data.
[0098] Step S405: In response to the inclusion of speech data in the media data, based on the speech loudness distribution results corresponding to the speech data.
[0099] Step S406: Based on the voice loudness metadata, adjust the playback loudness of the media data to obtain the target media data for playback.
[0100] The media data playback loudness processing method provided in this embodiment determines the target loudness equalization mode through the historical playback configuration information of the playback device. When the target loudness equalization mode is the voice equalization mode, the playback loudness of the media data is adjusted based on the voice loudness metadata. This enables the adjusted target media data to have an actual voice loudness that better matches expectations during playback, thereby helping to improve the user experience.
[0101] This embodiment provides a method for processing the playback loudness of media data, which can be used in the aforementioned mobile terminals, such as mobile phones and tablet computers. Figure 5 This is a flowchart of a media data playback loudness processing method according to an embodiment of the present disclosure, such as... Figure 5 As shown, the process includes the following steps:
[0102] Step S501: Obtain media data.
[0103] Step S502, in response to the media data including speech data, based on the speech loudness distribution result corresponding to the speech data.
[0104] Specifically, step S502 includes:
[0105] Step S5021: In response to the inclusion of voice data in the media data, determine the first duration of the media data and the second duration of the voice data respectively.
[0106] To determine the distribution of voice data within the media data, a first duration for the media data and a second duration for the voice data are determined. The first duration can be understood as the total playback time of the media data, and the second duration as the total playback time of the voice data. The second duration is less than or equal to the first duration.
[0107] Step S5022: If the ratio between the second duration and the first duration is greater than a preset threshold, then the speech loudness metadata of the speech data is determined based on the speech loudness distribution result corresponding to the speech data.
[0108] If the ratio between the second duration and the first duration is greater than a preset threshold, it indicates that there is relatively more speech data in the media data, and the loudness equalization processing of the speech data is effective. Therefore, to make the actual speech loudness of the speech data more balanced, the speech loudness metadata of the speech data is determined based on the speech loudness distribution results corresponding to the speech data. The preset threshold can be determined according to actual needs. For example, the preset threshold can be 15%.
[0109] Step S503: Based on the voice loudness metadata, adjust the playback loudness of the media data to obtain the target media data for playback.
[0110] The media data playback loudness processing method provided in this embodiment adjusts the playback loudness of the media data based on the voice loudness metadata when the ratio between the second duration and the first duration is determined to be greater than a preset threshold. This ensures that the loudness equalization performed on the voice data is effective, thereby guaranteeing the playback effect of the media data.
[0111] In some optional implementations, the above method further includes:
[0112] Step S504: If the ratio between the second duration and the first duration is less than or equal to a preset threshold, then the overall loudness distribution result is obtained based on the overall loudness distribution of the media data.
[0113] If the ratio between the second duration and the first duration is less than or equal to a preset threshold, it indicates that the speech data in the media data is relatively small. If loudness equalization processing is continued on the speech data, the effect will be minimal and it will be considered an ineffective process. Therefore, in order to ensure that the overall loudness of the media data is balanced, the overall loudness distribution of the media data is analyzed to obtain an overall loudness distribution result that reflects the overall loudness distribution of the media data.
[0114] Step S505: Based on the overall loudness metadata corresponding to the overall loudness distribution result, adjust the playback loudness of the media data to obtain the target media data for playback.
[0115] By using overall loudness metadata, the overall loudness distribution of media data can be clearly identified. This allows for a more balanced overall loudness when adjusting the playback loudness of media data, thereby improving the playback effect.
[0116] This is one or more specific application embodiments of the present disclosure. Figure 6 The process flow for media data processing is illustrated, which can include a parameter preparation stage and a streaming processing stage. The preparation stage includes determining the loudness gain, determining the parameters of the DRC curve, and determining the loudness compensation gain. The streaming processing stage includes: applying loudness gain processing to the media data using the loudness gain to obtain first media data; obtaining the DRC curve using the parameters of the DRC curve, and then applying dynamic range control processing to the first media data using the DRC curve to obtain second media data; applying loudness compensation gain to the second media data to obtain third media data; and finally, applying peak limiting to the third media data to obtain the target media data.
[0117] In determining the loudness gain, if the target equalization mode is speech equalization mode, the determined loudness gain is the speech loudness gain. If the target equalization mode is the default equalization mode, the determined loudness gain is the overall loudness gain.
[0118] By using the above-mentioned loudness processing method for media data playback, loudness equalization processing can be performed, making loudness adjustment more flexible. Targeted loudness equalization processing can be carried out on the overall data of voice data or media data, thereby effectively improving the playback effect of media data and ensuring that the final playback loudness meets user expectations, thus achieving the goal of enhancing the auditory experience.
[0119] This embodiment also provides a media data playback loudness processing device, which is used to implement the above embodiments and preferred embodiments, and will not be repeated as already described. As used below, the term "module" can be a combination of software and / or hardware that implements a predetermined function. Although the device described in the following embodiments is preferably implemented in software, hardware implementation, or a combination of software and hardware, is also possible and contemplated.
[0120] This embodiment provides a media data playback loudness processing device, such as... Figure 7 As shown, it includes:
[0121] The first acquisition module 701 is used to acquire media data;
[0122] The first processing module 702 is used to respond to the media data including speech data and to process the speech loudness distribution result corresponding to the speech data;
[0123] The second processing module 703 is used to adjust the playback loudness of media data based on voice loudness metadata in order to obtain target media data for playback.
[0124] In some optional implementations, the apparatus for determining the speech loudness distribution result corresponding to the speech data includes:
[0125] The first detection module is used to perform voice detection on the media data and identify the media data segments that correspond to the voice data.
[0126] The second detection module is used to analyze the loudness distribution of media data segments and obtain loudness distribution results;
[0127] The third detection module is used to determine the speech loudness distribution result corresponding to the speech data based on the loudness distribution result.
[0128] In some optional implementations, if there are multiple media data segments, the third detection module includes:
[0129] The first processing unit is used to sequentially fuse the current loudness distribution result with the previous loudness distribution result according to the order of multiple media data segments, and use the fused result as the previous loudness distribution result to be fused for the next loudness distribution result, so as to obtain the speech loudness distribution result corresponding to the speech data.
[0130] In some alternative implementations, the second processing module 703 includes:
[0131] The analysis module is used to obtain the overall loudness distribution results based on the overall loudness distribution of media data;
[0132] The second processing unit is used to determine the overall loudness metadata based on the overall loudness distribution results, so as to obtain the overall loudness of the media data;
[0133] The third processing unit is used to determine the actual loudness of the speech data through speech loudness metadata;
[0134] The fourth processing unit is used to determine the target loudness-to-dialogue ratio of the media data based on the difference between the overall loudness and the actual speech loudness.
[0135] The fifth processing unit is used to adjust the playback loudness of the media data based on the target loudness-to-dialogue ratio to obtain the target media data for playback.
[0136] In some alternative implementations, the fifth processing unit includes:
[0137] The first acquisition unit is used to acquire the target playback loudness of the voice data;
[0138] The parameter determination unit is used to determine the dynamic range control parameters of the media data based on the target playback loudness, the actual speech loudness, and the target loudness-to-dialogue ratio.
[0139] The adjustment unit is used to adjust the playback loudness of media data based on dynamic range control parameters to obtain target media data for playback.
[0140] In some optional implementations, the dynamic range control parameters include the dynamic range compression ratio, and the second execution unit includes:
[0141] The first determining unit is used to determine the first loudness compression ratio based on the ratio between the actual speech loudness and the target playback loudness.
[0142] The second determining unit is used to determine the second loudness compression ratio based on the ratio between the target loudness-to-dialogue ratio and the specified loudness-to-dialogue ratio.
[0143] The third determining unit is used to determine the dynamic range compression ratio based on the comparison result between the first loudness compression ratio and the second loudness compression ratio.
[0144] In some optional implementations, the third determining unit includes:
[0145] The third determining subunit is used to take the larger of the first loudness compression ratio and the second loudness compression ratio as the compression ratio of the dynamic range.
[0146] In some optional implementations, the dynamic range control parameters also include a static characteristic threshold, and the second execution unit further includes:
[0147] The fourth determining unit is used to determine the initial loudness of the actual speech loudness through speech loudness metadata;
[0148] The fifth determining unit is used to set the initial loudness as the static characteristic threshold.
[0149] In some optional implementations, the fifth processing unit further includes:
[0150] The second acquisition unit is used to acquire the reference speech loudness of the speech data;
[0151] The sixth processing unit is used to determine the speech loudness gain of the speech data based on the difference between the reference speech loudness and the actual speech loudness.
[0152] The seventh processing unit is used to adjust the playback loudness of the media data based on the speech loudness gain and dynamic range control parameters to obtain the target media data for playback.
[0153] In some alternative implementations, after acquiring the media data, the device further includes:
[0154] The second acquisition module is used to acquire the historical playback configuration information of the playback device, which is the device used to play the target media data.
[0155] The third processing module is used to determine the target loudness equalization mode based on the analysis results of historical playback configuration information;
[0156] The fourth processing module is used to determine whether the media data includes voice data in response to the target loudness equalization mode being voice equalization mode.
[0157] In some alternative implementations, the first processing module 702 includes:
[0158] The statistics module is used to determine the first duration of the media data and the second duration of the voice data, respectively.
[0159] The fifth processing module is used to determine the speech loudness metadata of the speech data based on the speech loudness distribution result corresponding to the speech data if the ratio between the second duration and the first duration is greater than a preset threshold.
[0160] In some alternative embodiments, the apparatus further includes:
[0161] The sixth processing module is used to obtain the overall loudness distribution result based on the overall loudness distribution of the media data if the ratio between the second duration and the first duration is less than or equal to a preset threshold.
[0162] The seventh processing module is used to adjust the playback loudness of the media data based on the overall loudness metadata corresponding to the overall loudness distribution results, so as to obtain the target media data for playback.
[0163] Further functional descriptions of the above modules and units are the same as those in the corresponding embodiments described above, and will not be repeated here.
[0164] In this embodiment, the media data playback loudness processing device is presented in the form of a functional unit. Here, a unit refers to an ASIC (Application Specific Integrated Circuit) circuit, a processor and memory that execute one or more software or fixed programs, and / or other devices that can provide the above functions.
[0165] This disclosure also provides an electronic device having the above-described features. Figure 7 The device shown is a loudness processing device for playing media data.
[0166] Please see Figure 8 , Figure 8 This is a schematic diagram of the structure of an electronic device provided in an optional embodiment of this disclosure, such as... Figure 8 As shown, the electronic device includes one or more processors 10, memory 20, and interfaces for connecting the components, including high-speed interfaces and low-speed interfaces. The components communicate with each other via different buses and can be mounted on a common motherboard or otherwise as required. The processors can process instructions executed within the electronic device, including instructions stored in or on memory to display graphical information of a GUI on external input / output devices (such as display devices coupled to the interfaces). In some alternative implementations, multiple processors and / or multiple buses can be used with multiple memories and multiple memory modules, if desired. Similarly, multiple electronic devices can be connected, each providing some of the necessary operations (e.g., as a server array, a group of blade servers, or a multiprocessor system). Figure 8 Take a processor 10 as an example.
[0167] Processor 10 may be a central processing unit, a network processor, or a combination thereof. Processor 10 may further include a hardware chip. The hardware chip may be an application-specific integrated circuit (ASIC), a programmable logic device (PLD), or a combination thereof. The programmable logic device may be a complex programmable logic device (CAMP), a field-programmable gate array (FPGA), a general-purpose array logic (GPA), or any combination thereof.
[0168] The memory 20 stores instructions executable by at least one processor 10 to cause at least one processor 10 to perform the method shown in the above embodiments.
[0169] The memory 20 may include a program storage area and a data storage area. The program storage area may store the operating system and applications required for at least one function; the data storage area may store data created based on the use of the electronic device. Furthermore, the memory 20 may include high-speed random access memory and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some alternative embodiments, the memory 20 may optionally include memory remotely located relative to the processor 10, and these remote memories may be connected to the electronic device via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.
[0170] The memory 20 may include volatile memory, such as random access memory; the memory may also include non-volatile memory, such as flash memory, hard disk or solid-state drive; the memory 20 may also include a combination of the above types of memory.
[0171] The electronic device also includes an input device 30 and an output device 40. The processor 10, memory 20, input device 30, and output device 40 can be connected via a bus or other means. Figure 8 Taking the example of a connection between China and Israel via a bus.
[0172] Input device 30 can receive input numerical or character information, and generate key signal inputs related to user settings and function control of the electronic device, such as a touch screen, keypad, mouse, trackpad, touchpad, joystick, one or more mouse buttons, trackball, joystick, etc. Output device 40 may include display devices, auxiliary lighting devices (e.g., LEDs), and haptic feedback devices (e.g., vibration motors). The aforementioned display devices include, but are not limited to, liquid crystal displays, light-emitting diodes, displays, and plasma displays. In some alternative embodiments, the display device may be a touch screen.
[0173] This disclosure also provides a computer-readable storage medium in which the methods described in this disclosure can be implemented in hardware or firmware, or implemented as recordable on a storage medium, or implemented as computer code originally stored on a remote storage medium or a non-transitory machine-readable storage medium and subsequently stored on a local storage medium after being downloaded over a network. Thus, the methods described herein can be processed by software stored on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. The storage medium may be a magnetic disk, optical disk, read-only memory, random access memory, flash memory, hard disk, or solid-state drive, etc.; further, the storage medium may also include combinations of the above types of memory. It is understood that computers, processors, microprocessor controllers, or programmable hardware include storage components capable of storing or receiving software or computer code that, when accessed and executed by the computer, processor, or hardware, implements the methods shown in the above embodiments.
[0174] A portion of this invention can be applied as a computer program product, such as computer program instructions, which, when executed by a computer, can invoke or provide the methods and / or technical solutions according to the invention through the operation of the computer. Those skilled in the art will understand that the forms in which computer program instructions exist in a computer-readable medium include, but are not limited to, source files, executable files, installation package files, etc. Correspondingly, the ways in which computer program instructions are executed by a computer include, but are not limited to: the computer directly executing the instructions, or the computer compiling the instructions and then executing the corresponding compiled program, or the computer reading and executing the instructions, or the computer reading and installing the instructions and then executing the corresponding installed program. Here, the computer-readable medium can be any available computer-readable storage medium or communication medium accessible to a computer.
[0175] It is understood that before using the technical solutions disclosed in the various embodiments of this disclosure, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this disclosure in an appropriate manner in accordance with relevant laws and regulations, and user authorization should be obtained.
[0176] For example, upon receiving a user's active request, a prompt message is sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the software or hardware, such as the electronic device, application, server, or storage medium performing the operations of this disclosed technical solution, based on the prompt message.
[0177] As an optional but non-limiting implementation, in response to a user's active request, sending a prompt message to the user can be done via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose "agree" or "disagree" to provide personal information to the electronic device.
[0178] It is understood that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of this disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of this disclosure.
[0179] Although embodiments of the present disclosure have been described in conjunction with the accompanying drawings, those skilled in the art can make various modifications and variations without departing from the spirit and scope of the present disclosure, and such modifications and variations all fall within the scope defined by the appended claims.
Claims
1. A method for processing the playback loudness of media data, characterized in that, The method includes: Acquire media data; In response to the inclusion of voice data in the media data, the voice loudness metadata of the voice data is determined based on the voice loudness distribution result corresponding to the voice data; Based on the voice loudness metadata, the playback loudness of the media data is adjusted to obtain the target media data for playback.
2. The method according to claim 1, characterized in that, The process of determining the loudness distribution result corresponding to the speech data includes: Perform speech detection on the media data to determine the media data segments that correspond to the speech data; Determine the loudness distribution of the media data segment to obtain the loudness distribution result; Based on the loudness distribution results, the loudness distribution results corresponding to the speech data are determined.
3. The method according to claim 2, characterized in that, If there are multiple media data segments, then determining the speech loudness distribution result corresponding to the speech data based on the loudness distribution result includes: According to the sequential order of the multiple media data segments, the current loudness distribution result is fused with the previous loudness distribution result in turn, and the fused result is used as the previous loudness distribution result to be fused for the next loudness distribution result, so as to obtain the speech loudness distribution result corresponding to the speech data.
4. The method according to claim 1, characterized in that, The step of adjusting the playback loudness of the media data based on the voice loudness metadata to obtain target media data for playback includes: Based on the overall loudness distribution of the media data, the overall loudness distribution result is obtained; Based on the overall loudness distribution results, overall loudness metadata is determined to obtain the overall loudness of the media data; The actual loudness of the speech data is determined using the speech loudness metadata. The target loudness-to-dialogue ratio of the media data is determined based on the difference between the overall loudness and the actual speech loudness. Based on the target loudness-to-dialogue ratio, the playback loudness of the media data is adjusted to obtain the target media data for playback.
5. The method according to claim 4, characterized in that, The step of adjusting the playback loudness of the media data based on the target loudness-to-dialogue ratio to obtain target media data for playback includes: Obtain the target playback loudness of the voice data; Based on the target playback loudness, the actual voice loudness, and the target loudness-to-dialogue ratio, determine the dynamic range control parameters of the media data; The playback loudness of the media data is adjusted based on the dynamic range control parameters to obtain the target media data for playback.
6. The method according to claim 5, characterized in that, The dynamic range control parameters include the dynamic range compression ratio. Determining the dynamic range control parameters of the media data based on the target playback loudness, the actual speech loudness, and the target loudness-to-dialogue ratio includes: The first loudness compression ratio is determined based on the ratio between the actual speech loudness and the target playback loudness. The second loudness compression ratio is determined based on the ratio between the target loudness-to-dialogue ratio and the specified loudness-to-dialogue ratio; Based on the comparison between the first loudness compression ratio and the second loudness compression ratio, the dynamic range compression ratio is determined.
7. The method according to claim 6, characterized in that, The step of determining the dynamic range compression ratio based on the comparison result between the first loudness compression ratio and the second loudness compression ratio includes: The loudness compression ratio, which is the larger of the first loudness compression ratio and the second loudness compression ratio, is taken as the compression ratio of the dynamic range.
8. The method according to claim 6 or 7, characterized in that, The dynamic range control parameters also include a static characteristic threshold. The determination of the dynamic range control parameters for the media data based on the target playback loudness, the actual speech loudness, and the target loudness-to-dialogue ratio further includes: The starting loudness of the actual speech loudness is determined using the speech loudness metadata. The initial loudness is used as the static characteristic threshold.
9. The method according to claim 5, characterized in that, The step of adjusting the playback loudness of the media data based on the target loudness-to-dialogue ratio to obtain target media data for playback further includes: Obtain the reference loudness of the speech data; The speech loudness gain of the speech data is determined based on the difference between the reference speech loudness and the actual speech loudness. Based on the speech loudness gain and the dynamic range control parameters, the playback loudness of the media data is adjusted to obtain the target media data for playback.
10. The method according to claim 1, characterized in that, After acquiring the media data, the method further includes: Obtain historical playback configuration information of the playback device, wherein the playback device is a device used to play the target media data; Based on the analysis results of the historical playback configuration information, the target loudness equalization mode is determined; In response to the target loudness equalization mode being a speech equalization mode, a process is performed to determine whether the media data includes the speech data.
11. The method according to claim 1, characterized in that, The step of determining the speech loudness metadata of the speech data based on the speech loudness distribution result corresponding to the speech data includes: The first duration of the media data and the second duration of the voice data are determined respectively; If the ratio between the second duration and the first duration is greater than a preset threshold, then the voice loudness metadata of the voice data is determined based on the voice loudness distribution result corresponding to the voice data.
12. The method according to claim 11, characterized in that, The method further includes: If the ratio between the second duration and the first duration is less than or equal to the preset threshold, then the overall loudness distribution result is obtained based on the overall loudness distribution of the media data; Based on the overall loudness metadata corresponding to the overall loudness distribution result, the playback loudness of the media data is adjusted to obtain the target media data for playback.
13. A media data playback loudness processing device, characterized in that, The device includes: The first acquisition module is used to acquire media data; The first processing module is configured to, in response to the inclusion of voice data in the media data, determine the voice loudness metadata of the voice data based on the voice loudness distribution result corresponding to the voice data; The second processing module is used to adjust the playback loudness of the media data based on the voice loudness metadata to obtain the target media data for playback.
14. An electronic device, characterized in that, include: A memory and a processor are communicatively connected, the memory stores computer instructions, and the processor executes the computer instructions to perform the playback loudness processing method for media data according to any one of claims 1 to 12.
15. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions for causing the computer to execute the playback loudness processing method for media data according to any one of claims 1 to 12.
16. A computer program product, characterized in that, Includes computer instructions for causing a computer to perform the playback loudness processing method for media data according to any one of claims 1 to 12.