Loudness processing method and device for media data, electronic equipment and storage medium
By obtaining historical media data of preset duration for loudness analysis, the problem of uneven loudness during media data playback is solved, and the timely update and balanced processing of loudness is achieved.
Patent Information
- Application Number
- CN202410153837.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-02-02
- Publication Date
- 2025-08-05
AI Technical Summary
During the media data playback, it is difficult for the prior art to achieve balanced processing of loudness, especially in live broadcast scenarios. As time goes by, loudness updates become less and less sensitive, resulting in the problem of loudness fluctuations.
By acquiring historical media data of preset duration for loudness analysis, the loudness metadata is obtained, and the loudness processing of the media data to be processed based on this, including loudness gain, dynamic compression and peak limit, ensuring the balance of loudness.
It improves the determination efficiency of loudness metadata, ensures the speed of loudness update, avoids the problem of loudness fluctuations, and realizes the balanced processing of loudness.
Smart Images

Figure CN120434451A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer technology, and in particular to a method, device, electronic device, and storage medium for processing loudness of media data. Background Art
[0002] During the media data playback process, in order to ensure a good playback effect, it is usually necessary to involve equalization processing of the playback loudness. Summary of the Invention
[0003] In view of this, the present disclosure provides a method, apparatus, electronic device, and storage medium for processing loudness of media data to solve the problem of balanced playback loudness.
[0004] In a first aspect, the present disclosure provides a method for processing loudness of media data, the method comprising:
[0005] Acquire the media data to be processed and historical media data of a preset duration, where the end position of the historical media data corresponds to the start position of the media data to be processed, and the preset duration is less than the cumulative duration of all processed media data;
[0006] performing loudness analysis based on the historical media data to obtain loudness metadata;
[0007] Loudness processing is performed on the media data to be processed based on the loudness metadata, and target media data is determined and played.
[0008] In a second aspect, the present disclosure provides a device for processing loudness of media data, the device comprising:
[0009] a media data acquisition module configured to acquire the media data to be processed and historical media data of a preset duration, wherein the end position of the historical media data corresponds to the start position of the media data to be processed, and the preset duration is less than the cumulative duration of all processed media data;
[0010] a loudness analysis module, configured to perform loudness analysis based on the historical media data to obtain loudness metadata;
[0011] The loudness processing module is configured to perform loudness processing on the media data to be processed based on the loudness metadata, and determine and play target media data.
[0012] In a third aspect, the present disclosure provides an electronic device comprising: a memory and a processor, the memory and the processor being communicatively connected to each other, the memory storing computer instructions, and the processor executing the computer instructions to thereby perform the method for loudness processing of media data according to the first aspect or any corresponding embodiment thereof.
[0013] In a fourth aspect, the present disclosure provides a computer-readable storage medium having computer instructions stored thereon, the computer instructions being used to cause a computer to execute the loudness processing method for media data according to the first aspect or any corresponding embodiment thereof.
[0014] The loudness processing method for media data provided by the disclosed embodiments obtains the media data to be processed and historical media data of a preset duration. The obtained historical media data is temporally adjacent to the media data to be processed. For media data loudness processing, the loudness of media data older than the current one has less impact on the current loudness. Therefore, loudness analysis is performed on the historical media data of the preset duration, rather than on all processed media data, to obtain loudness metadata. This method improves the efficiency of determining loudness metadata and ensures fast loudness updates. Loudness processing of the media data to be processed based on the obtained loudness metadata allows for timely loudness processing, thereby ensuring loudness balance. After media data has been played for a long time, loudness updates based on processed media data can become less sensitive due to the large amount of data processing, resulting in fluctuating loudness. Therefore, this method avoids this problem and effectively ensures loudness balance. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] In order to more clearly illustrate the specific embodiments of the present disclosure or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the specific embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present disclosure. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0016] Figure 1 is a flowchart of a method for processing loudness of media data according to an embodiment of the present disclosure;
[0017] Figure 2 Schematic diagram of a method for extracting loudness metadata in related art;
[0018] Figure 3 is a flowchart of another method for processing loudness of media data according to an embodiment of the present disclosure;
[0019] Figure 4 is a flowchart of another method for processing loudness of media data according to an embodiment of the present disclosure;
[0020] Figure 5 is a structural block diagram of a device for processing loudness of media data according to an embodiment of the present disclosure;
[0021] Figure 6Schematic diagram of the hardware structure of the electronic device according to the embodiment of the present disclosure. DETAILED DESCRIPTION
[0022] To make the purpose, technical solutions, and advantages of the embodiments of the present disclosure more clear, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the drawings in the embodiments of the present disclosure. Obviously, the described embodiments are part of the embodiments of the present disclosure, not all of the embodiments. Based on the embodiments of the present disclosure, all other embodiments obtained by those skilled in the art without making creative efforts shall fall within the scope of protection of the present disclosure.
[0023] It is understandable that before using the technical solutions disclosed in the various embodiments of this disclosure, the type, scope of use, usage scenarios, etc. of the personal information involved in this disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.
[0024] For example, in response to a user's active request, a prompt message is sent to the user to clearly inform the user that the operation requested will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the electronic device, application, server, storage medium, or other software or hardware that performs the operations of the disclosed technical solution based on the prompt message.
[0025] As an optional but non-limiting implementation, in response to receiving a user's active request, the prompt information may be sent to the user in the form of a pop-up window, in which the prompt information may be presented in text form. Furthermore, the pop-up window may also contain a selection control for the user to select "agree" or "disagree" to provide personal information to the electronic device.
[0026] It is understandable that the above notification and user authorization process are merely illustrative and do not limit the implementation of the present disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of the present disclosure.
[0027] It is understandable that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) must comply with the requirements of relevant laws, regulations and relevant provisions.
[0028] In related technologies, loudness processing of media data in live broadcast scenarios typically involves calculating loudness metadata. Loudness metadata includes short-term loudness, instantaneous loudness, and statistical loudness. Statistical loudness is calculated by analyzing the loudness of all media data captured during the live broadcast. However, as the live broadcast progresses, the volume of media data used for statistical loudness analysis increases, leading to less sensitive loudness updates and, consequently, less sensitive loudness gain. This can lead to fluctuating loudness during the live broadcast.
[0029] Based on this, the loudness processing method for media data provided in the embodiments of the present disclosure calculates loudness metadata based on a preset duration of historical media data after acquiring the media data to be processed. The end position of the historical media data corresponds to the start position of the media data to be processed. Loudness processing is then performed on the media data to be processed, combining the calculated loudness metadata to achieve loudness equalization.
[0030] Furthermore, in the loudness processing method provided in the related art, when a live broadcast just begins, the short-term loudness may not be available. At this time, the volume of the media data will be processed using the maximum loudness, which will increase the volume of the media data and amplify the background noise.
[0031] Based on this, the loudness processing method for media data provided by the embodiments of the present disclosure calculates the difference between the short-term loudness and the instantaneous loudness and the target loudness to obtain the corresponding compression ratio. The maximum of the two compression ratios is used as the current dynamic compression ratio to avoid amplifying the background noise.
[0032] It should be noted that the loudness processing method for media data provided in the embodiments of the present disclosure can be applied not only to live broadcast scenarios but also to on-demand scenarios, and no limitation is imposed on its application scenarios.
[0033] According to an embodiment of the present disclosure, an embodiment of a method for processing loudness of media data is provided. It should be noted that the steps shown in the flowcharts of the accompanying drawings can be executed in a computer system, such as a set of computer-executable instructions. Moreover, although a logical order is shown in the flowcharts, in some cases, the steps shown or described may be executed in a different order than shown.
[0034] In this embodiment, a method for processing the loudness of media data is provided, which can be used in electronic devices such as mobile phones and tablet computers. Figure 1 FIG. 1 is a flow chart of a method for processing loudness of media data according to an embodiment of the present disclosure. Figure 1 As shown, the process includes the following steps:
[0035] Step S101: Acquire the media data to be processed and historical media data of a preset duration.
[0036] The end position of the historical media data corresponds to the start position of the media data to be processed, and the preset duration is less than the cumulative duration of all processed media data.
[0037] The media data to be processed is the media data currently requiring loudness processing. The media data may be in the form of video data or audio data, with no limitation to the specific form. The duration of the media data to be processed can be set based on actual needs, for example, 100 ms. This means that loudness processing is performed every 100 ms.
[0038] The preset duration of historical media data is the processed media data, and the end position of the historical media data corresponds to the start position of the media data to be processed. For example, in a live broadcast scenario, the start position of the media data to be processed is 10 seconds after the live broadcast begins. Accordingly, the preset duration of historical media data is 10 seconds after the live broadcast begins.
[0039] The preset duration is less than the cumulative duration of all processed media data. If the cumulative duration of processed media data is 10S and the duration of the media data to be processed is 100ms, the preset duration of the historical media data can be 30ms, and so on.
[0040] It should be noted that, while obtaining the media data to be processed, loudness-related data delivered along with the media data to be processed can also be obtained. The loudness-related data includes, but is not limited to, loudness processing types, such as the default loudness mode and the anchor loudness mode. The default loudness mode processes loudness gains based on statistical loudness, while the anchor loudness mode processes loudness gains based on anchor loudness.
[0041] Step S102: loudness analysis is performed based on historical media data to obtain loudness metadata.
[0042] In related art, Figure 2 This illustrates how loudness metadata is processed. If loudness processing is performed on the media data to be processed every xms (e.g., 100ms), the loudness metadata for these 100ms of media data is calculated on the broadcast side or the content distribution network corresponding to the broadcast side and distributed along with the processed media data to the viewer side. Accordingly, the viewer side uses the obtained loudness metadata to perform loudness processing on the media data to be processed. In this case, the short-term loudness and instantaneous loudness included in the loudness metadata are related to the media data to be processed, while the statistical loudness is related to all media data in the live broadcast. However, as shown above, this processing method can result in untimely loudness updates.
[0043] In this embodiment, the loudness metadata is extracted based on historical media data. Since the historical media data corresponding to the loudness metadata in this embodiment is different from the extraction object in the related art, the obtained loudness metadata is different.
[0044] Specifically, in this embodiment, the loudness metadata of the preset duration is adjacent to the media data to be processed. For example, the media data to be processed is the media data from 100ms to 200ms after the start of the live broadcast, and the preset duration is 30ms. Based on this, the historical media data of the preset duration is the media data from 70ms to 100ms after the start of the live broadcast.
[0045] If the to-be-processed media data is acquired at preset time intervals, and historical media data of a preset duration is adjacent to the to-be-processed media data, then the historical media data of the preset duration is also acquired at preset time intervals. If the currently acquired historical media data of the preset duration is media data from 70ms to 100ms after the start of the live broadcast, then the next to-be-processed media data acquired is media data from 200ms to 300ms after the start of the live broadcast. Accordingly, the acquired historical media data of the preset duration is media data from 170ms to 200ms after the start of the live broadcast.
[0046] After acquiring historical media data for a preset duration, loudness calculation is performed based on the historical media data to obtain loudness metadata. Loudness metadata includes at least one of short-term loudness, instantaneous loudness, and statistical loudness. The specific content included is not limited herein and can be set based on actual needs.
[0047] Step S103 : loudness processing is performed on the media data to be processed based on the loudness metadata, and target media data is determined and played.
[0048] After step S102, loudness metadata corresponding to the media data to be processed is determined. Loudness processing of the media data to be processed is based on the loudness metadata obtained in step S102. Loudness processing includes loudness gain, dynamic compression, and peak limiting.
[0049] Specifically, the loudness gain is obtained by combining the difference between the currently calculated loudness and the target loudness. Dynamic compression parameters include dynamic compression ratio and threshold, etc. Peak limiting processes media data based on the peak threshold.
[0050] After loudness processing is performed on the media data to be processed based on the loudness metadata, target media data to be played is obtained, and the target media data is played.
[0051] The loudness processing method for media data provided in this embodiment obtains the media data to be processed and historical media data of a preset duration. The acquired historical media data is temporally adjacent to the media data to be processed. For media data loudness processing, the loudness of media data older than the current one has less impact on the current loudness. Therefore, loudness analysis is performed on the preset duration of historical media data, rather than all processed media data, to obtain loudness metadata. This approach improves the efficiency of determining loudness metadata and ensures rapid loudness updates. Loudness processing of the media data to be processed based on the acquired loudness metadata allows for timely loudness processing, thereby ensuring loudness balance. After media data has been played for a long time, loudness updates based on processed media data can become less sensitive due to the large amount of data processing, resulting in fluctuating loudness. Therefore, this approach avoids this problem and effectively ensures loudness balance.
[0052] In this embodiment, a method for processing the loudness of media data is provided, which can be used in electronic devices such as mobile phones and tablet computers. Figure 3 FIG. 1 is a flow chart of a method for processing loudness of media data according to an embodiment of the present disclosure. Figure 3 As shown, the process includes the following steps:
[0053] Step S301: Acquire the media data to be processed and historical media data of a preset duration.
[0054] The end position of the historical media data corresponds to the start position of the media data to be processed, and the preset duration is less than the cumulative duration of all processed media data. Figure 2 The detailed description of step S201 in the illustrated embodiment is omitted here.
[0055] Step S302: Perform loudness analysis based on historical media data to obtain loudness metadata. Figure 2 The detailed description of step S202 in the illustrated embodiment is omitted here.
[0056] Step S303: loudness processing is performed on the media data to be processed based on the loudness metadata, and target media data is determined and played.
[0057] Specifically, the above step S303 includes:
[0058] Step S3031: Obtain target loudness.
[0059] The target loudness may be set based on actual needs, or may be adaptively set based on the current playback environment and the performance of the playback device, etc., and is not limited here.
[0060] Step S3032: Determine the current loudness gain and the current dynamic compression parameter corresponding to the media data to be processed based on the loudness metadata and the target loudness.
[0061] Loudness metadata includes various types of loudness data, such as short-term loudness, instantaneous loudness, and statistical loudness. Based on the difference between different types of loudness data and the target loudness, the current loudness gain and current dynamic compression parameters corresponding to the nursing media data to be tested are obtained.
[0062] The current loudness gain is used to represent the difference between the current loudness and the target loudness. The current dynamic compression parameters include the dynamic compression ratio, the dynamic compression threshold, and the like.
[0063] In some optional implementations, determining the current loudness gain corresponding to the to-be-processed media data based on the loudness metadata and the target loudness in step S3032 includes:
[0064] Step a1: Obtain statistical loudness in loudness metadata.
[0065] Step a2: obtaining the current loudness gain based on the difference between the target loudness and the statistical loudness.
[0066] Statistical loudness is derived from historical media data over a preset duration. If the default loudness processing mode is used, the current loudness gain is calculated based on the difference between the target loudness and the statistical loudness. That is, the current loudness gain is GAIN = target loudness - statistical loudness.
[0067] Since the statistical loudness is obtained based on historical media data of a preset duration, the amount of calculation is small, so that the loudness gain can be updated in a timely manner to obtain the current loudness gain suitable for the current playback.
[0068] In some optional implementations, determining the current loudness gain corresponding to the to-be-processed media data based on the loudness metadata and the target loudness in step S3032 includes:
[0069] Step b1: Obtain anchor loudness in loudness metadata.
[0070] Step b2: Obtain the current loudness gain based on the difference between the target loudness and the anchor loudness.
[0071] Anchor loudness is determined based on the proportion of speech in historical media data. The specific calculation method depends on actual needs. It is sufficient to ensure that the anchor loudness is combined with the proportion of speech in historical media data when determining it.
[0072] If the loudness processing mode sent is the anchor loudness mode (anchor), the current loudness gain is obtained based on the difference between the target loudness and the anchor loudness, that is, the current loudness gain GAIN = target loudness - anchor loudness.
[0073] The anchor loudness is obtained based on the speech ratio, and thus can be determined based on the speech conditions in historical media data, with high accuracy.
[0074] In some optional implementations, determining the current dynamic compression parameter corresponding to the to-be-processed media data based on the loudness metadata and the target loudness in step S3032 includes:
[0075] Step c1: Obtain short-term loudness and instantaneous loudness in loudness metadata.
[0076] Step c2: determining a first difference between the short-term loudness and the target loudness, and a second difference between the instantaneous loudness and the target loudness.
[0077] Step c3: obtaining a first compression ratio based on a ratio of the first difference to a first preset value.
[0078] Step c4: obtaining a second compression ratio based on a ratio of the second difference to a second preset value, where the second preset value is greater than the first preset value.
[0079] Step c5: determining the maximum value of the first compression ratio and the second compression ratio as the current dynamic compression ratio, where the dynamic compression parameters include the current dynamic compression ratio.
[0080] Short-term loudness and momentary loudness are extracted from the loudness metadata. A first difference between the short-term loudness and the target loudness, and a second difference between the momentary loudness and the target loudness, are then calculated. Based on this, a first compression ratio = (short-term loudness - target loudness) / a first preset value, and a second compression ratio = (momentary loudness - target loudness) / a second preset value. The second preset value is greater than the first preset value. For example, the first preset value is 5, and the second preset value is 6.
[0081] The first compression ratio and the second compression ratio are compared to obtain a maximum value between the two, and the maximum value is determined as the current dynamic compression ratio. The dynamic compression parameter includes the current dynamic compression ratio.
[0082] In live broadcast scenarios, the short-term loudness has not been updated in the first 3 seconds of the live broadcast. At this time, when the maximum instantaneous loudness is updated, the maximum instantaneous loudness is used to calculate the second compression ratio. The compression degree is smaller than the short-term loudness to prevent the problem of excessive volume after the initial loudness processing.
[0083] Short-term loudness requires a certain amount of media data to be accumulated. This may be difficult to determine at the beginning of media playback. In this case, dynamic compression parameters can be adjusted based on instantaneous loudness. Because the second preset value is greater than the first, the compression level corresponding to the second compression ratio is lower than that corresponding to the first compression ratio. This prevents the initial loudness processing from causing excessive volume, specifically, excessive background noise.
[0084] In some optional implementations, determining the current dynamic compression parameter corresponding to the to-be-processed media data based on the loudness metadata and the target loudness in step S3032 further includes: determining a difference between the target loudness and a third preset value as a threshold value, the current dynamic compression parameter includes the threshold value, and the third preset value is greater than the second preset value.
[0085] For example, the third preset value is 20, and accordingly, the threshold value = target loudness - 20. Based on the calculation result of the threshold value, it is determined as the threshold value. The current dynamic compression parameter includes the threshold value.
[0086] Step S3033: loudness processing is performed on the media data to be processed based on the current loudness gain and the current dynamic compression parameter, and target media data is determined and played.
[0087] After obtaining the current loudness gain, the media data to be processed is processed using the loudness gain to obtain first media data. The first media data is then dynamically compressed using the current dynamic compression parameter to obtain second media data. The second media data is then peak-limited using the peak threshold to obtain target media data. After obtaining the target media data, it is played.
[0088] In some optional implementations, the above step S3033 includes:
[0089] Step d1 : Smoothing the current loudness gain and the current dynamic compression parameter to obtain a target loudness gain and a target dynamic compression parameter.
[0090] Step d2: performing loudness processing on the media data to be processed based on the target loudness gain and the target dynamic compression parameter to obtain target media data.
[0091] Step d3: play the target media data.
[0092] Because the changes between the historical loudness gain and dynamic compression parameters corresponding to the previous media data to be processed and the current loudness gain and dynamic compression parameters corresponding to the media data to be processed are discontinuous, there may be sudden changes. Therefore, to ensure a smooth transition between the loudness gain and dynamic compression parameters, the current loudness gain and dynamic compression parameters are smoothed to obtain the target loudness gain and target dynamic compression parameters.
[0093] Then, loudness processing is performed on the media data to be processed based on the target loudness gain and the target dynamic compression parameter to obtain target media data and play the target media data.
[0094] Since the loudness gains and dynamic compression parameters are discrete, loudness equalization can be better achieved through smoothing processing.
[0095] The processing methods for the loudness gain and dynamic compression parameters are similar. Here, taking the smoothing of the loudness gain as an example, the smoothing of the current loudness gain is described in detail to obtain the target loudness gain. That is, in some optional embodiments, the smoothing of the current loudness gain in step d1 to obtain the target loudness gain includes:
[0096] Step d11: Obtain historical loudness gains corresponding to historical media data and a preset number of sampling points corresponding to the media data to be processed.
[0097] Step d12: Determine the difference between the current loudness gain and the historical loudness gain.
[0098] Step d13: obtaining a loudness gain adjustment step size based on a ratio of the difference value to a preset number of sampling points.
[0099] Step d14: Obtain a target loudness gain based on the historical loudness gain and the loudness gain adjustment step size.
[0100] The historical loudness gain corresponding to the previous media data to be processed refers to the loudness gain used when performing loudness gain processing on the previous media data to be processed, and is referred to as the historical loudness gain. The preset number of sampling points corresponding to the media data to be processed refers to the number of sampling points used to sample the media data to be processed. Specifically, when performing loudness processing on the media data to be processed, the sampled media data is processed sequentially. Based on this, the media data to be processed is sampled based on the preset number of sampling points to obtain a sampling result, which is then subjected to loudness gain processing based on loudness gain adjustment compensation, thereby achieving a smooth transition in loudness.
[0101] Specifically, the difference between the current loudness gain and the historical loudness gain is calculated to obtain the loudness gain adjustment amount. The loudness gain adjustment step size is then determined based on the ratio of the loudness gain adjustment amount to the preset number of sampling points. In other words, the loudness gain adjustment amount for each loudness gain adjustment is determined based on the historical loudness gain and the loudness gain adjustment step size. This allows for a smooth transition from the historical loudness gain to the current loudness gain. The loudness gain obtained for each loudness gain adjustment of the sampled data is referred to as the target loudness gain. Smoothing is performed based on the preset number of sampling points to achieve gradual adjustment of the loudness gain.
[0102] The loudness processing method for media data provided in this embodiment determines the current loudness gain corresponding to the media data to be processed through loudness metadata, and also obtains the current dynamic compression parameter, thereby achieving adaptive adjustment of the loudness gain and dynamic compression parameter, thereby balancing loudness and dynamics, achieving maximum loudness while preserving dynamics to the greatest extent possible.
[0103] In this embodiment, a method for processing the loudness of media data is provided, which can be used in electronic devices such as mobile phones and tablet computers. Figure 4 FIG. 1 is a flow chart of a method for processing loudness of media data according to an embodiment of the present disclosure. Figure 4 As shown, the process includes the following steps:
[0104] Step S401: Acquire the media data to be processed and historical media data of a preset duration.
[0105] The end position of the historical media data corresponds to the start position of the media data to be processed, and the preset duration is less than the cumulative duration of all processed media data.
[0106] Specifically, the above step S401 includes:
[0107] Step S4011, obtaining the window length and preset step length of the sliding window.
[0108] The window length is the preset duration.
[0109] The acquisition of historical media data of a preset duration is achieved through a sliding window, wherein the window length of the sliding window is the preset duration, and the preset step length of the sliding window is the time interval between two consecutive acquisitions of the media data to be processed.
[0110] It should be noted that the time interval between two adjacent acquisitions of the to-be-processed media data may be customized or determined based on the length of the media data frame, etc., and is not limited here.
[0111] Step S4012: extract the acquired media data based on the window length of the sliding window and the preset step length to obtain historical media data.
[0112] Once acquired media data is processed, it can be stored sequentially based on the time it was acquired. A sliding window is then used to extract data from the stored media data to obtain historical media data of a preset duration. When the next piece of media data to be processed is acquired, the sliding window is moved according to a preset step size to obtain historical media data of the preset duration corresponding to the next piece of media data to be processed.
[0113] For example, the preset step size is 100ms and the preset duration is 30ms. The acquired media data to be processed is the media data from 100ms to 200ms of the live broadcast. The sliding window corresponds to the media data from 70ms to 100ms of the live broadcast, and the corresponding historical media data is obtained. After acquiring the media data from 200ms to 300ms of the live broadcast, the sliding window is controlled to move backward 100ms to obtain the media data from 170ms to 200ms of the live broadcast, and the corresponding historical media data is obtained.
[0114] Step S402: Perform loudness analysis based on historical media data to obtain loudness metadata. Figure 1 Step S202 of the illustrated embodiment will not be described in detail here.
[0115] Step S403: loudness processing is performed on the media data to be processed based on the loudness metadata, and the target media data is determined and played. Figure 1 Step S203 of the illustrated embodiment will not be described in detail here.
[0116] The loudness processing method for media data provided in this embodiment realizes the extraction of historical media data of a preset duration by means of a sliding window, thereby ensuring the extraction efficiency of historical media data of the preset duration.
[0117] As a specific application embodiment of the embodiment of the present disclosure, take the live broadcast scene as an example. The stream is pushed in real time on the anchor side, and the stream is pulled in real time on the viewer side. The media data generated by the anchor side is pushed to the viewer side in real time, and the viewer side performs loudness equalization processing on the acquired media data to be processed. Specifically, the viewer side processes 100ms of media data to be processed each time, obtains 30ms of historical media data from the acquired 100ms of media data to be processed through a sliding window, and uses this 30ms of historical media data to calculate the loudness metadata, and obtains the loudness metadata corresponding to each 100ms of media data to be processed. After obtaining the loudness metadata, loudness processing can be performed on each 100ms of media data to be processed, including loudness gain, dynamic compression and peak limiting. After the loudness processing, the target media data for playback is obtained, and then the target media data is played.
[0118] This embodiment also provides a device for processing the loudness of media data. This device is used to implement the aforementioned embodiments and preferred implementations, and details already described will not be repeated. As used below, the term "module" may refer to a combination of software and / or hardware that implements a predetermined function. Although the devices described in the following embodiments are preferably implemented using software, implementation using hardware, or a combination of software and hardware, is also possible and contemplated.
[0119] This embodiment provides a device for processing loudness of media data. Figure 5 As shown, including:
[0120] The media data acquisition module 501 is used to acquire the media data to be processed and historical media data of a preset duration, where the end position of the historical media data corresponds to the start position of the media data to be processed, and the preset duration is less than the cumulative duration of all processed media data.
[0121] The loudness analysis module 502 is configured to perform loudness analysis based on historical media data to obtain loudness metadata.
[0122] The loudness processing module 503 is configured to perform loudness processing on the media data to be processed based on the loudness metadata, and determine and play target media data.
[0123] In some optional implementations, the loudness processing module 503 includes:
[0124] The target loudness obtaining unit is configured to obtain the target loudness.
[0125] The current parameter determination unit is configured to determine a current loudness gain and a current dynamic compression parameter corresponding to the media data to be processed based on the loudness metadata and the target loudness.
[0126] The loudness processing unit is configured to perform loudness processing on the media data to be processed based on the current loudness gain and the current dynamic compression parameter, and determine and play the target media data.
[0127] In some optional implementations, the current parameter determination unit includes:
[0128] The statistical loudness acquisition subunit is used to obtain the statistical loudness in the loudness metadata.
[0129] The first current loudness gain determining subunit is configured to obtain a current loudness gain based on a difference between the target loudness and the statistical loudness.
[0130] In some optional implementations, the current parameter determination unit includes:
[0131] The anchor loudness acquisition subunit is used to obtain the anchor loudness in the loudness metadata. The anchor loudness is determined based on the speech ratio in the historical media data.
[0132] The second current loudness gain determination subunit is configured to obtain a current loudness gain based on a difference between the target loudness and the anchor loudness.
[0133] In some optional implementations, the current parameter determination unit includes:
[0134] The loudness acquisition subunit is used to obtain the short-term loudness and instantaneous loudness in the loudness metadata.
[0135] The difference determination subunit is configured to determine a first difference between the short-term loudness and the target loudness, and a second difference between the instantaneous loudness and the target loudness.
[0136] The first compression ratio determining subunit is configured to obtain a first compression ratio based on a ratio of the first difference value to a first preset value.
[0137] The second compression ratio determining subunit is configured to obtain a second compression ratio based on a ratio of the second difference to a second preset value, where the second preset value is greater than the first preset value.
[0138] The compression ratio determination subunit is used to determine the maximum value of the first compression ratio and the second compression ratio as the current dynamic compression ratio, and the dynamic compression parameters include the current dynamic compression ratio.
[0139] In some optional implementations, the current parameter determination unit further includes:
[0140] The threshold determination subunit is configured to determine a difference between the target loudness and a third preset value as a threshold, the current dynamic compression parameter includes the threshold, and the third preset value is greater than the second preset value.
[0141] In some optional implementations, the loudness processing unit includes:
[0142] The smoothing processing subunit is used to perform smoothing processing on the current loudness gain and the current dynamic compression parameter respectively to obtain the target loudness gain and the target dynamic compression parameter.
[0143] The loudness processing subunit is configured to perform loudness processing on the media data to be processed based on the target loudness gain and the target dynamic compression parameter to obtain target media data.
[0144] The playback subunit is used to play the target media data.
[0145] In some optional implementations, the smoothing processing subunit includes:
[0146] The sampling point number acquisition subunit is used to obtain the historical loudness gain corresponding to the previous media data to be processed and the preset number of sampling points corresponding to the media data to be processed.
[0147] The difference determination subunit is used to determine the difference between the current loudness gain and the historical loudness gain.
[0148] The adjustment step size determination subunit is used to obtain the loudness gain adjustment step size based on the ratio of the difference value to the preset number of sampling points.
[0149] The target loudness gain determination subunit is configured to obtain a target loudness gain based on a historical loudness gain and a loudness gain adjustment step size.
[0150] In some optional implementations, the media data acquisition module 501 includes:
[0151] The window length obtaining unit is used to obtain the window length and preset step length of the sliding window, where the window length is the preset duration.
[0152] The media data extraction unit is used to extract the acquired media data based on the window length of the sliding window and a preset step length to obtain historical media data. The preset step length corresponds to the time interval between two adjacent acquisitions of the media data to be processed.
[0153] The media data loudness processing device in this embodiment is presented in the form of a functional unit. The unit here refers to an ASIC (Application Specific Integrated Circuit), a processor and memory that executes one or more software or fixed programs, and / or other devices that can provide the above functions.
[0154] The further functional description of each of the above modules and units is the same as that of the above corresponding embodiments and will not be repeated here.
[0155] The present disclosure also provides an electronic device having the above Figure 5 The loudness processing device of media data is shown.
[0156] See also Figure 6 , Figure 6 is a structural diagram of an electronic device provided by an optional embodiment of the present disclosure, such as Figure 6 As shown, the electronic device includes: one or more processors 10, a memory 20, and interfaces for connecting various components, including high-speed interfaces and low-speed interfaces. The various components utilize different buses to communicate with each other and can be installed on a common mainboard or installed in other ways as needed. The processor can process instructions executed in the electronic device, including instructions stored in or on the memory to display graphical information of a GUI on an external input / output device (such as a display device coupled to the interface). In some optional embodiments, if necessary, multiple processors and / or multiple buses can be used together with multiple memories and multiple memories. Similarly, multiple electronic devices can be connected, and each device provides some necessary operations (for example, as a server array, a group of blade servers, or a multi-processor system). Figure 6 A processor 10 is taken as an example.
[0157] The processor 10 may be a central processing unit, a network processor, or a combination thereof. The processor 10 may further include a hardware chip. The hardware chip may be an application-specific integrated circuit, a programmable logic device, or a combination thereof. The programmable logic device may be a complex programmable logic device, a field programmable gate array, a general purpose array logic, or any combination thereof.
[0158] The memory 20 stores instructions that can be executed by at least one processor 10, so that the at least one processor 10 executes the method shown in the above embodiment.
[0159] The memory 20 may include a program storage area and a data storage area, wherein the program storage area may store an operating system and applications required for at least one function; the data storage area may store data created based on the use of the electronic device, etc. In addition, the memory 20 may include a high-speed random access memory, and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some optional embodiments, the memory 20 may optionally include a memory remotely located relative to the processor 10, and these remote memories may be connected to the electronic device via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0160] The memory 20 may include a volatile memory, such as a random access memory; the memory may also include a non-volatile memory, such as a flash memory, a hard disk or a solid-state drive; the memory 20 may also include a combination of the above types of memory.
[0161] The electronic device further includes an input device 30 and an output device 40. The processor 10, the memory 20, the input device 30 and the output device 40 may be connected via a bus or other means. Figure 6 The bus connection is taken as an example.
[0162] The input device 30 can receive input digital or character information and generate key signal input related to user settings and function control of the electronic device, such as a touch screen, a keypad, a mouse, a trackpad, a touch pad, an indicator stick, one or more mouse buttons, a trackball, a joystick, etc. The output device 40 can include a display device, an auxiliary lighting device (e.g., an LED), and a tactile feedback device (e.g., a vibration motor). The above-mentioned display device includes but is not limited to a liquid crystal display, a light emitting diode, a display, and a plasma display. In some optional embodiments, the display device can be a touch screen.
[0163] The embodiments of the present disclosure also provide a computer-readable storage medium. The above-mentioned method according to the embodiments of the present disclosure can be implemented in hardware, firmware, or implemented as a computer code that can be recorded in a storage medium, or implemented as a computer code that is originally stored in a remote storage medium or a non-temporary machine-readable storage medium and downloaded through a network and will be stored in a local storage medium, so that the method described herein can be stored in such software processing on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only storage memory, a random access memory, a flash memory, a hard disk or a solid-state drive, etc.; further, the storage medium can also include a combination of the above-mentioned types of memory. It can be understood that a computer, a processor, a microprocessor controller or programmable hardware includes a storage component that can store or receive software or computer code. When the software or computer code is accessed and executed by a computer, a processor or hardware, the method shown in the above embodiment is implemented.
[0164] Although the embodiments of the present disclosure have been described with reference to the accompanying drawings, those skilled in the art may make various modifications and variations without departing from the spirit and scope of the present disclosure, and such modifications and variations are all within the scope defined by the appended claims.
Claims
1. A method for processing loudness of media data, characterized in that: The method comprises: Acquire the media data to be processed and historical media data of a preset duration, where the end position of the historical media data corresponds to the start position of the media data to be processed, and the preset duration is less than the cumulative duration of all processed media data; performing loudness analysis based on the historical media data to obtain loudness metadata; Loudness processing is performed on the media data to be processed based on the loudness metadata, and target media data is determined and played.
2. The method according to claim 1, characterized in that The performing loudness processing on the to-be-processed media data based on the loudness metadata, and determining and playing target media data, includes: Get the target loudness; determining, based on the loudness metadata and the target loudness, a current loudness gain and a current dynamic compression parameter corresponding to the media data to be processed; Loudness processing is performed on the media data to be processed based on the current loudness gain and the current dynamic compression parameter, and the target media data is determined and played.
3. The method according to claim 2, characterized in that Determining a current loudness gain corresponding to the to-be-processed media data based on the loudness metadata and the target loudness includes: Obtaining statistical loudness in the loudness metadata; The current loudness gain is obtained based on a difference between the target loudness and the statistical loudness.
4. The method according to claim 2, characterized in that Determining a current loudness gain corresponding to the to-be-processed media data based on the loudness metadata and the target loudness includes: Obtaining an anchor loudness in the loudness metadata, where the anchor loudness is determined based on a speech ratio in the historical media data; The current loudness gain is obtained based on a difference between the target loudness and the anchor loudness.
5. The method according to claim 2, characterized in that Determining a current dynamic compression parameter corresponding to the to-be-processed media data based on the loudness metadata and the target loudness includes: Obtaining short-term loudness and instantaneous loudness from the loudness metadata; respectively determining a first difference between the short-term loudness and the target loudness, and a second difference between the instantaneous loudness and the target loudness; Obtaining a first compression ratio based on a ratio of the first difference to a first preset value; obtaining a second compression ratio based on a ratio of the second difference to a second preset value, wherein the second preset value is greater than the first preset value; The maximum value of the first compression ratio and the second compression ratio is determined as the current dynamic compression ratio, and the dynamic compression parameter includes the current dynamic compression ratio.
6. The method according to claim 2, characterized in that Determining a current dynamic compression parameter corresponding to the to-be-processed media data based on the loudness metadata and the target loudness further includes: A difference between the target loudness and a third preset value is determined as a threshold value, the current dynamic compression parameter includes the threshold value, and the third preset value is greater than the second preset value.
7. The method according to claim 2, characterized in that The performing loudness processing on the to-be-processed media data based on the current loudness gain and the current dynamic compression parameter, and determining and playing the target media data, includes: performing smoothing processing on the current loudness gain and the current dynamic compression parameter respectively to obtain a target loudness gain and a target dynamic compression parameter; performing loudness processing on the to-be-processed media data based on the target loudness gain and the target dynamic compression parameter to obtain target media data; Play the target media data.
8. The method according to claim 7, characterized in that Smoothing the current loudness gain to obtain the target loudness gain includes: Obtaining a historical loudness gain corresponding to the last media data to be processed and a preset number of sampling points corresponding to the media data to be processed; determining a difference between the current loudness gain and the historical loudness gain; Obtaining a loudness gain adjustment step size based on a ratio of the difference to the preset number of sampling points; The target loudness gain is obtained based on the historical loudness gain and the loudness gain adjustment step.
9. The method according to any one of claims 1 to 8, characterized in that Get historical media data for a preset duration, including: Obtaining a window length and a preset step length of a sliding window, wherein the window length is the preset duration; The acquired media data is extracted based on the window length of the sliding window and the preset step length to obtain the historical media data, wherein the preset step length corresponds to a time interval between two adjacent acquisitions of the media data to be processed.
10. A loudness processing device for media data, characterized in that: The device comprises: a media data acquisition module configured to acquire the media data to be processed and historical media data of a preset duration, wherein the end position of the historical media data corresponds to the start position of the media data to be processed, and the preset duration is less than the cumulative duration of all processed media data; a loudness analysis module, configured to perform loudness analysis based on the historical media data to obtain loudness metadata; The loudness processing module is configured to perform loudness processing on the media data to be processed based on the loudness metadata, and determine and play target media data.
11. An electronic device, characterized in that: include: A memory and a processor, wherein the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the loudness processing method for media data according to any one of claims 1 to 9 by executing the computer instructions.
12. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a computer to execute the loudness processing method for media data according to any one of claims 1 to 9.