Data processing method and apparatus, storage medium, electronic device, and program product
By obtaining the initial loudness parameters and reference information of the audio data, the audio processing parameters are determined, solving the problem of uniform loudness of audio data, realizing personalized audio data processing, and improving the user experience.
Patent Information
- Application Number
- PCT/CN2025/086871
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-04-17
- Filing Date
- 2025-04-02
- Publication Date
- 2025-10-23
AI Technical Summary
Existing audio data playback technology cannot uniformly process loudness according to the needs of different users, resulting in a poor user experience.
By acquiring the initial loudness parameters and reference information of the target audio data, the audio processing parameters, including the target loudness parameters and sound effect parameters, are determined, and personalized audio data processing is performed.
It achieves precise processing of audio data, meets the loudness requirements of different users, and improves the user's playback experience.
Smart Images

Figure CN2025086871_23102025_PF_FP_ABST
Abstract
Description
Data processing method and device, storage medium, electronic device, and program product
[0001] This application claims priority to Chinese Patent Application No. 202410466062.5, filed on April 17, 2024, the disclosure of which is incorporated herein in its entirety as part of the present application. TECHNICAL FIELD
[0002] The present disclosure relates to a data processing method, device, storage medium, electronic device and program product. BACKGROUND
[0003] With the development of computer technology, audio technology has also developed. For example, users can watch audio or post audio through some application programs. These application programs can be installed on various electronic devices, so that the electronic devices can serve as audio playback terminals. SUMMARY
[0004] This summary is provided to introduce a selection of concepts, which will be described in greater detail below in the detailed description section. This summary does not intend to identify key or essential features of the claimed technology, nor is it intended for use in limiting the scope of the claimed technology.
[0005] In a first aspect, the present disclosure provides a data processing method, comprising: obtaining an initial loudness parameter of target audio data and at least one reference information, the reference information being capable of affecting a required loudness of the target audio data; determining an audio processing parameter according to the initial loudness parameter and the at least one reference information, the audio processing parameter comprising: a target loudness parameter; and processing the target audio data according to the audio processing parameter.
[0006] In a second aspect, the present disclosure provides a data processing device, comprising: an obtaining module configured to obtain an initial loudness parameter of target audio data and at least one reference information, the reference information being capable of affecting a required loudness of the target audio data; a processing module configured to determine an audio processing parameter according to the initial loudness parameter and the at least one reference information, the audio processing parameter comprising: a target loudness parameter; and process the target audio data according to the audio processing parameter.
[0007] In a third aspect, the present disclosure provides a computer-readable medium having a computer program stored thereon, the program being executed by a processing device to implement the steps of the data processing method of the first aspect.
[0008] In a fourth aspect, the present disclosure provides an electronic device, comprising: a storage device having a computer program stored thereon; and a processing device configured to execute the computer program in the storage device to implement the steps of the data processing method according to the first aspect.
[0009] In a fifth aspect, the present disclosure provides a computer program product, comprising a computer program configured to implement the steps of the data processing method according to the first aspect when executed by a processor.
[0010] Other features and advantages of the present disclosure will be further described in the following detailed description. BRIEF DESCRIPTION OF DRAWINGS
[0011] The above and other features and advantages of the embodiments of the present disclosure will become more apparent by describing in detail some embodiments thereof with reference to the attached drawings in which:
[0012] Fig. 1 is a flow chart of a data processing method according to an embodiment of the present disclosure.
[0013] Fig. 2 is a decision diagram of an audio processing parameter according to an embodiment of the present disclosure.
[0014] Fig. 3 is a flow chart of an audio parameter decision according to an embodiment of the present disclosure.
[0015] Fig. 4 is a flow chart of an audio data processing according to an embodiment of the present disclosure.
[0016] Fig. 5 is a block diagram of an audio data processing according to an embodiment of the present disclosure.
[0017] Fig. 6 is an example diagram of an application scenario according to an embodiment of the present disclosure.
[0018] Fig. 7 is a structural block diagram of a data processing device according to an embodiment of the present disclosure.
[0019] Fig. 8 is a structural diagram of an electronic device suitable for implementing an embodiment of the present disclosure. DETAILED DESCRIPTION
[0020] Embodiments of the present disclosure will be described below in greater detail with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as being limited to the embodiments set forth herein, but rather these embodiments are provided so as to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are merely for exemplary purposes and are not intended to limit the scope of protection of the present disclosure.
[0021] It should be understood that each step recited in the method embodiments of the present disclosure can be performed in different orders and / or in parallel. In addition, the method embodiments can include additional steps and / or omit performing the steps shown. The scope of the present disclosure is not limited in this respect.
[0022] The term "comprising" and variations thereof as used herein are open-ended, and mean "including but not limited to". The term "based on" means "based, at least in part, on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Related terms shall be construed accordingly.
[0023] It should be noted that the terms "first", "second", and the like in the present disclosure are merely used to distinguish different devices, modules or units, and do not imply the order or interdependence of the functions performed by these devices, modules or units.
[0024] It should be noted that the terms "one", "multiple" in the present disclosure are illustrative and not restrictive, and those skilled in the art should understand that "one or more" should be understood unless otherwise explicitly indicated in the context.
[0025] The names of the messages or information exchanged between the devices in the embodiments of the present disclosure are merely for illustrative purposes, and are not intended to limit the scope of the messages or information.
[0026] It can be understood that, before using the technical solutions disclosed in the embodiments of the present disclosure, the type, use range, use scenario, etc. of the personal information involved in the present disclosure should be informed to the user and the authorization of the user should be obtained in a proper manner according to relevant laws and regulations.
[0027] For example, in response to receiving the active request of the user, prompt information is sent to the user to explicitly prompt the user that the operation requested to be performed will require obtaining and using the personal information of the user. Thus, the user can voluntarily choose whether to provide the personal information to the software or hardware such as electronic device, application program, server or storage medium, etc. performing the operation of the technical solutions of the present disclosure according to the prompt information.
[0028] As an optional but non-limiting implementation manner, in response to receiving the active request of the user, the manner of sending prompt information to the user may, for example, be a pop-up window manner, and the prompt information may, for example, be presented in the form of text in the pop-up window. In addition, the pop-up window may, for example, carry a selection control for the user to select "agree" or "disagree" to provide the personal information to the electronic device.
[0029] It can be understood that the above notification and user authorization obtaining process is only illustrative and does not limit the implementation of the present disclosure, and other methods that meet relevant laws and regulations can also be applied to the implementation of the present disclosure.
[0030] At the same time, it can be understood that the data involved in the technical solution (including but not limited to the data itself, the acquisition or use of the data) should comply with the requirements of relevant laws and regulations and relevant provisions.
[0031] The technical solution provided by the embodiment of the present disclosure can be applied to various audio data processing scenarios, in which the problem of loudness uniformity needs to be solved.
[0032] Audio data can be understood as sound data that can be heard by human ears. In some application scenarios, audio data exists alone. In other application scenarios, audio data and image data can form video data.
[0033] It can be understood that if the audio data does not exist alone, but forms video data with image data or more data, the processing of the audio data can also be directly regarded as the processing of the video data. Therefore, the technical solution of the embodiment of the present disclosure can also be applied to various video data processing scenarios.
[0034] Loudness uniformity can be understood as the loudness of multiple audio data played on the same audio data playing end being consistent. Since the loudness of different audio data can be different, if different audio data is directly played on the same playing end, it will cause the user to feel different loudness when watching different audio data. Therefore, the problem of loudness uniformity needs to be solved, and on the basis of loudness uniformity, the user of the audio data playing end will not feel different loudness, and the experience of the user can be guaranteed.
[0035] In the related art, the loudness equalization technology of the watching end (watching / playing end) is to process different audio data with different loudness to the same loudness (for example: -18, -16, etc. loudness value) through volume gain (Gain) and intensity limiting (Limiter) modules, so as to realize loudness equalization.
[0036] It can be understood that the watching end can be a client, which can be installed on an electronic device. Through the watching end, the user can play audio data, for example, the user plays a short video, the data corresponding to the short video includes audio data and picture data, and when the user plays the short video, the corresponding audio data will be played.
[0037] However, in the loudness equalization technology, the same set of algorithm parameters is used, which cannot meet the needs (preferences) of different users, and is not diverse, thereby resulting in poor user playback experience.
[0038] Considering that different users may have different needs, such as loudness preferences, dynamic range preferences, etc., or the needs change with time, space, etc., the processing of the audio data needs to meet the needs of the user.
[0039] For example, using the related technical solutions, the user's experience of playing audio may include that a user who frequently publishes videos complains that the published videos are over-processed in dynamic range when played after being published. A user who frequently watches short videos in speaker mode still encounters situations where the volume is large and small when watching videos, and needs to adjust the system volume to solve the problem. Different users have different needs for loudness, such as the elderly needing higher loudness and children needing lower loudness. Different video content has different optimal loudness and dynamic range compression ratios. In different power and different processor performance situations, the combination of loudness processing and other sound effects can be simplified or directly ignored to provide the most basic and optimal user experience.
[0040] Based on this, the embodiments of the present disclosure provide a technical solution that determines target loudness parameters in combination with loudness needs, and then processes audio data based on the target loudness needs, so that the processed audio data can also meet the corresponding needs, realizes accurate processing of audio data, and further improves the playback experience of audio data.
[0041] Please refer to FIG. 1, which shows a flowchart of a data processing method according to an embodiment of the present disclosure. The data processing method includes the following steps:
[0042] In step S11, the initial loudness parameters of the target audio data and at least one reference information are obtained.
[0043] In step S12, the audio processing parameters are determined according to the initial loudness parameters and the at least one reference information.
[0044] In step S13, the target audio data is processed according to the audio processing parameters.
[0045] In step S11, the reference information can affect the demand loudness of the target audio data. In some embodiments, the reference information can directly represent the demand loudness. In other embodiments, the reference information may not be able to directly represent the demand loudness, but the corresponding demand loudness can be determined through analysis and processing of the reference information. In addition, the reference information can represent the demand for the loudness, i.e., the demand for the value of the loudness, or it can represent the demand for the processing of the loudness, i.e., whether to process the loudness.
[0046] As an optional implementation, the at least one reference information can be at least one of the following information: a posture of a playback user of the audio data; preference information of the playback user of the audio data; device performance data of a playback device of the audio data, the device performance data including at least one of loudness playback performance data, battery performance data, and processor performance data; playback environment information of the audio data; and a voice proportion of the audio data.
[0047] It can be understood that the above information is only as an optional implementation, and can also be other information, for example, the requirements of elderly users and child users for loudness are different; the requirements of hearing-impaired users and healthy users for loudness are also different, and the like, which are not exemplified one by one in the embodiments of the present disclosure.
[0048] The posture of the playback user, for example, walking, lying, and the like. It can be understood that the electronic device usually has a sensor that can detect the posture of the user, such as a gyroscope, an accelerometer, and the like. On the basis of obtaining the authorization of the user, the posture information of the user can be directly obtained. When the user is in different postures, the requirements for loudness can be different, for example, the user can require a smaller loudness when lying, and the user can require a larger loudness when walking. Therefore, the posture of the user can be regarded as a reference information.
[0049] The preference information of the playback user can represent the preference of the user for the audio data, for example, a publishing frequency, a playback duration, a volume adjustment, and the like.
[0050] If the publishing frequency is greater than 50%, it represents that the requirement for loudness is higher; if the publishing frequency is less than 50%, it represents that the requirement for loudness is lower. If the playback duration is longer, it represents that the requirement for loudness is higher; if the playback duration is shorter, it represents that the requirement for loudness is lower. If the frequency of increasing the volume is higher than the frequency of decreasing the volume, it represents that the requirement for loudness is higher; if the frequency of increasing the volume is lower than the frequency of decreasing the volume, it represents that the requirement for loudness is lower. It should be noted that the preference information of the user is obtained with the permission of the user.
[0051] The device performance data of the playback device can analyze the requirement for loudness from the device level. The device performance data is the data that the device itself has.
[0052] The loudness playback performance data can represent the loudness playback capability of the device. For the same loudness of audio data, the higher the loudness playback capability, the higher the loudness of the audio data played by the device. Therefore, the requirement for loudness can also be analyzed through the loudness playback performance data.
[0053] The battery performance data can be the real-time power of the battery, the full charge of the battery, etc. When the power of the battery is high, the loudness can be processed; when the power of the battery is low, the loudness processing can not be realized. Therefore, through the battery performance data, the processing demand of the loudness can be analyzed.
[0054] The processor performance data can be the processing speed and efficiency of the processor. When the processor performance is good, for example, there is no lag, the loudness can be processed. When the processing performance is poor, for example, there is lag, the loudness processing can not be realized. Therefore, through the processor performance data, the processing demand of the loudness can be analyzed.
[0055] The playing environment information can be understood as the environment in which the playing device is currently located, for example, noisy environment, quiet environment, etc. Through the detection and distinction of the sound, the acquisition of the playing environment information can be realized. It can be understood that when the playing device is in a noisy environment, a higher loudness is required so that the user can hear clearly, so the loudness demand is higher. When the playing device is in a quiet environment, a lower loudness can be used. Therefore, through the playing environment information, the demand of the loudness can be determined.
[0056] The voice proportion can be understood as the proportion of the sound in the data. For example, for single audio data, the voice proportion can be 100%; or in the case of a silent part, the voice proportion is less than 100%; for audio data in video data, the voice proportion can not reach 100%. For audio data with higher voice proportion, loudness processing is required; for audio data with lower voice proportion, loudness processing can not be required. Therefore, through the voice proportion, the processing demand of the loudness can be determined. Through the analysis of the audio data and the video data, the voice proportion information can be determined.
[0057] In step S11, the target audio data can be audio data issued by the server. It can be understood that the audio data playing end corresponds to a corresponding server, and the server can transmit audio data to the playing end to make the playing end play.
[0058] In some embodiments, before the server issues the audio data, the audio data can be processed by the loudness equalization. The loudness equalization processing method used by the server is to process the loudness of the audio data to a fixed loudness value. Thus, by uniformly processing the video to a certain fixed loudness, the difference of the audio loudness is reduced under the premise of minimizing the loss of sound quality.
[0059] In some embodiments, the audio data issued by the server is also the audio data uploaded by the corresponding client. For example, assuming that the client is divided into a video publishing (creating) side and a video watching (consuming) side. Then, the user of the video publishing side publishes a video through a video publishing operation, the published video is uploaded to the server, and the server processes and then issues the video to the video watching side for playing. In addition, the video publishing side can also be one of the video watching sides.
[0060] Therefore, as an optional implementation, the data processing method further includes: in response to receiving a publishing request of the initial audio data, sending the initial audio data to the server to make the server process the initial audio data to obtain target audio data; receiving the target audio data issued by the server; and playing the processed target audio data.
[0061] It can be understood that this implementation corresponds to a complete process from audio data publishing to audio data playing, and the audio data playing side is the same as the audio data publishing side. Therefore, for the user publishing the audio data, the played audio data can match the user's demand, thereby improving the user experience.
[0062] Therefore, the technical solution provided by the embodiments of the present disclosure can be adapted to various application scenarios of audio data, and can make the processed audio data match the user's loudness demand, thereby improving the user experience.
[0063] In step S11, the initial loudness parameter can be information contained in data issued together with the target audio data, which can be understood as the current loudness parameter of the target audio data. For example, the server issues the target audio data and audio data metadata, and the metadata is data used to describe the audio data, which can include the initial loudness parameter.
[0064] In step S12, the audio processing parameter includes a target loudness parameter. The target loudness parameter can be understood as a loudness parameter determined to match the demand loudness according to the initial loudness parameter and at least one reference information.
[0065] In some embodiments, the audio processing parameter can include not only the target loudness parameter but also an audio effect parameter. Therefore, in step S13, the target audio data can be processed according to the target loudness parameter and the audio effect parameter.
[0066] As a first optional implementation, step S12 includes: determining at least one loudness compensation value according to at least one reference information, wherein each reference information corresponds to one loudness compensation value.
[0067] The target loudness parameter is determined according to the at least one loudness compensation value and the initial loudness parameter.
[0068] In this embodiment, the target loudness parameter can be represented as: the initial loudness parameter + at least one loudness compensation value, which can also be understood as a loudness adjustment offset.
[0069] For example, the sum of the at least one loudness compensation value can be represented as: x = f1(user posture) + f2(playback device performance) + f3(user preference) + f4(playback environment) + f5(voice proportion) +…+ fn(N); and the target loudness parameter can be represented as: x + x0. Wherein, x0 represents the initial loudness value, and f1-fn represents the corresponding relationship between the loudness compensation value and the reference information, so that based on the corresponding relationship, the loudness compensation value corresponding to each reference information can be determined; the corresponding relationship can be a conversion relationship, or a formula, an algorithm, etc. Then, summing the loudness compensation values can determine the sum x, and adding the initial loudness value to the sum x can determine the target loudness parameter.
[0070] In some embodiments, the sum of the at least one loudness compensation value is within the range of 0-2, wherein 0 and 2 are both available values.
[0071] As a second optional embodiment, step S12 can include: determining the audio processing parameter according to the preset decision model or the preset decision module, the initial loudness parameter and the at least one reference information, the preset decision model or the preset decision module being used to output the audio processing parameter according to the input initial loudness parameter and the at least one reference information.
[0072] It can be understood that both the target loudness parameter and the sound effect parameter can be determined using this embodiment.
[0073] In some embodiments, some prior data can be obtained first, the prior data including: a plurality of groups of audio parameters, each of which corresponds to an initial loudness parameter and at least one reference information. Using the prior data, the decision model can be trained. The trained decision model can directly output the audio processing parameter based on the input initial loudness parameter and the at least one reference information.
[0074] The decision model can be a neural network model, a random forest model, etc.
[0075] In some embodiments, the decision module also has the ability to output the audio processing parameter based on the input initial loudness parameter and the at least one reference information, like the decision model. For example, the decision module can be a module loaded with a decision model or an audio processing parameter determination algorithm.
[0076] Please refer to FIG. 2 below, which shows a decision diagram of an audio processing parameter according to an embodiment of the present disclosure. As shown in FIG. 2, the determination of the target loudness parameter and the sound effect parameter is realized through a sound effect chain decision module or model.
[0077] In FIG. 2, the user posture, the preference index (preference release / play), the device playback capability, the environment, the voice proportion, the power, and the performance index such as lag are input into the sound effect chain decision module or model, and the sound effect chain decision module or model outputs the target loudness and the sound effect chain configuration parameter. The sound effect chain configuration parameter, for example, is whether each node (processing module in the sound effect chain) is turned on, and the specific algorithm parameter, etc.
[0078] In some embodiments, the target audio data can be processed based on the target loudness parameter and the sound effect parameter, so that the processed audio data can meet the target loudness parameter and the sound effect parameter.
[0079] In another embodiment, the target loudness parameter can also be embodied in the finally determined sound effect parameter. Therefore, as an optional implementation, the audio processing parameter further includes an initial sound effect parameter. In this implementation, the initial sound effect parameter in the audio processing parameter can be the sound effect parameter determined according to the initial loudness parameter and at least one reference information.
[0080] Further, step S13 can include determining a target sound effect parameter according to the initial sound effect parameter and the target loudness parameter, and processing the target audio data according to the target sound effect parameter.
[0081] In this implementation, the integrated target sound effect parameter can be determined in combination with the initial sound effect parameter and the target loudness parameter, and then the target audio data is processed according to the target sound effect parameter. That is, the target loudness parameter can be controlled through the target sound effect parameter in the final processing.
[0082] As an optional implementation, the target sound effect parameter includes at least one of the gain parameter, the frequency response parameter, the dynamic range compression parameter, and the intensity limiting parameter. The definition or function of these parameters can refer to the mature technology in the field, and will not be specifically introduced here.
[0083] As a first optional implementation, processing the target audio data according to the target sound effect parameter includes sequentially processing the target audio data according to the frequency response parameter, the gain parameter, the dynamic range compression parameter, and the intensity limiting parameter.
[0084] In this implementation, when the frequency response parameter is involved, the processing of the frequency response parameter is performed first, and then the processing of the gain parameter, the dynamic range compression parameter and the loudness limiting parameter is performed in sequence; relatively comprehensive audio processing can be implemented, the playing effect of the audio data is ensured, and the playing experience is improved.
[0085] As a second optional implementation, the target audio data is processed according to the target sound effect parameter, including: processing the target audio data according to the gain parameter, the dynamic range compression parameter and the loudness limiting parameter in sequence.
[0086] In this implementation, when the frequency response parameter is not involved, the processing of the gain parameter, the dynamic range compression parameter and the loudness limiting parameter is performed in sequence; unnecessary audio processing process can be avoided, and the audio processing efficiency is improved.
[0087] In some embodiments, if the playing mode of the audio data is the earphone playing mode or the loudspeaker playing mode, the target sound effect parameter includes: the gain parameter, the frequency response parameter, the dynamic range compression parameter and the loudness limiting parameter; if the playing mode of the audio data is not the earphone playing mode and not the loudspeaker playing mode, the target sound effect parameter includes: the gain parameter, the dynamic range compression parameter and the loudness limiting parameter.
[0088] In this implementation, the frequency response parameter is involved in the case of the earphone playing mode or the loudspeaker playing mode; the frequency response parameter can not be involved in other playing modes.
[0089] In the embodiments of the present disclosure, when multiple sound effect parameters are involved, the audio data processing can be implemented through the sound effect processing chain.
[0090] Therefore, as an optional implementation, the audio processing parameter further includes at least one sound effect parameter, and the step S13 includes: creating an audio processing chain according to the at least one sound effect parameter, the audio processing chain including at least one audio processing module, each audio processing module corresponding to one sound effect parameter and the target loudness parameter; and processing the target audio data through the audio processing chain.
[0091] In combination with the foregoing introduction of the embodiments, the at least one sound effect parameter in this implementation can be an initial sound effect parameter, so that each audio processing module can perform audio processing according to one initial sound effect parameter and the target loudness parameter. Alternatively, the at least one sound effect parameter can also be a target sound effect parameter, which is determined in combination with the initial sound effect parameter and the target loudness parameter. Thus, each audio processing module can perform audio processing according to one target sound effect parameter.
[0092] It can be understood that the number of audio processing modules in the sound effect processing chain is consistent with the number of sound effect parameters. For example, if there are 4 sound effect parameters in total, 4 audio processing modules are created.
[0093] For example, the sound effect processing chain can sequentially include a frequency response module, a gain module, a dynamic range compression module, and an intensity limiting module.
[0094] In other embodiments, if the target loudness parameter is the same as the initial loudness parameter, it means that no loudness equalization processing is needed. At this time, the sound effect processing chain can not be created, that is, the audio processing process is not performed, and the audio data can be directly played. Through this implementation, in the case of poor device performance or no special loudness requirement of the user, meaningless audio processing process can be avoided, the audio processing efficiency is improved, and the device performance is guaranteed.
[0095] In some embodiments, the sound effect processing chain can be determined by the sound effect chain processing module according to the audio processing parameters determined in step S12.
[0096] In the embodiments of the present disclosure, step S11 and step S12 can be implemented by the pre-set decision model or the pre-set decision module introduced in the foregoing embodiments, and step S13 can be implemented by the sound effect chain processing module in this implementation. Through the combined application of the model or the module, the processing of the audio data can be quickly realized, the processing efficiency of the audio data is improved, and the accuracy of the audio data processing can also be guaranteed.
[0097] Please refer to FIG. 3, which shows an audio parameter decision flowchart according to an embodiment of the present disclosure. As shown in FIG. 3, first, the power and performance are judged. If the power is low or the performance is too poor, the loudness offset can be 0, and the client does not need to perform audio processing, that is, the sound effect processing chain is not created.
[0098] Then, the offsets corresponding to the posture, preference index (preference publishing / play), device playback capability, voice proportion, environment, etc. are calculated, and the target loudness is obtained by integrating the offsets.
[0099] Next, if the device frequency response is known, it means that the frequency response parameter needs to be calculated at present; and if the device frequency response is unknown, it means that the frequency response parameter does not need to be calculated.
[0100] In addition, the gain parameter, the dynamic range compression parameter, and the intensity limiting parameter also need to be calculated respectively. The calculation of the intensity limiting parameter can be combined with the target loudness parameter, and the calculation of the other two parameters can be combined with the target loudness parameter. Of course, it can also be combined with the target loudness parameter, which is not limited here.
[0101] Reference is made to FIG. 4, which shows an audio data processing flowchart of an embodiment of the present disclosure. As shown in FIG. 4, the decision module outputs target loudness parameters and sound effect (processing) chain configuration parameters; then, the sound effect chain is created. Further, the sound effect chain is processed; finally, the processed audio data is played.
[0102] Reference is made to FIG. 5, which shows an audio data processing block diagram of an embodiment of the present disclosure. As shown in FIG. 5, in actual application, the technical solution of the embodiment of the present disclosure can be divided into a parameter preparation stage and a streaming processing stage.
[0103] In the parameter preparation stage, the target loudness and various sound effect parameters are calculated according to the audio metadata carried by the video and the information related to the user end. The various audio parameters involve gain parameters, dynamic range compression parameters, and intensity limiting parameters. In the streaming processing stage, the audio data is first adjusted in gain according to the gain parameters, and then the loudness is normalized before entering the dynamic range compression module to adjust the loudness and dynamic range. Finally, the gain and limiting are adjusted by the intensity limiting module as a whole to prevent the explosion of sound. Finally, the processed audio is output. In the speaker playback mode, the frequency response module in the figure will start processing, and the server will issue a speaker gain compensation for the calculation of overall loudness equalization gain.
[0104] Reference is made to FIG. 6, which shows an example diagram of an application scenario suitable for an embodiment of the present disclosure. As shown in FIG. 6, it involves a creation side, a video cloud server, and a consumption side. The creation side and the consumption side can both serve as a video playback end.
[0105] In the video cloud server, the video is first uniformly processed to a certain platform loudness, and the difference in video loudness is reduced on the premise of minimizing the loss of sound quality. The video cloud server also achieves data processing by creating a sound effect chain, which includes a gain module (gain), a dynamic range compression module (drc), and an intensity limiting module (limiter).
[0106] In the video playback end, the video cloud server issues the transcoded video and the video loudness metadata to the client together, which will be processed by the decision module. The inputs of the processing include audio content labels (such as voice, music, news, etc.), current user preference representation indicators (such as daily average publishing rate and daily average playing time), current environment (background noise), device playback capability, and audio indicators. After processing, the gain (gain) of the loudness equalization, the dynamic and loudness adjustment parameters (DRC, Limiter parameters) are obtained. In addition, in some cases of speaker playback mode, the EQ processing (an EQ (frequency response) processing module is added before DRC) is started, which simulates the attenuation of the digital signal by the speaker. The non-speaker module does not start the frequency response processing.
[0107] A compensation gain is added to the gain of the subsequent loudness processing, which converts the performance energy attenuated by the previous frequency response processing into gain compensation to the subsequent loudness equalization link to ensure that the overall loudness equalization is achieved in the speaker playback mode. It is worth noting that the entire link supports real-time processing during short video playback. For example, after frequency response processing, the signal processed by the frequency response does not need to be scanned for loudness again on the client side. Instead, the loudness information of the source video and the speaker compensation gain are scanned through the cloud, directly delivered to the client, and the audio is directly processed through the entire audio chain during playback.
[0108] For the above-mentioned video cloud server platform loudness normalization link, the processing link order is fixed, such as gain-drc-limiter. After the platform loudness normalization, the audio with regular loudness is played on the consumption side, and the target loudness and dynamic range are fine-tuned according to more personalized parameters, which may also include frequency response processing. More personalized decisions are made on the client side through the audio decision module to determine which nodes and specific parameters are included in the audio chain required for playing the current video. This audio chain and parameter are automatically adjusted according to the user's brushing of different videos, user preferences, and power changes, so as to achieve the best user audio experience.
[0109] Through actual application, this personalized loudness processing scheme can improve the daily activity of the client and increase the user's stay time. Therefore, it can be seen that this technical scheme can significantly improve the user's playback experience.
[0110] Reference is made to FIG. 7, which shows a structural block diagram of a data processing apparatus according to an embodiment of the present disclosure. As shown in FIG. 7, the data processing apparatus includes:
[0111] The acquisition module 701 is configured to acquire an initial loudness parameter of target audio data and at least one reference information capable of affecting a required loudness of the target audio data. The processing module 702 is configured to determine an audio processing parameter according to the initial loudness parameter and the at least one reference information, the audio processing parameter including a target loudness parameter, and process the target audio data according to the audio processing parameter.
[0112] In some embodiments, the data processing apparatus further includes a sending module and a playing module. The sending module is configured to send the initial audio data to a server in response to receiving a publishing request of the initial audio data, so that the server processes the initial audio data to obtain the target audio data. The acquisition module 701 is further configured to receive the target audio data sent by the server. The playing module is configured to play the processed target audio data.
[0113] In some embodiments, the processing module 702 is further configured to determine at least one loudness compensation value according to the at least one reference information, wherein each reference information corresponds to one loudness compensation value; and determine the target loudness parameter according to the at least one loudness compensation value and the initial loudness parameter.
[0114] In some embodiments, the processing module 702 is further configured to determine a target sound effect parameter according to the initial sound effect parameter and the target loudness parameter; and process the target audio data according to the target sound effect parameter.
[0115] In some embodiments, the target sound effect parameter comprises at least one of a gain parameter, a frequency response parameter, a dynamic range compression parameter, and an intensity clipping parameter.
[0116] In some embodiments, the processing module 702 is further configured to process the target audio data according to the frequency response parameter, the gain parameter, the dynamic range compression parameter, and the intensity clipping parameter in sequence; or process the target audio data according to the gain parameter, the dynamic range compression parameter, and the intensity clipping parameter in sequence.
[0117] In some embodiments, if the playback mode of the audio data is a headphone playback mode or a speaker playback mode, the target sound effect parameter comprises a gain parameter, a frequency response parameter, a dynamic range compression parameter, and an intensity clipping parameter; and if the playback mode of the audio data is neither a headphone playback mode nor a speaker playback mode, the target sound effect parameter comprises a gain parameter, a dynamic range compression parameter, and an intensity clipping parameter.
[0118] In some embodiments, the processing module 702 is further configured to create an audio processing chain according to the at least one sound effect parameter, wherein the audio processing chain comprises at least one audio processing module, each of the at least one audio processing module corresponding to one sound effect parameter and the target loudness parameter; and process the target audio data through the audio processing chain.
[0119] In some embodiments, the processing module 702 is further configured to determine an audio processing parameter according to a preset decision model or a preset decision module, the initial loudness parameter, and the at least one reference information, wherein the preset decision model or the preset decision module is configured to output the audio processing parameter according to the input initial loudness parameter and the at least one reference information.
[0120] In some embodiments, the at least one reference information is at least one of the following information: a posture of a playing user of the audio data; preference information of the playing user of the audio data; device performance data of a playing device of the audio data, the device performance data including at least one of loudness playback performance data, battery performance data, and processor performance data; playing environment information of the audio data; and a voice proportion of the audio data.
[0121] Reference is made below to FIG. 8, which illustrates a structural diagram of an electronic device 800 suitable for use in implementing embodiments of the present disclosure. The terminal device in embodiments of the present disclosure can include, but is not limited to, a mobile terminal such as a mobile phone, a notebook computer, a digital broadcast receiver, a PDA (Personal Digital Assistant), a PAD (Tablet Personal Computer), a PMP (Portable Multimedia Player), a car terminal (e.g., a car navigation terminal), and the like, as well as a stationary terminal such as a digital TV, a desktop computer, and the like. The electronic device shown in FIG. 8 is merely an example, and should not impose any limitation on the functions and use range of embodiments of the present disclosure.
[0122] As shown in FIG. 8, the electronic device 800 can include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 801 that can perform various appropriate actions and processes according to programs stored in a read-only memory (ROM) 802 or loaded into a random access memory (RAM) 803 from a storage device 808. Various programs and data required for the operation of the electronic device 800 are also stored in the RAM 803. The processing device 801, the ROM 802, and the RAM 803 are connected to each other through a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.
[0123] Generally, the following devices can be connected to the I / O interface 805: an input device 806 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, and the like; an output device 807 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, and the like; a storage device 808 including, for example, a magnetic tape, a hard disk, and the like; and a communication device 809. The communication device 809 can allow the electronic device 800 to communicate with other devices wirelessly or via wires to exchange data. Although FIG. 8 shows the electronic device 800 with various devices, it should be understood that all of the shown devices are not required to be implemented or provided. More or fewer devices can be alternatively implemented or provided.
[0124] In particular, the processes described above with reference to the flowcharts can be implemented as a computer software program in accordance with embodiments of the present disclosure. For example, embodiments of the present disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program comprising program code for executing the methods illustrated by the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via the communication device 809, or installed from the storage device 808, or installed from the ROM 802. When the computer program is executed by the processing device 801, the above-described functions defined in the methods of the embodiments of the present disclosure are performed.
[0125] It should be noted that the computer-readable medium described above in the present disclosure can be a computer-readable signal medium or a computer-readable storage medium or any combination thereof. The computer-readable storage medium can be, for example but not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or apparatus, or any suitable combination thereof. More specific examples of the computer-readable storage medium can include, but are not limited to, an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In the present disclosure, the computer-readable storage medium can be any tangible medium that contains or stores a program used by or in connection with an instruction execution system, apparatus, or device. In the present disclosure, the computer-readable signal medium can include a data signal propagated in baseband or propagated as a carrier wave in a propagated data signal, in which the computer-readable program code is carried. Such a propagated data signal can take a variety of forms, including but not limited to electro-magnetic, optical, or any suitable combination thereof. The computer-readable signal medium can also be any computer-readable medium that can be used to carry or store program code for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to wire, cable, optical fiber, RF, etc., or any suitable combination thereof.
[0126] In some embodiments, the client can communicate using any currently known or future developed network protocol, such as HTTP (HyperText Transfer Protocol), and can be interconnected with digital data communications (e.g., a communications network) of any form or medium, including, for example, local area networks ("LANs"), wide area networks ("WANs"), internetworks (e.g., the Internet), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any currently known or future developed networks.
[0127] The computer readable medium described above can be included in the electronic device described above; or can exist separately from the electronic device and can be mounted in the electronic device.
[0128] The computer readable medium described above carries one or more programs, when the one or more programs are executed by the electronic device, cause the electronic device to: acquire an initial loudness parameter of target audio data and at least one reference information, the reference information being capable of affecting a required loudness of the target audio data; determine an audio processing parameter according to the initial loudness parameter and the at least one reference information, the audio processing parameter including: a target loudness parameter; and process the target audio data according to the audio processing parameter.
[0129] Computer program code for carrying out operations of the present disclosure can be written in any combination of one or more programming languages, including an object oriented programming language such as Java, Smalltalk, C++ or the like, and conventional procedural programming languages, such as the "C" programming language or similar programming languages. The program code can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network ("LAN") or a wide area network ("WAN"), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider).
[0130] The computer program product of the first aspect can include one or more non-transitory computer-readable media storing instructions that, when executed, cause one or more processors to perform the operations of the first aspect. The one or more non-transitory computer-readable media can include, for example, magnetic media such as one or more magnetic disks, magnetic tapes or cassettes; optical media such as one or more compact discs, optical discs or Blu-ray discs; magneto-optical media such as one or more floptical discs; solid state media such as one or more solid state drives or other flash memory arrays; or any suitable combination of these. The one or more non-transitory computer-readable media can be encoded with instructions that, when executed, cause one or more processors to perform the operations of the first aspect.
[0131] The modules described in the embodiments of the present disclosure can be implemented by software, or by hardware. In some cases, the name of a module does not constitute a limitation on the module itself. For example, a obtaining module can also be described as a module that obtains an initial loudness parameter of target audio data and at least one reference information.
[0132] The functions described above can be performed, at least in part, by one or more hardware logic components. For example, non-limiting examples of exemplary types of hardware logic components that can be used include field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), system-on-a-chip (SOCs), complex programmable logic devices (CPLDs), etc.
[0133] In the context of the present disclosure, a machine-readable medium can be a tangible medium that contains or stores a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium will include one or more lines of electrical connections, portable computer disks, hard disk drives, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or Flash memory), optical fibers, portable compact disc read-only memories (CD-ROMs), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0134] By the technical solution, the audio processing parameter is determined based on the initial loudness parameter of the target audio data and the at least one reference information, and the target audio data is processed according to the audio processing parameter. Since the reference information can affect the demand loudness of the audio data, the target loudness parameter in the audio processing parameter is a loudness parameter determined based on the initial loudness parameter and the demand loudness, so the target loudness parameter can match the loudness demand of the audio data. Therefore, the technical solution can determine the target loudness parameter matching the loudness demand, so that the audio data processed based on the target loudness parameter can also meet the corresponding demand, realize accurate processing of the audio data, and improve the playing experience of the audio data.
[0135] The above description is merely the preferred embodiments of the present disclosure and the explanation of the applied technical principles. Those skilled in the art should understand that the disclosed range of the present disclosure is not limited to the technical solutions formed by the specific combinations of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or equivalent features without departing from the above disclosed concept. For example, the technical solutions formed by replacing the above features with the technical features disclosed in the present disclosure (but not limited to) having similar functions.
[0136] In addition, although each operation is described in a particular order, this should not be understood as requiring the operations to be performed in the specific order shown or in a sequential order. In certain circumstances, multitasking and parallel processing can be advantageous. Similarly, although several implementation details are included in the above discussion, these should not be interpreted as limiting the scope of the present disclosure. Certain features described in the context of separate embodiments can also be combined in a single embodiment. Conversely, various features described in the context of a single embodiment can also be separated and implemented in multiple embodiments.
[0137] Although the subject matter has been described in language specific to structural features and / or methodological acts, it is to be understood that the subject defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are merely illustrative of specific ways to implement the claims. Regarding the devices in the above embodiments, the specific manner in which the various modules perform operations has been described in detail in the embodiments related to the method, and will not be described here in detail.
Claims
1. A data processing method, comprising: obtaining an initial loudness parameter of target audio data and at least one reference information, the reference information capable of affecting a required loudness of the target audio data; determining an audio processing parameter according to the initial loudness parameter and the at least one reference information, the audio processing parameter comprising a target loudness parameter; processing the target audio data according to the audio processing parameter.
2. The data processing method of claim 1, further comprising: in response to receiving a publishing request of initial audio data, sending the initial audio data to a server to make the server process the initial audio data to obtain the target audio data; receiving the target audio data issued by the server; playing the processed target audio data.
3. The data processing method according to claim 1 or 2, wherein, The determining of the audio processing parameter according to the initial loudness parameter and the at least one reference information comprises: determining at least one loudness compensation value according to the at least one reference information, wherein each reference information corresponds to one loudness compensation value; determining the target loudness parameter according to the at least one loudness compensation value and the initial loudness parameter.
4. The data processing method according to any one of claims 1 to 3, wherein, The audio processing parameter further comprises an initial sound effect parameter, and the processing of the target audio data according to the audio processing parameter comprises: determining a target sound effect parameter according to the initial sound effect parameter and the target loudness parameter; processing the target audio data according to the target sound effect parameter.
5. The data processing method of claim 4, wherein, The target sound effect parameter comprises at least one of a gain parameter, a frequency response parameter, a dynamic range compression parameter and an intensity clipping parameter.
6. The data processing method of claim 5, wherein, The processing of the target audio data according to the target sound effect parameter comprises: processing the target audio data according to the frequency response parameter, the gain parameter, the dynamic range compression parameter and the intensity clipping parameter in sequence; or processing the target audio data according to the gain parameter, the dynamic range compression parameter and the intensity clipping parameter in sequence.
7. The data processing method according to claim 5 or 6, wherein, If a playing mode of the audio data is an earphone playing mode or a loudspeaker playing mode, the target sound effect parameter comprises the gain parameter, the frequency response parameter, the dynamic range compression parameter and the intensity clipping parameter; if the playing mode of the audio data is neither the earphone playing mode nor the loudspeaker playing mode, the target sound effect parameter comprises the gain parameter, the dynamic range compression parameter and the intensity clipping parameter.
8. The data processing method according to any one of claims 1 to 3, wherein, The audio processing parameter further comprises at least one sound effect parameter, and the processing of the target audio data according to the audio processing parameter comprises: creating an audio processing chain according to the at least one sound effect parameter, the audio processing chain comprising at least one audio processing module, each audio processing module corresponding to one sound effect parameter and the target loudness parameter; processing the target audio data through the audio processing chain.
9. The data processing method according to claim 1 or 2, wherein, The determining of the audio processing parameter according to the initial loudness parameter and the at least one reference information comprises: According to a preset decision model or a preset decision module, the initial loudness parameter and the at least one reference information, an audio processing parameter is determined, the preset decision model or the preset decision module is used to output the audio processing parameter according to the input initial loudness parameter and the at least one reference information.
10. The data processing method according to any one of claims 1 to 9, wherein, The at least one reference information is at least one of the following information: A posture of a user playing the audio data; Preference information of the user playing the audio data; Device performance data of a device playing the audio data, the device performance data including at least one of loudness playback performance data, battery performance data and processor performance data; Playing environment information of the audio data; A voice proportion of the audio data.
11. A data processing apparatus, comprising: an acquisition module configured to acquire an initial loudness parameter of target audio data and at least one reference information, the reference information being capable of affecting a required loudness of the target audio data; a processing module configured to determine an audio processing parameter according to the initial loudness parameter and the at least one reference information, the audio processing parameter including a target loudness parameter; processing the target audio data according to the audio processing parameter.
12. A computer readable medium having stored thereon a computer program, the program being executed by a processing apparatus to implement the steps of the data processing method of any one of claims 1-10.
13. An electronic device, comprising: a storage device having stored thereon a computer program; a processing apparatus configured to execute the computer program in the storage device to implement the steps of the data processing method of any one of claims 1-10.
14. A computer program product comprising a computer program, the computer program being executed by a processor to implement the steps of the data processing method of any one of claims 1-10.
Citation Information
Patent Citations
Audio data processing method and device, electronic equipment and storage medium
CN110928518A
Audio processing method and device, electronic equipment and storage medium
CN114071315A
Audio playing method, electronic equipment and readable storage medium
CN115268828A
Audio playing method and device, medium and electronic equipment
CN116755653A
Vehicle-mounted sound effect automatic adjustment method, device and equipment and readable storage medium
CN119110221A