Volume adjustment method and device, electronic equipment, medium and product

By acquiring metadata of media data and using regression processing to establish a volume adjustment probability model, the playback volume of media data is automatically adjusted, solving the problem of users manually adjusting the volume and improving the playback effect and user experience.

CN121722348APending Publication Date: 2026-03-24BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-09-23
Publication Date
2026-03-24

AI Technical Summary

Technical Problem

In existing technologies, users need to manually adjust the volume when playing media data on the terminal, which affects the user experience.

Method used

By acquiring the metadata of media data, a volume adjustment probability model is established using regression processing. Based on the model, the volume adjustment probability is predicted, and the playback volume of the media data is automatically adjusted.

Benefits of technology

This reduces the probability of volume adjustment during media playback, improving playback quality and user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121722348A_ABST
    Figure CN121722348A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of computer processing, and discloses a volume adjusting method and device, electronic equipment, a medium and a product. The invention provides a volume adjusting method. The volume adjusting method comprises the following steps: acquiring target media data and media metadata corresponding to the target media data; obtaining a volume adjustment probability of playing according to a preset volume corresponding to the target media data, wherein the volume adjustment probability is obtained based on a regression processing result of the media metadata; and adjusting the preset volume based on the volume adjustment probability to determine the to-be-played volume of the target media data. According to the invention, the obtained volume to be played better conforms to the playing expectation, and the probability that the volume of the target media data is adjusted in the playing process is reduced, so that the purpose of improving the playing effect of the media data is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates to the technical field of computer processing, and in particular, to a volume adjustment method and device, electronic equipment, medium and product. BACKGROUND

[0002] In related technologies, when media data is played on a terminal, it is balanced to ensure the balance of the playback loudness on the terminal side. However, in the actual playback process, the actual volume of part of the media data still needs to be manually adjusted by the user, which affects the user's experience. SUMMARY

[0003] Therefore, the present disclosure provides a volume adjustment method, device, electronic equipment, medium and product to solve the problem that the volume needs to be manually adjusted when playing media data.

[0004] In a first aspect, the present disclosure provides a volume adjustment method, which comprises:

[0005] obtaining target media data and media metadata corresponding to the target media data;

[0006] obtaining a volume adjustment probability for playing the target media data according to a preset volume corresponding to the target media data, the volume adjustment probability being obtained based on a regression processing result of the media metadata;

[0007] adjusting the preset volume based on the volume adjustment probability to determine a to-be-played volume of the target media data.

[0008] In a second aspect, the present disclosure provides a volume adjustment device, which comprises:

[0009] a first obtaining module configured to obtain target media data and media metadata corresponding to the target media data;

[0010] a first processing module configured to obtain a volume adjustment probability for playing the target media data according to a preset volume corresponding to the target media data, the volume adjustment probability being obtained based on a regression processing result of the media metadata;

[0011] an adjusting module configured to adjust the preset volume based on the volume adjustment probability to determine a to-be-played volume of the target media data.

[0012] In a third aspect, the present disclosure provides an electronic equipment, which comprises a memory and a processor, the memory and the processor are communicatively connected, the memory stores computer instructions, and the processor executes the computer instructions to perform the volume adjustment method of the first aspect or any of the corresponding embodiments thereof.

[0013] In a fourth aspect, the present disclosure provides a computer-readable storage medium, having stored thereon computer instructions for causing a computer to execute the volume adjustment method of the first aspect or any of its possible implementation forms.

[0014] In a fifth aspect, the present disclosure provides a computer program product comprising computer instructions for causing a computer to execute the volume adjustment method of the first aspect or any of its possible implementation forms.

[0015] The volume adjustment method provided by the embodiment can determine the volume adjustment probability of the target media data played according to the corresponding preset volume based on the media metadata corresponding to the target media data, so that the process of volume adjustment is more targeted, and the obtained to-be-played volume is more consistent with the playing expectation, which can effectively reduce the probability of the volume of the target media data being adjusted in the playing process, thereby achieving the purpose of improving the playing effect of the media data. BRIEF DESCRIPTION OF DRAWINGS

[0016] In order to more clearly illustrate the technical solutions in the specific embodiments of the present disclosure or the prior art, the drawings needed in the specific embodiments or prior art description will be briefly introduced below. Obviously, the drawings in the following description are some embodiments of the present disclosure, and those skilled in the art can also obtain other drawings according to these drawings without creative labor.

[0017] Figure 1 is a flowchart of a volume adjustment method provided by an embodiment of the present disclosure;

[0018] Figure 2 is a flowchart of a training method of a target volume adjustment probability model provided by an embodiment of the present disclosure;

[0019] Figure 3 is a flowchart of another volume adjustment method provided by an embodiment of the present disclosure;

[0020] Figure 4 is a flowchart of still another volume adjustment method provided by an embodiment of the present disclosure;

[0021] Figure 5 is a structural block diagram of a volume adjustment device provided by an embodiment of the present disclosure;

[0022] Figure 6 is a hardware structure schematic diagram of an electronic device provided by an embodiment of the present disclosure. DETAILED DESCRIPTION

[0023] In order to make the purposes, technical solutions, and advantages of the embodiments of the present disclosure clearer, the technical solutions in the embodiments of the present disclosure will be described clearly and completely below with reference to the drawings in the embodiments of the present disclosure. Obviously, the described embodiments are only some but not all of the embodiments of the present disclosure. Based on the embodiments in the present disclosure, all other embodiments obtained by a person of ordinary skill in the art without creative work fall within the protection scope of the present disclosure.

[0024] In the related art, when media data is played on a terminal, it is balanced to ensure the balance of the playback loudness on the terminal side. However, in the actual playback process, the loudness corresponding to the playback volume of part of the media data does not meet the user's needs, and thus the user needs to manually increase or decrease the volume, thereby affecting the user's experience.

[0025] In view of this, the embodiments of the present disclosure provide a volume adjustment method, which can effectively reduce the number of times of volume adjustment of media data during playback, thereby achieving the purpose of improving the playback effect of media data.

[0026] According to the embodiments of the present disclosure, a volume adjustment method embodiment is provided. It should be noted that the steps shown in the flowchart of the drawings can be executed in a computer system such as a group of computer executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that here.

[0027] In the present embodiment, a volume adjustment method is provided, which can be used in electronic devices such as mobile phones, tablet computers, etc. Figure 1 The flowchart of the volume adjustment method according to the embodiments of the present disclosure is shown in FIG. 1, which includes the following steps: Figure 1 As shown in the figure, the flow includes the following steps:

[0028] Step S101, obtaining target media data and media metadata corresponding to the target media data.

[0029] The target media data can be understood as any media data that needs to be played. For example, the target media data can be audio or video that needs to be played.

[0030] In order to understand the related information of the target media data, the media metadata corresponding to the target media data is acquired at the same time the target media data is acquired, so as to determine the related media information of the target media data according to the media metadata, and then subsequent targeted processing can be carried out, and the reliability of volume adjustment is improved. The media metadata includes but is not limited to the following various media information of the target media data: corresponding media source loudness, loudness range, starting loudness of the loudness range, ending loudness of the loudness range, maximum instantaneous loudness, maximum short-time loudness, first low-frequency direct loudspeaker compensation gain (for example, 100hz), second low-frequency direct loudspeaker compensation gain (for example, 150hz), third low-frequency direct loudspeaker compensation gain (for example, 200hz), total playback time, and target loudness, etc.

[0031] In step S102, a volume adjustment probability of playing the target media data according to a preset volume is acquired.

[0032] The volume adjustment probability is obtained based on a regression processing result of the media metadata. The preset volume refers to an initial playback volume of the target media data before volume adjustment. The regression processing is a data analysis method for studying the relationship between two or more variables by establishing a mathematical model. By taking the volume adjustment probability as the dependent variable, the media metadata of the target media data as the independent variable, and the volume adjustment probability of playing the target media data according to different preset volumes as the dependent variable for regression processing, the relationship between the two can be fully explored, and based on the regression processing result of the media metadata, the volume demand corresponding to the target media data can be better evaluated, so that the probability of the target media data being played according to the preset volume and being adjusted can be analyzed based on the obtained regression processing result, and the situation of blindly adjusting the volume can be effectively avoided, and the rationality and accuracy of volume adjustment are improved.

[0033] In step S103, based on the volume adjustment probability, the preset volume is adjusted to determine the to-be-played volume of the target media data.

[0034] Based on the obtained volume adjustment probability, it can be determined whether to adjust the preset volume. If the probability is high, it indicates that the target media data needs to adjust the preset volume; if the probability is low, it indicates that the target media data does not need to adjust the preset volume.

[0035] Based on the volume adjustment probability, the volume is adjusted, so that the to-be-played volume determined after adjustment is more consistent with the playback expectation, and the probability of the target media data being adjusted during playback can be effectively reduced, thereby helping to improve the playback effect of the media data.

[0036] The volume adjustment method provided in the embodiment can determine the volume adjustment probability of the target media data based on the media metadata corresponding to the target media data, so that the volume adjustment process is more targeted, the obtained to-be-played volume is more in line with the playing expectation, the probability that the target media data is adjusted in the playing process can be effectively reduced, and the purpose of improving the playing effect of the media data is achieved.

[0037] In some optional embodiments, the volume adjustment probability is determined based on a preset target volume adjustment probability model, and the target volume adjustment probability model is trained based on regression processing of historical playing records of a plurality of media data samples. That is, a model is trained by performing regression processing on historical playing records of a large number of media data samples in advance, to learn the regression relationship between the media data under different preset volumes and the volume adjustment, and then obtain the target volume adjustment probability model that can be used to predict the volume adjustment probability, so that in actual application, the volume adjustment probability corresponding to the target media data can be determined through the pre-trained target volume adjustment probability model, the determination efficiency can be improved, the volume adjustment efficiency is promoted, the situation that the user manually adjusts the volume is reduced, and thus the volume adjustment performance of the system is improved.

[0038] In the embodiment, a training method of a target volume adjustment probability model is provided, which can be used in electronic devices such as mobile phones, tablet computers and the like, Figure 2 is a flowchart of the training method of the target volume adjustment probability model according to the embodiment of the disclosure, as shown in Figure 2 The flowchart includes the following steps:

[0039] In step S201, a plurality of media data samples and corresponding historical playing records are obtained.

[0040] The plurality of media data samples and the corresponding historical playing records are obtained, so that the volume adjustment of the corresponding media data samples in the past playing can be determined based on the historical playing records, and then the relationship between the media data under different playing volumes and the volume adjustment can be captured, so that the prediction accuracy of the volume adjustment probability can be improved in the subsequent process. The source of the plurality of media data samples can be local storage or cloud, which can be determined according to actual needs.

[0041] In some optional embodiments, the media data types corresponding to the plurality of media data samples include at least one type. For example, the media data types can include but are not limited to music, video, voice and the like. By using a plurality of types of media data samples for training, the model can learn the characteristics and rules of different types of media, so as to improve the accuracy of the volume adjustment probability prediction of different types of media, and meet the volume adjustment demand of the user for different types of media data.

[0042] In step S202, the loudness adjustment parameter and the volume adjustment probability parameter of the corresponding media data sample are determined based on the historical playback record.

[0043] According to the historical playback record, the historical playback situation of the media data sample can be determined, and then by analyzing the volume change in the playback record, the loudness adjustment parameter of the media data sample can be obtained, and the volume adjustment probability parameter corresponding to the volume adjustment situation under different playback volumes in the actual playback process can be determined.

[0044] The loudness adjustment parameter includes but is not limited to the following parameters: target loudness, final gain, limit compensation gain, dynamic range control compensation gain, final dynamic range control compression ratio, and maximum loudness range of the corresponding media data sample during playback. The volume adjustment probability parameter includes but is not limited to the first probability of the playback volume being increased and the second probability of the playback volume being decreased. In some optional implementation scenarios, the volume adjustment probability parameter can be obtained through online burying of the corresponding media data sample, thereby helping to improve the determination efficiency.

[0045] In some optional embodiments, the above step S202 includes:

[0046] In step a1, according to the historical playback record, the first playback times of the corresponding media data sample played according to the playback volume, the second playback times of the volume increased, and the third playback times of the volume decreased in the past are determined.

[0047] In step a2, according to the ratio of the second playback times to the first playback times, the first probability of the playback volume being increased in the past playback process of the media data sample is determined.

[0048] In step a3, according to the ratio of the third playback times to the first playback times, the second probability of the playback volume being decreased in the past playback process of the media data sample is determined.

[0049] In step a4, according to the first probability and the second probability, the volume adjustment probability parameter of the media data sample played according to the playback volume in the past is determined.

[0050] Specifically, the related volume adjustment record of the corresponding media data sample played in the past is extracted from the historical playback record, and then by analyzing the volume adjustment record, the first playback times of the media data sample played according to the playback volume, the second playback times of the volume increased, and the third playback times of the volume decreased can be counted respectively.

[0051] Dividing the second number of plays by the first number of plays can determine a first probability that the media data sample was played at an increased volume in the past at the play volume. Dividing the third number of plays by the first number of plays can determine a second probability that the media data sample was played at a decreased volume in the past at the play volume.

[0052] According to the first probability and the second probability, the volume adjustment of the media data sample played at the play volume in the past can be comprehensively analyzed, and then the volume adjustment probability parameter of the media data sample is obtained, and then subsequent model training is helpful for the model to learn the volume play preference of the media data. For example, the first probability and the second probability can be averaged or weighted averaged to comprehensively analyze and determine the volume adjustment probability parameter.

[0053] Step S203, based on the media metadata sample of the media data sample and the corresponding loudness adjustment parameter and the volume adjustment probability parameter, training the play volume adjustment probability model to obtain the target volume adjustment probability model.

[0054] Using the media metadata sample of all media data samples and the corresponding loudness adjustment parameter and the volume adjustment probability parameter, the play volume adjustment probability model is trained, and in the process of training, the parameters of the model are continuously adjusted to optimize the performance of the model, so that it can more accurately predict the volume adjustment probability, thereby obtaining the target volume adjustment probability model. The play volume adjustment probability model is constructed based on a regression model. The types of regression models include but are not limited to linear regression model, decision tree regression model, random forest regression model, support vector machine regression model, etc.

[0055] In some optional embodiments, the above step S203 includes:

[0056] Step b1, according to the media metadata sample of the media data sample, determining a first media parameter set of the media data sample in the past play process, so as to obtain an input parameter set corresponding to the media data sample in combination with the corresponding loudness adjustment parameter, the first probability and the second probability;

[0057] Step b2, inputting the input parameter set into the play volume adjustment probability model for regression processing to obtain an intermediate model;

[0058] Step b3, in response to the accuracy rate of the intermediate model being greater than or equal to a preset threshold, it is determined that the training is completed, and the intermediate model is taken as the target volume adjustment probability model.

[0059] Specifically, based on the media metadata sample of the media data sample, a first media parameter set is determined during its past playback process. The loudness adjustment parameters, the first probability, and the second probability are then combined with this first media parameter set to obtain the input parameter set corresponding to the media data sample. The first media parameter set may include the historical media source loudness, historical loudness range, the starting loudness of the historical loudness range, the ending loudness of the historical loudness range, the historical maximum instantaneous loudness, the historical maximum short-term loudness, the historical first low-frequency direct speaker compensation gain (e.g., 100 Hz), the historical second low-frequency direct speaker compensation gain (e.g., 150 Hz), the historical third low-frequency direct speaker compensation gain (e.g., 200 Hz), the historical total playback duration, and the historical target loudness, etc.

[0060] The input parameter set is fed into the playback volume adjustment probability model for regression processing. The purpose of regression processing is to predict the volume adjustment probability based on the input parameter set. By training on a large set of input parameters, the playback volume adjustment probability model can gradually learn the relationship between media metadata samples, loudness adjustment parameters, and volume adjustment probabilities, and continuously optimize the model's parameters to improve prediction accuracy. During training, the accuracy of intermediate models is continuously calculated. When the accuracy of the intermediate model is greater than or equal to a preset threshold, the model is considered to have completed training. At this point, the intermediate model is used as the target volume adjustment probability model and can be applied to actual volume adjustment scenarios to ensure that the target volume adjustment probability model has sufficient accuracy and reliability, and can provide effective guidance for actual volume adjustment.

[0061] The training method for the target volume adjustment probability model provided in this embodiment, by training a regression model based on historical playback records, enables the playback volume adjustment probability model to learn the relationship between media data and volume adjustment under different playback volumes during the training process. This allows the target volume adjustment probability model obtained after training to improve the accuracy of volume adjustment, thereby helping to reduce the occurrence of users manually adjusting the volume.

[0062] In some alternative implementations, media metadata includes various media information of the target media data across specified dimensions, thereby helping to improve the accuracy of determining the volume adjustment probability. For example, media metadata includes various media information of the target media data across 12 dimensions. This 12-dimensional media information includes: the corresponding media source loudness, loudness range, the starting loudness of the loudness range, the ending loudness of the loudness range, the maximum instantaneous loudness, the maximum short-term loudness, the 100Hz low-frequency direct speaker compensation gain, the 150Hz low-frequency direct speaker compensation gain, the 200Hz low-frequency direct speaker compensation gain, the total playback duration, the current playback duration, and the target loudness.

[0063] This embodiment provides a volume adjustment method that can be used in electronic devices such as mobile phones and tablets. Figure 3 This is a flowchart of a volume adjustment method according to an embodiment of the present disclosure, such as... Figure 3 As shown, the process includes the following steps:

[0064] Step S301: Obtain the target media data and the corresponding media metadata.

[0065] Step S302: Obtain the volume adjustment probability of playing according to the preset volume corresponding to the target media data.

[0066] Step S303: Adjust the preset volume based on the volume adjustment probability to determine the playback volume of the target media data.

[0067] In some optional implementations, step S303 above includes:

[0068] Step S3031: In response to the volume adjustment probability being greater than the first threshold, based on the volume adjustment records of the target media data when it was played in the past, determine the first target probability of the preset volume being increased and the second target probability of the volume being decreased.

[0069] The first threshold can be understood as the maximum probability value for determining that the preset volume does not need to be adjusted. When the volume adjustment probability is greater than the first threshold, it indicates that the volume will be adjusted when playing target media data at that preset volume. Therefore, when the volume adjustment probability is determined to be greater than the first threshold, based on the volume adjustment records, a first target probability of increasing the preset volume and a second target probability of decreasing it are determined respectively. This allows for subsequent adjustments to the preset volume based on the comparison between the first and second target probabilities, thereby improving the effectiveness of volume adjustment.

[0070] Step S3032: Based on the comparison result between the first target probability and the second target probability, determine the target adjustment strategy.

[0071] By comparing the first and second target probabilities, the adjustment direction can be clearly defined, making the resulting target adjustment strategy more reasonable and effective. For example, if the first target probability is greater than the second target probability, it means that during the past playback of the target media data, the probability of the preset volume being increased is greater than the probability of the preset volume being decreased. Therefore, when determining the subsequent target adjustment strategy, an adjustment strategy that increases the preset volume should be adopted. Conversely, if the first target probability is less than the second target probability, it means that during the historical playback of the target media data, the probability of the preset volume being increased is less than the probability of the preset volume being decreased. Therefore, when determining the subsequent target adjustment strategy, an adjustment strategy that decreases the preset volume should be adopted.

[0072] Step S3033: Adjust the preset volume according to the target adjustment strategy to determine the playback volume of the target media data.

[0073] After determining the target adjustment strategy, the preset volume is adjusted according to the target adjustment strategy so that the volume of the target media data to be played is more in line with the playback expectation. This can effectively reduce the probability that the volume of the target media data will be adjusted during playback, thereby helping to improve the playback effect of the media data.

[0074] In some optional implementations, step S303 above further includes:

[0075] Step S3034: In response to the volume adjustment probability being less than or equal to the first threshold, loudness equalization processing is performed on the target media data to determine the playback volume of the target media data.

[0076] In response to a volume adjustment probability less than or equal to a first threshold, indicating that the target media data has a low probability of volume adjustment when played at a preset volume during historical playback, loudness equalization is performed on the target media data to further optimize the audio effect of the target media data and obtain the playback volume of the target media data.

[0077] The volume adjustment method provided in this embodiment, after determining the probability of playing the target media data at a preset volume, uses different methods to determine the volume to be played of the target media data based on the comparison result of the volume adjustment probability and a first threshold. This makes the volume adjustment method more flexible and targeted, thereby effectively improving the playback quality of the target media data, reducing the probability of manual adjustment, and meeting the user's expectations.

[0078] In some optional examples, the process of determining the first target probability can be as follows: Based on the volume adjustment records, playback records where the target media data was played at a preset volume during historical playback can be selected. Then, based on these playback records, the number of times the preset volume was increased can be counted, thus obtaining the first adjustment quantity. The ratio between the first adjustment quantity and the first playback count is taken as the first target probability. Similarly, when determining the second target probability, based on the playback records, the number of times the preset volume was decreased can be counted, thus obtaining the second adjustment quantity. The ratio between the second adjustment quantity and the first playback count is taken as the second target probability.

[0079] In some optional implementations, the process of determining the target adjustment strategy based on the comparison between the first target probability and the second target probability includes: if the first difference between the first target probability and the second target probability is greater than a second threshold, then the target adjustment strategy is determined to be the first adjustment strategy. The second threshold can be understood as a threshold for determining whether volume increase processing is needed. The specific value of the second threshold can be set according to requirements; for example, the second threshold can be 0.4 or 0.5. When the first difference between the first target probability and the second target probability is greater than the second threshold, it indicates that during historical playback of the target media data, when played at a preset volume, the preset volume is mostly increased to alleviate the problem of insufficient volume provided by the preset volume affecting the auditory effect. Therefore, the first adjustment strategy for increasing the preset volume is selected as the target adjustment strategy so that subsequent volume adjustments can increase the preset volume to meet playback expectations, thereby improving the playback effect of the target media data.

[0080] In some alternative implementations, if the target adjustment strategy is the first adjustment strategy, the preset volume is adjusted according to the target adjustment strategy to determine the playback volume of the target media data, including: determining the maximum value of the dynamic range control curve according to the first adjustment strategy, and adjusting the preset volume based on the maximum value to obtain the playback volume of the target media data.

[0081] Specifically, the maximum value of the Dynamic Range Control (DRC) curve is determined according to the first adjustment strategy. Dynamic Range Control is a method of adjusting volume by compressing or expanding the dynamic range of an audio signal. To maximize the preset volume, the maximum value of the DRC curve is determined, and this maximum value is used to specifically adjust the preset volume, ensuring that the playback volume of the target media data meets playback expectations.

[0082] In some alternative implementations, the process of determining the target adjustment strategy based on the comparison between the first target probability and the second target probability further includes: if the first difference between the first target probability and the second target probability is less than or equal to a second threshold, then determining the target adjustment strategy based on the second difference between the second target probability and the first target probability. If the first difference between the first target probability and the second target probability is less than or equal to the second threshold, it indicates that during historical playback of the target media data, the preset volume was not frequently adjusted by increasing the volume. Therefore, to determine whether the preset volume was frequently adjusted by decreasing the volume during historical playback, the target adjustment strategy is determined based on the second difference between the second target probability and the first target probability, thereby improving the accuracy of determining the target adjustment strategy.

[0083] In some optional examples, if the second difference is greater than the third threshold, the target adjustment strategy is determined to be the second adjustment strategy, which is used to reduce the preset volume; if the second difference is less than or equal to the third threshold, the target adjustment strategy is determined to be the third adjustment strategy, which is used to compress the dynamic range control curve and adjust the preset volume based on the compressed dynamic range control curve.

[0084] Specifically, the third threshold can be understood as the threshold for determining whether volume reduction processing is needed. The specific value of the third threshold can be set according to requirements; for example, the third threshold can be 0.2 or 0.3. If the second difference is greater than the third threshold, it indicates that during the historical playback of the target media data, when played at the preset volume, the preset volume is mostly adjusted by reducing the volume to alleviate the problem of excessive volume affecting the listening effect. Therefore, the second adjustment strategy for reducing the preset volume is selected as the target adjustment strategy so that when the preset volume is adjusted according to the second adjustment strategy, the preset volume can be effectively reduced to meet the playback expectations. For example, the second adjustment strategy can refer to reducing the preset volume by adjusting the limit makeup gain or DRC compensation gain. Preferably, the limit makeup gain or DRC compensation gain can be adjusted by setting it to zero or adjusting the weight to achieve the purpose of reducing the preset volume, avoiding the situation where the audio signal is too strong or distorted due to excessive compensation gain, thereby improving the stability of volume adjustment.

[0085] If the second difference is less than or equal to the third threshold, it indicates that during the historical playback of the target media data, the preset volume may have increased or decreased. Therefore, the third adjustment strategy is selected as the target adjustment strategy to ensure a more stable playback volume when adjusting the preset volume according to the third adjustment strategy. Specifically, the third adjustment strategy compresses the dynamic range control curve and adjusts the preset volume based on the compressed dynamic range control curve, effectively preventing volume fluctuations during adjustment. This can improve the listening experience of the target media data to a certain extent, making the sound playback clearer and more comfortable, thus enhancing the playback quality of the target media data.

[0086] Determining the target adjustment strategy based on the comparison between the second difference and the third threshold allows for more flexible adaptation to different playback situations, improving the quality and effect of media data playback and meeting playback expectations.

[0087] As another or more specific application embodiments of this disclosure, the process of determining the playback volume of target media data is as follows: Figure 4 As shown, the specific process is as follows:

[0088] First, acquire the target media data and the corresponding media metadata, and input the media metadata into the preset target volume adjustment probability model to determine the volume adjustment probability of the target media data playing at the preset volume.

[0089] Secondly, given the determined volume adjustment probability, the volume adjustment probability is compared with a first threshold to determine whether the volume adjustment probability is greater than the first threshold.

[0090] If the volume adjustment probability is less than or equal to the first threshold, then loudness equalization processing is performed on the target media data to determine the playback volume of the target media data.

[0091] If the volume adjustment probability is greater than a first threshold, then based on the volume adjustment records during historical playback of target media data, a first target probability of increasing the preset volume and a second target probability of decreasing the preset volume are determined. It is then determined whether a first difference between the first and second target probabilities is greater than a second threshold. If the first difference is greater than the second threshold, the target adjustment strategy is determined to be the first adjustment strategy. According to the first adjustment strategy, the maximum value of the dynamic range control curve is determined, and the preset volume is adjusted based on this maximum value. The adjusted preset volume is then used to update the parameters for processing the target media data, and the updated parameters are used for loudness equalization processing to determine the playback volume of the target media data.

[0092] If the first difference between the first target probability and the second target probability is less than or equal to the second threshold, then it is determined whether the second difference between the second target probability and the first target probability is greater than the third threshold. If the second difference between the second target probability and the first target probability is greater than the third threshold, then the target adjustment strategy is determined to be the second adjustment strategy. According to the second adjustment strategy, the preset volume is reduced by adjusting the limit makeup gain or DRC compensation gain. Then, the adjusted preset volume is used to update the parameters for processing the target media data, and the updated parameters are used for loudness equalization processing to determine the playback volume of the target media data.

[0093] If the second difference between the second target probability and the first target probability is less than or equal to the third threshold, then the target adjustment strategy is determined to be the third adjustment strategy. According to the third adjustment strategy, the dynamic range control curve is compressed, and the preset volume is adjusted based on the compressed dynamic range control curve. Then, the adjusted preset volume is used to update the parameters for processing the target media data, and the updated parameters are used for loudness equalization processing, thereby determining the playback volume of the target media data.

[0094] Adjusting volume in the above way makes the process more scientific and targeted, resulting in a higher volume of the target media data that is more in line with expectations, and thus helps to improve the playback quality and effect of the target media data.

[0095] As another specific application embodiment of this disclosure, the target adjustment strategy can be obtained through a preset volume adjustment module. The process of processing target media data includes a parameter preparation stage and a streaming processing stage. The preparation stage includes determining loudness gain, determining the parameters of the DRC curve, and determining loudness compensation gain. In the preparation stage, based on the target media data and the corresponding media metadata, the volume adjustment probability of the target media data being played at a preset volume is determined in advance. In response to a volume adjustment probability greater than a first threshold, based on the target adjustment strategy determined by the volume adjustment module, the loudness gain, the parameters for determining the DRC curve, and the loudness compensation gain determined in the preparation stage are adjusted specifically to ensure that the loudness gain processing of the target media data in the streaming processing stage is more reasonable and the auditory effect better meets expectations.

[0096] In the streaming processing stage, loudness gain is applied to the target media data to obtain the first media data; the DRC curve is obtained using the parameters of the DRC curve, and then the first media data is processed by dynamic range control using the DRC curve to obtain the second media data; loudness compensation is applied to the second media data using loudness compensation gain to obtain the third media data; finally, peak limiting is applied to the third media data to obtain the target media data.

[0097] This embodiment also provides a volume adjustment device for implementing the above embodiments and preferred embodiments; details already described will not be repeated. As used below, the term "module" can refer to a combination of software and / or hardware that performs a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, hardware implementations, or a combination of software and hardware, are also possible and contemplated.

[0098] This embodiment provides a volume adjustment device, such as... Figure 5 As shown, it includes:

[0099] The first acquisition module 501 is used to acquire target media data and media metadata corresponding to the target media data;

[0100] The first processing module 502 is used to obtain the volume adjustment probability of playing according to the preset volume corresponding to the target media data. The volume adjustment probability is obtained based on the regression processing result of the media metadata.

[0101] The adjustment module 503 is used to adjust the preset volume and the volume of the target media data to be played based on the volume adjustment probability.

[0102] In some alternative implementations, the volume adjustment probability is determined based on a preset target volume adjustment probability model, which is trained by regression processing of historical playback records of multiple media data samples.

[0103] In some alternative implementations, the training apparatus for the target volume adjustment probability model includes:

[0104] The second acquisition module is used to acquire multiple media data samples and corresponding historical playback records;

[0105] The second processing module is used to determine the loudness adjustment parameters and volume adjustment probability parameters of the corresponding media data samples based on historical playback records.

[0106] The training module is used to train the playback volume adjustment probability model based on the media metadata samples of the media data samples and the corresponding loudness adjustment parameters and volume adjustment probability parameters, so as to obtain the target volume adjustment probability model. The playback volume adjustment probability model is constructed based on the regression model.

[0107] In some alternative implementations, the second processing module includes:

[0108] The first processing unit is used to determine, based on historical playback records, the first number of times the corresponding media data sample was played at the playback volume, the second number of times the volume was increased, and the third number of times the volume was decreased in the past.

[0109] The second processing unit is used to determine the first probability that the playback volume of the media data sample was increased during past playback based on the ratio of the second playback count to the first playback count.

[0110] The third processing unit is used to determine a second probability that the playback volume of the media data sample was reduced during past playback based on the ratio of the third playback count to the first playback count.

[0111] The fourth processing unit is used to determine the volume adjustment probability parameter of the media data sample being played according to the playback volume in the past, based on the first probability and the second probability.

[0112] In some alternative implementations, the training module includes:

[0113] The fifth processing unit is used to determine the first media parameter set of the media data sample in the past playback process based on the media metadata sample of the media data sample, so as to combine the corresponding loudness adjustment parameter, the first probability and the second probability to obtain the input parameter set corresponding to the media data sample.

[0114] The sixth processing unit is used to input the input parameter set into the playback volume adjustment probability model for regression processing to obtain an intermediate model;

[0115] The seventh processing unit is used to determine that training is complete when the accuracy of the intermediate model is greater than or equal to a preset threshold, and to use the intermediate model as the target volume adjustment probability model.

[0116] In some alternative implementations, the media data types corresponding to the multiple media data samples include at least one.

[0117] In some alternative implementations, media metadata includes various media information of the target media data under a specified dimension.

[0118] In some alternative implementations, the adjustment module 503 includes:

[0119] The eighth processing unit is used to respond to the volume adjustment probability being greater than the first threshold by determining, based on the volume adjustment records of the target media data when it was played in the past, a first target probability of the preset volume being increased and a second target probability of the volume being decreased.

[0120] The strategy determination unit is used to determine the target adjustment strategy based on the comparison result between the first target probability and the second target probability;

[0121] The volume adjustment unit is used to adjust the preset volume according to the target adjustment strategy in order to determine the playback volume of the target media data.

[0122] In some optional implementations, the strategy determination unit includes:

[0123] The first execution unit is configured to determine the target adjustment strategy as the first adjustment strategy if the first difference between the first target probability and the second target probability is greater than the second threshold. The first adjustment strategy is used to increase the preset volume.

[0124] In some alternative implementations, the volume control unit includes:

[0125] The second execution unit is used to determine the maximum value of the dynamic range control curve according to the first adjustment strategy, and adjust the preset volume based on the maximum value to obtain the playback volume of the target media data.

[0126] In some optional implementations, the strategy determination unit further includes:

[0127] The third execution unit is used to determine a target adjustment strategy based on the second difference between the second target probability and the first target probability if the first difference between the first target probability and the second target probability is less than or equal to the second threshold.

[0128] In some alternative implementations, the third execution unit includes:

[0129] The first determining unit is used to determine the target adjustment strategy as the second adjustment strategy if the second difference is greater than the third threshold. The second adjustment strategy is used to reduce the preset volume.

[0130] The second determining unit is used to determine the target adjustment strategy as the third adjustment strategy if the second difference is less than or equal to the third threshold. The third adjustment strategy is used to compress the dynamic range control curve and adjust the preset volume based on the compressed dynamic range control curve.

[0131] In some alternative implementations, the adjustment module 503 further includes:

[0132] The ninth processing module is used to perform loudness equalization processing on the target media data in response to the volume adjustment probability being less than or equal to the first threshold, and to determine the playback volume of the target media data.

[0133] Further functional descriptions of the above modules and units are the same as those in the corresponding embodiments described above, and will not be repeated here.

[0134] In this embodiment, the volume adjustment device is presented in the form of a functional unit. Here, a unit refers to an ASIC (Application Specific Integrated Circuit) circuit, a processor and memory that execute one or more software or fixed programs, and / or other devices that can provide the above functions.

[0135] This disclosure also provides an electronic device having the above-described features. Figure 5 The volume control shown.

[0136] Please see Figure 6 , Figure 6 This is a schematic diagram of the structure of an electronic device provided in an optional embodiment of this disclosure, such as... Figure 6 As shown, the electronic device includes one or more processors 10, memory 20, and interfaces for connecting the components, including high-speed interfaces and low-speed interfaces. The components communicate with each other via different buses and can be mounted on a common motherboard or otherwise as required. The processors can process instructions executed within the electronic device, including instructions stored in or on memory to display graphical information of a GUI on external input / output devices (such as display devices coupled to the interfaces). In some alternative implementations, multiple processors and / or multiple buses can be used with multiple memories and multiple memory modules, if desired. Similarly, multiple electronic devices can be connected, each providing some of the necessary operations (e.g., as a server array, a group of blade servers, or a multiprocessor system). Figure 6 Take a processor 10 as an example.

[0137] Processor 10 may be a central processing unit, a network processor, or a combination thereof. Processor 10 may further include a hardware chip. The hardware chip may be an application-specific integrated circuit (ASIC), a programmable logic device (PLD), or a combination thereof. The programmable logic device may be a complex programmable logic device (CAMP), a field-programmable gate array (FPGA), a general-purpose array logic (GPA), or any combination thereof.

[0138] The memory 20 stores instructions executable by at least one processor 10 to cause at least one processor 10 to perform the method shown in the above embodiments.

[0139] The memory 20 may include a program storage area and a data storage area. The program storage area may store the operating system and applications required for at least one function; the data storage area may store data created based on the use of the electronic device. Furthermore, the memory 20 may include high-speed random access memory and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some alternative embodiments, the memory 20 may optionally include memory remotely located relative to the processor 10, and these remote memories may be connected to the electronic device via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.

[0140] The memory 20 may include volatile memory, such as random access memory; the memory may also include non-volatile memory, such as flash memory, hard disk or solid-state drive; the memory 20 may also include a combination of the above types of memory.

[0141] The electronic device also includes an input device 30 and an output device 40. The processor 10, memory 20, input device 30, and output device 40 can be connected via a bus or other means. Figure 6 Taking the example of a connection between China and Israel via a bus.

[0142] Input device 30 can receive input numerical or character information, and generate key signal inputs related to user settings and function control of the electronic device, such as a touch screen, keypad, mouse, trackpad, touchpad, joystick, one or more mouse buttons, trackball, joystick, etc. Output device 40 may include display devices, auxiliary lighting devices (e.g., LEDs), and haptic feedback devices (e.g., vibration motors). The aforementioned display devices include, but are not limited to, liquid crystal displays, light-emitting diodes, displays, and plasma displays. In some alternative embodiments, the display device may be a touch screen.

[0143] This disclosure also provides a computer-readable storage medium in which the methods described in this disclosure can be implemented in hardware or firmware, or implemented as recordable on a storage medium, or implemented as computer code originally stored on a remote storage medium or a non-transitory machine-readable storage medium and subsequently stored on a local storage medium after being downloaded over a network. Thus, the methods described herein can be processed by software stored on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. The storage medium may be a magnetic disk, optical disk, read-only memory, random access memory, flash memory, hard disk, or solid-state drive, etc.; further, the storage medium may also include combinations of the above types of memory. It is understood that computers, processors, microprocessor controllers, or programmable hardware include storage components capable of storing or receiving software or computer code that, when accessed and executed by the computer, processor, or hardware, implements the methods shown in the above embodiments.

[0144] A portion of this invention can be applied as a computer program product, such as computer program instructions, which, when executed by a computer, can invoke or provide the methods and / or technical solutions according to the invention through the operation of the computer. Those skilled in the art will understand that the forms in which computer program instructions exist in a computer-readable medium include, but are not limited to, source files, executable files, installation package files, etc. Correspondingly, the ways in which computer program instructions are executed by a computer include, but are not limited to: the computer directly executing the instructions, or the computer compiling the instructions and then executing the corresponding compiled program, or the computer reading and executing the instructions, or the computer reading and installing the instructions and then executing the corresponding installed program. Here, the computer-readable medium can be any available computer-readable storage medium or communication medium accessible to a computer.

[0145] It is understood that before using the technical solutions disclosed in the various embodiments of this disclosure, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this disclosure in an appropriate manner in accordance with relevant laws and regulations, and user authorization should be obtained.

[0146] For example, upon receiving a user's active request, a prompt message is sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the software or hardware, such as the electronic device, application, server, or storage medium performing the operations of this disclosed technical solution, based on the prompt message.

[0147] As an optional but non-limiting implementation, in response to a user's active request, sending a prompt message to the user can be done via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose "agree" or "disagree" to provide personal information to the electronic device.

[0148] It is understood that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of this disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of this disclosure.

[0149] Although embodiments of the present disclosure have been described in conjunction with the accompanying drawings, those skilled in the art can make various modifications and variations without departing from the spirit and scope of the present disclosure, and such modifications and variations all fall within the scope defined by the appended claims.

Claims

1. A volume adjustment method, characterized in that, The method includes: Acquire the target media data and the media metadata corresponding to the target media data; Obtain the volume adjustment probability of playing according to the preset volume corresponding to the target media data, wherein the volume adjustment probability is obtained based on the regression processing result of the media metadata; Based on the volume adjustment probability, the preset volume is adjusted to determine the playback volume of the target media data.

2. The method according to claim 1, characterized in that, The volume adjustment probability is determined based on a preset target volume adjustment probability model, which is trained by regression processing of historical playback records of multiple media data samples.

3. The method according to claim 2, characterized in that, The training process of the target volume adjustment probability model includes: Acquire multiple media data samples and their corresponding historical playback records; Based on the historical playback records, the loudness adjustment parameters and volume adjustment probability parameters of the corresponding media data samples are determined; Based on the media metadata samples of the media data samples and the corresponding loudness adjustment parameters and volume adjustment probability parameters, a playback volume adjustment probability model is trained to obtain the target volume adjustment probability model, wherein the playback volume adjustment probability model is constructed based on a regression model.

4. The method according to claim 3, characterized in that, The step of determining the volume adjustment probability parameter for the corresponding media data sample based on the historical playback records includes: Based on the historical playback records, determine the first number of times the corresponding media data sample was played at the playback volume, the second number of times the volume was increased, and the third number of times the volume was decreased in the past. Based on the ratio of the second playback count to the first playback count, a first probability is determined that the playback volume of the media data sample was increased during past playback. Based on the ratio of the third playback count to the first playback count, a second probability is determined that the playback volume of the media data sample was reduced during past playback. Based on the first probability and the second probability, determine the volume adjustment probability parameter of the media data sample being played according to the playback volume in the past.

5. The method according to claim 4, characterized in that, The target volume adjustment probability model is obtained by training a playback volume adjustment probability model based on the media metadata sample, the corresponding loudness adjustment parameters, and the corresponding volume adjustment probability parameters, using the media data sample as the basis. Based on the media metadata sample of the media data sample, a first media parameter set of the media data sample during past playback is determined, and combined with the corresponding loudness adjustment parameter, the first probability and the second probability, the input parameter set corresponding to the media data sample is obtained; The input parameter set is input into the playback volume adjustment probability model for regression processing to obtain an intermediate model; If the accuracy of the intermediate model is greater than or equal to a preset threshold, then the training is determined to be complete, and the intermediate model is used as the target volume adjustment probability model.

6. The method according to claim 3, characterized in that, The media data types corresponding to the multiple media data samples include at least one type.

7. The method according to claim 1, characterized in that, The media metadata includes various media information of the target media data under a specified dimension.

8. The method according to claim 1, characterized in that, The step of adjusting the preset volume based on the volume adjustment probability to determine the playback volume of the target media data includes: In response to the volume adjustment probability being greater than a first threshold, based on the volume adjustment records of the target media data when it was played in the past, a first target probability of the preset volume being increased and a second target probability of the preset volume being decreased are determined respectively. Based on the comparison between the first target probability and the second target probability, a target adjustment strategy is determined; Adjust the preset volume according to the target adjustment strategy to determine the playback volume of the target media data.

9. The method according to claim 8, characterized in that, The step of determining the target adjustment strategy based on the comparison result between the first target probability and the second target probability includes: If the first difference between the first target probability and the second target probability is greater than the second threshold, then the target adjustment strategy is determined to be the first adjustment strategy, which is used to increase the preset volume.

10. The method according to claim 9, characterized in that, The step of adjusting the preset volume according to the target adjustment strategy and determining the playback volume of the target media data includes: Based on the first adjustment strategy, the maximum value of the dynamic range control curve is determined, and the preset volume is adjusted based on the maximum value to obtain the playback volume of the target media data.

11. The method according to claim 9, characterized in that, The step of determining the target adjustment strategy based on the comparison result between the first target probability and the second target probability further includes: If the first difference between the first target probability and the second target probability is less than or equal to the second threshold, then a target adjustment strategy is determined based on the second difference between the second target probability and the first target probability.

12. The method according to claim 11, characterized in that, The step of determining the target adjustment strategy based on the difference between the second target probability and the first target probability includes: If the second difference is greater than the third threshold, then the target adjustment strategy is determined to be the second adjustment strategy, which is used to reduce the preset volume. If the second difference is less than or equal to the third threshold, the target adjustment strategy is determined to be the third adjustment strategy. The third adjustment strategy is used to compress the dynamic range control curve and adjust the preset volume based on the compressed dynamic range control curve.

13. The method according to claim 8, characterized in that, The step of adjusting the preset volume based on the volume adjustment probability to determine the playback volume of the target media data further includes: In response to the volume adjustment probability being less than or equal to the first threshold, loudness equalization processing is performed on the target media data to determine the playback volume of the target media data.

14. A volume control device, characterized in that, The device includes: The first acquisition module is used to acquire target media data and media metadata corresponding to the target media data; The first processing module is used to obtain the volume adjustment probability of playing according to the preset volume corresponding to the target media data, wherein the volume adjustment probability is obtained based on the regression processing result of the media metadata; An adjustment module is used to adjust the preset volume and the playback volume of the target media data based on the volume adjustment probability.

15. An electronic device, characterized in that, include: A memory and a processor are communicatively connected, the memory storing computer instructions, and the processor executing the computer instructions to perform the volume adjustment method according to any one of claims 1 to 13.

16. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions for causing the computer to perform the volume adjustment method according to any one of claims 1 to 13.

17. A computer program product, characterized in that, Includes computer instructions for causing a computer to perform the volume adjustment method according to any one of claims 1 to 13.