A multi-source audio quality uniform optimization method and system
By optimizing the format, sampling rate, and loudness of multi-source audio through an audio optimization terminal, the problem of inconsistent listening experience caused by differences in parameters of different audio content is solved, thereby improving the uniformity of audio and the listening experience.
Patent Information
- Application Number
- CN202511366650.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-24
- Publication Date
- 2025-12-09
- Estimated Expiration
- 2045-09-24
AI Technical Summary
The audio content produced by different audio content production units has significantly different parameters, resulting in a poor listening experience for listeners on the same distribution platform.
The audio optimization terminal optimizes the format, sampling rate, and loudness of multiple audio sources to ensure consistency among the audio segments before splicing them together to generate a unified audio for playback.
It enhances the uniformity of audio and the listening experience for the audience, achieving unified optimization of multi-source audio.
Smart Images

Figure CN120872280B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of audio processing, in particular to a multi-source audio quality unified optimization method and system. BACKGROUND
[0002] Currently, the units of audio content manufacturing and distribution may not be the same; and the parameters of the audio content produced by different audio content manufacturing units are quite different (for example, there are great differences in audio format, sampling rate, volume and other parameters), which leads to poor listening experience of the audience when the multi-source audio content from different audio content manufacturing units is spliced and distributed on the same distribution platform. Therefore, there is an urgent need for a technical solution that can uniformly optimize multi-source audio with large parameter differences. SUMMARY
[0003] The main purpose of the present application is to provide a multi-source audio quality unified optimization method and system, which aims to solve the problem of the urgent need for a technical solution that can uniformly optimize multi-source audio with large parameter differences.
[0004] The technical solution provided by the present application is:
[0005] A multi-source audio quality unified optimization method applied to a multi-source audio quality unified optimization system; the system comprises an audio optimization terminal; the method comprises:
[0006] The audio optimization terminal acquires a plurality of audio segments that need to be spliced and mixed, and marks them as to-be-optimized audio segments;
[0007] The audio optimization terminal determines whether format optimization is needed;
[0008] If format optimization is needed, the audio optimization terminal determines a target format, and marks the to-be-optimized audio segments whose audio format is inconsistent with the target format as to-be-converted audio segments, and converts the audio format of the to-be-converted audio segments to the target format;
[0009] When the format optimization is performed or format optimization is not needed, the audio optimization terminal determines whether sampling rate optimization is needed;
[0010] If sampling rate optimization is needed, the audio optimization terminal determines a target sampling rate based on each to-be-optimized audio segment, and performs resampling processing on the to-be-optimized audio segments based on the target sampling rate, so that the sampling rates of the to-be-optimized audio segments are consistent;
[0011] When the sampling rate optimization is performed or sampling rate optimization is not needed, the audio optimization terminal determines whether loudness optimization is needed;
[0012] If loudness optimization is required, the audio optimization terminal obtains the average loudness value corresponding to each audio segment to be optimized, and performs loudness optimization on the audio segment to be optimized based on the average loudness value;
[0013] When loudness optimization is completed, or when loudness optimization is not required, the audio optimization terminal splices the audio segments to be optimized end to end in a preset order to obtain the audio to be played.
[0014] Preferably, if sampling rate optimization is required, the audio optimization terminal determines a target sampling rate based on each audio segment to be optimized, and performs resampling processing on the audio segments to be optimized based on the target sampling rate to ensure that the sampling rates of each audio segment to be optimized are consistent, including:
[0015] The audio optimization terminal acquires the sampling rate of each audio segment to be optimized and marks it as the original sampling rate;
[0016] The audio optimization terminal uses the most frequent original sampling rate among all the original sampling rates as the target sampling rate.
[0017] The audio optimization terminal marks audio segments to be optimized with an original sampling rate higher than the target sampling rate as first audio segments, and audio segments to be optimized with an original sampling rate lower than the target sampling rate as second audio segments;
[0018] The audio optimization terminal performs downsampling processing on the first audio segment to adjust the sampling rate of the first audio segment to the target sampling rate;
[0019] The audio optimization terminal performs upsampling processing on the second audio segment to adjust the sampling rate of the second audio segment to the target sampling rate.
[0020] Preferably, if loudness optimization needs to be performed, the audio optimization terminal obtains the average loudness value corresponding to each audio segment to be optimized, and performs loudness optimization on the audio segment to be optimized based on the average loudness value, including:
[0021] The audio optimization terminal obtains the average loudness value corresponding to each audio segment to be optimized, and obtains the preset standard loudness value;
[0022] The audio optimization terminal marks the audio segment to be optimized as the third audio segment, which is greater than the standard loudness value and whose difference from the standard loudness value is greater than a first preset value.
[0023] The audio optimization terminal marks the audio segment to be optimized as the fourth audio segment, which is less than the standard loudness value and whose absolute value of the difference from the standard loudness value is greater than the first preset value.
[0024] The audio optimization terminal obtains a difference value between the standard loudness value and an average loudness value of the third audio segment, and takes the difference value as a first loudness difference value, performs loudness attenuation on the third audio segment, and the attenuation amount is the first loudness difference value;
[0025] The audio optimization terminal obtains a difference value between the standard loudness value and an average loudness value of the fourth audio segment, and takes the difference value as a second loudness difference value, performs loudness gain on the fourth audio segment, and the gain amount is the second loudness difference value.
[0026] Preferably, the system further comprises an audio distribution terminal in communication connection with the audio optimization terminal, and a user terminal in communication connection with the audio distribution terminal; when the loudness optimization is performed or the loudness optimization is not needed to be performed, the audio optimization terminal splices the to-be-optimized audio segments at the beginning and the end in a preset order to obtain to-be-played audio, and then comprises:
[0027] The audio optimization terminal sends the to-be-played audio to the audio distribution terminal;
[0028] The audio distribution terminal sends the to-be-played audio to the user terminal;
[0029] The user terminal obtains a historical running log of a current user operation on the user terminal, and determines a target volume based on the historical running log;
[0030] The user terminal plays the to-be-played audio, and the playing volume is the target volume.
[0031] Preferably, the historical running log comprises an average playing volume corresponding to a use of the user terminal by the user in different time periods; the user terminal obtains a historical running log of a current user operation on the user terminal, and determines a target volume based on the historical running log, comprising:
[0032] The user terminal marks a time period in which a current time falls as a target time period;
[0033] The user terminal obtains an average playing volume corresponding to a use of the user terminal by the user in the target time period based on the historical running log, and takes the average playing volume as the target volume.
[0034] Preferably, the user terminal plays the to-be-played audio, and the playing volume is the target volume, and then comprises:
[0035] The user terminal judges whether an operation instruction for volume adjustment by the user is received within a first preset time length since the to-be-played audio starts to be played;
[0036] If yes, the user terminal marks the played audio as adjusted audio, obtains the volume difference between the volume after the volume adjustment and the volume before the volume adjustment, and marks the volume difference as the volume adjustment value corresponding to the adjusted audio;
[0037] The user terminal determines whether the number of the adjusted audios in the first preset number of the played audios is greater than or equal to a second preset number, wherein the second preset number is less than the first preset number;
[0038] If the number of the adjusted audios in the first preset number of the played audios is greater than or equal to the second preset number, the user terminal determines whether a first condition or a second condition is met, wherein the first condition is that the volume adjustment values corresponding to the adjusted audios in the first preset number of the played audios are all positive values and are all greater than the preset volume value, and the second condition is that the volume adjustment values corresponding to the adjusted audios in the first preset number of the played audios are all negative values and are all less than the opposite value of the preset volume value;
[0039] If the first condition is met, the user terminal increases the target volume, and the increasing amount is the preset volume value;
[0040] If the second condition is met, the user terminal decreases the target volume, and the decreasing amount is the preset volume value.
[0041] Preferably, the user terminal further comprises a microphone capable of collecting sound signals; the user terminal plays the audio to be played, and the playing volume is the target volume, and further comprises:
[0042] Before the user terminal starts to play the audio to be played, the user terminal divides the audio to be played into a plurality of sequentially adjacent sub-audio segments, and marks the first sub-audio segment, obtains the average loudness value of each first sub-audio segment, and marks the first loudness value;
[0043] The user terminal marks the first sub-audio segment with the first loudness value less than a second preset value as the target audio segment;
[0044] After the user terminal starts to play the audio to be played, the user terminal collects the sound signals of the surrounding environment during each time the target audio segment is played through the microphone, and obtains the average loudness value of the sound signals of the surrounding environment;
[0045] When the average loudness values of the sound signals of the surrounding environment in the past second preset time period are all less than a third preset value, and the current time is within a preset time period, the user terminal determines that the user is in a resting state;
[0046] When the user is in the resting state, the user terminal decreases the playing volume of the audio to be played by 1 every third preset time length.
[0047] Preferably, the user terminal further comprises a microphone capable of collecting sound signals; the user terminal plays the audio to be played at the target volume, and further comprises:
[0048] When the user terminal plays the audio to be played through the earphone and the current time is in a preset time period, the user terminal generates a monitoring instruction;
[0049] The user terminal collects the sound signals of the surrounding environment through the microphone based on the monitoring instruction and marks them as environmental sound signals;
[0050] The user terminal divides the environmental sound signals in the past fourth preset time length into a plurality of sequentially adjacent sub-audio segments and marks them as second sub-audio segments, obtains the average loudness value of each second sub-audio segment and marks it as a second loudness value;
[0051] The user terminal marks the second loudness value meeting a third condition in the past fourth preset time length as a first target value, wherein the third condition is that the two second loudness values adjacent to the first target value are both less than the first target value;
[0052] When the number of first target values is at least 2 and the interval time length between the collection time of adjacent two first target values is less than a fifth preset time length, the user terminal determines the second target value and the third target value corresponding to each first target value, wherein the second target value corresponding to the first target value is the second loudness value earlier than the first target value and equal to a fourth preset value, and the third target value corresponding to the first target value is the second loudness value later than the first target value and equal to the fourth preset value;
[0053] The user terminal marks the second loudness value between the collection time of the first target value and the corresponding second target value as the fourth target value corresponding to the first target value, and marks the second loudness value between the collection time of the first target value and the corresponding third target value as the fifth target value corresponding to the first target value;
[0054] The user terminal calculates the first slope and the second slope corresponding to each first target value:
[0055] ,
[0056] ,
[0057] In the formula, is the first slope corresponding to the jth first target value; is a second slope corresponding to the jth first target value; j is a positive integer, and 1≤j≤J, J is the number of first target values; is an (o+1)th third target value corresponding to the jth first target value, o is a positive integer, and 1≤o≤ , is the number of third target values corresponding to the jth first target value; is the duration of the second sub-audio segment; is a (q+1)th fifth target value corresponding to the jth first target value, q is a positive integer, and 1≤q≤ , is the number of fifth target values corresponding to the jth first target value;
[0058] When the standard deviation between the first slopes corresponding to each first target value, and the standard deviation between the second slopes corresponding to each first target value are all less than a preset threshold, the user terminal infers that the user is in a resting state;
[0059] When the user is in a resting state, the user terminal reduces the playback volume of the currently playing audio by 1 every third preset duration.
[0060] The application also proposes a multi-source audio quality unified optimization system, which applies a multi-source audio quality unified optimization method; the system comprises an audio optimization terminal.
[0061] The above technical solution can achieve the following beneficial effects:
[0062] The multi-source audio quality unified optimization method proposed by the application can uniformly optimize multi-source audio with large differences in parameters; first, a plurality of audio segments that need to be optimized are marked as to-be-optimized audio segments, then the to-be-optimized audio is sequentially optimized in format, sampling rate and loudness, and then the audio after the above optimization is spliced to obtain to-be-played audio, which can be sent to the user for listening, and the listening experience of the to-be-played audio after optimization is better, and the uniformity of the audio is better, thereby improving the user experience. BRIEF DESCRIPTION OF DRAWINGS
[0063] In order to more clearly illustrate the technical solutions in the embodiments of the application or the prior art, the following will briefly introduce the drawings needed to be used in the embodiments or prior art description. Obviously, the drawings in the following description are only some embodiments of the application, and for those skilled in the art, other drawings can be obtained from the structures shown in the drawings without creative labor.
[0064] Figure 1 is a flow step diagram of a first embodiment of a multi-source audio quality unified optimization method proposed by the application. Detailed Implementation
[0065] It should be understood that the specific embodiments described herein are for illustrative purposes only and are not intended to limit the scope of the invention.
[0066] This invention proposes a method and system for unified optimization of multi-source audio quality.
[0067] As attached Figure 1 As shown, in the first embodiment of the multi-source audio quality unified optimization method proposed in this invention, this embodiment is applied to a multi-source audio quality unified optimization system; the system includes an audio optimization terminal (e.g., a computer terminal); this embodiment includes the following steps:
[0068] Step S110: The audio optimization terminal acquires multiple audio segments that need to be spliced and mixed, and marks them as audio segments to be optimized.
[0069] Specifically, the audio segments to be optimized here are audio segments from multiple sources, and also audio segments that need to be optimized uniformly.
[0070] Step S120: The audio optimization terminal determines whether format optimization needs to be performed.
[0071] Specifically, there are currently many audio formats (such as MP3, WMA, WAV, etc.). When the formats of multiple audio files to be optimized are not unified, format optimization is required.
[0072] Step S130: If format optimization is required, the audio optimization terminal determines the target format (e.g., WAV), marks the audio segments to be optimized that are inconsistent with the target format as audio segments to be converted, and converts the audio format of the audio segments to be converted to the target format.
[0073] Step S140: When format optimization is completed, or when format optimization is not required, the audio optimization terminal determines whether sampling rate optimization needs to be performed.
[0074] Specifically, the mainstream audio sampling rates currently include 11025Hz, 22050Hz, 24000Hz, 44100Hz, and 48000Hz. When the sampling rates of the audio to be optimized are not the same, sampling rate optimization needs to be performed.
[0075] Step S150: If sampling rate optimization is required, the audio optimization terminal determines a target sampling rate (e.g., 24000Hz) based on each audio segment to be optimized, and performs resampling processing on the audio segments to be optimized based on the target sampling rate so that the sampling rate of each audio segment to be optimized is consistent.
[0076] Specifically, when the sampling rate optimization needs to be performed, the to-be-optimized audio segments are resampled (i.e., the sampling rate of the to-be-optimized audio is adjusted) so that the sampling rates of the to-be-optimized audio segments are consistent.
[0077] Step S160: When the sampling rate optimization is performed or the sampling rate optimization does not need to be performed, the audio optimization terminal determines whether the loudness optimization needs to be performed.
[0078] Specifically, the loudness is the strength of the sound, which is a subjective feeling of the size of the sound. The loudness is determined by the vibration amplitude of the sound signal. When the loudness difference between the to-be-optimized audios is large, the loudness optimization needs to be performed.
[0079] Step S170: If the loudness optimization needs to be performed, the audio optimization terminal obtains the average loudness values corresponding to the to-be-optimized audio segments, and performs the loudness optimization on the to-be-optimized audio segments based on the average loudness values, so that the loudness of the to-be-optimized audio segments tends to be consistent.
[0080] Step S180: When the loudness optimization is performed or the loudness optimization does not need to be performed, the audio optimization terminal splices the to-be-optimized audio segments in a preset order to obtain the to-be-broadcast audio.
[0081] Specifically, the preset order is determined by manual audio program directing, so that the complete to-be-broadcast audio is obtained. The to-be-broadcast audio is the audio after the unified optimization and splicing, and the user has a better listening experience.
[0082] The multi-source audio quality unified optimization method can unify the multi-source audios with large parameter differences. The multi-source audios to be optimized are marked as to-be-optimized audio segments, and then the to-be-optimized audios are sequentially optimized in the format, the sampling rate, and the loudness. The audios after the above optimization are spliced to obtain the to-be-broadcast audio. The to-be-broadcast audio can be sent to the user for listening, the listening experience of the to-be-broadcast audio after the optimization is better, the audio is more unified, and the user experience is improved.
[0083] In a second embodiment of the multi-source audio quality unified optimization method, based on the first embodiment, step S140 includes the following steps:
[0084] Step S210: The audio optimization terminal obtains the sampling rates of the to-be-optimized audio segments and marks them as original sampling rates.
[0085] Step S220: The audio optimization terminal takes the original sampling rate with the highest frequency among all the original sampling rates as a target sampling rate.
[0086] Specifically, for example, in the original sampling rates of the plurality of audio segments to be optimized, the frequency of 24000Hz appears the highest, and 24000Hz is taken as the target sampling rate.
[0087] Step S230: The audio optimization terminal marks the audio segment to be optimized with an original sampling rate higher than the target sampling rate as a first audio segment, and marks the audio segment to be optimized with an original sampling rate lower than the target sampling rate as a second audio segment.
[0088] Specifically, the first audio segment is an audio segment whose sampling rate needs to be reduced, and the second audio segment is an audio segment whose sampling rate needs to be increased.
[0089] Step S240: The audio optimization terminal performs down-sampling processing on the first audio segment to adjust the sampling rate of the first audio segment to the target sampling rate, specifically including the following steps:
[0090] Step S241: The audio optimization terminal sets the audio signal of the first audio segment as , the original sampling rate of the first audio segment is (unit: Hz), the original sampling period of the first audio segment is (unit: second), that is , the corresponding time , (n=0, 1, 2,...), n is the original sampling point index.
[0091] Step S242: The audio optimization terminal sets the target sampling rate as , and satisfies <, , the target sampling period is , sets the audio signal of the first audio segment after the down-sampling processing as , (p=0, 1, 2,...), p is the sampling point index after the down-sampling processing.
[0092] Step S243: The audio optimization terminal obtains a down-sampling value D, wherein, .
[0093] Step S244: The audio optimization terminal filters out the audio signal with a frequency higher than from the audio signal of the first audio segment based on a filter to obtain a transit audio segment.
[0094] Step S245: The audio optimization terminal retains one signal every D sampling points of the transit audio segment to complete the down-sampling processing, and the specific processing formula is:
[0095] ,
[0096] ,
[0097] wherein, is a transit audio segment, is a convolution; is a filter.
[0098] Step S250: The audio optimization terminal performs upsampling processing on the second audio segment to adjust the sampling rate of the second audio segment to the target sampling rate, specifically including the following steps:
[0099] Step S251: The audio optimization terminal sets the original audio signal of the second audio segment as , the original sampling rate of the second audio segment is (unit: Hz), the original sampling period of the second audio segment is (unit: second), that is, corresponds to time (n=0, 1, 2,...), n is the original sampling point index.
[0100] Step S252: The audio optimization terminal sets the target sampling rate as , and satisfies > , the target sampling period is , sets the audio signal of the second audio segment after upsampling processing as (k=0, 1, 2,...), k is the sampling point index after upsampling processing, and the time of the audio signal of the second audio segment after upsampling processing is .
[0101] Step S253: The audio optimization terminal obtains an upsampling value M, wherein, .
[0102] Step S254: The audio optimization terminal inserts (M-1) zero values between adjacent original sampling points of the second audio segment to obtain an intermediate signal , wherein z represents the inserted zero value, the original sampling point corresponds to , so as to retain the original audio information, and the position of the inserted zero value is .
[0103] Step S255: The audio optimization terminal performs filtering processing on the intermediate signal based on a filter (LPF) to obtain a final audio signal :
[0104] ,
[0105] wherein, is the unit impulse response of the filter, and satisfies: the cut-off frequency , m is a sampling point sequence.
[0106] Specifically, the embodiment provides a specific scheme for sampling rate optimization of the audio segment to be optimized.
[0107] In a third embodiment of the multi-source audio quality unified optimization method provided in the application, based on the first embodiment, step S160 comprises the following steps:
[0108] Step S310: The audio optimization terminal acquires the average loudness value corresponding to each audio segment to be optimized, and acquires a preset standard loudness value.
[0109] Specifically, the calculation formula of the loudness value of a certain sampling point of the audio to be optimized is:
[0110]
[0111] In the formula, is the loudness value, and abs(data) is the audio signal absolute value of a certain sampling point of the audio to be optimized; the average loudness value is obtained by averaging the loudness values of all sampling points of the audio to be optimized.
[0112] The standard loudness value herein is a loudness value that is more acceptable to the audience in the industry, for example, 30 dB.
[0113] Step S320: The audio optimization terminal marks the audio segment to be optimized, which is greater than the standard loudness value and has a difference greater than a first preset value (for example, 10 dB) from the standard loudness value, as a third audio segment.
[0114] Specifically, the third audio segment herein is the audio segment to be optimized with a larger loudness, which needs to be reduced.
[0115] Step S330: The audio optimization terminal marks the audio segment to be optimized, which is less than the standard loudness value and has an absolute value of the difference from the standard loudness value greater than the first preset value, as a fourth audio segment.
[0116] Specifically, the fourth audio segment herein is the audio segment to be optimized with a smaller loudness, which needs to be increased.
[0117] Step S340: The audio optimization terminal acquires the difference between the standard loudness value and the average loudness value of the third audio segment, and takes the difference as a first loudness difference (a negative value), performs loudness attenuation on the third audio segment, and the attenuation amount is the first loudness difference, and specifically comprises the following steps:
[0118] Step S341: The audio optimization terminal acquires the original amplitude of each sampling point of the third audio segment , n is the sampling point index of the third audio segment, and in digital audio, is a quantized value, the amplitude of each sampling point of the third audio segment is attenuated to :
[0119] ,
[0120] ,
[0121] wherein, is the first loudness difference value.
[0122] Step S350: The audio optimization terminal obtains a difference value between the standard loudness value and the average loudness value of the fourth audio segment, and takes the difference value as a second loudness difference value (a positive value), performs loudness gain on the fourth audio segment, and the gain amount is the second loudness difference value, including the following steps:
[0123] Step S351: The audio optimization terminal obtains the original amplitude of each sampling point of the fourth audio segment , n is the sampling point index of the fourth audio segment, and in digital audio, is a quantized value, the amplitude of each sampling point of the fourth audio segment is attenuated to :
[0124] ,
[0125] ,
[0126] wherein, is the second loudness difference value.
[0127] Specifically, the embodiment provides a specific scheme for performing loudness optimization on the audio segment to be optimized.
[0128] In a fourth embodiment of the multi-source audio quality unified optimization method provided in the application, based on the first embodiment, the system further comprises an audio distribution terminal (for example, a computer terminal with network communication capability) in communication connection with the audio optimization terminal, and a user terminal (for example, a smart phone terminal) in communication connection with the audio distribution terminal; step S180 further includes the following steps:
[0129] Step S410: The audio optimization terminal sends the audio to be played to the audio distribution terminal.
[0130] Step S420: The audio distribution terminal sends the audio to be played to the user terminal.
[0131] Step S430: The user terminal obtains a historical running log of the current user operating the user terminal, and determines a target volume based on the historical running log.
[0132] Specifically, the historical operation log of the user operating the user terminal can reflect the volume demand characteristics of the user for audio playing, and therefore the target volume is determined based on the historical operation log.
[0133] Specifically, the volume unit here is the gain percentage when the user terminal plays audio, for example, when the volume is 100%, the volume of the audio played by the user terminal is the largest, and when the volume is 0%, the user terminal is muted.
[0134] Step S440: The user terminal plays the audio to be played, and the volume of the playing is the target volume.
[0135] In a fifth embodiment of the multi-source audio quality unified optimization method proposed in the application, based on the fourth embodiment, the historical operation log includes the average playing volume corresponding to the user using the user terminal in different time periods; for example, in this embodiment, each day is divided into 12 time periods, and each time period corresponds to a duration of 2 hours, from 0-2, 2-4, and so on; the external environment faced in different time periods is different, so different playing volumes are used, and then when the user plays the audio to be played, based on the time period in which the current time is located, the playing volume suitable for being used (i.e. the target volume) can be determined; the user terminal obtains the historical operation log of the user operating the user terminal, and determines the target volume based on the historical operation log, including:
[0136] Step S510: The user terminal marks the time period in which the current time falls as the target time period.
[0137] Step S520: The user terminal obtains the average playing volume corresponding to the user using the user terminal in the target time period based on the historical operation log, and uses it as the target volume.
[0138] Specifically, for example, the current time when the user is ready to play the audio to be played is 7 am, and the target time period in which 7 am falls is 6-8 am; therefore, the average playing volume corresponding to the terminal in the time period of 6-8 am is used as the target volume.
[0139] In a sixth embodiment of the multi-source audio quality unified optimization method proposed in the application, based on the fourth embodiment, step S440 further includes the following steps:
[0140] Step S610: The user terminal determines whether an operation instruction for volume adjustment by the user is received within a first preset time period (for example, 3 seconds) since the audio to be played starts to be played.
[0141] Specifically, the operation instruction here can be triggered by the user pressing the volume adjustment button of the user terminal.
[0142] If yes, step S620 is performed: the user terminal marks the currently played audio to be played as adjusted audio, obtains the volume of the adjusted audio after the user performs volume adjustment, subtracts the volume before the volume adjustment, and marks the corresponding volume adjustment value of the adjusted audio.
[0143] Specifically, if yes, it indicates that the user starts to play the audio to be played through the user terminal, that is, the user adjusts the volume of the audio to be played, which indicates that the volume of the currently played audio to be played is not suitable for the user's listening experience. The volume adjustment value here is the amplitude of the user's volume adjustment (for example, 20%).
[0144] Step S630: The user terminal determines whether the number of adjusted audios in the first preset number (for example, 10) of past played audios to be played is greater than or equal to a second preset number (for example, 5), wherein the second preset number is less than the first preset number.
[0145] Specifically, if the number of adjusted audios in the 10 audios to be played that have been played in the past is greater than or equal to 5, it indicates that the user has adjusted the volume of the audio to be played multiple times in the past, and it can be inferred that the initial volume (that is, the aforementioned target volume) of the audio to be played played by the user terminal is not suitable for the current listening experience; therefore, the target volume needs to be adjusted.
[0146] Step S640: If greater than or equal to the second preset number, the user terminal determines whether a first condition or a second condition is met, wherein the first condition is that the volume adjustment values corresponding to the adjusted audios in the first preset number of past played audios to be played are all positive values (that is, each adjustment is to increase the volume), and are all greater than the preset volume value (for example, 30%); the second condition is that the volume adjustment values corresponding to the adjusted audios in the first preset number of past played audios to be played are all negative values (that is, each adjustment is to decrease the volume), and are all less than the opposite value of the preset volume value (for example, -30%).
[0147] Step S650: If the first condition is met, the user terminal increases the target volume, and the increase amount is the preset volume value.
[0148] Specifically, if the first condition is met, it indicates that the user has manually increased the playing volume of the audio to be played multiple times in the past, and it can be inferred that the user prefers to listen to audio with a larger volume, so the target volume is increased, and the increase amount is the aforementioned preset volume value (that is, 30%).
[0149] Step S660: If the second condition is met, the user terminal decreases the target volume, and the decrease amount is the preset volume value.
[0150] Specifically, if the second condition is satisfied, it indicates that the user has manually reduced the playing volume of the audio to be played multiple times in the past, and thus it can be inferred that the user prefers to listen to audio with a smaller volume. Therefore, the target volume is reduced, and the reduction amount is the aforementioned preset volume value (i.e., 30%).
[0151] In a seventh embodiment of the multi-source audio quality uniform optimization method provided in the application, based on the fourth embodiment, the user terminal further comprises a microphone capable of collecting sound signals; and step S440 further comprises the following steps:
[0152] Step S710: Before the user terminal starts playing the audio to be played, the user terminal divides the audio to be played into a plurality of sequentially adjacent sub-audio segments, and marks the first sub-audio segment. The average loudness value of each first sub-audio segment is obtained and marked as a first loudness value.
[0153] Specifically, the duration of each first sub-audio segment is the same, for example, 1 second.
[0154] Step S720: The user terminal marks the first sub-audio segment with a first loudness value less than a second preset value (for example, 20 dB) as a target audio segment.
[0155] Specifically, the target audio segment is a first sub-audio segment with lower loudness, for example, an audio segment corresponding to a human voice gap period in the audio to be played.
[0156] Step S730: After the user terminal starts playing the audio to be played, the user terminal collects the sound signals of the surrounding environment during each playing of the target audio segment through the microphone, and obtains the average loudness value of the sound signals of the surrounding environment.
[0157] Specifically, the sound signals of the surrounding environment collected by the microphone during each playing of the target audio segment can reflect the overall sound situation of the environment in which the user terminal is currently located, and thus the average loudness value of the sound signals of the surrounding environment can reflect the strength of the ambient sound during the playing of the target audio segment. The reason for collecting the sound signals of the surrounding environment only during each playing of the target audio segment is that the loudness of the target audio segment itself is not high, so the interference of the external environment sound signal strength will be smaller, that is, the average loudness value of the sound signals of the surrounding environment will be more realistic.
[0158] Step S740: When the average loudness values of the sound signals of the surrounding environment in the past second preset time length (for example, 10 minutes) are all less than a third preset value (30 dB), and the current time is within a preset time period (for example, 11 pm to 6 am), the user terminal infers that the user is in a resting state.
[0159] Specifically, when the average loudness values of the ambient sound signals in each of the past second preset time periods are all less than the third preset value, and the current time is in the preset time period, it indicates that the user is likely to listen to the to-be-played audio while preparing to sleep and rest.
[0160] Step S750: When the user is in the resting state, the user terminal decreases the playing volume of the to-be-played audio currently being played by 1 every third preset time period (for example, 3 minutes).
[0161] Specifically, when the user is in the resting state, the user terminal decreases the playing volume of the to-be-played audio currently being played by 1 every third preset time period, so as to assist the user in resting and improve the user experience.
[0162] In an eighth embodiment of the multi-source audio quality unified optimization method provided in the application, based on the fourth embodiment, the user terminal further comprises a microphone capable of collecting sound signals; and step S440 further comprises the following steps:
[0163] Step S810: When the user terminal plays the to-be-played audio through the earphone, and the current time is in the preset time period (for example, 11 pm to 6 am), the user terminal generates a monitoring instruction.
[0164] Specifically, when the user terminal plays the to-be-played audio through the earphone, and the current time is in the preset time period, it indicates that the user is more likely to listen to the to-be-played audio and prepare to rest, so the ambient sound signals are collected based on the microphone subsequently to determine whether the user has entered the resting state.
[0165] Step S820: The user terminal collects ambient sound signals through the microphone based on the monitoring instruction, and marks them as ambient sound signals.
[0166] Step S830: The user terminal divides the ambient sound signals in the past fourth preset time period (for example, 1 minute) into a plurality of sequentially adjacent sub-audio segments, and marks them as second sub-audio segments, obtains the average loudness value of each second sub-audio segment, and marks it as a second loudness value.
[0167] Specifically, the time length of each second sub-audio segment is the same, for example, 1 second.
[0168] Step S840: The user terminal marks the second loudness value meeting a third condition in the past fourth preset time period as a first target value, wherein the third condition is that the two second loudness values adjacent to the first target value are both less than the first target value.
[0169] Specifically, if the user falls asleep, the ambient sound signal here can reflect that the user has regular and uniform snoring, and regular snoring can be presented in loudness, specifically:
[0170] The second loudness values in each second audio segment in the ambient sound signal in the past fourth preset time length are compared in size to find a first target value in the second loudness values, and the two second loudness values adjacent to the first target value are both smaller than the first target value, that is, the first target value is the point with the maximum loudness, corresponding to the moment when the user's snoring is the loudest.
[0171] Step S850: When the number of first target values is at least two, and the interval length between the collection moments of adjacent two first target values is less than a fifth preset time length (5 seconds), the user terminal determines a second target value and a third target value corresponding to each first target value, wherein the second target value corresponding to the first target value is a second loudness value earlier than the first target value and equal to a fourth preset value (here, the fourth preset value is set to be the normal minimum loudness of snoring when a human body falls asleep, for example, 10 dB), and the third target value corresponding to the first target value is a second loudness value later than the first target value and equal to the fourth preset value.
[0172] Specifically, the fifth preset time length here is set to be the maximum normal interval length (for example, 5 seconds) between snoring after a human body falls asleep under normal circumstances; here, the premise is that the number of first target values is at least two, to ensure that the sample number of first target values is sufficient for subsequent analysis and judgment.
[0173] Specifically, in this step, the second target value corresponding to each first target value and the third target value are determined, and it is known that the second target value is the second loudness value corresponding to the start moment of the user's snoring action, and the third target value is the second loudness value corresponding to the end moment of the user's snoring action. The period formed by the second target value - the first target value - the third target value corresponds to a complete snoring action of the user.
[0174] Step S860: The user terminal marks the second loudness value between the first target value and the corresponding second target value at the collection moment as a fourth target value corresponding to the first target value, and marks the second loudness value between the first target value and the corresponding third target value at the collection moment as a fifth target value corresponding to the first target value.
[0175] Specifically, it is known that the fourth target value is the loudness value from the start of snoring to the maximum snoring loudness of the user, so the fourth target value should be constantly increasing with time; similarly, it is known that the fifth target value is the loudness value from the maximum snoring loudness to the end of snoring of the user, so the fifth target value should be constantly decreasing with time.
[0176] Step S870: The user terminal calculates the first slope and the second slope corresponding to each first target value.
[0177]
[0178]
[0179] wherein, is the first slope corresponding to the jth first target value; is the second slope corresponding to the jth first target value; j is a positive integer, and 1≤j≤J, J is the number of first target values (for example, 3, i.e. the user snores 3 times in the past fourth preset time length); is the (o+1)th third target value corresponding to the jth first target value, o is a positive integer, and 1≤o≤ is the number of third target values corresponding to the jth first target value; is the time length of the second sub-audio segment; is the (q+1)th fifth target value corresponding to the jth first target value, q is a positive integer, and 1≤q≤ is the number of fifth target values corresponding to the jth first target value.
[0180] Specifically, the first slope in this step can reflect the rate of change of the loudness in the process of the loudness of snoring sound continuously increasing in a snoring action of the user; the second slope can reflect the rate of change of the loudness in the process of the loudness of snoring sound continuously decreasing in a snoring action of the user; if the user has rested and fallen asleep, the snoring actions of the user should be regular and uniform, so the first slopes corresponding to each snoring action should be close, and the second slopes corresponding to each snoring action should also be close.
[0181] Step S880: When the standard deviation between the first slopes corresponding to each first target value, and the standard deviation between the second slopes corresponding to each first target value are all less than a preset threshold (for example, 0.5), the user terminal determines that the user is in a resting state.
[0182] Specifically, the standard deviation can reflect the degree of dispersion between parameters, and the smaller the standard deviation, the smaller the degree of dispersion between the corresponding parameters, i.e. the closer; therefore, when the standard deviation between the first slopes corresponding to each first target value, and the standard deviation between the second slopes corresponding to each first target value are all less than a preset threshold, it indicates that the loudness change corresponding to each snoring action of the user is close, and it can be determined that the user is in a resting state.
[0183] Step S890: When the user is in the resting state, the user terminal decreases the playing volume of the audio to be played which is currently playing by 1 every third preset time length.
[0184] The application further provides a multi-source audio quality unified optimization system applying the multi-source audio quality unified optimization method.
[0185] The above-mentioned serial numbers of the embodiments of the application are only for description, and do not represent the advantages and disadvantages of the embodiments.
[0186] The embodiments of the application are described above in combination with the drawings, but the application is not limited to the above-mentioned specific embodiments, and the above-mentioned specific embodiments are only illustrative but not restrictive. Those skilled in the art can make many forms under the inspiration of the application without departing from the purpose of the application and the scope protected by the claims, and these all belong to the protection of the application.
Claims
1. A method for unified optimization of multi-source audio quality, characterized in that, An application to a multi-source audio quality unified optimization system; the system includes an audio optimization terminal; the method includes: The audio optimization terminal acquires multiple audio segments that need to be spliced and mixed, and marks them as audio segments to be optimized; The audio optimization terminal determines whether format optimization needs to be performed; If format optimization is required, the audio optimization terminal determines the target format, marks the audio segments to be optimized that do not match the target format as audio segments to be converted, and converts the audio format of the audio segments to be converted to the target format; After format optimization is completed, or when format optimization is not required, the audio optimization terminal determines whether sampling rate optimization needs to be performed. If sampling rate optimization is required, the audio optimization terminal determines the target sampling rate based on each audio segment to be optimized, and performs resampling processing on the audio segments to be optimized based on the target sampling rate, so that the sampling rate of each audio segment to be optimized is consistent. When sampling rate optimization is completed, or when sampling rate optimization is not required, the audio optimization terminal determines whether loudness optimization needs to be performed. If loudness optimization is required, the audio optimization terminal obtains the average loudness value corresponding to each audio segment to be optimized, and performs loudness optimization on the audio segment to be optimized based on the average loudness value; When loudness optimization is completed, or when loudness optimization is not required, the audio optimization terminal splices the audio segments to be optimized end to end in a preset order to obtain the audio to be played. The system also includes an audio distribution terminal communicatively connected to the audio optimization terminal, and a user terminal communicatively connected to the audio distribution terminal; when loudness optimization is completed, or when loudness optimization is not required, the audio optimization terminal splices the audio segments to be optimized end-to-end in a preset order to obtain the audio to be played, and then further includes: The audio optimization terminal sends the audio to be played to the audio distribution terminal; The audio distribution terminal sends the audio to be played to the user terminal; The user terminal obtains the historical operation log of the current user's operation on the user terminal, and determines the target volume based on the historical operation log; The user terminal plays the audio to be played, and the volume of the playback is the target volume; The user terminal further includes a microphone capable of acquiring sound signals; the user terminal plays the audio to be played, and the volume of the playback is the target volume, and then further includes: When the user terminal plays the audio to be played through the headphones, and the current time is within a preset time period, the user terminal generates a monitoring command; The user terminal collects ambient sound signals through the microphone based on the monitoring command and marks them as ambient sound signals; The user terminal divides the ambient sound signal within the past fourth preset time period into multiple sequentially adjacent sub-audio segments and marks them as second sub-audio segments, obtains the average loudness value of each second sub-audio segment, and marks it as a second loudness value; The user terminal marks the second loudness value that meets the third condition within the past fourth preset time period as the first target value, wherein the third condition is: the two second loudness values adjacent to the first target value are both less than the first target value; When the number of the first target values is at least two, and the interval between the collection times of two adjacent first target values is less than a fifth preset time, the user terminal determines the second target value and the third target value corresponding to each first target value, wherein the second target value corresponding to the first target value is a second loudness value that is earlier than the first target value and equal to a fourth preset value, and the third target value corresponding to the first target value is a second loudness value that is later than the first target value and equal to a fourth preset value. The user terminal marks the second loudness value that is between the first target value and the corresponding second target value at the time of collection as the fourth target value corresponding to the first target value, and marks the second loudness value that is between the first target value and the corresponding third target value at the time of collection as the fifth target value corresponding to the first target value; The user terminal calculates the first slope and the second slope corresponding to each first target value: , , In the formula, The first slope corresponding to the j-th first target value; Let be the second slope corresponding to the j-th first objective value; j is a positive integer, and 1≤j≤J, where J is the number of first objective values; Let $\mathbf(j)$ be the $\mathbf(o+1)$ third objective value corresponding to the $\mathbf(j)$ first objective value, where $o$ is a positive integer and $1 \leq \mathbf(j)$. , This represents the number of third objective values corresponding to the j-th first objective value. The duration of the second sub-audio segment; Let be the (q+1)th fifth objective value corresponding to the j-th first objective value, where q is a positive integer and 1 ≤ q ≤ 1. , The number of fifth objective values corresponding to the j-th first objective value; When the standard value between the first slopes corresponding to each first target value and the standard deviation between the second slopes corresponding to each first target value are both less than a preset threshold, the user terminal presumes that the user is in a resting state. When the user is in a resting state, the user terminal reduces the playback volume of the currently playing audio once every third preset time interval.
2. The method for unified optimization of multi-source audio quality according to claim 1, characterized in that, If sampling rate optimization is required, the audio optimization terminal determines a target sampling rate based on each audio segment to be optimized, and performs resampling processing on the audio segments to be optimized based on the target sampling rate to ensure that the sampling rates of each audio segment to be optimized are consistent, including: The audio optimization terminal acquires the sampling rate of each audio segment to be optimized and marks it as the original sampling rate; The audio optimization terminal uses the most frequent original sampling rate among all the original sampling rates as the target sampling rate. The audio optimization terminal marks audio segments to be optimized with an original sampling rate higher than the target sampling rate as first audio segments, and audio segments to be optimized with an original sampling rate lower than the target sampling rate as second audio segments; The audio optimization terminal performs downsampling processing on the first audio segment to adjust the sampling rate of the first audio segment to the target sampling rate; The audio optimization terminal performs upsampling processing on the second audio segment to adjust the sampling rate of the second audio segment to the target sampling rate.
3. The method for unified optimization of multi-source audio quality according to claim 1, characterized in that, If loudness optimization is required, the audio optimization terminal obtains the average loudness value corresponding to each audio segment to be optimized, and performs loudness optimization on the audio segments to be optimized based on the average loudness value, including: The audio optimization terminal obtains the average loudness value corresponding to each audio segment to be optimized, and obtains the preset standard loudness value; The audio optimization terminal marks the audio segment to be optimized as the third audio segment, which is greater than the standard loudness value and whose difference from the standard loudness value is greater than a first preset value. The audio optimization terminal marks the audio segment to be optimized as the fourth audio segment, which is less than the standard loudness value and whose absolute value of the difference from the standard loudness value is greater than the first preset value. The audio optimization terminal obtains the difference between the standard loudness value and the average loudness value of the third audio segment, and uses it as the first loudness difference to attenuate the loudness of the third audio segment, with the attenuation amount being the first loudness difference. The audio optimization terminal obtains the difference between the standard loudness value and the average loudness value of the fourth audio segment, and uses it as the second loudness difference to perform loudness gain on the fourth audio segment, with the gain amount being the second loudness difference.
4. The method for unified optimization of multi-source audio quality according to claim 1, characterized in that, The historical operation log includes the average playback volume when the user uses the user terminal at different time periods. The user terminal obtains the historical operation log of the current user's operation on the user terminal, and determines the target volume based on the historical operation log, including: The user terminal marks the time period that the current moment falls into as the target time period; The user terminal obtains the average playback volume corresponding to the user's use of the user terminal during the target time period based on historical operation logs, and uses it as the target volume.
5. The method for unified optimization of multi-source audio quality according to claim 1, characterized in that, The user terminal plays the audio to be played, and the volume of the playback is the target volume, and then the following is also included: The user terminal determines whether it has received a volume adjustment command from the user within a first preset duration from the start of the audio to be played. If so, the user terminal marks the audio to be played this time as adjusted audio, obtains the volume of the adjusted audio after the user adjusts the volume minus the volume before the volume adjustment, and marks it as the volume adjustment value corresponding to the adjusted audio; The user terminal determines whether the number of adjusted audios in the first preset number of audios to be played in the past is greater than or equal to a second preset number, wherein the second preset number is less than the first preset number; If the quantity is greater than or equal to the second preset quantity, the user terminal determines whether the first condition or the second condition is met. The first condition is that the volume adjustment values corresponding to the adjusted audio in the first preset quantity of audio to be played in the past are all positive values and are all greater than the preset volume value. The second condition is that the volume adjustment values corresponding to the adjusted audio in the first preset quantity of audio to be played in the past are all negative values and are all less than the opposite value of the preset volume value. If the first condition is met, the user terminal increases the target volume by the amount of the increase, which is equal to the preset volume value. If the second condition is met, the user terminal will reduce the target volume by the amount of reduction, which is the preset volume value.
6. The method for unified optimization of multi-source audio quality according to claim 1, characterized in that, The user terminal also includes a microphone capable of collecting sound signals; The user terminal plays the audio to be played, and the volume of the playback is the target volume. This is preceded by: Before the user terminal starts playing the audio to be played, the user terminal divides the audio to be played into multiple sequentially adjacent sub-audio segments and marks them as the first sub-audio segments, obtains the average loudness value of each first sub-audio segment, and marks it as the first loudness value; The user terminal marks the first sub-audio segment whose first loudness value is less than the second preset value as the target audio segment; After the user terminal starts playing the audio to be played, the user terminal uses the microphone to collect the sound signal of the surrounding environment during each playback of the target audio segment, and obtains the average loudness value of the sound signal of the surrounding environment; When the average loudness value of the sound signals of the surrounding environment within the past second preset time period is less than the third preset value, and the current time is within the preset time period, the user terminal presumes that the user is in a resting state. When the user is in a resting state, the user terminal reduces the playback volume of the currently playing audio once every third preset time interval.
7. A multi-source audio quality unified optimization system, characterized in that, The system employs the multi-source audio quality unified optimization method as described in any one of claims 1-6; the system includes an audio optimization terminal.
Citation Information
Patent Citations
Volume adjusting method and device
CN104978166A
Multimedia file splicing method and device, equipment and medium
CN111182315A
Audio playing control method and device, equipment and medium
CN112243151A