A method and related device for detecting sound quality of high-resolution audio
By screening and analyzing the sampling rate, quantization accuracy and spectral characteristics of audio, high-resolution audio can be automatically identified, solving the problem of low manual recognition accuracy and achieving higher recognition accuracy and efficiency.
Patent Information
- Application Number
- CN202210461512.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-28
- Publication Date
- 2025-09-12
- Estimated Expiration
- 2042-04-28
AI Technical Summary
The existing technology relies on manual observation when recognizing high-resolution audio, resulting in low recognition accuracy.
By obtaining the sampling rate and quantization accuracy of the audio to be tested, filtering out audio that meets the preset standards, determining the effective spectrum height, filtering out audio with an effective spectrum height greater than the frequency threshold, and judging whether its high-frequency band energy shows a downward trend and there is no energy mutation, it is confirmed whether the audio is high-resolution audio.
The recognition accuracy of high-resolution audio is improved, the influence of human subjective will is avoided, and efficient recognition can be achieved on computer equipment.
Smart Images

Figure CN114822595B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of audio processing technology, and in particular to a sound quality detection method and related device for high-resolution audio. Background Art
[0002] High-resolution audio refers to audio files with a sampling rate greater than 44100Hz or a quantization accuracy greater than 16 bits.
[0003] In everyday life, there are often a large number of fake high-resolution audio files created by resampling low-quality audio, or audio files with abnormal spectra. Fake high-resolution audio refers to audio that meets the sampling rate and quantization accuracy requirements, but often lacks effective energy in the spectrum above 44100Hz. Audio files with abnormal spectra refer to audio files whose audio data has been corrupted due to improper operation during the audio recording or editing process.
[0004] Existing technologies generally rely on manual observation to identify high-resolution audio, but manual recognition often suffers from low recognition accuracy due to subjective differences. Summary of the Invention
[0005] The embodiments of the present invention provide a method and related apparatus for detecting the sound quality of high-resolution audio, which are used to improve the recognition accuracy of high-resolution audio.
[0006] A first aspect of an embodiment of the present application provides a method for detecting sound quality of high-resolution audio, comprising:
[0007] Get the sampling rate and quantization accuracy of the audio to be detected;
[0008] Filtering out a first audio whose sampling rate and quantization accuracy both meet preset standards from the audio to be detected;
[0009] Determine an effective spectrum height of the first audio, where the effective spectrum height is a frequency corresponding to a location of maximum audio energy;
[0010] Filtering out a second audio whose effective spectrum height is greater than a frequency threshold from the first audio;
[0011] If the overall energy trend of the preset high frequency band in the second audio presents a downward trend and no frequency point with a sudden energy change occurs in the preset high frequency band, the second audio is determined to be high-resolution audio.
[0012] Preferably, after selecting the first audio whose sampling rate and quantization accuracy meet the preset standards from the audio to be detected, and before obtaining the effective spectrum height of the first audio, the method further includes:
[0013] Obtaining a sampling rate of the first audio;
[0014] Classifying the first audio into a target audio category according to a plurality of preset sampling intervals, wherein each sampling interval corresponds to a different audio category;
[0015] If the sampling rate of the first audio is not equal to the resampling rate corresponding to the target category audio, the first audio is resampled according to the resampling rate corresponding to the target category audio, so that the resampling rate of the first audio is the same as the resampling rate corresponding to the target category audio, wherein different categories of audio correspond to different resampling rates.
[0016] Preferably, determining the effective spectrum height of the first audio comprises:
[0017] Frame-by-frame windowing of the first audio, and performing time domain to frequency domain conversion to obtain a frequency domain signal of each frame of audio in the first audio;
[0018] calculating the energy of each frame of audio in the first audio, so as to screen out valid frames in the first audio according to the energy of each frame of audio in the first audio and a first energy threshold;
[0019] Determine the effective spectral height of the effective frame in the first audio.
[0020] Preferably, selecting the second audio whose effective spectrum height is greater than a frequency threshold from the first audio includes:
[0021] Counting the number of first valid frames in the first audio whose effective spectrum height is greater than the frequency threshold;
[0022] If the ratio of the number of first valid frames in the first audio to the total number of valid frames in the first audio is greater than a first ratio threshold, the first audio is regarded as the second audio of the same target category.
[0023] After screening out valid frames in the first audio and before determining effective spectrum heights of the valid frames in the first audio, the method further includes:
[0024] If the ratio of the number of valid frames in the first audio to the total number of audio frames in the first audio is less than a second ratio threshold, it is determined that the first audio is non-high-resolution audio.
[0025] Preferably, determining the effective spectrum height of the effective frame in the first audio includes:
[0026] Divide the valid frames in the first audio into M frequency bands according to a preset frequency interval, each frequency band including N frequency points, where M is greater than or equal to 2 and N is greater than or equal to 1;
[0027] Calculate, according to the frequency band calculation range corresponding to the target audio category, the first-order frequency band energy E1 of each frequency band and the sum E2 of the energies of two adjacent first-order frequency bands along the frequency axis in each valid frame of the first audio, where different audio categories correspond to different frequency band calculation ranges;
[0028] The spectrum height of each valid frame in the first audio is determined according to E1 or E2 corresponding to each frequency band in each valid frame in the first audio and a rule for determining the effective spectrum height.
[0029] Preferably, the effective spectrum height determination rule includes:
[0030] An inflection point frequency band where the energy of each frame begins to decrease is determined, and the frequency of the inflection point frequency band is determined as the effective spectrum height of each frame.
[0031] Preferably, determining the effective spectrum height of each valid frame in the first audio according to E1 or E2 corresponding to each frequency band in each valid frame in the first audio and an effective spectrum height determination rule includes:
[0032] Obtaining a first frequency band corresponding to a maximum E2 value in each valid frame of the first audio;
[0033] If the E1 value of the first frequency band is greater than the second energy threshold, or the E2 value of the second frequency band adjacent to the first frequency band along the frequency axis is greater than the third energy threshold, the first frequency band is regarded as the inflection point frequency band, and the frequency of the inflection point frequency band is determined as the effective spectrum height of the effective frame, wherein the second energy threshold and the third energy threshold are used to characterize the energy variation range of the inflection point frequency band;
[0034] or,
[0035] If the energy of the first frequency band is less than the invalid energy threshold, the first frequency band is regarded as the inflection point frequency band, and the frequency of the inflection point frequency band is determined as the effective spectrum height of the effective frame.
[0036] Preferably, if the overall energy trend of the preset high frequency band in the second audio presents a downward trend, and no frequency point with a sudden energy change occurs in the preset high frequency band, then determining that the second audio is high-resolution audio includes:
[0037] Calculating, according to a plurality of preset center frequency bands and preset high frequency points in a preset high frequency band of the target category audio, a plurality of energy differences between the plurality of preset center frequency bands and the preset high frequency points for each frame of audio in the second audio, wherein the frequencies of the plurality of preset center frequency bands increase sequentially but are all less than the frequency of the preset high frequency points, and wherein different categories of audio correspond to different plurality of preset center frequency bands and different preset high frequency points;
[0038] counting a number of second valid frames in which the multiple energy difference values in the second audio are all greater than corresponding multiple thresholds, wherein the multiple thresholds respectively corresponding to the multiple energy difference values decrease as the frequencies of the multiple preset center frequency bands increase;
[0039] If the ratio of the second valid frame number to the total audio frame number in the second audio is not less than a third ratio threshold, it indicates that the overall energy trend of the preset high frequency band in the second audio presents a downward trend.
[0040] Preferably, if the overall energy trend of the preset high frequency band in the second audio presents a downward trend and no frequency point with a sudden energy change occurs in the preset high frequency band, then determining that the second audio is high-resolution audio further includes:
[0041] According to a preset high frequency band of the target audio category, obtaining the preset high frequency band of each audio frame in the second audio, wherein different high frequency band ranges are set for different audio categories;
[0042] Calculating an energy distribution interval of a preset high frequency band of each audio frame in the second audio;
[0043] Counting the number of third valid frames without abnormal frequency points within the energy distribution interval, wherein the abnormal frequency point is a frequency point whose frequency energy is greater than a critical value of the energy distribution interval;
[0044] If the ratio of the third number of valid frames to the total number of audio frames in the second audio is not less than a fourth ratio threshold, it indicates that no frequency point with sudden energy changes occurs in the preset high frequency band of the second audio;
[0045] Determine that the second audio is high-resolution audio.
[0046] Preferably, if the overall energy trend of the preset high frequency band in the second audio presents a downward trend and no frequency point with a sudden energy change occurs in the preset high frequency band, then determining that the second audio is high-resolution audio further includes:
[0047] Slide each audio frame of the second audio using a sliding window with a single frequency point as a step size, wherein the sliding window is at least larger than the frequency spectrum range of the single frequency point;
[0048] For each audio frame in the second audio, calculating the sum of first-order differences of energies of all frequency points in the sliding window;
[0049] If the sum of the first-order differences of the energies of all frequency points in the sliding window is negative and greater than a fourth energy threshold, determining that the audio frame in the second audio is a normal frame;
[0050] or,
[0051] If the sign of the sum of the first-order differences of the energies of all frequency points in the sliding window is positive and less than a fifth energy threshold, determining that the audio frame in the second audio is a normal frame;
[0052] Counting a proportion of normal frames in the second audio; if the proportion of normal frames in the second audio is not less than a fifth proportion threshold, indicating that no frequency point with a sudden energy change occurs in the preset high frequency band of the second audio;
[0053] Determine that the second audio is high-resolution audio.
[0054] Preferably, if the overall energy trend of the preset high frequency band in the second audio presents a downward trend and no frequency point with a sudden energy change occurs in the preset high frequency band, then determining that the second audio is high-resolution audio further includes:
[0055] Slide each audio frame of the second audio using a sliding window with a single frequency point as a step size, wherein the sliding window is at least larger than the frequency spectrum range of the single frequency point;
[0056] For each audio frame in the second audio, calculating the sum of the first-order differences of the energy of each frequency point in the sliding window;
[0057] If the sign of the sum of the first-order differences of the energies of each frequency point in the sliding window is negative and greater than a sixth energy threshold, determining that the audio frame in the second audio is a normal frame;
[0058] or,
[0059] If the sum of the first-order differences of the energies of each frequency point in the sliding window is positive and less than a seventh energy threshold, determining that the audio frame in the second audio is a normal frame;
[0060] Counting a proportion of normal frames in the second audio; if the proportion of normal frames in the second audio is not less than a sixth proportion threshold, indicating that no frequency point with sudden energy change occurs in the preset high frequency band of the second audio;
[0061] Determine that the second audio is high-resolution audio.
[0062] Preferably, before sliding through each audio frame in the second audio using a sliding window with a single frequency point as a step size, the method further includes:
[0063] A smoothing process is performed on each audio frame in the second audio using a smoothing filter.
[0064] A second aspect of an embodiment of the present application provides a high-resolution audio quality detection device, comprising:
[0065] An acquisition unit, used to obtain the sampling rate and quantization accuracy of the audio to be detected;
[0066] a first screening unit, configured to screen out, from the audio to be detected, a first audio whose sampling rate and quantization accuracy both meet preset standards;
[0067] a determining unit, configured to determine an effective spectrum height of the first audio, where the effective spectrum height is a frequency corresponding to a location with maximum energy of the audio;
[0068] a second filtering unit, configured to filter out, from the first audio, a second audio whose effective spectrum height is greater than a frequency threshold;
[0069] The determining unit is further configured to determine that the second audio is high-resolution audio if the overall energy trend of the preset high-frequency band in the second audio presents a downward trend and no frequency point with a sudden energy change occurs in the preset high-frequency band.
[0070] Preferably, the acquisition unit is further configured to:
[0071] After selecting a first audio whose sampling rate and quantization accuracy both meet preset standards from the audio to be detected, and before obtaining the effective spectrum height of the first audio, obtaining the sampling rate of the first audio;
[0072] The device further comprises:
[0073] a classification unit, configured to classify the first audio into a target category of audio according to a plurality of preset sampling intervals, wherein each sampling interval corresponds to a different category of audio;
[0074] A resampling unit is used to resample the first audio according to the resampling rate corresponding to the target category audio if the sampling rate of the first audio is not equal to the resampling rate corresponding to the target category audio, so that the resampling rate of the first audio is the same as the resampling rate corresponding to the target category audio, wherein different categories of audio correspond to different resampling rates.
[0075] Preferably, the determining unit is specifically configured to:
[0076] Frame-by-frame windowing of the first audio, and performing time domain to frequency domain conversion to obtain a frequency domain signal of each frame of audio in the first audio;
[0077] calculating the energy of each frame of audio in the first audio, so as to screen out valid frames in the first audio according to the energy of each frame of audio in the first audio and a first energy threshold;
[0078] Determine the effective spectral height of the effective frame in the first audio.
[0079] Preferably, the second screening unit is specifically used for:
[0080] Counting the number of first valid frames in the first audio whose effective spectrum height is greater than the frequency threshold;
[0081] If the ratio of the number of first valid frames in the first audio to the total number of valid frames in the first audio is greater than a first ratio threshold, the first audio is regarded as the second audio of the same target category.
[0082] Preferably, the determining unit is further configured to:
[0083] After filtering out valid frames in the first audio and before determining the effective spectrum height of the valid frames in the first audio, if the ratio of the number of valid frames in the first audio to the total number of audio frames in the first audio is less than a second ratio threshold, the first audio is determined to be non-high-resolution audio.
[0084] Preferably, the determining unit is specifically configured to:
[0085] Divide the valid frames in the first audio into M frequency bands according to a preset frequency interval, each frequency band including N frequency points, where M is greater than or equal to 2 and N is greater than or equal to 1;
[0086] Calculate, according to the frequency band calculation range corresponding to the target audio category, the first-order frequency band energy E1 of each frequency band and the sum E2 of the energies of two adjacent first-order frequency bands along the frequency axis in each valid frame of the first audio, where different audio categories correspond to different frequency band calculation ranges;
[0087] The spectrum height of each valid frame in the first audio is determined according to E1 or E2 corresponding to each frequency band in each valid frame in the first audio and a rule for determining the effective spectrum height.
[0088] Preferably, the effective spectrum height determination rule includes:
[0089] An inflection point frequency band where the energy of each frame begins to decrease is determined, and the frequency of the inflection point frequency band is determined as the effective spectrum height of each frame.
[0090] Preferably, the determining unit is specifically configured to:
[0091] Obtaining a first frequency band corresponding to a maximum E2 value in each valid frame of the first audio;
[0092] If the E1 value of the first frequency band is greater than the second energy threshold, or the E2 value of the second frequency band adjacent to the first frequency band along the frequency axis is greater than the third energy threshold, the first frequency band is regarded as the inflection point frequency band, and the frequency of the inflection point frequency band is determined as the effective spectrum height of the effective frame, wherein the second energy threshold and the third energy threshold are used to characterize the energy variation range of the inflection point frequency band;
[0093] or,
[0094] If the energy of the first frequency band is less than the invalid energy threshold, the first frequency band is regarded as the inflection point frequency band, and the frequency of the inflection point frequency band is determined as the effective spectrum height of the effective frame.
[0095] Preferably, the determining unit is specifically configured to:
[0096] Calculating, according to a plurality of preset center frequency bands and preset high frequency points in a preset high frequency band of the target category audio, a plurality of energy differences between the plurality of preset center frequency bands and the preset high frequency points for each frame of audio in the second audio, wherein the frequencies of the plurality of preset center frequency bands increase sequentially but are all less than the frequency of the preset high frequency points, and wherein different categories of audio correspond to different plurality of preset center frequency bands and different preset high frequency points;
[0097] counting a number of second valid frames in which the multiple energy difference values in the second audio are all greater than corresponding multiple thresholds, wherein the multiple thresholds respectively corresponding to the multiple energy difference values decrease as the frequencies of the multiple preset center frequency bands increase;
[0098] If the ratio of the second valid frame number to the total audio frame number in the second audio is not less than a third ratio threshold, it indicates that the overall energy trend of the preset high frequency band in the second audio presents a downward trend.
[0099] Preferably, the determining unit is specifically configured to:
[0100] According to a preset high frequency band of the target audio category, obtaining the preset high frequency band of each audio frame in the second audio, wherein different high frequency band ranges are set for different audio categories;
[0101] Calculating an energy distribution interval of a preset high frequency band of each audio frame in the second audio;
[0102] Counting the number of third valid frames without abnormal frequency points within the energy distribution interval, wherein the abnormal frequency point is a frequency point whose frequency energy is greater than a critical value of the energy distribution interval;
[0103] If the ratio of the third number of valid frames to the total number of audio frames in the second audio is not less than a fourth ratio threshold, it indicates that no frequency point with sudden energy changes occurs in the preset high frequency band of the second audio;
[0104] Determine that the second audio is high-resolution audio.
[0105] Preferably, the determining unit is specifically configured to:
[0106] Slide each audio frame of the second audio using a sliding window with a single frequency point as a step size, wherein the sliding window is at least larger than the frequency spectrum range of the single frequency point;
[0107] For each audio frame in the second audio, calculating the sum of first-order differences of energies of all frequency points in the sliding window;
[0108] If the sum of the first-order differences of the energies of all frequency points in the sliding window is negative and greater than a fourth energy threshold, determining that the audio frame in the second audio is a normal frame;
[0109] or,
[0110] If the sign of the sum of the first-order differences of the energies of all frequency points in the sliding window is positive and less than a fifth energy threshold, determining that the audio frame in the second audio is a normal frame;
[0111] Counting a proportion of normal frames in the second audio; if the proportion of normal frames in the second audio is not less than a fifth proportion threshold, indicating that no frequency point with a sudden energy change occurs in the preset high frequency band of the second audio;
[0112] Determine that the second audio is high-resolution audio.
[0113] Preferably, the determining unit is specifically configured to:
[0114] Slide each audio frame of the second audio using a sliding window with a single frequency point as a step size, wherein the sliding window is at least larger than the frequency spectrum range of the single frequency point;
[0115] For each audio frame in the second audio, calculating the sum of the first-order differences of the energy of each frequency point in the sliding window;
[0116] If the sign of the sum of the first-order differences of the energies of each frequency point in the sliding window is negative and greater than a sixth energy threshold, determining that the audio frame in the second audio is a normal frame;
[0117] or,
[0118] If the sum of the first-order differences of the energies of each frequency point in the sliding window is positive and less than a seventh energy threshold, determining that the audio frame in the second audio is a normal frame;
[0119] Counting a proportion of normal frames in the second audio; if the proportion of normal frames in the second audio is not less than a sixth proportion threshold, indicating that no frequency point with sudden energy change occurs in the preset high frequency band of the second audio;
[0120] Determine that the second audio is high-resolution audio.
[0121] Preferably, the device further comprises:
[0122] The pre-processing unit is configured to perform smoothing processing on each audio frame in the second audio using a smoothing filter before sliding the sliding window across each audio frame in the second audio with a single frequency point as a step size.
[0123] An embodiment of the present application further provides a computer device comprising a processor, which, when executing a computer program stored in a memory, is used to implement the high-resolution audio quality detection method described in the first aspect of the embodiment of the present application.
[0124] An embodiment of the present application further provides a computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, it is used to implement the sound quality detection method of high-resolution audio described in the first aspect of the embodiment of the present application.
[0125] It can be seen from the above technical solutions that the embodiments of the present invention have the following advantages:
[0126] In an embodiment of the present application, a preliminary screening is first performed on the audio to be detected to screen out a first audio whose sampling rate and quantization accuracy meet the preset standards. Then, based on the effective energy, a second audio whose effective spectrum height is greater than the frequency threshold is screened out from the first audio. Finally, when it is confirmed that the energy of the preset high-frequency band of the second audio shows a downward trend and there are no frequency points with sudden energy changes in the preset high-frequency band, the second audio is confirmed to be high-resolution audio. That is, this embodiment of the present application, on the one hand, implements the screening of high-resolution audio based on specific quantization standards, avoids the involvement of human subjective will, and improves the accuracy of high-resolution audio recognition. On the other hand, the algorithm of the recognition standard can also be set on a computer device, thereby improving the efficiency of high-resolution audio recognition. BRIEF DESCRIPTION OF THE DRAWINGS
[0127] Figure 1 This is a schematic diagram of an embodiment of a method for detecting high-resolution audio in an embodiment of the present application;
[0128] Figure 2 In the embodiment of this application Figure 1 The detailed steps of steps 103 and 104 in the embodiment are as follows:
[0129] Figure 3 In the embodiment of this application Figure 2 A refinement of step 203 in the embodiment;
[0130] Figure 4 Schematic diagram of effective intra-frame frequency bands and frequencies in an embodiment of the present application;
[0131] Figure 5 In the embodiment of this application Figure 1 A refinement of step 105 in the embodiment;
[0132] Figure 6 This is a schematic diagram of another embodiment of the high-resolution audio detection method in the embodiment of the present application;
[0133] Figure 7 This is a schematic diagram of abnormal frequency points appearing in a valid frame in an embodiment of the present application;
[0134] Figure 8 This is a schematic diagram of another embodiment of the high-resolution audio detection method in the embodiment of the present application;
[0135] Figure 9 This is a schematic diagram of another embodiment of the high-resolution audio detection method in the embodiment of the present application;
[0136] Figure 10 Schematic diagram of an embodiment of a high-resolution audio detection device in an embodiment of the present application. DETAILED DESCRIPTION
[0137] The embodiments of the present invention provide a method and related apparatus for detecting the sound quality of high-resolution audio, which are used to improve the recognition accuracy of high-resolution audio.
[0138] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present invention.
[0139] The terms "first," "second," "third," "fourth," and the like in the specification and claims of the present invention and in the accompanying drawings are used to distinguish similar objects and are not necessarily used to describe a particular order or precedence. It should be understood that the terms used in this manner are interchangeable where appropriate so that the embodiments described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "including" and "having," and any variations thereof, are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or apparatus comprising a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0140] High-resolution audio refers to audio files with a sampling rate greater than 44100Hz or a quantization accuracy greater than 16 bits. To address the problem of low recognition accuracy caused by manual recognition of high-resolution audio in the existing technology, this application proposes a sound quality detection method and related devices for high-resolution audio to improve the recognition accuracy of high-resolution audio.
[0141] For ease of understanding, the following describes the sound quality detection method of high-resolution audio in this application. Figure 1 , Figure 1 A schematic diagram of an embodiment of a method for detecting sound quality of high-resolution audio:
[0142] 101. Obtain the sampling rate and quantization accuracy of the audio to be detected;
[0143] Because the quality of high-resolution audio is higher than that of CD, high-resolution audio is generally an audio file with a sampling rate greater than 44100 Hz or a quantization accuracy greater than 16 bits. Therefore, in order to detect high-resolution audio, the embodiment of the present application needs to first obtain the sampling rate and quantization accuracy of the audio to be detected.
[0144] Specifically, the sampling rate refers to the number of times the recording device samples the sound signal within a unit of time (1S), and the unit is Hz (Hertz). The quantization process is to convert each sample with continuous amplitude values into a discrete value representation, and the quantization accuracy refers to the number of binary bits occupied by each sample. At the same time, the number of binary bits also reflects the accuracy of measuring the amplitude of the sound waveform. Among them, the greater the quantization accuracy, the better the sound quality.
[0145] In addition, the audio to be detected in the embodiment of the present application can be a segment of audio randomly intercepted from the audio being played, or it can be a segment of audio selected from an audio library. The specific method of obtaining the audio to be detected is not limited here.
[0146] 102. Filter out a first audio whose sampling rate and quantization accuracy meet preset standards from the audio to be detected;
[0147] After the sampling rate and quantization accuracy of the audio to be detected are obtained, the first audio whose sampling rate and quantization accuracy meet the preset standard can be screened out from the audio to be detected according to the preset standard of high-resolution audio.
[0148] Specifically, the preset standard in the embodiment of the present application is a sampling rate greater than 44100 Hz, and a quantization accuracy greater than 16 bits. However, it should be noted that as users' requirements for audio quality increase, the preset standard in this application can also be other custom standards with a sampling rate and quantization accuracy greater than 44100 Hz and 16 bits respectively. The specific definition of the preset standard is not limited here.
[0149] 103. Determine an effective spectrum height of the first audio, where the effective spectrum height is a frequency corresponding to a location with maximum audio energy.
[0150] After selecting the first audio whose sampling rate and quantization accuracy meet the standards from the audio to be tested, in order to identify false high-resolution audio, that is, audio that meets the sampling rate and quantization accuracy requirements but has no effective energy in the spectrum above 44100Hz, the embodiment of the present application also needs to determine the effective spectrum height of the first audio, where the effective spectrum height is the frequency corresponding to the maximum energy of the audio. Therefore, in the embodiment of the present application, determining the effective spectrum height of the first audio is to determine whether the high frequency band of the first audio has effective energy.
[0151] 104. Filter out, from the first audio, a second audio whose effective spectrum height is greater than a frequency threshold;
[0152] In order to ensure that the first audio has effective energy in the high frequency band, a frequency domain threshold may be preset to filter out the second audio whose effective spectrum height is greater than the frequency domain threshold from the first audio.
[0153] In this step, the embodiment of the present application filters out a second audio frequency whose effective spectrum height is greater than a frequency threshold from the first audio frequency. The effective spectrum height is the frequency corresponding to the maximum energy of the audio frequency. Therefore, in this step, the effective spectrum height of the second audio frequency is greater than the frequency threshold, which reflects that the effective energy in the high frequency band of the second audio frequency is greater than a preset energy threshold, that is, there is effective energy in the high frequency band.
[0154] Specifically, the process of how to filter out the second audio from the first audio will be described in the following embodiments and will not be repeated here.
[0155] 105. Determine whether the overall energy trend of the preset high frequency band in the second audio presents a downward trend, and whether there is no frequency point with sudden energy change in the preset high frequency band. If so, execute step 106; if not, execute step 107.
[0156] After the second audio is filtered out from the first audio, it is further determined whether the overall energy trend of the preset high-frequency band in the second audio shows a downward trend, and whether there is no frequency point with sudden energy changes in the preset high-frequency band. If so, step 106 is executed; if not, step 107 is executed.
[0157] After filtering out the second audio, because the energy of high-resolution audio gradually shows a downward trend as the frequency increases, and high-resolution audio generally does not have burrs (i.e., sharp sounds) with sudden energy changes, the embodiment of the present application further determines whether the overall energy trend of a preset high-frequency band in the second audio shows a downward trend after filtering out the second audio, and whether there are no frequency points with sudden energy changes in the preset high-frequency band.
[0158] Specifically, the process of determining whether the energy of the preset high frequency band in the second audio presents a downward trend and whether no energy mutation point occurs in the preset high frequency band will be described in detail in the following embodiments and will not be repeated here.
[0159] 106. If the energy of the preset high frequency band in the second audio shows a downward trend and no frequency point with a sudden energy change occurs in the preset high frequency band, determine that the second audio is high-resolution audio.
[0160] If the energy of the preset high frequency band in the second audio presents a downward trend and no frequency point with a sudden energy change occurs in the preset high frequency band, the second audio is determined to be high-resolution audio.
[0161] It should be noted that in this application, different high-frequency bands are generally set according to the sampling rate of the second audio. For example, when the sampling rate is equal to 48000Hz, the high-frequency band is generally preset to 20000 to 24000Hz, and when the sampling rate is greater than 48000Hz and less than or equal to 96000Hz, the high-frequency band is generally preset to 24000Hz to 48000Hz, and when the sampling rate is greater than 96000Hz, the high-frequency band is generally preset to 24000Hz to 48000Hz.
[0162] 107. If the energy of the preset high frequency band in the second audio does not show a decreasing trend, and / or a frequency point with a sudden energy change occurs in the preset high frequency band, determine that the second audio is non-high-resolution audio.
[0163] Furthermore, if the energy of the preset high frequency band in the second audio does not show a decreasing trend, and / or a frequency point with a sudden energy change occurs in the preset high frequency band, the second audio is determined to be non-high-resolution audio.
[0164] In an embodiment of the present application, a preliminary screening is first performed on the audio to be detected to screen out a first audio whose sampling rate and quantization accuracy meet the preset standards. Then, based on the effective energy, a second audio whose effective spectrum height is greater than the frequency threshold is screened out from the first audio. Finally, when it is confirmed that the energy of the preset high-frequency band of the second audio shows a downward trend and there are no frequency points with sudden energy changes in the preset high-frequency band, the second audio is confirmed to be high-resolution audio. That is, this embodiment of the present application, on the one hand, implements the screening of high-resolution audio based on specific quantization standards, avoids the involvement of human subjective will, and improves the accuracy of high-resolution audio recognition. On the other hand, the algorithm of the recognition standard can also be set on a computer device, thereby improving the efficiency of high-resolution audio recognition.
[0165] based on Figure 1In the described embodiment, between step 102 and step 103, in order to achieve efficient screening of the second audio from the first audio, the sampling rate and preset sampling interval of the first audio can also be obtained, and the sampling rate of the first audio can be classified into the target category audio according to the preset sampling interval, and when the sampling rate of the first audio is not equal to the resampling rate corresponding to the target category audio, the first audio is resampled according to the resampling rate corresponding to the target category audio, so that the resampling rate of the first audio is the same as the resampling rate corresponding to the target category audio, wherein different categories of audio correspond to different resampling rates.
[0166] Specifically, because the first audio can be audio of any frequency range with a sampling rate greater than 44100 Hz, in the process of screening the second audio from the first audio, the determination of the effective spectrum height needs to perform corresponding conversion according to the frequency range of the first audio. Assuming that there are multiple first audios, if the frequency range of each first audio is different, it is necessary to perform a corresponding conversion according to the frequency range of each first audio.
[0167] In order to simplify this process, that is, to achieve efficient screening of the second audio from the first audio, we can preset different sampling intervals and set the audio in different sampling intervals as different categories of audio, so as to determine the effective spectrum height of the first audio according to the effective spectrum height determination process of the corresponding category, thereby achieving efficient screening of the second audio from the first audio.
[0168] As an optional implementation, the division can be carried out in the following manner, such as setting the audio with a sampling rate equal to 48000Hz as the first category of audio; setting the audio with a sampling rate greater than 48000Hz and less than or equal to 96000Hz as the second category of audio; and setting the audio with a sampling rate greater than 96000Hz as the third category of audio. For the first category of audio, the sampling rate of the audio is not processed, for the second category of audio, the resampling rate of the audio is unified to 96000Hz, and for the third category of audio, the resampling rate of the audio is unified to 192000Hz, because according to Shannon's sampling theorem, only when the sampling rate is twice the signal bandwidth (signal frequency bandwidth) can the original continuous signal be completely reconstructed from the sampled sample. Therefore, this application needs to resample the second category of audio and the third category of audio to 96000Hz and 192000Hz respectively.
[0169] It should be noted that when defining the sampling interval and the audio category, they can be defined according to different standards, and there is no specific limitation on the definition method of the sampling interval and the audio category.
[0170] Based on the above embodiments, Figure 1 Steps 103 and 104 in the embodiment are described in detail. Figure 2 for Figure 1 The detailed steps of steps 103 and 104 in the embodiment are as follows:
[0171] 201. Divide the first audio into frames, perform windowing, and perform time domain to frequency domain conversion to obtain a frequency domain signal of each frame of the first audio.
[0172] Because speech signals are generally continuous time-domain signals, in order to realize digital processing of speech signals, the speech signals need to be discretized first to obtain discrete periodic frequency-domain signals.
[0173] Specifically, as a discretization method for the speech signal, the first audio frame can be windowed. Because the speech signal is unstable at the macro level but stable at the micro level, that is, has short-term stability, the speech signal can be divided into several short segments for processing, and each short segment is a frame.
[0174] During the digitization of speech signals, it is necessary to truncate long periods of signal, or window the speech signal. This allows non-periodic speech signals to exhibit the characteristics of a periodic function. However, the windowing process weakens the signal at both ends of a frame, so when framing, overlap between frames is required. The specific window function can be a Hamming window or a rectangular window, and the specific form of the window function is not limited here.
[0175] After completing the framing and windowing of the speech signal, the time domain and frequency domain conversion can be performed on each frame of the speech signal to convert the speech signal from a time domain signal to a frequency domain signal. As a specific conversion method between the time domain and the frequency domain, it can be Fourier transform, short-time Fourier transform, etc., and no specific restrictions are made here.
[0176] 202. Calculate the energy of each audio frame in the first audio, and screen out valid frames in the first audio based on the energy of each audio frame in the first audio and a first energy threshold.
[0177] After obtaining the frequency domain signal of each frame signal in the first audio, the energy of each frame audio is further calculated to screen out valid frames in the first audio according to the energy of each frame audio and a first energy threshold.
[0178] Specifically, the high-resolution audio must have effective energy in a preset high frequency band. Therefore, this step can calculate the energy of each frame of audio and screen out effective frames in the first audio based on the energy of each frame of audio and the first energy threshold.
[0179] The energy of each frame of audio is calculated based on the square of the amplitude of the signal. For digital signals, the energy of each frame of speech signal is the sum of the squares of the amplitudes of the signals at each frequency point.
[0180] 203. Determine the effective spectrum height of the effective frame in the first audio;
[0181] After the valid frames in the first audio are screened out, the effective spectrum height of the valid frames in the first audio is further determined. Specifically, the process of determining the effective spectrum height will be described in the following embodiments and will not be repeated here.
[0182] 204. Count the number of first valid frames in the first audio whose effective spectrum height is greater than the frequency threshold;
[0183] After obtaining the effective spectrum height of each effective frame in the first audio, the number of first effective frames in which the effective spectrum height of the effective frames in the first audio is greater than the frequency threshold is counted, and step 205 is performed according to the number of first effective frames.
[0184] Specifically, the frequency domain threshold in this step is generally 20000 Hz.
[0185] 205. If the ratio of the number of first valid frames in the first audio to the total number of valid frames in the first audio is greater than a first ratio threshold, consider the first audio as the second audio of the same target category.
[0186] If the ratio of the number of the first valid frames to the total number of valid frames in the first audio is greater than a first ratio threshold, the first audio is regarded as the second audio of the same target category.
[0187] Assume that the total number of valid frames in the first audio is 100, the number of first valid frames is 80, and the first ratio threshold is 60%. Since 80% is greater than 60%, the first audio is determined to be the second audio.
[0188] Specifically, because the target category to which the first audio belongs is classified according to the sampling rate of the first audio and the preset sampling interval in the above process, after the second audio is filtered out from the first audio, the second audio is also correspondingly regarded as the same target category. For example, when the first audio is the first category audio, the second audio is also the first category audio, and when the first audio is the second category audio, the second audio is also the second category audio, so as to facilitate further processing of the first audio and the second audio in the later stage.
[0189] In the embodiment of the present application, a process of filtering out the second audio from the first audio is described in detail, thereby improving the reliability of the process of filtering out the second audio from the first audio.
[0190] Further, in the implementation Figure 2In the process of the embodiment, that is, in the process of filtering out the second audio from the first audio, in order to save computational complexity, the ratio of the number of valid frames in the first audio to the total number of audio frames in the first audio can be counted after step 202 and before step 203. If the ratio is less than the second ratio threshold, the first audio is directly determined to be non-high-resolution audio, thereby saving the computational complexity of steps 203 to 205 and improving the detection efficiency of high-resolution audio.
[0191] based on Figure 2 The embodiment described above is described below. Figure 2 For a detailed description of step 203 in the embodiment, please refer to Figure 3 , Figure 3 for Figure 2 The detailed steps of step 203 in the embodiment are as follows:
[0192] 301. Divide valid frames in the first audio into M frequency bands according to a preset frequency interval, each frequency band including N frequency points, where M is greater than or equal to 2, and N is greater than or equal to 1;
[0193] Figure 2 After obtaining the frequency domain signal of the speech frame in the embodiment, each frame of the speech signal can be divided into M frequency bands according to a preset frequency interval, such as 100 Hz or 500 Hz, where each frequency band includes N frequency points, where M is greater than or equal to 2 and N is greater than or equal to 1.
[0194] For ease of understanding, Figure 4 A schematic diagram of a frame signal is given in which a valid frame is divided into four frequency bands, and each frequency band includes three frequency points.
[0195] 302. Calculate, according to the frequency band calculation range corresponding to the target audio type, the first-order frequency band energy E1 of each frequency band and the sum E2 of the energies of two adjacent first-order frequency bands along the frequency axis in each valid frame of the first audio. Different audio types correspond to different frequency band calculation ranges.
[0196] In order to facilitate the determination of the effective spectrum height of the valid frame in the first audio, the audio is classified according to the preset sampling interval in the aforementioned embodiment. For the sake of consistency of description, the aforementioned category classification standard and the setting range of the high-frequency band are used here. For example, audio with a sampling rate equal to 48,000 Hz is set as the first category of audio; audio with a sampling rate greater than 48,000 Hz and less than or equal to 96,000 Hz is set as the second category of audio; audio with a sampling rate greater than 96,000 Hz is set as the third category of audio; when the sampling rate is equal to 48,000 Hz, the high-frequency band is generally preset to 20,000 to 24,000 Hz, and when the sampling rate is greater than 48,000 Hz and less than or equal to 96,000 Hz, the high-frequency band is generally preset to 24,000 Hz to 48,000 Hz, and when the sampling rate is greater than 96,000 Hz, the high-frequency band is generally preset to 24,000 Hz to 48,000 Hz.
[0197] For the convenience of description, assuming that the first audio is the first category audio, the frequency band calculation range of the first category audio is up to 24000Hz. Figure 4 For each frequency point in , assume that they are marked as 1, 2, 3...12 frequency points, and the frequency bands are marked as the first frequency band, the second frequency band, the third frequency band, and the fourth frequency band, respectively. Calculate the energy of each frequency point (wherein the frequency range of the highest frequency point cannot be greater than 24000 Hz), where the energy of the frequency point is the square of the amplitude value of the frequency point signal.
[0198] After obtaining the energy of each frequency point, the energy of the corresponding frequency band is the average of the frequency point energies. For example, the energy of the first frequency band is the average of the energies of frequency points 1, 2, and 3, the energy of the second frequency band is the average of the energies of frequency points 4, 5, and 6, the third frequency band is the average of the energies of frequency points 7, 8, and 9, and the fourth frequency band is the average of the energies of frequency points 10, 11, and 12. Specifically, when calculating the frequency band energy, in order to avoid the impact of abnormal frequency points on the frequency band energy, it is also possible to remove one maximum value and one minimum value in each frequency band and then calculate the average value.
[0199] After obtaining the energy of each frequency band, the first-order frequency band energy E1 of the first frequency band is the difference between the energy of the second frequency band and the energy of the first frequency band. Assuming that the energy of the first frequency band is X1, the energy of the second frequency band is X2, the energy of the third frequency band is X3, and the energy of the fourth frequency band is X4, then E1 = X2-X1; the first-order frequency band energy E1 of the second frequency band is X3-X2, the first-order frequency band energy E1 of the third frequency band is X4-X3, and the sum of the energies of two adjacent first-order frequency bands E2 is X3-X1 and X4-X2 respectively.
[0200] 303 : Determine an effective spectrum height of each effective frame in the first audio according to E1 or E2 corresponding to each frequency band in each effective frame in the first audio and an effective spectrum height determination rule.
[0201] After obtaining E1 and E2 corresponding to each frequency band in each valid frame of the first audio, the effective spectrum height of each valid frame in the first audio is further determined according to E1 and E2 corresponding to each frequency band and an effective spectrum height determination rule.
[0202] Specifically, the effective spectrum height determination rule includes: determining an inflection point frequency band where the energy of the effective frame begins to decrease, and determining the frequency of the inflection point frequency band as the effective spectrum height of the effective frame.
[0203] The following describes the process of determining the inflection point frequency band where the effective frame energy begins to decrease, combining E1 and E2 corresponding to each frequency band:
[0204] Determine a first frequency band corresponding to a maximum E2 value in each valid frame of the first audio; if the E1 value of the first frequency band is greater than a second energy threshold, or the E2 value of a second frequency band adjacent to the first frequency band along the frequency axis is greater than a third energy threshold, then consider the first frequency band to be an inflection point frequency band, where the second energy threshold and the third energy threshold are used to characterize an energy range of the inflection point frequency band; or, if the energy of the first frequency band is less than an invalid energy threshold, then consider the first frequency band to be the inflection point frequency band, and determine the frequency of the inflection point frequency band as the effective spectrum height of the valid frame.
[0205] Assuming that in step 302, the maximum value of E2 is X3-X1, then the first frequency band corresponding to the maximum E2 value is Figure 4 If the E1 value (i.e. X2-X1) of the first frequency band is greater than the second energy threshold (the empirical value is 20dB), or Figure 4 The E2 (i.e. X4-X2) value of the second frequency band is greater than the third energy threshold (the empirical value is 40dB), then it is determined Figure 4 The first frequency band in is the inflection point frequency band in the valid frame, and the frequency of the first frequency band is the effective spectrum height of the valid frame.
[0206] When the above rules do not apply, that is, when the inflection point frequency band cannot be determined according to E1 and E2, it is further determined whether the energy of the first frequency band corresponding to the maximum E2 value is less than the invalid energy threshold. If so, the first frequency band is regarded as the inflection point frequency band, and the frequency of the inflection point frequency band is determined as the effective spectrum height of the valid frame.
[0207] The embodiment of the present application describes in detail the process of determining the effective spectrum height of each effective frame in the first audio, thereby improving the reliability of determining the effective spectrum height of each effective frame in the embodiment of the present application.
[0208] Based on the above embodiment, since the energy of the high-resolution audio in the preset high frequency band is decreasing, the embodiment of the present application describes the judgment process of whether the audio in the preset high frequency band of the second resolution audio is decreasing. Figure 5 , Figure 5 for Figure 1 The detailed steps of step 105 in the embodiment are as follows:
[0209] 501. Calculate, based on a plurality of preset center frequency bands and preset high frequency points in a preset high frequency band of the target audio category, a plurality of energy differences between the plurality of preset center frequency bands and the preset high frequency points for each frame of the second audio, wherein the frequencies of the plurality of preset center frequency bands increase sequentially but are all less than the frequency of the preset high frequency points, and different audio categories correspond to different plurality of preset center frequency bands and different preset high frequency points;
[0210] Because the energy of high-resolution audio in a preset high frequency band shows a downward trend, the embodiment of the present application describes the judgment process of whether the energy of a preset high frequency band in the second audio shows a downward trend. Based on the consistency with the description of the previous embodiment, the present application still adopts the following classification standard: audio with a sampling rate equal to 48000Hz is set as the first category of audio; audio with a sampling rate greater than 48000Hz and less than or equal to 96000Hz is set as the second category of audio; audio with a sampling rate greater than 96000Hz is set as the third category of audio.
[0211] In the embodiment of the present application, for the first type of audio, the preset high frequency band is 20,000 Hz to 24,000 Hz, and for the second and third types of audio, the preset high frequency band is from 24,000 Hz to 48,000 Hz. For the first type of audio, multiple preset center frequency bands are 18,000 Hz, 20,000 Hz, and 22,000 Hz, and the preset high frequency point is 24,000 Hz; for the second type of audio, multiple preset center frequency bands are 26,000 Hz, 36,000 Hz, and 42,000 Hz, and the preset high frequency point is 48,000 Hz; for the third type of audio, multiple preset center frequency bands are 26,000 Hz, 36,000 Hz, and 46,000 Hz, and the preset high frequency point is 48,000 Hz.
[0212] It should be noted that for multiple preset center frequency bands and preset high frequency points in different categories, users can customize them according to their actual measurement needs. The multiple preset center frequency bands and preset high frequency points in the above three types of audio are only a specific implementation method. There is no specific restriction on the values of the multiple preset center frequency bands and preset high frequency points in each type of audio.
[0213] Assuming that the second audio is the first type of audio, the first energy difference value from the frequency band centered on 18000 Hz to the frequency band centered on 24000 Hz for each valid frame in the second audio, the second energy difference value from the frequency band centered on 20000 Hz to the frequency band centered on 24000 Hz for each valid frame, and the third energy difference value from the frequency band centered on 22000 Hz to the frequency band centered on 24000 Hz for each valid frame are calculated respectively, and the relationship between the first energy difference value, the second energy difference value and the third energy difference value and the corresponding energy difference threshold value are further compared.
[0214] It is easy to understand that in order to reflect the downward trend of the energy in the high frequency band of the first audio, the three energy difference thresholds corresponding to the first energy difference, the second energy difference, and the third energy difference respectively show a decreasing trend.
[0215] If the second audio belongs to the second category of audio or the third category of audio, the same method is used to respectively calculate multiple energy differences between multiple preset center frequency bands and preset high frequency points.
[0216] 502. Count the number of second valid frames in the second audio in which the multiple energy difference values are all greater than corresponding multiple thresholds, wherein the multiple thresholds respectively corresponding to the multiple energy difference values decrease as the frequencies of the multiple preset center frequency bands increase;
[0217] After obtaining multiple energy differences between multiple preset high-frequency center frequency bands and preset high-frequency points in the second audio, the number of second valid frames in which the multiple energy differences in the second audio are all greater than corresponding multiple thresholds is further counted.
[0218] 503. Determine whether the ratio of the second valid frame number to the total audio frame number in the second audio is not less than a third ratio threshold; if so, execute step 504; if not, execute step 505.
[0219] After obtaining the number of second valid frames in the second audio for which multiple energy differences are all greater than the corresponding multiple thresholds, it is further determined whether the ratio of the number of second valid frames to the total number of audio frames in the second audio is not less than a third ratio threshold. If so, step 504 is executed; if not, step 505 is executed.
[0220] Specifically, the third ratio threshold here may be 60% or 70%, and the specific value of the third ratio threshold is not specifically limited here.
[0221] 504. If the ratio of the second valid frame number to the total audio frame number in the second audio is not less than a third ratio threshold, it indicates that the overall energy trend of the preset high frequency band in the second audio is decreasing.
[0222] If the ratio of the second valid frame number to the total audio frame number in the second audio is not less than the third ratio threshold, it indicates that the overall energy trend of the preset high frequency band in the second audio presents a downward trend.
[0223] 505. If the ratio of the second number of valid frames to the total number of audio frames in the second audio is less than a third ratio threshold, it indicates that the overall trend of energy in the preset high frequency band in the second audio does not show a downward trend.
[0224] If the ratio of the second valid frame number to the total audio frame number in the second audio is less than the third ratio threshold, it indicates that the overall energy trend of the preset high frequency band in the second audio does not show a downward trend.
[0225] In the embodiment of the present application, a process of determining whether the energy of the preset high frequency band in the second audio presents a downward trend is described in detail, thereby improving the reliability of the process of determining whether the energy of the preset high frequency band in the second audio presents a downward trend.
[0226] Based on the above embodiment, the sound quality detection method of high-resolution audio in the embodiment of the present application is described in detail below. Figure 6 , Figure 6 Another embodiment of the high-resolution audio quality detection method in the embodiment of this application is as follows:
[0227] 601. Obtain the sampling rate and quantization accuracy of the audio to be detected;
[0228] 602. Filter out a first audio whose sampling rate and quantization accuracy both meet preset standards from the audio to be detected;
[0229] 603. Determine an effective spectrum height of the first audio, where the effective spectrum height is a frequency corresponding to a location with maximum audio energy.
[0230] 604. Filter out, from the first audio, a second audio whose effective spectrum height is greater than a frequency threshold;
[0231] It should be noted that steps 601 to 604 in the embodiment of the present application are the same as Figure 1 The descriptions of steps 101 to 104 in the embodiment are similar and will not be repeated here.
[0232] 605. Calculate, based on the multiple preset center frequency bands and the preset high frequency points in the preset high frequency band of the target audio category, multiple energy differences between the multiple preset center frequency bands and the preset high frequency points for each frame of the second audio, wherein the frequencies of the multiple preset center frequency bands increase sequentially but are all lower than the frequency of the preset high frequency points, and different audio categories correspond to different multiple preset center frequency bands and different preset high frequency points;
[0233] 606. Count the number of second valid frames in the second audio in which the multiple energy difference values are all greater than corresponding multiple thresholds, wherein the multiple thresholds respectively corresponding to the multiple energy difference values decrease as the frequencies of the multiple preset center frequency bands increase;
[0234] 607 . Determine whether the ratio of the second valid frame number to the total audio frame number in the second audio is not less than a third ratio threshold; if so, execute step 608 ; otherwise, execute step 609 .
[0235] 608. If the ratio of the second number of valid frames to the total number of audio frames in the second audio is not less than a third ratio threshold, it indicates that the overall trend of energy in the preset high frequency band in the second audio is decreasing.
[0236] 609. If the ratio of the second number of valid frames to the total number of audio frames in the second audio is less than a third ratio threshold, it indicates that the overall trend of energy in the preset high frequency band in the second audio does not show a downward trend.
[0237] It should be noted that the description of steps 605 to 609 in the embodiment of the present application is the same as Figure 5 The descriptions of steps 501 to 505 in the embodiment are similar and will not be repeated here.
[0238] 610. Obtain the preset high frequency band of each audio frame in the second audio according to the preset high frequency band of the target audio category, where different high frequency band ranges are set for different audio categories;
[0239] The classification intervals and classification standards of the audio, as well as the preset high frequency bands in each type of audio are consistent with those described in the above embodiment and will not be repeated here.
[0240] Assuming that the second audio belongs to the first category of audio, the preset high frequency band of each audio frame in the second audio is obtained respectively according to the preset high frequency point of the target category audio (that is, the preset high frequency band of the first category of audio), where the preset high frequency band of the first category of audio is 20,000 Hz to 24,000 Hz. For the second and third categories of audio, the preset high frequency band is from 24,000 Hz to 48,000 Hz.
[0241] 611. Calculate an energy distribution interval of a preset high frequency band of each audio frame in the second audio;
[0242] After obtaining the preset high frequency band of each audio frame in the second audio, the energy distribution range of the preset high frequency band of each audio frame in the second audio frame is calculated respectively. Among them, the energy distribution range between 20000 Hz and 24000 Hz is calculated for the first type of audio, and the energy distribution range between 24000 Hz and 48000 Hz is calculated for the second and third types of audio.
[0243] 612. Count the number of third valid frames without abnormal frequency points within the energy distribution interval, where an abnormal frequency point is a frequency point whose energy is greater than a critical value of the energy distribution interval;
[0244] After obtaining the energy distribution interval of the preset high frequency band of each audio frame in the second audio, the number of third valid frames without abnormal frequency points in the energy distribution interval is further counted, wherein the abnormal frequency point is a frequency point whose frequency band energy is greater than the critical value of the energy distribution interval.
[0245] For example, assuming that the energy distribution range of the preset high frequency band of each audio frame of the second audio is [-75dB, -100dB], if there is a frequency point with energy greater than -75dB or energy less than -100dB within the energy range [-75dB, -100dB], then the frequency point is considered an abnormal frequency point.
[0246] 613. Determine whether the ratio of the third number of valid frames to the total number of audio frames in the second audio is not less than a fourth ratio threshold; if so, execute step 614; if not, execute step 615.
[0247] After obtaining the third valid frame number without abnormal frequency points in the energy distribution range, further determine whether the ratio of the third valid frame number to the total number of audio frames in the second audio is not less than a fourth ratio threshold. If so, execute step 614; if not, execute step 615.
[0248] Specifically, the fourth ratio threshold here may be 60% or 70%, and the specific value of the fourth ratio threshold is not specifically limited here.
[0249] At step 614 , if the ratio of the third number of valid frames to the total number of audio frames in the second audio is not less than a fourth ratio threshold, it indicates that no frequency point with a sudden energy change occurs in the preset high frequency band of the second audio, and the second audio is determined to be high-resolution audio.
[0250] After obtaining the number of third valid frames in the second audio, the ratio of the number of third valid frames to the total number of audio frames in the second audio is further obtained. If the ratio is not less than a fourth ratio threshold, it indicates that there is no frequency point with energy mutation in the preset high frequency band of the second audio, and the second audio is determined to be high-resolution audio.
[0251] 615. If the ratio of the third number of valid frames to the total number of audio frames in the second audio is less than a fourth ratio threshold, it indicates that a frequency point with a sudden energy change occurs in the preset high frequency band of the second audio, and the second audio is determined to be non-high-resolution audio.
[0252] After obtaining the number of third valid frames in the second audio, the ratio of the number of third valid frames to the total number of audio frames in the second audio is further obtained. If the ratio is less than the fourth ratio threshold, it indicates that a frequency point with a sudden energy change occurs within the preset high frequency band of the second audio, and the second audio is determined to be non-high-resolution audio.
[0253] In the embodiment of the present application, a process of determining whether there is an energy mutation frequency point within a preset high frequency band is described in detail, thereby improving the reliability of the process.
[0254] based on Figure 6 The above embodiment describes the process of determining whether there is a sudden energy change frequency point in a frame from the perspective of the frame. Figure 6 In the judgment process of the embodiment, it is assumed that a frequency point with a sudden energy change occurs at the end of the frame, such as Figure 7 When an abnormal frequency point is detected in the frame, the frequency point may be ignored due to the average method used in the energy calculation process. Therefore, the embodiment of the present application can also describe the abnormal frequency point judgment process from the perspective of the frequency band within the frame. Please refer to Figure 8 , Figure 8 Another embodiment of the high-resolution audio quality detection method in the embodiment of this application is as follows:
[0255] 801. Obtain the sampling rate and quantization accuracy of the audio to be detected;
[0256] 802. Filter out a first audio whose sampling rate and quantization accuracy both meet preset standards from the audio to be detected;
[0257] 803. Determine an effective spectrum height of the first audio, where the effective spectrum height is a frequency corresponding to a location with maximum audio energy.
[0258] 804. Filter out, from the first audio, a second audio whose effective spectrum height is greater than a frequency threshold;
[0259] It should be noted that steps 801 to 804 in the embodiment of the present application are the same as Figure 1 The descriptions of steps 101 to 104 in the embodiment are similar and will not be repeated here.
[0260] 805. Calculate, based on the plurality of preset center frequency bands and the preset high frequency points in the preset high frequency band of the target audio category, a plurality of energy differences between the plurality of preset center frequency bands and the preset high frequency points for each frame of the second audio, wherein the frequencies of the plurality of preset center frequency bands increase sequentially but are all less than the frequency of the preset high frequency points, and different audio categories correspond to different plurality of preset center frequency bands and different preset high frequency points;
[0261] 806. Count the number of second valid frames in the second audio in which the multiple energy difference values are all greater than corresponding multiple thresholds, wherein the multiple thresholds respectively corresponding to the multiple energy difference values decrease as the frequencies of the multiple preset center frequency bands increase;
[0262] 807 . Determine whether the ratio of the second valid frame number to the total audio frame number in the second audio is not less than a third ratio threshold; if so, execute step 808 ; otherwise, execute step 809 .
[0263] 808. If the ratio of the second number of valid frames to the total number of audio frames in the second audio is not less than a third ratio threshold, it indicates that the overall trend of energy in the preset high frequency band in the second audio is decreasing.
[0264] 809. If the ratio of the second number of valid frames to the total number of audio frames in the second audio is less than a third ratio threshold, it indicates that the overall trend of energy in the preset high frequency band in the second audio does not show a downward trend.
[0265] It should be noted that the description of steps 805 to 809 in the embodiment of the present application is the same as Figure 5 The descriptions of steps 501 to 505 in the embodiment are similar and will not be repeated here.
[0266] 810. Use a sliding window to slide across each audio frame of the second audio with a single frequency point as a step size, wherein the sliding window is at least larger than the frequency spectrum range of the single frequency point;
[0267] After confirming in step 808 that the energy of the preset high frequency band in the second audio shows a downward trend, a sliding window can be further used to slide across each audio frame in the second audio with a single frequency point as a step size, wherein the sliding window is at least larger than the spectral range of the single frequency point.
[0268] 811. For each audio frame in the second audio, calculate the sum of first-order differences of energy of all frequency points in the sliding window;
[0269] For ease of description, assuming that the size of the sliding window is a spectrum range of 3 frequency points, the sum of the first-order differences of all frequency points in the sliding window corresponds to the difference between the energy of the third frequency point and the energy of the first frequency point, the difference between the energy of the fourth frequency point and the energy of the second frequency point, the difference between the energy of the fifth frequency point and the third frequency point, and so on.
[0270] 812. Determine whether the sign of the sum of the first-order differences of the energies of all frequency points in the sliding window is negative and greater than the fourth energy threshold, or whether the sign of the sum of the first-order differences of the energies of all frequency points in the sliding window is positive and less than the fifth energy threshold. If so, execute step 813; if not, execute step 814.
[0271] If the sign of the sum of the first-order differences of the energies of all frequency points in the sliding window is negative and is greater than the fourth energy threshold, the audio frame in the second audio is determined to be a normal frame. Because the sign of the sum of the first-order differences of all frequency points in the sliding window is negative, it means that the energy of the frequency points in the sliding window shows a downward trend. If it is greater than the fourth energy threshold, it means that there are no frequency points in the sliding window whose energy suddenly drops, that is, there are no frequency points whose energy suddenly changes.
[0272] If the sign of the sum of the first-order differences of the energies of all frequency points in the sliding window is positive and is less than the fifth energy threshold, the audio frame in the second audio is determined to be a normal frame. This is because the sign of the sum of the first-order differences of all frequency points in the sliding window is positive, indicating that the energy of the frequency points in the sliding window shows an upward trend. If it is less than the fifth energy threshold, it means that although the energy of the frequency points in the sliding window shows an upward trend, the energy of the frequency points with increasing energy does not undergo a sudden change, that is, there is no frequency point with a sudden change in energy.
[0273] Because under the premise that the energy of the preset high frequency band of the frame shows a downward trend, the frequency point of the audio frame is allowed to have an energy spiral increase or spiral decrease, as long as there is no frequency point with sudden energy changes.
[0274] Therefore, the embodiment of the present application can determine whether the sign of the sum of the first-order differences of the energy of all frequency points in the sliding window is negative and greater than the fourth energy threshold, or the sign of the sum of the first-order differences of the energy of all frequency points in the sliding window is positive and less than the fifth energy threshold, and if so, execute step 813, if not, execute step 814.
[0275] 813. If the sign of the sum of the first-order differences of the energies of all frequency points in the sliding window is negative and greater than a fourth energy threshold, or if the sign of the sum of the first-order differences of the energies of all frequency points in the sliding window is positive and less than a fifth energy threshold, determine that the audio frame in the second audio is a normal frame;
[0276] If the sign of the sum of the first-order differences of the energies of all the frequency points in the sliding window is negative and greater than a fourth energy threshold, or if the sign of the sum of the first-order differences of the energies of all the frequency points in the sliding window is positive and less than a fifth energy threshold, it indicates that no frequency point with a sudden energy change occurs in the sliding window, and the audio frame in the second audio is determined to be a normal frame;
[0277] 814. If the sign of the sum of the first-order differences of the energies of all frequency points in the sliding window is negative and not greater than a fourth energy threshold, and / or if the sign of the sum of the first-order differences of the energies of all frequency points in the sliding window is positive and not less than a fifth energy threshold, determine that the audio frame in the second audio is an abnormal frame.
[0278] If the sign of the sum of the first-order differences of the energies of all frequency points in the sliding window is negative and is not greater than the fourth energy threshold, and / or if the sign of the sum of the first-order differences of the energies of all frequency points in the sliding window is positive and is not less than the fifth energy threshold, it indicates that a frequency point with a sudden change in energy has appeared in the sliding window, and the audio frame in the second audio is determined to be an abnormal frame.
[0279] 815 . Count the proportion of normal frames in the second audio and determine whether the proportion of normal frames in the second audio is not less than a fifth proportion threshold. If so, execute step 816 ; otherwise, execute step 817 .
[0280] After calculating the ratio of normal frames in the second audio, it is further determined whether the ratio of normal frames in the second audio is not less than a fifth ratio threshold. If so, step 816 is executed; otherwise, step 817 is executed.
[0281] 816. A frequency point indicating that no energy mutation occurs within the preset high frequency band of the second audio, and determining that the second audio is high-resolution audio.
[0282] A ratio of normal frames in the second audio is counted. If the ratio of normal frames in the second audio is not less than a fifth ratio threshold, it indicates that no frequency point with a sudden energy change occurs in the preset high frequency band of the second audio, and the second audio is determined to be high-resolution audio.
[0283] 817. Indicate a frequency point at which energy mutation occurs within a preset high frequency band of the second audio, and determine that the second audio is non-high-resolution audio.
[0284] A ratio of normal frames in the second audio is counted. If the ratio of normal frames in the second audio is less than a fifth ratio threshold, it indicates that a frequency point with a sudden energy change occurs in a preset high frequency band of the second audio, and the second audio is determined to be non-high-resolution audio.
[0285] In the embodiment of the present application, the process of not having an energy mutation frequency point within the preset high frequency band of the second audio is described from the perspective of the intra-frame frequency band, and the process of identifying the energy mutation frequency point from the perspective of the intra-frame frequency band is more accurate than the process of identifying the energy mutation frequency point from the perspective of the frame.
[0286] based on Figure 8 In order to achieve a finer granularity in identifying the energy mutation frequency point, the embodiment described above can also be used to describe the process of not having an energy mutation frequency point within the second audio preset high frequency point from the perspective of frequency point. Figure 9 , Figure 9 Another embodiment of the high-resolution audio quality detection method in the embodiment of this application is as follows:
[0287] 901. Obtain the sampling rate and quantization accuracy of the audio to be detected;
[0288] 902. Filter out a first audio whose sampling rate and quantization accuracy meet preset standards from the audio to be detected;
[0289] 903. Determine an effective spectrum height of the first audio, where the effective spectrum height is a frequency corresponding to a location with maximum audio energy.
[0290] 904. Filter out, from the first audio, a second audio whose effective spectrum height is greater than a frequency threshold;
[0291] It should be noted that steps 901 to 904 in the embodiment of the present application are the same as Figure 1 The descriptions of steps 101 to 104 in the embodiment are similar and will not be repeated here.
[0292] 905. Calculate, based on the plurality of preset center frequency bands and the preset high frequency points in the preset high frequency band of the target audio category, a plurality of energy differences between the plurality of preset center frequency bands and the preset high frequency points for each frame of the second audio, wherein the frequencies of the plurality of preset center frequency bands increase sequentially but are all less than the frequency of the preset high frequency points, and different audio categories correspond to different plurality of preset center frequency bands and different preset high frequency points;
[0293] 906. Count the number of second valid frames in the second audio in which the multiple energy difference values are all greater than corresponding multiple thresholds, wherein the multiple thresholds respectively corresponding to the multiple energy difference values decrease as the frequencies of the multiple preset center frequency bands increase;
[0294] 907. Determine whether the ratio of the second valid frame number to the total audio frame number in the second audio is not less than a third ratio threshold; if so, execute step 908; otherwise, execute step 909.
[0295] 908. If the ratio of the second number of valid frames to the total number of audio frames in the second audio is not less than a third ratio threshold, it indicates that the overall trend of energy in the preset high frequency band in the second audio is decreasing.
[0296] 909. If the ratio of the second number of valid frames to the total number of audio frames in the second audio is less than a third ratio threshold, it indicates that the overall trend of energy in the preset high frequency band in the second audio does not show a downward trend.
[0297] It should be noted that the description of steps 905 to 909 in the embodiment of the present application is the same as Figure 5 The descriptions of steps 501 to 505 in the embodiment are similar and will not be repeated here.
[0298] 910. Use a sliding window to slide across each audio frame of the second audio with a single frequency point as a step size, wherein the sliding window is at least larger than the frequency spectrum range of the single frequency point;
[0299] After confirming in step 908 that the energy of the preset high frequency band in the second audio shows a downward trend, a sliding window can be further used to slide across each audio frame in the second audio with a single frequency point as a step size, wherein the sliding window is at least larger than the spectral range of the single frequency point.
[0300] 911. For each audio frame in the second audio, calculate the sum of the first-order differences of the energy of each frequency point in the sliding window;
[0301] For the convenience of description, it is also assumed that the size of the sliding window is a spectrum range of 3 frequency points, which is different from Figure 8 In the embodiment, the sum of the first-order differences of all frequency points in the sliding window is calculated. Figure 9 It is the sum of the first-order differences of each frequency point of the statistical sliding, and the sum of the first-order differences of each frequency point in the sliding window is the energy difference between the second frequency point and the first frequency point, the energy difference between the third frequency point and the second frequency point, the energy difference between the fourth frequency point and the third frequency point, and so on.
[0302] 912. Determine whether the sign of the sum of the first-order differences of the energy of each frequency point in the sliding window is negative and greater than the sixth energy threshold; or whether the sign of the sum of the first-order differences of the energy of each frequency point in the sliding window is positive and less than the seventh energy threshold. If so, execute step 913; if not, execute step 914.
[0303] Because the sign of the sum of the first-order differences of the energy of each frequency point in the sliding window is negative and is greater than the sixth energy threshold, it means that the energy of the frequency points in the sliding window shows a downward trend. If it is greater than the sixth energy threshold, it means that there is no frequency point in the sliding window whose energy suddenly drops, that is, there is no frequency point whose energy suddenly changes.
[0304] If the sum of the first-order differences of the energies of each frequency point in the sliding window is positive and less than the seventh energy threshold, it means that the energy of the frequency points in the sliding window shows an upward trend, and if it is less than the seventh energy threshold, it means that there is no frequency point in the sliding window whose energy suddenly increases, that is, there is no frequency point whose energy suddenly changes.
[0305] Because under the premise that the overall energy trend of the preset high-frequency band within the frame shows a downward trend, the frequency points of the audio frame are allowed to have energy spiraling up or spiraling down, as long as there are no frequency points with sudden energy changes.
[0306] Therefore, the embodiment of the present application can determine whether the sign of the sum of the first-order differences of the energy of each frequency point in the sliding window is negative and greater than the sixth energy threshold; or whether the sign of the sum of the first-order differences of the energy of each frequency point in the sliding window is positive and less than the seventh energy threshold, and if so, execute step 913, and if not, execute step 914.
[0307] 913. If the sign of the sum of the first-order differences of the energies of each frequency point in the sliding window is negative and greater than a sixth energy threshold, or if the sign of the sum of the first-order differences of the energies of each frequency point in the sliding window is positive and less than a seventh energy threshold, determine that the audio frame in the second audio is a normal frame;
[0308] If the sign of the sum of the first-order differences of the energy of each frequency point in the sliding window is negative and greater than the sixth energy threshold, or if the sign of the sum of the first-order differences of the energy of each frequency point in the sliding window is positive and less than the seventh energy threshold, it is determined that the audio frame in the second audio is a normal frame.
[0309] 914. If the sign of the sum of the first-order differences of the energies of each frequency point in the sliding window is negative and is not greater than a sixth energy threshold, and / or if the sign of the sum of the first-order differences of the energies of each frequency point in the sliding window is positive and is not less than a seventh energy threshold, determine that the audio frame in the second audio is an abnormal frame.
[0310] If the sign of the sum of the first-order differences of the energy of each frequency point in the sliding window is negative and not greater than the sixth energy threshold, and / or if the sign of the sum of the first-order differences of the energy of each frequency point in the sliding window is positive and not less than the seventh energy threshold, then the audio frame in the second audio is determined to be an abnormal frame.
[0311] 915. Count the proportion of normal frames in the second audio and determine whether the proportion of normal frames in the second audio is not less than a sixth proportion threshold. If so, execute step 916; otherwise, execute step 917.
[0312] Because the proportion of normal frames in the second audio frame exceeds a certain threshold, it indicates that there is no frequency point with a sudden energy change in the high frequency band of the second audio frame, and step 916 is executed. If the proportion of normal frames in the second audio frame does not exceed the certain threshold, it indicates that there is a frequency point with a sudden energy change in the high frequency band of the second audio frame, and step 917 is executed.
[0313] It should be noted that the sixth ratio threshold here can be 50% or 60%, and the user can customize it according to actual conditions, and no specific limitation is made here.
[0314] 916. Indicate that there is no frequency point at which energy mutation occurs in the preset high frequency band of the second audio, and determine that the second audio is high-resolution audio.
[0315] If the ratio of normal frames in the second audio frame exceeds a sixth ratio threshold, it indicates that no frequency point with sudden energy change occurs in the high frequency band of the second audio frame, and the second audio is further determined to be high-resolution audio.
[0316] 917. Indicate a frequency point at which a sudden energy change occurs in the preset high frequency band of the second audio, and determine that the second audio is unresolved audio.
[0317] If the ratio of normal frames in the second audio frame does not exceed the sixth ratio threshold, it indicates that a frequency point with a sudden energy change occurs in the high frequency band of the second audio frame, and the second audio is further determined to be non-high-resolution audio.
[0318] In the embodiment of the present application, the process of not having an energy mutation frequency point in the preset high frequency band of the second audio is described from the perspective of the intra-frame frequency point, and the process of identifying the energy mutation frequency point from the perspective of the intra-frame frequency point is more accurate than the process of identifying the energy mutation frequency point from the perspective of the intra-frame frequency band.
[0319] Further, based on Figure 8 or Figure 9 In the embodiment described above, before using a sliding window to slide through each audio frame in the second audio in steps of a single frequency point, each audio frame in the second audio may be smoothed using a smoothing filter (such as a Savitzky-Golay filter, a mean filter, or a median filter) to filter out noise in the second audio.
[0320] The above describes the sound quality detection method of high-resolution audio in the embodiment of the present application. The following describes the sound quality detection device of high-resolution audio in the embodiment of the present application. Figure 10 , Figure 10 This is a schematic diagram of an embodiment of a high-resolution audio quality detection device in an embodiment of the present application:
[0321] An acquisition unit 1001 is configured to acquire a sampling rate and a quantization accuracy of the audio to be detected;
[0322] A first screening unit 1002 is configured to screen out first audio whose sampling rate and quantization accuracy meet preset standards from the audio to be detected;
[0323] A determining unit 1003 is configured to determine an effective spectrum height of the first audio, where the effective spectrum height is a frequency corresponding to a location with maximum audio energy;
[0324] The second filtering unit 1004 is configured to filter out, from the first audio, a second audio whose effective spectrum height is greater than a frequency threshold;
[0325] The determining unit 1003 is further configured to determine that the second audio is high-resolution audio if the overall energy trend of the preset high-frequency band in the second audio shows a downward trend and no frequency point with a sudden energy change occurs in the preset high-frequency band.
[0326] Preferably, the acquisition unit is further configured to:
[0327] After selecting a first audio whose sampling rate and quantization accuracy both meet preset standards from the audio to be detected, and before obtaining the effective spectrum height of the first audio, obtaining the sampling rate of the first audio;
[0328] The device further comprises:
[0329] a classification unit 1005, configured to classify the first audio into a target audio category according to a plurality of preset sampling intervals, wherein each sampling interval corresponds to a different audio category;
[0330] The resampling unit 1006 is used to resample the first audio according to the resampling rate corresponding to the target category audio if the sampling rate of the first audio is not equal to the resampling rate corresponding to the target category audio, so that the resampling rate of the first audio is the same as the resampling rate corresponding to the target category audio, wherein different categories of audio correspond to different resampling rates.
[0331] Preferably, the determining unit 1003 is specifically configured to:
[0332] Frame-by-frame windowing of the first audio, and performing time domain to frequency domain conversion to obtain a frequency domain signal of each frame of audio in the first audio;
[0333] calculating the energy of each frame of audio in the first audio, so as to screen out valid frames in the first audio according to the energy of each frame of audio in the first audio and a first energy threshold;
[0334] Determine the effective spectral height of the effective frame in the first audio.
[0335] Preferably, the second screening unit 1004 is specifically used to:
[0336] Counting the number of first valid frames in the first audio whose effective spectrum height is greater than the frequency threshold;
[0337] If the ratio of the number of first valid frames in the first audio to the total number of valid frames in the first audio is greater than a first ratio threshold, the first audio is regarded as the second audio of the same target category.
[0338] Preferably, the determining unit 1003 is further configured to:
[0339] After filtering out valid frames in the first audio and before determining the effective spectrum height of the valid frames in the first audio, if the ratio of the number of valid frames in the first audio to the total number of audio frames in the first audio is less than a second ratio threshold, the first audio is determined to be non-high-resolution audio.
[0340] Preferably, the determining unit 1003 is specifically configured to:
[0341] Divide the valid frames in the first audio into M frequency bands according to a preset frequency interval, each frequency band including N frequency points, where M is greater than or equal to 2 and N is greater than or equal to 1;
[0342] Calculate, according to the frequency band calculation range corresponding to the target audio category, the first-order frequency band energy E1 of each frequency band and the sum E2 of the energies of two adjacent first-order frequency bands along the frequency axis in each valid frame of the first audio, where different audio categories correspond to different frequency band calculation ranges;
[0343] The spectrum height of each valid frame in the first audio is determined according to E1 or E2 corresponding to each frequency band in each valid frame in the first audio and a rule for determining the effective spectrum height.
[0344] Preferably, the effective spectrum height determination rule includes:
[0345] An inflection point frequency band where the energy of each frame begins to decrease is determined, and the frequency of the inflection point frequency band is determined as the effective spectrum height of each frame.
[0346] Preferably, the determining unit 1003 is specifically configured to:
[0347] Obtaining a first frequency band corresponding to a maximum E2 value in each valid frame of the first audio;
[0348] If the E1 value of the first frequency band is greater than the second energy threshold, or the E2 value of the second frequency band adjacent to the first frequency band along the frequency axis is greater than the third energy threshold, the first frequency band is regarded as the inflection point frequency band, and the frequency of the inflection point frequency band is determined as the effective spectrum height of the effective frame, wherein the second energy threshold and the third energy threshold are used to characterize the energy variation range of the inflection point frequency band;
[0349] or,
[0350] If the energy of the first frequency band is less than the invalid energy threshold, the first frequency band is regarded as the inflection point frequency band, and the frequency of the inflection point frequency band is determined as the effective spectrum height of the effective frame.
[0351] Preferably, the determining unit 1003 is specifically configured to:
[0352] Calculating, according to a plurality of preset center frequency bands and preset high frequency points in a preset high frequency band of the target category audio, a plurality of energy differences between the plurality of preset center frequency bands and the preset high frequency points for each frame of audio in the second audio, wherein the frequencies of the plurality of preset center frequency bands increase sequentially but are all less than the frequency of the preset high frequency points, and wherein different categories of audio correspond to different plurality of preset center frequency bands and different preset high frequency points;
[0353] counting a number of second valid frames in which the multiple energy difference values in the second audio are all greater than corresponding multiple thresholds, wherein the multiple thresholds respectively corresponding to the multiple energy difference values decrease as the frequencies of the multiple preset center frequency bands increase;
[0354] If the ratio of the second valid frame number to the total audio frame number in the second audio is not less than a third ratio threshold, it indicates that the overall energy trend of the preset high frequency band in the second audio presents a downward trend.
[0355] Preferably, the determining unit 1003 is specifically configured to:
[0356] According to a preset high frequency band of the target audio category, obtaining the preset high frequency band of each audio frame in the second audio, wherein different high frequency band ranges are set for different audio categories;
[0357] Calculating an energy distribution interval of a preset high frequency band of each audio frame in the second audio;
[0358] Counting the number of third valid frames without abnormal frequency points within the energy distribution interval, wherein the abnormal frequency point is a frequency point whose frequency energy is greater than a critical value of the energy distribution interval;
[0359] If the ratio of the third number of valid frames to the total number of audio frames in the second audio is not less than a fourth ratio threshold, it indicates that no frequency point with sudden energy changes occurs in the preset high frequency band of the second audio;
[0360] Determine that the second audio is high-resolution audio.
[0361] Preferably, the determining unit 1003 is specifically configured to:
[0362] Slide each audio frame of the second audio using a sliding window with a single frequency point as a step size, wherein the sliding window is at least larger than the frequency spectrum range of the single frequency point;
[0363] For each audio frame in the second audio, calculating the sum of first-order differences of energies of all frequency points in the sliding window;
[0364] If the sum of the first-order differences of the energies of all frequency points in the sliding window is negative and greater than a fourth energy threshold, determining that the audio frame in the second audio is a normal frame;
[0365] or,
[0366] If the sign of the sum of the first-order differences of the energies of all frequency points in the sliding window is positive and less than a fifth energy threshold, determining that the audio frame in the second audio is a normal frame;
[0367] Counting a proportion of normal frames in the second audio; if the proportion of normal frames in the second audio is not less than a fifth proportion threshold, indicating that no frequency point with a sudden energy change occurs in the preset high frequency band of the second audio;
[0368] Determine that the second audio is high-resolution audio.
[0369] Preferably, the determining unit 1003 is specifically configured to:
[0370] Slide each audio frame of the second audio using a sliding window with a single frequency point as a step size, wherein the sliding window is at least larger than the frequency spectrum range of the single frequency point;
[0371] For each audio frame in the second audio, calculating the sum of the first-order differences of the energy of each frequency point in the sliding window;
[0372] If the sign of the sum of the first-order differences of the energies of each frequency point in the sliding window is negative and greater than a sixth energy threshold, determining that the audio frame in the second audio is a normal frame;
[0373] or,
[0374] If the sum of the first-order differences of the energies of each frequency point in the sliding window is positive and less than a seventh energy threshold, determining that the audio frame in the second audio is a normal frame;
[0375] Counting a proportion of normal frames in the second audio; if the proportion of normal frames in the second audio is not less than a sixth proportion threshold, indicating that no frequency point with sudden energy change occurs in the preset high frequency band of the second audio;
[0376] Determine that the second audio is high-resolution audio.
[0377] Preferably, the device further comprises:
[0378] The pre-processing unit 1007 is configured to perform smoothing processing on each audio frame in the second audio using a smoothing filter before sliding the sliding window across each audio frame in the second audio with a single frequency point as a step size.
[0379] It should be noted that the functions of each unit in the embodiment of the present application are the same as those in the embodiment of the present application. Figures 1 to 9 The description in the embodiment is similar and will not be repeated here.
[0380] In an embodiment of the present application, the first screening unit 1002 is first used to perform a preliminary screening on the audio to be detected to screen out the first audio whose sampling rate and quantization accuracy meet the preset standards. Then, based on the perspective of effective energy, the second screening unit 1004 is used to screen out the second audio whose effective spectrum height is greater than the frequency threshold from the first audio. Finally, when the determination unit 1003 confirms that the energy of the preset high-frequency band of the second audio shows a downward trend and there is no frequency point with energy mutation in the preset high-frequency band, the second audio is confirmed to be high-resolution audio. That is, in this embodiment of the present application, on the one hand, the screening of high-resolution audio is realized based on specific quantization standards, thereby avoiding the participation of human subjective will and improving the accuracy of high-resolution audio recognition. On the other hand, the algorithm of the recognition standard can also be set on a computer device, thereby improving the efficiency of high-resolution audio recognition.
[0381] The above describes the high-resolution audio quality detection device in the embodiment of the present application from a modular perspective. Next, the computer device in the embodiment of the present invention is described from the perspective of hardware processing:
[0382] The computer device is used to implement the function of a high-resolution audio quality detection device. In one embodiment of the present invention, the computer device includes:
[0383] processor and memory;
[0384] The memory is used to store computer programs, and when the processor is used to execute the computer programs stored in the memory, the following steps can be implemented:
[0385] Get the sampling rate and quantization accuracy of the audio to be detected;
[0386] Filtering out a first audio whose sampling rate and quantization accuracy both meet preset standards from the audio to be detected;
[0387] Determine an effective spectrum height of the first audio, where the effective spectrum height is a frequency corresponding to a location of maximum audio energy;
[0388] Filtering out a second audio whose effective spectrum height is greater than a frequency threshold from the first audio;
[0389] If the overall energy trend of the preset high frequency band in the second audio presents a downward trend and no frequency point with a sudden energy change occurs in the preset high frequency band, the second audio is determined to be high-resolution audio.
[0390] In some embodiments of the present invention, after selecting the first audio whose sampling rate and quantization accuracy meet preset standards from the audio to be detected, and before obtaining the effective spectrum height of the first audio, the processor may further be configured to implement the following steps:
[0391] Obtaining a sampling rate of the first audio;
[0392] Classifying the first audio into a target audio category according to a plurality of preset sampling intervals, wherein each sampling interval corresponds to a different audio category;
[0393] If the sampling rate of the first audio is not equal to the resampling rate corresponding to the target category audio, the first audio is resampled according to the resampling rate corresponding to the target category audio, so that the resampling rate of the first audio is the same as the resampling rate corresponding to the target category audio, wherein different categories of audio correspond to different resampling rates.
[0394] In some embodiments of the present invention, the processor may further be configured to implement the following steps:
[0395] Frame-by-frame windowing of the first audio, and performing time domain to frequency domain conversion to obtain a frequency domain signal of each frame of audio in the first audio;
[0396] calculating the energy of each frame of audio in the first audio, so as to screen out valid frames in the first audio according to the energy of each frame of audio in the first audio and a first energy threshold;
[0397] Determine the effective spectral height of the effective frame in the first audio.
[0398] In some embodiments of the present invention, the processor may further be configured to implement the following steps:
[0399] Counting the number of first valid frames in the first audio whose effective spectrum height is greater than the frequency threshold;
[0400] If the ratio of the number of first valid frames in the first audio to the total number of valid frames in the first audio is greater than a first ratio threshold, the first audio is regarded as the second audio of the same target category.
[0401] In some embodiments of the present invention, after screening out valid frames in the first audio and before determining effective spectrum heights of the valid frames in the first audio, the processor may further be configured to implement the following steps:
[0402] If the ratio of the number of valid frames in the first audio to the total number of audio frames in the first audio is less than a second ratio threshold, it is determined that the first audio is non-high-resolution audio.
[0403] In some embodiments of the present invention, the processor may further be configured to implement the following steps:
[0404] Divide the valid frames in the first audio into M frequency bands according to a preset frequency interval, each frequency band including N frequency points, where M is greater than or equal to 2 and N is greater than or equal to 1;
[0405] Calculate, according to the frequency band calculation range corresponding to the target audio type, the first-order frequency band energy E1 of each frequency band and the sum E2 of the energies of two adjacent first-order frequency bands along the frequency axis in each valid frame of the first audio, where different frequency band calculation ranges correspond to different audio types;
[0406] The spectrum height of each valid frame in the first audio is determined according to E1 or E2 corresponding to each frequency band in each valid frame in the first audio and a rule for determining the effective spectrum height.
[0407] In some embodiments of the present invention, the effective spectrum height determination rule includes:
[0408] An inflection point frequency band where the energy of each frame begins to decrease is determined, and the frequency of the inflection point frequency band is determined as the effective spectrum height of each frame.
[0409] In some embodiments of the present invention, the processor may further be configured to implement the following steps:
[0410] An inflection point frequency band where the energy of the valid frame begins to decrease is determined, and the frequency of the inflection point frequency band is determined as the effective spectrum height of the valid frame.
[0411] In some embodiments of the present invention, the processor may further be configured to implement the following steps:
[0412] Obtaining a first frequency band corresponding to a maximum E2 value in each valid frame of the first audio;
[0413] If the E1 value of the first frequency band is greater than the second energy threshold, or the E2 value of the second frequency band adjacent to the first frequency band along the frequency axis is greater than the third energy threshold, the first frequency band is regarded as the inflection point frequency band, and the frequency of the inflection point frequency band is determined as the effective spectrum height of the effective frame, wherein the second energy threshold and the third energy threshold are used to characterize the energy variation range of the inflection point frequency band;
[0414] or,
[0415] If the energy of the first frequency band is less than the invalid energy threshold, the first frequency band is regarded as the inflection point frequency band, and the frequency of the inflection point frequency band is determined as the effective spectrum height of the effective frame.
[0416] In some embodiments of the present invention, the processor may further be configured to implement the following steps:
[0417] Calculating, according to a plurality of preset center frequency bands and preset high frequency points in a preset high frequency band of the target category audio, a plurality of energy differences between the plurality of preset center frequency bands and the preset high frequency points for each frame of audio in the second audio, wherein the frequencies of the plurality of preset center frequency bands increase sequentially but are all less than the frequency of the preset high frequency points, and wherein different categories of audio correspond to different plurality of preset center frequency bands and different preset high frequency points;
[0418] counting a number of second valid frames in which the multiple energy difference values in the second audio are all greater than corresponding multiple thresholds, wherein the multiple thresholds respectively corresponding to the multiple energy difference values decrease as the frequencies of the multiple preset center frequency bands increase;
[0419] If the ratio of the second valid frame number to the total audio frame number in the second audio is not less than a third ratio threshold, it indicates that the overall energy trend of the preset high frequency band in the second audio presents a downward trend.
[0420] In some embodiments of the present invention, the processor may further be configured to implement the following steps:
[0421] According to a preset high frequency band of the target audio category, obtaining the preset high frequency band of each audio frame in the second audio, wherein different high frequency band ranges are set for different audio categories;
[0422] Calculating an energy distribution interval of a preset high frequency band of each audio frame in the second audio;
[0423] Counting the number of third valid frames without abnormal frequency points within the energy distribution interval, wherein the abnormal frequency point is a frequency point whose frequency energy is greater than a critical value of the energy distribution interval;
[0424] If the ratio of the third number of valid frames to the total number of audio frames in the second audio is not less than a fourth ratio threshold, it indicates that no frequency point with sudden energy changes occurs in the preset high frequency band of the second audio;
[0425] Determine that the second audio is high-resolution audio.
[0426] In some embodiments of the present invention, the processor may further be configured to implement the following steps:
[0427] Slide each audio frame of the second audio using a sliding window with a single frequency point as a step size, wherein the sliding window is at least larger than the frequency spectrum range of the single frequency point;
[0428] For each audio frame in the second audio, calculating the sum of first-order differences of energies of all frequency points in the sliding window;
[0429] If the sum of the first-order differences of the energies of all frequency points in the sliding window is negative and greater than a fourth energy threshold, determining that the audio frame in the second audio is a normal frame;
[0430] or,
[0431] If the sign of the sum of the first-order differences of the energies of all frequency points in the sliding window is positive and less than a fifth energy threshold, determining that the audio frame in the second audio is a normal frame;
[0432] Counting a proportion of normal frames in the second audio; if the proportion of normal frames in the second audio is not less than a fifth proportion threshold, indicating that no frequency point with a sudden energy change occurs in the preset high frequency band of the second audio;
[0433] Determine that the second audio is high-resolution audio.
[0434] In some embodiments of the present invention, the processor may further be configured to implement the following steps:
[0435] Slide each audio frame of the second audio using a sliding window with a single frequency point as a step size, wherein the sliding window is at least larger than the frequency spectrum range of the single frequency point;
[0436] For each audio frame in the second audio, calculating the sum of the first-order differences of the energy of each frequency point in the sliding window;
[0437] If the sign of the sum of the first-order differences of the energies of each frequency point in the sliding window is negative and greater than a sixth energy threshold, determining that the audio frame in the second audio is a normal frame;
[0438] or,
[0439] If the sum of the first-order differences of the energies of each frequency point in the sliding window is positive and less than a seventh energy threshold, determining that the audio frame in the second audio is a normal frame;
[0440] Counting a proportion of normal frames in the second audio; if the proportion of normal frames in the second audio is not less than a sixth proportion threshold, indicating that no frequency point with sudden energy change occurs in the preset high frequency band of the second audio;
[0441] Determine that the second audio is high-resolution audio.
[0442] In some embodiments of the present invention, before sliding the sliding window across each audio frame of the second audio with a single frequency point as a step size, the processor may further be configured to implement the following steps:
[0443] A smoothing process is performed on each audio frame in the second audio using a smoothing filter.
[0444] It can be understood that when the processor in the computer device described above executes the computer program, it can also realize the functions of the various units in the corresponding device embodiments described above, which will not be repeated here. Exemplarily, the computer program can be divided into one or more modules / units, and the one or more modules / units are stored in the memory and executed by the processor to complete the present invention. The one or more modules / units can be a series of computer program instruction segments that can perform specific functions, and the instruction segments are used to describe the execution process of the computer program in the high-resolution audio sound quality detection device. For example, the computer program can be divided into the various units in the above-mentioned high-resolution audio sound quality detection device, and each unit can realize the specific functions described in the above-mentioned corresponding high-resolution audio sound quality detection device.
[0445] The computer device may be a computing device such as a desktop computer, laptop, PDA, or cloud server. The computer device may include, but is not limited to, a processor and memory. Those skilled in the art will appreciate that a processor and memory are merely examples of computer devices and do not constitute a limitation of the computer device. The computer device may include more or fewer components, or a combination of certain components, or different components. For example, the computer device may also include input and output devices, network access devices, buses, and the like.
[0446] The processor may be a central processing unit (CPU), other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor, etc. The processor is the control center of the computer device and connects various parts of the entire computer device using various interfaces and lines.
[0447] The memory can be used to store the computer programs and / or modules, and the processor implements various functions of the computer device by running or executing the computer programs and / or modules stored in the memory, and calling the data stored in the memory. The memory can mainly include a program storage area and a data storage area, wherein the program storage area can store an operating system, at least one application required for a function, etc.; the data storage area can store data created based on the use of the terminal, etc. In addition, the memory can include a high-speed random access memory, and can also include a non-volatile memory, such as a hard disk, a memory, a plug-in hard disk, a smart memory card (Smart Media Card, SMC), a secure digital (Secure Digital, SD) card, a flash card (Flash Card), at least one disk storage device, a flash memory device, or other volatile solid-state storage device.
[0448] The present invention further provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the processor may be configured to perform the following steps:
[0449] Get the sampling rate and quantization accuracy of the audio to be detected;
[0450] Filtering out a first audio whose sampling rate and quantization accuracy both meet preset standards from the audio to be detected;
[0451] Determine an effective spectrum height of the first audio, where the effective spectrum height is a frequency corresponding to a location of maximum audio energy;
[0452] Filtering out a second audio whose effective spectrum height is greater than a frequency threshold from the first audio;
[0453] If the overall energy trend of the preset high frequency band in the second audio presents a downward trend and no frequency point with a sudden energy change occurs in the preset high frequency band, the second audio is determined to be high-resolution audio.
[0454] In some embodiments of the present invention, after selecting the first audio whose sampling rate and quantization accuracy meet preset standards from the audio to be detected, and before obtaining the effective spectrum height of the first audio, the processor may further be configured to implement the following steps:
[0455] Obtaining a sampling rate of the first audio;
[0456] Classifying the first audio into a target audio category according to a plurality of preset sampling intervals, wherein each sampling interval corresponds to a different audio category;
[0457] If the sampling rate of the first audio is not equal to the resampling rate corresponding to the target category audio, the first audio is resampled according to the resampling rate corresponding to the target category audio, so that the resampling rate of the first audio is the same as the resampling rate corresponding to the target category audio, wherein different categories of audio correspond to different resampling rates.
[0458] In some embodiments of the present invention, when a computer program stored in a computer-readable storage medium is executed by a processor, the processor may further be configured to implement the following steps:
[0459] Frame-by-frame windowing of the first audio, and performing time domain to frequency domain conversion to obtain a frequency domain signal of each frame of audio in the first audio;
[0460] calculating the energy of each frame of audio in the first audio, so as to screen out valid frames in the first audio according to the energy of each frame of audio in the first audio and a first energy threshold;
[0461] Determine the effective spectral height of the effective frame in the first audio.
[0462] In some embodiments of the present invention, when a computer program stored in a computer-readable storage medium is executed by a processor, the processor may further be configured to implement the following steps:
[0463] Counting the number of first valid frames in the first audio whose effective spectrum height is greater than the frequency threshold;
[0464] If the ratio of the number of first valid frames in the first audio to the total number of valid frames in the first audio is greater than a first ratio threshold, the first audio is regarded as the second audio of the same target category.
[0465] In some embodiments of the present invention, after screening out valid frames in the first audio and before determining the effective spectrum heights of the valid frames in the first audio, when a computer program stored in a computer-readable storage medium is executed by a processor, the processor may further be configured to implement the following steps:
[0466] If the ratio of the number of valid frames in the first audio to the total number of audio frames in the first audio is less than a second ratio threshold, it is determined that the first audio is non-high-resolution audio.
[0467] In some embodiments of the present invention, when a computer program stored in a computer-readable storage medium is executed by a processor, the processor may further be configured to implement the following steps:
[0468] Divide the valid frames in the first audio into M frequency bands according to a preset frequency interval, each frequency band including N frequency points, where M is greater than or equal to 2 and N is greater than or equal to 1;
[0469] Calculate, according to the frequency band calculation range corresponding to the target audio type, the first-order frequency band energy E1 of each frequency band and the sum E2 of the energies of two adjacent first-order frequency bands along the frequency axis in each valid frame of the first audio, where different frequency band calculation ranges correspond to different audio types;
[0470] The spectrum height of each valid frame in the first audio is determined according to E1 or E2 corresponding to each frequency band in each valid frame in the first audio and a rule for determining the effective spectrum height.
[0471] In some embodiments of the present invention, when a computer program stored in a computer-readable storage medium is executed by a processor, the processor may further be configured to implement the following steps:
[0472] An inflection point frequency band where the energy of the valid frame begins to decrease is determined, and the frequency of the inflection point frequency band is determined as the effective spectrum height of the valid frame.
[0473] In some embodiments of the present invention, the effective spectrum height determination rule includes:
[0474] An inflection point frequency band where the energy of each frame begins to decrease is determined, and the frequency of the inflection point frequency band is determined as the effective spectrum height of each frame.
[0475] In some embodiments of the present invention, when a computer program stored in a computer-readable storage medium is executed by a processor, the processor may further be configured to implement the following steps:
[0476] Obtaining a first frequency band corresponding to a maximum E2 value in each valid frame of the first audio;
[0477] If the E1 value of the first frequency band is greater than the second energy threshold, or the E2 value of the second frequency band adjacent to the first frequency band along the frequency axis is greater than the third energy threshold, the first frequency band is regarded as the inflection point frequency band, and the frequency of the inflection point frequency band is determined as the effective spectrum height of the effective frame, wherein the second energy threshold and the third energy threshold are used to characterize the energy variation range of the inflection point frequency band;
[0478] or,
[0479] If the energy of the first frequency band is less than the invalid energy threshold, the first frequency band is regarded as the inflection point frequency band, and the frequency of the inflection point frequency band is determined as the effective spectrum height of the effective frame.
[0480] In some embodiments of the present invention, when a computer program stored in a computer-readable storage medium is executed by a processor, the processor may further be configured to implement the following steps:
[0481] Calculating, according to a plurality of preset center frequency bands and preset high frequency points in a preset high frequency band of the target category audio, a plurality of energy differences between the plurality of preset center frequency bands and the preset high frequency points for each frame of audio in the second audio, wherein the frequencies of the plurality of preset center frequency bands increase sequentially but are all less than the frequency of the preset high frequency points, and wherein different categories of audio correspond to different plurality of preset center frequency bands and different preset high frequency points;
[0482] counting a number of second valid frames in which the multiple energy difference values in the second audio are all greater than corresponding multiple thresholds, wherein the multiple thresholds respectively corresponding to the multiple energy difference values decrease as the frequencies of the multiple preset center frequency bands increase;
[0483] If the ratio of the second valid frame number to the total audio frame number in the second audio is not less than a third ratio threshold, it indicates that the overall energy trend of the preset high frequency band in the second audio presents a downward trend.
[0484] In some embodiments of the present invention, when a computer program stored in a computer-readable storage medium is executed by a processor, the processor may further be configured to implement the following steps:
[0485] According to a preset high frequency band of the target audio category, obtaining the preset high frequency band of each audio frame in the second audio, wherein different high frequency band ranges are set for different audio categories;
[0486] Calculating an energy distribution interval of a preset high frequency band of each audio frame in the second audio;
[0487] Counting the number of third valid frames without abnormal frequency points within the energy distribution interval, wherein the abnormal frequency point is a frequency point whose frequency energy is greater than a critical value of the energy distribution interval;
[0488] If the ratio of the third number of valid frames to the total number of audio frames in the second audio is not less than a fourth ratio threshold, it indicates that no frequency point with sudden energy changes occurs in the preset high frequency band of the second audio;
[0489] Determine that the second audio is high-resolution audio.
[0490] In some embodiments of the present invention, when a computer program stored in a computer-readable storage medium is executed by a processor, the processor may further be configured to implement the following steps:
[0491] Slide each audio frame of the second audio using a sliding window with a single frequency point as a step size, wherein the sliding window is at least larger than the frequency spectrum range of the single frequency point;
[0492] For each audio frame in the second audio, calculating the sum of first-order differences of energies of all frequency points in the sliding window;
[0493] If the sum of the first-order differences of the energies of all frequency points in the sliding window is negative and greater than a fourth energy threshold, determining that the audio frame in the second audio is a normal frame;
[0494] or,
[0495] If the sign of the sum of the first-order differences of the energies of all frequency points in the sliding window is positive and less than a fifth energy threshold, determining that the audio frame in the second audio is a normal frame;
[0496] Counting a proportion of normal frames in the second audio; if the proportion of normal frames in the second audio is not less than a fifth proportion threshold, indicating that no frequency point with a sudden energy change occurs in the preset high frequency band of the second audio;
[0497] Determine that the second audio is high-resolution audio.
[0498] In some embodiments of the present invention, when a computer program stored in a computer-readable storage medium is executed by a processor, the processor may further be configured to implement the following steps:
[0499] Slide each audio frame of the second audio using a sliding window with a single frequency point as a step size, wherein the sliding window is at least larger than the frequency spectrum range of the single frequency point;
[0500] For each audio frame in the second audio, calculating the sum of the first-order differences of the energy of each frequency point in the sliding window;
[0501] If the sign of the sum of the first-order differences of the energies of each frequency point in the sliding window is negative and greater than a sixth energy threshold, determining that the audio frame in the second audio is a normal frame;
[0502] or,
[0503] If the sum of the first-order differences of the energies of each frequency point in the sliding window is positive and less than a seventh energy threshold, determining that the audio frame in the second audio is a normal frame;
[0504] Counting a proportion of normal frames in the second audio; if the proportion of normal frames in the second audio is not less than a sixth proportion threshold, indicating that no frequency point with sudden energy change occurs in the preset high frequency band of the second audio;
[0505] Determine that the second audio is high-resolution audio.
[0506] In some embodiments of the present invention, before sliding the sliding window across each audio frame of the second audio in steps of a single frequency point, when a computer program stored in a computer-readable storage medium is executed by a processor, the processor may further be configured to implement the following steps:
[0507] A smoothing process is performed on each audio frame in the second audio using a smoothing filter.
[0508] It is understood that if the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a corresponding computer-readable storage medium. Based on this understanding, the present invention implements all or part of the processes in the above-mentioned corresponding embodiment methods, and can also be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, it can implement the steps of the above-mentioned various method embodiments. The computer program includes computer program code, which can be in source code form, object code form, executable file or some intermediate form. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal and software distribution medium. It should be noted that the content contained in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, computer-readable media do not include electric carrier signals and telecommunication signals.
[0509] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0510] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.
[0511] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0512] In addition, the functional units in the various embodiments of the present invention may be integrated into a single processing unit, each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0513] As described above, the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that the technical solutions described in the above embodiments can still be modified, or some of the technical features thereof can be replaced by equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for detecting the sound quality of high-resolution audio, characterized in that: include: Get the sampling rate and quantization accuracy of the audio to be detected; Selecting, from the audio to be detected, a first audio whose sampling rate and quantization precision both meet preset criteria, wherein the preset criteria include a sampling rate of at least 44100 Hz and a quantization precision of at least 16 bits; Determine an effective spectrum height of the first audio, where the effective spectrum height is a frequency corresponding to a location of maximum audio energy; Filtering out a second audio frequency having an effective spectrum height greater than a frequency threshold from the first audio frequency, wherein the frequency threshold is no greater than 20,000 Hz; If the overall energy trend of a preset high-frequency band in the second audio shows a downward trend, and no frequency point with a sudden energy change occurs in the preset high-frequency band, then the second audio is determined to be high-resolution audio, and the preset high-frequency band includes 20,000 Hz to 24,000 Hz, or includes 24,000 Hz to 48,000 Hz.
2. The sound quality detection method according to claim 1, wherein: After selecting a first audio whose sampling rate and quantization accuracy both meet preset standards from the audio to be detected, and before obtaining the effective spectrum height of the first audio, the method further includes: Obtaining a sampling rate of the first audio; Classifying the first audio into a target audio category according to a plurality of preset sampling intervals, wherein each sampling interval corresponds to a different audio category; If the sampling rate of the first audio is not equal to the resampling rate corresponding to the target category audio, the first audio is resampled according to the resampling rate corresponding to the target category audio, so that the resampling rate of the first audio is the same as the resampling rate corresponding to the target category audio, wherein different categories of audio correspond to different resampling rates.
3. The sound quality detection method according to claim 2, wherein: The determining of the effective spectrum height of the first audio comprises: Frame-by-frame windowing of the first audio, and performing time domain to frequency domain conversion to obtain a frequency domain signal of each frame of audio in the first audio; calculating the energy of each frame of audio in the first audio, so as to screen out valid frames in the first audio according to the energy of each frame of audio in the first audio and a first energy threshold; Determine the effective spectral height of the effective frame in the first audio.
4. The sound quality detection method according to claim 3, wherein: Screening out a second audio whose effective spectrum height is greater than a frequency threshold from the first audio includes: Counting the number of first valid frames in the first audio whose effective spectrum height is greater than the frequency threshold; If the ratio of the number of first valid frames in the first audio to the total number of valid frames in the first audio is greater than a first ratio threshold, the first audio is regarded as the second audio of the same target category.
5. The sound quality detection method according to claim 4, characterized in that: After screening out valid frames in the first audio and before determining effective spectrum heights of the valid frames in the first audio, the method further includes: If the ratio of the number of valid frames in the first audio to the total number of audio frames in the first audio is less than a second ratio threshold, it is determined that the first audio is non-high-resolution audio.
6. The sound quality detection method according to claim 3, wherein: Determining the effective spectrum height of the effective frame in the first audio includes: Divide the valid frames in the first audio into M frequency bands according to a preset frequency interval, each frequency band including N frequency points, where M is greater than or equal to 2 and N is greater than or equal to 1; Calculate, according to the frequency band calculation range corresponding to the target audio category, the first-order frequency band energy E1 of each frequency band and the sum E2 of the energies of two adjacent first-order frequency bands along the frequency axis in each valid frame of the first audio, where different audio categories correspond to different frequency band calculation ranges; The spectrum height of each valid frame in the first audio is determined according to E1 or E2 corresponding to each frequency band in each valid frame in the first audio and a rule for determining the effective spectrum height.
7. The sound quality detection method according to claim 6, characterized in that: The effective spectrum height determination rule includes: An inflection point frequency band where the energy of each frame begins to decrease is determined, and the frequency of the inflection point frequency band is determined as the effective spectrum height of each frame.
8. The sound quality detection method according to claim 7, characterized in that: The determining, based on E1 or E2 corresponding to each frequency band in each valid frame of the first audio and a rule for determining the effective spectrum height, of each valid frame of the first audio includes: Obtaining a first frequency band corresponding to a maximum E2 value in each valid frame of the first audio; If the E1 value of the first frequency band is greater than the second energy threshold, or the E2 value of the second frequency band adjacent to the first frequency band along the frequency axis is greater than the third energy threshold, the first frequency band is regarded as the inflection point frequency band, and the frequency of the inflection point frequency band is determined as the effective spectrum height of the effective frame, wherein the second energy threshold and the third energy threshold are used to characterize the energy range of the inflection point frequency band; or, If the energy of the first frequency band is less than the invalid energy threshold, the first frequency band is regarded as the inflection point frequency band, and the frequency of the inflection point frequency band is determined as the effective spectrum height of the valid frame.
9. The sound quality detection method according to claim 2, wherein: If the overall energy trend of the preset high frequency band in the second audio presents a downward trend and no frequency point with sudden energy changes occurs in the preset high frequency band, determining that the second audio is high-resolution audio includes: Calculating, according to a plurality of preset center frequency bands and preset high frequency points in a preset high frequency band of the target category audio, a plurality of energy differences between the plurality of preset center frequency bands and the preset high frequency points for each frame of audio in the second audio, wherein the frequencies of the plurality of preset center frequency bands increase sequentially but are all less than the frequency of the preset high frequency points, and wherein different categories of audio correspond to different plurality of preset center frequency bands and different preset high frequency points; counting a number of second valid frames in which the multiple energy difference values in the second audio are all greater than corresponding multiple thresholds, wherein the multiple thresholds respectively corresponding to the multiple energy difference values decrease as the frequencies of the multiple preset center frequency bands increase; If the ratio of the second valid frame number to the total audio frame number in the second audio is not less than a third ratio threshold, it indicates that the overall energy trend of the preset high frequency band in the second audio presents a downward trend.
10. The sound quality detection method according to claim 9, characterized in that: If the overall energy trend of the preset high frequency band in the second audio presents a downward trend and no frequency point with sudden energy change occurs in the preset high frequency band, then determining that the second audio is high-resolution audio further includes: According to a preset high frequency band of the target audio category, obtaining the preset high frequency band of each audio frame in the second audio, wherein different high frequency band ranges are set for different audio categories; Calculating an energy distribution interval of a preset high frequency band of each audio frame in the second audio; Counting the number of third valid frames without abnormal frequency points within the energy distribution interval, wherein the abnormal frequency point is a frequency point whose frequency energy is greater than a critical value of the energy distribution interval; If the ratio of the third number of valid frames to the total number of audio frames in the second audio is not less than a fourth ratio threshold, it indicates that no frequency point with sudden energy changes occurs in the preset high frequency band of the second audio; Determine that the second audio is high-resolution audio.
11. The sound quality detection method according to claim 9, wherein: If the overall energy trend of the preset high frequency band in the second audio presents a downward trend and no frequency point with sudden energy change occurs in the preset high frequency band, then determining that the second audio is high-resolution audio further includes: Slide each audio frame of the second audio using a sliding window with a single frequency point as a step size, wherein the sliding window is at least larger than the frequency spectrum range of the single frequency point; For each audio frame in the second audio, calculating the sum of first-order differences of energies of all frequency points in the sliding window; If the sum of the first-order differences of the energies of all frequency points in the sliding window is negative and greater than a fourth energy threshold, determining that the audio frame in the second audio is a normal frame; or, If the sign of the sum of the first-order differences of the energies of all frequency points in the sliding window is positive and less than a fifth energy threshold, determining that the audio frame in the second audio is a normal frame; Counting a proportion of normal frames in the second audio; if the proportion of normal frames in the second audio is not less than a fifth proportion threshold, indicating that no frequency point with a sudden energy change occurs in the preset high frequency band of the second audio; Determine that the second audio is high-resolution audio.
12. The sound quality detection method according to claim 9, wherein: If the overall energy trend of the preset high frequency band in the second audio presents a downward trend and no frequency point with sudden energy change occurs in the preset high frequency band, then determining that the second audio is high-resolution audio further includes: Slide each audio frame of the second audio using a sliding window with a single frequency point as a step size, wherein the sliding window is at least larger than the frequency spectrum range of the single frequency point; For each audio frame in the second audio, calculating the sum of the first-order differences of the energy of each frequency point in the sliding window; If the sign of the sum of the first-order differences of the energies of each frequency point in the sliding window is negative and greater than a sixth energy threshold, determining that the audio frame in the second audio is a normal frame; or, If the sum of the first-order differences of the energies of each frequency point in the sliding window is positive and less than a seventh energy threshold, determining that the audio frame in the second audio is a normal frame; Counting a proportion of normal frames in the second audio; if the proportion of normal frames in the second audio is not less than a sixth proportion threshold, indicating that no frequency point with sudden energy change occurs in the preset high frequency band of the second audio; Determine that the second audio is high-resolution audio.
13. The sound quality detection method according to claim 11 or 12, characterized in that: Before sliding through each audio frame in the second audio using the sliding window with a single frequency point as a step size, the method further includes: A smoothing process is performed on each audio frame in the second audio using a smoothing filter.
14. A computer device comprising a processor, characterized in that: When executing the computer program stored in the memory, the processor is configured to implement the high-resolution audio sound quality detection method according to any one of claims 1 to 13.
15. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, it is used to implement the sound quality detection method for high-resolution audio according to any one of claims 1 to 13.
Citation Information
Patent Citations
True quality judging method and system for music
CN104103279A
Detection of clipping event in audio signals
US20170052758A1