Audio quality detection method and device and electronic equipment
After receiving tasks from the server, the terminal device plays and records audio, performs time-domain alignment and spectral difference analysis, which solves the accuracy and coverage problems of audio quality detection in the existing technology and realizes automated detection of audio quality across the entire link.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- HANGZHOU NETEASE CLOUD MUSIC TECH CO LTD
- Filing Date
- 2026-01-30
- Publication Date
- 2026-05-01
AI Technical Summary
In existing technologies, manual subjective listening detection solutions are costly and have a high false positive rate, and cannot cover all audio problems, while server-side static audio detection solutions cannot detect audio quality problems at the audio playback end.
The terminal device receives the detection task from the server, plays and records the audio to be detected, aligns the recorded audio and the audio to be detected in the time domain, performs spectral difference analysis, generates the detection result, and returns it to the server.
It enables automated detection of the entire audio playback process, improving the accuracy and reliability of audio quality detection.
Smart Images

Figure CN121963780A_ABST
Abstract
Description
Audio quality testing methods, devices and electronic equipment Technical Field
[0001] This disclosure relates to the field of quality testing technology, and in particular to an audio quality testing method, apparatus and electronic device. Background Technology
[0002] Audio quality detection technologies in related fields typically include manual subjective listening detection schemes and server-side static audio detection schemes. However, manual subjective listening detection schemes require significant manpower, and it is difficult for humans to identify all audio problems, resulting in a high false positive rate and narrow coverage. Server-side static audio detection schemes only perform audio quality detection on audio files and cannot cover audio quality detection at the audio playback end. Summary of the Invention
[0003] In view of this, the purpose of this disclosure is to provide an audio quality detection method, apparatus and electronic device for detecting the audio quality of an audio playback device.
[0004] In a first aspect, embodiments of this disclosure provide an audio quality detection method. The method is applied to a terminal device, which is communicatively connected to a server. The method includes: receiving a detection task sent by the server, the detection task including an audio to be detected; playing the audio to be detected and recording the played audio to be detected to obtain a recorded audio; performing time-domain alignment on the recorded audio and the audio to be detected to obtain a time-domain aligned audio, and performing spectral difference analysis on the time-domain aligned audio to obtain a detection result; and returning the detection result to the server.
[0005] Secondly, this disclosure also provides an audio quality detection device, which is installed on a terminal device and communicatively connected to a server. The device includes: a task receiving module for receiving a detection task sent by the server, wherein the detection task includes an audio to be detected; an audio recording module for playing the audio to be detected and recording the played audio to be detected to obtain a recorded audio; an audio analysis module for performing time-domain alignment on the recorded audio and the audio to be detected to obtain time-domain aligned audio, and performing spectral difference analysis on the time-domain aligned audio to obtain a detection result; and a result return module for returning the detection result to the server.
[0006] Thirdly, this disclosure provides an electronic device including a processor and a memory, the memory storing machine-executable instructions that can be executed by the processor, the processor executing the machine-executable instructions to implement the above-described audio quality detection method.
[0007] Fourthly, this disclosure provides a computer-readable storage medium storing computer-executable instructions that, when invoked and executed by a processor, cause the processor to implement the aforementioned audio quality detection method.
[0008] The embodiments disclosed herein bring the following beneficial effects: An audio quality detection method, apparatus, and electronic device provided by this disclosure firstly receive a detection task sent by a server, the detection task including the audio to be detected; then, the audio to be detected is played and the played audio is recorded to obtain recorded audio; next, the recorded audio and the audio to be detected are time-domain aligned to obtain time-domain aligned audio, and spectral difference analysis is performed on the time-domain aligned audio to obtain a detection result; then, the detection result is returned to the server. In this method, when the audio to be detected indicated by the detection task is played through the terminal device, the played audio is recorded, and the audio to be detected is used as a reference audio. By performing spectral difference analysis and frame-by-frame comparison on the reference audio and the recorded audio, automated detection of audio quality problems in the entire audio playback chain is achieved, improving the accuracy and reliability of audio quality detection.
[0009] Other features and advantages of this disclosure will be set forth in the following description and will be apparent in part from the description or may be learned by practicing the disclosure. The objects and other advantages of this disclosure are realized and obtained through the structures particularly pointed out in the description, claims and drawings.
[0010] To make the above-mentioned objects, features and advantages of this disclosure more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. Attached Figure Description
[0011] To more clearly illustrate the technical solutions in the specific embodiments of this disclosure or the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this disclosure. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0012] Figure 1 is a flowchart of an audio quality detection method provided in an embodiment of this disclosure; Figure 2 is an architecture diagram of an audio quality detection system provided in an embodiment of this disclosure; Figure 3 is a schematic diagram of an audio recording provided in an embodiment of this disclosure; Figure 4 is a flowchart of an audio quality detection method provided in an embodiment of this disclosure; Figure 5 is a structural schematic diagram of an audio quality detection device provided in an embodiment of this disclosure; Figure 6 is a structural schematic diagram of an electronic device provided in an embodiment of this disclosure. Detailed Implementation
[0013] To make the objectives, technical solutions, and advantages of the embodiments of this disclosure clearer, the technical solutions of the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this disclosure, and not all embodiments. The components of the embodiments of this disclosure described and shown in the accompanying drawings can generally be arranged and designed in various different configurations.
[0014] Therefore, the following detailed description of the embodiments of this disclosure provided in the accompanying drawings is not intended to limit the scope of the claimed disclosure, but merely to illustrate selected embodiments of the disclosure. All other embodiments obtained by those skilled in the art based on the embodiments of this disclosure without inventive effort are within the scope of protection of this disclosure.
[0015] The relevant technologies offer two audio quality testing solutions: Solution 1, a manual subjective listening test. This solution relies on experienced professional listeners in an anechoic chamber that meets specified standards, using professional monitoring equipment to play the audio to be tested (covering different sound quality modes and sound effect combinations). The listener judges whether there are problems such as noise, stuttering, distortion, or phase shift through auditory perception. The testing process involves: establishing a "sound quality type - sound effect - device" test matrix, manually playing and testing each item, recording the problems and related parameters, and generating a test report. However, this method cannot exhaustively cover all combinations of "sound quality mode × sound effect × mainstream devices." Previously identified detailed sound quality issues are often missed due to the inability to effectively cover them. Furthermore, it has low sensitivity to hidden problems such as high-frequency noise and minor stuttering, with an average discovery cycle of 15 days, requiring user complaints to trigger investigations. Additionally, full-chain verification of new devices requires 3-5 working days, and differences in hearing thresholds among different individuals lead to a high misjudgment rate; for example, novices may easily overlook weak noise.
[0016] Option 2, the server-side static audio detection solution, focuses on the massive amount of audio materials stored on the server (including lossless, high-definition, and standard-definition audio at various quality levels, covering all genres such as pop, classical, and rock). It addresses quality issues such as noise, distortion, and damage that may occur during the production, transcoding, storage, and transmission of audio files. By extracting the audio's own characteristic parameters and using preset detection logic, it achieves rapid screening and judgment of audio material quality, ensuring that audio files uploaded to the server meet the platform's sound quality standards and avoiding user listening experience problems caused by inherent audio quality defects. However, this method only detects the audio file itself and cannot cover dynamic processes such as edge decoding, NPU (Neural Processing Unit) super-resolution, sound effect overlay, and limiter processing. For example, lossless audio that is acceptable on the server may introduce distortion after edge processing. Furthermore, music often contains many electronic elements that enhance rhythm and atmosphere; these sounds are easily detected as noise or abnormalities in a no-reference mode.
[0017] To address the aforementioned issues, this disclosure provides an audio quality detection method, apparatus, and electronic device. This technology can be applied to audio quality detection scenarios in any audio end-side playback link.
[0018] To facilitate understanding of the embodiments of this disclosure, an audio quality detection method is first described in detail. This method is applied to a terminal device that is connected to a server. As shown in Figure 1, the method includes the following specific steps: Step S102, receiving a detection task sent by the server, wherein the detection task includes the audio to be detected.
[0019] In practical implementation, the aforementioned terminal device is an electronic device used for playing and detecting audio quality. This electronic device can be a mobile phone, computer, or other similar device. When a user needs to perform quality detection on the audio to be detected, the server sends a detection task for the audio to be detected to the terminal device.
[0020] Step S104: Play the audio to be tested and record the played audio to be tested to obtain the recorded audio.
[0021] After receiving the detection task, the terminal device needs to obtain the audio to be detected as indicated by the detection task, and play the audio to be detected through the player set on the terminal device. During the playback of the audio to be detected, the recording device set on the terminal device records the played audio to be detected in real time, thus obtaining the recorded audio corresponding to the audio to be detected.
[0022] Step S106: Perform time-domain alignment on the recorded audio and the audio to be detected to obtain the time-domain aligned audio, and perform spectral difference analysis on the time-domain aligned audio to obtain the detection result.
[0023] After obtaining the recorded audio corresponding to the audio to be detected, the recorded audio and the audio to be detected are aligned in the time domain to obtain the time-aligned audio; wherein, the recorded audio and the audio to be detected in the time-aligned audio have a one-to-one alignment mapping relationship in each audio frame.
[0024] After obtaining the time-domain aligned audio, spectral difference analysis, time-domain content consistency checks, and volume fluctuation checks can be performed. The results, derived from this multi-dimensional analysis, are then output as detection results. These results include whether any audio quality issues exist for the audio being tested, what those issues are, and where the problematic audio is located. These audio quality issues include, but are not limited to: silence, stuttering, dropped frames, overlap, redundancy, static, and fluctuating volume.
[0025] Step S108: Return the detection results to the server.
[0026] After the terminal device generates the detection result, it sends the detection result to the server so that the server can adjust the audio to be detected based on the detection result.
[0027] This disclosure provides an audio quality detection method that, when playing the audio to be detected indicated by a detection task through a terminal device, records the played audio and uses the audio to be detected as a reference audio. By performing spectral difference analysis and frame-by-frame comparison on the reference audio and the recorded audio, the method achieves automated detection of audio quality problems throughout the entire audio playback chain, thereby improving the accuracy and reliability of audio quality detection.
[0028] The following examples describe the method for obtaining detection results.
[0029] Specifically, the process of performing time-domain alignment on the recorded audio and the audio to be detected to obtain the time-domain aligned audio, and then performing spectral difference analysis on the time-domain aligned audio to obtain the detection result can be achieved through the following steps 10-11.
[0030] Step 10: Perform time-domain alignment on the recorded audio and the audio to be detected to obtain the time-domain aligned audio.
[0031] When performing temporal alignment between the recorded audio and the audio to be detected, features are first extracted from both audio and recorded audio to obtain the audio features of the audio to be detected and the audio features of the recorded audio. Then, the feature similarity between the audio features of the audio to be detected and the audio features of the recorded audio is determined. Based on the feature similarity, the alignment position of the audio to be detected and the recorded audio in the temporal domain is determined. Based on the alignment position, the recorded audio and the audio to be detected are temporally aligned to obtain the temporally aligned audio.
[0032] The aforementioned audio features may include Mel-frequency cepstral coefficients and / or time-domain waveforms. Mel-frequency cepstral coefficients are an audio feature extraction algorithm that simulates human hearing characteristics; this disclosure uses them to generate unique audio fingerprints to achieve precise alignment between the audio to be detected and the recorded audio. The time-domain waveform of the audio is a visual representation of the sound signal in the time dimension, describing how the amplitude of the audio signal (such as changes in sound pressure level) changes over time.
[0033] In practical implementation, after obtaining the audio features, starting from the first audio frame of the audio to be detected and the first audio frame of the recorded audio, the feature similarity between the audio features of the audio to be detected and the audio features of the recorded audio can be analyzed frame by frame. When the feature similarity meets the preset conditions, the audio frame corresponding to the feature similarity that meets the preset conditions is determined as the alignment position, and the audio to be detected and the recorded audio are temporally aligned starting from the alignment position to obtain the temporally aligned audio.
[0034] In an optional embodiment, the specific process of determining the feature similarity between the audio features of the audio to be detected and the audio features of the recorded audio, and determining the alignment position of the audio to be detected and the recorded audio in the time domain based on the feature similarity, may include: sliding a first sliding window of a first preset duration across the audio to be detected and the recorded audio with a preset step size, and determining the feature similarity between the audio features of the audio to be detected and the audio features of the recorded audio within the first sliding window during the sliding process; if the feature similarity is greater than a preset similarity threshold, determining the position of the first sliding window within the recorded audio as the alignment position of the audio to be detected and the recorded audio in the time domain.
[0035] In practice, the specific duration corresponding to the first preset duration can be determined according to the development needs. For example, the first preset duration can be set to 60 seconds or 40 seconds, the sliding step size of the first sliding window can be set to 20ms or 30ms, etc., and the preset similarity threshold can be determined according to the user's settings, for example, set to 500 or 400, etc.
[0036] Specifically, starting from the first audio frame in the audio to be detected and the recorded audio, a first sliding window of a first preset duration is slid across the audio to be detected and the recorded audio with a preset step size. During the sliding of the first sliding window, the feature similarity between the audio features of the audio to be detected and the audio features of the recorded audio within the first sliding window is determined in real time. If the feature similarity is greater than a preset similarity threshold, an alignment position is found, and the sliding of the first sliding window is stopped. The position of the first sliding window within the recorded audio is determined as the alignment position of the audio to be detected and the recorded audio in the time domain. If the feature similarity is not greater than the preset similarity threshold, the sliding of the first sliding window continues until an alignment position is found or the first sliding window reaches the position of the last audio frame of the recorded audio.
[0037] In one optional embodiment, before feature extraction is performed on the audio to be detected and the recorded audio respectively, it is also necessary to standardize the audio to be detected so that the audio to be detected and the recorded audio have the same format, sampling rate, bit depth and number of channels.
[0038] In practical implementation, the temporal alignment of the audio to be detected and the recorded audio can be achieved using the DTW (Dynamic Time Warping) algorithm, which generates an alignment mapping relationship between the two audio files. If the audio to be detected and the recorded audio cannot achieve temporal alignment, it indicates that the recorded audio has audio quality issues such as silence or stuttering.
[0039] Step 11: Perform temporal content consistency detection and volume fluctuation detection on the temporally aligned audio to obtain the detection results.
[0040] After aligning the audio to be detected and the recorded audio in the temporal domain, the differences between the two audio can be calculated frame by frame. This means performing temporal content consistency detection and volume fluctuation detection on the temporally aligned audio. Specifically, audio temporal content consistency is used to detect audio quality issues such as dropped frames, stuttering, overlap, and redundancy; volume fluctuation detection is used to detect audio quality issues such as fluctuating audio volume.
[0041] The types of audio quality problems that can be detected according to the embodiments of this disclosure include a variety, as shown in Table 1.
[0042] Table 1
[0043] In practical implementation, the above audio content difference comparison refers to the above temporal consistency detection, and the above volume consistency analysis refers to the above volume fluctuation detection.
[0044] In an optional embodiment, the specific process of performing temporal content consistency detection and volume fluctuation detection on the temporally aligned audio to obtain the detection results can be achieved through the following steps 20-22: Step 20, calculate the content difference between the recorded audio and the audio to be detected in the temporally aligned audio frame by frame to obtain the content difference detection result.
[0045] In practical implementation, the first step is to locate the starting point of the temporally aligned audio. Then, a segment of recorded audio of the same length as the audio to be detected is extracted from the temporally aligned audio to achieve precise audio alignment. Next, the content differences between the recorded audio and the audio to be detected are calculated frame by frame in the extracted temporally aligned audio. Audio segments with content differences are recorded, generating a content difference detection result. This result includes the audio segments with content differences and the corresponding audio quality issues.
[0046] In an optional embodiment, the specific process of calculating the content difference between the recorded audio and the audio to be detected in the temporally aligned audio frame by frame to obtain the content difference detection result may include: calculating the bit error rate between the recorded audio and the audio to be detected in the temporally aligned audio frame by frame to obtain the bit error rate corresponding to each frame of audio; and generating the content difference detection result based on the bit error rate corresponding to each frame of audio.
[0047] The Bit Error Rate (BER), also known as the bit error rate or bit error rate, is a core indicator for measuring the reliability of digital communication systems. It is defined as the ratio of the number of bits that err during transmission to the total number of bits transmitted.
[0048] In practical implementation, a bit error rate greater than a preset ratio threshold indicates that there is a content difference in the audio of the corresponding audio frame; the preset ratio threshold can be determined according to the research and development needs, for example, the preset ratio threshold can be set to 0.15 or 0.3, etc.
[0049] In an optional embodiment, the specific process of generating content difference detection results based on the bit error rate corresponding to each audio frame may include: determining the target audio frame in the temporally aligned audio whose bit error rate is greater than a preset ratio threshold, and determining the duration of the target audio frame; generating content difference detection results based on the target audio and the duration of the target audio.
[0050] In practice, the audio segments with audio quality problems can be obtained based on the duration of the target audio, and the audio quality problems can be determined based on the audio segments. Then, the audio segments and the corresponding audio quality problems are recorded in the content difference detection results.
[0051] In practice, stuttering occurs when a portion of the audio is lost within a certain time window. This can be detected by calculating the content difference between the audio to be detected and the recorded audio. The time window here can be the time window corresponding to the duration mentioned above. Some special audio processing techniques are very similar to clipping, but by comparing the content differences between the audio to be detected and the pre-recorded audio, it can be determined whether the problem is caused by special processing or playback issues.
[0052] Step 21: If the content difference detection result indicates that there is no content difference, calculate the volume difference between the recorded audio and the audio to be detected in the time-domain aligned audio by using a sliding window method to obtain the volume difference detection result.
[0053] In practical implementation, volume fluctuation detection needs to be performed only after ensuring that there are no differences in the audio content of the time-domain aligned audio, so as to eliminate the interference of content differences on volume detection. In practical applications, a second sliding window of a second preset duration can be used to slide over the time-domain aligned audio with a preset step size. During the sliding process, the volume difference value of the time-domain aligned audio within the second sliding window is determined. If the volume difference value is less than a preset difference threshold, it is determined that there is a volume difference in the time-domain aligned audio within the second sliding window. Based on the volume difference value of the time-domain aligned audio, the volume difference detection result is determined.
[0054] The specific duration corresponding to the aforementioned second preset duration can be determined according to R&D needs. For example, the second preset duration can be set to 60 seconds or 40 seconds, the sliding step size of the second sliding window can be set to 20ms or 30ms, etc., and the aforementioned preset difference threshold can be determined according to user settings, for example, set to 16dB or 18dB, etc. The second preset duration corresponding to the second sliding window should not be set too long or too short. If the second preset duration is set too short, it is easily affected by instantaneous fluctuations; if the second preset duration is set too long, it will miss the acquisition of rapid volume changes.
[0055] Specifically, starting from the first audio frame of the time-aligned audio, a second sliding window of a second preset duration slides across the time-aligned audio at preset step sizes. During this sliding, the volume difference between the audio to be detected and the recorded audio within the second sliding window is determined in real time. If the volume difference is less than a preset difference threshold, a volume difference is confirmed between the audio to be detected and the recorded audio within the second sliding window. The first sliding window continues to slide until the second sliding window reaches the last audio frame of the time-aligned audio. Then, a volume difference detection result is generated based on the volume difference value. This result includes the audio segment with the volume difference and the audio quality issues present in that segment. These audio quality issues may include, but are not limited to: fluctuating volume (abnormal dynamic range), excessive compression by the limiter leading to volume attenuation, and gain imbalance introduced by certain audio effect plugins.
[0056] The specific process for determining the volume difference value of the time-domain aligned audio within the second sliding window may include: calculating the RMS energy corresponding to the audio to be detected and the recorded audio within the second sliding window after time-domain alignment, then calculating the difference between the RMS energy corresponding to the audio to be detected and the recorded audio, and determining the difference value as the volume difference value.
[0057] Step 22: Based on the content difference detection results and volume difference detection results, the final detection results are obtained.
[0058] By integrating the content difference detection results and the volume difference detection results, the final detection result can be obtained. This result includes audio quality problem description information and audio segments with differences. Among them, the audio segments with differences include: the audio position with differences in the audio to be detected, the audio position with differences in the recorded audio, and the duration of the audio segment with differences (i.e., the target audio mentioned above).
[0059] In an optional embodiment, the above detection task further includes detection conditions, which can be determined according to user needs. That is, the user can dynamically adjust the detection conditions of the terminal device according to their needs. These detection conditions are used to indicate the data acquisition strategy of the terminal device. Based on the detection conditions, multiple data acquisition switches and storage paths on the terminal device can be precisely controlled, avoiding storage pressure caused by full data acquisition. Each of the multiple data acquisition switches corresponds to a different data acquisition point.
[0060] Table 2 shows the relevant descriptions of the seven data collection points provided in this disclosure.
[0061] Table 2
[0062] The term "downmix" refers to a type of audio processing technology. Its core function is to convert multi-channel audio signals (such as 5.1 or 7.1 channels) into a format with fewer channels (such as two-channel stereo or mono) to suit the hardware limitations or specific needs of playback devices.
[0063] Based on the above description, when the terminal device receives a detection task, it can determine the audio acquisition point according to the detection conditions indicated by the detection task; when playing the audio to be detected, the audio to be detected is acquired at the audio acquisition point to obtain the acquired audio. The audio acquisition point can be any one or more of the aforementioned multiple acquisition data points.
[0064] After obtaining the collected audio, the final detection results can be obtained based on the content difference detection results and the volume difference detection results. Then, audio analysis can be performed on the collected audio based on the content difference detection results and the volume difference detection results to identify audio quality problems and target audio points in the collected audio. Based on the audio quality problems and target audio points, the detection results are generated.
[0065] This disclosure establishes key data acquisition points across the entire audio processing chain, covering the complete signal path from decoding to hardware output. Data is captured at each acquisition point according to a unified timing specification, ensuring accurate comparison and problem localization in subsequent analysis stages.
[0066] The following examples describe another method for audio quality detection.
[0067] Specifically, the above detection task also includes detection location and audio set to be detected; based on this, the specific process of playing the audio to be detected may include: determining the audio to be detected corresponding to the detection location from the audio set to be detected, and playing the audio to be detected.
[0068] In practice, the server sets up a set of audio files to be tested, which includes a large number of audio files. When the terminal device receives the test task from the server, it needs to retrieve the audio files to be tested from the set of audio files according to the test position indicated by the test task, and play the audio files to be tested in order to perform audio quality testing on the audio files to be tested.
[0069] Figure 2 shows the architecture of an audio quality detection system provided in this embodiment. The server sends a detection task to the terminal device via an HTTP API. The detection task includes a standardized music test set (i.e., the audio set to be detected) and test cases. The test cases include detection location, detection conditions, and detection expectations. The terminal device receives and executes the detection task, collects audio data, analyzes and processes the data, and uploads the detection results to the server. The server receives and stores the detection results.
[0070] When the terminal device receives a detection task from the server, it needs to acquire the audio to be detected from the audio set to be detected according to the detection location indicated by the detection task. It then plays the audio to be detected through playback content, thereby acquiring the audio corresponding to the audio acquisition point indicated by the detection task via the playback kernel. Finally, it records the played audio to be detected through the sound card to obtain the recorded audio. The acquired and recorded audio are input into the audio analysis module. The audio analysis module uses the audio to be detected as a reference file, aligns the reference file and the recorded audio to obtain time-domain aligned audio, and then performs spectral difference analysis on the time-domain aligned audio to obtain the detection result. When the detection result indicates that the audio to be detected has an audio quality problem, the detection result is sent to the server. Specifically, when the audio to be detected has an audio quality problem, the difference between the acquired audio before and after the audio acquisition point can be used to accurately locate the audio quality problem.
[0071] Figure 3 shows a schematic diagram of an audio recording method provided in this embodiment. The audio to be tested can be played through surround speakers 1, surround speakers 2, and a subwoofer. The played audio is recorded using a sound card to obtain the recorded audio. Automated acquisition is achieved based on SoundDevice, resulting in a real audio recording played by the terminal device. Figure 4 shows a flowchart of an audio quality detection method provided in this embodiment. In Figure 4, the music app is set on the terminal device. The server sends JSON test cases (i.e., the aforementioned detection task) to the terminal device. The GCDWebServer on the terminal device parses and executes the JSON test cases, configures the playback parameters of the music app, plays the audio to be tested, and records the played audio through the sound card to obtain the recorded audio. Temporal alignment and spectral difference analysis are performed on the recorded audio and the audio to be tested to obtain the detection result, which is then reported to the server.
[0072] Corresponding to the above method embodiment, this disclosure embodiment also provides an audio quality detection device, which is installed on a terminal device and is communicatively connected to a server. As shown in FIG5, the device includes: a task receiving module 50, used to receive a detection task sent by the server, wherein the detection task includes the audio to be detected.
[0073] The audio recording module 51 is used to play the audio to be detected and record the played audio to be detected, thus obtaining the recorded audio.
[0074] The audio analysis module 52 is used to perform time-domain alignment on the recorded audio and the audio to be detected to obtain the time-domain aligned audio, and to perform spectral difference analysis on the time-domain aligned audio to obtain the detection result.
[0075] The result return module 53 is used to return the detection results to the server.
[0076] The aforementioned audio quality detection device records the audio being played when the detection task instruction is played through a terminal device, and uses the audio to be detected as a reference audio. By performing spectral difference analysis and frame-by-frame comparison on the reference audio and the recorded audio, it achieves automated detection of audio quality problems throughout the entire audio playback chain, thereby improving the accuracy and reliability of audio quality detection.
[0077] Furthermore, the aforementioned audio analysis module 52 is used to: perform temporal alignment on the recorded audio and the audio to be detected to obtain temporally aligned audio; and perform temporal content consistency detection and volume fluctuation detection on the temporally aligned audio to obtain detection results.
[0078] Furthermore, the aforementioned audio analysis module 52 is also used to: extract features from the audio to be detected and the recorded audio respectively to obtain the audio features of the audio to be detected and the audio features of the recorded audio; determine the feature similarity between the audio features of the audio to be detected and the audio features of the recorded audio; determine the alignment position of the audio to be detected and the recorded audio in the time domain based on the feature similarity; and perform time domain alignment on the recorded audio and the audio to be detected based on the alignment position to obtain the time-domain aligned audio.
[0079] Furthermore, the aforementioned audio features include: Mel frequency cepstral coefficients and / or time-domain waveforms.
[0080] Furthermore, the aforementioned audio analysis module 52 is also used to: slide a first sliding window of a first preset duration over the audio to be detected and the recorded audio with a preset step size, and determine the feature similarity between the audio features of the audio to be detected and the audio features of the recorded audio within the first sliding window during the sliding process; if the feature similarity is greater than a preset similarity threshold, determine the position of the first sliding window within the recorded audio as the alignment position of the audio to be detected and the recorded audio in the time domain.
[0081] Furthermore, the aforementioned audio analysis module 52 is also used to: calculate the content difference between the recorded audio and the audio to be detected in the temporally aligned audio frame by frame to obtain a content difference detection result; if the content difference detection result indicates that there is no content difference, calculate the volume difference between the recorded audio and the audio to be detected in the temporally aligned audio by using a sliding window to obtain a volume difference detection result; and obtain the final detection result based on the content difference detection result and the volume difference detection result.
[0082] Furthermore, the audio analysis module 52 is also used to: calculate the bit error rate between the recorded audio and the audio to be detected in the temporally aligned audio frame by frame, to obtain the bit error rate corresponding to each frame of audio; and generate content difference detection results based on the bit error rate corresponding to each frame of audio.
[0083] Furthermore, the bit error rate exceeding the preset ratio threshold indicates that there is a content difference in the audio of the corresponding audio frame; based on this, the audio analysis module 52 is also used to: determine the target audio frame in the temporally aligned audio whose bit error rate exceeds the preset ratio threshold, and determine the duration of the target audio frame; and generate a content difference detection result based on the target audio and the duration of the target audio.
[0084] Furthermore, the aforementioned audio analysis module 52 is also used to: slide a second sliding window of a second preset duration over the time-domain aligned audio with a preset step size, and determine the volume difference value of the time-domain aligned audio within the second sliding window during the sliding process; if the volume difference value is less than a preset difference threshold, determine that there is a volume difference in the time-domain aligned audio within the second sliding window; and determine the volume difference detection result based on the volume difference value of the time-domain aligned audio.
[0085] Furthermore, the aforementioned detection task also includes detection conditions; based on this, the aforementioned device further includes a point acquisition module, used to: determine audio acquisition points according to the detection conditions; when playing the audio to be detected, acquire the played audio to be detected at the audio acquisition points to obtain the acquired audio; based on this, the aforementioned audio analysis module 52 is also used to: perform audio analysis on the acquired audio based on the content difference detection results and the volume difference detection results, determine the audio quality problems and the target audio points in the acquired audio where audio quality problems exist; and generate detection results based on the audio quality problems and the target audio points.
[0086] Furthermore, the above detection task also includes a detection location and a set of audio to be detected; based on this, the audio recording module 51 is used to: determine the audio to be detected corresponding to the detection location from the set of audio to be detected, and play the audio to be detected.
[0087] The audio quality detection device provided in this disclosure has the same implementation principle and technical effect as the aforementioned method embodiment. For the sake of brevity, any parts not mentioned in the device embodiment can be referred to the corresponding content in the aforementioned method embodiment.
[0088] This embodiment also provides an electronic device, as shown in FIG6. The electronic device includes a processor and a memory. The memory stores computer-executable instructions that can be executed by the processor. The processor executes the computer-executable instructions to implement the above-described audio quality detection method. The electronic device can be a server or a terminal device.
[0089] Specifically, the above-mentioned audio quality detection method is applied to a terminal device, which is connected to a server. The method includes: receiving a detection task sent by the server, the detection task including the audio to be detected; playing the audio to be detected and recording the played audio to be detected to obtain the recorded audio; performing time-domain alignment on the recorded audio and the audio to be detected to obtain the time-domain aligned audio, and performing spectral difference analysis on the time-domain aligned audio to obtain the detection result; and returning the detection result to the server.
[0090] The aforementioned audio quality detection method records the audio being played when the audio to be detected, as indicated by the detection task, and uses the audio to be detected as a reference audio. By performing spectral difference analysis and frame-by-frame comparison on the reference audio and the recorded audio, it achieves automated detection of audio quality problems throughout the entire audio playback chain, thereby improving the accuracy and reliability of audio quality detection.
[0091] In an optional embodiment, the steps of performing time-domain alignment on the recorded audio and the audio to be detected to obtain time-domain aligned audio, and performing spectral difference analysis on the time-domain aligned audio to obtain detection results include: performing time-domain alignment on the recorded audio and the audio to be detected to obtain time-domain aligned audio; performing time-domain content consistency detection and volume fluctuation detection on the time-domain aligned audio to obtain detection results.
[0092] In an optional embodiment, the step of performing temporal alignment on the recorded audio and the audio to be detected to obtain temporally aligned audio includes: extracting features from the audio to be detected and the recorded audio respectively to obtain audio features of the audio to be detected and audio features of the recorded audio; determining the feature similarity between the audio features of the audio to be detected and the audio features of the recorded audio; determining the temporal alignment position of the audio to be detected and the recorded audio based on the feature similarity; and performing temporal alignment on the recorded audio and the audio to be detected based on the alignment position to obtain temporally aligned audio.
[0093] In an optional embodiment, the above-mentioned audio features include: Mel frequency cepstral coefficients and / or time-domain waveforms.
[0094] In an optional embodiment, the step of determining the feature similarity between the audio features of the audio to be detected and the audio features of the recorded audio, and determining the alignment position of the audio to be detected and the recorded audio in the time domain based on the feature similarity, includes: sliding a first sliding window of a first preset duration across the audio to be detected and the recorded audio with a preset step size, and determining the feature similarity between the audio features of the audio to be detected and the audio features of the recorded audio in the first sliding window during the sliding process; if the feature similarity is greater than a preset similarity threshold, determining the position of the first sliding window within the recorded audio as the alignment position of the audio to be detected and the recorded audio in the time domain.
[0095] In an optional embodiment, the steps of performing temporal content consistency detection and volume fluctuation detection on the temporally aligned audio to obtain detection results include: calculating the content difference between the recorded audio and the audio to be detected in the temporally aligned audio frame by frame to obtain a content difference detection result; if the content difference detection result indicates that there is no content difference, calculating the volume difference between the recorded audio and the audio to be detected in the temporally aligned audio using a sliding window method to obtain a volume difference detection result; and obtaining the final detection result based on the content difference detection result and the volume difference detection result.
[0096] In an optional embodiment, the step of calculating the content difference between the recorded audio and the audio to be detected in the temporally aligned audio frame by frame to obtain the content difference detection result includes: calculating the bit error rate between the recorded audio and the audio to be detected in the temporally aligned audio frame by frame to obtain the bit error rate corresponding to each frame of audio; and generating the content difference detection result based on the bit error rate corresponding to each frame of audio.
[0097] In an optional embodiment, a bit error rate greater than a preset ratio threshold indicates that there is a content difference in the audio of the corresponding audio frame; based on this, the step of generating a content difference detection result according to the bit error rate corresponding to each audio frame includes: determining the target audio frame in the temporally aligned audio whose bit error rate is greater than the preset ratio threshold, and determining the duration of the target audio frame; generating a content difference detection result based on the target audio and the duration of the target audio.
[0098] In an optional embodiment, the step of calculating the volume difference between the recorded audio and the audio to be detected in the time-domain aligned audio using a sliding window to obtain a volume difference detection result includes: sliding a second sliding window of a second preset duration over the time-domain aligned audio with a preset step size, and determining the volume difference value of the time-domain aligned audio within the second sliding window during the sliding process; if the volume difference value is less than a preset difference threshold, determining that there is a volume difference in the time-domain aligned audio within the second sliding window; and determining the volume difference detection result based on the volume difference value of the time-domain aligned audio.
[0099] In an optional embodiment, the above detection task further includes detection conditions; based on this, the above method further includes: determining audio acquisition points according to the detection conditions; when playing the audio to be detected, acquiring the played audio to be detected at the audio acquisition points to obtain the acquired audio; based on this, the step of obtaining the final detection result based on the content difference detection result and the volume difference detection result includes: performing audio analysis on the acquired audio based on the content difference detection result and the volume difference detection result to determine the audio quality problem and the target audio points in the acquired audio where the audio quality problem exists; generating the detection result based on the audio quality problem and the target audio points.
[0100] In an optional embodiment, the detection task further includes a detection location and a set of audio to be detected; based on this, the step of playing the audio to be detected includes: determining the audio to be detected corresponding to the detection location from the set of audio to be detected, and playing the audio to be detected.
[0101] Furthermore, the electronic device shown in Figure 6 also includes a bus 102 and a communication interface 103, with the processor 101, the communication interface 103, and the memory 100 connected via the bus 102.
[0102] The memory 100 may include high-speed random access memory (RAM) and may also include non-volatile memory, such as at least one disk storage device. Communication between this system network element and at least one other network element is achieved through at least one communication interface 103 (which can be wired or wireless), such as the Internet, wide area network, local area network, metropolitan area network, etc. The bus 102 may be an ISA bus, PCI bus, or EISA bus, etc. The bus can be divided into address bus, data bus, control bus, etc. For ease of illustration, only a single bidirectional arrow is used in Figure 6, but this does not indicate that there is only one bus or one type of bus.
[0103] Processor 101 may be an integrated circuit chip with signal processing capabilities. In implementation, each step of the above method can be completed by the integrated logic circuitry in the hardware of processor 101 or by instructions in software form. The processor 101 can be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it can also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this disclosure. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of this disclosure can be directly manifested as execution by a hardware decoding processor, or execution by a combination of hardware and software modules in the decoding processor. The software module can reside in a readily available storage medium in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, or registers. This storage medium is located in memory 100, and processor 101 reads information from memory 100 and, in conjunction with its hardware, completes the steps of the method described in the foregoing embodiments.
[0104] This disclosure also provides a computer-readable storage medium storing computer-executable instructions. When these computer-executable instructions are called and executed by a processor, they cause the processor to implement the above-described audio quality detection method. For specific implementation details, please refer to the method embodiments, which will not be repeated here.
[0105] Specifically, the above-mentioned audio quality detection method is applied to a terminal device, which is connected to a server. The method includes: receiving a detection task sent by the server, the detection task including the audio to be detected; playing the audio to be detected and recording the played audio to be detected to obtain the recorded audio; performing time-domain alignment on the recorded audio and the audio to be detected to obtain the time-domain aligned audio, and performing spectral difference analysis on the time-domain aligned audio to obtain the detection result; and returning the detection result to the server.
[0106] The aforementioned audio quality detection method records the audio being played when the audio to be detected, as indicated by the detection task, and uses the audio to be detected as a reference audio. By performing spectral difference analysis and frame-by-frame comparison on the reference audio and the recorded audio, it achieves automated detection of audio quality problems throughout the entire audio playback chain, thereby improving the accuracy and reliability of audio quality detection.
[0107] In an optional embodiment, the steps of performing time-domain alignment on the recorded audio and the audio to be detected to obtain time-domain aligned audio, and performing spectral difference analysis on the time-domain aligned audio to obtain detection results include: performing time-domain alignment on the recorded audio and the audio to be detected to obtain time-domain aligned audio; performing time-domain content consistency detection and volume fluctuation detection on the time-domain aligned audio to obtain detection results.
[0108] In an optional embodiment, the step of performing temporal alignment on the recorded audio and the audio to be detected to obtain temporally aligned audio includes: extracting features from the audio to be detected and the recorded audio respectively to obtain audio features of the audio to be detected and audio features of the recorded audio; determining the feature similarity between the audio features of the audio to be detected and the audio features of the recorded audio; determining the temporal alignment position of the audio to be detected and the recorded audio based on the feature similarity; and performing temporal alignment on the recorded audio and the audio to be detected based on the alignment position to obtain temporally aligned audio.
[0109] In an optional embodiment, the above-mentioned audio features include: Mel frequency cepstral coefficients and / or time-domain waveforms.
[0110] In an optional embodiment, the step of determining the feature similarity between the audio features of the audio to be detected and the audio features of the recorded audio, and determining the alignment position of the audio to be detected and the recorded audio in the time domain based on the feature similarity, includes: sliding a first sliding window of a first preset duration across the audio to be detected and the recorded audio with a preset step size, and determining the feature similarity between the audio features of the audio to be detected and the audio features of the recorded audio in the first sliding window during the sliding process; if the feature similarity is greater than a preset similarity threshold, determining the position of the first sliding window within the recorded audio as the alignment position of the audio to be detected and the recorded audio in the time domain.
[0111] In an optional embodiment, the steps of performing temporal content consistency detection and volume fluctuation detection on the temporally aligned audio to obtain detection results include: calculating the content difference between the recorded audio and the audio to be detected in the temporally aligned audio frame by frame to obtain a content difference detection result; if the content difference detection result indicates that there is no content difference, calculating the volume difference between the recorded audio and the audio to be detected in the temporally aligned audio using a sliding window method to obtain a volume difference detection result; and obtaining the final detection result based on the content difference detection result and the volume difference detection result.
[0112] In an optional embodiment, the step of calculating the content difference between the recorded audio and the audio to be detected in the temporally aligned audio frame by frame to obtain the content difference detection result includes: calculating the bit error rate between the recorded audio and the audio to be detected in the temporally aligned audio frame by frame to obtain the bit error rate corresponding to each frame of audio; and generating the content difference detection result based on the bit error rate corresponding to each frame of audio.
[0113] In an optional embodiment, a bit error rate greater than a preset ratio threshold indicates that there is a content difference in the audio of the corresponding audio frame; based on this, the step of generating a content difference detection result according to the bit error rate corresponding to each audio frame includes: determining the target audio frame in the temporally aligned audio whose bit error rate is greater than the preset ratio threshold, and determining the duration of the target audio frame; generating a content difference detection result based on the target audio and the duration of the target audio.
[0114] In an optional embodiment, the step of calculating the volume difference between the recorded audio and the audio to be detected in the time-domain aligned audio using a sliding window to obtain a volume difference detection result includes: sliding a second sliding window of a second preset duration over the time-domain aligned audio with a preset step size, and determining the volume difference value of the time-domain aligned audio within the second sliding window during the sliding process; if the volume difference value is less than a preset difference threshold, determining that there is a volume difference in the time-domain aligned audio within the second sliding window; and determining the volume difference detection result based on the volume difference value of the time-domain aligned audio.
[0115] In an optional embodiment, the above detection task further includes detection conditions; based on this, the above method further includes: determining audio acquisition points according to the detection conditions; when playing the audio to be detected, acquiring the played audio to be detected at the audio acquisition points to obtain the acquired audio; based on this, the step of obtaining the final detection result based on the content difference detection result and the volume difference detection result includes: performing audio analysis on the acquired audio based on the content difference detection result and the volume difference detection result to determine the audio quality problem and the target audio points in the acquired audio where the audio quality problem exists; generating the detection result based on the audio quality problem and the target audio points.
[0116] In an optional embodiment, the detection task further includes a detection location and a set of audio to be detected; based on this, the step of playing the audio to be detected includes: determining the audio to be detected corresponding to the detection location from the set of audio to be detected, and playing the audio to be detected.
[0117] The aforementioned function, when implemented as a software functional unit and sold or used as an independent product, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this disclosure, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, a terminal device, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this disclosure. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0118] In the description of this disclosure, it should be noted that the terms "center," "upper," "lower," "left," "right," "vertical," "horizontal," "inner," and "outer," etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings, and are only for the convenience of describing this disclosure and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of this disclosure. Furthermore, the terms "first," "second," and "third" are used for descriptive purposes only and should not be construed as indicating or implying relative importance.
[0119] Finally, it should be noted that the above-described embodiments are merely specific implementations of this disclosure, used to illustrate the technical solutions of this disclosure, and not to limit it. The protection scope of this disclosure is not limited thereto. Although this disclosure has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that any person skilled in the art can still modify or easily conceive of changes to the technical solutions described in the foregoing embodiments, or make equivalent substitutions for some of the technical features, within the scope of the technology disclosed in this disclosure; and these modifications, changes, or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this disclosure, and should all be covered within the protection scope of this disclosure. Therefore, the protection scope of this disclosure should be determined by the protection scope of the claims.
Claims
1. An audio quality detection method, characterized in that, The method is applied to a terminal device that is communicatively connected to a server. The method includes: receiving a detection task sent by the server, the detection task including an audio to be detected; playing the audio to be detected and recording the played audio to obtain a recorded audio; performing time-domain alignment on the recorded audio and the audio to be detected to obtain a time-domain aligned audio, and performing spectral difference analysis on the time-domain aligned audio to obtain a detection result; and returning the detection result to the server.
2. The method according to claim 1, characterized in that, The steps of performing time-domain alignment on the recorded audio and the audio to be detected to obtain time-domain aligned audio, and performing spectral difference analysis on the time-domain aligned audio to obtain detection results include: performing time-domain alignment on the recorded audio and the audio to be detected to obtain time-domain aligned audio; performing time-domain content consistency detection and volume fluctuation detection on the time-domain aligned audio to obtain detection results.
3. The method according to claim 2, characterized in that, The step of performing temporal alignment between the recorded audio and the audio to be detected to obtain temporally aligned audio includes: extracting features from the audio to be detected and the recorded audio respectively to obtain audio features of the audio to be detected and audio features of the recorded audio; determining the feature similarity between the audio features of the audio to be detected and the audio features of the recorded audio; determining the temporal alignment position of the audio to be detected and the recorded audio based on the feature similarity; and performing temporal alignment between the recorded audio and the audio to be detected based on the alignment position to obtain temporally aligned audio.
4. The method according to claim 3, characterized in that, The step of determining the feature similarity between the audio features of the audio to be detected and the audio features of the recorded audio, and determining the alignment position of the audio to be detected and the recorded audio in the time domain based on the feature similarity, includes: sliding a first sliding window of a first preset duration on the audio to be detected and the recorded audio with a preset step size, and determining the feature similarity between the audio features of the audio to be detected and the audio features of the recorded audio in the first sliding window during the sliding process; if the feature similarity is greater than a preset similarity threshold, determining the position of the first sliding window within the recorded audio as the alignment position of the audio to be detected and the recorded audio in the time domain.
5. The method according to claim 2, characterized in that, The steps of performing temporal content consistency detection and volume fluctuation detection on the temporally aligned audio to obtain detection results include: calculating the content difference between the recorded audio and the audio to be detected in the temporally aligned audio frame by frame to obtain a content difference detection result; if the content difference detection result indicates that there is no content difference, calculating the volume difference between the recorded audio and the audio to be detected in the temporally aligned audio using a sliding window method to obtain a volume difference detection result; and obtaining the final detection result based on the content difference detection result and the volume difference detection result.
6. The method according to claim 5, characterized in that, The step of calculating the content difference between the recorded audio and the audio to be detected in the temporally aligned audio frame by frame to obtain the content difference detection result includes: calculating the bit error rate between the recorded audio and the audio to be detected in the temporally aligned audio frame by frame to obtain the bit error rate corresponding to each frame of audio; and generating the content difference detection result based on the bit error rate corresponding to each frame of audio.
7. The method according to claim 6, characterized in that, The bit error rate being greater than a preset ratio threshold indicates that there is a content difference in the audio of the corresponding audio frame; the step of generating a content difference detection result based on the bit error rate corresponding to each audio frame includes: determining the target audio frame in the temporally aligned audio whose bit error rate is greater than the preset ratio threshold, and determining the duration of the target audio frame; generating a content difference detection result based on the target audio and the duration of the target audio.
8. An audio quality detection device, characterized in that, The device is installed on a terminal device, which is communicatively connected to a server. The device includes: a task receiving module for receiving a detection task sent by the server, the detection task including an audio to be detected; an audio recording module for playing the audio to be detected and recording the played audio to be detected to obtain a recorded audio; an audio analysis module for performing time-domain alignment on the recorded audio and the audio to be detected to obtain time-domain aligned audio, and performing spectral difference analysis on the time-domain aligned audio to obtain a detection result; and a result return module for returning the detection result to the server.
9. An electronic device, characterized in that, The device includes a processor and a memory, the memory storing computer-executable instructions that can be executed by the processor, the processor executing the computer-executable instructions to implement the audio quality detection method according to any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions, which, when invoked and executed by a processor, cause the processor to implement the audio quality detection method according to any one of claims 1-7.