Audio quality detection method and device, equipment, storage medium and product
By performing standardized preprocessing and specific model segmentation on audio files, the problem of low accuracy in traditional audio quality detection in complex scenarios is solved, achieving high-accuracy audio quality detection applicable to various scenarios.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- 浙江省中波发射管理中心
- Filing Date
- 2025-12-31
- Publication Date
- 2026-04-21
AI Technical Summary
Traditional audio quality detection methods rely on fixed parameter thresholds, making it difficult to achieve high-precision audio quality detection in complex scenarios.
By acquiring the audio file to be processed and performing standardized preprocessing, and using a pre-trained audio segmentation model and anomaly type prediction model, audio segmentation and anomaly type prediction are performed to determine the audio quality.
It improves the accuracy of audio quality detection and is applicable to various scenarios, including broadcast, membership and online audio. It realizes the integration from segmentation to instantaneous detection, captures instantaneous anomalies and identifies cumulative risks at the minute level.
Smart Images

Figure CN121905221A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of audio data processing technology, and in particular to an audio quality detection method, apparatus, device, storage medium, and product. Background Technology
[0002] With the explosive growth of audio data in communication, media, and other scenarios, the demand for real-time and accurate audio quality monitoring is becoming increasingly urgent. Traditional speech quality detection methods typically rely on fixed parameter thresholds, such as fixed decibel and energy thresholds, as hard indicators. This approach suffers a significant drop in accuracy when dealing with complex scenarios, resulting in low precision in detecting audio data quality in complex environments. Summary of the Invention
[0003] This invention provides an audio quality testing method, apparatus, device, storage medium, and product to improve the accuracy of audio quality testing.
[0004] According to one aspect of the present invention, an audio quality detection method is provided, the method comprising:
[0005] The audio file to be processed is obtained, and the audio file to be processed is subjected to standardized preprocessing according to the preset processing configuration parameters to obtain a standardized audio frame sequence.
[0006] The standardized audio frame sequence is input into a pre-trained audio segmentation model to obtain at least one audio segmentation sequence output by the model and its corresponding audio type, start timestamp, and end timestamp.
[0007] Audio segmentation sequences of non-speech type and audio containing interference speech type are identified as audio segmentation sequences to be analyzed, and the audio segmentation sequences to be analyzed are divided into sequence windows to obtain at least one audio segment to be analyzed;
[0008] Each of the audio segments to be analyzed is input into a pre-trained audio anomaly type prediction model to obtain the audio anomaly type corresponding to each of the audio segments to be analyzed, as output by the model.
[0009] The audio quality of the audio file to be processed is determined based on the audio anomaly type corresponding to each of the audio segments to be analyzed.
[0010] According to another aspect of the present invention, an audio quality detection device is provided, the device comprising:
[0011] The audio acquisition module is used to acquire the audio file to be processed and perform standardized preprocessing on the audio file to be processed according to preset processing configuration parameters to obtain a standardized audio frame sequence.
[0012] The audio sequence prediction module is used to input the standardized audio frame sequence into a pre-trained audio segmentation model to obtain at least one audio segmentation sequence output by the model and its corresponding audio type, start timestamp and end timestamp respectively.
[0013] The audio segment generation module is used to determine audio segmentation sequences with non-speech and speech-interference types as audio segmentation sequences to be analyzed, and to divide the audio segmentation sequences to be analyzed into sequence windows to obtain at least one audio segment to be analyzed.
[0014] An anomaly type determination module is used to input each of the audio segments to be analyzed into a pre-trained audio anomaly type prediction model, and obtain the audio anomaly type corresponding to each of the audio segments to be analyzed output by the model.
[0015] The audio quality determination module is used to determine the audio quality of the audio file to be processed based on the audio anomaly type corresponding to each of the audio segments to be analyzed.
[0016] According to another aspect of the present invention, an electronic device is provided, the electronic device comprising:
[0017] At least one processor; and
[0018] A memory communicatively connected to the at least one processor; wherein,
[0019] The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the audio quality detection method according to any embodiment of the present invention.
[0020] According to another aspect of the present invention, a computer-readable storage medium is provided, the computer-readable storage medium storing computer instructions for causing a processor to execute and implement the audio quality detection method according to any embodiment of the present invention.
[0021] According to another aspect of the present invention, a computer program product is provided, comprising a computer program that, when executed by a processor, implements the audio quality detection method described in any embodiment of the present invention.
[0022] The technical solution of this invention involves inputting a standardized audio frame sequence into a pre-trained audio segmentation model to obtain at least one audio segmentation sequence output by the model, along with its corresponding audio type, start timestamp, and end timestamp. Audio segmentation sequences with non-speech and speech-containing types are identified as the audio segmentation sequences to be analyzed. The audio segmentation sequences to be analyzed are then divided into sequence windows to obtain at least one audio segment to be analyzed. Each audio segment to be analyzed is input into a pre-trained audio anomaly type prediction model to obtain the audio anomaly type corresponding to each audio segment output by the model. Based on the audio anomaly type corresponding to each audio segment to be analyzed, the audio quality of the audio file to be processed is determined. Using a specific model for audio sequence segmentation improves the accuracy of audio sequence segmentation and sequence type prediction. The integrated approach from segmentation to instantaneous detection captures instantaneous anomalies and identifies cumulative risks at the minute level, improving the accuracy of audio quality detection. It is applicable to multiple scenarios, including broadcast, membership, and online audio, achieving scenario adaptability and flexibility for audio quality detection methods.
[0023] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description
[0024] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0025] Figure 1A This is a flowchart of an audio quality detection method provided in Embodiment 1 of the present invention;
[0026] Figure 1B This is a schematic diagram of the model structure of an audio segmentation model provided in Embodiment 1 of the present invention;
[0027] Figure 2 This is a flowchart of an audio quality detection method provided in Embodiment 2 of the present invention;
[0028] Figure 3 This is a schematic diagram of an audio quality detection device according to Embodiment 3 of the present invention;
[0029] Figure 4 This is a schematic diagram of the structure of an electronic device that implements the audio quality detection method of this invention. Detailed Implementation
[0030] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.
[0031] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0032] Example 1
[0033] Figure 1A This is a flowchart of an audio quality detection method provided in Embodiment 1 of the present invention. This embodiment is applicable to situations where audio data in scenarios such as broadcasting, speeches, or live streaming is being tested for audio quality. The method can be executed by an audio quality detection device, which can be implemented in hardware and / or software and can be configured in an electronic device. Figure 1A As shown, the method includes:
[0034] S110. Obtain the audio file to be processed, and perform standardized preprocessing on the audio file to be processed according to the preset processing configuration parameters to obtain a standardized audio frame sequence.
[0035] S120. Input the standardized audio frame sequence into the pre-trained audio segmentation model to obtain at least one audio segmentation sequence output by the model and its corresponding audio type, start timestamp and end timestamp.
[0036] S130. The audio segmentation sequences with non-speech type and audio containing interference speech type are identified as audio segmentation sequences to be analyzed, and the audio segmentation sequences to be analyzed are divided into sequence windows to obtain at least one audio segment to be analyzed.
[0037] S140. Input each audio segment to be analyzed into the pre-trained audio anomaly type prediction model to obtain the audio anomaly type corresponding to each audio segment to be analyzed output by the model.
[0038] S150. Determine the audio quality of the audio file to be processed based on the audio anomaly type corresponding to each audio segment to be analyzed.
[0039] The audio file to be processed can be an audio file from any scenario, such as a live broadcast or voice call in real-time scenarios, or an offline audio file stored locally or in the cloud. The file format of the audio file to be processed can include a variety of formats. For example, for offline scenarios, the audio file to be processed can be an audio file in formats such as MPS (Music Production Suite), MAV (Multi-channel Audio and Video), and FLAC (Free Lossless Audio Codec); for real-time scenarios, the audio file to be processed can be a real-time audio stream using protocols such as RTMP (Real-Time Messaging Protocol) and FLV (Flash Video).
[0040] The processing configuration parameters can be preset by relevant technical personnel according to actual needs. For example, the processing configuration parameters may include sampling rate, frame length and frame shift for frame processing, and number of audio channels.
[0041] In one optional embodiment, the audio file to be processed is subjected to standardized preprocessing according to preset processing configuration parameters to obtain a standardized audio frame sequence, including: format encoding the audio file to be processed to obtain a format-encoded audio file; parameter normalization of the format-encoded audio file according to the preset sampling rate and preset number of channels in the processing configuration parameters to obtain a normalized audio file; frame segmentation of the normalized audio file according to the preset frame length and preset frame shift in the processing configuration parameters to obtain a frame-segmented audio frame sequence; and frame verification of the frame-segmented audio frame sequence to obtain a standardized audio frame sequence that passes the verification.
[0042] Specifically, existing decoding tools are used to encode audio files of different formats. This involves converting the audio files into uncompressed PCM (Pulse Code Modulation) raw data and eliminating format encoding differences to obtain format-encoded audio files. A sinusoidal interpolation algorithm is then used to uniformly convert the format-encoded audio files based on the preset sampling rate and preset number of channels in the processing configuration parameters; this is called parameter normalization. For example, the preset sampling rate can be set to 16kHz or 18000Hz, and the preset number of channels can be set to mono. Therefore, the format-encoded audio files are uniformly converted to a 16kHz sampling rate and mono format, ensuring audio parameter consistency and obtaining a normalized audio file.
[0043] For example, the preset frame length and preset frame shift can be preset by relevant technical personnel. For instance, the preset frame length can be 20ms (320 sampling points), and the preset frame shift can be 10ms (160 sampling points). Accordingly, the normalized audio file is segmented into frames according to the 20ms frame length and 10ms frame shift to obtain a segmented audio frame sequence. Frame verification is then performed on the segmented audio frame sequence to filter out invalid frames such as all-zero frames and frames with abnormal amplitude values, thereby verifying the integrity and validity of the frame sequence and obtaining a standardized audio frame sequence that passes the verification.
[0044] A standardized audio frame sequence is input into a pre-trained audio segmentation model to obtain at least one audio segmentation sequence output by the model, along with its corresponding audio type, start timestamp, and end timestamp. For example... Figure 1B The diagram shows the model structure of an audio segmentation model. The audio segmentation model includes a feature extraction layer, residual block layers (ResNetBlocks), and fully connected layers.
[0045] Accordingly, the standardized audio frame sequence is input into a pre-trained audio segmentation model to obtain at least one audio segmentation sequence output by the model, along with its corresponding audio type, start timestamp, and end timestamp. This process includes: inputting the standardized audio frame sequence into the feature extraction layer of the pre-trained audio segmentation model for basic feature extraction to obtain basic audio features; inputting the basic audio features into a residual block layer for deep feature learning through multiple residual connections to obtain deep audio features; and inputting the deep audio features into a fully connected layer for feature mapping and classification to obtain at least one audio segmentation sequence and its corresponding audio type, start timestamp, and end timestamp.
[0046] Specifically, the standardized audio sequence is input into the feature extraction layer of the pre-trained audio segmentation model. This layer includes convolutional operations and batch normalization to extract the basic spectral features of the audio. The basic audio features are then input into the residual block layer, which consists of multiple cascaded residual blocks. Each residual block contains a two-layer linear transformation and skip connections, effectively avoiding the gradient vanishing problem of deep networks, thereby capturing deep audio features. The deep audio features are then input into the fully connected layer for dimension mapping and feature fusion, and finally output classification probability and timestamp information to obtain at least one audio sequence and its corresponding audio type, start timestamp, and end timestamp.
[0047] In the feature extraction process, audio feature extraction includes multi-dimensional features: (1) temporal features: RMS (Root Mean Square) energy statistics and zero-crossing rate; (2) frequency domain features: bass (20-250Hz), midrange (250-4000Hz), and treble (4000-8000Hz) power distribution; (3) MFCC (Mel-Frequency Cepstral Coefficients) features: mean and standard deviation of 40-dimensional Mel-frequency cepstral coefficients; (4) Mel (Mel Scale) spectral features: energy distribution of 128 Mel frequency bands; (5) spectral features: spectral center, bandwidth, flatness, roll-off point, and chromaticity features; (6) musical features: tempo, harmonic / percussion ratio, spectral contrast, and tonnertz features. These multi-dimensional features can comprehensively characterize audio content and improve the accuracy of segmentation and recognition.
[0048] Audio types can include non-speech, normal speech, and speech with interference. For example, if the normalized audio frame sequence is a speech sequence with a total length of 210 seconds, the results after processing by the audio segmentation model are: Audio Segmentation Sequence 1, corresponding to the normal speech type, with a start timestamp of 0s and an end timestamp of 60s; Audio Segmentation Sequence 2, corresponding to the non-speech type, with a start timestamp of 60s and an end timestamp of 120s; Audio Segmentation Sequence 3, corresponding to the normal speech type, with a start timestamp of 120s and an end timestamp of 150s; and Audio Segmentation Sequence 4, corresponding to the speech with interference type, with a start timestamp of 150s and an end timestamp of 210s. It is understandable that the audio length can be truncated and padded according to actual needs, for example, truncated and padded to 1-5 seconds.
[0049] In an optional embodiment, in addition to using the above-described audio segmentation model to generate the audio segmentation sequence and determine the audio type, start timestamp, and end timestamp of the audio segmentation sequence, this embodiment also provides another processing method, which pre-sets parameters, including: high energy threshold and low energy threshold, high zero-crossing rate threshold and low zero-crossing rate threshold, spectral flatness threshold, power frequency interference frequency (50Hz or 60Hz), and peak energy threshold of blast sound.
[0050] The root mean square energy per frame (RMSE) and zero-crossing rate (ZCR) of a standardized audio frame sequence are determined. The frames are converted into a spectrum using Fast Fourier Transform (FFT), and spectral flatness is calculated. It should be noted that speech frames have low spectral flatness, while noise frames have high spectral flatness. The energy proportion in key frequency ranges, such as the energy proportion around 50Hz and the energy proportion in high-frequency bands (above 2kHz), is also considered. If the frame RMSE is not less than a preset high energy threshold, and the ZCR is not greater than a preset high ZCR threshold and the spectral flatness is not greater than a preset spectral flatness threshold, then the corresponding sequence frame is marked as a normal speech frame. If the frame RMSE is not greater than a preset low energy threshold, or the ZCR is not less than a preset low ZCR threshold, or the spectral flatness is not less than a preset spectral flatness threshold, then the corresponding sequence frame is marked as a non-speech frame. Frames falling between these two thresholds are marked as ambiguous frames. For ambiguous frames, a neighboring frame state correction is applied. If the three frames before and after the ambiguous frame are all speech frames, then it is determined to be a speech frame; otherwise, it is a non-speech frame. Three or more consecutive frames of the same type are aggregated into a speech segment, and the segmentation results of speech frames and non-speech frames are output.
[0051] Calculate the energy percentage of frequencies around 50Hz or 60Hz in a speech segment. If the energy percentage is not less than 15%, it is marked as a speech segment that may contain current noise interference. Calculate the frame energy abrupt change value in a speech segment, specifically the ratio between the energy of the current frame and the energy of the previous frame. If the abrupt change value is not less than 5 and the peak energy is not less than the peak threshold for blast noise energy, it is marked as a speech segment that may contain blast noise interference. Calculate the energy percentage of the low-frequency band (below 200Hz) in a speech segment. If the energy percentage is not less than 60% and the duration is not less than 3 frames, it is marked as a speech segment that may contain abnormal bass interference.
[0052] Based on the audio segmentation sequence and its corresponding audio type, start timestamp, and end timestamp determined by the above method, and combined with the audio segmentation sequence and its corresponding audio type, start timestamp, and end timestamp determined by the above audio segmentation model method, the results of the two determination methods are integrated, taking into account both the model dimension and the algorithm dimension, thereby improving the accuracy of determining the audio segmentation sequence and its corresponding audio type, start timestamp, and end timestamp.
[0053] Audio segments of non-speech and those containing interfering speech types are identified as the audio segments to be analyzed. These segments are then divided into sequence windows to obtain at least one audio segment to be analyzed. The time detection window for sequence division can be preset by relevant technical personnel according to actual needs. For example, the time detection window can be set to 5 seconds, and the sliding step size can be set to 1 second. That is, the audio segment to be analyzed is divided into detection windows according to a 5-second window and a 1-second sliding step size. Window segments shorter than 5 seconds are retained as complete windows. The windows are then bound to the audio segments to be analyzed and the timestamp information.
[0054] Each audio segment to be analyzed is input into a pre-trained audio anomaly type prediction model, and the model outputs the audio anomaly type corresponding to each audio segment. The audio anomaly type prediction model is used to predict the audio anomaly type of the audio segments, such as electrical noise, popping sounds, silence, and bass anomalies. The audio anomaly type prediction model can be pre-trained by relevant technical personnel. This embodiment also provides a model training method for the audio anomaly type prediction model. The model training method for the audio anomaly type prediction model is as follows:
[0055] Historical audio segments within a given time period are acquired and tagged with type labels to obtain the true anomaly types corresponding to the historical audio segments. The historical audio segments and their corresponding true anomaly types are then input into a pre-selected network model to obtain the predicted anomaly types output by the model. Based on the true anomaly types of the historical audio segments and the preset anomaly types, the network model is trained until the preset model training termination condition is met, resulting in an audio anomaly type prediction model.
[0056] The historical audio segments can be speech segments predicted over historical time periods, or audio segments pre-acquired by relevant technicians from different scenarios. The type tags for historical audio segments can include electrical noise, popping sounds, silence, and abnormal bass.
[0057] Historical audio segments and their corresponding true anomaly types are input into a pre-selected network model to obtain the predicted anomaly type output by the model. The pre-selected network model can be RetNet101 (a residual network with 101 layers). This network model effectively addresses the degradation problem of deep networks through residual learning, enabling the extraction of deep-level audio features. Based on the true anomaly types of historical audio segments and the preset anomaly types, a current loss value is determined for the current iteration time period based on a preset loss function. The network model is then trained based on the current loss value until the current loss value reaches a set loss threshold, or the current loss value stabilizes, or the current iteration count reaches a set iteration count threshold, thus obtaining the audio anomaly type prediction model.
[0058] The audio quality of the audio file to be processed is determined based on the audio anomaly type corresponding to each audio segment to be analyzed. For example, the presence of persistent noise or persistent silence can be determined based on the audio anomaly type corresponding to each audio segment to be analyzed; if such phenomena exist, it indicates that the audio quality is low.
[0059] During training, class weights are automatically calculated to address imbalance issues, and model training parameters such as the learning rate can be preset by relevant technical personnel.
[0060] The technical solution of this invention involves inputting a standardized audio frame sequence into a pre-trained audio segmentation model to obtain at least one audio segmentation sequence output by the model, along with its corresponding audio type, start timestamp, and end timestamp. Audio segmentation sequences with non-speech and speech-containing types are identified as the audio segmentation sequences to be analyzed. The audio segmentation sequences to be analyzed are then divided into sequence windows to obtain at least one audio segment to be analyzed. Each audio segment to be analyzed is input into a pre-trained audio anomaly type prediction model to obtain the audio anomaly type corresponding to each audio segment output by the model. Based on the audio anomaly type corresponding to each audio segment to be analyzed, the audio quality of the audio file to be processed is determined. Using a specific model for audio sequence segmentation improves the accuracy of audio sequence segmentation and sequence type prediction. The integrated approach from segmentation to instantaneous detection captures instantaneous anomalies and identifies cumulative risks at the minute level, improving the accuracy of audio quality detection. It is applicable to multiple scenarios, including broadcast, membership, and online audio, achieving scenario adaptability and flexibility for audio quality detection methods.
[0061] Example 2
[0062] Figure 2 This is a flowchart of an audio quality detection method provided in Embodiment 2 of the present invention. This embodiment is an optimization and improvement based on the above technical solutions.
[0063] Furthermore, the step "determine the audio quality of the audio file to be processed based on the audio anomaly type corresponding to each audio segment to be analyzed" is refined to "determine at least one audio anomaly scenario corresponding to the audio file to be processed based on at least one matching rule in the preset audio anomaly scenario rule base, according to the audio anomaly type and segment time window corresponding to each audio segment to be analyzed; determine the audio quality of the audio file to be processed based on each audio anomaly scenario." This improves the method for determining the audio quality of the audio file to be processed.
[0064] It should be noted that for parts not described in detail in the embodiments of the present invention, please refer to the descriptions in other embodiments. For example... Figure 2As shown, the method includes the following specific steps:
[0065] S210. Obtain the audio file to be processed, and perform standardized preprocessing on the audio file to be processed according to the preset processing configuration parameters to obtain a standardized audio frame sequence.
[0066] S220. Input the standardized audio frame sequence into the pre-trained audio segmentation model to obtain at least one audio segmentation sequence output by the model and its corresponding audio type, start timestamp and end timestamp.
[0067] S230. The audio segmentation sequences with non-speech type and audio containing interference speech type are identified as the audio segmentation sequences to be analyzed, and the audio segmentation sequences to be analyzed are divided into sequence windows to obtain at least one audio segment to be analyzed.
[0068] S240. Input each audio segment to be analyzed into the pre-trained audio anomaly type prediction model to obtain the audio anomaly type corresponding to each audio segment to be analyzed output by the model.
[0069] S250. Based on the audio anomaly type and segment time window corresponding to each audio segment to be analyzed, determine at least one audio anomaly scenario corresponding to the audio file to be processed based on at least one matching rule in the preset audio anomaly scenario rule base.
[0070] S260. Determine the audio quality of the audio file to be processed based on each audio anomaly scenario.
[0071] The audio anomaly scenario rule base includes several matching rules that can be pre-defined by technical personnel. These rules may include continuous noise rules, prolonged silence rules, continuous bass anomaly rules, and custom rules. Continuous noise rules can be defined as detecting the number of electrical hums or popping sounds within 10 seconds, verifying if there are at least 5 consecutive triggers. Prolonged silence rules can be defined as merging consecutive silence windows, verifying if the total duration is at least 1 minute. Continuous bass anomaly rules can be defined as merging consecutive bass anomaly windows, verifying if the total duration is at least 3 minutes. Custom rules can be defined by the user, setting configuration conditions to complete the compatibility verification.
[0072] Specifically, based on the audio anomaly type and segment time window corresponding to each audio segment to be analyzed, the instantaneous detection results are stored in a time-series database according to audio identifiers and timestamps to construct a traceable detection data pool. The preset scenarios in the rule base are converted into time-series analysis logic, including time window statistics, consecutive counts, and duration verification. Rule matching is performed between the instantaneous audio segments in the detection data pool and at least one matching rule in the audio anomaly scenario rule base to determine at least one audio anomaly scenario corresponding to the audio file to be processed, such as persistent noise, prolonged silence, and persistent bass anomalies.
[0073] The audio quality of the audio file to be processed is determined based on each audio anomaly scenario. For example, if there are many matching scenarios for an audio anomaly scenario, the audio quality of the audio file to be processed will be lower; if there are few matching scenarios for an audio anomaly scenario, the audio quality of the audio file to be processed will be higher.
[0074] In one optional embodiment, determining the audio quality of the audio file to be processed based on each audio anomaly scenario includes: determining a first quality detection result based on the scene feature weight parameters corresponding to each audio anomaly scenario; and determining a second quality detection result based on the number of audio anomaly scenarios; and determining the audio quality of the audio file to be processed based on the first quality detection result and the second quality detection result.
[0075] It should be noted that since the impact of audio anomalies on audio quality varies, different scene feature weight parameters can be assigned to different audio anomaly scenarios. For example, the scene feature weight parameter for a continuous noise scenario is set to 0.5, the scene feature weight parameter for a long period of silence is set to 0.3, and the scene feature weight parameter for a continuous bass anomaly is set to 0.2.
[0076] Based on each audio anomaly scenario, and using the corresponding scenario feature weight parameters, a first quality detection result is determined. For example, if the audio anomaly scenarios are a continuous noise scenario and a long period of silence scenario, the quality detection score corresponding to the first quality detection result is 0.7.
[0077] The second quality detection result is determined based on the number of audio abnormality scenarios. For example, if the audio abnormality scenarios are continuous noise scenarios and long-term silence scenarios, the number of scenarios is 2. The quality detection score corresponding to the second quality detection result is the ratio between the number of scenarios and the total number of scenarios. Where the total number of scenarios is 3, the quality detection score corresponding to the second quality detection result is 0.67.
[0078] Based on the first and second quality detection results, the audio quality of the audio file to be processed is determined. Continuing the previous example, the total quality detection score of the audio file to be processed is the average of the quality detection scores corresponding to the first and second quality detection results, which is 0.69. According to the mapping relationship between the total quality detection score and the quality level, the audio quality of the audio file to be processed is determined. For example, a total quality detection score greater than 0.9 corresponds to high quality; a total quality detection score not greater than 0.9 and not less than 0.7 corresponds to medium quality; a total quality detection score less than 0.7 and not less than 0.5 corresponds to low-to-medium quality; and a total quality detection score less than 0.5 corresponds to low quality. Therefore, continuing the previous example, the audio quality of the audio file to be processed is low-to-medium quality.
[0079] This embodiment's technical solution determines at least one audio anomaly scenario corresponding to the audio file to be processed based on at least one matching rule in a preset audio anomaly scenario rule base, according to the audio anomaly type and segment time window corresponding to each audio segment to be analyzed. Based on each audio anomaly scenario, the audio quality of the audio file to be processed is determined. In the process of determining the audio quality, at least one matching rule is comprehensively considered, which achieves accurate determination of audio anomaly scenarios and improves the accuracy of audio quality determination.
[0080] Example 3
[0081] Figure 3 This is a schematic diagram of an audio quality detection device provided in Embodiment 3 of the present invention. The audio quality detection device provided in this embodiment of the present invention is applicable to situations where audio data in scenarios such as broadcasting, speeches, or live streaming is being detected for audio quality. This audio quality detection device can be implemented in hardware and / or software, such as… Figure 3 As shown, the device includes: an audio acquisition module 301, an audio sequence prediction module 302, an audio segment generation module 303, an anomaly type determination module 304, and an audio quality determination module 305. Among them,
[0082] The audio acquisition module 301 is used to acquire the audio file to be processed and perform standardized preprocessing on the audio file to be processed according to preset processing configuration parameters to obtain a standardized audio frame sequence.
[0083] The audio sequence prediction module 302 is used to input the standardized audio frame sequence into a pre-trained audio segmentation model to obtain at least one audio segmentation sequence output by the model and its corresponding audio type, start timestamp and end timestamp respectively.
[0084] The audio segment generation module 303 is used to determine audio segmentation sequences with non-speech and speech-interference types as audio segmentation sequences to be analyzed, and to divide the audio segmentation sequences to be analyzed into sequence windows to obtain at least one audio segment to be analyzed.
[0085] The anomaly type determination module 304 is used to input each of the audio segments to be analyzed into a pre-trained audio anomaly type prediction model to obtain the audio anomaly type corresponding to each of the audio segments to be analyzed output by the model.
[0086] The audio quality determination module 305 is used to determine the audio quality of the audio file to be processed based on the audio anomaly type corresponding to each of the audio segments to be analyzed.
[0087] The technical solution of this invention involves inputting a standardized audio frame sequence into a pre-trained audio segmentation model to obtain at least one audio segmentation sequence output by the model, along with its corresponding audio type, start timestamp, and end timestamp. Audio segmentation sequences with non-speech and speech-containing types are identified as the audio segmentation sequences to be analyzed. The audio segmentation sequences to be analyzed are then divided into sequence windows to obtain at least one audio segment to be analyzed. Each audio segment to be analyzed is input into a pre-trained audio anomaly type prediction model to obtain the audio anomaly type corresponding to each audio segment output by the model. Based on the audio anomaly type corresponding to each audio segment to be analyzed, the audio quality of the audio file to be processed is determined. Using a specific model for audio sequence segmentation improves the accuracy of audio sequence segmentation and sequence type prediction. The integrated approach from segmentation to instantaneous detection captures instantaneous anomalies and identifies cumulative risks at the minute level, improving the accuracy of audio quality detection. It is applicable to multiple scenarios, including broadcast, membership, and online audio, achieving scenario adaptability and flexibility for audio quality detection methods.
[0088] Optionally, the audio quality determination module 305 includes:
[0089] The audio abnormal scene determination unit is used to determine at least one audio abnormal scene corresponding to the audio file to be processed based on at least one matching rule in the preset audio abnormal scene rule base, according to the audio abnormal type and segment time window corresponding to each of the audio segments to be analyzed.
[0090] The audio quality determination unit is used to determine the audio quality of the audio file to be processed based on each of the aforementioned audio abnormality scenarios.
[0091] Optional, audio quality determination unit, specifically used for:
[0092] Based on each of the aforementioned audio anomaly scenarios, and using the scene feature weight parameters corresponding to each of the aforementioned audio anomaly scenarios, a first quality detection result is determined; and,
[0093] The second quality detection result is determined based on the number of audio abnormality scenarios.
[0094] The audio quality of the audio file to be processed is determined based on the first quality detection result and the second quality detection result.
[0095] Optionally, the audio segmentation model includes a feature extraction layer, a residual block layer, and a fully connected layer based on a residual network; correspondingly, the audio sequence prediction module 302 is specifically used for:
[0096] The standardized audio frame sequence is input into the feature extraction layer of a pre-trained audio segmentation model to extract basic features and obtain basic audio features.
[0097] The basic audio features are input into the residual block layer, and deep feature learning is performed through multiple residual connections to obtain deep audio features.
[0098] The deep audio features are input into a fully connected layer for feature mapping and classification, resulting in at least one audio segmentation sequence and its corresponding audio type, start timestamp, and end timestamp.
[0099] Optionally, the training method for the audio anomaly type prediction model is as follows:
[0100] Obtain historical audio segments within a historical time period, and label the historical audio segments with type tags to obtain the actual anomaly type corresponding to the historical audio segments;
[0101] The historical audio segments and their corresponding actual anomaly types are input into a pre-selected network model to obtain the predicted anomaly types output by the model.
[0102] Based on the actual anomaly types and preset anomaly types of the historical audio segments, the network model is trained until the preset model training termination condition is met, thus obtaining an audio anomaly type prediction model.
[0103] Optionally, the audio acquisition module 301 is used for:
[0104] The audio file to be processed is format-encoded to obtain a format-encoded audio file;
[0105] Based on the preset sampling rate and preset number of channels in the processing configuration parameters, the format-encoded audio file is normalized to obtain a normalized audio file;
[0106] Based on the preset frame length and preset frame shift in the processing configuration parameters, the normalized audio file is divided into frames to obtain a framed audio frame sequence.
[0107] Frame verification is performed on the segmented audio frame sequence to obtain a standardized audio frame sequence that passes the verification.
[0108] The audio quality detection device provided in this embodiment of the invention can execute the audio quality detection method provided in any embodiment of the invention, and has the corresponding functional modules and beneficial effects of the method.
[0109] Example 4
[0110] Figure 4 A schematic diagram of an electronic device 40 that can be used to implement embodiments of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (e.g., helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.
[0111] like Figure 4 As shown, the electronic device 40 includes at least one processor 41 and a memory, such as a read-only memory (ROM) 42 or a random access memory (RAM) 43, communicatively connected to the at least one processor 41. The memory stores computer programs executable by the at least one processor. The processor 41 can perform various appropriate actions and processes based on the computer program stored in the ROM 42 or loaded from storage unit 48 into the RAM 43. The RAM 43 may also store various programs and data required for the operation of the electronic device 40. The processor 41, ROM 42, and RAM 43 are interconnected via a bus 44. An input / output (I / O) interface 45 is also connected to the bus 44.
[0112] Multiple components in electronic device 40 are connected to I / O interface 45, including: input unit 46, such as keyboard, mouse, etc.; output unit 47, such as various types of monitors, speakers, etc.; storage unit 48, such as disk, optical disk, etc.; and communication unit 49, such as network card, modem, wireless transceiver, etc. Communication unit 49 allows electronic device 40 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0113] Processor 41 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 41 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 41 performs the various methods and processes described above, such as audio quality detection methods.
[0114] In some embodiments, the audio quality detection method may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 48. In some embodiments, part or all of the computer program may be loaded and / or installed on electronic device 40 via ROM 42 and / or communication unit 49. When the computer program is loaded into RAM 43 and executed by processor 41, one or more steps of the audio quality detection method described above may be performed. Alternatively, in other embodiments, processor 41 may be configured to perform the audio quality detection method by any other suitable means (e.g., by means of firmware).
[0115] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0116] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0117] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.
[0118] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0119] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or middleware components (e.g., application servers), or frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.
[0120] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.
[0121] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and no limitation is imposed herein.
[0122] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.
Claims
1. An audio quality detection method, characterized in that, include: The audio file to be processed is obtained, and the audio file to be processed is subjected to standardized preprocessing according to the preset processing configuration parameters to obtain a standardized audio frame sequence. The standardized audio frame sequence is input into a pre-trained audio segmentation model to obtain at least one audio segmentation sequence output by the model and its corresponding audio type, start timestamp, and end timestamp. Audio segmentation sequences of non-speech type and audio containing interference speech type are identified as audio segmentation sequences to be analyzed, and the audio segmentation sequences to be analyzed are divided into sequence windows to obtain at least one audio segment to be analyzed; Each of the audio segments to be analyzed is input into a pre-trained audio anomaly type prediction model to obtain the audio anomaly type corresponding to each of the audio segments to be analyzed, as output by the model. The audio quality of the audio file to be processed is determined based on the audio anomaly type corresponding to each of the audio segments to be analyzed.
2. The method according to claim 1, characterized in that, The step of determining the audio quality of the audio file to be processed based on the audio anomaly type corresponding to each of the audio segments to be analyzed includes: Based on the audio anomaly type and segment time window corresponding to each of the audio segments to be analyzed, at least one audio anomaly scenario corresponding to the audio file to be processed is determined based on at least one matching rule in the preset audio anomaly scenario rule base. The audio quality of the audio file to be processed is determined based on each of the aforementioned audio anomaly scenarios.
3. The method according to claim 2, characterized in that, Determining the audio quality of the audio file to be processed based on each of the aforementioned audio anomaly scenarios includes: Based on each of the aforementioned audio anomaly scenarios, and using the scene feature weight parameters corresponding to each of the aforementioned audio anomaly scenarios, a first quality detection result is determined; and, The second quality detection result is determined based on the number of audio abnormality scenarios. The audio quality of the audio file to be processed is determined based on the first quality detection result and the second quality detection result.
4. The method according to claim 1, characterized in that, The audio segmentation model includes a feature extraction layer, a residual block layer, and a fully connected layer based on a residual network; correspondingly, the step of inputting the standardized audio frame sequence into the pre-trained audio segmentation model to obtain at least one audio segmentation sequence output by the model and its corresponding audio type, start timestamp, and end timestamp includes: The standardized audio frame sequence is input into the feature extraction layer of a pre-trained audio segmentation model to extract basic features and obtain basic audio features. The basic audio features are input into the residual block layer, and deep feature learning is performed through multiple residual connections to obtain deep audio features. The deep audio features are input into a fully connected layer for feature mapping and classification, resulting in at least one audio segmentation sequence and its corresponding audio type, start timestamp, and end timestamp.
5. The method according to claim 1, characterized in that, The training method for the audio anomaly type prediction model is as follows: Obtain historical audio segments within a historical time period, and label the historical audio segments with type tags to obtain the actual anomaly type corresponding to the historical audio segments; The historical audio segments and their corresponding actual anomaly types are input into a pre-selected network model to obtain the predicted anomaly types output by the model. Based on the actual anomaly types and preset anomaly types of the historical audio segments, the network model is trained until the preset model training termination condition is met, thus obtaining an audio anomaly type prediction model.
6. The method according to claim 1, characterized in that, The step of performing standardized preprocessing on the audio file to be processed according to preset processing configuration parameters to obtain a standardized audio frame sequence includes: The audio file to be processed is format-encoded to obtain a format-encoded audio file; Based on the preset sampling rate and preset number of channels in the processing configuration parameters, the format-encoded audio file is normalized to obtain a normalized audio file; Based on the preset frame length and preset frame shift in the processing configuration parameters, the normalized audio file is divided into frames to obtain a framed audio frame sequence. Frame verification is performed on the segmented audio frame sequence to obtain a standardized audio frame sequence that passes the verification.
7. An audio quality detection device, characterized in that, include: The audio acquisition module is used to acquire the audio file to be processed and perform standardized preprocessing on the audio file to be processed according to preset processing configuration parameters to obtain a standardized audio frame sequence. The audio sequence prediction module is used to input the standardized audio frame sequence into a pre-trained audio segmentation model to obtain at least one audio segmentation sequence output by the model and its corresponding audio type, start timestamp and end timestamp respectively. The audio segment generation module is used to determine audio segmentation sequences with non-speech and speech-interference types as audio segmentation sequences to be analyzed, and to divide the audio segmentation sequences to be analyzed into sequence windows to obtain at least one audio segment to be analyzed. An anomaly type determination module is used to input each of the audio segments to be analyzed into a pre-trained audio anomaly type prediction model, and obtain the audio anomaly type corresponding to each of the audio segments to be analyzed output by the model. The audio quality determination module is used to determine the audio quality of the audio file to be processed based on the audio anomaly type corresponding to each of the audio segments to be analyzed.
8. An electronic device, characterized in that, The electronic device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the audio quality detection method according to any one of claims 1-6.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that cause a processor to execute the audio quality detection method according to any one of claims 1-6.
10. A computer program product, characterized in that, The computer program product includes a computer program that, when executed by a processor, implements the audio quality detection method according to any one of claims 1-6.