Audio signal anomaly monitoring method, device, equipment, medium and program product
By analyzing multiple preset indicators of the audio signal and setting the threshold range, the problem of low efficiency in audio signal quality assessment in the existing technology is solved, and efficient and accurate monitoring of audio signal anomalies is achieved.
Patent Information
- Application Number
- CN202210241730.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-11
- Publication Date
- 2025-09-09
- Estimated Expiration
- 2042-03-11
AI Technical Summary
The efficiency of analyzing audio signal quality in the existing technology is low, resulting in inaccurate and inefficient video conference call quality assessment.
By obtaining the audio signal at the current moment in the target scene, analyzing multiple preset signal monitoring indicators, including abnormal signals and abnormal operations, using a neural network model to calculate the analysis results of the signal monitoring indicators, determining the quality quantitative index value of the audio signal, and setting a threshold range to determine whether the audio signal has an abnormality.
It realizes comprehensive and accurate analysis of audio signal quality, improves the efficiency of audio signal anomaly monitoring, and can detect and accurately judge the abnormal situation of audio signals in real time.
Smart Images

Figure CN114627897B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of audio processing technology, and in particular to a method, apparatus, device, medium, and program product for monitoring audio signal anomalies. Background Art
[0002] With the continuous development of Internet technology, remote communication is becoming more and more widely used in various fields. Many companies choose video conferencing systems for remote communication. During the call, the audio signal may be subject to various interferences or failures, which may affect the call quality of the video conference.
[0003] In related technologies, when the call quality of a video conference is poor, staff evaluate the quality of the audio signal during the call based on historical experience to determine the cause of the poor call quality of the video conference and make improvements.
[0004] However, the related art is less efficient in analyzing the quality of audio signals. Summary of the Invention
[0005] Based on this, it is necessary to provide an audio signal anomaly monitoring method, device, equipment, medium and program product that can improve the efficiency of analyzing audio signal quality in order to address the above technical problems.
[0006] In a first aspect, the present application provides a method for monitoring anomalies in an audio signal, the method comprising:
[0007] Get the audio signal at the current moment in the target scene;
[0008] Analyze multiple preset signal monitoring indicators of the audio signal to obtain analysis results of each signal monitoring indicator;
[0009] Determine the quality quantitative index value of the audio signal based on the analysis results of each signal monitoring index;
[0010] If the quality quantification index value of the audio signal is not within a preset threshold range, it is determined that an abnormality exists in the audio signal.
[0011] In one embodiment, the signal monitoring indicators include at least abnormal signals and abnormal operations; multiple preset monitoring indicators of the audio signal are analyzed to obtain analysis results of each monitoring indicator, including:
[0012] Obtaining signal parameter values of an audio signal;
[0013] If the signal monitoring indicator is an abnormal signal indicator, then analyzing whether there is an abnormal signal in each frame of the audio signal according to the signal parameter value of the audio signal to obtain a signal analysis result;
[0014] If the signal monitoring indicator is an abnormal operation indicator, then the signal parameter value of the audio signal is analyzed to determine whether there is an abnormal operation in each frame of the audio signal to obtain an operation analysis result.
[0015] In one embodiment, the abnormal signal is an interrupt signal; and analyzing each frame of the audio signal to determine whether the abnormal signal exists, and obtaining a signal analysis result, includes:
[0016] Perform frame processing on the audio signal to obtain the peak-to-peak value and spectrum maximum value of the audio signal in each frame;
[0017] If the peak-to-peak level is within the first preset level peak threshold range, and the spectrum maximum is within the first preset spectrum threshold range, it is determined that the signal analysis result is that an interrupt signal exists in the audio signal.
[0018] In one embodiment, the abnormal signal is a howling signal; and analyzing each frame of the audio signal to determine whether the abnormal signal exists, and obtaining a signal analysis result, includes:
[0019] The audio signal is divided into frames to obtain the peak-to-average value of the audio signal level of each frame and the second-order phase change between two consecutive frames;
[0020] Obtaining the number of consecutive frames that meet the howling condition; the howling condition is that the level peak average value is within a second preset level peak threshold range, and the second-order phase change between two consecutive frames is within a preset second-order phase change threshold range;
[0021] If the number of consecutive frames is greater than a preset threshold, it is determined that the signal analysis result is that a howling signal exists in the audio signal.
[0022] In one embodiment, the abnormal signal is an echo signal, and analyzing whether there is an abnormal signal in each frame of the audio signal to obtain a signal analysis result includes:
[0023] The audio signal is divided into frames to obtain the maximum peak value of the power cepstrum of each frame of the audio signal;
[0024] If the maximum peak value of the power cepstrum is within a preset power cepstrum threshold range, it is determined that the signal analysis result is that an echo signal exists in the audio signal.
[0025] In one embodiment, the abnormal signal is an interference signal; and analyzing each frame of the audio signal to determine whether the abnormal signal exists, and obtaining a signal analysis result, includes:
[0026] Perform frame processing on the audio signal to obtain the peak-to-peak level and sound source loudness of each frame of the audio signal;
[0027] If the peak-to-peak level is within the second preset level peak threshold range, and the sound source loudness is within the preset loudness threshold range, it is determined that the signal analysis result is that an interference signal exists in the audio signal.
[0028] In one embodiment, analyzing whether there is abnormal operation in each frame of audio signal to obtain the operation analysis result includes:
[0029] Obtain the sound source audio signal when the sound source is activated in the target scene;
[0030] Analyze multiple signal monitoring indicators of the sound source audio signal to obtain a quality quantification value of the sound source audio signal;
[0031] If the quality quantization value of the sound source audio signal is within a preset threshold range, it is determined that the operation analysis result is abnormal in the startup sound source operation.
[0032] In one embodiment, determining a quality quantification index value of an audio signal based on analysis results of each signal monitoring index includes:
[0033] Determine the quality quantification index value of each audio evaluation index based on the analysis results of each signal monitoring index;
[0034] The weighted average of the quality quantification index values of the various audio evaluation indicators is calculated to obtain the quality quantification index value of the audio signal.
[0035] In a second aspect, the present application further provides an audio signal anomaly monitoring device, the device comprising:
[0036] An acquisition module is used to obtain the audio signal at the current moment in the target scene;
[0037] An analysis module is used to analyze multiple preset signal monitoring indicators of the audio signal and obtain analysis results of each signal monitoring indicator;
[0038] A first determination module is used to determine a quality quantification index value of the audio signal based on the analysis results of each signal monitoring index;
[0039] The second determining module is configured to determine that an abnormality exists in the audio signal when the quality quantification index value of the audio signal is not within a preset threshold range.
[0040] In a third aspect, the present application further provides a computer device, which includes a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, all contents of the method embodiment of the first aspect are implemented.
[0041] In a fourth aspect, the present application further provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, all the contents of the method embodiment of the first aspect are implemented.
[0042] In a fifth aspect, the present application also provides a computer program product, which includes a computer program. When the computer program is executed by a processor, all the contents of the method embodiment of the first aspect are implemented.
[0043] The above-mentioned audio signal anomaly monitoring method, device, equipment, medium and program product, the method obtains the audio signal at the current moment in the target scene, analyzes multiple preset signal monitoring indicators of the audio signal, obtains the analysis results of each signal monitoring indicator, and then determines the quality quantification index value of the audio signal based on the analysis results of each signal monitoring indicator. If the quality quantification index value of the audio signal is not within the preset threshold range, it is determined that the audio signal is abnormal. The method uses multiple preset signal monitoring indicators to analyze, making the quality quantification index value of the audio signal more comprehensive and more accurate, and then by setting a preset threshold range, it can more accurately determine whether the audio signal is abnormal; at the same time, by obtaining the audio signal at the current moment in the target scene in real time, the audio signal can be monitored in real time, thereby improving the efficiency of analyzing the quality of the audio signal. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] Figure 1 A diagram showing an application environment of a method for monitoring anomalies in an audio signal according to an embodiment;
[0045] Figure 2 1 is a flow chart of a method for monitoring anomalies in an audio signal according to an embodiment;
[0046] Figure 3 1 is a flow chart of a method for monitoring anomalies in an audio signal according to an embodiment;
[0047] Figure 4 1 is a flow chart of a method for monitoring anomalies in an audio signal according to an embodiment;
[0048] Figure 5 1 is a flow chart of a method for monitoring anomalies in an audio signal according to an embodiment;
[0049] Figure 6 1 is a flow chart of a method for monitoring anomalies in an audio signal according to an embodiment;
[0050] Figure 7 1 is a flow chart of a method for monitoring anomalies in an audio signal according to an embodiment;
[0051] Figure 81 is a flow chart of a method for monitoring anomalies in an audio signal according to an embodiment;
[0052] Figure 9 1 is a flow chart of a method for monitoring anomalies in an audio signal according to an embodiment;
[0053] Figure 10 1 is a flow chart of a method for monitoring anomalies in an audio signal according to an embodiment;
[0054] Figure 11 1 is a flow chart of a method for monitoring anomalies in an audio signal according to an embodiment;
[0055] Figure 12 FIG. 4 is a structural block diagram of an audio signal abnormality monitoring device in one embodiment. DETAILED DESCRIPTION
[0056] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.
[0057] The audio signal anomaly monitoring method provided in the embodiment of the present application can be applied to Figure 1 The application environment shown in FIG. This application environment includes computer devices, which may include, but are not limited to, various personal computers, laptops, smartphones, tablets, IoT devices, and portable wearable devices. IoT devices may include smart speakers, smart TVs, smart air conditioners, and smart car devices. Portable wearable devices may include smart watches, smart bracelets, and head-mounted devices.
[0058] In one embodiment, Figure 2 As shown, a method for monitoring abnormality of an audio signal is provided, which is applied to Figure 1 The computer device in the example is used to illustrate the process, including the following steps:
[0059] S201: Acquire the audio signal of the target scene at the current moment.
[0060] The target scene may be one branch scene or multiple branch scenes.
[0061] Optionally, the computer device can be connected to the audio acquisition device via Bluetooth, or the computer device can be connected to the audio acquisition device via wireless network communication technology (Wireless Fidelity, WIFI). This embodiment does not limit the connection method between the computer device and the audio acquisition device.
[0062] Furthermore, it is understood that the audio capture device can acquire the current audio signal in real time, and the audio capture device transmits the acquired audio signal to the computer device, so that the computer device can acquire the audio signal in the target scene in real time. For example, when the target scene is sub-venue 1, the computer device can acquire the audio signal of sub-venue 1 in real time; when the target scene is sub-venues 1 to 5, the computer device can acquire the audio signals of sub-venues 1 to 5 in real time.
[0063] S202: Analyze multiple preset signal monitoring indicators of the audio signal to obtain analysis results of each signal monitoring indicator.
[0064] Among them, the preset signal monitoring indicators can be abnormal operation indicators and abnormal signal indicators.
[0065] Specifically, the computer device may input an audio signal into a preset neural network model, calculate multiple preset signal monitoring indicators in the audio signal through the neural network model, and output analysis results for each signal monitoring indicator. The analysis results for each signal monitoring indicator may be scores for the multiple preset signal monitoring indicators, or the computer device may determine whether the multiple preset signal monitoring indicators are within a normal range, and the analysis results may be normal or abnormal for the multiple preset signal monitoring indicators.
[0066] S203: Determine a quality quantification index value of the audio signal according to the analysis results of each signal monitoring index.
[0067] Specifically, when the analysis result of each signal monitoring indicator is a score for multiple preset signal monitoring indicators, the computer device can calculate the average value of the scores of multiple preset signal monitoring indicators, and use the average value of the scores of multiple preset signal monitoring indicators as the quality quantification index value of the audio signal, or the computer device can obtain the historical weight values of multiple preset signal monitoring indicators, calculate the weighted average value of the scores of multiple preset signal monitoring indicators, and use the weighted average value of the scores of multiple preset signal monitoring indicators as the quality quantification index value of the audio signal.
[0068] Furthermore, it can be understood that when the analysis results of each signal monitoring indicator are that multiple preset signal monitoring indicators are normal or abnormal, the computer device can combine the historical weight values of multiple preset signal monitoring indicators to determine the normal number and abnormal number of multiple preset signal monitoring indicators, and determine the percentage value of the normal number of multiple preset signal monitoring indicators to the total number as the quality quantitative index value of the audio signal.
[0069] S204: If the quality quantification index value of the audio signal is not within a preset threshold range, it is determined that an abnormality exists in the audio signal.
[0070] Specifically, the preset threshold range can be set based on historical experience, and the computer device can determine whether the quality quantization index value of the audio signal is within the preset threshold range. When the quality quantization index value of the audio signal is within the preset threshold range, the audio signal is not abnormal; when the quality quantization index value of the audio signal is not within the preset threshold range, it is determined that the audio signal is abnormal. For example, when the preset threshold range is set based on historical experience between 80 and 100 points, when the quality quantization value of the audio signal is 90, the audio signal is not abnormal; when the quality quantization value of the audio signal is 75, the audio signal is abnormal.
[0071] In the above-mentioned audio signal anomaly monitoring method, the method obtains the audio signal at the current moment in the target scene, analyzes multiple preset signal monitoring indicators of the audio signal, obtains the analysis results of each signal monitoring indicator, and then determines the quality quantification indicator value of the audio signal based on the analysis results of each signal monitoring indicator. If the quality quantification indicator value of the audio signal is not within the preset threshold range, it is determined that the audio signal is abnormal. This method uses multiple preset signal monitoring indicators for analysis, making the quality quantification indicator value of the audio signal more comprehensive and more accurate. By setting the preset threshold range, it can more accurately determine whether the audio signal is abnormal. At the same time, by obtaining the audio signal at the current moment in the target scene in real time, the audio signal can be monitored in real time, thereby improving the efficiency of analyzing the quality of the audio signal.
[0072] Figure 3 A flowchart of an audio signal anomaly monitoring method provided in an embodiment of the present application. The embodiment of the present application involves a signal monitoring indicator including at least abnormal signals and abnormal operations; an optional implementation method is to analyze multiple preset monitoring indicators of the audio signal to obtain the analysis results of each monitoring indicator. Figure 2 Based on the embodiment shown, Figure 3 As shown, the above S202 may include the following steps:
[0073] S301: Obtain signal parameter values of an audio signal.
[0074] Specifically, the signal parameter values may include peak-to-peak levels, spectral values, second-order phase change between two consecutive frames, and peak values of power cepstrum. The computer device may obtain the signal parameter values of the audio signal in real time using relevant algorithms. For example, the computer device may obtain the peak-to-peak level value of each audio signal using a Gaussian function, or obtain the maximum value of the spectrum of each audio signal using a Fourier transform.
[0075] S302: If the signal monitoring indicator is an abnormal signal indicator, analyze whether there is an abnormal signal in each frame of the audio signal according to the signal parameter value of the audio signal to obtain a signal analysis result.
[0076] Specifically, abnormal signal indicators may include interruption signals, howling signals, echo signals, and interference signals. The computer device can analyze each frame of the audio signal based on peak-to-peak level, spectrum value, second-order phase change between two consecutive frames, and peak value of the power cepstrum to determine whether an abnormal signal is present in the audio signal. When the signal parameter values meet the preset abnormal signal conditions, an abnormal signal is present in each frame of the audio signal; when the signal parameter values do not meet the preset abnormal signal conditions, no abnormal signal is present in each frame of the audio signal.
[0077] S303: If the signal monitoring indicator is an abnormal operation indicator, analyze whether there is an abnormal operation in each frame of the audio signal according to the signal parameter value of the audio signal to obtain an operation analysis result.
[0078] Specifically, abnormal operation refers to the failure to start the audio capture device as required during the process of capturing audio data from the target scene, resulting in abnormalities in the captured audio data. The computer device can analyze each frame of the audio signal using the aforementioned signal parameter values to determine whether abnormal operation has occurred in the audio signal.
[0079] In the above-mentioned audio signal anomaly monitoring method, the method obtains the signal parameter value of the audio signal. If the signal monitoring indicator is an abnormal signal indicator, the method analyzes whether there is an abnormal signal in each frame of the audio signal based on the signal parameter value of the audio signal to obtain a signal analysis result. If the signal monitoring indicator is an abnormal operation indicator, the method analyzes whether there is an abnormal operation in each frame of the audio signal based on the signal parameter value of the audio signal to obtain an operation analysis result. The monitoring indicators in this method are divided into abnormal signal indicators and abnormal operation indicators. The parameter values of the audio signal can be used to judge the two indicators separately, which can comprehensively determine whether there is an abnormality in the audio signal, making the obtained analysis results more comprehensive.
[0080] Figure 4 A flowchart of an audio signal abnormality monitoring method provided in an embodiment of the present application. The embodiment of the present application involves an optional implementation method in which the abnormal signal is an interrupt signal; analyzing whether there is an abnormal signal in each frame of audio signal to obtain a signal analysis result. Figure 3 Based on the embodiment shown, Figure 4 As shown, the above S302 may include the following steps:
[0081] S401 , performing frame processing on the audio signal to obtain the peak-to-peak value of the audio signal level and the maximum value of the frequency spectrum phase change of each frame.
[0082] Specifically, the computer device may segment the audio signal to obtain multiple small audio signal segments, with each segment being considered a frame. The audio signals of each frame may overlap. The computer device may fit each audio signal frame using a Gaussian function with a variable reference to obtain the peak-to-peak value of each audio signal frame. Furthermore, the computer device may calculate each audio signal frame using a Fourier transform to obtain the maximum value of the frequency spectrum of each audio signal frame.
[0083] S402: If the peak-to-peak level is within a first preset peak level threshold range, and the spectrum maximum is within a first preset spectrum threshold range, it is determined that the signal analysis result is that an interrupt signal exists in the audio signal.
[0084] The interruption signal refers to a situation where, during the audio signal acquisition process of the audio acquisition device, the audio signal of the current frame acquired cannot be connected with the audio signal of the previous frame due to some reason.
[0085] Specifically, the computer device can determine whether there is an interruption signal in the audio signal based on the peak-to-peak value of the level and the maximum value of the spectrum. The first preset level peak threshold range and the first preset spectrum threshold range are both set based on historical experience. The computer device can determine whether the peak-to-peak value of the audio signal is within the first preset level peak threshold range. The peak-to-peak value of the audio signal is within the first preset level peak threshold range, and the maximum value of the spectrum of the audio signal is within the first preset spectrum threshold range. When both the peak-to-peak value of the audio signal and the maximum value of the spectrum of the audio signal meet the conditions at the same time, it is determined that there is an interruption signal in the audio signal.
[0086] In the above-mentioned audio signal anomaly monitoring method, the method frames the audio signal to obtain the peak-to-peak level and spectral maximum of each frame. If the peak-to-peak level falls within a first preset level peak threshold range, and the spectral maximum falls within a first preset spectral threshold range, the signal analysis result determines that an interruption signal is present in the audio signal. This method simultaneously determines the peak-to-peak level and spectral maximum of the audio signal, using these two indicators to more accurately determine whether an interruption signal exists in the audio signal, thereby making the resulting determination more accurate.
[0087] Figure 5 A flowchart of an audio signal abnormality monitoring method provided in an embodiment of the present application. The embodiment of the present application involves an optional implementation method in which the abnormal signal is a howling signal; analyzing whether there is an abnormal signal in each frame of audio signal to obtain a signal analysis result. Figure 3 Based on the embodiment shown, Figure 5 As shown, the above S302 may include the following steps:
[0088] S501 , performing frame processing on the audio signal to obtain the peak-to-average value of the audio signal level of each frame and the second-order phase change between two consecutive frames.
[0089] Among them, the second-order phase change refers to the change in signal phase before and after the audio signal passes through the system.
[0090] Specifically, the computer device can divide the audio signal into multiple small segments of audio signals, and regard each small segment as a frame. There are overlapping parts between the audio signals of each frame. The computer device can use a Gaussian function with a variable reference to fit the audio signal of each frame to obtain the peak-to-peak value of the audio signal of each frame, and the computer device can obtain the second-order phase change between two consecutive frames through the Fourier function.
[0091] S502, obtaining the number of consecutive frames that meet the howling condition; the howling condition is that the level peak average value is within a second preset level peak threshold range, and the second-order phase change between two consecutive frames is within a preset second-order phase change threshold range.
[0092] Specifically, the computer device can determine whether there is a howling signal in the audio signal based on the peak-to-peak level and the second-order phase change between two consecutive frames. The second preset level peak threshold range and the preset second-order phase change threshold range are both set based on historical experience. The computer device can calculate the number of consecutive frame audio signals that meet the above two conditions and determine the number as the number of consecutive frame audio data that meet the howling condition.
[0093] S503: If the number of consecutive frames is greater than a preset threshold, determine that the signal analysis result is that a howling signal exists in the audio signal.
[0094] Howling is a common abnormal sound signal in video conferencing. Howling occurs when a microphone converts analog audio signals into digital signals when picking up audio. This processed audio signal is then amplified through speakers and other amplification equipment. This amplified audio signal inevitably feeds back into the microphone, creating this abnormal signal. Howling signals are primarily categorized as single-frequency howling caused by unsaturated single-frequency oscillation and saturated howling caused by self-oscillation.
[0095] Specifically, the preset threshold value is obtained through historical experience. After the number of consecutive frames that meet the howling condition is determined in step S502, the computer device can compare the number of consecutive frames that meet the howling condition with the preset threshold value. When the number of consecutive frames that meet the howling condition is greater than the preset threshold value, a howling signal exists in the audio signal; when the number of consecutive frames that meet the howling condition is less than or equal to the preset threshold value, no howling signal exists in the audio signal.
[0096] In the above-mentioned audio signal anomaly monitoring method, the method frames the audio signal, obtains the peak-to-average value of each frame of the audio signal and the second-order phase change between two consecutive frames, and determines the number of consecutive frames that meet the howling condition. The howling condition is that the peak-to-average value is within a second preset peak-to-peak threshold range, and the second-order phase change between two consecutive frames is within a preset second-order phase change threshold range. If the number of consecutive frames is greater than the preset threshold value, the signal analysis result is determined to indicate the presence of a howling signal in the audio signal. This method uses both the peak-to-average value of the audio signal and the second-order phase change between two consecutive frames to make a more accurate judgment on whether a howling signal exists in the audio signal, thereby making the judgment result more accurate.
[0097] Figure 6 A flowchart of an audio signal abnormality monitoring method provided in an embodiment of the present application. The embodiment of the present application involves an optional implementation method in which the abnormal signal is an echo signal; analyzing whether there is an abnormal signal in each frame of audio signal to obtain a signal analysis result. Figure 3 Based on the embodiment shown, Figure 6 As shown, the above S302 may include the following steps:
[0098] S601 : performing frame processing on the audio signal to obtain the maximum peak value of the power cepstrum of each frame of the audio signal.
[0099] Specifically, the computer device can divide the audio signal into multiple small segments of audio signals, and regard each small segment as a frame. There are overlapping parts between the audio signals of each frame. The computer device performs power spectrum calculation and filtering on the audio signal to obtain a weighted sinusoidal signal, and then obtains the maximum peak value of the power inverse spectrum of each frame of audio signal through logarithm calculation and power spectrum transformation.
[0100] S602: If the maximum peak value of the power cepstrum is within a preset power cepstrum threshold range, it is determined that the signal analysis result is that an echo signal exists in the audio signal.
[0101] During a video conference, echo signals are divided into direct echo and indirect echo based on the path they are generated. Direct echo occurs when the audio signal, after being emitted by the local sound reinforcement equipment, directly re-enters the microphone without any reflection path. Direct echo is characterized by short latency and is affected by the audio signal energy, the sound reinforcement equipment playback volume, the microphone's sensitivity and directionality, and the distance and angle between the sound reinforcement equipment and the microphone. Indirect echo occurs when the audio signal emitted by the sound reinforcement equipment reflects one or more times off walls, the ground, or other different paths before entering the microphone. Indirect echo is characterized by long latency, high jitter, and is easily affected by the venue environment.
[0102] Specifically, the preset power cepstrum threshold range is obtained through historical experience. The computer device can compare the maximum peak value of the power cepstrum with the preset power cepstrum threshold range. When the maximum peak value of the power cepstrum is within the preset power cepstrum threshold range, an echo signal exists in the audio signal; when the maximum peak value of the power cepstrum is not within the preset power cepstrum threshold range, no echo signal exists in the audio signal.
[0103] Furthermore, it is understandable that the computer device can input an audio signal with an echo signal into an adaptive filter, use the adaptive filter to simulate the path of the echo signal, process the echo signal, subtract the echo estimation value from the processed audio signal, and obtain a residual signal, that is, an audio signal without an echo signal, but this signal may also include an echo residual signal, and then perform residual echo cancellation on the signal to complete the adaptive echo cancellation process, that is, eliminate the echo in the audio signal.
[0104] In the aforementioned audio signal anomaly monitoring method, the method frames the audio signal and obtains the maximum peak value of the power cepstrum of each frame. If the maximum peak value of the power cepstrum falls within a preset power cepstrum threshold, the signal analysis result indicates the presence of an echo signal in the audio signal. This method uses the maximum peak value of the power cepstrum to more accurately determine whether an echo signal is present in the audio signal.
[0105] Figure 7 A flowchart of an audio signal anomaly monitoring method provided in an embodiment of the present application. The embodiment of the present application involves an optional implementation method in which the abnormal signal is an interference signal; analyzing whether there is an abnormal signal in each frame of audio signal to obtain a signal analysis result. Figure 3 Based on the embodiment shown, Figure 7 As shown, the above S302 may include the following steps:
[0106] S701 : Perform frame processing on the audio signal to obtain the peak-to-peak level and sound source loudness of each frame of the audio signal.
[0107] Specifically, the computer device can divide the audio signal into multiple small segments of audio signals, and regard each small segment as a frame. There are overlapping parts between the audio signals of each frame. The computer device can use a Gaussian function with a variable reference to fit the audio signals of each frame to obtain the peak-to-peak value of the audio signal of each frame. In addition, the computer device can determine the loudness of the sound source of each frame of audio signal by the vibration amplitude of the sound.
[0108] S702: If the peak-to-peak level is within a second preset peak-to-peak level threshold range, and the sound source loudness is within a preset loudness threshold range, it is determined that the signal analysis result is that an interference signal exists in the audio signal.
[0109] Among them, the interference signal can be monitored by the peak-to-peak level and the loudness of the sound source.
[0110] Specifically, the second preset level peak threshold range and the preset loudness threshold range are obtained through historical experience. The computer device may compare the peak-to-peak level value with the second preset level peak threshold range, and simultaneously compare the loudness of the sound source with the preset loudness threshold range. When the peak-to-peak level value is within the second preset level peak threshold range and an echo signal is present in the audio signal, no interference signal is present in the audio signal; when the peak-to-peak level value is not within the second preset level peak threshold range, or when no echo signal is present in the audio signal, or the peak-to-peak level value is not within the second preset level peak threshold range and no echo signal is present in the audio signal, no interference signal is present in the audio signal.
[0111] In the above-mentioned audio signal anomaly monitoring method, the method frames the audio signal to obtain the peak-to-peak level and sound source loudness of each frame. If the peak-to-peak level is within a second preset level peak threshold range, and the sound source loudness is within a preset loudness threshold range, the signal analysis result determines that an interference signal is present in the audio signal. This method simultaneously determines the presence of an interference signal using both the peak-to-peak level and the sound source loudness. The peak-to-peak level can accurately determine the presence of an interference signal in the audio signal, while the sound source loudness can accurately determine the presence of a loudness exceeding the limit in the audio signal. Using these two indicators, the presence of an interference signal in the audio signal can be more accurately determined.
[0112] Figure 8 The present invention provides a flowchart of an audio signal anomaly monitoring method. The present invention relates to an optional implementation method for analyzing whether there is abnormal operation in each frame of audio signal and obtaining the operation analysis result. Figure 3 Based on the embodiment shown, Figure 8 As shown, the above S303 may include the following steps:
[0113] S801: Acquire a sound source audio signal when a sound source is activated in a target scene.
[0114] Specifically, since the audio signal may be abnormal due to operational errors when the audio acquisition device is started to record the target scene, when the audio acquisition device collects the sound source audio signal in the target scene, the audio acquisition device sends the collected sound source audio signal in the target scene to the computer device, and the computer device can obtain the sound source audio signal when the sound source is started in the target scene.
[0115] S802: Analyze multiple signal monitoring indicators of the sound source audio signal to obtain a quality quantization value of the sound source audio signal.
[0116] Among them, multiple signal monitoring indicators of the sound source audio signal may include signal-to-noise ratio, harmonic distortion, crosstalk, intermodulation distortion, noise level and frequency response. The signal-to-noise ratio is the logarithm of the ratio of signal power to noise power in the audio signal. The signal-to-noise ratio is directly proportional to the effect of the audio signal, that is, the higher the signal-to-noise ratio, the better the audio signal quality; harmonic distortion is the change caused by the audio amplifier device to the signal waveform. The harmonic distortion is inversely proportional to the performance of the audio amplifier device. The smaller the harmonic distortion, the better the performance of the audio amplifier device; crosstalk is the interference between different signal lines. Mutual inductance and mutual capacitance are likely to occur between two signal lines. Noise will be generated between signal lines. In video conferencing systems, the longest signal lines are more likely to cause crosstalk. Intermodulation distortion refers to the mutual interference between different signal frequencies. It is caused by the mutual modulation of signals when power amplifiers and other equipment amplify the audio. It is the distortion of the sum and difference of the excitation signals, and the mutual interference between different signal frequencies. The noise level is mainly used to measure the background noise level of the sound amplification equipment when there is no signal input or output. It is used to evaluate whether the sound amplification equipment itself will introduce large noise. The frequency response is a parameter that measures the change in output signal amplitude based on the standard frequency signal level.
[0117] Specifically, the computer device may input a sound source audio signal into a preset neural network model, calculate multiple preset signal monitoring indicators in the sound source audio signal through the neural network model, and output analysis results for each signal monitoring indicator. The analysis results for each signal monitoring indicator may be scores for the multiple preset signal monitoring indicators, or the computer device may determine whether the multiple preset signal monitoring indicators are within a normal range, and the analysis results may be normal or abnormal for the multiple preset signal monitoring indicators.
[0118] S803: If the quality quantization value of the sound source audio signal is within a preset threshold range, determine that the operation analysis result is abnormal operation of the starting sound source.
[0119] Specifically, the preset threshold range can be set based on historical experience, and the computer device can determine whether the quality quantification index value of the sound source audio signal is within the preset threshold range. When the quality quantification index value of the sound source audio signal is within the preset threshold range, there is no abnormality in the sound source audio signal; when the quality quantification index value of the sound source audio signal is not within the preset threshold range, it is determined that there is an abnormality in the sound source audio signal.
[0120] In the above-mentioned audio signal anomaly monitoring method, the method obtains the sound source audio signal when the sound source is activated in the target scene, analyzes multiple signal monitoring indicators of the sound source audio signal, and obtains a quality quantification value of the sound source audio signal. If the quality quantification value of the sound source audio signal is within a preset threshold range, the operation analysis result is determined to be an abnormal operation of the activation sound source. By obtaining the sound source audio signal when the sound source is activated and accurately processing the sound source audio signal using multiple signal monitoring indicators, the method makes the judgment result of the activation sound source more accurate.
[0121] Figure 9 The present invention provides a flowchart of an audio signal anomaly monitoring method. The present invention relates to an optional implementation method for determining the quality quantification index value of an audio signal based on the analysis results of each signal monitoring index. Figure 2 Based on the embodiment shown, Figure 9 As shown, the above S203 may include the following steps:
[0122] S901: Determine the quality quantification index value of each audio evaluation index according to the analysis results of each signal monitoring index.
[0123] Specifically, when the analysis result of each signal monitoring indicator is a score for multiple preset signal monitoring indicators, the computer device can use the score of multiple preset signal monitoring indicators as the quality quantification index value of each audio evaluation indicator; when the analysis result of each signal monitoring indicator is that multiple preset signal monitoring indicators are normal or abnormal, the computer device can determine the percentage of the normal number of multiple preset signal monitoring indicators to the total number as the quality quantification index value of the audio signal.
[0124] S902: Calculate the weighted average of the quality quantization index values of the various audio evaluation indicators to obtain the quality quantization index value of the audio signal.
[0125] Specifically, the computer device can calculate the weighted value of the quality quantification index value of each audio evaluation index based on the historical weight of each audio evaluation index, divide the weighted value of the quality quantification index value of each audio evaluation index by the total number of audio evaluation indicators, and determine the calculation result as the quality quantification index value of the audio signal.
[0126] In the aforementioned audio signal anomaly monitoring method, the method determines the quality quantification index value of each audio evaluation index based on the analysis results of each signal monitoring index, and then calculates the weighted average of the quality quantification index values of each audio evaluation index to obtain the quality quantification index value of the audio signal. By calculating the weighted average of the quality quantification values of multiple audio evaluation indicators, this method makes the obtained quality quantification index value of the audio signal more accurate.
[0127] In one embodiment, in order to facilitate understanding by those skilled in the art, the audio signal abnormality monitoring method is described in detail below. Figure 10 As shown, the method may include:
[0128] S1001, obtaining the audio signal of the target scene at the current moment;
[0129] S1002, performing frame processing on the audio signal to obtain the peak-to-peak value and the maximum value of the spectrum of each frame of the audio signal;
[0130] S1003: If the peak-to-peak level is within a first preset peak level threshold range, and the spectrum maximum is within a first preset spectrum threshold range, determining that the signal analysis result indicates that an interrupt signal exists in the audio signal;
[0131] S1004, performing frame processing on the audio signal to obtain the peak-to-average value of the audio signal level of each frame and the second-order phase change between two consecutive frames;
[0132] S1005, obtaining the number of consecutive frames that meet the howling condition;
[0133] S1006, if the number of consecutive frames is greater than a preset threshold, determining that the signal analysis result indicates that a howling signal exists in the audio signal;
[0134] S1007, performing frame processing on the audio signal to obtain the maximum peak value of the power cepstrum of each frame of the audio signal;
[0135] S1008, if the maximum peak value of the power cepstrum is within a preset power cepstrum threshold range, determining that the signal analysis result is that an echo signal exists in the audio signal;
[0136] S1009, performing frame processing on the audio signal to obtain the peak-to-peak value of the audio signal level and the sound source loudness of each frame;
[0137] S1010: If the peak-to-peak level is within a second preset peak level threshold range, and the sound source loudness is within a preset loudness threshold range, determining that the signal analysis result indicates that an interference signal exists in the audio signal;
[0138] S1011, obtaining a sound source audio signal when a sound source is activated in a target scene;
[0139] S1012, analyzing multiple signal monitoring indicators of the sound source audio signal to obtain a quality quantification value of the sound source audio signal;
[0140] S1013, if the quality quantization value of the sound source audio signal is within a preset threshold range, determining that the operation analysis result is abnormal operation of the starting sound source;
[0141] S1014, determining a quality quantification index value of each audio evaluation index based on the analysis results of each signal monitoring index;
[0142] S1015, calculating a weighted average of the quality quantization index values of the various audio evaluation indicators to obtain a quality quantization index value of the audio signal;
[0143] S1016: If the quality quantification index value of the audio signal is not within the preset threshold range, it is determined that an abnormality exists in the audio signal.
[0144] It should be noted that for the descriptions in S1001-S1016 above, reference can be made to the relevant descriptions in the above embodiments, and the effects are similar, so this embodiment will not be repeated here.
[0145] Furthermore, it is understandable that Figure 11 A flow chart showing the voice conversion method is provided. The main venue in the video conferencing system sends a network command to each branch venue through a network protocol. After each branch venue receives the network command from the main venue, the microphone of each branch venue obtains the audio data and related parameters of each branch venue, and sends the audio data and related parameters of the microphone of each branch venue to the main venue. The audio data is analyzed using audio quality evaluation indicators to determine whether there are abnormal signals in the audio data. The evaluation results are then stored in a database and displayed on the interface of the video conferencing system.
[0146] In the above-mentioned audio signal abnormality monitoring method, the method obtains the audio signal at the current moment in the target scene, performs frame processing on the audio signal, obtains the level peak-to-peak value and the spectrum maximum value of each frame of the audio signal, and if the level peak-to-peak value is within a first preset level peak threshold range, and the spectrum maximum value is within the first preset spectrum threshold range, it is determined that the signal analysis result is that there is an interruption signal in the audio signal, the audio signal is framed, the level peak average value of each frame of the audio signal and the second-order phase change between two consecutive frames are obtained, and the number of consecutive frames that meet the howling condition is obtained. If the number of consecutive frames is greater than the preset threshold value, it is determined that the signal analysis result is that there is a howling signal in the audio signal, the audio signal is framed, the maximum peak value of the power cepstrum of each frame of the audio signal is obtained, and if the maximum peak value of the power cepstrum is within the preset power cepstrum threshold range, it is determined that the signal analysis result is that there is an echo signal in the audio signal. The method comprises the steps of: obtaining a signal signal, performing frame processing on the audio signal, obtaining a peak-to-peak level and a sound source loudness of the audio signal of each frame, if the peak-to-peak level is within a second preset level peak threshold range, and the sound source loudness is within a preset loudness threshold range, determining that the signal analysis result is that an interference signal exists in the audio signal, obtaining the sound source audio signal when the sound source is started in the target scene, analyzing multiple signal monitoring indicators of the sound source audio signal, and obtaining a quality quantization value of the sound source audio signal, if the quality quantization value of the sound source audio signal is within a preset threshold range, determining that the operation analysis result is that the operation of starting the sound source is abnormal, determining the quality quantization index value of each audio evaluation index based on the analysis result of each signal monitoring index, calculating the weighted average of the quality quantization index values of each audio evaluation index, and obtaining the quality quantization index value of the audio signal, if the quality quantization index value of the audio signal is not within the preset threshold range, determining that the audio signal is abnormal. This method uses multiple preset signal monitoring indicators for analysis, making the quality quantification index values of audio signals more comprehensive and accurate. By setting a preset threshold range, it can more accurately determine whether there are abnormalities in the audio signal. At the same time, by acquiring the audio signal at the current moment in the target scene in real time, the audio signal can be monitored in real time, improving the efficiency of analyzing the audio signal quality.
[0147] It should be understood that, although the steps in the flowcharts of the above embodiments are shown in sequence as indicated by the arrows, these steps are not necessarily performed in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and these steps can be performed in other orders. Moreover, at least a portion of the steps in the flowcharts of the above embodiments may include multiple steps or multiple stages, and these steps or stages are not necessarily performed at the same time, but can be performed at different times. The execution order of these steps or stages is not necessarily to be performed in sequence, but can be performed in turn or alternately with other steps or at least a portion of steps or stages in other steps.
[0148] Based on the same inventive concept, embodiments of the present application also provide an audio signal anomaly monitoring device for implementing the aforementioned audio signal anomaly monitoring method. The solution provided by this device is similar to the solution described in the aforementioned method. Therefore, the specific limitations of one or more of the following embodiments of the audio signal anomaly monitoring device can be found in the aforementioned limitations of the audio signal anomaly monitoring method and will not be further elaborated here.
[0149] In one embodiment, Figure 12 As shown, an audio signal anomaly monitoring device is provided, comprising: an acquisition module 11, an analysis module 12, a first determination module 13 and a second determination module 14, wherein:
[0150] An acquisition module 11 is used to acquire the audio signal at the current moment in the target scene;
[0151] An analysis module 12 is configured to analyze a plurality of preset signal monitoring indicators of the audio signal and obtain analysis results of each signal monitoring indicator;
[0152] A first determining module 13 is configured to determine a quality quantification index value of the audio signal based on the analysis results of each signal monitoring index;
[0153] The second determining module 14 is configured to determine that an abnormality exists in the audio signal when the quality quantification index value of the audio signal is not within a preset threshold range.
[0154] The audio signal anomaly monitoring device provided in this embodiment can execute the above method embodiment. Its implementation principles and technical effects are similar and will not be repeated here.
[0155] In one embodiment, the analysis module includes: an acquisition unit, a first analysis unit, and a second analysis unit, wherein:
[0156] An acquisition unit, configured to acquire a signal parameter value of an audio signal;
[0157] The first analysis unit is configured to analyze whether there is an abnormal signal in each frame of the audio signal according to the signal parameter value of the audio signal when the signal monitoring indicator is an abnormal signal indicator, and obtain a signal analysis result;
[0158] The second analyzing unit is configured to analyze whether there is an abnormal operation in each frame of the audio signal according to the signal parameter value of the audio signal when the signal monitoring indicator is an abnormal operation indicator, and obtain an operation analysis result.
[0159] The audio signal anomaly monitoring device provided in this embodiment can execute the above method embodiment. Its implementation principles and technical effects are similar and will not be repeated here.
[0160] Optionally, the above-mentioned first analysis unit is specifically used to perform frame processing on the audio signal to obtain the peak-to-peak value and the spectrum maximum value of each frame of the audio signal; if the peak-to-peak value of the level is within the first preset level peak threshold range, and the spectrum maximum value is within the first preset spectrum threshold range, it is determined that the signal analysis result is that there is an interruption signal in the audio signal.
[0161] The audio signal anomaly monitoring device provided in this embodiment can execute the above method embodiment. Its implementation principles and technical effects are similar and will not be repeated here.
[0162] Optionally, the above-mentioned first analysis unit is specifically used to perform frame processing on the audio signal, obtain the peak-to-peak average value of the audio signal of each frame and the second-order phase change between two consecutive frames; obtain the number of consecutive frames that meet the howling condition; the howling condition is that the peak-to-peak average value of the audio signal is within the second preset peak-to-peak threshold range, and the second-order phase change between two consecutive frames is within the preset second-order phase change threshold range; if the number of consecutive frames is greater than the preset threshold value, it is determined that the signal analysis result is that a howling signal exists in the audio signal.
[0163] The audio signal anomaly monitoring device provided in this embodiment can execute the above method embodiment. Its implementation principles and technical effects are similar and will not be repeated here.
[0164] Optionally, the first analysis unit is specifically configured to perform frame processing on the audio signal to obtain the maximum peak value of the power cepstrum of each frame of the audio signal; if the maximum peak value of the power cepstrum is within a preset power cepstrum threshold range, the signal analysis result is determined to be that an echo signal exists in the audio signal.
[0165] The audio signal anomaly monitoring device provided in this embodiment can execute the above method embodiment. Its implementation principles and technical effects are similar and will not be repeated here.
[0166] Optionally, the first analysis unit is specifically configured to perform frame processing on the audio signal to obtain the peak-to-peak value of the audio signal in each frame and the loudness of the sound source; if the peak-to-peak value is within a second preset peak-to-peak value threshold range, and the loudness of the sound source is within a preset loudness threshold range, it is determined that the signal analysis result indicates that an interference signal exists in the audio signal.
[0167] The audio signal anomaly monitoring device provided in this embodiment can execute the above method embodiment. Its implementation principles and technical effects are similar and will not be repeated here.
[0168] Optionally, the above-mentioned second analysis unit is specifically used to obtain the sound source audio signal when the sound source is started in the target scene; analyze multiple signal monitoring indicators of the sound source audio signal to obtain the quality quantization value of the sound source audio signal; if the quality quantization value of the sound source audio signal is within a preset threshold range, it is determined that the operation analysis result is an abnormal operation of starting the sound source.
[0169] The audio signal anomaly monitoring device provided in this embodiment can execute the above method embodiment. Its implementation principles and technical effects are similar and will not be repeated here.
[0170] In one embodiment, the first determination module includes: determining the quality quantification index value of each audio evaluation index based on the analysis results of each signal monitoring index; calculating the weighted average of the quality quantification index values of each audio evaluation index to obtain the quality quantification index value of the audio signal.
[0171] The audio signal anomaly monitoring device provided in this embodiment can execute the above method embodiment. Its implementation principles and technical effects are similar and will not be repeated here.
[0172] Each module in the above-mentioned audio signal anomaly monitoring device can be implemented in whole or in part through software, hardware, or a combination thereof. Each module can be embedded in or independent of a processor in a computer device in hardware form, or can be stored in a memory in the computer device in software form, so that the processor can call and execute the corresponding operations of each module.
[0173] In one embodiment, a computer device is provided, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, all contents in all method embodiments are implemented.
[0174] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, all contents in all method embodiments are implemented.
[0175] In one embodiment, a computer program product is provided, comprising a computer program, which implements all contents of all method embodiments when executed by a processor.
[0176] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties.
[0177] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, database or other media used in the embodiments provided in this application may include at least one of non-volatile and volatile memory. Non-volatile memory may include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory may include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The database involved in the various embodiments provided herein may include at least one of a relational database and a non-relational database. Non-relational databases may include, but are not limited to, distributed databases based on blockchains. The processor involved in the various embodiments provided herein may be, but are not limited to, a general-purpose processor, a central processing unit, a graphics processing unit, a digital signal processor, a programmable logic unit, a data processing logic unit based on quantum computing, and the like.
[0178] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0179] The above embodiments merely illustrate several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present invention. It should be noted that a person skilled in the art may make various modifications and improvements without departing from the spirit of the present invention, all of which fall within the scope of protection of the present application. Therefore, the scope of protection of the present application shall be determined by the appended claims.
Claims
1. A method for monitoring anomalies in audio signals, characterized in that: The method comprises: Acquire an audio signal at a current moment in a target scene, and obtain a signal parameter value of the audio signal; the audio signal is acquired by an audio acquisition device; If the signal monitoring indicator of the audio signal is an abnormal signal indicator, analyzing whether there is an abnormal signal in each frame of the audio signal according to the signal parameter value of the audio signal to obtain a signal analysis result; if the signal monitoring indicator of the audio signal is an abnormal operation indicator, obtaining a sound source audio signal when the sound source is started in the target scene, analyzing multiple signal monitoring indicators of the sound source audio signal to obtain a quality quantization value of the sound source audio signal, and if the quality quantization value of the sound source audio signal is within a preset threshold range, determining that the operation analysis result is an abnormal operation of the starting sound source; Determining a quality quantization index value of each audio evaluation index based on the signal analysis result and the operation analysis result, and calculating a weighted average of the quality quantization index values of each audio evaluation index to obtain a quality quantization index value of the audio signal; If the quality quantification index value of the audio signal is not within a preset threshold range, determining that the audio signal is abnormal; The abnormal signal is an interruption signal; and the analyzing whether there is an abnormal signal in each frame of the audio signal according to the signal parameter value of the audio signal to obtain a signal analysis result includes: performing frame processing on the audio signal, obtaining a peak-to-peak value and a spectrum maximum value of the audio signal in each frame, and if the peak-to-peak value is within a first preset peak-to-peak value range and the spectrum maximum value is within a first preset spectrum threshold range, determining that the signal analysis result is that an interruption signal is present in the audio signal.
2. The method according to claim 1, characterized in that The abnormal signal is a howling signal; and the analyzing whether there is an abnormal signal in each frame of the audio signal to obtain a signal analysis result includes: Performing frame processing on the audio signal to obtain the peak-to-average value of the audio signal level in each frame and the second-order phase change between two consecutive frames; Obtaining the number of consecutive frames that meet a howling condition; the howling condition being that the level peak average is within a second preset level peak threshold range, and the second-order phase change between the two consecutive frames is within a preset second-order phase change threshold range; If the number of consecutive frames is greater than a preset threshold, it is determined that the signal analysis result is that the howling signal exists in the audio signal.
3. The method according to claim 1, characterized in that The abnormal signal is an echo signal, and the analyzing whether there is an abnormal signal in each frame of the audio signal to obtain a signal analysis result includes: Performing frame processing on the audio signal to obtain the maximum peak value of the power cepstrum of the audio signal in each frame; If the maximum peak value of the power cepstrum is within a preset power cepstrum threshold range, it is determined that the signal analysis result is that an echo signal exists in the audio signal.
4. The method according to claim 1, wherein The abnormal signal is an interference signal; and the analyzing whether there is an abnormal signal in each frame of the audio signal to obtain a signal analysis result includes: Performing frame processing on the audio signal to obtain the peak-to-peak value of the audio signal level and the sound source loudness of each frame; If the peak-to-peak level is within a second preset level peak threshold range, and the sound source loudness is within a preset loudness threshold range, it is determined that the signal analysis result is that an interference signal exists in the audio signal.
5. An audio signal abnormality monitoring device, characterized in that: The device comprises: An acquisition module, configured to acquire an audio signal at a current moment in a target scene and obtain a signal parameter value of the audio signal; the audio signal is acquired by an audio acquisition device; an analysis module for, if the signal monitoring indicator of the audio signal is an abnormal signal indicator, analyzing whether there is an abnormal signal in each frame of the audio signal according to the signal parameter value of the audio signal, and obtaining a signal analysis result; if the signal monitoring indicator of the audio signal is an abnormal operation indicator, obtaining a sound source audio signal when the sound source is started in the target scene, analyzing multiple signal monitoring indicators of the sound source audio signal, and obtaining a quality quantization value of the sound source audio signal; if the quality quantization value of the sound source audio signal is within a preset threshold range, determining that the operation analysis result is an abnormal operation of the starting sound source; a first determining module, configured to determine a quality quantization index value of each audio evaluation index based on the signal analysis result and the operation analysis result, and calculate a weighted average of the quality quantization index values of the respective audio evaluation indexes to obtain a quality quantization index value of the audio signal; a second determining module, configured to determine that an abnormality exists in the audio signal if the quality quantification index value of the audio signal is not within a preset threshold range; The abnormal signal is an interruption signal; and the analyzing whether there is an abnormal signal in each frame of the audio signal according to the signal parameter value of the audio signal to obtain a signal analysis result includes: performing frame processing on the audio signal, obtaining a peak-to-peak value and a spectrum maximum value of the audio signal in each frame, and if the peak-to-peak value is within a first preset peak-to-peak value range and the spectrum maximum value is within a first preset spectrum threshold range, determining that the signal analysis result is that an interruption signal is present in the audio signal.
6. The device according to claim 5, characterized in that The abnormal signal is a howling signal; and the analysis module is specifically configured to: Performing frame processing on the audio signal to obtain the peak-to-average value of the audio signal level in each frame and the second-order phase change between two consecutive frames; Obtaining the number of consecutive frames that meet a howling condition; the howling condition being that the level peak average is within a second preset level peak threshold range, and the second-order phase change between the two consecutive frames is within a preset second-order phase change threshold range; If the number of consecutive frames is greater than a preset threshold, it is determined that the signal analysis result is that the howling signal exists in the audio signal.
7. The device according to claim 5, characterized in that The abnormal signal is an echo signal; the analysis module is specifically used to: Performing frame processing on the audio signal to obtain the maximum peak value of the power cepstrum of the audio signal in each frame; If the maximum peak value of the power cepstrum is within a preset power cepstrum threshold range, it is determined that the signal analysis result is that an echo signal exists in the audio signal.
8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 4 are implemented.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 4 are implemented.
10. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 4 are implemented.
Citation Information
Patent Citations
Fluency determination method and device, electronic equipment and storage medium
CN111124868A
Video conference system and audio quality diagnosis method thereof
CN111641799A