Spraying detection method, computer equipment and storage medium
By applying Kalman filtering technology and energy state fusion in audio signal frames, the energy threshold is dynamically adjusted, and the misjudgment problem of spraying detection in complex environments is solved, and the accuracy of spraying detection is improved.
Patent Information
- Application Number
- CN202510715123.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-30
- Publication Date
- 2025-08-15
AI Technical Summary
The existing spray detection methods are prone to misjudgment in complex audio environments, resulting in low accuracy of spray detection.
By obtaining the actual energy state of the audio signal frame, and using Kalman filtering technology for prediction and fusion, combining the proportion of short-time signal energy and high-frequency energy, the energy threshold is dynamically adjusted for spraying and spray detection.
Effectively overcome environmental noise and audio fluctuations interference, and improve the accuracy of spraying microphone detection.
Smart Images

Figure CN120496563A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of audio processing technology, and in particular to a method for detecting a microphone ejection, a computer device, a storage medium, and a computer program product. Background Art
[0002] With the development of audio processing technology, a technology for recording songs using singing software has emerged. Users usually need to use audio capture equipment for recording during the singing process. However, during the singing recording process, if the audio capture equipment is placed too close, it will cause popping sound, resulting in the occurrence of microphone popping, which in turn leads to the destruction of the sound quality of the recording. Therefore, during the recording process, microphone popping detection is required to avoid the microphone popping phenomenon.
[0003] Traditionally, microphone pop-out detection is typically implemented by setting a fixed energy threshold. By monitoring the energy level of the audio signal, a microphone pop-out is detected when the energy level exceeds the preset threshold. However, this method is prone to misjudgment in complex audio environments and cannot accurately identify microphone pop-out. Consequently, existing microphone pop-out detection methods have low accuracy. Summary of the Invention
[0004] Based on this, it is necessary to provide a method, device, computer equipment, computer-readable storage medium and computer program product for detecting a sprayed microphone that can improve the accuracy of the detection of a sprayed microphone in order to address the above technical problems.
[0005] In a first aspect, the present application provides a method for detecting wheat spraying, comprising:
[0006] Obtaining an actual energy state of each audio signal frame among a plurality of audio signal frames included in the audio signal to be detected; wherein the actual energy state is an energy state obtained by performing energy detection on the audio signal frame;
[0007] When the current audio signal frame is the first audio signal frame of the multiple audio signal frames, performing comprehensive processing on actual energy states of at least one audio signal frame among the multiple audio signal frames to obtain a target energy state of the current audio signal frame;
[0008] When the current audio signal frame is any audio signal frame after the first audio signal frame of the multiple audio signal frames, predicting the energy state of the current audio signal frame by using the target energy state of the previous audio signal frame to obtain the predicted energy state of the current audio signal frame, and fusing the actual energy state of the current audio signal frame with the predicted energy state to obtain the target energy state of the current audio signal frame;
[0009] The target energy state of the current audio signal frame is compared with a preset energy state threshold to obtain a microphone ejection detection result of the current audio signal frame.
[0010] In one embodiment, the fusing the actual energy state of the current audio signal frame and the predicted energy state to obtain the target energy state of the current audio signal frame includes: performing observation generation processing on the actual energy state of the current audio signal frame to obtain the observed energy state of the current audio signal frame; obtaining the Kalman gain of the current audio signal frame; the Kalman gain is used to describe the trust priority between the observed energy state and the predicted energy state; and using the Kalman gain of the current audio signal frame as a fusion weight, performing weighted fusion processing on the observed energy state of the current audio signal frame and the predicted energy state of the current audio signal frame to obtain the target energy state of the current audio signal frame.
[0011] In one embodiment, obtaining the Kalman gain of the current audio signal frame includes: obtaining covariance update information of an audio signal frame previous to the current audio signal frame; the covariance update information is used to describe the uncertainty of a target energy state; using the covariance update information of the audio signal frame previous to the current audio signal frame to predict covariance prediction information of the current audio signal frame; and performing Kalman gain update processing using the covariance prediction information of the current audio signal frame to obtain the Kalman gain of the current audio signal frame.
[0012] In one embodiment, the covariance update information of the previous audio signal frame of the current audio signal frame is obtained by the following steps: performing covariance information update processing on the covariance prediction information of the previous audio signal frame of the current audio signal frame and the Kalman gain of the previous audio signal frame of the current audio signal frame to obtain the covariance update information of the current audio signal frame.
[0013] In one embodiment, the observation generation processing of the actual energy state of the current audio signal frame to obtain the observed energy state of the current audio signal frame includes: obtaining pre-constructed observation information and observation noise information; and inputting the observation information, observation noise information, and the actual energy state of the current audio signal frame into a pre-constructed observation model to obtain the observed energy state of the current audio signal frame.
[0014] In one embodiment, predicting the energy state of the current audio signal frame using the target energy state of the previous audio signal frame of the current audio signal frame to obtain the predicted energy state of the current audio signal frame includes: obtaining pre-constructed signal energy state transition information and process noise information; and inputting the signal energy state transition information, process noise information, and the target energy state of the previous audio signal frame of the current audio signal frame into a pre-constructed state transition model to obtain the predicted energy state of the current audio signal frame.
[0015] In one embodiment, the comprehensively processing the actual energy states of at least one audio signal frame among the multiple audio signal frames to obtain the target energy state of the current audio signal frame includes: obtaining multiple actual energy states that meet a preset condition from the actual energy states of the multiple audio signal frames; and taking an average of the multiple actual energy states that meet the preset condition as the target energy state of the current audio signal frame.
[0016] In one embodiment, obtaining the actual energy state of each audio signal frame in multiple audio signal frames included in the audio signal to be detected includes: obtaining the actual short-time signal energy and the actual high-frequency energy ratio of each audio signal frame in the multiple audio signal frames; the actual short-time signal energy is the sum of the signal energies of each frequency corresponding to the audio signal frame, and the actual high-frequency energy ratio is the ratio between the sum of the high-frequency signal energies of the audio signal frame and the actual short-time signal energy of the audio signal frame; and splicing the actual short-time signal energy and the actual high-frequency energy ratio of each of the audio signal frames to construct the actual energy state of each audio signal frame.
[0017] In one embodiment, comparing the target energy state of the current audio signal frame with a preset energy state threshold to obtain a microphone popping detection result of the current audio signal frame includes: obtaining a target short-time signal energy and a target high-frequency energy ratio of the current audio signal frame from the target energy state of the current audio signal frame; and determining that a microphone popping phenomenon occurs in the current audio signal frame when the target short-time signal energy exceeds a preset energy threshold and the target high-frequency energy ratio exceeds a preset ratio threshold.
[0018] In a second aspect, the present application further provides a device for detecting a wheat spray, comprising:
[0019] An actual energy acquisition module is configured to acquire an actual energy state of each audio signal frame in a plurality of audio signal frames included in the audio signal to be detected; wherein the actual energy state is an energy state obtained by performing energy detection on the audio signal frame;
[0020] a first target energy acquisition module, configured to, when a current audio signal frame is the first audio signal frame of the multiple audio signal frames, perform comprehensive processing on actual energy states of at least one audio signal frame among the multiple audio signal frames to obtain a target energy state of the current audio signal frame;
[0021] a second target energy acquisition module, configured to, when a current audio signal frame is any audio signal frame subsequent to a first audio signal frame of the multiple audio signal frames, predict an energy state of the current audio signal frame by using a target energy state of an audio signal frame previous to the current audio signal frame to obtain a predicted energy state of the current audio signal frame, and fuse an actual energy state of the current audio signal frame with the predicted energy state to obtain a target energy state of the current audio signal frame;
[0022] The signal microphone ejection detection module is configured to compare the target energy state of the current audio signal frame with a preset energy state threshold to obtain a microphone ejection detection result of the current audio signal frame.
[0023] In a third aspect, the present application further provides a computer device comprising a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of the method described in any one of the embodiments of the first aspect when executing the computer program.
[0024] In a fourth aspect, the present application further provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the method described in any one of the embodiments of the first aspect.
[0025] In a fifth aspect, the present application also provides a computer program product, comprising a computer program, which, when executed by a processor, implements the steps of the method described in any one of the embodiments of the first aspect.
[0026] The above-mentioned method, apparatus, computer equipment, storage medium and computer program product for microphone pop-up detection obtain the actual energy state of each audio signal frame of the audio signal to be detected, and when the current audio signal frame is the first frame, the actual energy state of at least one audio signal frame can be used for comprehensive processing to obtain the target energy state of the current audio signal frame. If the current audio signal frame is any frame after the first frame, the target energy state of the previous audio signal frame can be used to obtain the predicted energy state of each current audio signal frame, thereby fusing the actual energy state and the predicted energy state to obtain the target energy state of each audio signal frame, and then using the target energy state of each audio signal frame and the preset energy state threshold to perform microphone pop-up detection. In this way, the interference of environmental noise and audio fluctuations can be effectively overcome, thereby improving the accuracy of microphone pop-up detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following briefly introduces the drawings required for use in the embodiments or related technical descriptions. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0028] Figure 1 1 is a schematic flow chart of a wheat spray detection method according to an embodiment;
[0029] Figure 2 A schematic diagram of a process for obtaining target energy state information in one embodiment;
[0030] Figure 3 1 is a flow chart of obtaining the Kalman gain in one embodiment;
[0031] Figure 4 2 is a block diagram of a wheat spray detection device according to an embodiment;
[0032] Figure 5 FIG. 1 is a diagram showing the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION
[0033] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.
[0034] In one embodiment, Figure 1 As shown, a method for detecting a microphone being sprayed is provided. This embodiment uses the method applied to a terminal as an example. It is understood that the method can also be applied to a server, or to a system including a terminal and a server, and implemented through interaction between the terminal and the server. In this embodiment, the method includes the following steps:
[0035] Step S101 : obtaining an actual energy state of each audio signal frame in a plurality of audio signal frames included in an audio signal to be detected; wherein the actual energy state is an energy state obtained by performing energy detection on the audio signal frame.
[0036] The audio signal to be detected refers to the audio signal for microphone ejection detection. For example, it can be a recording signal captured by an audio capture device while a user sings. Multiple audio signal frames can be obtained by segmenting the audio signal to be detected. For example, multiple audio signal frames can be obtained by segmenting the audio signal to be detected with a frame length of 30 milliseconds and a frame shift of 15 milliseconds. The actual energy state refers to the signal energy state obtained by energy detection of the audio signal frame. The data format of the signal energy state can be a signal energy state vector.
[0037] Specifically, after collecting the recorded signal as the audio signal to be detected, the terminal may further perform frame processing on the audio signal to be detected to obtain multiple audio signal frames, and then obtain the actual energy state of each audio signal frame.
[0038] Step S102 : When the current audio signal frame is the first audio signal frame of multiple audio signal frames, perform comprehensive processing on the actual energy state of at least one audio signal frame among the multiple audio signal frames to obtain a target energy state of the current audio signal frame.
[0039] The current audio signal frame may refer to any one of multiple audio signal frames, and the target energy state may refer to an estimated value of the signal energy state of the current audio signal frame. The signal energy state may also be a signal energy state vector and may be used to implement microphone ejection detection. If the current audio signal frame is the first frame among multiple audio signal frames, the estimated value may be obtained by comprehensively processing the actual energy state of at least one of the multiple audio signal frames collected. For example, the actual energy state of the current audio signal frame may be directly used as the target energy state of the current audio signal frame, or the average or median of the actual energy states of several consecutive frames may be used as the target energy state of the current audio signal frame.
[0040] Specifically, if the current audio signal frame is the first audio signal frame of multiple audio signal frames, the terminal may perform comprehensive processing on the actual energy state of at least one audio signal frame of the multiple audio signal frames to obtain the target energy state of the current audio signal frame.
[0041] Step S103: When the current audio signal frame is any audio signal frame after the first audio signal frame of the multiple audio signal frames, the target energy state of the previous audio signal frame of the current audio signal frame is used to predict the energy state of the current audio signal frame to obtain the predicted energy state of the current audio signal frame, and the actual energy state of the current audio signal frame and the predicted energy state are fused to obtain the target energy state of the current audio signal frame.
[0042] The previous audio signal frame refers to the previous audio signal frame arranged before the current audio signal frame. If the current audio signal frame is any audio signal frame after the first frame, then there must be a previous audio signal frame. For example, if the current audio signal frame is the second frame among multiple audio signal frames, then the previous audio signal frame can be the first frame among multiple audio signal frames. Similarly, if the current audio signal frame is the third frame among multiple audio signal frames, then the previous audio signal frame can be the second frame among multiple audio signal frames. The predicted energy state refers to the signal energy state predicted based on the target energy state of the previous frame, and the prediction method can be implemented through the Kalman filter iteration method. Similar to the actual energy state, the predicted energy state can also be in the form of a signal energy state vector.
[0043] Specifically, if a current audio signal frame among multiple audio signal frames is not the first frame, the terminal can first collect the target energy state of the previous audio signal frame of the audio signal frame, and then use the target energy state of the previous audio signal frame to predict the signal energy state of the current audio signal frame as the predicted energy state of the current audio signal frame. Then, the predicted energy state of the current audio signal frame is combined with the actual energy state of the current audio signal frame to obtain the target energy state of the current audio signal frame. For example, the current audio signal frame is the second frame among multiple audio signal frames. The terminal can collect the target energy information of the first frame, and then use the target energy state of the first frame to obtain the predicted energy state of the second frame. Then, the actual energy state and the predicted energy state of the second frame can be fused to calculate the target energy state of the second frame. Similarly, the target energy information of each subsequent audio signal frame can be obtained.
[0044] Step S104: Compare the target energy state of the current audio signal frame with a preset energy state threshold to obtain a microphone ejection detection result of the current audio signal frame.
[0045] The preset energy state threshold is an energy state threshold set in advance for realizing microphone pop-up detection. For example, when the target energy state exceeds the threshold, it is determined that the current audio signal frame has a microphone pop-up phenomenon. Specifically, after obtaining the target energy state of each current audio signal frame, the target energy state can be compared with the preset energy state threshold to perform microphone pop-up detection, thereby obtaining the microphone pop-up detection result of each current audio signal frame.
[0046] In the above-mentioned method for detecting microphone pop-up, the actual energy state of each audio signal frame of the audio signal to be detected is obtained, and when the current audio signal frame is the first frame, the actual energy state of at least one audio signal frame can be used for comprehensive processing to obtain the target energy state of the current audio signal frame. If the current audio signal frame is any frame after the first frame, the target energy state of the previous audio signal frame can be used to obtain the predicted energy state of each current audio signal frame, thereby fusing the actual energy state and the predicted energy state to obtain the target energy state of each audio signal frame, and then using the target energy state of each audio signal frame and the preset energy state threshold to perform microphone pop-up detection. In this way, the interference of environmental noise and audio fluctuations can be effectively overcome, thereby improving the accuracy of microphone pop-up detection.
[0047] In one embodiment, Figure 2 As shown, step S103 may further include:
[0048] Step S201 : performing observation generation processing on the actual energy state of the current audio signal frame to obtain the observed energy state of the current audio signal frame.
[0049] The observed signal energy state information refers to the signal energy state observation value of the current audio signal frame, or it can be a signal energy state vector. This observation value can be used to characterize the system state-related data directly obtained by the measuring device. After obtaining the actual energy state of the current audio signal frame, the terminal can also use a pre-constructed observation model to observe and generate the actual energy state, thereby obtaining the observed energy state of the current audio signal frame.
[0050] Step S202: Acquire the Kalman gain of the current audio signal frame; the Kalman gain is used to describe the trust priority between the observed energy state and the predicted energy state;
[0051] Step S203 : Using the Kalman gain of the current audio signal frame as a fusion weight, weighted fusion processing is performed on the observed energy state of the current audio signal frame and the predicted energy state of the current audio signal frame to obtain a target energy state of the current audio signal frame.
[0052] Among them, Kalman gain refers to the weight ratio between the observed energy state and the predicted energy state when updating the state estimation. It is used to measure the weight distribution between the observed energy state and the predicted energy state. The Kalman gain reflects the trust between the observed energy state and the predicted energy state. The larger the value, the more trust is placed in the observed energy state; the smaller the value, the more trust is placed in the observed energy state.
[0053] Specifically, after obtaining the observed energy state, the terminal can also obtain the Kalman gain of the current audio signal frame, and then use the Kalman gain to perform weighted fusion on the observed energy state and the predicted energy state to calculate the target energy state of the current audio signal frame.
[0054] For example, the target energy state of the current audio signal frame, i.e., the kth frame It can be calculated by the following formula:
[0055]
[0056] in, represents the target energy state of the kth frame, represents the predicted energy state of the kth frame, represents the Kalman gain of the kth frame, It represents the observed energy state of the kth frame, and Represents the pre-constructed observation information, that is, the observation matrix, which can be:
[0057]
[0058] In this embodiment, the observed energy state can also be obtained by the actual energy state of the current audio signal frame, and the Kalman gain of the current audio signal frame is used to perform weighted fusion of the observed energy state and the predicted energy state to obtain the target energy state of the current audio signal frame. In this way, the accuracy of obtaining the target energy state can be improved.
[0059] Furthermore, if Figure 3 As shown, step S202 may further include:
[0060] Step S301: obtaining covariance update information of the previous audio signal frame of the current audio signal frame; the covariance update information is used to describe the uncertainty of the target energy state;
[0061] Step S302 : Using the covariance update information of the previous audio signal frame of the current audio signal frame, covariance prediction information of the current audio signal frame is predicted.
[0062] Among them, the covariance update information can be a covariance matrix, which can be used to describe the relationship between multidimensional random variables. Its elements represent the covariance between random variables, and thus can describe the uncertainty of the target energy state. The covariance prediction information refers to the covariance prediction value of the current audio signal frame obtained by predicting the covariance update information of the previous audio signal frame.
[0063] Specifically, the terminal may first collect covariance information of the previous audio signal frame, and then use the covariance information of the previous audio signal frame to predict the covariance of the current audio signal frame, that is, obtain covariance prediction information of the current audio signal frame.
[0064] For example, the covariance prediction information of the current audio signal frame, that is, the kth frame It can be calculated by the following formula:
[0065]
[0066] in, represents the covariance prediction information of the k-th frame, Represents the pre-constructed signal energy state transfer information, which can be a state transfer matrix. represents the covariance update information of the k-1th frame, represents the covariance matrix of the pre-constructed process noise. Among them, the state transfer matrix It can be:
[0067]
[0068] For relatively stable audio signals and Can be set between 0.9 and 1.1, and It can be set between -0.1 and 0.1 to describe the natural change of audio features from one moment to the next.
[0069] The covariance matrix of the process noise is It can be:
[0070]
[0071] According to prior experiments, and The initial setting is between 0.01 and 0.1. and It is set between -0.01 and 0.01.
[0072] Step S303 : performing Kalman gain update processing using the covariance prediction information of the current audio signal frame to obtain the Kalman gain of the current audio signal frame.
[0073] Finally, the terminal can use the obtained covariance prediction information of the current audio signal frame to implement the iterative update of the Kalman gain, thereby calculating the Kalman gain of the current audio signal frame, such as the Kalman gain of the kth frame. It can be calculated by the following formula:
[0074]
[0075] in, represents the Kalman gain of the kth frame, represents the covariance prediction information of the k-th frame, represents the pre-constructed observation information, namely the observation matrix, and Denotes the covariance matrix of the observation noise. And the covariance matrix of the observation noise can be:
[0076]
[0077] and It can be initially set between 0.05 and 0.2 based on the error estimation of feature extraction. and Settings can be between -0.05 and 0.05.
[0078] In this embodiment, the covariance prediction information of the current audio signal frame can be obtained by using the covariance update information of the previous audio signal frame, and the Kalman gain can be iteratively updated using the covariance prediction information. In this way, the accuracy of the Kalman gain calculation can be improved.
[0079] In addition, the covariance update information of the previous audio signal frame of the current audio signal frame is obtained by the following steps: updating the covariance information of the covariance prediction information of the previous audio signal frame of the current audio signal frame and the Kalman gain of the previous audio signal frame of the current audio signal frame to obtain the covariance update information of the current audio signal frame.
[0080] The covariance update information of the previous audio signal frame of the current audio signal frame can be achieved through the following steps, namely, first collecting the covariance prediction information of the previous audio signal frame of the current audio signal frame and the Kalman gain of the previous audio signal frame, and then using the covariance prediction information and Kalman gain of the previous audio signal frame to obtain the covariance update information of the current audio signal frame.
[0081] Similarly, after obtaining the Kalman gain of the current audio signal frame, the Kalman gain can also be used to update the covariance information of the current audio signal frame, that is, the covariance prediction information of the current audio signal frame and the Kalman gain are used to obtain the covariance update information of the current audio signal frame, so as to obtain the covariance prediction information of the next audio signal frame of the current audio signal frame through the covariance update information of the current audio signal frame, thereby obtaining the Kalman gain of the next audio signal frame, and realizing the iterative update of the Kalman gain and the covariance information.
[0082] For example, if the current audio signal frame is the second frame among multiple audio signal frames, the terminal can collect the covariance update information of the previous audio signal frame of the current audio signal frame, that is, collect the covariance update information of the first frame, and then use the covariance update information of the first frame to obtain the covariance prediction information of the second frame, and further calculate the Kalman gain of the second frame. In addition, the Kalman gain of the second frame can be used to obtain the covariance update information of the second frame, and the covariance prediction information and Kalman gain of the third frame can be obtained by iterating in this way. The Kalman gain of each audio signal frame can be obtained.
[0083] Moreover, the covariance information of the current audio signal frame, i.e. the kth frame It can be calculated by the following formula:
[0084]
[0085] in, represents the covariance information of the kth frame, represents the identity matrix, represents the Kalman gain of the kth frame, represents the pre-constructed observation information, namely the observation matrix, and Represents the covariance prediction information of the k-th frame.
[0086] In this embodiment, the covariance prediction information and Kalman gain of the previous audio signal frame can also be used to obtain the covariance update information of the current audio signal frame, thereby achieving iterative update of the covariance information. In this way, the accuracy of the covariance information can be improved.
[0087] In addition, step S201 may further include: obtaining pre-constructed observation information and observation noise information; inputting the observation information, observation noise information and the actual energy state of the current audio signal frame into the pre-constructed observation model to obtain the observation energy state of the current audio signal frame.
[0088] The pre-constructed observation information may refer to a pre-constructed observation matrix , the observation noise information can be represented by zero-mean Gaussian noise. In this embodiment, the observed energy state can be obtained by processing the actual energy state through an observation model. The observation model can be a certain observation equation, and the observation equation can include a pre-constructed observation matrix and observation noise.
[0089] For example, the observed energy state of the current audio signal frame, i.e., the kth frame It can be represented by the following observation equation:
[0090]
[0091] in, represents the observed energy state of the kth frame, Represents the pre-constructed observation information, that is, the observation matrix, represents the actual energy state of the kth frame, and Represents the observation noise information.
[0092] In this embodiment, the observed energy state of the current audio signal frame may also be obtained through the observation model, and this approach may improve the accuracy of obtaining the observed energy state.
[0093] In one embodiment, step S103 may further include: obtaining pre-constructed signal energy state transition information and process noise information; inputting the signal energy state transition information, process noise information, and the target energy state of the previous audio signal frame of the current audio signal frame into a pre-constructed state transition model to obtain a predicted energy state of the current audio signal frame.
[0094] The pre-constructed signal energy state transfer information may refer to a pre-constructed state transfer matrix , the process noise information can be represented by zero-mean Gaussian noise. In this embodiment, the predicted energy state can be obtained by processing the target energy state of the previous audio signal frame through a state transition model. The state transition model can be a certain state variance, and the state equation can include a pre-constructed state transition matrix and process noise.
[0095] For example, the predicted energy state of the current audio signal frame, i.e., the kth frame It can be calculated by the following formula:
[0096]
[0097] in, represents the predicted energy state of the kth frame, Represents the pre-constructed signal energy state transfer information, that is, the state transfer matrix, represents the target energy state of the k-1th frame, and Represents process noise information.
[0098] In this embodiment, the predicted energy state of the current audio signal frame may also be obtained through a state equation, and this approach may improve the accuracy of obtaining the predicted energy state.
[0099] In one embodiment, step S102 may further include: acquiring multiple actual energy states that meet a preset condition from the actual energy states of the multiple audio signal frames; and taking an average value of the multiple actual energy states that meet the preset condition as the target energy state of the current audio signal frame.
[0100] However, if the current audio signal frame is the first frame among multiple audio signal frames, since it is impossible to use the target energy state information of the previous audio signal frame to obtain the predicted energy state of the current audio signal frame, and then obtain the target energy state of the current audio signal frame, an initial energy state can be set and used as the target energy state of the first frame.
[0101] The set initial energy state information can be obtained by satisfying preset conditions, such as the average value of actual energy state information corresponding to the previous preset number of audio signal frames. The preset number is 5, that is, the average value of the actual energy states of the previous 5 audio signal frames is used as the target energy state of the first frame.
[0102] Specifically, if the current audio signal frame is the first frame, the terminal can first extract the actual energy state of the audio signal frame that meets the preset conditions, for example, obtain the actual energy state of the first 5 frames, and calculate the average value of the actual energy state of the first 5 frames, and then use the average value as the target energy state of the first frame.
[0103] In this embodiment, if the current audio signal frame is the first frame, the average value of the actual energy states corresponding to the audio signal frames that meet the preset conditions can also be used as the target energy state of the current audio signal frame. In this way, the accuracy of obtaining the target energy state of the first frame can be improved.
[0104] In addition, step S101 may further include: obtaining the actual short-time signal energy and the actual high-frequency energy ratio of each audio signal frame in the multiple audio signal frames; the actual short-time signal energy is the sum of the signal energies of each frequency corresponding to the audio signal frame, and the actual high-frequency energy ratio is the ratio between the sum of the high-frequency signal energies of the audio signal frame and the actual short-time signal energy of the audio signal frame; splicing the actual short-time signal energy and the actual high-frequency energy ratio of each of the audio signal frames to construct the actual energy state of each of the audio signal frames.
[0105] The actual short-time signal energy may refer to the short-time signal energy actually detected in each audio signal frame, that is, the sum of the signal energies of each frequency corresponding to the audio signal frame, while the actual high-frequency energy ratio refers to the high-frequency energy ratio actually detected in each audio signal frame, which refers to the ratio between the sum of the high-frequency signal energy of the audio signal frame and the actual short-time signal energy of the audio signal frame.
[0106] In this embodiment, the actual energy state of each audio signal frame may include the actual short-time signal energy and the actual high-frequency energy ratio. Similarly, the predicted energy state and the target energy state also have two components, namely the short-time signal energy and the high-frequency energy ratio, and the same type of components can be processed separately, that is, the predicted energy state includes the predicted short-time signal energy and the predicted high-frequency energy ratio, and the target energy state includes the target short-time signal energy and the target high-frequency energy ratio. The target short-time signal energy is calculated based on the actual short-time signal energy and the predicted short-time signal energy, and the target high-frequency energy ratio is calculated based on the actual high-frequency energy ratio and the predicted high-frequency energy ratio.
[0107] For example, the actual energy state of the current audio signal frame, that is, the kth frame It can be expressed as:
[0108]
[0109] in, represents the actual short-time signal energy of the kth frame, and Indicates the actual high-frequency energy ratio of the k-th frame.
[0110] In addition, the short-time signal energy of each audio signal frame can be calculated using the following formula:
[0111]
[0112] in, is the sample sequence of the nth audio signal frame, L is the number of sample points corresponding to the frame length, and the energy intensity of the audio signal in the frame can be measured by calculating the sum of the squares of the sample values in the frame.
[0113] At the same time, the energy proportion of the high-frequency part (2kHz - 8kHz) of each frame of audio can be calculated. First, the audio signal is converted from the time domain to the frequency domain through the fast Fourier transform (FFT) to obtain the spectrum. , let the total energy of the high frequency part be:
[0114]
[0115] and Corresponding to the indexes of 2kHz and 8kHz in the frequency domain, the proportion of high-frequency energy in the frame is:
[0116]
[0117] This calculation analyzes the energy distribution in the high-frequency portion of an audio signal.
[0118] In this embodiment, the actual short-time signal energy and the actual high-frequency energy ratio of each audio signal frame can also be obtained, and the actual energy state of each audio signal frame can be constructed by using the actual short-time signal energy and the actual high-frequency energy ratio. In this way, the completeness of the actual energy state can be improved.
[0119] Furthermore, step S104 may further include: obtaining the target short-time signal energy and target high-frequency energy ratio of the current audio signal frame from the target energy state of the current audio signal frame; when the target short-time signal energy exceeds a preset energy threshold and the target high-frequency energy ratio exceeds a preset ratio threshold, determining that the current audio signal frame has a microphone popping phenomenon.
[0120] Since the actual energy state can be composed of the actual short-time signal energy and the actual high-frequency energy ratio, the target energy state generated based on the actual energy state can also be composed of the target short-time signal energy and the target high-frequency energy ratio, where the target short-time signal energy represents the short-time signal energy in the target energy state, and the target high-frequency energy ratio is the high-frequency energy ratio in the target energy state.
[0121] Specifically, after obtaining the target energy state of the current audio signal frame, the terminal can also extract the target short-term signal energy and target high-frequency energy ratio from the target energy state information. It can then determine whether the target short-term signal energy exceeds a preset energy threshold and whether the target high-frequency energy ratio exceeds a preset ratio threshold. If both exceed the threshold, the terminal can determine that the current audio signal frame has experienced microphone popping.
[0122] For example, the target energy state information of the current audio signal frame, i.e., the kth frame It can be expressed as:
[0123]
[0124] in, represents the target short-time signal energy of the kth frame, and Indicates the target high-frequency energy ratio of the k-th frame.
[0125] Finally, if the short-term energy Exceeding the set energy threshold (Assuming 10 to 20 times the average short-term energy of normal audio) and the proportion of high-frequency energy Above the set ratio threshold (between 0.3 - 0.6), it is determined that wheat spraying occurs.
[0126] In this embodiment, after obtaining the target energy state of each audio signal frame, the target short-time signal energy and target high-frequency energy ratio of each audio signal frame can also be extracted therefrom. Only when the target short-time signal energy exceeds the preset energy threshold and the target high-frequency energy ratio exceeds the preset ratio threshold, it is determined that the current audio signal frame has a microphone popping phenomenon. In this way, the microphone popping phenomenon can be comprehensively judged in combination with the short-time energy and high-frequency energy ratio of the audio, effectively reducing misjudgment, and thus the accuracy of microphone popping detection can be further improved.
[0127] In one embodiment, a Kalman filter-based method for detecting microphone popping is provided. This method utilizes Kalman filtering technology to frame the audio signal and extract short-term energy and high-frequency energy ratio features. This model, comprised of specific state and observation equations, is then constructed. State estimates are then iteratively updated to accurately determine microphone popping. Compared to traditional methods based on simple energy thresholds, this method effectively overcomes interference from ambient noise and audio fluctuations, improving detection accuracy and ensuring audio quality. This method is crucial for subsequent processing such as speech recognition and audio editing.
[0128] This embodiment can be applied to various audio recording devices and software, such as professional recording equipment, smartphone recording functions, audio editing software, etc. During use, the user performs audio recording operations as usual. After the recording is completed, the system automatically runs the microphone spray detection program based on Kalman filtering. If the microphone spray phenomenon is detected, the time point or time period when the microphone spray occurs can be displayed to the user in the form of a mark or prompt information on the software interface, which is convenient for the user to quickly locate and deal with the problem and improve the audio recording quality. For example, on the timeline of the audio editing software, the microphone spray part is marked with a line segment of a special color, and repair suggestions or tool links are provided. This can be achieved specifically through the following process:
[0129] The front end is responsible for collecting and preprocessing audio data and transmitting the audio signal to the back end in real time. During the collection process, digital conversion is performed according to the set sampling rate (such as 44.1kHz) and quantization accuracy (such as 16 bits).
[0130] The backend first frames the input audio signal, setting the frame length to 30 milliseconds and the frame shift to 15 milliseconds. This step is to divide the continuous audio stream into small segments, so that the characteristics of each segment can be analyzed later. Then, the short-term energy of each frame is calculated using the following formula:
[0131]
[0132] in, is the sample sequence of the nth audio signal frame, L is the number of sample points corresponding to the frame length, and by calculating the sum of the squares of the sample values within the frame, the energy intensity of the audio signal in that frame can be measured. At the same time, the energy proportion of the high-frequency part (2kHz - 8kHz) of each frame of audio can also be calculated. First, the audio signal is converted from the time domain to the frequency domain through the fast Fourier transform (FFT) to obtain the spectrum , let the total energy of the high frequency part be:
[0133]
[0134] and Corresponding to the indexes of 2kHz and 8kHz in the frequency domain, the proportion of high-frequency energy in the frame is:
[0135]
[0136] This calculation can analyze the energy distribution of the high-frequency part of the audio signal, because the proportion of high-frequency energy often changes abnormally when the microphone is playing.
[0137] Then build the Kalman filter model, the state vector Contains short-term energy and high-frequency energy ratio :
[0138]
[0139] Equation of state , where the state transfer matrix is:
[0140]
[0141] For relatively stable audio signals and Can be set between 0.9 and 1.1, and It can be set between -0.1 and 0.1 to describe the natural change of audio features from one moment to the next. is the process noise, assuming it is zero-mean Gaussian noise, the covariance matrix is:
[0142]
[0143] According to prior experiments, and The initial setting is between 0.01 and 0.1. and It is set between -0.01 and 0.01.
[0144] Observation equation , where the observation matrix is:
[0145]
[0146] is the observation noise, also assumed to be zero-mean Gaussian noise, and the covariance matrix is:
[0147]
[0148] and It can be initially set between 0.05 and 0.2 based on the error estimation of feature extraction. and Settings can be between -0.05 and 0.05.
[0149] During initialization, the average value of the first 5 to 10 frames of audio features is taken as the initial state estimate , the initial covariance matrix:
[0150]
[0151] in According to the initial fluctuation range of short-time energy set between 0.1 and 1, The initial fluctuation range is set between 0.01 and 0.1 according to the high-frequency energy ratio.
[0152] In the Kalman filter iteration, the prediction step is performed first, and the state prediction value is calculated according to the state equation:
[0153]
[0154] And the covariance predicted value:
[0155]
[0156] Then the update step is performed to calculate the Kalman gain:
[0157]
[0158] Then update the state estimate:
[0159]
[0160] And the covariance estimate:
[0161]
[0162] Finally, if the short-term energy Exceeding the set energy threshold (set to 10 to 20 times the average short-term energy of normal audio) and the proportion of high-frequency energy Above the set ratio threshold (between 0.3 and 0.6), it is determined that wheat spraying occurs.
[0163] It should be understood that, although the various steps in the flowcharts involved in the various embodiments described above are displayed in sequence according to the instructions of the arrows, these steps are not necessarily executed in sequence in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and these steps can be executed in other orders. Moreover, at least a portion of the steps in the flowcharts involved in the various embodiments described above can include multiple steps or multiple stages, and these steps or stages are not necessarily executed and completed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a portion of steps or stages in other steps.
[0164] Based on the same inventive concept, embodiments of the present application also provide a wheat spray detection device for implementing the aforementioned wheat spray detection method. The solution provided by this device is similar to the solution described in the aforementioned method. Therefore, the specific limitations of one or more wheat spray detection device embodiments provided below can be found in the aforementioned limitations of the wheat spray detection method and will not be further elaborated here.
[0165] In one embodiment, Figure 4 As shown, a microphone spray detection device is provided, including: an actual energy acquisition module 401, a first target energy acquisition module 402, a second target energy acquisition module 403 and a signal microphone spray detection module 404, wherein:
[0166] The actual energy acquisition module 401 is configured to acquire the actual energy state of each audio signal frame in a plurality of audio signal frames included in the audio signal to be detected; wherein the actual energy state is the energy state obtained by performing energy detection on the audio signal frame;
[0167] a first target energy acquisition module 402 configured to, when the current audio signal frame is the first audio signal frame of the multiple audio signal frames, comprehensively process the actual energy states of at least one audio signal frame among the multiple audio signal frames to obtain a target energy state of the current audio signal frame;
[0168] a second target energy acquisition module 403 for, when the current audio signal frame is any audio signal frame subsequent to the first audio signal frame of the multiple audio signal frames, predicting the energy state of the current audio signal frame using the target energy state of the previous audio signal frame of the current audio signal frame to obtain the predicted energy state of the current audio signal frame, and fusing the actual energy state of the current audio signal frame with the predicted energy state to obtain the target energy state of the current audio signal frame;
[0169] The signal microphone pop-up detection module 404 is configured to compare the target energy state of the current audio signal frame with a preset energy state threshold to obtain a microphone pop-up detection result of the current audio signal frame.
[0170] Each module in the aforementioned wheat spray detection device can be implemented in whole or in part through software, hardware, or a combination thereof. Each module can be embedded in or independent of a processor in a computer device in hardware form, or can be stored in a computer device memory in software form, so that the processor can call and execute the corresponding operations of each module.
[0171] In an exemplary embodiment, a computer device is provided. The computer device may be a terminal, and its internal structure diagram may be as shown in FIG. Figure 5 As shown. The computer device includes a processor, memory, an input / output interface, a communication interface, a display unit, and an input device. The processor, memory, and input / output interface are connected via a system bus, and the communication interface, display unit, and input device are connected to the system bus via the input / output interface. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores an operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with external terminals via wired or wireless means, and the wireless means can be implemented via Wi-Fi, mobile cellular networks, NFC (near-field communication), or other technologies. When executed by the processor, the computer program implements a method for detecting a microphone spray. The display unit of the computer device is used to form a visually visible image and can be a display screen, a projection device, or a virtual reality imaging device. The display screen can be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device can be a touch layer covering the display screen, or a button, trackball or touchpad set on the computer device casing, or an external keyboard, touchpad or mouse.
[0172] Those skilled in the art will understand that Figure 5The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.
[0173] In one embodiment, a computer device is further provided, including a memory and a processor. The memory stores a computer program, and the processor implements the steps in the above method embodiments when executing the computer program.
[0174] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps in the above-mentioned method embodiments are implemented.
[0175] In one embodiment, a computer program product is provided, including a computer program, which implements the steps in the above method embodiments when executed by a processor.
[0176] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant regulations.
[0177] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the above-mentioned embodiments. In particular, any reference to memory, database, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The databases involved in the various embodiments provided herein may include at least one of a relational database and a non-relational database. Non-relational databases may include, but are not limited to, distributed databases based on blockchains. The processors involved in the various embodiments provided herein may be, but are not limited to, general-purpose processors, central processing units (CPUs), graphics processing units (GPUs), digital signal processors (DSPs), programmable logic devices (PLDs), data processing logic devices based on quantum computing, and the like.
[0178] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0179] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present application. It should be noted that a person of ordinary skill in the art may make various modifications and improvements without departing from the spirit of the present application, and these modifications and improvements fall within the scope of protection of the present application. Therefore, the scope of protection of the present application shall be determined by the appended claims.
Claims
1. A method for detecting wheat spraying, characterized in that: The method comprises: Obtaining an actual energy state of each audio signal frame among a plurality of audio signal frames included in the audio signal to be detected; wherein the actual energy state is an energy state obtained by performing energy detection on the audio signal frame; When the current audio signal frame is the first audio signal frame of the multiple audio signal frames, performing comprehensive processing on actual energy states of at least one audio signal frame among the multiple audio signal frames to obtain a target energy state of the current audio signal frame; When the current audio signal frame is any audio signal frame after the first audio signal frame of the multiple audio signal frames, predicting the energy state of the current audio signal frame by using the target energy state of the previous audio signal frame to obtain the predicted energy state of the current audio signal frame, and fusing the actual energy state of the current audio signal frame with the predicted energy state to obtain the target energy state of the current audio signal frame; The target energy state of the current audio signal frame is compared with a preset energy state threshold to obtain a microphone ejection detection result of the current audio signal frame.
2. The method according to claim 1, characterized in that The fusing the actual energy state of the current audio signal frame and the predicted energy state to obtain the target energy state of the current audio signal frame includes: performing observation generation processing on the actual energy state of the current audio signal frame to obtain an observed energy state of the current audio signal frame; Obtaining a Kalman gain of the current audio signal frame; the Kalman gain is used to describe the trust priority between the observed energy state and the predicted energy state; The Kalman gain of the current audio signal frame is used as a fusion weight to perform weighted fusion processing on the observed energy state of the current audio signal frame and the predicted energy state of the current audio signal frame to obtain a target energy state of the current audio signal frame.
3. The method according to claim 2, characterized in that The obtaining of the Kalman gain of the current audio signal frame includes: Obtaining covariance update information of an audio signal frame previous to the current audio signal frame, wherein the covariance update information is used to describe uncertainty of a target energy state; Predicting covariance prediction information of the current audio signal frame by using covariance update information of an audio signal frame previous to the current audio signal frame; Kalman gain updating processing is performed using the covariance prediction information of the current audio signal frame to obtain the Kalman gain of the current audio signal frame.
4. The method according to claim 3, characterized in that The covariance update information of the previous audio signal frame of the current audio signal frame is obtained by the following steps: Performing covariance information updating processing on the covariance prediction information of the previous audio signal frame of the current audio signal frame and the Kalman gain of the previous audio signal frame of the current audio signal frame to obtain covariance update information of the current audio signal frame.
5. The method according to claim 2, characterized in that The performing observation generation processing on the actual energy state of the current audio signal frame to obtain the observed energy state of the current audio signal frame includes: Obtaining pre-constructed observation information and observation noise information; The observation information, the observation noise information, and the actual energy state of the current audio signal frame are input into a pre-constructed observation model to obtain the observation energy state of the current audio signal frame.
6. The method according to claim 1, characterized in that The predicting the energy state of the current audio signal frame by using the target energy state of the previous audio signal frame of the current audio signal frame to obtain the predicted energy state of the current audio signal frame includes: Obtain pre-constructed signal energy state transfer information and process noise information; The signal energy state transition information, process noise information, and the target energy state of the previous audio signal frame of the current audio signal frame are input into a pre-constructed state transition model to obtain a predicted energy state of the current audio signal frame.
7. The method according to claim 1, characterized in that The comprehensively processing the actual energy state of at least one audio signal frame among the multiple audio signal frames to obtain the target energy state of the current audio signal frame includes: Acquire, from the actual energy states of the plurality of audio signal frames, a plurality of actual energy states that meet a preset condition; An average value of the multiple actual energy states that meet the preset conditions is used as the target energy state of the current audio signal frame.
8. The method according to any one of claims 1 to 7, characterized in that The obtaining of the actual energy state of each audio signal frame in the plurality of audio signal frames included in the audio signal to be detected includes: Obtaining an actual short-time signal energy and an actual high-frequency energy ratio of each audio signal frame in the multiple audio signal frames; the actual short-time signal energy is the sum of signal energies of each frequency corresponding to the audio signal frame, and the actual high-frequency energy ratio is the ratio between the sum of the high-frequency signal energy of the audio signal frame and the actual short-time signal energy of the audio signal frame; The actual short-time signal energy and the actual high-frequency energy proportion of each audio signal frame are spliced to construct the actual energy state of each audio signal frame.
9. The method according to claim 8, characterized in that The comparing the target energy state of the current audio signal frame with a preset energy state threshold to obtain a microphone ejection detection result of the current audio signal frame includes: acquiring, from the target energy state of the current audio signal frame, a target short-time signal energy and a target high-frequency energy ratio of the current audio signal frame; When the target short-time signal energy exceeds a preset energy threshold and the target high-frequency energy ratio exceeds a preset ratio threshold, it is determined that the current audio signal frame has a microphone popping phenomenon.
10. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 9 are implemented.
11. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 9 are implemented.
12. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 9 are implemented.