Game team voice noise reduction method, device and medium based on dynamic threshold mechanism
Through the speech noise reduction method of dynamic threshold mechanism, the problem of inaccurate noise filtering under different ambient noise conditions is solved, and efficient noise filtering and voice fidelity in game teaming voice scenes is realized, improving the user experience.
Patent Information
- Application Number
- CN202510480623.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-17
- Publication Date
- 2025-08-22
- Estimated Expiration
- 2045-04-17
AI Technical Summary
The prior art is difficult to dynamically and accurately filter noise under different ambient noise conditions, resulting in a decrease in speech clarity and poor user experience.
The speech noise reduction method based on the dynamic threshold mechanism is adopted. By collecting and dividing audio signal frames, the combination of initial threshold and dynamic threshold is used, combining environmental noise reference values and adjustment factors, the threshold is adjusted in real time to filter noise, including multi-stage filtering, frequency domain analysis and machine learning-assisted threshold adjustment.
Accurately filter noise in different noise environments, improve the clarity and stability of voice signals, and improve user experience.
Smart Images

Figure CN120032662B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of voice noise reduction, and in particular to a method, device and medium for game team voice noise reduction based on a dynamic threshold mechanism. Background Art
[0002] With the rapid development of online games and esports, more and more players need to communicate and collaborate in real time through voice. However, in actual use, players often encounter various noise interferences such as mechanical noise, human voices, and wind noise, which reduces the clarity of players' voices and communication efficiency. To improve the quality of team voice chat, effective voice noise reduction is necessary for different noise scenarios.
[0003] Current speech noise reduction technologies typically rely on fixed thresholds to make preliminary judgments on audio data, or simply perform simple noise suppression using basic algorithms at the hardware level. However, these fixed thresholds or single algorithms struggle to balance noise filtering and speech fidelity when noise levels fluctuate significantly, and are prone to missed or misjudgment. This is especially true in noisy internet cafes or outdoor environments, where players' voices are often drowned out by the noise or over-filtered, resulting in a poor user experience. Therefore, how to dynamically and accurately filter noise under varying ambient noise conditions has become a pressing technical challenge that existing technologies must address. Summary of the Invention
[0004] In view of this, the embodiments of the present application provide a game team voice noise reduction method, device and medium based on a dynamic threshold mechanism to solve the problem that the existing technology is prone to missed judgment or misjudgment, and is unable to dynamically and accurately filter noise under different environmental noise conditions, resulting in poor user experience.
[0005] In a first aspect of an embodiment of the present application, a method for reducing game team voice noise based on a dynamic threshold mechanism is provided, comprising: collecting an audio signal containing player voice and ambient noise, and dividing the audio signal into multiple time frames; converting the amplitude or decibel value of the time frame to obtain an energy value corresponding to each time frame; comparing the energy value with a set initial threshold, and if the energy value is lower than the initial threshold, marking it as a noise frame or a low-priority frame; if the energy value is higher than the initial threshold, retaining it as a voice frame for further processing; when it is detected that the continuous time frames with energy values lower than the initial threshold exceed a preset time period, performing statistics and smoothing on the energy values within the preset time period to obtain an ambient noise reference value; calculating a dynamic threshold based on the ambient noise reference value and a preset adjustment factor, and updating the dynamic threshold in real time when the ambient noise reference value changes over time; comparing the energy value with the dynamic threshold, and if the energy value is lower than the dynamic threshold, filtering the corresponding time frame as a noise frame; and if the energy value is higher than the dynamic threshold, retaining the corresponding time frame as a valid voice frame; encoding the time frames retained after filtering by the dynamic threshold to generate a final game team voice signal.
[0006] The second aspect of the embodiment of the present application provides a game team voice noise reduction device based on a dynamic threshold mechanism, including: an acquisition module for acquiring an audio signal containing player voice and environmental noise, and dividing the audio signal into multiple time frames; a conversion module for converting the amplitude or decibel value of the time frame to obtain an energy value corresponding to each time frame; a first comparison module for comparing the energy value with a set initial threshold value, if the energy value is lower than the initial threshold value, it is marked as a noise frame or a low priority frame, if the energy value is higher than the initial threshold value, it is retained as a voice frame to be further processed; a processing module for detecting that the energy value is lower than the initial threshold value for a continuous period of time. When a frame exceeds a preset time period, the energy value within the preset time period is statistically analyzed and smoothed to obtain an ambient noise reference value; a calculation module is used to calculate a dynamic threshold based on the ambient noise reference value and a preset adjustment factor, and to update the dynamic threshold in real time as the ambient noise reference value changes over time; a second comparison module is used to compare the energy value with the dynamic threshold; if the energy value is lower than the dynamic threshold, the corresponding time frame is filtered as a noise frame; if the energy value is higher than the dynamic threshold, the corresponding time frame is retained as a valid voice frame; a generation module is used to encode the time frame retained after filtering through the dynamic threshold to generate the final game team voice signal.
[0007] According to a third aspect of an embodiment of the present application, an electronic device is provided, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the steps of the above method are implemented when the processor executes the computer program.
[0008] According to a fourth aspect of the embodiments of the present application, a computer-readable storage medium is provided, which stores a computer program. When the computer program is executed by a processor, the steps of the above method are implemented.
[0009] At least one of the above technical solutions adopted in the embodiments of the present application can achieve the following beneficial effects:
[0010] By collecting audio signals containing player voices and environmental noise, and dividing the audio signals into multiple time frames; converting the amplitude or decibel value of the time frame to obtain the energy value corresponding to each time frame; comparing the energy value with the set initial threshold, if the energy value is lower than the initial threshold, it is marked as a noise frame or a low-priority frame, if the energy value is higher than the initial threshold, it is retained as a voice frame for further processing; when it is detected that the continuous time frames with energy values lower than the initial threshold exceed the preset time period, the energy values within the preset time period are statistically analyzed and smoothed to obtain the environmental noise reference value; based on the environmental noise reference value and the preset adjustment factor, the dynamic threshold is calculated, and when the environmental noise reference value changes over time, the dynamic threshold is updated in real time; comparing the energy value with the dynamic threshold, if the energy value is lower than the dynamic threshold, the corresponding time frame is filtered as a noise frame, if the energy value is higher than the dynamic threshold, the corresponding time frame is retained as a valid voice frame; encoding the time frame retained after filtering by the dynamic threshold to generate the final game team voice signal. This application can reduce the phenomenon of missed or misjudgment of noise, realize dynamic and accurate noise filtering under different environmental noise conditions, and improve user experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the embodiments or descriptions of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0012] Figure 1 1 is a flow chart of a method for reducing game team voice noise based on a dynamic threshold mechanism provided by an embodiment of the present application;
[0013] Figure 2 This is a schematic diagram of the structure of a game team voice noise reduction device based on a dynamic threshold mechanism provided by an embodiment of the present application;
[0014] Figure 3 It is a structural diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0015] In the following description, specific details such as specific system structures and techniques are provided for purposes of illustration rather than limitation to facilitate a thorough understanding of the embodiments of the present application. However, it will be apparent to those skilled in the art that the present application may be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid obscuring the description of the present application with unnecessary detail.
[0016] The technical problem to be solved by this application is how to effectively reduce voice noise in game team voice scenarios, especially to cope with the constantly changing background noise in different environments, so as to ensure the clarity and comfort of players' voice calls.
[0017] To this end, this application provides a method for reducing game team voice noise based on a dynamic threshold mechanism. The main technical implementation ideas of this solution include the following two points:
[0018] First, an algorithm converts the collected sound decibel information into energy values. Based on a set threshold, it filters out sound signals with lower energy values (corresponding to low noise or invalid speech). This operation can initially filter out background noise and reduce redundancy in subsequent processing.
[0019] Second, to address the difficulty of adapting fixed thresholds to different environments, a dynamic threshold mechanism has been introduced. This algorithm detects changes in ambient noise in real time and dynamically adjusts the threshold, automatically setting the appropriate threshold in both quiet and noisy environments. This improves the accuracy and adaptability of noise filtering, further enhancing the quality of player team voice chat.
[0020] In addition, in order to further improve the adaptability of the solution to noise under different environmental noise conditions, this application introduces a number of improved technologies on the basis of the above-mentioned main technical implementation ideas, including multi-stage filtering, frequency domain analysis, machine learning assisted threshold adjustment, etc., thus forming a more complete intelligent noise filtering solution.
[0021] The contents of the technical solution of this application are described in detail below with reference to the accompanying drawings and specific embodiments.
[0022] Figure 1 This is a flow chart of the method for reducing the noise of game team voice based on the dynamic threshold mechanism provided by the embodiment of the present application. Figure 1 As shown, the game team voice noise reduction method based on the dynamic threshold mechanism may specifically include:
[0023] S101, collecting an audio signal including a player's voice and environmental noise, and dividing the audio signal into multiple time frames;
[0024] S102, converting the amplitude or decibel value of the time frame to obtain the energy value corresponding to each time frame;
[0025] S103, comparing the energy value with a set initial threshold. If the energy value is lower than the initial threshold, the frame is marked as a noise frame or a low-priority frame. If the energy value is higher than the initial threshold, the frame is retained as a speech frame for further processing.
[0026] S104, when it is detected that the continuous time frames in which the energy value is lower than the initial threshold value exceed the preset time period, the energy value in the preset time period is statistically analyzed and smoothed to obtain an environmental noise reference value;
[0027] S105, calculating a dynamic threshold based on the ambient noise reference value and a preset adjustment factor, and updating the dynamic threshold in real time when the ambient noise reference value changes over time;
[0028] S106, comparing the energy value with a dynamic threshold. If the energy value is lower than the dynamic threshold, the corresponding time frame is filtered as a noise frame. If the energy value is higher than the dynamic threshold, the corresponding time frame is retained as a valid speech frame.
[0029] S107 , encoding the time frames retained after filtering by the dynamic threshold to generate a final game team voice signal.
[0030] In some embodiments, an audio signal including a player's voice and ambient noise is collected and divided into a plurality of time frames, including:
[0031] When the player uses a single microphone, the single-channel audio data is directly obtained as the analog voice signal;
[0032] When players use headsets or external multi-microphone devices, multiple channels of audio data are acquired in parallel and signal fusion or beamforming is performed on the multiple channels of audio data to obtain analog voice signals;
[0033] Convert the collected analog voice signal into a digital signal and quantize the digital signal to obtain digital audio that conforms to a predetermined audio coding format;
[0034] At the hardware level or in the audio driver, use the automatic gain control or acoustic echo cancellation module to perform preliminary noise reduction or gain processing on the digital audio;
[0035] The digital audio after preliminary noise reduction or gain processing is divided into multiple time frames according to a preset time window, so as to perform initial threshold comparison and dynamic threshold update according to the energy value of the time frame.
[0036] Specifically, in some cases, players use a single-microphone device, such as an external USB microphone or a built-in microphone. In this case, the system directly captures the analog voice signal from a single channel and converts it into a digital signal using an analog-to-digital converter (ADC) module. The sampling rate of the digital signal can be 16kHz or 48kHz to meet the quality requirements of real-time voice communication. During the conversion process, the system quantizes the signal to conform to the PCM (Pulse Code Modulation) format or other predetermined audio coding format for subsequent signal processing.
[0037] In another scenario, players use headsets or external multi-microphone devices, such as high-end gaming headsets or independent multi-microphone arrays. In this case, the system can acquire audio data from multiple channels in parallel and, after digital signal conversion, perform signal fusion or beamforming on the multi-channel audio data to improve voice signal quality. Signal fusion can employ a delay-and-sum algorithm to weightedly sum the signals from different microphones, improving the signal-to-noise ratio (SNR) of the player's voice. Alternatively, algorithms such as MVDR (Multi-Volume Detection and Recognition) or adaptive beamforming can dynamically adjust microphone weights based on the noise environment, enhancing the voice signal in the target direction while reducing background noise in other directions.
[0038] After completing audio signal acquisition and conversion, the system performs preliminary noise reduction or gain processing on the digital audio at the hardware or audio driver level to ensure the stability and clarity of the voice signal. For example, the system can use automatic gain control (AGC) to automatically adjust the gain of the input signal to ensure that the voice volume is always at an appropriate level. In addition, the system can apply acoustic echo cancellation (AEC) algorithms to eliminate echo interference caused by the sound from the speaker returning to the microphone, thereby improving voice clarity.
[0039] Furthermore, the system can perform noise suppression (NS) on the audio signal. For example, spectral subtraction or adaptive filtering can be used to reduce the impact of background noise on the speech signal. Furthermore, spatial noise suppression technology can be combined to differentiate sound signals from different directions based on the spatial information of a multi-microphone array, effectively suppressing ambient noise such as keyboard tapping and mouse clicks.
[0040] After the aforementioned preprocessing, the digital audio data is divided into multiple time frames according to preset time windows (such as 20ms or 40ms). This process can be integrated with the game engine or underlying audio driver to ensure real-time game voice communication. The frame-processed audio data serves as input for subsequent energy value calculations, initial threshold comparisons, and dynamic threshold updates, providing the foundation for further speech noise reduction and optimization.
[0041] In some embodiments, converting the amplitude or decibel value of the time frame to obtain the energy value corresponding to each time frame includes:
[0042] Obtaining the amplitude of the digital audio sampling point in each time frame, squaring the amplitude of each digital audio sampling point in turn and accumulating the obtained results to obtain a value used to represent the time domain energy of the time frame;
[0043] Alternatively, the audio signal corresponding to the time frame is converted into a decibel value, and then, according to a preset mapping relationship, the decibel value is converted into a numerical value for representing the energy or power of the time frame.
[0044] Specifically, the digitized audio signal is divided into time frames of fixed length, such as 20ms or 40ms per frame. After framing, each time frame contains several audio sampling points, each of which corresponds to a numerical value representing the audio amplitude at that moment.
[0045] In one approach, the time-domain energy value can be calculated directly based on the audio amplitude within a time frame. Specifically, the system sequentially obtains the amplitude of all sampling points within the time frame and performs a nonlinear transformation on each amplitude to better reflect the overall energy of the signal. The transformed amplitudes are then accumulated to obtain the overall energy value for the time frame. This method effectively reflects the strength of the audio signal within the frame and is suitable for time-domain energy analysis.
[0046] Alternatively, the audio signal can be converted into decibel values, and then, using a preset mapping relationship, these decibel values can be converted into numerical values representing the energy or power of the time frame. Specifically, the system first converts the audio signal within the time frame into decibels, making the energy values more consistent with the human ear's auditory perception characteristics. Then, based on the preset mapping relationship, the decibel values are matched to standard energy or power levels to obtain a more representative time frame energy value. This method can more intuitively represent the relative intensity of the audio signal and is suitable for energy normalization processing in different devices and environments.
[0047] Regardless of the method used to calculate the energy value, the obtained time frame energy value can be used for subsequent noise reduction processing such as threshold comparison and dynamic threshold adjustment to improve the clarity and stability of the speech signal.
[0048] In some embodiments, the energy value is compared with a set initial threshold. If the energy value is lower than the initial threshold, it is marked as a noise frame or a low priority frame. If the energy value is higher than the initial threshold, it is retained as a speech frame for further processing.
[0049] Specifically, the system sets a fixed initial threshold based on historical data, hardware characteristics, or factory calibration results. This initial threshold measures whether the energy level of a time frame is sufficient to indicate that the frame contains valid speech signals. Specifically, this threshold can be adjusted based on factors such as microphone sensitivity and background noise levels on different devices, making it suitable for various hardware environments.
[0050] In some examples, after the system calculates the energy value for a certain time frame, it compares the energy value with a set initial threshold:
[0051] If the energy value falls below an initial fixed threshold, the time frame is considered to contain only ambient noise or silence, and the frame is marked as a noise frame or a low-priority frame. These marked frames can be discarded or used as candidate data for subsequent analysis to reduce interference from invalid signals on voice communications.
[0052] If the energy value is higher than the initial fixed threshold, it is considered that the time frame may contain a valid speech signal, so the frame is retained and further processed as a candidate speech frame, such as dynamic threshold adjustment, frequency domain analysis, etc., to improve the accuracy of speech recognition.
[0053] Furthermore, to adapt the fixed threshold to different usage scenarios, the system can adapt it to different audio devices or user environments. For example, in a noisy internet cafe, the initial threshold can be set slightly higher to avoid misidentifying background noise as speech; while in a quiet home environment, the initial threshold can be set lower to ensure that even weak speech signals are effectively captured.
[0054] Through the above-mentioned embodiments, this embodiment can effectively distinguish noise frames from speech frames based on a fixed threshold, providing a basis for subsequent dynamic threshold adjustment and advanced noise reduction processing, and ensuring the clarity and stability of the speech signal.
[0055] In some embodiments, when it is detected that the continuous time frames in which the energy value is lower than the initial threshold value exceed a preset time period, the energy value within the preset time period is statistically analyzed and smoothed to obtain an environmental noise reference value, including:
[0056] When it is detected that the continuous time frames with energy values lower than the initial threshold value exceed the preset time period, it is determined that the player has not spoken, and the continuous time frames within the preset time period are regarded as environmental noise frames, and the energy values of the environmental noise frames are collected;
[0057] The energy values of the ambient noise frames are counted, and the noise energy values in different time periods are continuously accumulated or averaged using a smoothing algorithm to obtain an ambient noise reference value used to characterize the ambient noise level.
[0058] Specifically, the system continuously monitors the energy of the audio signal and determines whether the energy level of a time frame falls below a preset initial threshold. If the energy levels of multiple consecutive time frames are below the initial threshold for a period exceeding a preset time period (e.g., 500ms or 1s), the system deems that the player was silent during that period and that these low-energy time frames were primarily composed of ambient noise. Therefore, these consecutive time frames are marked as ambient noise frames and their energy values are collected for subsequent calculations.
[0059] Next, the system counts and processes the collected ambient noise frame energy values to obtain a more stable ambient noise reference value. During the statistical process, the system can use the following methods:
[0060] Sliding window calculation: The energy values of the environmental noise frames in the most recent period (such as 1 second) are stored in a buffer, and the average or median value of these values is calculated to smooth out short-term fluctuating noise data.
[0061] Weighted filtering: assign different weights to noise data in different time periods, for example, give higher weights to the most recent noise frame to improve the ability to respond to noise changes.
[0062] Exponential smoothing method: recursively calculate historical environmental noise values so that newer data have a greater impact on the reference value, thereby enhancing the system's adaptability to changes in noise levels.
[0063] This method allows the system to dynamically estimate the overall ambient noise level, obtaining a more stable and accurate ambient noise reference value. This reference value can be used in subsequent dynamic threshold calculations, allowing the system to adjust voice detection standards based on real-time noise levels, thereby improving the stability and clarity of game voice communications.
[0064] Furthermore, this embodiment is adaptable to different usage environments. For example, in a quiet home environment, the system detects a low ambient noise reference value, and the subsequent speech detection threshold is automatically lowered to ensure that low-volume speech can still be correctly recognized. In a noisy internet cafe or outdoor environment, the ambient noise reference value is higher, and the system will accordingly increase the speech detection threshold to reduce background noise interference.
[0065] In some embodiments, a dynamic threshold is calculated based on an ambient noise reference value and a preset adjustment factor, and the dynamic threshold is updated in real time as the ambient noise reference value changes over time, including:
[0066] According to the environmental noise reference value and in combination with the preset adjustment factor, a dynamic threshold for distinguishing noise frames from speech frames is calculated using a predefined algorithm or mapping rule;
[0067] When it is monitored that the ambient noise reference value changes over time, the output of the algorithm or mapping rule is adaptively adjusted according to the adjustment factor, and the dynamic threshold is controlled to be synchronously increased or decreased so that the updated dynamic threshold matches the real-time fluctuating ambient noise level.
[0068] Specifically, the system uses the aforementioned method to calculate an ambient noise reference value, which represents the background noise level in the current environment. When the ambient noise is high, the reference value is higher; when the environment is quieter, the reference value is lower. The system combines this reference value with a preset adjustment factor to calculate a dynamic threshold, which is used to distinguish between noise frames and speech frames.
[0069] When calculating the dynamic threshold, the system uses preset mapping rules or algorithmic models to adapt the threshold to varying ambient noise levels. For example, the system can employ a nonlinear mapping function to calculate a dynamic threshold adapted to the current noise level based on an ambient noise reference value. As the ambient noise level changes, the system adjusts the calculated dynamic threshold based on the set adjustment factor to maintain smooth fluctuations and avoid false positives caused by rapid noise fluctuations.
[0070] When the system detects that the ambient noise reference value increases over time (for example, the player's environment becomes noisier, such as in Internet cafes, stations, etc.), it will simultaneously increase the dynamic threshold to reduce false detection of low-energy noise, ensuring that only voice signals with obvious intensity can pass through the noise reduction system, thereby improving voice quality and preventing environmental noise from interfering with normal communication.
[0071] Conversely, when the system detects that the ambient noise reference value decreases over time (for example, the player moves from a noisy environment into a quiet room, or the surrounding background noise decreases), the system automatically lowers the dynamic threshold to enhance the ability to capture low-volume voices, ensuring that the system can still accurately recognize valid voice signals when the player speaks softly or picks up voices from a distance.
[0072] Furthermore, to prevent frequent changes in the dynamic threshold due to short-term noise fluctuations, the system can set a minimum adjustment step size or a smoothing strategy. For example, when the ambient noise fluctuates slightly, the system can slowly adjust the dynamic threshold to prevent the speech detection system from overreacting to brief noise fluctuations. However, if the ambient noise level changes dramatically (such as when suddenly entering a noisy environment), the system can quickly adjust the dynamic threshold to adapt to the new noise environment and improve speech intelligibility.
[0073] Through this embodiment, the system can adaptively adjust the dynamic threshold based on the real-time changes in environmental noise, thereby providing a stable, high-quality game voice communication experience in different environments.
[0074] In some embodiments, the energy value is compared with a dynamic threshold. If the energy value is lower than the dynamic threshold, the corresponding time frame is filtered as a noise frame. If the energy value is higher than the dynamic threshold, the corresponding time frame is retained as a valid speech frame.
[0075] Specifically, the system calculates the energy value of each time frame based on the aforementioned method and obtains a dynamic threshold for the current environment. When the energy value of a time frame is below the dynamic threshold, the system deems the frame to be primarily composed of ambient noise and marks it as a noise frame or invalid frame, filtering it to reduce noise interference. Conversely, if the energy value of a time frame is above the dynamic threshold, the system deems the frame to likely contain speech signals and processes it as a candidate speech frame.
[0076] In some examples, in order to improve the accuracy of the determination, this embodiment may also combine a fixed threshold and a dynamic threshold to jointly determine the frame processing strategy. For example:
[0077] Category 1: Energy value is lower than a fixed threshold → directly judged as a noise frame and filtered.
[0078] Category 2: Energy values higher than the dynamic threshold → directly determined as valid speech frames, retained for subsequent encoding and transmission.
[0079] Category 3: Energy values between the fixed and dynamic thresholds indicate that the frame may contain weak speech or background noise. The system then performs further frequency domain analysis, such as using a fast Fourier transform (FFT) to extract the frame's spectral features and determine whether it contains the primary frequency components of the speech signal (e.g., 300Hz to 3400Hz). If the frame's spectral features match those of speech, the frame is retained; otherwise, it is filtered out as a noise frame.
[0080] Furthermore, to reduce the impact of short-term noise fluctuations on speech detection, the system can perform temporal continuity analysis on time frames. For example, if the energy value of a time frame is below the dynamic threshold, but the energy values of multiple adjacent time frames before and after it are all above the dynamic threshold, the system can infer that the frame may be a valid speech signal based on the coherence of the speech signal and thus retain the frame rather than directly filtering it out.
[0081] Through the method of this embodiment, the system can accurately distinguish between voice frames and noise frames in different noise environments, improve the intelligibility of voice signals, and reduce the interference of background noise, ensuring the clarity and stability of game voice communication.
[0082] In some embodiments, after comparing the energy value to the dynamic threshold, the method further comprises:
[0083] Perform frequency domain analysis on valid speech frames to convert time domain signals into frequency domain information. If non-speech frequencies or narrowband noise concentrated in a specific frequency band are detected, filtering is applied to the non-speech frequencies or narrowband noise.
[0084] Set corresponding multi-level energy thresholds in different frequency bands, and make differentiated judgments on the energy values of each frequency band based on the concentrated distribution characteristics of noise types in specific frequency bands;
[0085] The same audio frame or multiple consecutive audio frames are comprehensively analyzed based on short-term energy and long-term energy, and instantaneous impulse noise and continuous background noise are identified and filtered separately.
[0086] Specifically, in this embodiment, after comparing the energy value with the dynamic threshold, the valid voice frame is further subjected to frequency domain analysis, multi-level energy threshold setting, and short-time / long-time energy joint judgment to improve the system's ability to recognize and filter different types of noise, thereby optimizing the clarity and stability of game voice communication.
[0087] After completing the initial energy screening, some time frames may still contain complex environmental noise, such as electromagnetic interference, mechanical noise, or other non-speech frequency components. To further improve speech quality, the system performs frequency domain analysis on these valid speech frames to detect and remove specific noise components.
[0088] Specifically, the system first performs a Fast Fourier Transform (FFT) on each valid speech frame, converting the time-domain signal into a frequency-domain signal and analyzing its spectral distribution. Game speech is typically concentrated in the 300Hz to 3400Hz voice frequency band. If it detects non-speech frequencies with significant energy concentration (such as 50Hz power supply noise, keyboard tapping, or wind noise), the system can use the following methods to process it:
[0089] Notch Filter: For stable narrowband noise (such as power supply noise), a band-stop filter is used to remove specific frequency components while maintaining the integrity of the voice signal as much as possible.
[0090] Adaptive Filtering: If the noise frequency is detected to change over time, an adaptive filter can be used to adjust the noise suppression strategy in real time to adapt to different environmental noise patterns.
[0091] Through the above frequency domain conversion and filtering processing, non-speech components in the speech signal can be effectively removed, thereby improving the purity of the final speech signal.
[0092] Furthermore, different types of noise often have different spectral distribution characteristics, for example:
[0093] Wind noise is usually concentrated in the lower frequency band (20Hz~300Hz);
[0094] Mechanical noise (e.g., fan noise, road noise) may be distributed over multiple narrow frequency bands;
[0095] Keyboard tapping sounds usually contain higher-frequency components (1000Hz~5000Hz);
[0096] Explosive sounds can cover a wide frequency spectrum.
[0097] In some examples, to more accurately suppress these noises, the system sets different energy thresholds in different frequency bands and makes differentiated judgments for specific noise types. For example:
[0098] In the low-frequency area below 300 Hz, if the energy value exceeds the set low-frequency threshold, it may be wind noise. In this case, the gain of the low-frequency signal can be reduced or low-frequency filtering can be performed directly.
[0099] In the high-frequency region of 1000Hz to 5000Hz, if the energy value is too high and discontinuous short pulse signals appear, it can be determined as keyboard tapping sound and noise reduction processing is performed.
[0100] In the mid-frequency region (300Hz~3400Hz), if the energy value meets the characteristics of the speech band, it can be identified as a speech signal and is retained first.
[0101] By setting different energy thresholds in this way, the system can classify various types of noise more finely and improve the accuracy of noise filtering.
[0102] Furthermore, in order to further improve the stability of the speech signal, the system combines short-time energy analysis and long-time energy analysis to identify different types of noise patterns and perform corresponding processing.
[0103] Short-term energy analysis: Suitable for detecting sudden noises such as keyboard clicks and blasts. This type of noise typically manifests as a sharp increase in energy over a short period of time. When the system detects that the energy value of a time frame is significantly higher than that of the preceding and following frames and does not conform to the duration characteristics of a normal speech signal, it determines that the frame may contain sudden noise and performs additional noise reduction processing.
[0104] Long-term energy analysis: Suitable for detecting persistent noise, such as air conditioning, fan noise, or background vocalizations. The system calculates the smoothed average energy across multiple time frames to determine whether a noise is persistent. If the energy value in a frequency band remains high for an extended period with minimal variability, the system identifies persistent noise and implements appropriate noise reduction strategies, such as band attenuation or adaptive noise suppression.
[0105] In addition, the system can combine the analysis results of short-term energy and long-term energy to form a more robust filtering mechanism. For example:
[0106] If the short-term energy fluctuations in a certain time frame are large, but the long-term energy is relatively stable, it may be sudden noise and impulse noise suppression can be performed.
[0107] If both the short-term energy and the long-term energy of a time frame are high, it may be that the background noise is increasing. In this case, the dynamic threshold can be increased to reduce noise interference.
[0108] The frequency domain analysis, multi-level energy threshold setting, and short-term / long-term energy joint judgment in this embodiment can all be used together to achieve more accurate noise suppression. For example, the system can:
[0109] First, perform dynamic threshold screening to exclude obvious noise frames;
[0110] Perform frequency domain analysis on the speech frames after preliminary screening to remove narrowband noise and specific frequency interference;
[0111] Use multi-level energy thresholds to perform refined filtering for noise types in different frequency bands;
[0112] Combining short-term and long-term energy analysis, it is finally determined whether to retain a speech frame.
[0113] Through the steps of the above embodiment, this embodiment can effectively reduce noise interference in different environments, improve the clarity of game voice communication, and enable players to obtain a good voice interaction experience in noisy or quiet environments.
[0114] In some embodiments, the method further comprises:
[0115] Based on the speech and noise data collected in different environments, the noise frames and valid speech frames are annotated, and the mapping relationship between environmental noise characteristics and thresholds is trained using a machine learning algorithm to obtain the corresponding machine learning model;
[0116] During actual operation, whenever a new ambient noise reference value is detected, the ambient noise reference value is input into the machine learning model for inference and combined with the current dynamic threshold calculated based on the adjustment factor to output the optimized dynamic threshold;
[0117] The model parameters of the machine learning model are updated in real time using the collected continuously changing noise environment data.
[0118] Specifically, in this embodiment, in order to further improve the accuracy and environmental adaptability of the dynamic threshold, a machine learning model is combined to model the mapping relationship between voice signals and noise characteristics, and the model is used for real-time inference and adaptive adjustment in actual operation, thereby optimizing the noise filtering effect and improving the stability and clarity of game voice communication.
[0119] First, the system needs to establish a noise analysis and dynamic threshold calculation model based on machine learning. To this end, the system collects speech and noise data in different environments, including but not limited to:
[0120] a quiet home environment (low background noise);
[0121] Noisy internet cafes and game halls (high background noise);
[0122] Outdoor environments, such as parks or streets (complex dynamic noise);
[0123] Echoing or reverberant rooms (environments with different acoustic properties).
[0124] Furthermore, during the data collection process, the system pre-processes the audio data and divides it into time frames. Then, a manual or automatic labeling system classifies each time frame and labels it as:
[0125] Noise frame: contains only environmental noise and no obvious speech components.
[0126] Valid voice frames: Contains player voices and has a high signal-to-noise ratio.
[0127] Suspicious frames: Contain weak speech or mixed noise and require further analysis.
[0128] Furthermore, after the labeling is completed, the system uses a machine learning algorithm to train the mapping relationship between noise characteristics and thresholds, so that the model can automatically predict the optimal dynamic threshold based on the noise characteristics. The machine learning model applicable to this embodiment includes:
[0129] Neural Networks (RNN, CNN, Transformer): Suitable for complex speech and noise pattern analysis, capable of extracting deep features and performing intelligent noise reduction adjustments.
[0130] Gaussian mixture model (GMM): Suitable for probabilistic modeling of noise and speech signals and calculating the optimal classification threshold.
[0131] Support Vector Machine (SVM): Suitable for small sample learning and can achieve threshold optimization with a small amount of training data.
[0132] After training is completed, the system obtains a set of model parameters. The model can predict the optimal dynamic threshold based on the environmental noise reference value, so that it can adapt to different noise environments.
[0133] In actual operation, whenever a new ambient noise reference value is detected, the system inputs the reference value into the trained machine learning model and optimizes it in combination with the current dynamic threshold calculation method. The specific process is as follows:
[0134] 1) Noise level detection:
[0135] The system continuously monitors the ambient noise level, calculates the ambient noise reference value, and determines whether the current dynamic threshold needs to be adjusted.
[0136] 2) Model inference:
[0137] The current environmental noise reference value is input into the machine learning model, and the system automatically calculates the optimal dynamic threshold or threshold range suitable for the current environment.
[0138] This inference process combines historical training data, enabling the system to adjust thresholds based on past experience and adapt to various complex environments.
[0139] 3) Threshold fusion:
[0140] The optimal threshold calculated by the machine learning model is fused with the current dynamic threshold calculated based on the adjustment factor.
[0141] A weighted average or adaptive adjustment algorithm can be used to ensure that the final threshold is consistent with the current noise environment and does not cause misjudgment due to sudden noise changes.
[0142] 4) Real-time applications:
[0143] The optimized dynamic threshold finally calculated is applied to the game voice processing system to filter subsequent time frames, so that the system can always operate under the best noise suppression conditions.
[0144] For example, when a player moves from a quiet room into a noisy internet cafe, the system detects a significant increase in ambient noise levels. The machine learning model predicts a higher dynamic threshold to avoid misidentifying ambient noise as valid speech signals. When the player returns to a quiet environment, the system automatically lowers the dynamic threshold to ensure that even when the player speaks softly, their speech can still be recognized.
[0145] Furthermore, in order to enable the system to adapt to different users and environments in the long term, an online learning mechanism is adopted to continuously optimize the performance of the machine learning model.
[0146] During long-term operation, the system continuously collects new environmental noise samples and stores them in local or cloud databases. This data can be used to expand the training set, enabling the model to adapt to complex situations such as different devices, different users, and different environments.
[0147] Through incremental training or federated learning, the model is gradually optimized without affecting the real-time user experience. If the dynamic threshold calculation is detected to be poor in certain specific environments (such as special echo environments or extreme noise environments), the system can adjust the adjustment factor or model parameters to improve adaptability.
[0148] This embodiment allows users to manually adjust the voice detection sensitivity and transmits this adjustment data back to the server for model optimization. For example, if a user repeatedly lowers the dynamic threshold in an internet cafe environment, the system can learn this habit and automatically lower the threshold in similar environments.
[0149] This embodiment combines a machine learning model with a dynamic threshold calculation method to achieve intelligent speech noise reduction optimization. Compared with the traditional fixed threshold method, it has the following advantages:
[0150] Traditional methods use fixed formulas to calculate dynamic thresholds, which are difficult to adapt to sudden changes in the environment. Machine learning models can adaptively adjust thresholds based on historical data to improve noise reduction effects in different environments.
[0151] Machine learning models can be trained based on large amounts of noisy environment data, making the calculated dynamic thresholds more accurate and reducing misjudgments.
[0152] Through the online learning mechanism, the system can continuously optimize the model parameters so that it can maintain a high recognition accuracy rate in the long term.
[0153] Adapt to various complex environments and ensure that players can get clear and stable voice communication experience in different scenarios such as Internet cafes, outdoors, and at home.
[0154] Through the method of this embodiment, the game voice noise reduction system can dynamically adjust the threshold, intelligently adapt to the environment, and continuously optimize the model, so that players can enjoy high-quality voice communication in different noise environments and avoid communication barriers caused by noise interference or misjudgment.
[0155] According to the technical solutions of the above embodiments of the present application, the present application has at least the following advantages:
[0156] Highly adaptable: Based on a dynamic threshold that changes with the ambient noise level, it can achieve stable and appropriate noise filtering effects in noisy or quiet environments.
[0157] High Accuracy: Through multi-stage filtering, frequency domain analysis, and optional machine learning-assisted mechanisms, different types of noise can be more effectively suppressed and filtered.
[0158] Strong real-time performance: Dynamic adjustment of the threshold can be completed within a time scale of milliseconds to seconds, meeting the real-time communication requirements of game voice and ensuring a good interactive experience for players.
[0159] Scalability: The solution can be integrated with hardware noise reduction modules (such as AEC / AGC), multi-microphone beamforming, and cloud-based machine learning platforms. It has a wide range of applications, including but not limited to VR / AR gaming, mobile network voice chat, and remote conferencing systems.
[0160] The following are device embodiments of the present application, which can be used to implement the method embodiments of the present application. For details not disclosed in the device embodiments of the present application, please refer to the method embodiments of the present application.
[0161] Figure 2 This is a schematic diagram of the structure of the game team voice noise reduction device based on the dynamic threshold mechanism provided by the embodiment of the present application. Figure 2 As shown, the game team voice noise reduction device based on the dynamic threshold mechanism includes:
[0162] The acquisition module 201 is used to acquire an audio signal containing player voice and environmental noise, and divide the audio signal into multiple time frames;
[0163] The conversion module 202 is used to convert the amplitude or decibel value of the time frame to obtain the energy value corresponding to each time frame;
[0164] A first comparison module 203 is configured to compare the energy value with a set initial threshold value. If the energy value is lower than the initial threshold value, the frame is marked as a noise frame or a low priority frame. If the energy value is higher than the initial threshold value, the frame is retained as a speech frame for further processing.
[0165] The processing module 204 is configured to, when it is detected that the continuous time frames in which the energy value is lower than the initial threshold value exceed the preset time period, perform statistics and smoothing processing on the energy value within the preset time period to obtain an environmental noise reference value;
[0166] A calculation module 205 is configured to calculate a dynamic threshold based on an ambient noise reference value and a preset adjustment factor, and to update the dynamic threshold in real time as the ambient noise reference value changes over time;
[0167] A second comparison module 206 is configured to compare the energy value with a dynamic threshold. If the energy value is lower than the dynamic threshold, the corresponding time frame is filtered as a noise frame. If the energy value is higher than the dynamic threshold, the corresponding time frame is retained as a valid speech frame.
[0168] The generation module 207 is used to encode the time frame retained after filtering by the dynamic threshold to generate a final game team voice signal.
[0169] In some embodiments, Figure 2 The acquisition module 201 directly obtains single-channel audio data as an analog voice signal when the player uses a single microphone; when the player uses a headset or an external multi-microphone device, multiple channels of audio data are obtained in parallel, and the multiple channels of audio data are subjected to signal fusion or beamforming to obtain an analog voice signal; the collected analog voice signal is converted into a digital signal, and the digital signal is quantized to obtain digital audio that conforms to a predetermined audio coding format; at the hardware level or in the audio driver, the automatic gain control or acoustic echo cancellation module is used to perform preliminary noise reduction or gain processing on the digital audio; the digital audio after preliminary noise reduction or gain processing is divided into multiple time frames according to a preset time window, so as to perform initial threshold comparison and dynamic threshold update according to the energy value of the time frame.
[0170] In some embodiments, Figure 2 The conversion module 202 obtains the amplitude of the digital audio sampling point in each time frame, squares the amplitude of each digital audio sampling point in turn, and accumulates the obtained results to obtain a numerical value used to represent the time domain energy of the time frame; or converts the audio signal corresponding to the time frame into a decibel value, and converts the decibel value into a numerical value used to represent the energy or power of the time frame according to a preset mapping relationship.
[0171] In some embodiments, Figure 2 When the processing module 204 detects that the continuous time frames with energy values lower than the initial threshold value exceed the preset time period, it determines that the player has not spoken, and regards the continuous time frames within the preset time period as environmental noise frames, and collects the energy values of the environmental noise frames; the energy values of the environmental noise frames are counted, and the noise energy values in different time periods are continuously accumulated or averaged using a smoothing algorithm to obtain an environmental noise reference value for representing the environmental noise level.
[0172] In some embodiments, Figure 2 The calculation module 205 calculates the dynamic threshold for distinguishing noise frames from speech frames based on the ambient noise reference value and in combination with a preset adjustment factor, using a predefined algorithm or mapping rule; when it is monitored that the ambient noise reference value changes over time, the output of the algorithm or mapping rule is adaptively adjusted according to the adjustment factor, and the dynamic threshold is controlled to be synchronously increased or decreased so that the updated dynamic threshold matches the real-time fluctuating ambient noise level.
[0173] In some embodiments, Figure 2The analysis module 208 performs frequency domain analysis on the valid speech frames and converts the time domain signal into frequency domain information. If non-speech frequencies or narrow-band noise concentrated in a specific frequency band are detected, filtering is applied to the non-speech frequencies or narrow-band noise. Corresponding multi-level energy thresholds are set in different frequency bands, and the energy values of each frequency band are differentiated according to the concentrated distribution characteristics of the noise type in the specific frequency band. The same audio frame or multiple consecutive audio frames are comprehensively analyzed based on the short-time energy and long-time energy, and instantaneous impulse noise and continuous background noise are identified and filtered separately.
[0174] In some embodiments, Figure 2 The training module 209 marks noise frames and valid speech frames based on the speech and noise data collected in different environments, and uses a machine learning algorithm to train the mapping relationship between environmental noise characteristics and thresholds to obtain a corresponding machine learning model; during actual operation, whenever a new environmental noise reference value is detected, the environmental noise reference value is input into the machine learning model for inference, and combined with the current dynamic threshold calculated according to the adjustment factor to output the optimized dynamic threshold; the model parameters of the machine learning model are updated in real time using the collected continuously changing noise environment data.
[0175] It should be understood that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.
[0176] Figure 3 Schematic diagram of the structure of the electronic device 3 provided in the embodiment of the present application. Figure 3 As shown, the electronic device 3 of this embodiment includes: a processor 301, a memory 302, and a computer program 303 stored in the memory 302 and executable on the processor 301. When the processor 301 executes the computer program 303, the steps of the above-mentioned method embodiments are implemented. Alternatively, when the processor 301 executes the computer program 303, the functions of the modules / units in the above-mentioned device embodiments are implemented.
[0177] For example, computer program 303 may be divided into one or more modules / units, which are stored in memory 302 and executed by processor 301 to implement the present application. One or more modules / units may be a series of computer program instruction segments capable of performing specific functions, and the instruction segments are used to describe the execution process of computer program 303 in electronic device 3.
[0178] The electronic device 3 may be a desktop computer, a notebook, a PDA, a cloud server or other electronic device. The electronic device 3 may include but is not limited to a processor 301 and a memory 302. Those skilled in the art will understand that Figure 3 It is only an example of electronic device 3 and does not constitute a limitation of electronic device 3. It may include more or fewer components than shown in the figure, or a combination of certain components, or different components. For example, the electronic device may also include input and output devices, network access devices, buses, etc.
[0179] The processor 301 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor.
[0180] Memory 302 can be an internal storage unit of electronic device 3, such as a hard drive or memory of electronic device 3. Memory 302 can also be an external storage device of electronic device 3, such as a plug-in hard drive, a Smart Media Card (SMC), a Secure Digital (SD) card, a flash memory card, etc. Furthermore, memory 302 can include both an internal storage unit of electronic device 3 and an external storage device. Memory 302 is used to store computer programs and other programs and data required by the electronic device. Memory 302 can also be used to temporarily store data that has been output or is about to be output.
[0181] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the division of the above-mentioned functional units and modules is used as an example for illustration. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiment can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of software functional units. In addition, the specific names of the functional units and modules are only for the convenience of distinguishing each other, and are not used to limit the scope of protection of this application. The specific working process of the units and modules in the above-mentioned system can refer to the corresponding process in the aforementioned method embodiment, and will not be repeated here.
[0182] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described or recorded in detail in a certain embodiment, reference can be made to the relevant description of other embodiments.
[0183] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0184] In the embodiments provided in this application, it should be understood that the disclosed apparatus / computer equipment and methods can be implemented in other ways. For example, the apparatus / computer equipment embodiments described above are merely schematic. For example, the division of modules or units is merely a logical function division. In actual implementation, there may be other division methods. Multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed may be through some interfaces, indirect coupling or communication connection of the apparatus or unit, which may be electrical, mechanical or other forms.
[0185] Units described as separate components may or may not be physically separate, and components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0186] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0187] If the integrated module / unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the present application can implement all or part of the processes in the above-mentioned embodiment method by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, it can implement the steps of each of the above-mentioned method embodiments. The computer program may include computer program code, which may be in source code form, object code form, executable file, or some intermediate form. Computer-readable media may include: any entity or device capable of carrying computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium.
[0188] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present application, and should all be included in the scope of protection of the present application.
Claims
1. A method for reducing game team voice noise based on a dynamic threshold mechanism, characterized in that: include: Collecting an audio signal containing a player's voice and ambient noise, and dividing the audio signal into a plurality of time frames; Converting the amplitude or decibel value of the time frame to obtain an energy value corresponding to each time frame; Comparing the energy value with a set initial threshold, if the energy value is lower than the initial threshold, marking it as a noise frame or a low priority frame; if the energy value is higher than the initial threshold, retaining it as a speech frame to be further processed; When it is detected that the continuous time frames in which the energy value is lower than the initial threshold value exceed a preset time period, recursively calculating the energy values of the continuous time frames using an exponential smoothing method to obtain an environmental noise reference value; A dynamic threshold is calculated based on the ambient noise reference value and a preset adjustment factor, and the dynamic threshold is updated in real time when the ambient noise reference value changes over time; wherein the system sets a minimum adjustment step or a smooth adjustment strategy, and when the ambient noise changes slightly, the system adjusts the dynamic threshold slowly, and if the ambient noise changes drastically, the system adjusts the dynamic threshold quickly; Comparing the energy value with the dynamic threshold, if the energy value is lower than the dynamic threshold, filtering the corresponding time frame as a noise frame, and if the energy value is higher than the dynamic threshold, retaining the corresponding time frame as a valid speech frame; wherein, if the energy value of a time frame is between the initial threshold and the dynamic threshold, performing a fast Fourier transform to extract the spectrum characteristics of the time frame, and determining whether it contains the main frequency components of the speech signal; if the spectrum characteristics of the time frame meet the speech characteristics, retaining the time frame, otherwise filtering the time frame as a noise frame; The time frames retained after filtering by the dynamic threshold are encoded to generate a final game team voice signal.
2. The method according to claim 1, characterized in that The collecting of an audio signal including a player's voice and environmental noise, and dividing the audio signal into a plurality of time frames, includes: When the player uses a single microphone, the single-channel audio data is directly obtained as the analog voice signal; When players use headsets or external multi-microphone devices, multiple channels of audio data are acquired in parallel and signal fusion or beamforming is performed on the multiple channels of audio data to obtain analog voice signals; Converting the collected analog voice signal into a digital signal and quantizing the digital signal to obtain digital audio that complies with a predetermined audio coding format; At the hardware level or in the audio driver, using an automatic gain control or acoustic echo cancellation module to perform preliminary noise reduction or gain processing on the digital audio; The digital audio after the preliminary noise reduction or gain processing is divided into multiple time frames according to a preset time window, so as to perform initial threshold comparison and dynamic threshold update according to the energy values of the time frames.
3. The method according to claim 1, characterized in that The converting the amplitude or decibel value of the time frame to obtain the energy value corresponding to each time frame includes: Obtaining the amplitude of the digital audio sampling point in each time frame, squaring the amplitude of each digital audio sampling point in turn and accumulating the obtained results to obtain a value used to represent the time domain energy of the time frame; Alternatively, the audio signal corresponding to the time frame is converted into a decibel value, and the decibel value is converted into a numerical value for representing the energy or power of the time frame according to a preset mapping relationship.
4. The method according to claim 1, wherein When it is detected that the continuous time frame in which the energy value is lower than the initial threshold exceeds a preset time period, the energy value within the preset time period is statistically analyzed and smoothed to obtain an environmental noise reference value, including: When it is detected that the continuous time frames in which the energy value is lower than the initial threshold value exceed the preset time period, it is determined that the player has not spoken, the continuous time frames within the preset time period are regarded as environmental noise frames, and the energy values of the environmental noise frames are collected; The energy values of the environmental noise frames are statistically analyzed, and the noise energy values in different time periods are continuously accumulated or averaged using a smoothing algorithm to obtain an environmental noise reference value for characterizing the environmental noise level.
5. The method according to claim 1, characterized in that The calculating the dynamic threshold based on the ambient noise reference value and a preset adjustment factor, and updating the dynamic threshold in real time when the ambient noise reference value changes over time, includes: Calculating a dynamic threshold for distinguishing noise frames from speech frames based on the environmental noise reference value and in combination with a preset adjustment factor using a predefined algorithm or mapping rule; When it is monitored that the ambient noise reference value changes over time, the output of the algorithm or mapping rule is adaptively adjusted according to the adjustment factor, and the dynamic threshold is controlled to be synchronously increased or decreased so that the updated dynamic threshold matches the real-time fluctuating ambient noise level.
6. The method according to claim 1, characterized in that After comparing the energy value with the dynamic threshold, the method further includes: Performing frequency domain analysis on the valid speech frame to convert the time domain signal into frequency domain information, and if non-speech frequencies or narrowband noise concentrated in a specific frequency band are detected, applying filtering processing to the non-speech frequencies or narrowband noise; Set corresponding multi-level energy thresholds in different frequency bands, and make differentiated judgments on the energy values of each frequency band based on the concentrated distribution characteristics of noise types in specific frequency bands; The same audio frame or multiple consecutive audio frames are comprehensively analyzed based on short-term energy and long-term energy, and instantaneous impulse noise and continuous background noise are identified and filtered separately.
7. The method according to claim 1, characterized in that The method further comprises: Based on the speech and noise data collected in different environments, the noise frames and valid speech frames are annotated, and the mapping relationship between environmental noise characteristics and thresholds is trained using a machine learning algorithm to obtain the corresponding machine learning model; During actual operation, whenever a new ambient noise reference value is detected, the ambient noise reference value is input into the machine learning model for inference, and combined with the current dynamic threshold value calculated based on the adjustment factor to output an optimized dynamic threshold value; The model parameters of the machine learning model are updated in real time using the collected continuously changing noise environment data.
8. A game team voice noise reduction device based on a dynamic threshold mechanism, characterized in that: include: An acquisition module, configured to acquire an audio signal containing player voice and ambient noise, and divide the audio signal into multiple time frames; a conversion module, configured to convert the amplitude or decibel value of the time frame to obtain an energy value corresponding to each time frame; a first comparison module, configured to compare the energy value with a set initial threshold, and if the energy value is lower than the initial threshold, mark the frame as a noise frame or a low-priority frame; and if the energy value is higher than the initial threshold, retain the frame as a speech frame to be further processed; a processing module, configured to, when detecting that the continuous time frames in which the energy value is lower than the initial threshold value exceed a preset time period, recursively calculate the energy values of the continuous time frames using an exponential smoothing method to obtain an environmental noise reference value; a calculation module for calculating a dynamic threshold based on the ambient noise reference value and a preset adjustment factor, and updating the dynamic threshold in real time as the ambient noise reference value changes over time; wherein the system sets a minimum adjustment step size or a smooth adjustment strategy, and when the ambient noise changes slightly, the system adjusts the dynamic threshold slowly; if the ambient noise changes drastically, the system adjusts the dynamic threshold quickly; a second comparison module, configured to compare the energy value with the dynamic threshold; if the energy value is lower than the dynamic threshold, the corresponding time frame is filtered as a noise frame; if the energy value is higher than the dynamic threshold, the corresponding time frame is retained as a valid speech frame; wherein, if the energy value of a time frame is between the initial threshold and the dynamic threshold, a fast Fourier transform is performed to extract the spectral features of the time frame, and determine whether it contains the main frequency components of the speech signal; if the spectral features of the time frame meet the speech features, the time frame is retained; otherwise, the time frame is filtered as a noise frame; The generation module is used to encode the time frame retained after filtering through the dynamic threshold to generate a final game team voice signal.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Method for enhancing clarity of speech signals
CN107967918A
Sound source processing method and related device
CN117079661A