Game team voice noise reduction method and device based on multi-channel signal processing

Through the signal processing of multi-channel microphone array, efficient voice noise reduction and enhancement in complex noise environments are achieved, solving the problem of poor single-channel sound pickup effect, and improving the speech clarity and real-timeness in game teams.

CN120032661BActive Publication Date: 2025-08-22QINGFENG (BEIJING) TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510480610.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-17
Publication Date
2025-08-22
Estimated Expiration
2045-04-17

AI Technical Summary

Technical Problem

In multiplayer games, single-channel sound pickup is difficult to effectively suppress complex noise, resulting in a decrease in clarity and intelligibility of voice chats, which cannot meet the needs of high-quality audio and real-time team communication.

Method used

The audio signal is obtained using a multi-channel microphone array, through synchronous sampling and digitization processing, frame division, energy threshold filtering, sound source positioning and beamforming, combined with multi-channel noise reduction processing, suppress background noise and enhance target voice.

Benefits of technology

It improves the noise suppression effect and voice real-time in complex noise environments, improves voice quality and stability, and ensures clear communication in game teams.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120032661B_ABST
    Figure CN120032661B_ABST
Patent Text Reader

Abstract

The present application provides a method and device for game team voice noise reduction based on multi-channel signal processing. The method includes: obtaining a multi-channel audio signal, and synchronously sampling and digitally processing the multi-channel audio signal; dividing the processed multi-channel audio signal into frames, and determining the corresponding energy value according to the decibel value of each frame of the audio signal; filtering or attenuating the audio signal with an energy value lower than the energy threshold, and locating the sound source according to the multi-channel audio signal after preliminary noise reduction; performing beamforming and multi-channel noise reduction processing on the multi-channel audio signal after preliminary noise reduction according to the target voice direction information; performing voice activation detection and gain control on the voice signal obtained after beamforming and multi-channel noise reduction processing, and outputting the optimized voice signal to the game communication module. The present application can improve the noise suppression effect in complex noise environments, improve the real-time performance and stability of voice, and improve voice quality.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of voice processing technology, and in particular to a method and device for reducing game team voice noise based on multi-channel signal processing. Background Art

[0002] In multiplayer games, players typically communicate and collaborate through real-time voice communication. However, in noisy environments, single-channel voice pickup often struggles to effectively suppress various interfering noises, such as wind noise, keyboard tapping, and background music. This results in reduced clarity and intelligibility in voice chat, directly impacting the gaming experience. With advancements in hardware, modern devices are now equipped with multi-microphone arrays, which provide more dimensional signal information for voice noise reduction and enhancement.

[0003] Existing single-channel speech noise reduction technologies primarily include short-time energy analysis, spectral subtraction, and adaptive filtering. While these methods are somewhat effective in relatively simple noise environments, they often struggle to achieve both clarity and real-time performance in gaming scenarios, where noise types vary and intensity fluctuates. Furthermore, single-channel processing cannot leverage multi-channel time difference, phase, and amplitude information for more effective target sound source localization and noise suppression, making it difficult to meet the demands for high-quality audio and real-time team communication.

[0004] Therefore, in multiplayer game voice applications, how to fully utilize the signal processing capabilities of multi-microphone arrays while taking into account stable noise suppression effects and real-time processing speed has become a technical problem that needs to be solved urgently. Summary of the Invention

[0005] In view of this, the embodiments of the present application provide a game team voice noise reduction method and device based on multi-channel signal processing to solve the problems existing in the prior art, such as poor noise suppression effect in complex noise environments, reduced real-time performance and stability of voice, and reduced voice quality.

[0006] In a first aspect of an embodiment of the present application, a game team voice noise reduction method based on multi-channel signal processing is provided, comprising: obtaining a multi-channel audio signal using a plurality of microphone arrays arranged on a game terminal, and synchronously sampling and digitally processing the multi-channel audio signal; dividing the processed multi-channel audio signal into frames, and determining the corresponding energy value according to the decibel value of each frame of the audio signal; filtering or attenuating the audio signal with an energy value lower than the energy threshold according to an adaptively determined energy threshold to obtain a multi-channel audio signal after preliminary noise reduction; locating the sound source according to the multi-channel audio signal after preliminary noise reduction using the time delay estimation or amplitude difference estimation between each microphone channel to obtain the target voice direction information; performing beamforming and multi-channel noise reduction processing on the multi-channel audio signal after preliminary noise reduction according to the target voice direction information to suppress background noise and enhance the target voice; performing voice activation detection and gain control on the voice signal obtained after beamforming and multi-channel noise reduction processing, and outputting the optimized voice signal to the game communication module.

[0007] According to a second aspect of an embodiment of the present application, a game team voice noise reduction device based on multi-channel signal processing is provided, comprising: an acquisition module for acquiring a multi-channel audio signal using a plurality of microphone arrays provided on a game terminal, and synchronously sampling and digitizing the multi-channel audio signal; a division module for performing frame division on the processed multi-channel audio signal and determining a corresponding energy value based on the decibel value of each frame of the audio signal; a noise reduction module for filtering or attenuating audio signals having energy values ​​below the energy threshold according to an adaptively determined energy threshold, to obtain a multi-channel audio signal after preliminary noise reduction; a positioning module for performing sound source positioning based on the multi-channel audio signal after preliminary noise reduction using time delay estimation or amplitude difference estimation between each microphone channel to obtain target voice direction information; a suppression module for performing beamforming and multi-channel noise reduction processing on the multi-channel audio signal after preliminary noise reduction based on the target voice direction information to suppress background noise and enhance the target voice; and an output module for performing voice activation detection and gain control on the voice signal obtained after the beamforming and multi-channel noise reduction processing, and outputting the optimized voice signal to the game communication module.

[0008] According to a third aspect of an embodiment of the present application, an electronic device is provided, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the steps of the above method are implemented when the processor executes the computer program.

[0009] According to a fourth aspect of the embodiments of the present application, a computer-readable storage medium is provided, which stores a computer program. When the computer program is executed by a processor, the steps of the above method are implemented.

[0010] At least one of the above technical solutions adopted in the embodiments of the present application can achieve the following beneficial effects:

[0011] By utilizing a plurality of microphone arrays provided on a game terminal to obtain a multi-channel audio signal, and synchronously sampling and digitally processing the multi-channel audio signal; dividing the processed multi-channel audio signal into frames, and determining the corresponding energy value according to the decibel value of each frame of the audio signal; filtering or attenuating the audio signal with an energy value lower than the energy threshold according to the adaptively determined energy threshold, and obtaining a multi-channel audio signal after preliminary noise reduction; based on the multi-channel audio signal after preliminary noise reduction, the time delay estimation or amplitude difference estimation between each microphone channel is used to locate the sound source and obtain the target voice direction information; based on the target voice direction information, beamforming and multi-channel noise reduction processing are performed on the multi-channel audio signal after preliminary noise reduction to suppress background noise and enhance the target voice; the voice signal obtained after beamforming and multi-channel noise reduction processing is subjected to voice activation detection and gain control, and the optimized voice signal is output to the game communication module. This application can improve the noise suppression effect in a complex noise environment, improve the real-time performance and stability of voice, and improve voice quality. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the embodiments or descriptions of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0013] Figure 1 1 is a flow chart of a method for reducing game team voice noise based on multi-channel signal processing provided by an embodiment of the present application;

[0014] Figure 2 2 is a schematic diagram of the structure of a game team voice noise reduction device based on multi-channel signal processing provided by an embodiment of the present application;

[0015] Figure 3 It is a structural diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0016] In the following description, specific details such as specific system structures and techniques are provided for purposes of illustration rather than limitation to facilitate a thorough understanding of the embodiments of the present application. However, it will be apparent to those skilled in the art that the present application may be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid obscuring the description of the present application with unnecessary detail.

[0017] In multiplayer online games, real-time voice chat is often provided to facilitate communication and collaboration between players. However, when players are in a noisy environment or the device's sound pickup is poor, various noises such as wind noise, keyboard tapping, and background music often interfere with the quality and efficiency of team communication.

[0018] To reduce background noise interference and improve the clarity and intelligibility of voice chats, a more comprehensive and effective voice noise reduction and enhancement technology is needed that can maintain good voice quality in extreme or complex environments while also balancing real-time performance and system resource usage.

[0019] In the past, common speech noise reduction and enhancement methods mainly relied on single-channel processing methods, that is, using the audio signal collected by a single microphone to perform speech enhancement and noise suppression. Typical single-channel noise reduction methods include:

[0020] Short-time energy or spectral subtraction: by estimating the background noise and attenuating it in the frequency domain;

[0021] Voice Activity Detection (VAD): Detects speech and noise intervals, and then strongly reduces the noise interval;

[0022] Adaptive filtering: Filters audio signals based on a specific noise model.

[0023] Single-channel processing can be effective in simple, noisy environments. However, as users' demands for communication quality and environmental variability increase (such as in noisy esports venues or public places), the sound pickup and noise reduction capabilities of a single microphone are often insufficient, making it difficult to ensure high-quality voice communication.

[0024] At the same time, modern smart devices and gaming consoles are beginning to be equipped with multiple microphones or even microphone arrays. These hardware capabilities provide richer signal information for multi-channel voice processing. For example, different channels may have different phases, delays, and amplitudes of sound arrival, which can help the system locate the sound source and achieve more accurate noise suppression.

[0025] In multiplayer game voice scenarios, common problems focus on the following aspects:

[0026] Severe background noise interference: Gamers may be in a noisy environment (such as an internet cafe or a competition venue), and single-channel noise reduction is limited in effectiveness.

[0027] High real-time and stability requirements: Game communication requires instant feedback. Excessive latency can lead to communication breakdowns. Furthermore, the type and intensity of noise may vary at any time.

[0028] Increased environmental complexity: In the same game scene, user devices vary in form (mobile phones, computers, headphones, etc.) and are scattered across different areas, requiring compatibility with multiple hardware and complex sound collection methods;

[0029] Voice quality and bandwidth utilization: Voice signals must be transmitted as clearly and without distortion as possible within limited bandwidth, while preventing large amounts of invalid or pure noise data from occupying network resources.

[0030] Therefore, how to make full use of the multi-channel microphone array to achieve better noise filtering and voice enhancement and reduce the noise occupation during transmission is the key technical problem that this application focuses on and aims to solve.

[0031] Before describing the embodiments of the present application in detail, the following describes the system architecture involved in the present application in actual scenarios. The game team voice noise reduction system based on multi-channel signal processing involved in the present application in actual scenarios may include the following:

[0032] 1. Multi-channel microphone acquisition module

[0033] Multiple microphone arrays are configured on the user's gaming terminal (or audio terminal) to synchronously collect ambient sound and voice signals.

[0034] Each microphone channel outputs a corresponding analog audio signal, which is then digitized (analog-to-digital conversion) in the audio front end.

[0035] 2. Audio preprocessing module

[0036] Perform conventional pre-processing operations such as DC offset removal, anti-aliasing filtering, and sampling rate conversion on the digitized audio signal.

[0037] Time-synchronize multi-channel signals to ensure that each channel signal has a consistent time domain reference during subsequent processing.

[0038] 3. Decibel to energy value conversion and threshold filtering module

[0039] The pre-processed multi-channel audio signal is divided into frames (such as 20ms or 10ms per frame) and the decibel value corresponding to each frame signal is calculated.

[0040] Convert decibel values ​​into energy values ​​(such as power spectrum density, short-time energy, etc.) to more intuitively reflect the sound intensity.

[0041] Set a dynamic or static energy threshold to filter or attenuate frames below the threshold to initially reduce the interference of background noise.

[0042] 4. Multi-channel signal processing and core algorithm module

[0043] Sound source localization: Based on multi-channel microphone array processing technology, the position or direction of the main sound source (such as player voice) is determined through methods such as time difference estimation, amplitude difference estimation, or phase difference estimation.

[0044] Beamforming: Uses sound source localization information to enhance speech signals in the main direction and suppress noise in other directions.

[0045] Multi-channel noise reduction: Based on beamforming, multi-channel Wiener filtering, MVDR (Minimum Variance Distortionless Response) filtering, or neural network methods (such as deep learning speech enhancement networks) are combined to further reduce noise in the time and frequency domains of multi-channel signals.

[0046] 5. Post-processing and output module

[0047] The initially purified speech output by the multi-channel signal processing module is then post-processed with voice activation detection (VAD) and speech gain control (AGC).

[0048] The optimized voice signal is output to the game's voice communication module or server to achieve real-time noise suppression and high-quality voice calls.

[0049] The contents of the technical solution of this application are described in detail below with reference to the accompanying drawings and specific embodiments.

[0050] Figure 1 Schematic diagram of the process of the game team voice noise reduction method based on multi-channel signal processing provided by the embodiment of the present application. Figure 1 As shown, the game team voice noise reduction method based on multi-channel signal processing may specifically include:

[0051] S101, using a plurality of microphone arrays provided on a game terminal to obtain a multi-channel audio signal, and synchronously sampling and digitally processing the multi-channel audio signal;

[0052] S102, dividing the processed multi-channel audio signal into frames, and determining a corresponding energy value according to the decibel value of each frame of the audio signal;

[0053] S103, filtering or attenuating audio signals having energy values ​​lower than the energy threshold according to the adaptively determined energy threshold to obtain a multi-channel audio signal after preliminary noise reduction;

[0054] S104, based on the multi-channel audio signal after preliminary noise reduction, using the time delay estimation or amplitude difference estimation between each microphone channel to perform sound source localization and obtain target speech direction information;

[0055] S105, performing beamforming and multi-channel noise reduction processing on the multi-channel audio signal after the preliminary noise reduction according to the target voice direction information to suppress background noise and enhance the target voice;

[0056] S106 , performing voice activation detection and gain control on the voice signal obtained after beamforming and multi-channel noise reduction processing, and outputting the optimized voice signal to the game communication module.

[0057] In some embodiments, synchronously sampling and digitizing a multi-channel audio signal includes:

[0058] Performing analog-to-digital conversion on analog audio signals output by multiple microphone arrays using a common sampling clock;

[0059] Performing time stamp alignment or interpolation processing on the converted multi-channel digital audio signal to make each channel consistent on the same time axis;

[0060] At least one pre-processing operation of direct current offset removal, anti-aliasing filtering, or sampling rate conversion is performed on the aligned multi-channel digital audio signal to obtain a processed multi-channel audio signal.

[0061] Specifically, in some examples, a user's gaming device or communication terminal is equipped with four microphone arrays, arranged near the four corners of the device, forming a regular spatial distribution structure. The spacing and positional relationship between each microphone array is known before shipment and recorded in the device's array configuration information file. This information file can be used for subsequent delay estimation and sound source localization.

[0062] First, in the four microphone arrays, each array integrates several microphones (such as a linear array or a circular array), and ensures that the physical coordinates of each microphone are accurately known.

[0063] The analog output of each array is connected to a common sampling clock through an independent analog-to-digital converter (ADC), ensuring that each channel is sampled at the same time.

[0064] Furthermore, a common sampling clock signal is fed into the ADC modules of the four microphone arrays in parallel, providing a unified and synchronized sampling reference for all microphone channels.

[0065] When the user interacts with the game voice, each microphone array receives the external sound signal at the same time, and converts the analog signal into digital form under a common sampling clock to generate a corresponding multi-channel digital audio stream.

[0066] Furthermore, after receiving the digital audio signals output by the four arrays, the system first accurately timestamps each channel (usually using a high-precision hardware timer) to identify the capture moment of each audio sample.

[0067] If there are slight timing deviations in the digital audio signals collected by each channel due to differences in device hardware or transmission delays, the audio data of certain channels will be compensated or aligned through interpolation to correct all channels to the same reference timeline.

[0068] After this step, all channels can accurately reflect the sound environment information at the same time point at each sampling moment.

[0069] Furthermore, the pre-processing operations may include the following:

[0070] DC offset removal: Statistical analysis is performed on the digital audio data of each channel to remove possible baseline drift or DC components, preventing subsequent algorithms from being interfered with by irrelevant offsets.

[0071] Anti-aliasing filtering: Filters digital audio based on the preset target sampling rate and Nyquist frequency range to remove invalid or aliased signals outside the frequency band higher than half the sampling rate;

[0072] Sampling rate conversion (if necessary): In practical applications, to meet the optimal frequency band requirements of subsequent noise reduction or beamforming algorithms, it may be necessary to convert from the initial sampling rate to the predetermined target sampling rate. For example, converting from 48kHz to 16kHz to balance real-time performance and bandwidth limitations;

[0073] Possible software or hardware delays can also be further corrected in this step to ensure that the multi-channel data used by subsequent algorithms are consistent in the time domain.

[0074] Furthermore, after the above processing, the system can obtain a multi-channel audio signal that is completely aligned with the common sampling clock and has been pre-processed, corresponding to each microphone channel in the four microphone arrays.

[0075] The multi-channel audio signal is packaged and stored or transmitted in real time to subsequent speech processing modules (such as frame division and energy value calculation, threshold filtering, beamforming, noise reduction, etc.), providing a reliable data basis for more accurate sound source localization and noise suppression.

[0076] According to the method of the above-mentioned embodiment, this application, based on the multi-microphone array hardware configuration, uses a common sampling clock and timestamp alignment or interpolation to realize the synchronous sampling and digital processing of multi-channel audio signals, and further combines steps such as DC bias removal, anti-aliasing filtering and sampling rate conversion to provide highly consistent and high-quality multi-channel input signals for subsequent beamforming and noise reduction algorithms, thereby effectively improving the acoustic processing accuracy and noise reduction efficiency in game team voice scenarios.

[0077] In some embodiments, dividing the processed multi-channel audio signal into frames and determining the corresponding energy value according to the decibel value of each frame of the audio signal includes:

[0078] The synchronized multi-channel audio signal is divided into frames according to a predetermined frame length and frame shift to obtain a plurality of short-time audio frames;

[0079] performing amplitude or power spectrum analysis on each short-time audio frame to calculate a decibel value of the short-time audio frame;

[0080] Based on the decibel value, the corresponding energy value is determined according to a predetermined mapping rule to provide reference parameters for subsequent adaptive energy threshold filtering and multi-channel noise reduction processing.

[0081] Specifically, a multi-channel audio signal is first acquired after synchronous sampling and digitization. Preprocessing operations such as DC offset removal, anti-aliasing filtering, or sampling rate conversion are completed for each microphone channel, and the time axis of each channel is kept consistent.

[0082] In some examples, the system divides the audio data of each channel into frames with a fixed frame length of 10ms and a frame shift of 5ms, thereby obtaining multiple short-duration audio frames. For example, if the sampling rate is 16kHz, each 10ms frame length corresponds to 160 sampling points; a new frame division begins after a 5ms shift, ensuring that there is some overlap between adjacent frames.

[0083] If necessary, other frame length and frame shift combinations (such as 20ms frame length and 10ms frame shift) can be used to adapt to different application requirements. In actual operation, the number of frames obtained by each channel remains the same, ensuring consistency in subsequent multi-channel fusion or comparison.

[0084] Furthermore, for the obtained short-time audio frame, a short-time Fourier transform (STFT) can be used or statistical analysis can be performed directly in the time domain to obtain a quantitative index representing the signal strength of the frame.

[0085] In this embodiment, one of the following methods is preferably adopted to calculate the strength of the short-time audio frame:

[0086] Root mean square (RMS): The overall amplitude level of the frame signal is obtained by summing the squares of the sampling points within the frame, taking the average, and then taking the square root.

[0087] Power spectral density (PSD): After performing STFT on the frame signal, the signal power distribution is integrated or weighted statistically in the frequency domain to obtain a comprehensive value reflecting the energy of each frequency component.

[0088] The obtained quantitative results such as RMS or PSD are converted into corresponding decibel values ​​(dB).

[0089] In order to facilitate subsequent adaptive energy threshold filtering and multi-channel noise reduction processing, in this embodiment, the decibel value is further mapped or converted into an energy value.

[0090] In some examples, the conversion method can be predefined based on system requirements. For example:

[0091] The decibel value is mapped to the range of [0,1] through a lookup table or linear function to represent the relative energy level of the current frame; or the exponential form of the decibel value is directly retained to obtain a value similar to "short-time power" or "power spectrum integral".

[0092] Compared with decibel values, energy values ​​or power values ​​are often easier to be received and processed by beamforming algorithms, multi-channel Wiener filters, or neural network models; in subsequent steps, the energy values ​​of different frames can be dynamically compared or thresholds can be adaptively set to improve the accuracy of noise filtering.

[0093] Furthermore, decibel and energy values ​​are generated for each microphone channel frame sequence in the same manner as above. If multi-channel fusion is required later, the decibel distribution or energy distribution of each frame in each channel can be established at this stage to provide more reference for the subsequent adaptive threshold determination.

[0094] The average decibel value, variance or maximum value can also be calculated at the multi-channel level to more comprehensively estimate the environmental noise and speech signal strength.

[0095] Through the framing and energy conversion steps of this embodiment, the system can accurately grasp the signal strength of each frame on a short time scale and form an energy value representation that is convenient for subsequent processing.

[0096] When subsequent algorithms (such as threshold filtering or multi-channel noise reduction) analyze these energy values, effective detection of noise intervals and precise positioning of speech intervals can be achieved.

[0097] This can better suppress invalid noise and highlight key voice information in multiplayer gaming scenarios, providing clearer, low-latency voice feedback for real-time communication between players.

[0098] According to the method of the above-mentioned embodiment, in the process of framing and converting decibel values ​​into energy values ​​of multi-channel audio signals, the present application fully considers the impact of short-time Fourier transform, frame length and frame shift selection, and energy mapping on subsequent noise reduction algorithms, so that the implementation method of this part can not only ensure the calculation accuracy, but also lay a solid data foundation for subsequent adaptive threshold filtering and multi-channel noise reduction.

[0099] In some embodiments, filtering or attenuating audio signals having energy values ​​lower than the energy threshold according to the adaptively determined energy threshold to obtain a multi-channel audio signal after preliminary noise reduction includes:

[0100] The energy threshold is adaptively determined based on real-time analysis of the background noise environment or historical energy statistics, wherein the energy threshold is increased when an increase in the ambient noise level is detected, and is decreased when a decrease in the ambient noise level is detected.

[0101] The energy value of each frame of audio signal is compared with the energy threshold. When the energy value is lower than the energy threshold, the current frame of audio signal is attenuated to a preset low amplitude, replaced with a background noise model, or weighted according to a predetermined coefficient.

[0102] The multi-channel audio signal after filtering or attenuation is used as the multi-channel audio signal after preliminary noise reduction.

[0103] Specifically, after completing frame division and energy value calculation for the multi-channel audio signal, the system will dynamically analyze the background noise intensity based on the distribution of energy values ​​in recent frames, historical energy statistics, or environmental noise monitoring data of the current device.

[0104] Based on the analysis results, an energy threshold T is adaptively determined: when the overall noise level is detected to be increasing, the system automatically increases the energy threshold T; when the noisy environment becomes quiet, the energy threshold T is automatically lowered.

[0105] The specific implementation can be based on a predefined update strategy or machine learning algorithm, such as setting a noise detection time window, regularly calculating the average energy level and comparing it with the established noise model, and adjusting the threshold in real time.

[0106] Furthermore, for the short-term audio frames of each microphone channel on the same time axis, the energy value is obtained respectively, and the energy value is compared with the threshold T, where:

[0107] When the energy value is higher than or equal to the threshold T, it is considered as a frame containing speech or strong information components, and usually maintains its original amplitude;

[0108] When the energy value is lower than the threshold T, the system determines that the frame is mainly a noise interval. To prevent invalid noise from affecting subsequent speech enhancement, the frame can be attenuated by performing one or more of the following methods:

[0109] Attenuate to preset low amplitude: reduce the overall audio amplitude of the frame to the set extremely low level;

[0110] Replace with background noise model: If a background noise model matching the current environment has been established, the original frame is replaced with the model;

[0111] Predetermined coefficient weighting processing: The frame signal is attenuated with a fixed or adaptive weighting factor to retain a small amount of noise information and avoid excessive distortion.

[0112] Furthermore, to ensure synchronization across multiple channels, the system simultaneously performs consistent time-segment analysis and processing on multiple channels when attenuating or filtering frames. If the energy value of one channel is significantly above a threshold, while the energy values ​​of other channels are below, the system can either attenuate the noisier channel individually or make a comprehensive assessment based on the average energy of all channels.

[0113] After the aforementioned filtering or attenuation operations, most of the pure noise frames in the multi-channel audio signal are weakened or replaced, resulting in a preliminarily noise-reduced multi-channel audio output. This output can be directly fed into subsequent sound source localization, beamforming, or multi-channel noise reduction processing stages, laying the foundation for more accurate speech enhancement.

[0114] Through the adaptive threshold determination and filtering method of this embodiment, the system can quickly increase the threshold when the ambient noise increases, thereby attenuating more noise frames; when the noise level decreases, the threshold is lowered accordingly, retaining more frames that may contain subtle speech components.

[0115] This embodiment can significantly reduce the interference of low-energy noise in multiplayer game scenarios while retaining relatively clear voice frame information, providing cleaner input for subsequent multi-channel signal processing, and improving the overall voice communication quality and real-time performance.

[0116] In some embodiments, based on the multi-channel audio signal after preliminary noise reduction, the time delay estimation or amplitude difference estimation between each microphone channel is used to perform sound source localization and obtain the target speech direction information, including:

[0117] Performing phase cross-correlation or amplitude difference analysis on the multi-channel audio signal after preliminary noise reduction to determine the arrival time difference or energy difference of the target speech between the microphone channels;

[0118] Combined with the geometric distribution information of the microphone array, the time difference or energy difference is matched and calculated with the array layout parameters to obtain the direction or coordinate information of the target speech;

[0119] The direction or coordinate information is used as the target speech direction information for subsequent beamforming and multi-channel noise reduction processing.

[0120] Specifically, after completing energy threshold filtering or attenuation of the original audio signal, a relatively clean multi-channel audio signal with reduced noise components (hereinafter referred to as "the multi-channel audio signal after preliminary noise reduction") can be obtained.

[0121] This signal still maintains the time synchronization relationship corresponding to each microphone channel and can be directly used for further delay estimation or amplitude difference analysis.

[0122] Furthermore, the system processes the multi-channel audio signal after preliminary noise reduction at the frame level, preferably using the generalized cross-correlation phase transform (GCC-PHAT) algorithm to measure the time difference of arrival (TDoA) of the same speech frame at different microphone channels.

[0123] If the noise environment or microphone layout is suitable for estimation based on amplitude difference, the signal amplitudes between multiple channels can be directly compared to infer the energy ratio of the sound source between each channel.

[0124] The obtained time difference or amplitude difference is the basic parameter for subsequent calculation of the direction and position of the sound source, avoiding the influence of noise interference on the accuracy.

[0125] Furthermore, in this embodiment, the geometric distribution information of the microphone array includes the relative position and spacing of each microphone and the overall shape of the array (such as a linear array, a circular array, etc.), which are all recorded in the array parameter configuration when the device leaves the factory.

[0126] By matching the time difference / energy difference obtained by GCC-PHAT or amplitude difference analysis with the geometric distribution information, the direction or coordinates of the target speech in space can be calculated through various positioning algorithms (such as triangulation positioning or least squares method). For example, the direction information can be expressed in the form of azimuth, pitch angle, etc.; the coordinate information can be used to obtain a more accurate sound source location in a two-dimensional or three-dimensional coordinate system.

[0127] Furthermore, the system outputs the target speech direction information (or sound source coordinates) based on the above matching calculation results as a reference for subsequent signal processing modules.

[0128] In gaming devices, if the microphone array moves with the player's head or terminal, updated positioning results for different time periods can be obtained in real time, ensuring that the beamforming and multi-channel noise reduction modules can always enhance the correct target direction.

[0129] This directional information can be subsequently read or called to perform beamforming weight setting, noise suppression strategy adjustment, etc., to further improve the clarity of the target speech.

[0130] This embodiment uses this sound source localization embodiment to accurately determine the direction or location of the player's voice. Even if a small amount of noise remains after initial noise reduction, the positioning accuracy can be significantly improved by leveraging multi-channel information and array geometry.

[0131] Accurate sound source localization results can significantly reduce noise interference in game team voice scenarios, creating better signal conditions for subsequent beamforming and multi-channel noise reduction, making voice communication between players clearer and more immersive.

[0132] In some embodiments, beamforming and multi-channel noise reduction processing are performed on the multi-channel audio signal after preliminary noise reduction based on the target speech direction information to suppress background noise and enhance the target speech, including:

[0133] Based on the target speech direction information, a beamforming coefficient is determined for enhancing the target direction signal, and the multi-channel audio signal after preliminary noise reduction is input into the beamforming module. In the beamforming module, a beamforming output signal is generated based on the beamforming coefficient, in which the target direction is enhanced and noise in other directions is suppressed;

[0134] Attenuating residual noise using a multi-channel noise reduction algorithm based on the beamforming output signal, wherein the multi-channel noise reduction algorithm includes at least one of the following methods: a statistical model-based filtering method, a deep learning-based speech enhancement method, or a minimum variance distortionless response filtering method;

[0135] The output signal after beamforming and multi-channel noise reduction processing is used as the enhanced speech signal for subsequent voice activity detection and gain control.

[0136] Specifically, the system first determines the azimuth and elevation angles, or specific 2D / 3D coordinates, of the target speech based on analysis of the multi-channel audio signals after preliminary noise reduction (e.g., time difference information obtained using the GCC-PHAT algorithm and information about the microphone array geometry). This target speech direction information serves as an important input parameter for beamforming and multi-channel noise reduction in this embodiment.

[0137] Furthermore, the system assigns corresponding beamforming weights or filter coefficients to each microphone channel based on the target speech direction information.

[0138] If adaptive filtering is used, the beamforming module will dynamically adjust the filtering weights based on the spectral characteristics and signal-to-noise ratio of the real-time noise environment, thereby achieving optimal signal enhancement in the direction of the target speech. If a fixed filter design is used, the directional filter coefficients are pre-set when the device leaves the factory or during the application initialization phase.

[0139] By configuring the above coefficients, the speech signal from the target direction can be enhanced in the spatial domain while suppressing background noise from other directions.

[0140] Furthermore, the multi-channel audio signal after preliminary noise reduction is input into the beamforming module, and signal synthesis and weighting processing are performed according to the calculated beamforming coefficients.

[0141] In the beamforming output signal, the target speech is spatially highlighted, while noise from other directions is partially attenuated, thereby improving the accuracy and efficiency of subsequent noise reduction algorithms.

[0142] Furthermore, based on the beamforming output, the residual noise is more finely attenuated to further improve voice quality. Depending on the application requirements, one or more of the following noise reduction methods can be flexibly selected:

[0143] MVDR filtering: minimum variance distortionless response filter, which preserves the integrity of the target speech by minimizing the output power of signals outside the target direction;

[0144] Multi-channel Wiener filtering: Statistically estimate the noise components in the frequency domain, combine the prior speech model and noise model, calculate the Wiener filter weights, and reduce the noise components;

[0145] Neural network noise reduction: Input the amplitude spectrum, phase spectrum and other features of multi-channel audio into a deep learning network (such as a time-frequency domain convolutional network, RNN or other speech enhancement models) to restore speech and output clearer speech frames.

[0146] Multi-channel input helps to make full use of the complementary information of time difference and amplitude difference between channels, and achieve better speech enhancement effect even in a noisy background environment.

[0147] Finally, the enhanced speech signal obtained through beamforming and multi-channel noise reduction is output. This output not only contains a more prominent and cleaner target speech, but also retains the main features and clarity of the original speech.

[0148] Further processing such as voice activation detection and gain control can be performed subsequently, so that users can obtain clear and high-quality real-time voice communication experience in game team scenarios.

[0149] By combining target speech direction information with a multi-channel noise reduction algorithm, this embodiment significantly reduces background noise interference and effectively improves the clarity and intelligibility of speech signals. Compared to single-channel noise reduction, the multi-channel solution described in this embodiment performs better in complex noise environments and enables the real-time communication required for gaming scenarios without significantly increasing processing latency.

[0150] In some embodiments, performing voice activation detection and gain control on a voice signal obtained after beamforming and multi-channel noise reduction processing, and outputting the optimized voice signal to a game communication module, includes:

[0151] Perform voice activation detection on the speech signal obtained after beamforming and multi-channel noise reduction processing to determine whether the current frame is a speech frame;

[0152] Maintaining the strength of the speech signal detected as a speech frame, and performing skip coding, attenuation, or replacing with a background noise model on the speech signal detected as a non-speech frame;

[0153] After completing voice activity detection, automatic gain control is performed on the speech frames to maintain a stable and consistent target speech volume in different noise environments and after beamforming processing;

[0154] The voice signal processed by voice activation detection and gain control is used as the optimized audio output and sent to the game communication module for real-time voice communication.

[0155] Specifically, after beamforming and multi-channel noise reduction are completed on the multi-channel audio signal, a relatively clean speech output is obtained. At this time, the system performs frame-level analysis on this output and uses the VAD algorithm to determine whether the current frame contains valid speech components.

[0156] Furthermore, if the frame is determined to be a speech frame, it is retained in the subsequent process; if it is determined to be a pure noise frame, the system may perform one or more of the following operations:

[0157] Skip coding: The frame is considered as invalid information and is not subjected to speech coding or transmission;

[0158] Attenuation processing: Reduce the signal amplitude within the frame to an extremely low level to prevent noise interference;

[0159] Replace with background noise model: If the system has established an environmental background noise model, you can use this model to replace the original signal to make the overall listening experience smoother.

[0160] Furthermore, after completing VAD, the retained speech frames may be affected by factors such as beamforming and dynamic noise suppression, exhibiting different amplitudes or energy distributions. To maintain a consistent and stable listening experience in the game scene, automatic gain control needs to be performed on the speech frames. For example, this may include the following operations:

[0161] Real-time detection of speech amplitude: By calculating the root mean square value (RMS) or power level of the current frame, it is determined whether the frame signal is too strong or too weak;

[0162] Adaptive gain adjustment: Based on the detection results, the gain factor is increased or decreased to keep the voice volume within a reasonable range in different environments or processing stages;

[0163] Smoothing and normalization: When processing multiple frames continuously, the gain adjustment value can be smoothed to prevent sudden changes in the user's hearing experience.

[0164] After voice activation detection and gain control, the system packages the optimized voice signal and outputs it to the game communication module. This module typically performs speech encoding (such as Opus or AAC) and sends the encoded audio data to the game server or other player terminals.

[0165] In non-real-time demand scenarios, the gain-processed voice data can also be used for local playback, cloud recording, or subsequent statistical analysis, ensuring that users can trace back and reproduce high-quality audio content.

[0166] This embodiment combines VAD and AGC to effectively reduce the transmission and bandwidth consumption of useless noise frames while maintaining a stable volume for real voice frames in various environments, significantly improving the clarity and comfort of voice communication between gamers. When used in noisy environments or with multiple people talking simultaneously, it further conserves system resources while preserving critical voice information, ensuring real-time and smooth in-game communication.

[0167] The following are device embodiments of the present application, which can be used to implement the method embodiments of the present application. For details not disclosed in the device embodiments of the present application, please refer to the method embodiments of the present application.

[0168] Figure 2 This is a schematic diagram of the structure of the game team voice noise reduction device based on multi-channel signal processing provided by the embodiment of the present application. Figure 2 As shown, the game team voice noise reduction device based on multi-channel signal processing includes:

[0169] An acquisition module 201 is configured to acquire multi-channel audio signals using a plurality of microphone arrays provided on a gaming terminal, and to perform synchronous sampling and digital processing on the multi-channel audio signals;

[0170] A division module 202 is configured to divide the processed multi-channel audio signal into frames and determine a corresponding energy value according to the decibel value of each frame of the audio signal;

[0171] The noise reduction module 203 is configured to filter or attenuate audio signals having energy values ​​lower than the energy threshold according to the adaptively determined energy threshold, thereby obtaining a multi-channel audio signal after preliminary noise reduction.

[0172] The localization module 204 is configured to perform sound source localization based on the multi-channel audio signal after preliminary noise reduction by using the time delay estimation or amplitude difference estimation between the microphone channels to obtain the target speech direction information;

[0173] Suppression module 205, configured to perform beamforming and multi-channel noise reduction processing on the multi-channel audio signal after preliminary noise reduction according to the target speech direction information, so as to suppress background noise and enhance the target speech;

[0174] The output module 206 is used to perform voice activation detection and gain control on the voice signal obtained after beamforming and multi-channel noise reduction processing, and output the optimized voice signal to the game communication module.

[0175] In some embodiments, Figure 2 The acquisition module 201 uses a common sampling clock to perform analog-to-digital conversion on the analog audio signals output by the multiple microphone arrays; performs time stamp alignment or interpolation processing on the converted multi-channel digital audio signals to make each channel consistent on the same time axis; performs at least one preprocessing operation of removing DC bias, anti-aliasing filtering or sampling rate conversion on the aligned multi-channel digital audio signals to obtain a processed multi-channel audio signal.

[0176] In some embodiments, Figure 2 The division module 202 divides the synchronized multi-channel audio signal into frames according to a predetermined frame length and frame shift to obtain multiple short-time audio frames; performs amplitude or power spectrum analysis on each short-time audio frame to calculate the decibel value of the short-time audio frame; and determines the corresponding energy value based on the decibel value according to a predetermined mapping rule to provide reference parameters for subsequent adaptive energy threshold filtering and multi-channel noise reduction processing.

[0177] In some embodiments, Figure 2 The noise reduction module 203 adaptively determines an energy threshold based on real-time analysis of the background noise environment or historical energy statistics, wherein the energy threshold is increased when an increase in the ambient noise level is detected, and the energy threshold is lowered when a decrease in the ambient noise level is detected; the energy value of each frame of the audio signal is compared with the energy threshold, and when the energy value is lower than the energy threshold, the current frame of the audio signal is attenuated to a preset low amplitude, replaced with a background noise model, or weighted according to a predetermined coefficient; and the multi-channel audio signal after filtering or attenuation processing is used as the multi-channel audio signal after preliminary noise reduction.

[0178] In some embodiments, Figure 2The positioning module 204 performs phase cross-correlation or amplitude difference analysis on the multi-channel audio signal after preliminary noise reduction to determine the arrival time difference or energy difference of the target speech between the microphone channels; combines the geometric distribution information of the microphone array, matches the time difference or energy difference with the array layout parameters, and obtains the direction or coordinate information of the target speech; and uses the direction or coordinate information as the target speech direction information for subsequent beamforming and multi-channel noise reduction processing.

[0179] In some embodiments, Figure 2 The suppression module 205 determines the beamforming coefficient for enhancing the target direction signal based on the target speech direction information, and inputs the multi-channel audio signal after preliminary noise reduction into the beamforming module, and generates a beamforming output signal in which the target direction is enhanced and the noise in other directions is suppressed according to the beamforming coefficient in the beamforming module; based on the beamforming output signal, the residual noise is weakened by using a multi-channel noise reduction algorithm, wherein the multi-channel noise reduction algorithm includes at least one of the following methods: a filtering method based on a statistical model, a speech enhancement method based on deep learning, or a minimum variance distortionless response filtering method; the output signal after beamforming and multi-channel noise reduction processing is used as the enhanced speech signal for subsequent speech activation detection and gain control.

[0180] In some embodiments, Figure 2 The output module 206 maintains the strength of the speech signal detected as a speech frame, and performs skip coding, attenuation, or replacement with a background noise model for the speech signal detected as a non-speech frame; after completing voice activation detection, automatic gain control is performed on the speech frame to maintain the target speech volume stable and consistent in different noise environments and after beamforming processing; the speech signal after voice activation detection and gain control processing is used as the optimized audio output and sent to the game communication module for real-time voice communication.

[0181] It should be understood that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.

[0182] Figure 3 Schematic diagram of the structure of the electronic device 3 provided in the embodiment of the present application. Figure 3 As shown, the electronic device 3 of this embodiment includes: a processor 301, a memory 302, and a computer program 303 stored in the memory 302 and executable on the processor 301. When the processor 301 executes the computer program 303, the steps of the above-mentioned method embodiments are implemented. Alternatively, when the processor 301 executes the computer program 303, the functions of the modules / units in the above-mentioned device embodiments are implemented.

[0183] For example, computer program 303 may be divided into one or more modules / units, which are stored in memory 302 and executed by processor 301 to implement the present application. One or more modules / units may be a series of computer program instruction segments capable of performing specific functions, and the instruction segments are used to describe the execution process of computer program 303 in electronic device 3.

[0184] The electronic device 3 may be a desktop computer, a notebook, a PDA, a cloud server or other electronic device. The electronic device 3 may include but is not limited to a processor 301 and a memory 302. Those skilled in the art will understand that Figure 3 It is only an example of electronic device 3 and does not constitute a limitation of electronic device 3. It may include more or fewer components than shown in the figure, or a combination of certain components, or different components. For example, the electronic device may also include input and output devices, network access devices, buses, etc.

[0185] The processor 301 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor.

[0186] Memory 302 can be an internal storage unit of electronic device 3, such as a hard drive or memory of electronic device 3. Memory 302 can also be an external storage device of electronic device 3, such as a plug-in hard drive, a Smart Media Card (SMC), a Secure Digital (SD) card, a flash memory card, etc. Furthermore, memory 302 can include both an internal storage unit of electronic device 3 and an external storage device. Memory 302 is used to store computer programs and other programs and data required by the electronic device. Memory 302 can also be used to temporarily store data that has been output or is about to be output.

[0187] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the division of the above-mentioned functional units and modules is used as an example for illustration. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiment can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of software functional units. In addition, the specific names of the functional units and modules are only for the convenience of distinguishing each other, and are not used to limit the scope of protection of this application. The specific working process of the units and modules in the above-mentioned system can refer to the corresponding process in the aforementioned method embodiment, and will not be repeated here.

[0188] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described or recorded in detail in a certain embodiment, reference can be made to the relevant description of other embodiments.

[0189] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0190] In the embodiments provided in this application, it should be understood that the disclosed apparatus / computer equipment and methods can be implemented in other ways. For example, the apparatus / computer equipment embodiments described above are merely schematic. For example, the division of modules or units is merely a logical function division. In actual implementation, there may be other division methods. Multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed may be through some interfaces, indirect coupling or communication connection of the apparatus or unit, which may be electrical, mechanical or other forms.

[0191] Units described as separate components may or may not be physically separate, and components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0192] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.

[0193] If the integrated module / unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the present application can implement all or part of the processes in the above-mentioned embodiment method by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, it can implement the steps of each of the above-mentioned method embodiments. The computer program may include computer program code, which may be in source code form, object code form, executable file, or some intermediate form. Computer-readable media may include: any entity or device capable of carrying computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium.

[0194] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present application, and should all be included in the scope of protection of the present application.

Claims

1. A method for reducing noise in game team voice based on multi-channel signal processing, characterized in that: include: Acquire multi-channel audio signals using a plurality of microphone arrays provided on the gaming terminal, and synchronously sample and digitize the multi-channel audio signals; Divide the processed multi-channel audio signal into frames, and determine the corresponding energy value according to the decibel value of each frame of the audio signal; filtering or attenuating the audio signal having an energy value lower than the energy threshold according to the adaptively determined energy threshold to obtain a multi-channel audio signal after preliminary noise reduction; Based on the multi-channel audio signal after the preliminary noise reduction, the sound source is localized by using the time delay estimation or amplitude difference estimation between the microphone channels to obtain the target speech direction information; Performing beamforming and multi-channel noise reduction processing on the multi-channel audio signal after the preliminary noise reduction according to the target speech direction information to suppress background noise and enhance the target speech; Perform voice activation detection and gain control on the voice signal obtained after beamforming and multi-channel noise reduction processing, and output the optimized voice signal to the game communication module; The method of performing sound source localization based on the multi-channel audio signal after preliminary noise reduction by using time delay estimation or amplitude difference estimation between microphone channels to obtain target speech direction information includes: Performing phase cross-correlation or amplitude difference analysis on the multi-channel audio signal after the preliminary noise reduction to determine the arrival time difference or energy difference of the target speech between the microphone channels; Combined with the geometric distribution information of the microphone array, the time difference or energy difference is matched and calculated with the array layout parameters to obtain the direction or coordinate information of the target speech; The direction or coordinate information is used as target speech direction information for subsequent beamforming and multi-channel noise reduction processing.

2. The method according to claim 1, characterized in that The synchronous sampling and digital processing of the multi-channel audio signal includes: Performing analog-to-digital conversion on analog audio signals output by the plurality of microphone arrays using a common sampling clock; Performing time stamp alignment or interpolation processing on the converted multi-channel digital audio signal to make each channel consistent on the same time axis; At least one pre-processing operation of direct current offset removal, anti-aliasing filtering, or sampling rate conversion is performed on the aligned multi-channel digital audio signal to obtain a processed multi-channel audio signal.

3. The method according to claim 1, characterized in that The step of dividing the processed multi-channel audio signal into frames and determining a corresponding energy value according to a decibel value of each frame of the audio signal includes: The synchronized multi-channel audio signal is divided into frames according to a predetermined frame length and frame shift to obtain a plurality of short-time audio frames; Performing amplitude or power spectrum analysis on each short-time audio frame to calculate a decibel value of the short-time audio frame; Based on the decibel value, a corresponding energy value is determined according to a predetermined mapping rule, so as to provide a reference parameter for subsequent adaptive energy threshold filtering and multi-channel noise reduction processing.

4. The method according to claim 1, wherein The step of filtering or attenuating the audio signal having an energy value lower than the energy threshold according to the adaptively determined energy threshold to obtain a multi-channel audio signal after preliminary noise reduction includes: Adaptively determining the energy threshold based on real-time analysis of the background noise environment or historical energy statistics, wherein the energy threshold is increased when an increase in the ambient noise level is detected, and is decreased when a decrease in the ambient noise level is detected; comparing the energy value of each frame of audio signal with the energy threshold, and when the energy value is lower than the energy threshold, performing attenuation on the current frame of audio signal to a preset low amplitude, replacing it with a background noise model, or performing weighted processing according to a predetermined coefficient; The multi-channel audio signal after the filtering or attenuation processing is used as the multi-channel audio signal after the preliminary noise reduction.

5. The method according to claim 1, wherein The method further comprises performing beamforming and multi-channel noise reduction processing on the multi-channel audio signal after the preliminary noise reduction according to the target voice direction information to suppress background noise and enhance the target voice, including: Determining beamforming coefficients for enhancing a target direction signal based on the target speech direction information, and inputting the multi-channel audio signal after preliminary noise reduction into a beamforming module, wherein the beamforming module generates a beamforming output signal in which the target direction is enhanced and noise in other directions is suppressed according to the beamforming coefficients; Attenuating residual noise using a multi-channel noise reduction algorithm based on the beamforming output signal, wherein the multi-channel noise reduction algorithm includes at least one of the following methods: a statistical model-based filtering method, a deep learning-based speech enhancement method, or a minimum variance distortionless response filtering method; The output signal after the beamforming and multi-channel noise reduction processing is used as the enhanced speech signal for subsequent speech activity detection and gain control.

6. The method according to claim 1, characterized in that The method of performing voice activation detection and gain control on the voice signal obtained after beamforming and multi-channel noise reduction processing, and outputting the optimized voice signal to the game communication module, includes: Perform voice activation detection on the speech signal obtained after beamforming and multi-channel noise reduction processing to determine whether the current frame is a speech frame; Maintaining the strength of the speech signal detected as a speech frame, and performing skip coding, attenuation, or replacing with a background noise model on the speech signal detected as a non-speech frame; After completing voice activity detection, automatic gain control is performed on the speech frames to maintain a stable and consistent target speech volume in different noise environments and after beamforming processing; The voice signal processed by voice activation detection and gain control is used as the optimized audio output and sent to the game communication module for real-time voice communication.

7. A game team voice noise reduction device based on multi-channel signal processing, characterized in that: include: An acquisition module, configured to acquire multi-channel audio signals using a plurality of microphone arrays provided on the gaming terminal, and to synchronously sample and digitize the multi-channel audio signals; A division module is used to divide the processed multi-channel audio signal into frames and determine the corresponding energy value according to the decibel value of each frame of the audio signal; a noise reduction module, configured to filter or attenuate audio signals having energy values ​​lower than the energy threshold according to an adaptively determined energy threshold, to obtain a multi-channel audio signal after preliminary noise reduction; A positioning module is used to perform sound source localization based on the multi-channel audio signal after preliminary noise reduction by using the time delay estimation or amplitude difference estimation between each microphone channel to obtain the target voice direction information; a suppression module, configured to perform beamforming and multi-channel noise reduction processing on the multi-channel audio signal after the preliminary noise reduction according to the target voice direction information, so as to suppress background noise and enhance the target voice; The output module is used to perform voice activation detection and gain control on the voice signal obtained after beamforming and multi-channel noise reduction processing, and output the optimized voice signal to the game communication module; The positioning module is used to perform phase cross-correlation or amplitude difference analysis on the multi-channel audio signal after the preliminary noise reduction to determine the arrival time difference or energy difference of the target speech between each microphone channel; Combined with the geometric distribution information of the microphone array, the time difference or energy difference is matched and calculated with the array layout parameters to obtain the direction or coordinate information of the target speech; the direction or coordinate information is used as the target speech direction information for subsequent beamforming and multi-channel noise reduction processing.

8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 6 are implemented.

9. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.

Citation Information

Patent Citations

  • Multi-microphone array beamforming signal enhancement method and device

    CN119811408A