Communication noise reduction method of industrial telephone set and industrial telephone set

By determining the direction parameters of the sound source and adaptive noise reduction processing, the problem of poor call quality in high-noise environments of industrial telephones is solved, and efficient communication under complex background noise is achieved.

CN120475095APending Publication Date: 2025-08-12J&R TECHNOLOGY LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510824209.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-18
Publication Date
2025-08-12

AI Technical Summary

Technical Problem

Existing industrial telephones are difficult to effectively suppress complex background noise in high noise environments, resulting in poor call quality and affecting information exchange efficiency and security response capabilities.

Method used

By obtaining user communication audio data to determine the sound source direction parameters, constructing microphone array configuration parameters, and adaptive noise reduction processing is performed in combination with the noise type of ambient audio data, including beamforming and filtering strategy optimization, dynamically adjusting the reception direction and gain of the microphone array to enhance voice signals and suppress noise.

Benefits of technology

Effectively weaken background interference in high noise environments, improve call quality and communication reliability, ensure the clarity and stability of communication content, and avoid misinformation and low-quality calls.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120475095A_ABST
    Figure CN120475095A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of communication equipment, in particular to a communication noise reduction method for an industrial telephone and the industrial telephone, and the method comprises the steps: obtaining environment audio data and user communication audio data; determining a user sound source direction parameter according to the user communication audio data; constructing configuration parameters of a microphone array based on the user sound source direction parameters; acquiring first enhanced voice data in the sound source direction of the user according to the configuration parameters; determining a noise type corresponding to the environment audio data according to the environment audio data; and based on the noise type, performing noise reduction processing on the first enhanced voice data to obtain noise-reduced audio data. According to the invention, complex background noise can be effectively suppressed in a high-noise industrial environment, so that the reliability and practicability of communication are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of communication equipment, and in particular to a communication noise reduction method for an industrial telephone and the industrial telephone. Background Art

[0002] Currently, most industrial telephones primarily reduce ambient noise interference through physical isolation and basic noise filtering algorithms, such as by adding soundproofing covers, shock-absorbing layers, or sealed structures. However, due to the high continuity, complex frequencies, and variable directions of high-intensity industrial noise, a single acoustic structure or static filtering parameters often cannot effectively address this complex background noise. Consequently, actual calls still suffer from issues such as fuzzy voices and significant noise interference, impacting the efficiency of information exchange and safety response capabilities of on-site personnel. Therefore, there is room for improvement. Summary of the Invention

[0003] The present application provides a communication noise reduction method for an industrial telephone and an industrial telephone, which can effectively suppress complex background noise in a high-noise industrial environment, thereby improving the reliability and practicality of communication.

[0004] In the first aspect, the above-mentioned invention object of the present application is achieved through the following technical solutions:

[0005] A method for reducing communication noise of an industrial telephone, the method comprising:

[0006] Obtain environmental audio data and user communication audio data;

[0007] Determining a user sound source direction parameter based on the user communication audio data;

[0008] Constructing configuration parameters of a microphone array based on the user sound source direction parameters;

[0009] Acquire first enhanced voice data in the direction of the user's voice source according to the configuration parameters;

[0010] Determining, based on the ambient audio data, a noise type corresponding to the ambient audio data;

[0011] Based on the noise type, noise reduction processing is performed on the first enhanced speech data to obtain noise-reduced audio data.

[0012] By adopting the above technical solution, the user sound source direction parameters are obtained and determined based on the user communication audio data, and the spatial position information of the user communication audio relative to the microphone array can be obtained, so that the microphone array configuration parameters are constructed based on the sound source direction, and the receiving direction and receiving gain of the microphone array can be adjusted to achieve spatial focus enhancement of the user communication audio data, and then the first enhanced voice data is subjected to noise reduction processing through the environmental audio data and the corresponding noise type, and the appropriate noise reduction strategy can be adaptively selected, so that different types of background interference can be effectively weakened in a high-noise industrial environment, the call quality can be improved, and the practicality and reliability of industrial telephone communications can be improved.

[0013] In a preferred example, the present application may be further configured as follows: determining the user speaking direction parameter based on the user communication audio data specifically includes:

[0014] Acquire multiple directional channel signals from the user communication audio data;

[0015] Based on a direction estimation algorithm, the multiple directional channel signals are analyzed to calculate the user sound source direction parameter.

[0016] By adopting the above technical solution, by obtaining multiple directional channel signals from the user communication audio data and analyzing the directional channel signals based on the direction estimation algorithm, it is possible to infer the incident angle of the user sound source relative to the array, thereby calculating the user sound source direction parameter.

[0017] In a preferred example, the present application may be further configured as follows: constructing the configuration parameters of the microphone array based on the user sound source direction parameter specifically includes:

[0018] Determining a main beam direction angle of the microphone array based on the user sound source direction parameter;

[0019] Determining a stable state of user communication based on the user audio data, and dynamically determining a beam width of the microphone array according to the stable state of user communication;

[0020] Configuration parameters of the microphone array are generated according to the main beam direction angle and beam width.

[0021] By adopting the above technical solution, the configuration parameters of the microphone array are constructed based on the user's sound source direction parameters, including the setting of the main beam direction angle and the dynamic adjustment of the beam width. This dynamic adjustment is performed according to the stable state of the user's communication, allowing the microphone array to adapt to the communication scenario more flexibly and ensure the best voice capture effect.

[0022] In a preferred example, the present application may be further configured as follows: determining the noise type corresponding to the ambient audio data according to the ambient audio data specifically includes:

[0023] Performing spectrum analysis on the ambient audio data by using Fourier transform to obtain corresponding spectrum energy distribution characteristic information;

[0024] The spectrum energy distribution characteristic information is input into a preset noise classification model to obtain the noise type corresponding to the environmental audio data.

[0025] By adopting the above technical solution and performing Fourier transform spectrum analysis on the ambient audio data, the spectrum energy distribution characteristic information can be extracted, and the characteristic information is input into the preset noise classification model, which can accurately determine the noise type corresponding to the ambient audio data and realize the discrimination of the noise type.

[0026] In a preferred example, the present application may be further configured as follows: performing noise reduction processing on the first enhanced speech data based on the noise type to obtain noise-reduced audio data specifically includes:

[0027] Based on the noise type, obtaining corresponding filtering strategy parameters from a preset filtering strategy mapping table;

[0028] The first enhanced speech data is subjected to noise reduction processing according to the filtering strategy parameters to obtain the noise-reduced audio data.

[0029] By adopting the above technical solution, the filtering strategy parameters corresponding to the noise type are obtained from the preset filtering strategy mapping table, and then the first enhanced voice data is subjected to noise reduction processing according to the filtering strategy parameters. It is possible to select appropriate algorithms and parameters for different types of noise for optimization processing, thereby improving the adaptability to noise changes in complex industrial environments, thereby effectively improving the communication quality.

[0030] In a preferred example, the present application can be further configured as follows: the filtering strategy parameters include noise reduction processing methods corresponding to different noise types, specifically including:

[0031] When the noise type is steady-state noise, using a Wiener filtering algorithm to perform frequency domain smoothing processing on the first enhanced speech data;

[0032] When the noise type is impulse noise, performing spectrum subband energy adjustment on the first enhanced speech data using spectral subtraction;

[0033] When the noise type is mixed noise, a combination of adaptive bandpass filtering and gating strategy is used to perform segmented filtering processing on the first enhanced speech data.

[0034] By adopting the above technical solution and providing different noise reduction processing methods according to different noise types, the best noise reduction effect can be achieved in a highly complex noise environment, and high-quality audio can be stably output.

[0035] In a preferred example, the present application may be further configured as follows: after obtaining the configuration parameters, adjusting the beam direction and beam width of the microphone array;

[0036] The microphone array is controlled to perform beamforming processing, weighted processing is performed on the multiple directional channel signals, and the first enhanced voice data is output.

[0037] By adopting the above technical solution, after obtaining the configuration parameters of the microphone array, by adjusting the beam direction and beam width, and controlling the microphone array to perform beamforming processing, multiple directional channel signals can be effectively weighted, and the first enhanced voice data can be output, which can suppress non-target directional noise, thereby improving the signal-to-noise ratio of the target audio and enhancing the overall call effect.

[0038] In a preferred example, the present application may be further configured as follows: after obtaining the noise-reduced audio data, performing a signal-to-noise ratio test on the noise-reduced audio data, and determining whether the noise-reduced audio data meets an output condition based on the signal-to-noise ratio test result;

[0039] If the signal-to-noise ratio detection result is up to standard, it is determined that the noise-reduced audio data meets the output condition, and the noise-reduced audio data is sent to the target communication device.

[0040] By adopting the above technical solution, by performing signal-to-noise ratio detection after noise reduction processing to evaluate the quality of the noise-reduced audio data, it is possible to ensure that only noise-reduced data that meets the conditions is sent to the target communication device, thereby ensuring the clarity and stability of the communication content and avoiding mistransmission and low-quality calls.

[0041] In a preferred example, the present application may be further configured as follows: if the signal-to-noise ratio detection result is not up to standard, determining that the noise reduction audio data does not meet the output condition;

[0042] When the signal-to-noise ratio detection result is not up to standard, optimizing the filtering strategy parameters corresponding to the noise type until the optimized filtering strategy parameters enable the corresponding noise reduction audio data to meet the output conditions;

[0043] After the noise reduction audio data meets the output condition, the filtering strategy mapping table is updated according to the optimized filtering strategy parameters.

[0044] By adopting the above technical solution, if the signal-to-noise ratio detection result does not meet the standard, it is convenient to further optimize the filtering strategy parameters corresponding to the noise type to improve the noise reduction effect. The optimized filtering strategy parameters will be updated to the filtering strategy mapping table, thereby improving the adaptability to different communication environments.

[0045] Secondly, the above-mentioned purpose of this application is achieved through the following technical solutions:

[0046] A computer device includes a microphone array, an audio acquisition module, a communication sending module, a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the communication noise reduction method are implemented.

[0047] In summary, this application includes at least one of the following beneficial technical effects:

[0048] 1. Obtain and determine the user sound source direction parameters based on the user communication audio data, and be able to obtain the spatial position information of the user communication audio relative to the microphone array, thereby constructing the microphone array configuration parameters based on the sound source direction, and being able to adjust the receiving direction and receiving gain of the microphone array, thereby achieving spatial focus enhancement of the user communication audio data, and then performing noise reduction processing on the first enhanced voice data through the environmental audio data and the corresponding noise type, and being able to adaptively select the appropriate noise reduction strategy, thereby effectively weakening different types of background interference in a high-noise industrial environment, improving call quality, and thus improving the practicality and reliability of industrial telephone communications;

[0049] 2. By performing signal-to-noise ratio testing after noise reduction processing to evaluate the quality of the noise-reduced audio data, it is possible to ensure that only noise-reduced data that meets the requirements is sent to the target communication device, thereby ensuring the clarity and stability of the communication content and avoiding mistransmission and low-quality calls. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] Figure 1 This is a flow chart of an implementation method for reducing communication noise of an industrial telephone in one embodiment of the present application;

[0051] Figure 2 It is a structural diagram of an industrial telephone in one embodiment of the present application. DETAILED DESCRIPTION

[0052] The present application is further described in detail below with reference to the accompanying drawings.

[0053] In one embodiment, the present application discloses a method for reducing communication noise of an industrial telephone, which specifically includes the following steps:

[0054] S10: Acquire environmental audio data and user communication audio data.

[0055] Specifically, the original audio input signal is received from each channel in the microphone array, and each analog audio signal is sampled and analog-to-digital converted to obtain digital ambient audio data and user communication audio data. For example, the audio signal of each channel in the microphone array is sampled at a sampling rate of 16kHz, and short-time frame processing is performed in a 20ms frame length and 10ms frame shift manner. The short-time energy of each frame signal is calculated, and the short-time energy is obtained by dividing the sum of the squares of the frame samples by the frame length. If the short-time energy in consecutive frames is greater than the set voice threshold (such as -35dB) and lasts for more than 200ms, it is determined to be user communication audio data, otherwise it is ambient audio data. For example, if the energy of the 10th to 25th frames is continuously higher than the threshold, it is marked as a voice segment, and the 1st to 9th frames have low energy and are marked as background segments. The final output data includes two parts: user communication audio data and ambient audio data.

[0056] S20: Determine the user sound source direction parameter according to the user communication audio data.

[0057] Specifically, step S20 includes:

[0058] S21: Acquire multiple directional channel signals from user communication audio data.

[0059] Specifically, the audio frames of all channels within the user communication time period are synchronously extracted and combined into multi-channel audio frame groups. Each group contains different channel data in the same time period to form a directional channel signal set. For example, if a certain segment contains 8 microphone channels, an 8×N matrix is output, and each row is a continuous speech frame sequence of a channel, thereby obtaining multiple directional channel signals.

[0060] S22: Based on the direction estimation algorithm, multiple directional channel signals are analyzed to calculate the user sound source direction parameter.

[0061] Specifically, the direction estimation algorithm uses the generalized cross-correlation-phase transform (GCC-PHAT) algorithm to extract the time delay of multiple directional channel signals, calculate the time delay difference between channels through the cross-correlation algorithm, and infer the incident angle of the sound source to the center of the array based on the microphone spacing and sound speed in the array. For example, if the time delays between channels 1, 2, and 3 are 0μs, 60μs, and 118μs, respectively, the microphone spacing is 2cm, and the sound speed is 340m / s, then a single-channel delay of 60μs corresponds to a sound source at 90° to the array. Based on the three-channel incremental relationship matching template, it can be determined that the current sound source comes from the positive lateral direction of the array on the right side, that is, the user sound source direction is 90°, and this angle is output as the user sound source direction parameter.

[0062] For example, the sound source direction can be obtained by matching the templates in the three-channel incremental relationship shown in Table 1, as shown in Table 1:

[0063] Sound source angle Ideal delay sequence (channel 1 → channel 2 → channel 3) 0° (directly ahead) 0μs, 0μs, 0μs (no delay) 30° right front 0μs, 29.4μs, 58.8μs 45° right 0μs, 41.6μs, 83.3μs 60° right posterior 0μs, 50μs, 100μs 90° due right 0μs, 58.8μs, 117.6μs

[0064] S30: Constructing configuration parameters of the microphone array based on the user sound source direction parameters.

[0065] Specifically, step S30 includes:

[0066] S31: Determine the main beam direction angle of the microphone array based on the user sound source direction parameter.

[0067] Specifically, the sound source direction parameters are mapped to a predefined main beam direction set. According to the minimum supported angular resolution, such as one level every 5°, the current direction is rounded to the nearest main beam direction. For example, if the direction is estimated to be 62°, the main beam direction is selected as 60°, and the angle and corresponding number are output.

[0068] S32: Based on the user audio data, determine the user communication stability state, and dynamically determine the beam width of the microphone array according to the user communication stability state.

[0069] Specifically, the stability of the current user's call status is determined based on the duration and intensity variation range of the continuous speech segments in the user's audio data. If the speech duration within 500ms exceeds 80% and the energy variation is less than 10dB, it is judged to be a stable call state and the beam width is set to ±20°, otherwise it is set to ±45°. For example, if the continuous speech activity exceeds 40 frames and the fluctuation is only 5dB, it is judged to be a stable state and the narrow beam mode is used.

[0070] S33: Generate configuration parameters of the microphone array according to the main beam direction angle and beam width.

[0071] Specifically, based on the determined main beam direction angle and beam width, the corresponding control parameter set is selected from the array control template library, including the gain weight and time delay item of each channel, and the complete configuration parameters are constructed. For example, the main beam direction is 60° and the beam width is ±20°, and the channel delay sequence is [0.0ms, 0.2ms, 0.4ms], and the gain is [1.0, 0.9, 0.8], thereby obtaining the configuration parameters of the microphone array.

[0072] S40: Acquire first enhanced voice data in the direction of the user's sound source according to the configuration parameters.

[0073] Specifically, after obtaining the configuration parameters of the microphone array, beamforming processing is performed based on the time delay and gain weight set for each channel, and multiple channel signals are phase-aligned and weighted-summed according to the target direction to form a directional voice signal output. The output signal is the first enhanced voice data, which has high energy concentration and interference suppression effects. For example, when the main beam direction is 60°, the delay settings of channels 1, 2, and 3 are 0ms, 0.2ms, and 0.4ms, respectively, and the gain weights are 1.0, 0.9, and 0.8. After processing, the voice component of the synthetic signal in the target direction is enhanced, and the background noise in the non-target direction is suppressed, so that the output enhanced voice has a higher signal-to-noise ratio and clarity.

[0074] S50: Determine, according to the ambient audio data, the noise type corresponding to the ambient audio data.

[0075] Specifically, step S50 includes:

[0076] S51: Performing spectrum analysis on the ambient audio data by using Fourier transform to obtain corresponding spectrum energy distribution feature information.

[0077] Specifically, each frame of ambient audio data is processed by fast Fourier transform to extract the power spectrum density information in the range of 0-8kHz, and the spectral energy is grouped by frequency band to form a feature vector. For example, in a frame of spectrum, the low-frequency (0-500Hz) energy accounts for 80%, and the mid- and high-frequency bands are almost 0. The spectral characteristics of the frame are obtained as "low-frequency concentrated", which serves as the input feature of the subsequent classification model.

[0078] S52: Inputting the spectrum energy distribution feature information into a preset noise classification model to obtain the noise type corresponding to the ambient audio data.

[0079] Specifically, the eigenvectors in the spectral energy distribution feature information are weighted mapped through the Mel filter group to generate a Mel spectrum representation as the input of the preset noise classification model. The noise classification model adopts the audio classification model VGGish, and the input is the Mel spectrum feature of each frame, thereby outputting a multi-category noise label probability distribution, and selecting the category with the highest probability as the recognition result. For example, if the model outputs the classification probability of "steady-state noise: 0.87, impulse noise: 0.08, mixed noise: 0.05", the current frame is determined to be steady-state noise. After continuously judging multiple frames, the final result is output, thereby obtaining the noise type corresponding to the environmental audio data, such as steady-state noise, impulse noise or mixed noise.

[0080] S60: Based on the noise type, perform noise reduction processing on the first enhanced speech data to obtain noise-reduced audio data.

[0081] Specifically, step S60 includes:

[0082] S61: Based on the noise type, corresponding filtering strategy parameters are obtained from a preset filtering strategy mapping table.

[0083] Specifically, the current noise type is used as an index key to perform a lookup in the filtering strategy mapping table to obtain the corresponding processing strategy item for that noise type. Strategy parameters include filter type, frequency band setting, gain control factor, and threshold. For example, if the current noise type is "impulse noise," the strategy found is "spectral subtraction," with a window length of 512 points and a threshold of -20dB. These are used as parameters for subsequent noise reduction processing.

[0084] S62: Perform noise reduction processing on the first enhanced speech data according to the filtering strategy parameters to obtain noise-reduced audio data.

[0085] Specifically, the first enhanced speech data is subjected to a noise reduction operation according to the strategy parameters. The specific filtering method, filter coefficients, and processing path are determined based on the filtering strategy parameters. Signal denoising, spectral energy adjustment, or amplitude suppression are performed, and the processed result is output. The result is the noise-suppressed, noise-reduced audio data. For example, when using spectral subtraction, the spectral energy of each frame of the first enhanced speech data is first calculated, and the background noise estimated spectrum is subtracted from it. A spectral threshold is applied to avoid negative spectral effects. Finally, the audio frame is reconstructed to output clear speech. For example, if high-frequency impulse noise is significant in a certain data segment, the sudden spikes will be completely eliminated after processing, and the speech will be smooth and recognizable.

[0086] Specifically, step S62 includes:

[0087] S621: When the noise type is steady-state noise, use a Wiener filter algorithm to perform frequency domain smoothing processing on the first enhanced speech data.

[0088] Specifically, when the noise type is steady-state noise, Wiener filtering is performed. First, the power spectral density of each frame of the first enhanced speech data is calculated. At the same time, the average background noise spectrum estimate is constructed based on the non-speech frames in the previously extracted environmental audio data. On this basis, the signal-to-noise ratio of each frequency point is calculated, and a frequency domain gain function is generated. The gain function is then applied to the spectrum of the current speech frame. Finally, an inverse transform is performed to output the smoothed speech data. For example, against the background of continuous machine roar, the average noise spectrum is stable. After Wiener filtering, the low-frequency noise is significantly suppressed, and the speech component is clearer.

[0089] S622: When the noise type is impulse noise, perform spectrum subband energy adjustment processing on the first enhanced speech data using spectral subtraction.

[0090] Specifically, when the noise type is impulse noise, spectral subtraction is performed. First, each frame of the first enhanced speech data is converted into a frequency domain signal through fast Fourier transform, and the power value of each frequency point is extracted. Then, based on the average noise spectrum calculated from the non-speech frame in the previous ambient audio, energy subtraction is performed according to the frequency point. The amplitude lower limit threshold is applied to the subtracted spectrum to avoid negative spectrum. Finally, the noise-reduced speech frame is reconstructed through inverse transform. For example, in the background of periodic pulse sound emitted by workshop equipment, it is detected that the energy of the 1kHz-2kHz frequency band in the speech frame is abnormally high. The burst energy in this frequency band can be removed through spectral subtraction, and the speech signal becomes continuous and smooth.

[0091] S623: When the noise type is mixed noise, perform segmented filtering processing on the first enhanced speech data by combining adaptive bandpass filtering with a gating strategy.

[0092] Specifically, when the noise type is mixed noise, a combination of adaptive bandpass filtering and gating strategy is performed, and voice activity detection is performed on each frame of the first enhanced voice data. If it is detected as a voice frame, the bandpass filter is enabled to retain the 300Hz to 3400Hz frequency band to suppress high and low frequency interference. If it is detected as a non-voice frame, the gating strategy is enabled to limit the overall energy of the frame to a set upper limit such as -45dB to reduce the sudden interference of residual background noise. For example, in a factory scene, there are both fan roars and occasional metal collisions. The voice segment retains the core voice information through the bandpass, and the blank segment effectively weakens the non-voice impact sound through the energy threshold strategy, so that the voice after noise reduction is more continuous and the listening experience is more natural.

[0093] In one embodiment, after the configuration parameters are obtained, the beam direction and beam width of the microphone array are adjusted.

[0094] The microphone array is controlled to perform beamforming processing, multiple directional channel signals are weighted, and first enhanced voice data is output.

[0095] Specifically, according to the constructed configuration parameters, adjustment instructions are sent to the array control module that controls the microphone array to set the delay compensation value and gain weight of each microphone channel so that the main beam direction of the array is toward the target sound source. At the same time, the beam width is dynamically set according to the stability of the user's voice. For example, when the user is talking stably, the main beam is set to 60° and the beam width is ±20°, and the delays of channels 2, 3, and 4 are adjusted to 0ms, 0.2ms, and 0.4ms respectively, and combined with gain values of 1.0, 0.9, and 0.8, to achieve the array's enhanced reception performance in the target direction and suppressed in the non-target direction.

[0096] More specifically, the beam control parameters are loaded on the directional channel signals in sequence according to the channel numbers. A delay compensation operation is first performed on each channel signal to align its phase. The signals are then multiplied by the gain weights and weighted superposition is performed to generate a single-channel output signal with directional enhancement characteristics, which is the first enhanced voice data. For example, when the main direction of the array is 45°, the delays of channels 1-4 are set to 0ms, 0.15ms, 0.3ms, and 0.45ms, respectively. The signals are multiplied by the gain coefficients of 1.0, 0.95, 0.9, and 0.85 and then superimposed and output, effectively improving the voice energy in that direction and suppressing the interference of surrounding background noise.

[0097] In one embodiment, after the noise-reduced audio data is obtained, a signal-to-noise ratio test is performed on the noise-reduced audio data, and whether the noise-reduced audio data meets the output condition is determined based on the signal-to-noise ratio test result.

[0098] If the signal-to-noise ratio test result is up to standard, the noise-reduced audio data is determined to meet the output conditions and the noise-reduced audio data is sent to the target communication device.

[0099] Specifically, the noise reduction audio data is segmented by frame, and the signal energy and background noise residual energy of each frame are counted separately. The current signal-to-noise ratio value is calculated by comparing the mean square value of the speech activity frame and the non-speech frame. If the value is greater than the preset output threshold, such as 15dB, the current speech quality is considered to meet the standard. For example, in a certain noise reduction speech segment, the average energy of the speech frame is -25dB, and the non-speech segment is -45dB. The calculated SNR is about 20dB, which meets the output requirements.

[0100] More specifically, when the current signal-to-noise ratio detection result is higher than the set threshold, the noise-reduced voice data that meets the standard will be encapsulated into a standard communication protocol format and sent to the target terminal device through a wired or wireless network module. For example, after a voice segment with a signal-to-noise ratio of 21dB is detected as meeting the standard, it will be encoded into the G.711 format and pushed to the connected industrial communication gateway through the serial port interface to achieve stable and clear voice communication output.

[0101] In one embodiment, if the signal-to-noise ratio detection result is not up to standard, it is determined that the noise reduction audio data does not meet the output condition.

[0102] When the signal-to-noise ratio detection result is not up to standard, the filtering strategy parameters corresponding to the noise type are optimized until the optimized filtering strategy parameters make the corresponding noise reduction audio data meet the output conditions.

[0103] After the noise reduction audio data meets the output conditions, the filtering strategy mapping table is updated according to the optimized filtering strategy parameters.

[0104] Specifically, after completing the noise reduction process and detecting the signal-to-noise ratio, if the calculated result is lower than the set output threshold, the audio segment is marked as not meeting the output conditions and is not sent for communication. After confirming that the current signal-to-noise ratio does not meet the standard, the currently used filtering strategy entry is matched according to the noise type, and the parameters in the strategy are directionally fine-tuned. For example, the noise estimation intensity in the spectral subtraction method is increased by 2dB, and the lower limit of the threshold is increased by 1 level. After each parameter modification, the noise reduction process and signal-to-noise ratio detection are re-executed, and the cycle is repeated until the signal-to-noise ratio is greater than the output threshold. For example, the initial SNR is 12dB, and it is increased to 15.7dB after 3 rounds of optimization. The optimization process is terminated when the output conditions are met.

[0105] For example, after the optimized noise reduction processing result meets the signal-to-noise ratio output condition, the filtering strategy parameters generated in the final optimization process are extracted and written into the filtering strategy mapping table corresponding to the current noise type, overwriting the original parameter record, so that when the same noise type is encountered in the future, the group of verified optimal strategies that have been met are directly called. For example, the current pulse noise uses a spectral subtraction coefficient of 0.8 and an energy threshold of -40dB after optimization. The system writes this parameter group under the "pulse noise" strategy item as the default call plan for the next time.

[0106] It should be understood that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.

[0107] In one embodiment, an industrial telephone is provided, comprising a microphone array, an audio acquisition module, a communication transmission module, a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the following steps are performed:

[0108] Obtain environmental audio data and user communication audio data;

[0109] Determine the user's sound source direction parameters based on the user's communication audio data;

[0110] Based on the user's sound source direction parameters, the configuration parameters of the microphone array are constructed;

[0111] Acquire first enhanced voice data in the direction of the user's voice source according to the configuration parameters;

[0112] Determining, based on the ambient audio data, a noise type corresponding to the ambient audio data;

[0113] Based on the noise type, noise reduction processing is performed on the first enhanced speech data to obtain noise-reduced audio data.

[0114] The above-described embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present application, and should all be included in the scope of protection of the present application.

Claims

1. A method for reducing communication noise of an industrial telephone, characterized in that: The communication noise reduction method comprises: Obtain environmental audio data and user communication audio data; Determining a user sound source direction parameter based on the user communication audio data; Constructing configuration parameters of a microphone array based on the user sound source direction parameters; Acquire first enhanced voice data in the direction of the user's voice source according to the configuration parameters; Determining, based on the ambient audio data, a noise type corresponding to the ambient audio data; Based on the noise type, noise reduction processing is performed on the first enhanced speech data to obtain noise-reduced audio data.

2. The communication noise reduction method according to claim 1, characterized in that: The determining of the user speaking direction parameter according to the user communication audio data specifically includes: Acquire multiple directional channel signals from the user communication audio data; Based on a direction estimation algorithm, the multiple directional channel signals are analyzed to calculate the user sound source direction parameter.

3. The communication noise reduction method according to claim 1, wherein: The configuration parameters of the microphone array are constructed based on the user sound source direction parameters, specifically including: Determining a main beam direction angle of the microphone array based on the user sound source direction parameter; Determining a stable state of user communication based on the user audio data, and dynamically determining a beam width of the microphone array according to the stable state of user communication; Configuration parameters of the microphone array are generated according to the main beam direction angle and beam width.

4. The communication noise reduction method according to claim 1, wherein: The determining, based on the ambient audio data, the noise type corresponding to the ambient audio data specifically includes: Performing spectrum analysis on the ambient audio data by using Fourier transform to obtain corresponding spectrum energy distribution characteristic information; The spectrum energy distribution characteristic information is input into a preset noise classification model to obtain the noise type corresponding to the environmental audio data.

5. The communication noise reduction method according to claim 1, wherein: The performing noise reduction processing on the first enhanced speech data based on the noise type to obtain noise-reduced audio data specifically includes: Based on the noise type, obtaining corresponding filtering strategy parameters from a preset filtering strategy mapping table; The first enhanced speech data is subjected to noise reduction processing according to the filtering strategy parameters to obtain the noise-reduced audio data.

6. The communication noise reduction method according to claim 5, characterized in that: The filtering strategy parameters include noise reduction processing methods corresponding to different noise types, specifically including: When the noise type is steady-state noise, using a Wiener filtering algorithm to perform frequency domain smoothing processing on the first enhanced speech data; When the noise type is impulse noise, performing spectrum subband energy adjustment processing on the first enhanced speech data by using spectral subtraction; When the noise type is mixed noise, a combination of adaptive bandpass filtering and gating strategy is used to perform segmented filtering processing on the first enhanced speech data.

7. The communication noise reduction method according to claim 2, characterized in that: After obtaining the configuration parameters, adjusting the beam direction and beam width of the microphone array; The microphone array is controlled to perform beamforming processing, weighted processing is performed on the multiple directional channel signals, and the first enhanced voice data is output.

8. The communication noise reduction method according to claim 5, characterized in that: After obtaining the noise-reduced audio data, performing a signal-to-noise ratio test on the noise-reduced audio data, and determining whether the noise-reduced audio data meets an output condition according to a result of the signal-to-noise ratio test; If the signal-to-noise ratio detection result is up to standard, it is determined that the noise-reduced audio data meets the output condition, and the noise-reduced audio data is sent to the target communication device.

9. The communication noise reduction method according to claim 8, characterized in that: If the signal-to-noise ratio detection result is not up to standard, it is determined that the noise reduction audio data does not meet the output condition; When the signal-to-noise ratio detection result is not up to standard, optimizing the filtering strategy parameters corresponding to the noise type until the optimized filtering strategy parameters enable the corresponding noise reduction audio data to meet the output conditions; After the noise reduction audio data meets the output condition, the filtering strategy mapping table is updated according to the optimized filtering strategy parameters.

10. An industrial telephone comprising a microphone array, an audio acquisition module, a communication transmission module, a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the steps of the communication noise reduction method according to any one of claims 1 to 9 are implemented.