An audio output method and device based on a hearing protection device and an electronic device

By using the audio output method of the hearing protection device, audio signals are collected using digital and analog microphones. Combined with audio extraction models and active noise reduction technology, the problem of the hearing protection device being unable to isolate low-frequency noise in a strong noise environment is solved. This achieves effective audio signal extraction and noise reduction processing, ensuring hearing protection and information reception for on-site personnel.

CN122120667APending Publication Date: 2026-05-29HANGZHOU HIKVISION DIGITAL TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
HANGZHOU HIKVISION DIGITAL TECHNOLOGY CO LTD
Filing Date
2024-11-27
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

Existing hearing protection devices cannot effectively isolate low-frequency noise in noisy environments, leading to hearing damage. They may also weaken effective audio signals, such as ambient sounds and speech.

Method used

The audio output method of the hearing protection device uses digital and analog microphones to collect audio signals. Combined with a pre-trained audio extraction model and active noise reduction technology, effective audio signals are extracted and actively denoised through frequency domain signal conversion and convolutional neural network processing.

Benefits of technology

While providing hearing protection, it ensures that on-site personnel can receive effective audio signals from the outside world, reduces noise signals, and has a simple and lightweight structure.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122120667A_ABST
    Figure CN122120667A_ABST
Patent Text Reader

Abstract

The embodiment of the application provides an audio output method and device based on a hearing protection device and electronic equipment, and relates to the technical field of audio extraction. The audio output method based on the hearing protection device comprises the following steps: receiving a first audio signal, a second audio signal and a third audio signal subjected to physical noise reduction by an ear cover; performing extraction processing on the received first audio signal based on a predetermined characteristic audio extraction mode to obtain a target audio signal; performing active noise reduction processing on the target audio signal based on the received second audio signal and the third audio signal to obtain a target audio signal subjected to noise reduction processing; and outputting the target audio signal subjected to noise reduction processing through a loudspeaker of the ear cover. It can be seen that the present scheme can provide hearing protection for on-site personnel in a strong noise environment while ensuring that the on-site personnel can receive audio that belongs to effective audio signals from the outside world.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of audio extraction technology, and in particular to an audio output method, device and electronic device based on a hearing protection device. Background Technology

[0002] In noisy environments, personnel typically protect their hearing by wearing hearing protection devices that provide physical sound insulation. However, existing hearing protection devices only offer limited protection against high-frequency noise, with very limited effectiveness against low-frequency noise, which can still damage the hearing of those present.

[0003] However, if low-frequency sounds are completely isolated, some effective audio signals may be attenuated as well. Effective audio signals can be ambient sounds, voice, and other audio signals that the wearer can receive. Ambient sounds can be alarm sounds or the sound of large mobile machinery approaching, and voice can be information communicated between on-site personnel and the outside world.

[0004] Currently, there is an urgent need for an audio output method based on hearing protection devices, so as to provide hearing protection for on-site personnel in a high-noise environment while ensuring that on-site personnel can receive audio signals that are valid from the outside world. Summary of the Invention

[0005] The purpose of this application is to provide an audio output method, device, and electronic device based on a hearing protection device, so as to provide hearing protection for on-site personnel in a high-noise environment while ensuring that on-site personnel can receive valid audio signals from the outside world. The specific technical solution is as follows:

[0006] In a first aspect, embodiments of this application provide an audio output method based on a hearing protection device, applied to the control unit of the hearing protection device. The hearing protection device includes an earmuff with a predetermined physical noise reduction function. A digital microphone and a feedforward microphone belonging to the analog microphone are disposed on the outer side of the earmuff, and a speaker and a feedback microphone belonging to the analog microphone are disposed inside the earmuff. The method includes:

[0007] The system receives a first audio signal acquired by the digital microphone, a second audio signal acquired by the feedforward microphone, and a third audio signal acquired by the feedback microphone and subjected to physical noise reduction by the earcups.

[0008] Based on a predetermined feature audio extraction method, the received first audio signal is processed to extract effective audio signals to obtain a target audio signal; wherein, the predetermined feature audio extraction method includes: signal processing based on a pre-trained audio extraction model for extracting effective audio signals; the audio extraction model is composed of a frequency domain signal conversion module, a first encoder, a liquid neural network (LNN), a first decoder, and an inverse processing module corresponding to the frequency domain signal conversion module connected in series, wherein the first encoder and the first decoder both belong to a convolutional neural network (CNN);

[0009] Based on the received second and third audio signals, the target audio signal is subjected to active noise reduction processing to obtain the noise-reduced target audio signal.

[0010] The noise-reduced target audio signal is output through the speakers of the earcups.

[0011] Secondly, embodiments of this application provide an audio output device based on a hearing protection device, applied to the control unit of the hearing protection device. The hearing protection device includes earmuffs with a predetermined physical noise reduction function. A digital microphone and a feedforward microphone belonging to the analog microphone are disposed on the outer side of the earmuffs, and a speaker and a feedback microphone belonging to the analog microphone are disposed inside the earmuffs. The device includes:

[0012] The receiving module is used to receive the first audio signal acquired by the digital microphone, the second audio signal acquired by the feedforward microphone, and the third audio signal acquired by the feedback microphone after physical noise reduction by the earcups.

[0013] An extraction processing module is used to extract effective audio signals from a received first audio signal based on a predetermined feature audio extraction method to obtain a target audio signal. The predetermined feature audio extraction method includes signal processing based on a pre-trained audio extraction model for extracting effective audio signals. The audio extraction model is composed of a frequency domain signal conversion module, a first encoder, a liquid neural network (LNN), a first decoder, and an inverse processing module corresponding to the frequency domain signal conversion module connected in series. Both the first encoder and the first decoder belong to convolutional neural networks (CNNs).

[0014] An active noise reduction processing module is used to perform active noise reduction processing on the target audio signal based on the received second audio signal and third audio signal to obtain the noise-reduced target audio signal.

[0015] The output module is used to output the noise-reduced target audio signal through the speaker of the earcups.

[0016] Thirdly, embodiments of this application provide an electronic device, including:

[0017] Memory, used to store computer programs;

[0018] When the processor executes a program stored in memory, it implements any of the above-described audio output methods based on the hearing protection device.

[0019] Fourthly, embodiments of this application provide a computer-readable storage medium storing a computer program, which, when executed by a processor, implements any of the above-described audio output methods based on a hearing protection device.

[0020] Beneficial effects of the embodiments in this application:

[0021] The audio output method based on a hearing protection device provided in this application embodiment can extract and process the received first audio signal for effective audio signals based on a predetermined feature audio extraction method to obtain a target audio signal. Based on the received second and third audio signals, active noise reduction processing is performed on the target audio signal to obtain a noise-reduced target audio signal. This noise-reduced target audio signal is then output through the speaker of the earcup. Therefore, this application embodiment extracts and processes the effective audio signal and reduces noise signals through active noise reduction, thereby providing hearing protection for on-site personnel while ensuring they can receive effective information from the outside world. Furthermore, the audio extraction model for extracting and processing the effective audio signal in this application can include a frequency domain signal conversion module, a first encoder, a liquid neural network (LNN), a first decoder, and an inverse processing module corresponding to the frequency domain signal conversion module. Compared to related technologies, the audio extraction model in this application is not composed of a large neural network model; therefore, the structure of the audio extraction model in this application is simpler and more lightweight.

[0022] Of course, implementing any product or method of this application does not necessarily require achieving all of the advantages described above at the same time. Attached Figure Description

[0023] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other embodiments can be obtained based on these drawings.

[0024] Figure 1 A flowchart illustrating an audio output method based on a hearing protection device, provided for an embodiment of this application;

[0025] Figure 2(a) is a schematic diagram of the effect of an audio extraction model provided in an embodiment of this application;

[0026] Figure 2(b) is a schematic diagram illustrating the working principle of an audio extraction model provided in an embodiment of this application;

[0027] Figure 3 A schematic diagram illustrating the principle of active noise reduction provided in an embodiment of this application;

[0028] Figure 4 A schematic diagram illustrating the principle of determining H^(n) provided in an embodiment of this application;

[0029] Figure 5 A schematic diagram illustrating the principle of eliminating periodic signals, provided for an embodiment of this application;

[0030] Figure 6 This is a schematic diagram illustrating the effect of a signal cancellation model provided in an embodiment of this application;

[0031] Figure 7(a) is a schematic diagram of the structure of a correction processing model provided in an embodiment of this application;

[0032] Figure 7(b) is a schematic diagram of the effect of a correction processing model provided in an embodiment of this application;

[0033] Figure 8 This application provides a schematic diagram of the structure of various models in a hearing protection device.

[0034] Figure 9 This is a schematic diagram illustrating a training process for an audio extraction model, provided as an embodiment of this application.

[0035] Figure 10(a) is a schematic diagram of the structure of the hearing protection device provided in the embodiment of this application;

[0036] Figure 10(b) is a schematic diagram of the working principle of a hearing protection device provided in an embodiment of this application;

[0037] Figure 10(c) is a schematic diagram of the working principle of another hearing protection device provided in the embodiment of this application;

[0038] Figure 11 A schematic diagram of the structure of an audio output device based on a hearing protection device provided in this application embodiment;

[0039] Figure 12 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0040] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art based on this application are within the scope of protection of this application.

[0041] First, a brief introduction to some technical terms used in the embodiments of this application:

[0042] Analog microphone: Also known as AMIC (Analog Microphone) or analog microphone, it is a traditional microphone that converts sound signals into electrical signals through analog circuitry.

[0043] Digital microphone: Also known as DMIC (Digital Microphone) or digital microphone, a digital microphone is based on an analog microphone, but converts the acquired electrical signals into digital signals.

[0044] Analog-to-digital converter (ADC): A device that converts continuous analog signals into discrete digital signals.

[0045] A digital-to-analog converter (DAC) is a device that converts discrete digital signals into continuous analog signals.

[0046] Feedforward microphone: Also known as FF_AMIC or AMIC (shell), used to collect external audio signals.

[0047] Feedback microphone: Also known as FB_AMIC or AMIC (internal cavity), it is used to collect audio signals after physical noise reduction inside the earcups.

[0048] Active noise cancellation: also known as ANC, is a noise cancellation technology mainly used in headphone noise cancellation.

[0049] Real part information and imaginary part information: These are information specific to complex numbers, for example, complex numbers. x is called the real part of the complex number z, and y is called the imaginary part of the complex number z. Therefore, the information located at position x can be called the real part information, and the information located at position y can be called the imaginary part information.

[0050] Furthermore, to better understand this solution, a brief introduction to the relevant technologies is provided below:

[0051] In some industrial scenarios, the signal strength of effective external information is much lower than that of environmental noise. Simply using a neural network model for audio extraction will result in low signal strength of the extracted audio signal and poor extraction effect. Furthermore, the extracted audio signal may be blurry and difficult to identify.

[0052] Based on the problems described above, this application provides an audio output method, apparatus, and electronic device based on a hearing protection device.

[0053] The following describes the audio output method based on a hearing protection device provided in the embodiments of this application.

[0054] The audio output method based on a hearing protection device provided in this application embodiment can be applied to the control unit of the hearing protection device. The hearing protection device also includes an earmuff with a predetermined physical noise reduction function. A digital microphone and a feedforward microphone belonging to the analog microphone are arranged on the outside of the earmuff. A speaker and a feedback microphone belonging to the analog microphone are also arranged inside the earmuff. Specifically, the shape of the hearing protection device can be an earmuff with multiple devices arranged inside and outside the earmuff. Its shape can also be an earphone. This application embodiment does not specifically limit the specific form of the hearing protection device.

[0055] Furthermore, any analog microphone, digital microphone, and speaker from the related technologies can be used as components in the hearing protection device of this application embodiment, and this application embodiment does not specifically limit this. In addition, the control unit of the hearing protection device can also be considered as the central processing unit (CPU) or core control unit of the hearing protection device, and this application embodiment does not specifically limit this.

[0056] In addition, the on-site personnel described in this application, that is, the personnel in the environment with noisy signals, can also be considered as the "users" of the hearing protection device. This application does not specifically limit this.

[0057] It should be noted that the frequency domain signal conversion module and the corresponding inverse processing module included in the audio extraction model described in the embodiments of this application do not limit the audio extraction model to physically containing these two hardware components. The terms "unit" and "module" in the embodiments of this application can refer to hardware components, software programs, or the necessary carriers for the operation of software programs. With technological evolution, the steps or functions required to execute these "units" or "modules" (e.g., steps described after "a unit or module is used for") may be performed using hardware, software, or a combination of both. Any hardware, software, or a combination of hardware and software capable of executing the steps required by the "unit" can be considered a "unit" or "module" in the embodiments of this application.

[0058] The following is a brief introduction to the application scenarios of the audio output method based on hearing protection devices provided in the embodiments of this application:

[0059] The application scenarios of this application embodiment can be noisy scenarios, such as industrial scenarios or shooting range scenarios. Specifically, industrial scenarios can be foundries, textile workshops, etc. This application embodiment does not specifically limit the application scenarios.

[0060] One method for audio output based on a hearing protection device is applied to the control unit of the hearing protection device. The hearing protection device includes earmuffs with a predetermined physical noise reduction function. A digital microphone and a feedforward microphone (which is an analog microphone) are disposed on the outer side of the earmuffs, and a speaker and a feedback microphone (which is an analog microphone) are disposed inside the earmuffs. The method includes:

[0061] The system receives a first audio signal acquired by the digital microphone, a second audio signal acquired by the feedforward microphone, and a third audio signal acquired by the feedback microphone and subjected to physical noise reduction by the earcups.

[0062] Based on a predetermined feature audio extraction method, the received first audio signal is processed to extract effective audio signals to obtain a target audio signal; wherein, the predetermined feature audio extraction method includes: signal processing based on a pre-trained audio extraction model for extracting effective audio signals; the audio extraction model is composed of a frequency domain signal conversion module, a first encoder, a liquid neural network (LNN), a first decoder, and an inverse processing module corresponding to the frequency domain signal conversion module connected in series, wherein the first encoder and the first decoder both belong to a convolutional neural network (CNN);

[0063] Based on the received second and third audio signals, the target audio signal is subjected to active noise reduction processing to obtain the noise-reduced target audio signal.

[0064] The noise-reduced target audio signal is output through the speakers of the earcups.

[0065] The audio output method based on a hearing protection device provided in this application embodiment can extract and process the received first audio signal for effective audio signals based on a predetermined feature audio extraction method to obtain a target audio signal. Based on the received second and third audio signals, active noise reduction processing is performed on the target audio signal to obtain a noise-reduced target audio signal. This noise-reduced target audio signal is then output through the speaker of the earcup. Therefore, this application embodiment extracts and processes the effective audio signal and reduces noise signals through active noise reduction, thereby providing hearing protection for on-site personnel while ensuring they can receive effective information from the outside world. Furthermore, the audio extraction model for extracting and processing the effective audio signal in this application can include a frequency domain signal conversion module, a first encoder, a liquid neural network (LNN), a first decoder, and an inverse processing module corresponding to the frequency domain signal conversion module. Compared to related technologies, the audio extraction model in this application is not composed of a large neural network model; therefore, the structure of the audio extraction model in this application is simpler and more lightweight.

[0066] The following describes, with reference to the accompanying drawings, an audio output method based on a hearing protection device provided in an embodiment of this application.

[0067] like Figure 1 As shown in the illustration, this application provides an audio output method based on a hearing protection device, applied to the control unit of the hearing protection device. The hearing protection device includes earmuffs with a predetermined physical noise reduction function. A digital microphone and a feedforward microphone (which is an analog microphone) are disposed on the outer side of the earmuffs, and a speaker and a feedback microphone (which is an analog microphone) are disposed inside the earmuffs. The method includes:

[0068] S101, receive the first audio signal acquired by the digital microphone, receive the second audio signal acquired by the feedforward microphone, and receive the third audio signal acquired by the feedback microphone after physical noise reduction by the earcups;

[0069] It is understood that the first audio signal is the audio signal about the outside world collected by the digital microphone, the second audio signal is the audio signal about the outside world collected by the feedforward microphone, and the third audio signal is the audio signal about the inside of the earcups collected by the feedback microphone after physical noise reduction; wherein, physical noise reduction reduces the transmission of noise signals through the absorption and isolation properties of materials, and this application embodiment does not specifically limit this.

[0070] It should be emphasized that both the first and second audio signals are undenoised audio signals, and therefore may contain both noise and valid audio signals. The third audio signal, however, is a physically denoised audio signal, and therefore may or may not contain noise; this embodiment does not specifically limit this. Furthermore, both the second and third audio signals are audio signals used in active noise reduction processing, which will be described in detail in subsequent steps and will not be elaborated upon here.

[0071] S102, based on a predetermined feature audio extraction method, the received first audio signal is processed to extract the effective audio signal to obtain the target audio signal;

[0072] The predetermined feature audio extraction method includes: signal processing based on a pre-trained audio extraction model for extracting effective audio signals; the audio extraction model is composed of a frequency domain signal conversion module, a first encoder, a liquid neural network (LNN), a first decoder, and an inverse processing module corresponding to the frequency domain signal conversion module connected in series; the first encoder and the first decoder both belong to convolutional neural networks (CNNs).

[0073] It is understood that the predetermined feature audio extraction method can also be considered as inputting the received first audio signal into a pre-trained audio extraction model for extracting effective audio signals, so that the audio extraction model performs signal processing on the first audio signal. The audio extraction model consists of multiple modules connected serially, specifically including: a frequency domain signal conversion module, a first encoder, a Liquid Neural Network (LNN), a first decoder, and an inverse processing module corresponding to the frequency domain signal conversion module. Furthermore, the audio extraction model can also be called a feature signal extraction unit or a signal extraction unit; this application embodiment does not specifically limit this terminology.

[0074] Specifically, the frequency domain signal conversion module can be used to perform STFT (short-time Fourier transform), thereby converting the input first audio signal into a frequency domain signal. The real and imaginary parts of the frequency domain signal are then input into the first encoder. The first encoder performs feature encoding processing on the received real and imaginary parts, respectively. The encoded features corresponding to the real and imaginary parts are then passed to the first decoder via a liquid neural network (LNN). The first decoder decodes the received encoded features and passes the decoded features to the inverse processing module corresponding to the frequency domain signal conversion module. This inverse processing module performs ISTFT (inverse short-time Fourier transform), converting the frequency domain signal into a time domain signal, thereby obtaining the target audio signal. Furthermore, both the first encoder and the first decoder belong to Convolutional Neural Networks (CNNs). CNNs can be based on multi-head local self-attention mechanisms to achieve feature encoding and decoding. Multi-head local attention is a widely used attention mechanism that obtains the attention distribution of the input sequence by running multiple independent attention mechanisms in parallel. In a multi-head attention mechanism, the input sequence first passes through three different linear transformation layers to obtain three vectors: Query, Key, and Value. The transformed vectors are then divided into several "heads," each with its own independent Query, Key, and Value matrices. The local attention mechanism based on the audio extraction model in this application can be 6-head (x6). Additionally, the LNN can have 32 neurons with random connection paths; this embodiment does not specifically limit this.

[0075] To better understand the structure of the audio extraction model, the following description is provided with reference to the accompanying figures, as shown in Figures 2(a) and 2(b):

[0076] The first audio signal input to the audio extraction model in Figure 2(a) can be represented as: Y(k) = S(k) + F1(k) + F2(k) + ... + Fn(k). The target audio signal output by the audio extraction model consists of two audio signals. The first audio signal is a non-speech signal, which can be represented as ∑Fn(k), specifically F1(k) + F2(k) + ... + Fn(k). The other audio signal is a valid audio signal, which can be represented as S(k).

[0077] The frequency domain signal conversion module in Figure 2(b) performs a short-time Fourier transform on the first audio signal to obtain a frequency domain signal. The real and imaginary parts of the frequency domain signal are then input into the first encoder, which is also called the Encoder. The Encoder performs feature encoding on the received real and imaginary parts, which can be understood as feature extraction. The LNN then transmits the encoded features corresponding to the real and imaginary parts to the first decoder to transmit the information. The first decoder, also called the Decoder, decodes the received features, which can be considered as the process of generating information. The decoded features are then transmitted to the inverse processing module corresponding to the frequency domain signal conversion module. The inverse processing module can convert the frequency domain signal into a time domain signal through an inverse short-time Fourier transform to obtain the target audio signal.

[0078] For clarity, the specific training process of the audio extraction model will be described in subsequent embodiments and will not be elaborated upon here. Furthermore, it should be emphasized that this application does not limit the specific structure of any component in the aforementioned audio extraction model, or any of the subsequent mentioned correction processing model, diffusion model, etc. Any structure capable of achieving the required functionality of a component can be applied to this application.

[0079] S103, based on the received second and third audio signals, perform active noise reduction processing on the target audio signal to obtain the noise-reduced target audio signal;

[0080] It is understandable that the target audio signal can be used as the desired signal, the second audio signal acquired by the feedforward microphone can be used as the reference signal of the adaptive filter, and the target audio signal after noise reduction can be composed of the desired signal and the noise cancellation signal; wherein, the noise cancellation signal is determined based on the second audio signal and the third audio signal.

[0081] To better understand the process of active noise reduction, the following explanation is provided with reference to the accompanying diagrams. Figure 3 As shown:

[0082] The first audio signal acquired by the digital microphone can be extracted and processed to obtain the target audio signal. This process can also be called feature audio extraction. H(n) is the transfer function of the hardware circuit. In this circuit, the transfer function H(n) is used to transfer the target audio signal to the computing unit ∑. The feedforward microphone can acquire the second audio signal, which is represented by X(n). H^(n) is the estimate of H(n), and W(n) is the parameter of the adaptive filter. The feedforward microphone can send the second audio signal X(n) into the adaptive filter and H^(n) respectively. Through H^(n), the second audio signal can be transmitted to the ANC. The ANC can be considered as an active noise reduction unit. Then, the ANC sends the noise-reduced audio signal to the adaptive filter. The adaptive filter can process the second audio signal. The audio signals X(n) and ANC are filtered. During the filtering process, W(n) can be adjusted in real time to obtain the filtered audio signal. Then, based on H(n), it is transmitted to the calculation unit ∑. The feedback microphone can collect the third audio signal based on the physical soundproofing structure and transmit the third audio signal to the calculation unit ∑. Then, the calculation unit ∑ can sum the target audio signal transmitted by H(n) and the third audio signal transmitted by the feedback microphone, and then subtract the filtered audio signal transmitted by H(n) to obtain the noise-reduced target audio signal Z(n). The noise-reduced target audio signal Z(n) can be output to the speaker or input to ANC. In addition, in this process, the error can also be calculated, that is, the error is calculated, thereby improving the effect of active noise reduction.

[0083] It is important to emphasize that H(n) and H^(n) are both fixed hardware circuits. The determination of H(n) and H^(n) will be combined with... Figure 4 To introduce, such as Figure 4 As shown:

[0084] White noise is collected by a feedforward microphone and divided into three paths. The first path is directly output to the speaker via H(n), and then collected by the feedback microphone and input to the DSP (Digital Signal Processor). The second path is output to an adaptive filter via H^(n) for convolution operation, and the difference is calculated with the audio signal collected by the feedback microphone. That is, the difference is calculated based on the computational unit ∑. When the difference approaches 0, the purpose of white noise cancellation is achieved. The third path is input to the ANC as reference information and as input to the cancellation algorithm. The result of the above difference can also be input to the ANC, and the gradient descent algorithm is used to iterate over H^(n).

[0085] The specific calculation process is as follows:

[0086] At time n, calculate the output F(n) of the adaptive filter by calculating e(n) = y(n) - F(n), where y(n) is the output of the feedback microphone and e(n) is the difference between the two. If the result is not close to 0, adjust the parameters of the adaptive filter using the gradient descent algorithm. Then, at n = n + 1, recalculate time n and calculate the output F(n) of the adaptive filter. Repeat this iterative calculation until... Approaching 0, that is, e(n) approaches 0.

[0087] Of course, the active noise reduction processing described above is only an example. Any active noise reduction technology in the related technologies can be applied to the embodiments of this application, and the embodiments of this application do not specifically limit it.

[0088] S104, the target audio signal after noise reduction is output through the speaker of the earcups.

[0089] It is understandable that the speaker in the earcups of the hearing protection device can output the noise-reduced target audio signal. The target audio signal does not contain noise signals, only valid audio signals, thus providing hearing protection for on-site personnel while ensuring that they can receive valid information from the outside world.

[0090] The audio output method based on a hearing protection device provided in this application embodiment can extract and process the received first audio signal for effective audio signals based on a predetermined feature audio extraction method to obtain a target audio signal. Based on the received second and third audio signals, active noise reduction processing is performed on the target audio signal to obtain a noise-reduced target audio signal. This noise-reduced target audio signal is then output through the speaker of the earcup. Therefore, this application embodiment extracts and processes the effective audio signal and reduces noise signals through active noise reduction, thereby providing hearing protection for on-site personnel while ensuring they can receive effective information from the outside world. Furthermore, the audio extraction model for extracting and processing the effective audio signal in this application can include a frequency domain signal conversion module, a first encoder, a liquid neural network (LNN), a first decoder, and an inverse processing module corresponding to the frequency domain signal conversion module. Compared to related technologies, the audio extraction model in this application is not composed of a large neural network model; therefore, the structure of the audio extraction model in this application is simpler and more lightweight.

[0091] Optionally, in another embodiment, before signal processing based on the pre-trained audio extraction model for extracting effective audio signals, the predetermined feature audio extraction method further includes:

[0092] Step A1: Perform periodic signal cancellation processing on the received first audio signal;

[0093] Accordingly, the signal processing based on the pre-trained audio extraction model for extracting effective audio signals includes step B1:

[0094] Step B1 involves performing signal processing on the eliminated audio signal based on a pre-trained audio extraction model for extracting valid audio signals.

[0095] It is understandable that, for step A1, the periodic signal is usually a noise signal generated by a machine, that is, a periodic noise signal. Therefore, after receiving the first audio signal, the periodic signal in the first audio signal can be directly eliminated to obtain the eliminated audio signal.

[0096] It is understandable that, for step B1, the predetermined feature audio extraction method can be: based on an audio extraction model, performing signal processing on the eliminated audio signal. It is important to emphasize that not all scenarios contain periodic signals (periodic noise signals). If no periodic signals exist in the current scenario, then eliminating periodic signals is unnecessary. For example, in a shooting range scenario, there are no periodic signals, so eliminating periodic signals is unnecessary. For example, in another implementation, the process of eliminating periodic signals...

[0097] It can be executed by a separation processing model, which can also be called a separation processing unit. This application does not specifically limit this. Similarly, if there is no periodic signal in the current scenario, then there is no need to set up a separation processing model in the hearing protection device, or the separation processing model does not work.

[0098] As can be seen, the embodiments of this application can perform periodic signal elimination processing on the first audio signal collected by the digital microphone, thereby separating the periodic signal and separating the effective audio signal submerged in the periodic signal, improving the noise reduction effect, and providing a basis for providing hearing protection for on-site personnel while ensuring that on-site personnel can receive effective information from the outside world.

[0099] In one implementation, step A1, which involves performing periodic signal cancellation processing on the received first audio signal, includes steps A11-A13:

[0100] Step A11: Delay the received first audio signal for a predetermined duration to obtain the delayed audio signal;

[0101] The scheduled duration must be greater than 30ms, but 100ms is usually acceptable. This application does not impose any specific restrictions on this.

[0102] It is understandable that after the delay processing, the effective audio signal in the delayed audio signal is unrelated to the first audio signal before the delay; however, since periodic signals have a period, the periodic signal in the delayed audio signal is still related to the first audio signal before the delay.

[0103] Step A12: The periodic signal in the delayed audio signal is filtered by a preset filter to obtain a periodic signal;

[0104] It is understandable that an autocorrelation function is set in the preset filter. The autocorrelation function is used to indicate the audio signal that is correlated with the first audio signal. Therefore, by filtering the periodic signal in the delayed audio signal through the preset filter, a periodic signal can be obtained.

[0105] Step A13: Eliminate the periodic signal in the acquired first audio signal.

[0106] It is understandable that the periodic signal obtained by filtering can be used to separate and eliminate the periodic signal in the acquired first audio signal. However, the audio signal after elimination cannot be directly confirmed as a valid audio signal. Step A13 can only eliminate periodic noise signals. Non-periodic noise signals may still exist in the audio signal after elimination. This application embodiment does not specifically limit this.

[0107] To better understand the process of eliminating periodic signals, the following explanation is provided with reference to the accompanying diagram. Figure 5 As shown:

[0108] Figure 5 This can be considered a schematic diagram of the working principle of the separation processing model. The signal at point a is the first audio signal input to the separation processing model, and the first audio signal can be characterized as follows: ,in, For periodic signals generated by mechanical vibration in a mechanically noisy environment, PB represents the periodic frequency band. This is a broadband signal, which can also be considered an effective audio signal. BB represents Broad Band, and k can be considered as time. The signal at point a can be delayed, with a delay duration of... For example, it can be 100ms. The signal at point b is the delayed audio signal, which can be represented as x(k), x(k) = y(k- x(k) can be considered as the delayed signal of y(k). FIR characterizes the filter. The signal at point c output by the filter is a periodic signal. By calculating the unit ∑, the difference between the signal at point a and the signal at point c can be calculated to obtain the signal at point d. The signal at point d is the audio signal with the periodic signal eliminated.

[0109] As can be seen, the embodiments of this application can filter the periodic signal in the delayed audio signal through delay processing to obtain the periodic signal, and eliminate the periodic signal in the first audio signal collected, thereby improving the noise reduction effect. This also provides a basis for providing hearing protection for on-site personnel while ensuring that on-site personnel can receive effective information from the outside world.

[0110] Optionally, in another embodiment, the output of the audio extraction model includes both speech and non-speech signals;

[0111] The pre-trained audio extraction model for extracting effective audio signals, after signal processing of the eliminated audio signal, further includes the predetermined feature audio extraction method as follows:

[0112] Step C1: Perform specified correction processing on the speech signal in the output of the audio extraction model;

[0113] Accordingly, the target audio signal includes two audio signals, wherein one of the two audio signals is the non-speech signal in the output result of the audio extraction model, and the other signal is the speech signal obtained by performing specified correction processing on the speech signal in the output result of the audio extraction model.

[0114] The specified correction process can correct the speech signal in the output of the audio extraction model, making the speech signal in the output of the audio extraction model clearer. This application does not specifically limit this process.

[0115] It is understandable that after the first audio signal is input into the audio extraction model, the audio extraction model can output two signals, one of which is a speech signal and the other is a non-speech signal. Then, after extracting the signal from the eliminated audio signal, the speech signal in the output of the audio extraction model can be subject to specified correction processing. The speech signal after specified correction processing can be used as one audio signal in the target audio signal, while the other audio signal in the target audio signal is the non-speech signal in the output of the audio extraction model.

[0116] As can be seen, the embodiments of this application can perform specified correction processing on the speech signal in the output result of the audio extraction model, making the speech signal in the output result of the audio extraction model clearer, and using the speech signal obtained by specified correction processing as one audio signal in the target audio signal, thereby improving the noise reduction effect, enhancing the clarity of the audio signal, and improving the user experience.

[0117] For example, in one implementation, step C1 involves performing specified correction processing on the speech signal in the output of the audio extraction model, including steps C11-C12:

[0118] Step C11: Perform speech content correction processing on the speech signal in the output result of the audio extraction model;

[0119] Understandably, the first step in the specified correction process can be to perform speech content correction processing on the speech signal in the output of the audio extraction model. Speech content correction processing can solve the problems of missing words and unclear words in the speech signal. For example, the content represented by the speech signal in the output of the audio extraction model is: Device 1. Obviously, the content represented by this speech signal is missing words, so speech content correction processing can be performed on this speech signal. The content represented by the speech signal after speech content correction processing is: Please turn on device 1.

[0120] For clarity, the process of correcting the speech content in the output of the audio extraction model will be described in subsequent embodiments and will not be elaborated here.

[0121] Step C12: Perform signal elimination processing on the speech signal after speech content correction processing for the specified type of signal;

[0122] The specified type of signal includes the voice signal emitted by the wearer of the hearing protection device and / or the signal of a predetermined category of ambient sound.

[0123] It is understood that after performing speech content correction processing on the speech signal in the output of the audio extraction model, the speech signal after speech content correction processing can also be processed to eliminate signals of a specified type. The eliminated signals of a specified type can be the speech signal emitted by the wearer of the hearing protection device and / or the signal of a predetermined category of ambient sound, or can be considered as residual audio signals of human voice / environment. This application embodiment does not specifically limit this; wherein, the signal of the predetermined category of ambient sound is an invalid audio signal, such as the gunshot signal in a shooting range scene, and the non-periodic noise signal emitted by machinery in an industrial scene. This application embodiment does not specifically limit this.

[0124] Furthermore, the step of eliminating a specified type of signal from the speech signal after speech content correction processing can be performed by a signal cancellation model, which can also be called a signal cancellation unit or a signal cancellation processing unit. This application does not specifically limit this. It should be emphasized that the structure of the signal cancellation model is consistent with the audio extraction model described above. The only difference is that the encoder and decoder in the signal cancellation model use a four-head local attention mechanism, and the LNN can have only 16 neurons. Therefore, the specific structure of the signal cancellation model will not be elaborated further. Also, the signal cancellation model is pre-trained, and the training process is consistent with that of the audio extraction model. The specific training process will be described in subsequent embodiments and will not be elaborated further here.

[0125] To better understand the relevant content regarding signal cancellation models, the following explanation is provided with reference to the accompanying diagrams, such as... Figure 6 As shown:

[0126] The speech signal after speech content correction processing of the input signal cancellation model is S'(k), and the speech signal after cancellation processing output by the signal cancellation model is S^(k).

[0127] As can be seen, the specified correction process in the embodiments of this application may include speech content correction processing of the speech signal and elimination processing of specified types of signals. Speech content correction processing can make the audio signal clearer, and elimination processing can make the audio signal free of residual audio signals, thereby improving the noise reduction effect and improving the user experience.

[0128] Optionally, in another embodiment, the speech content correction processing of the speech signal in the output of the audio extraction model includes step C111:

[0129] Step C111: Based on the pre-trained correction processing model for speech content correction, perform speech content correction processing on the speech signal in the output of the audio extraction model.

[0130] The correction processing model consists of a spectrum conversion unit, a second encoder, a gated loop unit, a preset diffusion model, and a second decoder.

[0131] The spectrum conversion unit is used to perform spectrum conversion on the speech signal in the output of the audio extraction model to obtain the target spectrum.

[0132] The second encoder is used to extract features from the target spectrum to obtain current feature information;

[0133] The gated loop unit is used to acquire historical feature information and input the historical feature information and the current feature information into the diffusion model; wherein, the historical feature information is the feature information obtained at the moment before the feature advance time of the current feature information;

[0134] The diffusion model is used to generate parameters to be utilized based on the historical feature information and the current feature information, and input the parameters to be utilized, the historical feature information and the current feature information into the second decoder; wherein, the parameters to be utilized are the convolution parameters to be utilized by the second decoder;

[0135] The second decoder is used to generate a speech signal after speech content correction processing based on the parameters to be used, historical feature information and current feature information.

[0136] It is understood that the speech signal from the output of the audio extraction model can be directly input into a pre-trained correction processing model for speech content correction, so that the correction processing model can perform speech content correction processing on the speech signal. Specifically, the correction processing model may include a spectrum conversion unit, a second encoder, a gated recurrent unit, a preset diffusion model, and a second decoder. Both the second encoder and the second decoder belong to Convolutional Neural Networks (CNNs). The gated recurrent unit can also be called a GRU (Gated Recurrent Unit), which is a variant of a recurrent neural network. The preset diffusion model can also be called a small diffusion model; this application does not specifically limit its usage. Furthermore, the correction processing model can also be called a correction processing unit; this application's embodiments do not specifically limit its usage.

[0137] It is important to emphasize that after the speech signal from the audio extraction model is input into the correction processing model, the spectrum conversion unit in the correction processing model can perform spectrum conversion on the speech signal to obtain the target spectrum. For example, in one implementation, the spectrum conversion is specifically a Mellin transform, an integral transform with a power function as its kernel; this application does not specifically limit this implementation. Furthermore, the second encoder is used to extract features from the target spectrum to obtain current feature information. The second encoder has the same function as the first encoder described above, so it will not be elaborated further here.

[0138] It is important to emphasize that the gated recurrent unit is used to introduce historical feature information during the speech content correction process and input the historical feature information and current feature information into the diffusion model. Since the speech signal has temporal continuity, it can be considered that the speech signals of the preceding and following sequences are dependent on each other. Therefore, the preceding sequence can be used to predict the following sequence. The gated recurrent unit has time series prediction properties. In this way, the gated recurrent unit can introduce historical feature information and input it and current feature information into the diffusion model at the same time.

[0139] It is important to emphasize that the diffusion model can generate the convolutional parameters to be used by the second decoder based on historical and current feature information. The diffusion model can also be tuned for the convolutional parameters used by the second decoder. In another implementation, the diffusion model may not exist, and the output of the gated recurrent unit can be directly input to the second decoder. This application does not specifically limit this approach. Furthermore, the second decoder can generate a speech signal after speech content correction processing based on the parameters to be used, historical feature information, and current feature information. Since the second decoder references historical feature information, the generated speech signal after speech content correction processing is more accurate.

[0140] To better understand the above-mentioned corrective treatment model, the following description is provided with reference to the accompanying figures, as shown in Figures 7(a) and 7(b):

[0141] Figure 7(a) illustrates the working principle of the correction processing model. The correction processing model includes a spectrum conversion unit, a second encoder, a gated loop unit, a diffusion model, and a second decoder. The spectrum conversion unit can convert the speech signal in the output of the audio extraction model to obtain the target spectrum and input it into the second encoder. The second encoder can extract features from the target spectrum to obtain the current feature information and input it into the gated loop unit. The gated loop unit can acquire historical feature information and input the historical feature information and the current feature information into the diffusion model. The diffusion model can generate parameters to be used based on the historical feature information and the current feature information, and input the parameters to be used, the historical feature information, and the current feature information into the second decoder. The second decoder can generate the speech signal after speech content correction processing based on the parameters to be used, the historical feature information, and the current feature information, thereby realizing speech content correction processing.

[0142] Figure 7(b) illustrates the effect of the correction processing model. Problematic audio signals (blurred, missing words, elision) can be input into the correction processing model, and the correction processing model can output clear audio signals.

[0143] As can be seen, the embodiments of this application can perform speech content correction processing on the speech signal in the output result of the audio extraction model, making the corrected speech signal clearer and eliminating the problems of blurriness, missing words, and swallowing sounds, thereby improving the noise reduction effect and enhancing the user experience.

[0144] Optionally, in another embodiment, the specified correction processing of the speech signal in the output of the audio extraction model further includes step D1:

[0145] Step D1 involves adjusting the pitch of the speech signal after the cancellation of the specified type of signal.

[0146] It is understood that the specified correction process can also perform pitch adjustment processing on the speech signal after the elimination process, thereby correcting the speech pitch distortion caused during the elimination process of the specified type of signal. In addition, the pitch adjustment processing can also be performed by a pitch adjustment model, which can also be called an EQ (Equalizer) model or an EQ unit. This application embodiment does not specifically limit this.

[0147] To better understand the models described above, the following explanation is provided in conjunction with the accompanying diagrams, such as... Figure 8 As shown:

[0148] The digital microphone inputs the first audio signal to a separation processing model, which performs periodic signal elimination processing. The eliminated audio signal is then input to an audio extraction model, which extracts the valid audio signal, resulting in speech and non-speech signals. The speech signal is then input to a correction processing model for speech content correction. The speech signal after speech content correction is input to a signal elimination model for elimination of a specified signal type. Finally, the speech signal after elimination of the specified signal type is input to a pitch adjustment model for pitch adjustment, resulting in one speech signal from the target audio signal. Furthermore, the models described above can also be considered as models used to execute feature audio extraction algorithms.

[0149] For example, in another implementation, all the models described above can be set in the hearing protection device. However, in some special scenarios, they may not be turned on or may not operate. For instance, in a shooting range scenario, there are no periodic signals, so the separation processing model does not need to be turned on. In this case, the first audio signal can be directly input to the audio extraction model. Furthermore, after the hearing protection device is turned on, during the initialization phase, it can be determined which models can be turned on / operate based on a pre-stored configuration file. The configuration file is pre-configured by the user to indicate the models that need to be turned on / operate in the current scenario; this embodiment does not specifically limit this.

[0150] As can be seen, the embodiments of this application can perform tone adjustment processing on the eliminated speech signal, ensuring that the output audio signal does not change pitch, improving the clarity of the output audio signal, and enhancing the user experience.

[0151] Optionally, in another embodiment, before performing active noise reduction processing on the target audio signal based on the received second and third audio signals to obtain the noise-reduced target audio signal, steps E1-E3 are further included:

[0152] Step E1: Calculate the specified feature information of the target audio signal as the target feature information;

[0153] It is understood that, before performing active noise reduction processing, specified feature information of the target audio signal can be calculated as target feature information. This specified feature information can be energy attribute information of the target audio signal, such as energy entropy, signal amplitude, short-time energy, zero-crossing rate, etc. This application embodiment does not specifically limit this. Of course, any method for calculating the above-mentioned target feature information is applicable to this application embodiment, and this application embodiment does not specifically limit this.

[0154] For example, in one implementation, during the process of calculating the specified feature information of the target audio signal, the target audio signal can be sliced, and the specified feature information can be calculated for each slice of audio signal. Based on the feature information of each slice of audio signal, the target feature information can be determined.

[0155] Step E2: Perform a matching analysis on the target feature information and the reference feature information to obtain the matching analysis result; wherein, the reference feature information is the specified feature information of the predetermined noise signal;

[0156] It should be emphasized that the reference feature information can be the specified feature information of the predetermined noise signal in the current scene, that is, the energy attribute information of the predetermined noise signal in the current scene; the predetermined noise signal is related to the current scene. For example, in the shooting range scene, the predetermined noise signal can be the gunshot signal; in the industrial scene, the predetermined noise signal can be the noise signal emitted by machinery.

[0157] It is understood that matching analysis can be performed on the specified feature information of the target audio signal and the specified feature information of the predetermined noise signal in the current scene to obtain the matching analysis result. The matching analysis result can represent whether they match or not. The matching analysis can analyze whether the specified feature information of the target audio signal is similar to the specified feature information of the predetermined noise signal in the current scene. If the similarity exceeds a threshold, the target feature information is considered to match the reference feature information. Specifically, the threshold can be 0.98. If the similarity exceeds this threshold, the target feature information is considered to match the reference feature information. This application embodiment does not specifically limit the threshold.

[0158] Step E3: If the matching analysis result shows that the target feature information does not match the reference feature information, then the step of performing active noise reduction processing on the target audio signal based on the received second and third audio signals to obtain the noise-reduced target audio signal is executed.

[0159] Understandably, if the matching analysis results show that the target feature information does not match the reference feature information, that is, the similarity between the target feature information and the reference feature information is lower than the threshold, the step of actively denoising the target audio signal based on the received second and third audio signals can be performed to obtain the denoised target audio signal. Conversely, if the matching analysis results show that the target feature information matches the reference feature information, that is, the similarity between the target feature information and the reference feature information is higher than the threshold, it can be considered that the target audio signal still contains noise signals. In this case, subsequent active denoising processing will not be performed to prevent damage to the hearing of on-site personnel after the audio signal is output.

[0160] Furthermore, the process of steps E1-E3 can also be called "energy output determination", and this application embodiment does not specifically limit it.

[0161] As can be seen, the embodiments of this application can perform matching analysis on target feature information and reference feature information to obtain matching analysis results. When the matching analysis results show that the target feature information and the reference feature information do not match, the step of actively denoising the target audio signal based on the received second audio signal and third audio signal can be executed to obtain the denoised target audio signal. This can provide hearing protection for on-site personnel in a strong noise environment without damaging their hearing.

[0162] Alternatively, in another embodiment, such as Figure 9 As shown, the training process of the audio extraction model includes:

[0163] S901, Obtain the sample training set;

[0164] The training set of samples includes a first sample audio signal and a ground truth audio signal corresponding to the first sample audio signal; the first sample audio signal is a synthesized audio signal consisting of a reference audio signal that is a valid audio signal and an environmental noise signal, and the ground truth audio signal corresponding to the first sample audio signal is a reference audio signal used when synthesizing the first sample audio signal.

[0165] It is understandable that the sample training set may contain the first sample audio signal and the ground truth audio signal corresponding to the first sample audio signal. The first sample audio signal may contain the reference audio signal of the valid audio signal and the environmental noise signal. The reference audio signal of the valid audio signal may include speech signals and non-speech signals, while the environmental noise signal may be noise signals in various scenarios. The reference audio signal of the valid audio signal can be used as the ground truth corresponding to the first sample audio signal.

[0166] For example, when establishing a sample training set, a first dataset corresponding to environmental noise signals can be established, and a second dataset corresponding to reference audio signals belonging to valid audio signals can be established. The data entries of the first dataset and the second dataset are the same. Each environmental noise signal in the first dataset can be synthesized one-to-one with each reference audio signal in the second dataset to obtain the sample training set. The data entries in the sample training set are also the same as the data entries in the first dataset and the second dataset.

[0167] S902, the first sample audio signal in the sample training set is input into the audio extraction model, so that the frequency domain signal conversion module in the audio extraction model converts the first sample audio signal into a frequency domain signal, and inputs the real part information and imaginary part information of the frequency domain signal into the first encoder, so that the first encoder performs feature encoding processing on the received real part information and imaginary part information respectively, and transmits the encoded features corresponding to the real part information and the encoded features corresponding to the imaginary part information to the first decoder through the liquid neural network (LNN), so that the first decoder decodes the received encoded features and transmits the decoded features to the inverse processing module to convert them into a time domain signal, thereby obtaining the output result corresponding to the first sample audio signal;

[0168] It is understandable that the current audio extraction model is the audio extraction model to be trained, but the processing flow of the internal modules is consistent with that of the trained audio extraction model, so it will not be elaborated on here. Therefore, inputting the first sample audio signal from the training set into the audio extraction model will cause the audio extraction model to output the result corresponding to the first sample audio signal.

[0169] S903, based on the true audio signal of the first sample audio signal and the corresponding output result of the first sample audio signal, the model loss of the audio extraction model is calculated using a pre-set first loss function;

[0170] It is understandable that by substituting the true audio signal of the first sample audio signal and the corresponding output result of the first sample audio signal into the first loss function set in advance, the loss of the audio extraction model can be calculated.

[0171] For example, in one implementation, the first loss function includes:

[0172] ;

[0173] Where L is the loss function, y is the output result corresponding to the first sample audio signal, y' is the true audio signal of the first sample audio signal, and δ and e are both natural numbers with a value of 2.71.

[0174] Understandably, based on the absolute value of the difference between the output result corresponding to the first sample audio signal and the true audio signal of the first sample audio signal, different expressions are used to calculate the loss of the audio extraction model. When the absolute value of the difference between the output result corresponding to the first sample audio signal and the true audio signal of the first sample audio signal is less than or equal to δ, the appropriate expression is selected. To calculate the model loss of the audio extraction model; conversely, to select an expression. To calculate the model loss of the audio extraction model.

[0175] S904, when it is determined that the audio extraction model has not converged based on the model loss, the model parameters of the audio extraction model are adjusted, and the process returns to the step of obtaining the sample training set.

[0176] Understandably, if the audio extraction model fails to converge based on the model loss, it can be assumed that there is a significant difference between the ground truth audio signal of the first sample audio signal and the corresponding output result. In this case, the model parameters of the audio extraction model can be adjusted, and the process can return to the step of obtaining the sample training set until the audio extraction model converges based on the model loss, that is, the difference between the ground truth audio signal of the first sample audio signal and the corresponding output result is small. At this point, the training of the audio extraction model can be considered complete.

[0177] For example, in one implementation, the correction processing model can also be pre-trained, and the training process may include steps F1-F4:

[0178] Step F1: Obtain the target sample training set; wherein, the target sample training set contains the second sample audio signal and the ground truth audio signal corresponding to the second sample audio signal; the second sample audio signal is an audio signal with blurred audio and missing words, and the ground truth audio signal corresponding to the second sample audio signal is a clear audio signal corresponding to the second sample audio signal;

[0179] Step F2: The second sample audio signal in the target sample training set is input into the correction processing model so that the correction processing model outputs the output result corresponding to the second sample audio signal.

[0180] Step F3: Based on the true audio signal of the second sample audio signal and the corresponding output result of the second sample audio signal, calculate the model loss of the correction processing model using a pre-set first loss function;

[0181] Step F4: If the model loss indicates that the correction model has not converged, adjust the model parameters of the correction model and return to the step of obtaining the sample training set.

[0182] It is understandable that steps F1-F4 are similar to the training process of the audio extraction model described above, and the processing process of the correction model in step F2 is also similar to that described in step C111 above, so they will not be elaborated on here.

[0183] For example, in another implementation, the signal cancellation model can also be pre-trained. Since the modules in the signal cancellation model are similar to those in the audio extraction model, the training process for the signal cancellation model can be found in the training process of the audio extraction model described above, so it will not be elaborated on here.

[0184] As can be seen, the embodiments of this application can be pre-trained for the audio extraction model so that the trained audio extraction model can extract and process effective audio signals, thereby providing hearing protection for on-site personnel while ensuring that on-site personnel can receive effective information from the outside world.

[0185] Based on the above method embodiments, in order to more clearly understand the solution, the structure and working principle of the hearing protection device are described below, as shown in Figures 10(a), 10(b), and 10(c):

[0186] Figure 10(a) shows the structure of the hearing protection device. The left earcup contains a data acquisition DMIC, a speaker, a feedforward microphone, and a feedback microphone. The data acquisition DMIC is the digital microphone described above. The right earcup contains a DSP, an active noise cancellation on-chip system, a speaker, Bluetooth, a volume control, a feedforward microphone, and a feedback microphone. Each of the models described in the above embodiment can be located in the DSP, where data processing is performed. The active noise cancellation on-chip system is used for active noise cancellation processing.

[0187] Figure 10(b) is a schematic diagram of the working principle of the hearing protection device. The external sound source can contain effective audio signals and environmental noise signals. Based on physical sound insulation, part of the audio signal can be output to the human ear. The digital microphone can also collect the external sound source to obtain the first audio signal, and extract the first audio signal to obtain the target audio signal. Then, the target audio signal is actively denoised, and the denoised signal is output through the speaker. The feedback microphone can collect the audio signal output by the speaker, and can also collect the audio signal after physical sound insulation and input them to the active noise reduction unit for active noise reduction.

[0188] Figure 10(c) is a schematic diagram of the working principle of the hearing protection device, but the working principle in Figure 10(c) is more detailed. The ambient sound source includes human voice and noise. Both the AMIC (shell) and DMIC (shell) can collect ambient audio. The ambient audio can also be physically denoised through the noise reduction physical structure and output to the inner cavity of the earmuff (human ear). The AMIC (inner cavity) is the feedback microphone described above, the AMIC (shell) is the feedforward microphone described above, and the DMIC (shell) is the digital microphone described above. The DMIC (shell) can perform AGC (Automatic Gain Control) on the first audio signal collected, and extract the audio signal after automatic gain control to obtain the target audio signal. The target audio signal is energy output determined. When the target feature information of the target audio signal does not match the reference feature information, the target audio signal is output to the volume regulator so that the volume regulator can control the volume, and the audio signal after volume control is output to the active noise reduction unit. The active noise cancellation unit can receive three inputs: one is the audio signal from the volume control, another is the audio signal from inside the earcup collected by the AMIC (inner cavity), which has been converted into a digital signal by the ADC, and the third is the output from the AMIC (outer shell). The AMIC (outer shell) can convert the collected audio signal (analog signal) into a data signal by the ADC and output the digital signal to the active noise cancellation unit. After receiving the three inputs, the active noise cancellation unit can perform active noise cancellation to obtain the noise-reduced audio signal (digital signal), and then convert the noise-reduced audio signal into an analog signal by the DAC, and output it to the inner cavity of the earcup through the speaker, that is, to the ear.

[0189] As can be seen, the embodiments of this application extract and process the effective audio signal, and reduce the noise signal through active noise reduction, thereby providing hearing protection for on-site personnel while ensuring that they can receive effective information from the outside world. Furthermore, the audio extraction model in this application for extracting and processing the effective audio signal may include a frequency domain signal conversion module, a first encoder, a liquid neural network (LNN), a first decoder, and an inverse processing module corresponding to the frequency domain signal conversion module. Compared to related technologies, the audio extraction model in this application is not composed of a large neural network model; therefore, the structure of the audio extraction model in this application is simpler and more lightweight.

[0190] Based on the above method embodiments, such as Figure 11As shown, this application provides an audio output device based on a hearing protection device, applied to the control unit of the hearing protection device. The hearing protection device includes earmuffs with a predetermined physical noise reduction function. A digital microphone and a feedforward microphone belonging to the analog microphone are disposed on the outside of the earmuffs, and a speaker and a feedback microphone belonging to the analog microphone are disposed inside the earmuffs. The device includes:

[0191] The receiving module 1110 is used to receive the first audio signal acquired by the digital microphone, the second audio signal acquired by the feedforward microphone, and the third audio signal acquired by the feedback microphone and physically denoised by the earcups.

[0192] The extraction processing module 1120 is used to extract effective audio signals from the received first audio signal based on a predetermined feature audio extraction method to obtain a target audio signal. The predetermined feature audio extraction method includes signal processing based on a pre-trained audio extraction model for extracting effective audio signals. The audio extraction model is composed of a frequency domain signal conversion module, a first encoder, a liquid neural network (LNN), a first decoder, and an inverse processing module corresponding to the frequency domain signal conversion module connected in series. Both the first encoder and the first decoder belong to convolutional neural networks (CNNs).

[0193] The active noise reduction processing module 1130 is used to perform active noise reduction processing on the target audio signal based on the received second audio signal and third audio signal to obtain the noise-reduced target audio signal.

[0194] The output module 1140 is used to output the noise-reduced target audio signal through the speaker of the earcups.

[0195] Optionally, before signal processing, the predetermined feature audio extraction method based on the pre-trained audio extraction model for extracting effective audio signals further includes:

[0196] For the received first audio signal, perform periodic signal cancellation processing;

[0197] The signal processing based on the pre-trained audio extraction model for extracting valid audio signals includes:

[0198] Signal processing is performed on the eliminated audio signal based on a pre-trained audio extraction model for extracting valid audio signals.

[0199] Optionally, the output of the audio extraction model includes both speech and non-speech signals.

[0200] The pre-trained audio extraction model for extracting effective audio signals, after signal processing of the eliminated audio signal, further includes the predetermined feature audio extraction method as follows:

[0201] The speech signal in the output of the audio extraction model is subjected to specified correction processing;

[0202] Accordingly, the target audio signal includes two audio signals, wherein one of the two audio signals is the non-speech signal in the output result of the audio extraction model, and the other signal is the speech signal obtained by performing specified correction processing on the speech signal in the output result of the audio extraction model.

[0203] Optionally, the speech signal in the output of the audio extraction model is subjected to specified correction processing, including:

[0204] The speech signal in the output of the audio extraction model is subjected to speech content correction processing;

[0205] The speech signal after speech content correction is processed to eliminate signals of a specified type.

[0206] The specified type of signal includes the voice signal emitted by the wearer of the hearing protection device and / or the signal of a predetermined category of ambient sound.

[0207] Optionally, the specified correction processing of the speech signal in the output of the audio extraction model further includes:

[0208] The speech signal after the elimination of a specified type of signal is subjected to pitch adjustment processing.

[0209] Optionally, the training process of the audio extraction model includes:

[0210] Obtain a sample training set; wherein the sample training set includes a first sample audio signal and a ground truth audio signal corresponding to the first sample audio signal; the first sample audio signal is a synthesized audio signal based on a reference audio signal belonging to a valid audio signal and an environmental noise signal, and the ground truth audio signal corresponding to the first sample audio signal is a reference audio signal used when synthesizing the first sample audio signal;

[0211] The first sample audio signal in the training set is input into the audio extraction model, so that the frequency domain signal conversion module in the audio extraction model converts the first sample audio signal into a frequency domain signal, and inputs the real part and imaginary part of the frequency domain signal into the first encoder, so that the first encoder performs feature encoding processing on the received real part and imaginary part information respectively, and transmits the encoded features corresponding to the real part information and the encoded features corresponding to the imaginary part information to the first decoder through the liquid neural network (LNN), so that the first decoder decodes the received encoded features and transmits the decoded features to the inverse processing module to convert them into a time domain signal, thereby obtaining the output result corresponding to the first sample audio signal;

[0212] Based on the true audio signal of the first sample audio signal and the corresponding output result of the first sample audio signal, the model loss of the audio extraction model is calculated using a pre-set first loss function.

[0213] If the audio extraction model is determined to have failed to converge based on the model loss, the model parameters of the audio extraction model are adjusted, and the process returns to the step of obtaining the sample training set.

[0214] Optionally, the first loss function includes:

[0215] ;

[0216] Where L is the loss function, y is the output result corresponding to the first sample audio signal, y' is the true audio signal of the first sample audio signal, and δ and e are both natural numbers with a value of 2.71.

[0217] Optionally, the process of eliminating periodic signals for the received first audio signal includes:

[0218] The received first audio signal is delayed for a predetermined duration to obtain the delayed audio signal;

[0219] The periodic signal in the delayed audio signal is filtered by a preset filter to obtain a periodic signal.

[0220] The periodic signal in the first acquired audio signal is eliminated.

[0221] Optionally, the step of performing speech content correction processing on the speech signal in the output of the audio extraction model includes:

[0222] Based on a pre-trained correction processing model for speech content correction, speech content correction processing is performed on the speech signal in the output of the audio extraction model.

[0223] The correction processing model consists of a spectrum conversion unit, a second encoder, a gated loop unit, a preset diffusion model, and a second decoder.

[0224] The spectrum conversion unit is used to perform spectrum conversion on the speech signal in the output of the audio extraction model to obtain the target spectrum.

[0225] The first encoder is used to extract features from the target spectrum to obtain current feature information;

[0226] The gated loop unit is used to acquire historical feature information and input the historical feature information and the current feature information into the diffusion model; wherein, the historical feature information is the feature information obtained at the moment before the feature advance time of the current feature information;

[0227] The diffusion model is used to generate parameters to be utilized based on the historical feature information and the current feature information, and input the parameters to be utilized, the historical feature information and the current feature information into the second decoder; wherein, the parameters to be utilized are the convolution parameters to be utilized by the second decoder;

[0228] The second decoder is used to generate a speech signal after speech content correction processing based on the parameters to be used, historical feature information and current feature information.

[0229] Optionally, before performing active noise reduction processing on the target audio signal based on the received second and third audio signals to obtain the noise-reduced target audio signal, the method further includes:

[0230] Calculate the specified feature information of the target audio signal as the target feature information;

[0231] The target feature information and the reference feature information are matched and analyzed to obtain the matching analysis result; wherein, the reference feature information is the specified feature information of the predetermined noise signal;

[0232] If the matching analysis result indicates that the target feature information does not match the reference feature information, then the step of performing active noise reduction processing on the target audio signal based on the received second and third audio signals to obtain the noise-reduced target audio signal is executed.

[0233] In the technical solution of this application, the operations of obtaining, storing, using, processing, transmitting, providing and disclosing user personal information are all carried out with the user's authorization.

[0234] This application also provides an electronic device, such as... Figure 12 As shown, it includes:

[0235] Memory 1201 is used to store computer programs;

[0236] When the processor 1202 executes the program stored in the memory 1201, it implements the above-described audio output method based on the hearing protection device.

[0237] Furthermore, the aforementioned electronic device may also include a communication bus and / or a communication interface, with the processor 1202, the communication interface, and the memory 1201 communicating with each other via the communication bus.

[0238] The communication bus mentioned in the above electronic devices can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into address bus, data bus, control bus, etc. For ease of illustration, only one thick line is used to represent it in the diagram, but this does not mean that there is only one bus or one type of bus.

[0239] The communication interface is used for communication between the aforementioned electronic devices and other devices.

[0240] The memory may include random access memory (RAM) or non-volatile memory (NVM), such as at least one disk storage device. Optionally, the memory may also be at least one storage device located remotely from the aforementioned processor.

[0241] The processors mentioned above can be general-purpose processors, including central processing units (CPUs), network processors (NPs), etc.; they can also be digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.

[0242] In another embodiment provided in this application, a computer-readable storage medium is also provided, which stores a computer program that, when executed by a processor, implements the above-described audio output method based on a hearing protection device.

[0243] In another embodiment provided in this application, a computer program product containing instructions is also provided, which, when run on a computer, causes the computer to execute the above-described audio output method based on a hearing protection device.

[0244] In the above embodiments, implementation can be achieved entirely or partially through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented entirely or partially in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a solid-state drive (SSD), etc.

[0245] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0246] The various embodiments in this specification are described in a related manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the system embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions of the method embodiments.

[0247] The above description is merely a preferred embodiment of this application and is not intended to limit the scope of protection of this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application are included within the scope of protection of this application.

Claims

1. An audio output method based on a hearing protection device, characterized in that, A control unit for a hearing protection device, the hearing protection device comprising earmuffs with a predetermined physical noise reduction function, a digital microphone and a feedforward microphone belonging to the analog microphone disposed on the outer side of the earmuffs, and a speaker and a feedback microphone belonging to the analog microphone disposed inside the earmuffs; the method comprising: The system receives a first audio signal acquired by the digital microphone, a second audio signal acquired by the feedforward microphone, and a third audio signal acquired by the feedback microphone and subjected to physical noise reduction by the earcups. Based on a predetermined feature audio extraction method, the received first audio signal is processed to extract effective audio signals to obtain a target audio signal; wherein, the predetermined feature audio extraction method includes: signal processing based on a pre-trained audio extraction model for extracting effective audio signals; the audio extraction model is composed of a frequency domain signal conversion module, a first encoder, a liquid neural network (LNN), a first decoder, and an inverse processing module corresponding to the frequency domain signal conversion module connected in series, wherein the first encoder and the first decoder both belong to a convolutional neural network (CNN); Based on the received second and third audio signals, the target audio signal is subjected to active noise reduction processing to obtain the noise-reduced target audio signal. The noise-reduced target audio signal is output through the speakers of the earcups.

2. The method according to claim 1, characterized in that, Before signal processing, the predetermined feature audio extraction method based on the pre-trained audio extraction model for extracting effective audio signals further includes: For the received first audio signal, perform periodic signal cancellation processing; The signal processing based on the pre-trained audio extraction model for extracting valid audio signals includes: Signal processing is performed on the eliminated audio signal based on a pre-trained audio extraction model for extracting valid audio signals.

3. The method according to claim 2, characterized in that, The output of the audio extraction model includes two signals: speech signal and non-speech signal. The pre-trained audio extraction model for extracting effective audio signals, after signal processing of the eliminated audio signal, further includes the predetermined feature audio extraction method as follows: The speech signal in the output of the audio extraction model is subjected to specified correction processing; Accordingly, the target audio signal includes two audio signals, wherein one of the two audio signals is the non-speech signal in the output result of the audio extraction model, and the other signal is the speech signal obtained by performing specified correction processing on the speech signal in the output result of the audio extraction model.

4. The method according to claim 3, characterized in that, The audio signal in the output of the audio extraction model is subjected to specified correction processing, including: The speech signal in the output of the audio extraction model is subjected to speech content correction processing; The speech signal after speech content correction is processed to eliminate signals of a specified type. The specified type of signal includes the voice signal emitted by the wearer of the hearing protection device and / or the signal of a predetermined category of ambient sound.

5. The method according to claim 4, characterized in that, The specified correction process for the speech signal in the output of the audio extraction model also includes: The speech signal after the elimination of a specified type of signal is subjected to pitch adjustment processing.

6. The method according to claim 1, characterized in that, The training process of the audio extraction model includes: Obtain a sample training set; wherein the sample training set includes a first sample audio signal and a ground truth audio signal corresponding to the first sample audio signal; the first sample audio signal is a synthesized audio signal based on a reference audio signal belonging to a valid audio signal and an environmental noise signal, and the ground truth audio signal corresponding to the first sample audio signal is a reference audio signal used when synthesizing the first sample audio signal; The first sample audio signal in the training set is input into the audio extraction model, so that the frequency domain signal conversion module in the audio extraction model converts the first sample audio signal into a frequency domain signal, and inputs the real part and imaginary part of the frequency domain signal into the first encoder, so that the first encoder performs feature encoding processing on the received real part and imaginary part information respectively, and transmits the encoded features corresponding to the real part information and the encoded features corresponding to the imaginary part information to the first decoder through the liquid neural network (LNN), so that the first decoder decodes the received encoded features and transmits the decoded features to the inverse processing module to convert them into a time domain signal, thereby obtaining the output result corresponding to the first sample audio signal; Based on the true audio signal of the first sample audio signal and the corresponding output result of the first sample audio signal, the model loss of the audio extraction model is calculated using a pre-set first loss function. If the audio extraction model is determined to have failed to converge based on the model loss, the model parameters of the audio extraction model are adjusted, and the process returns to the step of obtaining the sample training set.

7. The method according to claim 6, characterized in that, The first loss function includes: ; Where L is the loss function, y is the output result corresponding to the first sample audio signal, y' is the true audio signal of the first sample audio signal, and δ and e are both natural numbers.

8. The method according to claim 2, characterized in that, The process of eliminating periodic signals from the received first audio signal includes: The received first audio signal is delayed for a predetermined duration to obtain the delayed audio signal; The periodic signal in the delayed audio signal is filtered by a preset filter to obtain a periodic signal. The periodic signal in the first acquired audio signal is eliminated.

9. The method according to claim 4, characterized in that, The speech content correction processing of the speech signal in the output of the audio extraction model includes: Based on a pre-trained correction processing model for speech content correction, speech content correction processing is performed on the speech signal in the output of the audio extraction model. The correction processing model consists of a spectrum conversion unit, a second encoder, a gated loop unit, a preset diffusion model, and a second decoder. The spectrum conversion unit is used to perform spectrum conversion on the speech signal in the output of the audio extraction model to obtain the target spectrum. The first encoder is used to extract features from the target spectrum to obtain current feature information; The gated loop unit is used to acquire historical feature information and input the historical feature information and the current feature information into the diffusion model; wherein, the historical feature information is the feature information obtained at the moment before the feature advance time of the current feature information; The diffusion model is used to generate parameters to be utilized based on the historical feature information and the current feature information, and input the parameters to be utilized, the historical feature information and the current feature information into the second decoder; wherein, the parameters to be utilized are the convolution parameters to be utilized by the second decoder; The second decoder is used to generate a speech signal after speech content correction processing based on the parameters to be used, historical feature information and current feature information.

10. The method according to claim 1, characterized in that, Before performing active noise reduction processing on the target audio signal based on the received second and third audio signals to obtain the noise-reduced target audio signal, the method further includes: Calculate the specified feature information of the target audio signal as the target feature information; The target feature information and the reference feature information are matched and analyzed to obtain the matching analysis result; wherein, the reference feature information is the specified feature information of the predetermined noise signal; If the matching analysis result indicates that the target feature information does not match the reference feature information, then the step of performing active noise reduction processing on the target audio signal based on the received second and third audio signals to obtain the noise-reduced target audio signal is executed.

11. An audio output device based on a hearing protection device, characterized in that, A control unit for a hearing protection device, the hearing protection device comprising earmuffs with a predetermined physical noise reduction function, a digital microphone and a feedforward microphone belonging to the analog microphone disposed on the outer side of the earmuffs, and a speaker and a feedback microphone belonging to the analog microphone disposed inside the earmuffs; the device includes: The receiving module is used to receive the first audio signal acquired by the digital microphone, the second audio signal acquired by the feedforward microphone, and the third audio signal acquired by the feedback microphone after physical noise reduction by the earcups. An extraction processing module is used to extract effective audio signals from a received first audio signal based on a predetermined feature audio extraction method to obtain a target audio signal. The predetermined feature audio extraction method includes signal processing based on a pre-trained audio extraction model for extracting effective audio signals. The audio extraction model is composed of a frequency domain signal conversion module, a first encoder, a liquid neural network (LNN), a first decoder, and an inverse processing module corresponding to the frequency domain signal conversion module connected in series. Both the first encoder and the first decoder belong to convolutional neural networks (CNNs). An active noise reduction processing module is used to perform active noise reduction processing on the target audio signal based on the received second audio signal and third audio signal to obtain the noise-reduced target audio signal. The output module is used to output the noise-reduced target audio signal through the speaker of the earcups.

12. An electronic device, characterized in that, include: Memory, used to store computer programs; A processor, when executing a program stored in memory, implements the method described in any one of claims 1-10.

13. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the method described in any one of claims 1-10.