Noise reduction processing method and device of earphone and earphone
By acquiring feedback and feedforward microphone signals from the headphones, determining the channel transfer function, and performing inverse cancellation processing, the problem of blockage effect and external sound cancellation when wearing headphones is solved, achieving an effective noise reduction effect.
Patent Information
- Application Number
- CN202411984666.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-30
- Publication Date
- 2025-11-25
- Estimated Expiration
- 2044-12-30
AI Technical Summary
In existing technologies, when users wear headphones, the low-frequency energy of their voice increases due to the blockage effect, resulting in a muddy sound. Furthermore, feedback noise reduction methods can cancel out external ambient sounds, making it difficult to hear external sounds.
By acquiring signals from the feedback microphone and the feedforward microphone, the actual secondary channel transfer function is determined, the effective sound signal is determined based on the signal differences, and phase inversion cancellation processing is performed to output the target sound signal.
While preserving external ambient sounds, it eliminates the blockage effect and low-frequency noise, improves noise reduction, and avoids canceling out external ambient sounds.
Smart Images

Figure CN119893367B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of headphone technology, and more specifically, to a noise reduction processing method, device, and headphone for headphones. Background Technology
[0002] When a user wears headphones, part of the sound travels through the air to the ear canal and part travels through the Eustachian tube to the eardrum. Due to passive sound insulation, the sound transmitted through the air to the eardrum lacks high frequencies and amplifies low-frequency information, resulting in the headphone wearer hearing their own voice with a lot of low-frequency energy and a muddy sound, which is known as the blockage effect.
[0003] In related technologies, the blockage effect is eliminated by adding feedback noise reduction, which mainly eliminates the low-frequency energy. That is, an external sound signal is collected by a feedforward microphone, and the collected external sound signal is processed to generate an anti-phase sound wave to eliminate low-frequency energy. The anti-phase sound wave is then played through a speaker to eliminate the blockage effect. However, this method not only reduces the low-frequency energy of the user's own speech, but also cancels out the low-frequency energy of the external environmental signal listened to by the feedforward microphone, causing the user to not be able to hear the external environmental sound clearly. Summary of the Invention
[0004] One objective of this application is to provide a new technical solution for noise reduction processing of headphones, in order to solve the problem in the prior art where users cannot hear ambient sounds clearly due to the addition of feedback noise reduction that mainly eliminates low-frequency energy.
[0005] According to a first aspect of the present invention, a noise reduction processing method for headphones is provided, comprising:
[0006] Acquire the first sound signal collected by the feedback microphone and the second sound signal collected by the feedforward microphone;
[0007] The actual secondary channel transfer function is determined based on the first audio signal and the second audio signal;
[0008] If the actual secondary channel transfer function is not equal to the reference secondary channel transfer function, a valid audio signal is determined based on the second audio signal and the reference secondary channel transfer function.
[0009] Based on the effective sound signal, the first sound signal is subjected to phase inversion cancellation processing to obtain and output the first target sound signal.
[0010] Optionally, the step of performing phase cancellation processing on the first sound signal based on the effective sound signal to obtain and output the first target sound signal includes:
[0011] Based on the valid sound signal, the first sound signal is subjected to signal separation processing to obtain a third sound signal; wherein, the third sound signal does not include the valid sound signal;
[0012] The third sound signal is inverted to obtain the first target sound signal;
[0013] The first target sound signal is output through a speaker.
[0014] Optionally, determining the actual secondary channel transfer function based on the first audio signal and the second audio signal includes:
[0015] The second sound signal is filtered to obtain a filtered second sound signal;
[0016] The actual secondary channel transfer function is determined based on the filtered second audio signal and the first audio signal.
[0017] The step of determining the valid audio signal based on the second audio signal and the reference secondary channel transfer function includes:
[0018] The effective audio signal is determined based on the filtered second audio signal and the reference secondary channel transfer function.
[0019] Optionally, the step of filtering the second audio signal to obtain a filtered second audio signal includes:
[0020] Identify the target scene category corresponding to the second sound signal;
[0021] Based on the target scene category, determine the corresponding target filtering parameters;
[0022] The second sound signal is filtered according to the target filtering parameters to obtain the filtered second sound signal.
[0023] Optionally, the identification of the target scene category corresponding to the second sound signal includes:
[0024] Feature extraction is performed on the second sound signal to obtain sound signal features;
[0025] The sound signal features are input into the scene recognition model to obtain the target scene category corresponding to the sound signal features.
[0026] Optionally, determining the target filtering parameters corresponding to the target scene category based on the target scene category includes:
[0027] Obtain historical recognition data of the scene recognition model; wherein, the historical recognition data includes the estimated scene category and the actual scene category of the scene recognition model corresponding to each historical scene recognition;
[0028] The model evaluation value is obtained based on the estimated scene category and the actual scene category in the historical identification data;
[0029] If the model evaluation value meets the preset conditions, the target filtering parameters corresponding to the target scene category are determined.
[0030] Optionally, the step of filtering the second audio signal according to the target filtering parameters to obtain the filtered second audio signal includes:
[0031] If the recording duration is greater than or equal to the preset duration, the second sound signal is filtered according to the target filtering parameters to obtain the filtered second sound signal.
[0032] According to a second aspect of the present invention, a noise reduction processing device for headphones is also provided, comprising:
[0033] The acquisition module is used to acquire the first sound signal collected by the feedback microphone and the second sound signal collected by the feedforward microphone;
[0034] The determining module is used to determine the actual secondary channel transfer function based on the first audio signal and the second audio signal; and to determine the valid audio signal based on the second audio signal and the reference secondary channel transfer function when the actual secondary channel transfer function is not equal to the reference secondary channel transfer function.
[0035] The output module is used to perform phase cancellation processing on the first sound signal based on the effective sound signal to obtain and output the first target sound signal.
[0036] According to a third aspect of the present invention, a noise reduction processing apparatus for headphones is also provided, comprising a memory and a processor, the memory being configured to store executable instructions; the processor being configured to operate under the control of the instructions to perform the method as described in the first aspect.
[0037] According to a third aspect of the present invention, an earphone is also provided, including a feedback microphone, a feedforward microphone, and a noise reduction processing device for the earphone as described in the second or third aspect;
[0038] The feedback microphone collects the first sound signal and sends it to the noise reduction processing device of the earphone;
[0039] The feedforward microphone collects the second sound signal and transmits the second sound signal to the noise reduction processing device of the earphone.
[0040] One beneficial effect of this invention is that by acquiring a first sound signal collected by a feedback microphone and a second sound signal collected by a feedforward microphone, and determining the actual secondary channel transfer function based on the first and second sound signals, and when the actual secondary channel transfer function is not equal to the reference secondary channel transfer function, determining the effective sound signal based on the second sound signal and the reference secondary channel transfer function, and performing phase cancellation processing on the first sound signal based on the effective sound signal, a first target sound signal is obtained and output. This method can remove low-frequency noise from human speech amplified by passive sound insulation while retaining the effective components (i.e., retaining ambient sound) in the first sound signal collected by the feedback microphone, and can also avoid canceling ambient sound collected by the feedforward microphone, thus eliminating the blockage effect and improving the noise reduction effect. Attached Figure Description
[0041] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments of the invention and, together with their description, serve to explain the principles of the invention.
[0042] Figure 1 This is a schematic diagram of the hardware structure of the headphones according to an embodiment of this application;
[0043] Figure 2 This is a flowchart illustrating a noise reduction method for headphones according to an embodiment of the present invention.
[0044] Figure 3 This is a schematic diagram of the sound propagation path of an earphone according to an embodiment of the present invention;
[0045] Figure 4 This is a schematic diagram of the structure of a noise reduction processing device for headphones according to an embodiment of the present invention;
[0046] Figure 5 This is a schematic diagram of the structure of a noise reduction processing device for headphones according to another embodiment of the present invention;
[0047] Figure 6 This is a schematic diagram of the structure of an earphone according to an embodiment of the present invention. Detailed Implementation
[0048] Various exemplary embodiments of the present invention will now be described in detail with reference to the accompanying drawings. It should be noted that, unless otherwise specifically stated, the relative arrangement, numerical expressions, and values of the components and steps set forth in these embodiments do not limit the scope of the invention.
[0049] The following description of at least one exemplary embodiment is merely illustrative and is in no way intended to limit the invention or its application or use.
[0050] Techniques, methods, and equipment known to those skilled in the art may not be discussed in detail, but where appropriate, such techniques, methods, and equipment should be considered part of the specification.
[0051] In all the examples shown and discussed herein, any specific values should be interpreted as merely exemplary and not as limitations. Therefore, other examples of exemplary embodiments may have different values.
[0052] It should be noted that similar labels and letters in the following figures indicate similar items; therefore, once an item is defined in one figure, it does not need to be discussed further in subsequent figures.
[0053] <Hardware Configuration>
[0054] Figure 1 This is a block diagram of the hardware configuration of the earphone 100 according to an embodiment of this application.
[0055] like Figure 1 As shown, the headphones 100 include a feedback microphone 1000, a feedforward microphone 2000, and a noise reduction processing device 3000 for the headphones.
[0056] The feedback microphone 1000 is located inside the earphone and is used to collect the first sound signal. The feedback microphone is communicatively connected to the noise reduction processing unit 3000 of the earphone to send the first sound signal it has collected to the noise reduction processing unit 3000 of the earphone.
[0057] In this embodiment, refer to Figure 1 As shown, the feedback microphone 1000 may include a processor 1100, a memory 1200, a microphone 1300, a communication device 1400, etc.
[0058] Processor 1100 may be a mobile processor. Memory 1200 includes, for example, ROM (Read-Only Memory), RAM (Random Access Memory), non-volatile memory such as a hard disk, etc. Microphone 1300 can input voice information. Communication device 1400 is capable of wired or wireless communication, and may include short-range communication devices, such as any device that performs short-range wireless communication based on short-range wireless communication protocols such as Hilink, WiFi (IEEE 802.11), Mesh, Bluetooth, ZigBee, Thread, Z-Wave, NFC, UWB, LiFi, etc. Communication device 1400 may also include long-range communication devices, such as any device that performs WLAN, GPRS, 2G / 3G / 4G / 5G long-range communication.
[0059] In this embodiment, the memory 1200 of the feedback microphone 1000 is used to store instructions for controlling the processor 1100 to operate in order to at least execute the method performed by the feedback microphone according to any embodiment of the present invention. Those skilled in the art can design instructions based on the disclosed scheme of the present invention. How the instructions control the processor to operate is well known in the art and will not be described in detail here.
[0060] The feedforward microphone 2000 collects the second sound signal and sends the second sound signal to the noise reduction processing device 3000 of the headphones.
[0061] In this embodiment, refer to Figure 1 As shown, the feedforward microphone 2000 may include a processor 2100, a memory 2200, a microphone 2300, a communication device 2400, etc.
[0062] Processor 2100 may be a mobile processor. Memory 2200 includes, for example, ROM (Read-Only Memory), RAM (Random Access Memory), and non-volatile memory such as a hard disk. Microphone 2300 can input voice information. Communication device 2400 is capable of wired or wireless communication, for example. Communication device 2400 may include short-range communication devices, such as any device that performs short-range wireless communication based on short-range wireless communication protocols such as Hilink, WiFi (IEEE 802.11), Mesh, Bluetooth, ZigBee, Thread, Z-Wave, NFC, UWB, and LiFi. Communication device 2400 may also include long-range communication devices, such as any device that performs WLAN, GPRS, or 2G / 3G / 4G / 5G long-range communication.
[0063] In this embodiment, the memory 2200 of the feedforward microphone 2000 is used to store instructions for controlling the processor 2100 to operate to at least execute the method performed by the feedforward microphone according to any embodiment of the present invention. Those skilled in the art can design the instructions according to the disclosed scheme of the present invention. How the instructions control the processor to operate is well known in the art and will not be described in detail here.
[0064] The noise reduction processing device 3000 of the headphones is used to perform noise reduction processing based on the second sound signal and the first sound signal.
[0065] In this embodiment, refer to Figure 1 As shown, the noise reduction processing device 3000 for headphones may include a processor 3100, a memory 3200, a communication device 3300, a speaker 3400, etc.
[0066] Processor 3100 may be a mobile processor. Memory 3200 includes, for example, ROM (Read-Only Memory), RAM (Random Access Memory), and non-volatile memory such as a hard disk. Communication device 3300 may be capable of wired or wireless communication. Communication device 3300 may include short-range communication devices, such as any device that performs short-range wireless communication based on short-range wireless communication protocols such as Hilink, WiFi (IEEE 802.11), Mesh, Bluetooth, ZigBee, Thread, Z-Wave, NFC, UWB, and LiFi. Communication device 3300 may also include long-range communication devices, such as any device that performs WLAN, GPRS, or 2G / 3G / 4G / 5G long-range communication. Speaker 3400 outputs voice information.
[0067] In this embodiment, the memory 3200 of the headphone noise reduction processing device 3000 is used to store instructions for controlling the processor 3100 to operate to at least execute the method performed by the headphone noise reduction processing device according to any embodiment of the present invention. Those skilled in the art can design instructions according to the disclosed scheme of the present invention. How the instructions control the processor to operate is well known in the art and will not be described in detail here.
[0068] Despite Figure 1 The present invention illustrates multiple devices of the noise reduction processing device 3000 for headphones; however, the present invention may refer to only some of these devices. For example, the noise reduction processing device 3000 for headphones may refer only to the memory 3200 and the processor 3100.
[0069] <Method Implementation>
[0070] Figure 2 This is a schematic flowchart of a noise reduction processing method for headphones according to an embodiment of this application. The method is applied to a noise reduction processing device for headphones.
[0071] according to Figure 2 As shown, the noise reduction processing method for headphones in this embodiment may include the following steps S2100 to S2400:
[0072] Step S2100: Acquire the first sound signal collected by the feedback microphone and the second sound signal collected by the feedforward microphone.
[0073] In this embodiment, a feedforward microphone is positioned outside the earphone to collect external sound signals, i.e., the second sound signal. This second sound signal includes ambient noise and the wearer's voice. A feedback microphone is positioned near the eardrum and opposite the speaker to collect a first sound signal. The first sound signal includes the residual sound after passive sound insulation of the external sound signal and the sound played by the speaker after processing the second sound signal collected by the feedforward microphone through a monitoring filter.
[0074] For example, such as Figure 3 As shown, the noise reduction processing device of the headphones can be characterized by the control module H(s) and the storage reference secondary channel transfer function G(s) in the figure. The external sound signal is x(t), and the second sound signal collected by the feedforward microphone can be expressed as x(t)*R(s), where R(s) represents the transfer function of the external sound signal from the environment to the position of the feedforward microphone. The sound signal at the eardrum can be expressed as: e v (t)=x(t)*P(s)+x(t)*R(s)*T(s)*G v (s). Where P(s) is the transfer function of the external sound signal after passive sound insulation to the human ear position, T(s) is the transfer function of the monitoring filter, and G v (s) is the sound transfer function from the speaker to the ear. The sound signal collected by the feedback microphone is d(t), which is the first sound signal. This first sound signal includes the sound emitted by the speaker that is transmitted to the feedback microphone position and the sound signal transmitted to the feedback microphone position after the external sound signal is passively soundproofed.
[0075] It should be noted that P(s), R(s), G(s), and G... v (s) can all be obtained through actual measurement in the test environment. The test environment is a pressure field environment, and there are no special restrictions on the simulated human head used in the test.
[0076] Step S2200: Determine the actual secondary channel transfer function based on the first audio signal and the second audio signal.
[0077] In this embodiment, the actual secondary channel transfer function is... Figure 3 In It can refer to the transfer function of the sound signal output by the speaker being transmitted to the location of the feedback microphone.
[0078] In this embodiment, the output signal of the secondary channel can be determined based on the first audio signal, the input signal of the secondary channel can be determined based on the second audio signal, and then the actual transfer function of the secondary channel can be determined based on the output signal and the input signal of the secondary channel.
[0079] For example, the first audio signal can be used as the output signal of the secondary channel, and the signal output from the second audio signal after passing through the monitoring filter can be used as the input signal of the secondary channel. The ratio of the output signal to the input signal can be calculated to obtain the actual secondary channel transfer function. Alternatively, the secondary channel transfer model can be determined based on the output signal and the input signal, and this model is the actual secondary channel transfer function.
[0080] In some embodiments, determining the actual secondary channel transfer function based on the first audio signal and the second audio signal in step S2200 includes: steps S2200.1 to S2200.2.
[0081] Step S2200.1: Filter the second sound signal to obtain the filtered second sound signal.
[0082] In this embodiment, as Figure 3 As shown, the second audio signal is x(t)*R(s). After being filtered by a monitoring filter, the filtered second audio signal is obtained, which is x(t)*R(s)*T(s). The filtering process can include noise cancellation, frequency response adjustment, dynamic range compression, signal enhancement, etc., and is not limited here. Furthermore, these filtering processes can be characterized by the transfer function T(s) of the monitoring filter. The transfer function T(s) of the monitoring filter can be a transfer model.
[0083] Those skilled in the art should understand that the specific filtering methods described here are well-known in the field and will not be elaborated upon here.
[0084] Step S2200.2: Determine the actual secondary channel transfer function based on the filtered second audio signal and the first audio signal.
[0085] In this embodiment, the ratio of the first audio signal to the filtered second audio signal can be calculated as the actual secondary channel transfer function. Alternatively, the first audio signal and the filtered second audio signal can be input into a basic secondary channel transfer model to obtain the actual secondary channel transfer function. In this case, the actual secondary channel transfer function obtained is the model, and no limitation is made here.
[0086] According to the embodiments of this application, by directly using the filtered second audio signal as the input signal of the secondary channel to calculate the actual secondary channel transfer function, the complexity of the calculation can be reduced and the efficiency of noise reduction processing can be improved.
[0087] Step S2300: If the actual secondary channel transfer function is not equal to the reference secondary channel transfer function, a valid audio signal is determined based on the second audio signal and the reference secondary channel transfer function.
[0088] In this embodiment, the reference secondary channel transfer function G(s) can be a secondary channel transfer function set before the headphones leave the factory. It can be obtained by playing a known signal (such as white noise or a swept signal) in a controlled environment and measuring the frequency response between the speaker output and the feedback microphone input.
[0089] Actual secondary channel transfer function This is an approximation of the reference secondary channel transfer function G(s). If the actual secondary channel transfer function is equal to the reference secondary channel transfer function, that is, the secondary channel estimation is accurate, it means that the inverted sound wave generated after processing the second sound signal collected by the feedforward microphone will completely cancel the noise (i.e., the low-frequency noise amplified by passive sound insulation when a person speaks), and the first sound signal collected by the feedback microphone does not include noise components not captured by the feedforward microphone. At this time, the inverted sound wave generated after processing the second sound signal collected by the feedforward microphone is directly output through the speaker. The filtered second sound signal (i.e., the signal the user wants to hear) output by the monitoring filter after processing the second sound signal collected by the feedforward microphone does not enter the noise reduction processing device of the headphones for further processing.
[0090] If the actual secondary channel transfer function is not equal to the reference secondary channel transfer function, i.e., the estimation is inaccurate, the inverted sound wave output after processing the second sound signal acquired by the feedforward microphone cannot completely cancel the low-frequency noise. The first sound signal includes noise components not captured by the feedforward microphone (i.e., low-frequency noise amplified by passive sound insulation from human speech) and effective sound components (i.e., the effective sound signal). In order to cancel the noise components not captured by the feedforward microphone in the first sound signal, it is necessary to first remove the effective sound components (i.e., the effective sound signal) from the first sound signal.
[0091] Those skilled in the art will understand that the specific process of generating an antiphase sound wave by processing the second sound signal acquired by the feedforward microphone is well known in the art and will not be described in detail here. For example, the second sound signal may be separated to obtain a noise signal, and then an antiphase sound wave with the opposite phase and the same amplitude as the noise signal may be generated.
[0092] In this embodiment, the effective sound signal can be calculated using the second sound signal and the actual secondary channel transfer function. The effective sound signal is the estimated value of the effective sound components in the first sound signal acquired by the feedback microphone.
[0093] In some embodiments, determining the valid audio signal in step S2300 based on the second audio signal and the reference secondary channel transfer function includes:
[0094] The effective audio signal is determined based on the filtered second audio signal and the reference secondary channel transfer function.
[0095] In this embodiment, the filtered second sound signal x'(t) has a larger proportion of effective sound components than the second sound signal (x(t)*T(s)), except for some noise components. Therefore, the effective sound signal can be calculated by multiplying the filtered second sound signal and the reference secondary channel transfer function.
[0096] Step S2400: Based on the effective sound signal, perform phase inversion cancellation processing on the first sound signal to obtain and output the first target sound signal.
[0097] In this embodiment, the first sound signal includes noise components not captured by the feedforward microphone and effective sound components (i.e., the effective sound signal). The effective sound signal can be removed from the first sound signal to obtain the noise component signal not captured by the feedforward microphone, namely, the low-frequency noise amplified by passive sound insulation of human speech, referred to as the third sound signal. Then, the inverted sound wave of this third sound signal is generated and output through a speaker, thereby canceling out the low-frequency noise amplified by passive sound insulation in the first sound signal. This inverted sound wave of the third sound signal is the first target sound signal. The first target sound signal is... Figure 3 The example shown is y(t).
[0098] In some embodiments, step S2400, which involves performing phase cancellation processing on the first sound signal based on the valid sound signal to obtain and output the first target sound signal, includes steps S2400.1 to S2400.3.
[0099] Step S2400.1: Based on the valid sound signal, perform signal separation processing on the first sound signal to obtain the third sound signal.
[0100] In this embodiment, the effective sound signal is removed from the first sound signal to obtain the third sound signal. The third sound signal is the noise component that was not captured by the feedforward microphone, that is, the low-frequency noise of a person speaking through passive sound insulation amplification.
[0101] Step S2400.2: Invert the third sound signal to obtain the first target sound signal.
[0102] In this embodiment, the first target sound signal and the third sound signal have opposite phases but the same amplitude.
[0103] Step S2400.3: Send the first target sound signal to the speaker so that the speaker can play the first target sound signal.
[0104] According to an embodiment of this application, by performing signal separation processing on the first sound signal based on the effective sound signal to obtain a third sound signal, performing phase inversion processing on the third sound signal to obtain a first target sound signal, and sending the first target sound signal to a speaker for the speaker to play the first target sound signal, it is possible to remove noise signals in the first sound signal that are not captured by the feedforward microphone, that is, low-frequency noise amplified by passive sound insulation when a person speaks, and retain effective sound components, that is, retain external environmental sounds, which helps to improve the noise reduction effect and ensure that the sound heard by the user is as little as possible affected by noise interference.
[0105] According to an embodiment of this application, a first sound signal acquired by a feedback microphone and a second sound signal acquired by a feedforward microphone are obtained. Based on the first and second sound signals, an actual secondary channel transfer function is determined. If the actual secondary channel transfer function is not equal to a reference secondary channel transfer function, a valid sound signal is determined based on the second sound signal and the reference secondary channel transfer function. Based on the valid sound signal, the first sound signal is subjected to phase cancellation processing to obtain and output a first target sound signal. This method can remove low-frequency noise from human speech amplified by passive sound insulation while retaining the effective components (i.e., retaining ambient sound) in the first sound signal acquired by the feedback microphone. It also avoids canceling ambient sound acquired by the feedforward microphone, eliminating the blockage effect and improving noise reduction performance.
[0106] Traditional headphones typically use fixed filtering parameters for their monitoring filters, which cannot dynamically adjust based on the characteristics of ambient noise. This approach has several limitations. For example, in quiet environments, if the headphones' transparency mode uses fixed filtering parameters, ambient sounds (including background noise) may be excessively amplified or enhanced. In such cases, users may perceive the ambient sounds transmitted by the headphones as too strong or inappropriately amplified, which can affect listening comfort and even cause discomfort in some situations. Conversely, in noisy environments, the amount of noise transmitted cannot be effectively controlled, thus impacting the user experience.
[0107] The inventors proposed a method for adaptively configuring filter parameters based on the scene, which is executed by the headphone's noise reduction processing device. The details are as follows:
[0108] In some embodiments, step S2200.1, filtering the second sound signal to obtain a filtered second sound signal, includes steps SA1 to SA3.
[0109] Step SA1: Identify the target scene category corresponding to the second sound signal.
[0110] In this embodiment, the target scene category corresponding to the second sound signal is identified from the preset scene categories.
[0111] The target scene category corresponding to the second sound signal is identified using machine learning methods such as Support Vector Machine (SVM), Random Forest, Convolutional Neural Network (CNN), and Recurrent Neural Network (RNN). This target scene category can be, for example, a quiet outdoor environment, a busy street, or an industrial environment; there are no specific limitations here.
[0112] Those skilled in the art should understand that the preset scene categories can be set by technicians according to their needs. There is no limitation on specific scene categories here, nor is there a limitation on which preset scene category the target scene category should be.
[0113] In some embodiments, identifying the target scene category corresponding to the second sound signal in step SA1 includes steps SA11 and SA12.
[0114] Step SA11: Extract features from the second sound signal to obtain sound signal features.
[0115] In this embodiment, features can be extracted using a convolutional neural network. Specifically, the second audio signal can be input into a CNN, and the one-dimensional convolutional layers of the CNN automatically extract the local temporal features of the second audio signal. The pooling layers of the CNN can reduce the spatial dimension of the features, extract more abstract features, and increase the model's noise resistance. Other methods can also be used to extract features from the second audio signal to obtain audio signal features; this is not limited here.
[0116] Step SA12: Input the sound signal features into the scene recognition model to obtain the target scene category corresponding to the sound signal features.
[0117] In this embodiment, there are multiple sound signal features, and the scene recognition model is a multi-classification model. This multi-classification model can classify multiple sound signal features and determine the scene category corresponding to each sound signal feature. The scene categories corresponding to these multiple sound signal features can contain only one or more. In the case of multiple categories, the scene category with the highest score is determined as the target scene category.
[0118] For example, after extracting the sound signal features of the second sound signal through the convolutional and pooling layers of a CNN network, the learned high-level features are integrated through a fully connected layer and then output to a softmax layer for classification. The softmax layer converts the output of the fully connected layer into a probability distribution; that is, the softmax layer outputs a probability value for each scene category, representing the probability that the input sound signal feature belongs to that scene category. This probability value is between [0,1], and the sum of all probability values is 1. For example, if the softmax layer outputs 0.8 for the "traffic noise" scene, this means that the scene recognition model has 80% confidence that the current sound feature signal belongs to the traffic noise scene. Finally, the scene recognition model selects the scene category with the highest probability value as the final classification result, i.e., the target scene category.
[0119] Step SA2: Determine the corresponding target filtering parameters based on the target scene category.
[0120] In this embodiment, a predefined correspondence between scene categories and filtering parameters is established, meaning that a set of filtering parameters corresponds to each scene category.
[0121] In some embodiments, step SA2, which determines the target filtering parameters corresponding to the target scene category based on the target scene category, includes steps SA21 to SA23.
[0122] Step SA21: Obtain the historical recognition data of the scene recognition model.
[0123] In this embodiment, the scene recognition results of the scene recognition model for the previous N frames can be used as historical recognition data. This historical recognition data includes the model-predicted scene category and the actual scene category corresponding to each historical scene recognition (i.e., each frame).
[0124] Step SA22: Obtain the model evaluation value based on the estimated scene category and the actual scene category in the historical identification data.
[0125] In this embodiment, the model evaluation value can be at least one of the following: accuracy, precision, recall, confusion matrix, and recognition standard deviation, without limitation.
[0126] In some embodiments, the model evaluation values include at least one of accuracy and recognition standard deviation.
[0127] For example, the accuracy rate can be calculated by using the model's predicted scene category and the actual scene category from historical identification data to determine the proportion of samples correctly predicted by the model to the total number of samples.
[0128] For example, the model can estimate the scene category and the actual scene category based on historical recognition data, and then determine the recognition standard deviation. The recognition standard deviation quantifies the magnitude of the deviation between the recognition result and the average recognition result, i.e., the degree of fluctuation or dispersion of the result.
[0129] Step SA23: If the model evaluation value meets the preset conditions, determine the target filtering parameters corresponding to the target scene category.
[0130] In some embodiments, the model evaluation value includes accuracy, and the preset condition may be that the accuracy is greater than or equal to an accuracy threshold.
[0131] In other embodiments, the model evaluation values include accuracy and recognition standard deviation, and the preset conditions may be that the accuracy is greater than or equal to an accuracy threshold and the recognition standard deviation is less than or equal to a standard deviation threshold.
[0132] For example, if the model evaluation values are accuracy and recognition standard deviation, and the accuracy calculated based on historical recognition data is greater than the accuracy threshold and the recognition standard deviation is less than the standard deviation threshold, then the scene recognition model is considered to have a stable recognition result. In this case, the target filter parameters corresponding to the current recognition result (i.e., the target scene category) are switched. Otherwise, the current filter parameters remain unchanged.
[0133] According to the embodiments of this application, by determining the target filtering parameters corresponding to the target scene category when the model evaluation value meets the preset conditions, the problem of inappropriate configuration of filtering parameters caused by model instability can be avoided, thereby improving the reliability and accuracy of adaptive configuration of filtering parameters.
[0134] Step SA3: Filter the second audio signal according to the target filtering parameters to obtain the filtered second audio signal.
[0135] In this embodiment, a filter is configured according to the target filtering parameters, and the second sound signal is filtered by the filter after configuring the target filtering parameters to obtain the filtered second sound signal.
[0136] In some cases, scene recognition models may struggle to accurately determine the current environment or scene, leading to frequent changes in filter parameters. This ambiguous judgment can result in inconsistent and unstable user experience. To avoid unstable scene recognition results due to ambiguous judgments by the scene recognition model, which in turn leads to repeated changes in the configuration of filter parameters and a poor user experience, this application introduces a "cooler" mechanism. That is, after configuring the filter parameters, reconfiguration of the filter parameters is prohibited for a certain period of time (i.e., a preset duration).
[0137] Based on this, in some embodiments, step SA3 involves filtering the second sound signal according to the target filtering parameters to obtain the filtered second sound signal, including steps SA31 and SA32.
[0138] Step SA31: If the target filtering parameters are inconsistent with the current filter configuration parameters, obtain the recording duration.
[0139] In this embodiment, after determining the target filtering parameters corresponding to the target scene category, it is determined whether the currently determined target filtering parameters are consistent with the currently configured filtering parameters of the filter. If they are consistent, the target filtering parameters are not sent to the filter; that is, the configured filtering parameters of the filter remain unchanged. If they are inconsistent, the recording duration is obtained to determine whether to send the currently determined target filtering parameters to the filter. The recording duration is the time interval between the most recent sending of the target filtering parameters and the currently determined target filtering parameters; that is, the recording duration is the time between the most recent switching of the filter's configured filtering parameters based on the target filter parameters and the currently determined target filter parameters. When the recording duration is less than a preset duration, it indicates that the filter is still in the cooling-off period, and the currently determined target filtering parameters are not sent to the filter; that is, the filter does not configure the currently determined target filtering parameters.
[0140] Step SA32: If the recording duration is greater than or equal to the preset duration, the second sound signal is filtered according to the target filtering parameters to obtain the filtered second sound signal.
[0141] In this embodiment, when the recording duration is greater than or equal to the preset duration, the currently determined target filtering parameters are sent to the filter so that the filter can configure its parameters according to the currently determined target filtering parameters and perform filtering processing on the second sound signal to obtain the filtered second sound signal.
[0142] <Device Embodiment>
[0143] Figure 4 This is a schematic diagram of the structure of a noise reduction processing device for headphones according to an embodiment of this application.
[0144] like Figure 4 As shown, the noise reduction processing device 400 for headphones includes an acquisition module 410, a determination module 420, and an output module 430.
[0145] The acquisition module 410 is used to acquire the first sound signal collected by the feedback microphone and the second sound signal collected by the feedforward microphone;
[0146] The determining module 420 is used to determine the actual secondary channel transfer function based on the first audio signal and the second audio signal; and to determine the valid audio signal based on the second audio signal and the reference secondary channel transfer function when the actual secondary channel transfer function is not equal to the reference secondary channel transfer function.
[0147] The output module 430 is used to perform phase cancellation processing on the first sound signal according to the effective sound signal to obtain and output the first target sound signal.
[0148] Optionally, the output module 430 is specifically configured to perform signal separation processing on the first sound signal according to the valid sound signal to obtain a third sound signal; wherein the third sound signal does not include the valid sound signal; perform phase inversion processing on the third sound signal to obtain the first target sound signal; and send the first target sound signal to the speaker so that the speaker can play the first target sound signal.
[0149] Optionally, the determining module 420 is specifically used to filter the second sound signal to obtain a filtered second sound signal; determine the actual secondary channel transfer function based on the filtered second sound signal and the first sound signal; and determine the effective sound signal based on the filtered second sound signal and the reference secondary channel transfer function.
[0150] Optionally, the determining module 420 is specifically used to identify the target scene category corresponding to the second sound signal; determine the corresponding target filtering parameters according to the target scene category; and perform filtering processing on the second sound signal according to the target filtering parameters to obtain the filtered second sound signal.
[0151] Optionally, the determining module 420 is specifically used to extract features from the second sound signal to obtain sound signal features; and input the sound signal features into the scene recognition model to obtain the target scene category corresponding to the sound signal features.
[0152] Optionally, the determining module 420 is specifically used to acquire historical recognition data of the scene recognition model; wherein, the historical recognition data includes the estimated scene category and the actual scene category of the scene recognition model corresponding to each historical scene recognition; based on the estimated scene category and the actual scene category in the historical recognition data, a model evaluation value is obtained; and if the model evaluation value meets preset conditions, the target filtering parameters corresponding to the target scene category are determined.
[0153] Optionally, the determining module 420 is specifically used to obtain the recording duration when the target filtering parameters are inconsistent with the filtering parameters configured in the current filter; and to perform filtering processing on the second sound signal according to the target filtering parameters when the recording duration is greater than or equal to the preset duration, so as to obtain the filtered second sound signal.
[0154] Figure 5 A noise reduction processing apparatus 500 for headphones according to another embodiment of this application includes a memory 510 and a processor 520. The memory 510 is used to store executable instructions; the processor 520 is used to operate under the control of the instructions to perform the method described in any of the above method embodiments.
[0155] How instructions control the processor to perform operations is well known in the art, and therefore will not be described in detail here.
[0156] Figure 6 This is a schematic diagram of the structure of an earphone according to an embodiment of this application.
[0157] like Figure 6 As shown, the headphone 600 includes a feedback microphone 610, a feedforward microphone 620, and a noise reduction processing device 630 for the headphone.
[0158] The feedback microphone 610 collects the first sound signal and sends it to the noise reduction processing device 630 of the earphone;
[0159] The feedforward microphone 620 collects the second sound signal and transmits the second sound signal to the noise reduction processing device 630 of the headphones.
[0160] The noise cancellation processing device 630 of the headphones can be as follows: Figure 5 or Figure 4 The apparatus shown.
[0161] This invention can be a system, method, and / or computer program product. A computer program product may include a computer-readable storage medium having computer-readable program instructions loaded thereon for causing a processor to implement various aspects of the invention.
[0162] Computer-readable storage media can be tangible devices capable of holding and storing instructions for use by an instruction execution device. Computer-readable storage media can be, for example—but not limited to—electrical storage devices, magnetic storage devices, optical storage devices, electromagnetic storage devices, semiconductor storage devices, or any suitable combination thereof. More specific examples (a non-exhaustive list) of computer-readable storage media include: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disc read-only memory (CD-ROM), digital multifunction disc (DVD), memory sticks, floppy disks, mechanical encoding devices, such as punch cards or recessed protrusions storing instructions thereon, and any suitable combination thereof. The computer-readable storage media used herein are not to be construed as transient signals themselves, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through waveguides or other transmission media (e.g., light pulses through fiber optic cables), or electrical signals transmitted through wires.
[0163] The computer-readable program instructions described herein can be downloaded from computer-readable storage media to various computing / processing devices, or downloaded via a network, such as the Internet, local area network, wide area network, and / or wireless network, to an external computer or external storage device. The network may include copper transmission cables, fiber optic transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards them to the computer-readable storage media in the respective computing / processing device.
[0164] The computer program instructions used to perform the operations of this invention may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages such as Smalltalk, C++, etc., and conventional procedural programming languages such as the "C" language or similar programming languages. The computer-readable program instructions may be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or may be connected to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, electronic circuitry, such as programmable logic circuitry, field-programmable gate arrays (FPGAs), or programmable logic arrays (PLAs), is personalized by utilizing state information from the computer-readable program instructions. This electronic circuitry can execute the computer-readable program instructions to implement various aspects of the invention.
[0165] Various aspects of the present invention are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.
[0166] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processor of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner; thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.
[0167] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.
[0168] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction containing one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions marked in the blocks may occur in a different order than those marked in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions. It will be known to those skilled in the art that implementation in hardware, implementation in software, and implementation using a combination of software and hardware are equivalent.
[0169] The various embodiments of the present invention have been described above. These descriptions are exemplary and not exhaustive, and are not limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is chosen to best explain the principles, practical application, or technical improvements to the embodiments in the market, or to enable others skilled in the art to understand the embodiments disclosed herein. The scope of the invention is defined by the appended claims.
Claims
1. A noise reduction processing method for headphones, characterized in that, include: Acquire the first sound signal collected by the feedback microphone and the second sound signal collected by the feedforward microphone; The actual secondary channel transfer function is determined based on the first audio signal and the second audio signal; When the actual secondary channel transfer function is not equal to the reference secondary channel transfer function, a valid sound signal is determined based on the second sound signal and the reference secondary channel transfer function; wherein, when the actual secondary channel transfer function is not equal to the reference secondary channel transfer function, the first sound signal includes noise components not captured by the feedforward microphone and valid sound components, and the valid sound signal is an estimated value of the valid sound components in the first sound signal acquired by the feedback microphone; Based on the effective sound signal, the first sound signal is subjected to phase inversion cancellation processing to obtain and output the first target sound signal; The step of performing phase cancellation processing on the first sound signal based on the effective sound signal to obtain and output the first target sound signal includes: Based on the effective sound signal, the first sound signal is subjected to signal separation processing to obtain a third sound signal; wherein, the third sound signal is the remaining noise component in the first sound signal after removing the effective sound signal; The third sound signal is inverted to obtain the first target sound signal; The first target sound signal is sent to the speaker so that the speaker can play the first target sound signal.
2. The method according to claim 1, characterized in that, The step of determining the actual secondary channel transfer function based on the first audio signal and the second audio signal includes: The second sound signal is filtered to obtain a filtered second sound signal; The actual secondary channel transfer function is determined based on the filtered second audio signal and the first audio signal. The step of determining the valid audio signal based on the second audio signal and the reference secondary channel transfer function includes: The effective audio signal is determined based on the filtered second audio signal and the reference secondary channel transfer function.
3. The method according to claim 2, characterized in that, The step of filtering the second audio signal to obtain the filtered second audio signal includes: Identify the target scene category corresponding to the second sound signal; Based on the target scene category, determine the corresponding target filtering parameters; The second sound signal is filtered according to the target filtering parameters to obtain the filtered second sound signal.
4. The method according to claim 3, characterized in that, The target scene category corresponding to the second sound signal includes: Feature extraction is performed on the second sound signal to obtain sound signal features; The sound signal features are input into the scene recognition model to obtain the target scene category corresponding to the sound signal features.
5. The method according to claim 3, characterized in that, The step of determining the target filtering parameters corresponding to the target scene category based on the target scene category includes: Obtain historical recognition data of the scene recognition model; wherein, the historical recognition data includes the estimated scene category and the actual scene category of the scene recognition model corresponding to each historical scene recognition; The model evaluation value is obtained based on the estimated scene category and the actual scene category in the historical identification data; If the model evaluation value meets the preset conditions, the target filtering parameters corresponding to the target scene category are determined.
6. The method according to claim 3, characterized in that, The step of filtering the second audio signal according to the target filtering parameters to obtain the filtered second audio signal includes: If the target filtering parameters are inconsistent with the current filter configuration parameters, the recording duration is obtained; If the recording duration is greater than or equal to the preset duration, the second sound signal is filtered according to the target filtering parameters to obtain the filtered second sound signal.
7. A noise reduction processing device for headphones, characterized in that, include: The acquisition module is used to acquire the first sound signal collected by the feedback microphone and the second sound signal collected by the feedforward microphone; The determining module is used to determine the actual secondary channel transfer function based on the first audio signal and the second audio signal; When the actual secondary channel transfer function is not equal to the reference secondary channel transfer function, a valid sound signal is determined based on the second sound signal and the reference secondary channel transfer function; wherein, when the actual secondary channel transfer function is not equal to the reference secondary channel transfer function, the first sound signal includes noise components not captured by the feedforward microphone and valid sound components, and the valid sound signal is an estimated value of the valid sound components in the first sound signal acquired by the feedback microphone; The output module is used to perform phase inversion cancellation processing on the first sound signal according to the effective sound signal to obtain and output the first target sound signal; The output module is specifically used to perform signal separation processing on the first sound signal based on the effective sound signal to obtain a third sound signal; wherein, the third sound signal is the remaining noise component in the first sound signal after removing the effective sound signal; The third sound signal is inverted to obtain the first target sound signal; The first target sound signal is sent to the speaker so that the speaker can play the first target sound signal.
8. A noise reduction processing device for headphones, characterized in that, It includes a memory and a processor, the memory being used to store executable instructions; the processor being used to operate under the control of the instructions to perform the method as described in any one of claims 1 to 6.
9. An earphone, characterized in that, Includes a feedback microphone, a feedforward microphone, and a noise reduction processing device for the headphones as described in claim 7 or 8; The feedback microphone collects the first sound signal and sends it to the noise reduction processing device of the earphone; The feedforward microphone collects the second sound signal and transmits the second sound signal to the noise reduction processing device of the earphone.
Citation Information
Patent Citations
Audio signal processing method and device, earphone equipment and storage medium
CN116528099A
Sound signal processing method and earphone equipment
CN116709116A