Noise cancelling partner device for direct communication
The electronic device uses a reference signal from a partner device to selectively exclude specific voices from noise cancellation, addressing the issue of sound isolation in noisy environments by allowing important voices to be heard while attenuating ambient noise.
Patent Information
- Application Number
- PCT/EP2024/087428
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-21
- Filing Date
- 2024-12-19
- Publication Date
- 2025-06-26
AI Technical Summary
Existing noise cancellation technologies often result in sound isolation, which can be undesirable as it filters out important voices, such as service personnel or conversation partners, in environments like airplanes.
An electronic device with circuitry configured to perform noise cancellation on an audio signal based on a reference signal received from a partner device, allowing specific voices to be excluded from noise cancellation while attenuating ambient noise.
Enables users to hear specific voices, such as service personnel or conversation partners, while canceling out ambient noise, thereby improving communication in noisy environments without achieving complete sound isolation.
Smart Images

Figure EP2024087428_26062025_PF_FP_ABST
Abstract
Description
[0001] NOISE CANCELLING PARTNER DEVICE FOR DIRECT COMMUNICATION
[0002] TECHNICAL FIELD
[0003] The present disclosure generally pertains to audio signal processing, in particular, an electronic device and a method for noise cancellation.
[0004] TECHNICAL BACKGROUND
[0005] Noise cancelling devices and techniques, for example, active noise cancellation, are generally known. Headphones, in particular, are well-known to implement active noise cancellation. Therein, a loudspeaker may emit an anti-noise signal to compensate for the surrounding sound, also called ambient noise. Typically, all surrounding sound is filtered-out in this way, which leads to a state of sound isolation for the user.
[0006] Although there exist techniques for noise cancellation, it is generally desirable to improve on existing techniques.
[0007] SUMMARY
[0008] According to a first aspect the disclosure provides an electronic device comprising circuitry configured to perform noise cancellation on an audio signal based on a reference signal received from a partner device.
[0009] According to a second aspect the disclosure provides a method comprising: performing noise cancellation on an audio signal based on a reference signal received from a partner device.
[0010] Further aspects are set forth in the dependent claims, the drawings and the following description.
[0011] BRIEF DESCRIPTION OF THE DRAWINGS
[0012] Embodiments are explained by way of example with respect to the accompanying drawings, in which:
[0013] Fig. 1 schematically shows an example of an active noise cancellation device used for noise cancellation;
[0014] Fig. 2 illustrates a general approach of an active noise cancellation process or system;
[0015] Fig. 3 illustrates an electronic device for noise cancellation partnered with a partner device supplying a partner ID signal for steering the noise cancellation; Fig. 4 illustrates a noise cancellation process and system that rely on a difference signal obtained by comparing audio obtained from two partnered devices;
[0016] Fig. 5 illustrates in more detail the source-specific active noise cancellation based on two partnered devices while preserving the speech signal from the remote device as shown in Fig. 4;
[0017] Fig. 6 illustrates a process of determination of the adaptive filter E(z) which is configured to approximate the transfer function / ?(z) of the transmission path from the partner device to the electronic device;
[0018] Fig. 7 schematically shows another example of a noise cancellation process and system, here based on source separation;
[0019] Fig. 8 schematically shows a general approach of source separation;
[0020] Fig. 9 illustrates a network of electronic devices and partner devices connected to a server;
[0021] Fig. 10 illustrates a user creating a user account for accessing a database of speech models for purposes of noise cancellation as illustrated in Figs. 7 and 8;
[0022] Fig. 11 schematically illustrates an embodiment of an electronic device comprising circuitry for noise cancellation; and
[0023] Fig. 12 schematically illustrates an embodiment of a method for noise cancellation.
[0024] DETAILED DESCRIPTION OF EMBODIMENTS
[0025] Before a detailed description of the embodiments under reference of Fig. 1 is given, general explanations are made.
[0026] The embodiments disclose an electronic device comprising circuitry configured to perform noise cancellation on an audio signal based on a reference signal received from a partner device.
[0027] Thus, a device is proposed that enables to hear voices from specific people and cancels the rest of the noise. In this way, the use of noise cancellation does not result in filtering out these peoples’ voices. The noise cancellation does not lead to a state of sound isolation for the user, as a total sound isolation may not be desired by the user. For example, during airline travels, the device may enable the user to hear service personnel in an airplane, without disabling noise cancellation. Also, people traveling together may be able to have a conversation, which is not possible with total sound isolation. The electronic device may be a wearable electronic device, such as headphones, earphones (e.g., in-ear headphones), smart glasses etc., a smartphone, a laptop computer, a tablet computer, a personal computer, or the like.
[0028] The electronic device may include one or more microphones, for example, one or more microphones for every ear. The microphone(s) may be used to capture the audio signal.
[0029] The audio signal may for example include speech uttered by a user of the partner device and noise, e.g., environmental noise. Multiple microphones may be used to capture the audio signal to provide more complex audio information about the environment.
[0030] Also, one or more error microphones for measuring the effectiveness of the noise cancellation may be included in the electronic device, such as microphones facing towards the user’s ear for measuring the remaining sound reaching the ear after noise cancellation. One error microphone for each ear may be included in the electronic device.
[0031] Circuitry may include one or more entities capable of processing audio signals, such as a CPU (central processing unit), GPU (graphics processing unit), FPGA (field programmable gate array), an application-specific integrated circuit (ASIC), and / or any suitable kind of programmable microprocessor or integrated circuit and / or any other kind of processor or the like.
[0032] A functionality of the circuitry may be specified at least in part by a hardware configuration of the circuity and / or at least in part by software stored on or provided to the circuitry, wherein the software may include instructions executed by the circuitry.
[0033] Circuitry may also include a storage, a memory (RAM, ROM or the like) that, for example, stores (e.g., temporary) data or signals that are generated for processing audio signals. The memory may include a flipflop, a latch, a static random-access memory (SRAM), an embedded dynamic random-access memory (eDRAM) or the like. The circuitry may include input means (mouse, keyboard, microphone, etc.), output means (display (e.g., liquid crystal, (organic) light emitting diode, etc.), loudspeakers, etc., a (wireless) interface, etc., as it is generally known for electronic devices (computers, smartphones, etc.). The circuitry may include a communication unit for receiving signals (e.g., the reference signal), and / or for outputting signals. The communication unit may include a processor pin, a peripheral component interconnect (PCI) interface, a universal serial bus (USB) interface, or the like.
[0034] The partner device may include any one of the features of the electronic device and may also include a circuitry as included in the electronic device. In this way the partner device may function as an electronic device as described herein and vice versa. The partner device supplies the reference signal to the electronic device. The partner device may thus be used in the identification of the source that is to be excepted from the noise cancelling.
[0035] The partner device may be partnered to the electronic device. Partnering may refer to any type of communication exchange or coupling of the two devices. Partnering may also occur within a group of more than two devices. Wireless communication technologies, for example Bluetooth, near-field communication (NFC), pinging or the like may be used for partnering. Partnering may be conducted by activating a key, e.g., by pressing or touching the key or by voice activating the key etc., on one device or on both devices. Alternatively, for example, for near field communication (NFC), touching the devices, for example combined with activating a key may be used.
[0036] For partnering, a transmitting device, for example, the partner device, may first transmit a request for connection to the other device or devices with which to partner, for example, to the electronic device. The transmission of the request may be based on a user input. For that purpose, the transmitting device, e.g., the partner device, may be oriented by the user of the device towards the receiving device, e.g., electronic device, or towards the user of the receiving device. The selection of the receiving device by the user may be based on a list which identifies all devices in a predetermined vicinity. The user can then consult the list of devices that are nearby and select one or more of them. Alternatively, the transmitting device may send a general purpose request which can be picked up by all devices, for example, by all devices that are within a pre-determined distance to the transmitting device or by the device of the closest distance. By orienting the transmission device towards the intended receiving device, the selection of the device may occur.
[0037] To indicate to the respective user of the intended receiving device that a request was sent and / or from which device the request was sent, an optical signal may be used, for example an LED may light up. The user of the receiving device may then visually confirm that a request was sent and / or from which the device, that is, from which user the request was sent.
[0038] The request for connection may be transmitted together with or as an acoustic notice. Thereby the user of the receiving device may be notified that a request was sent. Also, an acoustic signal may be generated upon receipt of the request for connection, e.g., to indicate that a partnering request was received. An optical signal may be used by the receiving device, e.g., electronic device, to signal to the user of the transmitting device, e.g., partner device, that the request for connection was received. The receiving device, e.g., electronic device may transmit an acceptance of the request for connection to the transmitting device, e.g., partner device, for example after the user agreed to the connection by a corresponding input.
[0039] Different optical signals may be used for indicating the request was received and the request was accepted.
[0040] The partner device may be the receiving device and the electronic device may be the transmitting device or vice versa. Therefore, any feature described in this specification regarding the communication between devices may be conducted by any of the devices that are to be partnered, e.g., the partner device or the electronic device.
[0041] The partner device and / or the electronic device may be a special service device for noise cancellation with automatic partnering. For example, service personnel in an airplane may distribute such special service devices to passengers that automatically partner to a corresponding special service device, e.g., partner device, of the service personnel. Using these special service devices, the passengers may profit from the noise cancellation in that noise can be attenuated or cancelled, but that the voice of the service personnel can still be heard.
[0042] The electronic device may include a key or similar for changing the functionality of noise cancellation, such, as switching between different types of noise cancellation and / or switching the noise cancellation on or off. For example, the special service devices offered to the passengers may include a key for switching from the noise cancellation that may let the voice of the service personnel through to, for example, typical noise cancelling, wherein all sound is cancelled.
[0043] Partnering may be based on a predetermined distance threshold, and the switching from noise cancellation to typical noise cancelling of all sound may also be based on the distance threshold. That is, if the distance of the two devices is below a predetermined distance threshold, partnering may either be enabled, for example, in combination with a key activation on one or both (all) devices, or automatically enabled without any other action. Also, a notice may be generated after a request is received and / or after partnering is enabled for any one or all of the partnering devices. For example, an audio notice indicating a partnering option and / or partnering request, possibly in combination with a confirmation of the user, e.g., via a key activation, or the like, may be generated. For example, if the distance of both devices exceeds a predetermined threshold, the noise cancellation according to any of the present embodiments may be disabled and / or automatically switched off, for example to regular noise cancellation of all sound or even to no noise cancellation at all. The reference signal provides information that characterizes or identifies the source that is to be excluded from the noise cancelling. For example, the reference signal may include information related to the partner device. Alternatively, or additionally the reference signal may include information related to a source that is excluded from the noise cancellation. The reference signal may be associated with the source that is excluded from the noise cancellation. In this way, it may be possible to cancel only that part of the audio signal, which is not referenced by the reference signal, i.e., which is not based on the source indicated by the reference signal. That is, that part of the audio signal which originates from the source that is referenced by the reference signal, may be excluded from the noise cancellation.
[0044] The reference signal may indicate a distance and / or a direction of the source (e.g., the partner device and / or the user of the partner device). For example, the reference signal may identify the source (e.g., the partner device and / or the user of the partner device) based on the distance and / or the direction.
[0045] The reference signal may be source specific. Source specific may, for example, refer to the partner device and / or the user of the partner device being the source. For example, the source of the source specific reference signal may refer to the source of that part of the audio signal which is excluded from the noise cancellation. Additionally, or alternatively the source may refer to the source that supplies the information on which basis the exclusion of part of the audios signal from the noise cancellation occurs. For example, in some embodiments the user of the partner device, i.e., the partner device, may supply a reference audio signal, which is used in identifying which part of the audio signal is excluded from the noise cancellation.
[0046] The reference signal may indicate a source generating part of the audio signal, wherein the part of the audio signal generated by the indicated source may not be cancelled by the noise cancellation. In this way the indicated source may be excepted from the noise cancellation.
[0047] The noise cancellation may include a source separation of the audio signal. The source separation may be configured to separate the audio signal into a speech signal and a noise signal.
[0048] The source separation may be implemented based on a blind source separation or on a trained model. The trained model may be a trained speech model, trained to separate speech of a specific person from the audio signal.
[0049] The source separation may include multiple trained models trained to separate speech of one specific person from the rest of the audio signal. In this way the source separation may include multiple trained models tailored to different persons, for example family members or friends of the user. The trained model may be trained based on one or more speech samples of the person whose speech is to be separated from an audio signal via source separation. The trained model may be a machine learning algorithm and may, for example, be based on a support vector machine (SVM), an artificial neural network, e.g., a convolutional neural network (CNN) or multimodal Al model, e.g., transformer model, or the like.
[0050] The source separation of the audio signal may be based on the reference signal received from the partner device. For example, if the source separation includes multiple source separations as subunits as explained above, the reference signal may identify which of the multiple source separations is to be used.
[0051] The noise cancellation may include anti-noise signal generation circuitry configured to generate an anti-noise signal based on the noise signal. The anti-noise signal as obtained from the antinoise signal generation may for example be determined by conventional noise cancellation techniques such a phase-inversion.
[0052] The anti-noise signal generation may be configured to generate the anti-noise signal based on only the noise signal, excluding the speech signal obtained from source separation from the noise cancellation.
[0053] In some embodiments, the reference signal may be a partner ID. Partner ID may be anything that is capable of identifying the partner device and / or the user wearing the partner device. For example, a unique number, a string of characters, or any combination thereof. The partner ID may, therefore, indicate the source of the source-specific reference signal. For example, the source may be a specific person, for example, the user of the partner device or another person, whose voice or speech is then excluded from the noise cancellation.
[0054] The source separation may therefore determine based on the received partner ID whose speech shall be source separated from the audio signal. For example, if the source separation includes multiple source separation models, each model may have an ID which identifies the person whose speech can be separated by the model. By comparing the partner ID to the IDs of the model, a match may be found and the model matching to the partner ID may be selected as the source separation that is implemented. In this way the speech from the user of the partner device can be separated from the audio signal based on source separation.
[0055] Therefore, if the anti-noise signal is based on the noise signal, which may be the remaining signal after source separating the speech of the partner device user from the audio signal, the speech will not be noise cancelled, but only the remaining signal, i.e., the noise signal, will be cancelled. Thus, the user of the electronic device may be able to hear the speech of the partner device user while the rest of the noise is cancelled.
[0056] If the reference signal indicates a source generating an audio signal, the source may refer to the source generating the speech signal of the audio signal which is separated from the noise signal.
[0057] In alternative embodiments, the reference signal may be a reference audio signal received from the partner device. The reference audio signal may for example be captured by a microphone or more than one microphone of the partner device. This reference audio signal may for example include speech uttered by a user, e.g., the user of the partner device, and noise, e.g., environmental noise. The reference signal may alternatively, or additionally, include information that is derived from the reference audio signal captured by the microphone of the partner device. That is, instead of the reference audio signal being the basis for the noise cancellation, derived features or any additional relevant information may be used as a basis for the noise cancellation. In this way, features derived from the reference audio signal may describe the information required for steering the noise cancelling mechanism, e.g., the information required for excluding part of the audio signal from the noise cancellation.
[0058] The audio signal and the reference audio signal may differ in regard to magnitude of the captured speech included in the respective signals due to the physical distance of the electronic device and partner device, i.e., the distance of the respective microphones capturing the audio signal and the reference audio signal, to the speaker uttering the speech.
[0059] The noise cancellation may include a difference calculation configured to determine a difference signal based on the audio signal and the reference audio signal.
[0060] The noise cancellation may include a noise cancellation configured to determine an anti-noise signal based on the difference signal.
[0061] The noise cancellation may be further configured to determine the anti-noise signal based on an error signal. The error signal may be captured by the error microphone of the electronic device as described above.
[0062] The noise cancellation may be based on an adaptive filter which operates on the audio signal and generates the anti-noise signal.
[0063] The noise cancellation may include a second adaptive filter which is configured to determine, based on the difference signal, a filtered difference signal.
[0064] The second adaptive filter may be configured to approximate a transfer function of the transmission path from the partner device to the electronic device. The approximated transfer function of the transmission path from the partner device to the electronic device may include a reconstruction of direction of arrival. Directional of arrival may refer to the direction of arrival of the sound or soundwaves following the transmission path from the partner device to the electronic device.
[0065] The reconstruction of direction of arrival may be based on the geometric configuration of the electronic device and the partner device. Alternatively, the reconstruction of direction of arrival may be based on the geometric configuration of the respective microphones of the partner device and the electronic device.
[0066] The reconstruction of direction of arrival may be based on the geometric configuration of the electronic device and the partner device relative to each other and / or based on the geometric configuration of the respective microphone(s) of the electronic device and the microphone(s) of the partner device relative to each other.
[0067] The geometric configuration may be determined based on localization, such as GPS, triangulation of electromagnetic waves, etc.
[0068] The second adaptive filter may be obtained by recursively subjecting the refence audio signal to the second adaptive filter to obtain a filtered refence audio signal and determining a second difference signal between the audio signal and the filtered reference audio signal.
[0069] The adaptive filter may be optimized based on the filtered difference signal and the error signal in a way that the error signal is minimized.
[0070] The noise cancellation may be further configured to determine an added signal based on the filtered difference signal and the error signal.
[0071] Optimizing the adaptive filter may be based on the added signal.
[0072] The noise cancellation and any feature of the noise cancellation may be based on a learning model. The learning model may be a machine learning algorithm and may, for example, be based on a support vector machine (SVM), an artificial neural network, e.g., a convolutional neural network (CNN) or multimodal Al model, e.g., transformer model, or the like.
[0073] Any feature described above in regard to the electronic device may be extended to the partner device and vice versa. Thus, the users of both devices may speak to each other while cancelling the remaining noise of the environment. This may include that the respective own voice of the user is not cancelled. In this way oneself and the conversation partner can be heard as is the case in a natural conversation. Any feature described above may also extend to a multi-device approach of a group of users, each user having their own electronic device, which is also a partner device, wherein the whole group can speak to each other while cancelling the rest of the noise in the environment, i.e., the voice of all partnered users, and possibly their own voice as well, may be heard by all other users.
[0074] Some embodiments pertain to a method comprising: performing noise cancellation on an audio signal based on a reference signal received from a partner device.
[0075] The method may correspond to the electronic device and circuitry described above, and the method may accordingly exhibit any feature described above with respect to the electronic device and / or circuitry and / or any suitable feature described below with respect to any one of the figures. The method may be performed by the circuitry and / or by the electronic device described above.
[0076] The methods as described herein are also implemented in some embodiments as a computer program causing a computer and / or a processor to perform the method, when being carried out on the computer and / or processor. In some embodiments, also a non-transitory computer- readable recording medium is provided that stores therein a computer program product, which, when executed by a processor, such as the processor described above, causes the methods described herein to be performed.
[0077] Active noise cancellation
[0078] Active noise cancellation (ANC), also known as noise cancellation (NC), or active noise reduction (ANR), is a known method for reducing unwanted noise by the addition of an antinoise signal specifically designed to cancel the noise.
[0079] Human speech typically consists of frequencies in the range of 250 Hz up to 2 kHz. Noise canceling devices may therefore use two mechanisms. For low frequencies the active noise control procedures with adaptive filters as described in the above figure can be used. For higher frequencies above 1 kHz active noise cancellation may not work as well or may be difficult to implement. For higher frequencies acoustic isolation with ear muffs may therefore be used. The ear muffs may act as a low pass filter and suppress high frequency noise.
[0080] Fig. 1 schematically shows an example of an active noise cancellation device used for noise cancellation. An active noise cancellation device is used to cancel noise 51 (e.g., ambient noise of the environment), that arrives at the ears of user 5. Noise 51 is captured by microphone 28a of electronic device 44a. Microphone 28a may be any type of microphone (e.g., 107, Fig. 11), e.g., a single microphone or a microphone included in a microphone array. Electronic device 44a comprises active noise cancellation 25. The captured noise 51 is processed in active noise cancellation 25 to generate an anti -noise signal 48, e.g., a 180°-phase inverted audio signal to noise 1 arriving at the ear of the user 5. A loudspeaker 27 of electronic device 44a emits the antinoise signal 48 and thereby attenuates or cancels the noise 51 arriving at the ear of user 5. Electronic device 44a also includes an error microphone (see 57 in Fig. 2) located between the loudspeaker 27 and the ear of user 5. The error microphone is directed towards the ear of user 5 and it captures any remaining signal after noise cancellation. The captured signal, as error signal 29, is fed back to active noise cancellation 25. Active noise cancellation 25 optimizes the generation of anti-noise signal 26 in a way that the error signal 29 is minimized. The process of active noise cancellation 22 is explained in more detail in Fig. 2.
[0081] The setup of Fig. 1 is an example noise cancellation device 44a, including an in-ear headset and a necklace, wherein microphone 28a is located at the necklace. A noise canceler which may include the active noise cancellation 25 may be located within the necklace and / or within the in- ear headset. It is noted that the active noise cancellation 25 is not limited to any specific division of function into a specific unit. It is noted, that, alternatively, the electronic device 44a may also be an over-ear headphone which may include all microphones and a noise canceler which may include the active noise cancellation 25. Also in the case of the over-ear headphone, it is noted that the active noise cancellation 25 is not limited to any specific division of function into a specific unit.
[0082] Fig. 2 illustrates a general approach of an active noise cancellation process or system. The active noise cancellation system includes a microphone 28a that is configured to capture noise 51 (any audio signals, such as ambient noise of the environment). The active noise cancellation system further includes a speaker / amplifier 27 for emitting anti-noise signals 48. The active noise cancellation system further includes an error microphone 57 for capturing the remaining signal after noise cancellation. Microphone 28a, speaker / amplifier 27 and error microphone 57 are all included in an active noise cancellation device (e.g., electronic device 44a, Fig. 1) worn by a user (e.g., user 5, Fig. 1), which attenuates unwanted noise at the user’s ear.
[0083] As illustrated in Fig. 2, noise 51 (e.g., audio signal 1, Figs. 1 to 4) follows two paths, the audio signal path and the noise cancelling path. Both paths converge at 56 to cancel each other out.
[0084] Along the audio signal path, noise 51 in the form of an audio signal is influenced by any physical structures (room, etc.) that are within the audio signal path. The effect of these structures is described by a transfer function P(z) which is typically called the “plant transfer function”. This plant transfer function P(z) describes the acoustic transfer function (all physical structures) between the microphone 28a and the speaker 27. In other words, the plant transfer function P(z) represents the path the noise 51 takes through the physical transmission system starting from the microphone 28a to the point where noise is cancelled out, which occurs just behind the speaker 27 at the user’ s ear.
[0085] Noise 51 also travels along the noise cancelling path. Noise 51 is captured by microphone 28a and is thereby transferred from the acoustic domain to the electronic domain. This transfer is described by a transfer function M z which describes the effect that microphone 28a (and any amplifier and any A / D converter related to microphone 28a) has on the noise 51. The noise as captured by microphone 28a is then subjected to active noise cancellation (25 in Fig. 1). The effect of noise cancellation is represented by adaptive filter H(z)' . After noise cancellation as effected by adaptive filter H z), the resulting signal is played back through speaker 27. The effect of speaker 27 (and any amplifier and / or D / A converter involved) is characterized by the transfer function S(z). Both acoustic signals, i.e., the noise as transferred by the audio signal path and the noise as transferred by the noise cancelling path, interfere at 56. The interference should ideally result in that the noise is canceled out.
[0086] The error microphone 57 picks up the audio signal after interference at 56. Similar to the transfer function M z which describes the effect that microphone 28a has on the audio signal, the transfer function O(z) describes the effect that the capture process at error microphone 57 has on the audio signal. The signal captured by error microphone 57 (output of transfer function O(z)) is fed back in a feedback loop as input into the adaptive filter H z . The adaptive filter H z') (which represents the effect of the noise cancellation algorithms) is optimized via the error signal in a way that the error signal is minimized, i.e., the noise cancels out at 56 (the position of the user’s ear). That is, by the feedback loop, H (z)is optimized in a way that the combination of M z), H z) and S(z) approximates P(z).
[0087] The adaptive filter H z') itself may for example be implemented according to the skilled person’s general knowledge, e.g., by a Least Mean Square (LMS) algorithm or some variants of it such as normalized LMS or Filtered-X LMS or the like.
[0088] Partnering of noise cancellation devices
[0089] Fig. 3 illustrates an electronic device for noise cancellation partnered with a partner device supplying a partner ID signal for steering the noise cancellation.
[0090] A user 5 wearing an electronic device 44a that is used for noise cancellation is sitting in an airplane. User 5 uses electronic device 44a for active noise cancellation, in particular for cancelling the airplane noise. User 6, travelling together with user 5 and sitting on the chair next to user 5 in the airplane, talks to user 5.
[0091] Electronic device 44a for noise cancellation at user 5, which includes earphones, loudspeaker 27 and microphone 28a, is partnered with partner device 44b of user 6, which includes earphones and a microphone 28b as well as a loudspeaker 27. Thus, the two devices 44a, 44b are coupled to each other. User 6 is talking to user 5, therefore, speech 24 of user 6 reaches user 5, i.e., the ears of the listening user 6, via soundwaves. The soundwaves of speech 24 of user 6 are also captured by microphone 28a of electronic device 44a as part of an audio signal (e.g., audio signal 1 of Fig. 2), also including ambient noise, on which basis noise cancellation is performed. That is, the audio signal captured by microphone 28a includes speech 24 (e.g., speech 24 of Figs. 5-7) of user 6 and noise (e.g., noise 51, Figs. 5-7).
[0092] In the embodiments described below in more detail, noise cancellation is performed only on the noise part of the captured audio signal, but not on the speech of the user 6 of the partner device 44b. Thus, the user 5 can hear the speech 24 of user 6, but the noise (e.g., 51, Figs. 5-7), e.g., ambient noise, is cancelled out. The same may also apply in the other direction. That is, both devices 44a, 44b may conduct noise cancellation of the noise (e.g., 51 of Figs. 5-7) without cancelling the speech (e.g., 51, Figs. 5-7) of each user 5, 6. Also, not only the voice of the respective partner, but also the own voice of each user may be spared from being cancelled.
[0093] For partnering the two devices 44a and 44b, wireless communication technologies, for example Bluetooth, near-field communication (NFC), pinging or the like may be used. For that purpose, pairing mode may be enabled in the two devices and then the partnering is initiated, for example via Bluetooth. Partnering may be enabled based on activating a key for both devices, for example a key included in the electronic device, such as pressing a button, or the like. Alternatively, partnering may be initiated by touching the two devices, for example based on NFC.
[0094] Another mode of getting in contact may be to ping a device and ask for connection. The device that sends the ping may indicate this by an optical signal, so that the receiver of the ping can see who is requesting the connection. The receiver user (e.g., user 5) can then decide whether to connect with the other user (e.g., user 5). The ping may be an acoustic signal.
[0095] Optical signaling, but also any other communication technology, e.g., wireless communication technology, may be used by user 6 to indicate to user 5 that they want to speak to user 5 and that source-specific noise cancellation for their person should be initiated. In this way user 5 knows that they should, for example, switch from general noise cancellation of all audio signals to a noise cancellation that will let the speech of user 6 through. Also, optical signaling may be used to indicate the enabled pairing mode and / or the acceptance of the request for partnering.
[0096] For example, the receiver, i.e., user 5, of the request for connection / partnering signal from the speaker, e.g., user 6, may be selected by user 6 pointing with their partner device 44b in the direction of the electronic device 44a of user 5, which in case of headphones may include a head movement, optionally also including a simultaneous key activation. In addition, the currently selected electronic device 44a may indicate the selection via optical signaling, e.g., to avoid confusion in selecting a partner device.
[0097] For example, partner device 44b sends a signal via a signaling unit 14b to electronic device 44a to ask for connection. Electronic device 44a receives and accepts the request for connection and indicates this with one or more optical signals via their own signaling unit 14a. The optical signal may differ based on whether the signal was merely received or also accepted. Any one of the optical signals may include an acoustic signal. To confirm the partnering any one or both devices may display an optical signal via signaling unit 14a, 14b, which may differ from the optical signals 14a, 14b for requesting a connection, receiving a connection and / or accepting a connection. The signal sent as a partnering request may also be a partner ID signal 21a of Fig. 7.
[0098] Noise cancellation based on difference signal obtained from partnered devices
[0099] In Fig. 4, a noise cancellation process and system are described that rely on a difference signal obtained by comparing audio obtained from two partnered devices.
[0100] The noise cancellation process and system described here is based on the observation that microphone 28b of partner device 44b (see also Fig. 3), that is worn by user 6 also captures the speech (24, Fig. 3) uttered by user 6 and the noise of the environment. In particular, the partnered devices 44a and 44b (see also 44a and 44b of Fig. 3) receive essentially the same noise (e.g., any airplane noise) with their respective microphones 28a and 28b. That is, when determining a difference signal (via the difference calculation 31 of source specific active noise cancellation 61a), for example by subtracting the audio signal 1 captured by microphone 28a from the audio signal 20 captured by microphone 28b, the noise essentially cancels out.
[0101] As microphone 28b is located very close to the speaking user 6, the speech (e.g., 24, Fig. 3) of user 6, as captured by microphone 28b, is, typically, louder than as it is captured at microphone 28a of user 5 which is located farther away from the speaking user 6. There is, accordingly, a difference in the respective speech signals. When determining a difference signal by subtracting the audio signal 1 captured by microphone 28a from the audio signal 20 captured by microphone 28b, the speech does not cancel out. Thus, by determining the difference based on difference calculation 31 based on the signals at microphone 28a and microphone 28b the speech of user 6 may be separated from the noise. The resulting difference signal 21b of the difference calculation 31 can then be input, to the active noise cancellation 30 of source specific active noise cancellation 61a, which is explained in more detail in Fig. 6 below. As the audio signal 20 used as reference signal for noise cancellation is obtained from the partnered device of user 6 who utters the speech and thus is the source of the signal, it is a source specific signal.
[0102] The active noise cancellation 30 generates an anti-noise signal 26 for cancelling only the noise arriving via soundwaves at the ear of user 5. Anti-noise signal 26, e.g., a 180°-phase inverted signal to the noise arriving at the ear of user 5, is emitted by loudspeaker 27 of electronic device 44a and thereby attenuates or cancels the noise arriving at the ear of user 5. The speech is of user 6 is not cancelled out and thus heard by user 5 unimpeded.
[0103] Similar to the active noise cancellation described in Figs. 1 and 2 above, the source specific active noise cancellation 61a of electronic device 44a also receives audio from an error microphone (57 in Fig. 5) located between the loudspeaker 27 and the ear of user 5, which is directed towards the ear of user 5 and which captures any remaining signal after noise cancellation, i.e., error signal 29. Error signal 29 is fed back to active noise cancellation 30 of source specific noise cancellation 61a to optimize the generation of anti -noise signal 26 in a way that noise is cancelled, and speech is heard unimpeded.
[0104] The difference calculation 31 may be conducted in the partner device 44b. In that case, the reference signal (e.g., audio reference signal) may refer to the output of the difference calculation 31, which is the difference signal 21b. Alternatively, difference calculation 31 may be conducted in the electronic device 44a. In that case, the reference signal (e.g., audio reference signal) may refer to the audio signal 20.
[0105] Fig. 5 illustrates in more detail the source specific active noise cancellation 61a based on two partnered devices while preserving the speech signal from the remote device as shown in Fig. 4.
[0106] Fig. 5 shows two microphones, the local microphone 28a and the remote microphone 28b as well as speaker / amplifier 27 and error microphone 57. The remote microphone 28b is the microphone close to the speaker (user 6 of Figs. 3 and 4) and the local microphone 28a is the microphone far from the speaker (microphone of user 5 of Figs. 3 and 4). Local microphone 63, speaker / amplifier 27 and error microphone 57 are all included in the electronic device 44a worn by user 5 (see Figs. 3, 4). As described with regard to Fig. 4 above, noise 51 is assumed to be captured equally by both microphones 28a and 28b, respectively. The same noise 51 is thus captured by microphones 28a and 28b and is thereby transferred from the acoustic domain to the electronic domain. This transfer is described by the transfer functions M z which describe the effect that microphones 28a and 28b (and any related amplifier and / or any A / D converter) have on the noise 51.
[0107] At 71 a difference signal 21b is obtained based on the signal captured by local microphone 28a and based on the signal captured by remote microphone 28b. For example, at 71, a phase- inverted signal of the signal captured by local microphone 28a is added to the signal captured by remote microphone 28b in order to cancel the noise received at both microphones 28a, 28b.
[0108] However, as local microphone 28a is located farther away from the speaker (user 6 in Fig. 4) than remote microphone 28b, the speech 24 of the speaker arriving at local microphone 28b is influenced by the transmission path the speech 24 takes before it arrives at local microphone 28a. The effect of this transmission path is represented by transfer function / ?(z). / ?(z) is the acoustic transfer function between the remote microphone 28b and the local microphone 28a and includes aspects such as reproduction of reverberation and direction of arrival as well as the difference in loudness. Reverberation refers to the fact that more sound reflections occur if speech originates from a bigger distance. In other words, compared to speech 24 captured by remote microphone 28b, speech 24 captured by local microphone 28a includes more reverberation due to more reflections occurring on the way to local microphone 28a located farther away. Similarly, loudness difference refers to the fact that sound magnitude or loudness is lesser if speech originates from a greater distance. In other words, compared to speech 24 captured by remote microphone 28b, speech 24 captured by local microphone 28a is less loud due to intensity of the sound being dispersed on the way to local microphone 28a located farther away from the speaker. Direction of arrival refers to the natural phase delay of sound between the left and right ear of humans, often represented in headphones via a phase delay between left and right channel, and which allows humans to precisely locate the direction a sound, e.g., speech comes from. For example, speech 24, which from the perspective of electronic device 44a and microphone 28a is based on the voice of remote user 6 located far away, should appear to arrive from the correct direction of the user 6 located far away.
[0109] Additionally, the reconstruction of direction of arrival may be based on known geometric configuration of the devices (here, electronic device 44a and partner device 44b, Figs. 3, 4) of the respective microphone 28a and microphone 28b as well as the geometric configuration of the devices relative to each other. The geometric configuration of the two devices relative to each other may be determined based on any localization method known to the skilled person, such as GPS, triangulation of electromagnetic waves and the like. As a result, the direction of arrival of speech 24 may be reconstructed more precisely. That is, although, direction of arrival is already included in R(z), the orientation of both devices may also be used. In this way, the direction of arrival of speech 24 may be manipulated in a way that it sounds even more natural, e.g., appears more precisely from the direction of the speaker / source of speech 24, or an erroneous direction of arrival may be corrected.
[0110] As the noise 51 of both microphones 28a and 28b is the same and speech 24 is the only difference, the difference signal 21b obtained at 71 comprises essentially the speech 24 captured by the remote microphone 28b. This difference signal 21b obtained at 71 is subjected to adaptive filter E(z). As described in more detail below with regard to Fig. 6, adaptive filter E(z) is configured to approximate / ?(z) which represent the path that sound takes when travelling from user 6 to user 5. As E(z) approximates / ?(z), the output of adaptive filter E(z), i.e., the filtered difference signal 59 substantially corresponds to the speech signal 24 as it is perceived by local microphone 28a.
[0111] Additionally, the noise 51 and the speech 24 travel along the acoustic path through the transmission system to the ear of user 5. This transmission is represented by plant transfer function P(z). That is, P(z) represents the path the noise 51 and the speech 24 take through the transmission system on their way to the point where noise is cancelled, which occurs just behind speaker 27.
[0112] The effect of noise cancellation (30 in Fig. 4) is represented by adaptive filter H(z)' . Adaptive filter H z') operates on the signal captured by local microphone 28a of user 5. After noise cancellation as effected by adaptive filter H z), the resulting signal is played back through speaker 27. The effect of speaker 27 (and any amplifier and / or D / A converter involved) is characterized by the transfer function S(z). Both acoustic signals, i.e., the noise and speech as transferred by the audio signal path and the signal as transferred by the noise cancelling path, interfere at the location 67 of the ear of local user 5.
[0113] At 67 the audio signal including noise 51 and speech 24 via the transmission system represented by P(z) is partly cancelled by the emitted anti-noise signal as represented by S(z) in a way that only the noise 51 is cancelled and speech 24 remains. That is, the interference ideally results in that the noise is canceled out, but speech of remote user 6 remains.
[0114] The error microphone 57 picks up the audio signal after interference at 67. Similar to the transfer function M z which describes the effect that microphones 28a has on the audio signal, the transfer function O(z) describes the effect that the capture process at error microphone 57 has on the audio signal. The signal captured by error microphone 57 (output of transfer function O(z)) is fed back in a feedback loop as input into the adaptive filter H z') as error signal 29.
[0115] The adaptive filter H z') (which represents the effect of the noise cancellation algorithm) is optimized via the error signal 29 and the filtered difference signal 59 (the output of adaptive filter E(z)) in a way that the error signal is minimized, i.e. the noise cancels out at 67 (the position of the user’s ear) but the speech 24 does not cancel out at 67. To achieve this, in the example of Fig. 5, the filtered difference signal 59 is subtracted from the error signal 29 at 68 and the subtracted signal is provided to adaptive filter H z') as a feedback signal. In this way, the adaptive filter H (z)is optimized in a way that the error signal 29 and the output of E(z) are equal, and therefore, to output an anti-noise signal only for cancelling the noise 5, but not speech 24.
[0116] The optimization process described above may be implemented by a Least Mean Square (LMS) algorithm or variants of it (see e.g., https: / / en.wikipedia.org / wiki / Least_mean_squares_filter; Ardekani, Iman Tabatabaei, and Waleed H. Abdulla. "FxLMS-based Active Noise Control: A Quick Review.” APSIPA ASC (2011); Widrow, Bernard, et al. "Adaptive noise cancelling: Principles and applications." Proceedings of the IEEE 63.12 (1975): 1692-1716; Milosevic, Aleksandar, and Urs Schaufelberger. "Active noise control." University of Applied Sciences Rapperswil HSR (2005))
[0117] The mechanism described above may be extended to a multi -mi crophone approach. For example the multi -mi crophone approach may be implemented in a headset with a left side and a right side microphone. That is, in case of the electronic device being headphones, earphones or the like, both the left and right side microphone of the electronic device may be utilized for capturing the audio signal. Likewise, in case of the partner device being headphones, earphones or the like, both the left and right side microphone of the partner device may be utilized for capturing the reference audio signal.
[0118] The multi -mi crophone approach may result in improved estimation of accuracy of the transfer function.
[0119] Also, in case the electronic and partner devices include ear muffs, for example in case of headphones, the ear muffs may act as a low pass filter and suppress high frequencies. For recovery, of the higher frequencies in the system as explained above, the high frequencies in speech 24 may be added to the speaker / amplifier 66 output in order to preserve the natural speech frequency spectrum. Thus, to preserve the higher frequencies of the natural speech spectrum heard by user 5, which for human speech typically lies in the range of 250 Hz up to 2 kHz, high frequencies may be added at S(z).
[0120] Fig. 6 illustrates a process of determination of the adaptive filter E(z) which is configured to approximate the transfer function / ?(z) of the transmission path from the partner device to the electronic device. As Fig. 5 described above, Fig. 6 shows local microphone 28a (of user 5) and remote microphone 28b (of user 6 who utters speech). As described with regard to Figs. 4 and 5 above, noise 51 is assumed to be captured equally by both microphones 28a and 28b, respectively. However, as local microphone 28a is located farther away from the speaker (user 6 in Fig. 4) than remote microphone 28b, the speech 24 of the speaker arriving at local microphone 28a is affected by the transmission path from the partner device to the local device. The effect of this transmission path is represented by transfer function / ?(z) as explained in Fig. 5 above.
[0121] The respective transfer functions M z represent the effect of microphones 28a and 28b when transferring the signal from the acoustic domain to the digital signal. Thus M(z) of the remote microphone 28b effects noise 51 and the unchanged speech 24, whereas M z of the local microphone 28a effects noise 51 and a changed speech signal as represented by / ?(z).
[0122] The signal captured by remote microphone 28b is subjected to adaptive filter E(z), which approximates the acoustic transfer function / ?(z). At 64, the output 59b (e.g., filtered reference audio signal) of the adaptive filter E(z) is used for cancelling out the speech component within the signal obtained at the local microphone 28a. If E (z) well approximates the transfer function / ?(z) of the transmission path from the partner device to the electronic device, then, ideally, the speech signals from the two paths, i.e., from the acoustic path of speech 24 through the transmission system to the local microphone 28a represented by / ?(z) and M(z , and the electric path speech 24 takes via the remote microphone represented by M z and adapted by adaptive filter E(z) cancel out at 64.
[0123] At 64, a feedback error signal 65 is determined based on the signal as captured by remote microphone 28b and based on the signal as captured by local microphone 28a as obtained after filtering with E z . The error signal 65 obtained at 64 is fed back to the adaptive filter E(z) for optimization in an optimization loop. The Least Mean Square (LMS) algorithm or variants of it may be used to implement the optimization of adaptive filter F^z). After having determined adaptive filter E(z) in his optimization process, a copy of (z) is used in the system described in Figs. 4 and 5 above to approximate the transfer function / ?(z) of the transmission path from the partner device to the electronic device.
[0124] Noise cancellation based on source separation Fig. 7 schematically shows another example of a noise cancellation process and system, here based on source separation.
[0125] As in Fig. 3 and 4 described above in more detail, a user 5 wearing an electronic device 44a that is used for noise cancellation is located in a noisy environment, e.g., in an airplane. User 5 uses electronic device 44a for active noise cancellation, in particular for cancelling the noise. User 6, a fellow travelling together with user 5 and sitting on the chair next to user 5 in the airplane, talks to user 5.
[0126] User 6 is wearing an electronic device, called partner device 44b, that is communicatively coupled with the electronic device 44a of user 5. The partner device 44b of user 6 communicates a partner ID 21a as a reference signal to electronic device 44a of user 5. This partner ID 21a is indicative of user 6. It indicates to electronic device 44a of user 5 that user 6 is present in the vicinity of user 5. In particular, partner ID 21a of user 6 indicates to electronic device 44a of user 5 that speech of user 6 should not be cancelled out the via noise cancellation performed by electronic device 44a of user 5. As it carries information about the source of the speech uttered by user 6, partner ID 21a is a source specific signal.
[0127] Electronic device 44a includes earphones including a microphone 28a and a loudspeaker 27. Microphone 28a captures audio signal 1 which includes speech uttered by user 6 and noise. The partner ID signal 21a indicates to electronic device 44a that user 6 is the person that is the speaker of the speech in audio signal 1. This audio signal 1 captured by microphone 28a is then input to source separation 22.
[0128] Source separation 22 includes three source separation models, namely source separation A, source separation B and source separation C, each configured to separate the speech 24 of one specific person from the rest of the signal, i.e., from the noise 51 (e.g., including speech of other persons). Details of source separation technologies that can be applied by source separation A, source separation B and source separation C are described in more detail below, with regard to Fig. 8.
[0129] Source separation A, B and C may for example operate according to the principle described with regard to Fig. 1 above. Here, source separation A is configured to separate the speech of user 6 (source A) from noise. Source separation A may for example maintain a pretrained speech model that allows for separation of speech uttered by user 6. This speech model may have been exchanged by user 5 and 6 in advance. Likewise, source separation B is configured to separate speech of another person (source B, not shown in Fig. 2) from noise and source separation C is configured to separate speech of yet another person (source C, also not shown in Fig. 2) from noise.
[0130] Each of the source separation models is associated with a respective identifier identifying the person that the model is capable of separating. Here, source separation A that is configured to separate the speech of user 6 (source A) from noise is associated with partner ID 21a. Source separation models B and C are associated with other IDs (not shown in Fig. 2).
[0131] Based on the source specific partner ID 21a received from partner device 44b, source separation 22 is configured to select which of source separation A, source separation B and source separation C is applied to the audio signal. As partner ID 21a is associated with source separation A, source separation 22 selects source separation A for application on audio signal 1. Source separation 22 applies source separation A to the audio signal 1 to obtain speech 24 of user 6 separated (as a separated source, e.g., 2a in Fig. 1) from noise 51 (see the residual signal 3 in Fig. 1). The noise 51 includes the remaining signal of the audio signal, apart from the speech 24, in particular the noise (airplane engines and speech of other persons in the cabin) captured by microphone 28a. After source separation 22, the noise 51 is separated from the speech 24.
[0132] Based on the noise 51 an anti-noise signal 26 is generated via anti-noise signal generation 25. Anti-noise signal generation 25 may for example operate on any conventional means known to the skilled person, e.g., algorithms that generate a signal that will either phase shift or invert the polarity of the original signal. The anti-noise signal 26 generated by anti-noise signal generation 25 is emitted by the loudspeaker 27 of the earphones of electronic device 44a of user 5. Accordingly, the noise 51 is cancelled out by the anti -noise signal 26 at the ear of user 5. As the speech 24 is not cancelled by the anti-noise signal 26, the speech 24 remains in audio signal 1 that user 5 hears. In other words, only noise 51 of audio signal 1 is cancelled, but not the speech 24 of user 6, who is a fellow of user 5.
[0133] In the embodiment described above, source separation A maintains a pre-trained speech model that is capable of separating the speech of user 6 -a fellow of user 5- from noise. As a modification of this example, a stewardess may supply electronic device 44a to user 5 that includes in source separation 22 as source separation subunit a source separation C that comprises a model for separating the stewardess’ voice, so that when the stewardess transmits her partner ID signal to electronic device 44a speech from the stewardess is separated from noise and the speech of the stewardess is not cancelled out. In the embodiment described above, a source separation 22 using pre-training is implemented to separate audio signal 1 into sources A, B or C. However, in alternative embodiments, instead of a pretrained source separation 22, blind source separation may be implemented.
[0134] In the embodiment described above, source separation 22 is described as maintaining pretrained speech models that allow for separation of speech uttered by user 6. In alternative embodiments, blind source separation techniques might be alternatively applied which do not need a pretrained speech model.
[0135] Source separation may alternatively be based on two steps. First by separating human voice from other noise of the environment, for example, by known source separation techniques. Second, by separating the speaker of a specific person based on volume and different voice characteristics, for example due to the mechanical connection of the device with the speakers. Both steps may be based on their own learned models, for example, trained in advance and / or adjusted during runtime. The source separation 22 and any source separation described in this specification may be conducted in the remote device, e.g., the partner device 44b. In that case, the reference signal may refer to the output of the source separation 22, which is the noise 51. Alternatively, the source separation 22 may be conducted in the local device, e.g., electronic device 44a. Thus, the reference signal may refer to the partner ID 21a.
[0136] Fig. 8 schematically shows a general approach of source separation. Source separation (also called “demixing”) is performed which decomposes a source audio signal 1 including audio from multiple audio sources Source 1, Source 2, . . ., Source K (e.g., voice of person A, voice of person B, noise of airplane etc.) into “separations”, here into source estimates 2a-2d, wherein K is an integer number and denotes the number of audio sources. As the separation of the audio source signal may be imperfect, for example, due to the mixing of the audio sources, a residual signal 3 (r(n)) is generated in addition to the separated audio source signals 2a-2d. The residual signal may for example represent a difference between the input audio content and the sum of all separated audio source signals. The audio signal emitted by each audio source is represented in the input audio content 1 by its respective recorded sound waves. For input audio content having more than one audio channel, such as stereo or surround sound input audio content, also a spatial information for the audio sources is typically included or represented by the input audio content, e.g., by the proportion of the audio source signal included in the different audio channels. The separation of the source audio signal 1 into separated audio source signals 2a-2d and a residual 3 may be performed on the basis of blind source separation or other techniques which are able to separate audio sources. For example, instead of blind source separation, pre-trained speech models may be used for source separation.
[0137] Technical details about source separation process described in Fig. 8 above are known to the skilled person. An exemplifying technique for performing blind source separation is for example disclosed in European patent application EP 3 201 917, or by Uhlich, Stefan, et al. “Improving music source separation based on deep neural networks through data augmentation and network blending.” 2017 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2017. There also exist programming toolkits for performing source separation, such as Open-Unmix, DEMUCS, Spleeter, Asteroid, or the like which allow the skilled person to perform a source separation process as described in Fig. 1 above.
[0138] Implementation
[0139] Fig. 9 illustrates a network of electronic devices and partner devices connected to a server. Electronic device 44a of user 5 is partnered to electronic device 44b of user 6. Both devices are connected to network 50 and via network 50 to server 41. Server 41 includes a database (e.g., 45 of Fig. 10) maintain speech models used for source separations on which basis an audio signal can be decomposed into different sources (see Figs. 7 and 8), for example, source separation A, source separation B or source separation C of source separation 22 of Fig.7. The speech models maintained by server 41 each correspond to a source indicating the source of the speech (24, Fig. 7) that is to be separated from the noise (51, Fig. 7). Therefore, if source separated noise cancellation is performed as described in Fig. 7 the electronic device 44a can access the speech model by accessing the database (45 in Fig. 10) of the server 41 and use it in source separation 22 of Fig. 7. A database of partner IDs (e.g., 21a in Fig. 7) maintained in server 41 is associated with the respective speech models maintained by server 41. That is, each speech model of the database may correspond to an ID and / or to a source-specific audio pattern and the source separation may be selected from the database by electronic device 44a based on the partner ID (21a, Fig. 7), or based on matching the reference signal (for example, a reference audio signal, or features derived from the reference audio signal) to the source-specific audio pattern. Also, new speech models may be created in the database of server 41 that may later be used for source separation, for example, based on audio samples of a specific user, which may be supplied by electronic device 44a or partner device 44b. Also, multiple users may access the database of server 41, for example via one or more than one device, e.g., devices 42 and 43.
[0140] For example, database of server 41 may be accessed via electronic device 42, e.g., which may include a smartphone, of user 5 and partner device 43 of user 6, which may also include a smartphone. In this way users 5 and 6 may access the database 41 for purposes of transmitting audio samples for generating new source separations or for purposes of accessing source separations for source specific noise cancellation (see Figs. 7 and 8), not only via electronic device 44a and partner device 44b, but also via electronic device 42 and partner device 43. Also, electronic device 42 may correspond to or be included in electronic device 44a (corresponding to electronic device 44a of Figs. 3, 4 and 7) and partner device 43 may correspond to or be included in partner device 44b (corresponding to partner device 44b of Figs. 3,4 and 7).
[0141] Alternatively, the database of speech models (e.g., database 45 of Fig. 10) may be included in any one of the electronic devices 44a, 42 or partner devices 44b, 43 instead of server 41 and may be accessed based on partnering by other devices (electronic devices 44a, 42 or partner devices 44b, 43), for example via network 40.
[0142] Fig. 10 illustrates a user creating a user account for accessing a database of speech models for purposes of noise cancellation as illustrated in Figs. 7 and 8. User 6 uses partner device 42 to generate user account 47, which is stored on server 41. User account 47 includes a user ID, which may be used to send a partner ID signal (e.g., 21a, Fig. 7) to a partnered device (e.g., partner device 44b of user 6, wherein partner device 44b takes the function of electronic device 44a as described in Fig. 2). User account 47 also has access to database 45 including speech models for source separation (22 of Fig. 7) of server 41. User 5 may have access to user account 47 and therefore to database 45 via multiple different devices, for example, electronic device 44a (see Fig. 9) or electronic device 42. Thus, access to database 45 for purposes of generating new speech models in database 45, shared access to database 45 with other users or devices and access to database 45 for purposes of noise cancellation as described in Fig. 7 may be organized via user account 47 and / or similar user accounts 47 of other users.
[0143] Fig. 11 schematically illustrates an embodiment of an electronic device comprising circuitry for noise cancellation. The electronic device 100 may be an electronic device 44a or 44b of Figs. 3,4, 7 and 9. The electronic device may be headphones, earphones, a terminal computer, a smartphone, a tablet, a laptop or other mobile device, such as smart glasses, or the like. The electronic device 100 includes a CPU 101 as processor. Additionally, or alternatively, other computation hardware, such as GPU, TPU, DSP etc. may be used. The electronic device 100 further includes camera(s) 106, microphone(s) 107 and loudspeaker(s) 108 that are connected to the processor 101. The processor 101 may for example implement the noise cancellation of Figs. 2 to 7. The processor 101 may for example, be configured to implement the source separation 22 (Fig. 7), the generation of the anti-noise signal 25, 30 (Figs. 4, 7), the partnering of electronic devices 44a and 44b (e.g., Fig. 3), the active noise cancellation (61a, 61b, 30, Figs. 4,7). The microphone 107 may be configured to receive any kind of audio signal. The microphone 107 may be the remote microphone (28b, Figs. 3-6) or the local microphone (28a, Figs. 5- 7).Microphone 107 may be an error microphone (e.g., 57 of Fig. 5). The camera 106 may be one or more cameras, such as an RGB camera, and IR camera, a ToF camera, for example, an iToF or dTof, an event-based camera or the like.
[0144] The electronic device 100 further includes a user interface 109 that is connected to the processor 101. This user interface 109 acts as a man-machine interface and enables a dialogue between a user and the electronic device 100. For example, a user may make configurations to the system using this user interface 109. For example, a user may initialize or access user account 47 of Fig. 10 through the user interface 109.
[0145] The electronic device 100 further includes a Bluetooth interface 104, a WLAN interface 105 and a NFC 111 interface. These units 104, 105, 111 act as I / O interfaces for data communication with external devices (e.g., partner device 44b), such as for partnering as explained for example in Fig. 3. An ethernet interface may also be possible for outside communication. Alternatively, an LED lamp may be included in electronic device 100 for partnering and / or pinging an external device, as explained in Fig. 3. Also, for example, additional loudspeakers, microphones (e.g., error microphone 57, local microphone 28a, microphone 28b of partner device 44b), and cameras, e.g., a ToF camera, RGB camera or an event-based camera, with WLAN, Bluetooth or NFC connection may be coupled to the processor 101 via these interfaces 104, 105, 111.
[0146] The electronic device 100 further includes a data storage 102 and a data memory 103 (here a RAM). The data memory 103 is arranged to temporarily store or cache data or computer instructions for processing by the processor 101. The data storage 102 is arranged as a long-term storage, which may be obtained via the processor 101. Data storage 102 may for example store, database 45 of Figs. 8 and 9.
[0147] The connection between the processor 101 and the camera 106 may include a camera serial interface (CSI). The CSI is an interface between a camera 106 and a host processor 101. Thus, control signals and data from the processor 101 to the camera 106 as well as from the camera 106 to the processor 101 may be sent.
[0148] Furthermore, the electronic device 100 includes an artificial intelligence (Al) processor 110. The Al processor 110 may include a graphics processing unit (GPU) and / or a tensor processing unit 20 (TPU). The Al processor 110 may be configured to execute an Al model (e.g., an artificial neural network), for example, for performing the source separation 22 as explained in Figs. 1, 2 and 4.
[0149] Fig. 12 schematically illustrates an embodiment of a method for noise cancellation. In 51 the audio signal (e.g., 1 of Fig. 4) is acquired. In 52 the reference signal (e.g., 21a of Fig. 7; 20 Figs. 4-6) is received from a partner device (e.g., 44b, Fig. 3- 4, 7). In 53 noise cancellation is performed on the audio signal based on the received reference signal.
[0150] It should be recognized that the embodiments describing methods (e.g., Fig. 12), may describe methods with an exemplary ordering of method steps. The specific ordering of method steps is however given for illustrative purposes only and should not be construed as binding. For example, the order of 51 and 52 of Fig. 12 may be exchanged. Other changes of the ordering of method steps may be apparent to the skilled person.
[0151] Please note that the present disclosure is not limited to any specific division of functions in specific units. For instance, the, source specific ANC, source separation, difference calculation, active noise cancellation or anti-noise signal generation could be implemented by a respective programmed processor, field programmable gate array (FPGA) and the like.
[0152] A method for controlling an electronic device, such as electronic device 100 is described, for example, under reference of Figs. 4 to 8, 12. The method can also be implemented as a computer program causing a computer and / or a processor, such as processor 101 discussed above, to perform the method, when being carried out on the computer and / or processor. In some embodiments, also a non-transitory computer-readable recording medium is provided that stores therein a computer program product, which, when executed by a processor, such as the processor described above, causes the method described to be performed.
[0153] All units and entities described in this specification and claimed in the appended claims can, if not stated otherwise, be implemented as integrated circuit logic, for example on a chip, and functionality provided by such units and entities can, if not stated otherwise, be implemented by software.
[0154] In so far as the embodiments of the disclosure described above are implemented, at least in part, using software-controlled data processing apparatus, it will be appreciated that a computer program providing such software control and a transmission, storage or other medium by which such a computer program is provided are envisaged as aspects of the present disclosure.
[0155] Note that the present technology can also be configured as described below. [1] An electronic device (44a, 42, 100) comprising circuitry configured to perform noise cancellation (61a, 61b) on an audio signal (1) based on a reference signal (21a, 21b, 20) received from a partner device (44b, 43).
[0156] [2] The electronic device (44a, 42, 100) of [1], wherein the reference signal (21a, 20, 21b) is source specific.
[0157] [3] The electronic device (44a, 42, 100) of any one of [1] to [2], wherein the reference signal indicates a source generating part of the audio signal (1), wherein the part of the audio signal (1) generated by the indicated source is not cancelled by the noise cancellation.
[0158] [4] The electronic device (44a, 42, 100) of any one of [1] to [3], wherein the noise cancellation (61a, 61b) comprises a source separation (22) of the audio signal (1).
[0159] [5] The electronic device (44a, 42, 100) of [4], wherein the source separation (22) is configured to separate the audio signal (1) into a speech signal (24, 58) and a noise signal (51).
[0160] [6] The electronic device (44a, 42, 100) of [4] or [5], wherein the source separation (22) of the audio signal (1) is based on the reference signal (21a, 21b, 20) received from the partner device (44b, 43).
[0161] [7] The electronic device (44a, 42, 100) of any one of [5] to [6], wherein the noise cancellation (61b) comprises an anti -noise signal generation (25) configured to generate an antinoise signal (26) based on the noise signal (51).
[0162] [8] The electronic device (44a, 42, 100) of [7], wherein the anti -noise signal generation (25) is configured to generate the anti-noise signal (26) based on only the noise signal (51), excluding the speech signal (24, 58) obtained from source separation from the noise cancellation.
[0163] [9] The electronic device (44a, 42, 100) of any one of [1] to [8], wherein the reference signal is a partner ID (21a).
[0164]
[0010] The electronic device of any one of [4] to [9], wherein the source generates the speech signal (24) of the audio signal (1) which is separated from the noise signal (51).
[0165]
[0011] The electronic device (44a, 42, 100) of any one of [1] to
[0010] , wherein the reference signal (20) is a reference audio signal received from the partner device (44b).
[0166]
[0012] The electronic device (44a, 42, 100) of
[0011] , wherein the noise cancellation (61a) comprises a difference calculation (31) configured to determine a difference signal (21b) based on the audio signal (1) and the reference audio signal (20).
[0013] The electronic device (44a, 42, 100) of
[0012] , wherein the noise cancellation (61a) comprises a noise cancellation (61a, 30) configured to determine an anti-noise signal (26) based on the difference signal (21b).
[0167]
[0014] The electronic device (44a, 42, 100) of [7] to
[0013] , wherein the noise cancellation (61a, 30) is further configured to determine the anti-noise signal (26) based on an error signal (29).
[0168]
[0015] The electronic device (44a, 42, 100) of any one of [7] to
[0014] , wherein the noise cancellation (61a, 30) is based on an adaptive filter (ff(z)) which operates on the audio signal (1) and generates the anti-noise signal (26).
[0169]
[0016] The electronic device (44a, 42, 100) of any one of
[0014] to
[0015] , wherein the noise cancellation (61a, 30) comprises a second adaptive filter (E(z)) which is configured to determine, based on the difference signal (21b), a filtered difference signal (59).
[0170]
[0017] The electronic device (44a, 42, 100) of
[0016] , wherein the second adaptive filter (E(z)) is configured to approximate a transfer function ( / ?(z)) of the transmission path from the partner device (44b, 43) to the electronic device (44a, 42, 100).
[0171]
[0018] The electronic device (44a, 42, 100) of
[0017] , wherein the approximated transfer function (( / ?(z))) of the transmission path from the partner device (44b, 43) to the electronic device (44a, 42, 100) includes a reconstruction of direction of arrival.
[0172]
[0019] The electronic device (44a, 42, 100) of
[0018] , wherein the reconstruction of direction of arrival is based on a geometric configuration of the electronic device (44a, 42, 100) and the partner device (44b, 43).
[0173]
[0020] The electronic device (44a, 42, 100) of
[0018] or
[0019] , wherein the reconstruction of direction of arrival is based on a geometric configuration of the electronic device (44a, 42, 100) and the partner device (44b, 43) relative to each other.
[0174]
[0021] The electronic device (44a, 42, 100) of
[0020] , wherein the geometric configuration of the electronic device (44a, 42 100) and the partner device (44b, 43) relative to each other is determined based on localization.
[0175]
[0022] The electronic device (44a, 42, 100) of any one of
[0016] to
[0021] , wherein the second adaptive filter (E(z)) is obtained by recursively subjecting the refence audio signal (20) to the second adaptive filter (E(z)) to obtain a filtered refence audio signal (59b) and determine a second difference signal between the audio signal (1) and the filtered reference audio signal (59b).
[0023] The electronic device (44a, 42, 100) of any one of
[0016] to 22, wherein the adaptive filter (H (z)) is optimized based on the filtered difference signal (59) and the error signal (29) in a way that the error signal (29) is minimized.
[0176]
[0024] The electronic device (44a, 42, 100) of any one of
[0016] to
[0023] , wherein the noise cancellation (30) is further configured to determine an added signal based on the filtered difference signal (59) and the error signal (29).
[0177]
[0025] The electronic device (44a, 42, 100) of
[0023] or
[0024] , wherein optimizing the adaptive filter (H (z)) is based on the added signal.
[0178]
[0026] A method comprising: performing noise cancellation (61a, 61b) on an audio signal (1) based on a reference signal (21a, 20) received from a partner device (44b, 43).
[0179]
[0027] The method of
[0026] , wherein the reference signal (21a, 20) is source specific.
[0180]
[0028] The method (44a, 42, 100) of any one of
[0026] to
[0028] , wherein the reference signal indicates a source generating part of the audio signal (1), wherein the part of the audio signal (1) generated by the indicated source is not cancelled by the noise cancellation.
[0181]
[0029] The method (44a, 42, 100) of any one of
[0026] to
[0028] , wherein the noise cancellation (61a, 61b) comprises a source separation (22) of the audio signal (1).
[0182]
[0030] The method (44a, 42, 100) of
[0029] , wherein the source separation (22) is configured to separate the audio signal (1) into a speech signal (24, 58) and a noise signal (51).
[0183]
[0031] The method (44a, 42, 100) of
[0029] or
[0030] , wherein the source separation (22) of the audio signal (1) is based on the reference signal (21a, 21b, 20) received from the partner device (44b, 43).
[0184]
[0032] The method (44a, 42, 100) of any one of
[0030] to
[0031] , wherein the noise cancellation (61b) comprises an anti -noise signal generation (25) configured to generate an anti -noise signal (26) based on the noise signal (51).
[0185]
[0033] The method (44a, 42, 100) of
[0032] , wherein the anti-noise signal generation (25) is configured to generate the anti-noise signal (26) based on only the noise signal (51), excluding the speech signal (24, 58) obtained from source separation from the noise cancellation.
[0186]
[0034] The method (44a, 42, 100) of any one of
[0026] to
[0033] , wherein the reference signal is a partner ID (21a).
[0187]
[0035] The method of any one of
[0029] to 34], wherein the source generates the speech signal (24) of the audio signal (1) which is separated from the noise signal (51).
[0036] The method (44a, 42, 100) of any one of
[0026] to
[0035] , wherein the reference signal (20) is a reference audio signal received from the partner device (44b).
[0188]
[0037] The method (44a, 42, 100) of
[0036] , wherein the noise cancellation (61a) comprises a difference calculation (31) configured to determine a difference signal (21b) based on the audio signal (1) and the reference audio signal (20).
[0189]
[0038] The method (44a, 42, 100) of
[0037] , wherein the noise cancellation (61a) comprises a noise cancellation (61a, 30) configured to determine an anti-noise signal (26) based on the difference signal (21b).
[0190]
[0039] The method (44a, 42, 100) of
[0032] to
[0038] , wherein the noise cancellation (61a, 30) is further configured to determine the anti-noise signal (26) based on an error signal (29).
[0191]
[0040] The method (44a, 42, 100) of any one of
[0032] to
[0039] , wherein the noise cancellation (61a, 30) is based on an adaptive filter (H (z)) which operates on the audio signal (1) and generates the anti-noise signal (26).
[0192]
[0041] The method (44a, 42, 100) of any one of
[0039] to
[0040] , wherein the noise cancellation (61a, 30) comprises a second adaptive filter (E(z)) which is configured to determine, based on the difference signal (21b), a filtered difference signal (59).
[0193]
[0042] The method (44a, 42, 100) of
[0041] , wherein the second adaptive filter (E(z)) is configured to approximate a transfer function ( / ?(z)) of the transmission path from the partner device (44b, 43) to the method (44a, 42, 100).
[0194]
[0043] The method (44a, 42, 100) of
[0042] , wherein the approximated transfer function (( / ?(z))) of the transmission path from the partner device (44b, 43) to the method (44a, 42, 100) includes a reconstruction of direction of arrival.
[0195]
[0044] The method (44a, 42, 100) of
[0043] , wherein the reconstruction of direction of arrival is based on a geometric configuration of the method (44a, 42, 100) and the partner device (44b, 43).
[0196]
[0045] The method (44a, 42, 100) of
[0043] or
[0044] , wherein the reconstruction of direction of arrival is based on a geometric configuration of the method (44a, 42, 100) and the partner device (44b, 43) relative to each other.
[0197]
[0046] The method (44a, 42, 100) of
[0045] , wherein the geometric configuration of the method (44a, 42 100) and the partner device (44b, 43) relative to each other is determined based on localization.
[0047] The method (44a, 42, 100) of any one of
[0041] to
[0046] , wherein the second adaptive filter (E(z)) is obtained by recursively subjecting the refence audio signal (20) to the second adaptive filter (E(z)) to obtain a filtered refence audio signal (59b) and determine a second difference signal between the audio signal (1) and the filtered reference audio signal (59b).
[0048] The method (44a, 42, 100) of any one of
[0041] to
[0047] , wherein the adaptive filter (H (z)) is optimized based on the filtered difference signal (59) and the error signal (29) in a way that the error signal (29) is minimized.
[0198]
[0049] The method (44a, 42, 100) of any one of
[0041] to
[0048] , wherein the noise cancellation (30) is further configured to determine an added signal based on the filtered difference signal (59) and the error signal (29).
[0199]
[0050] The method (44a, 42, 100) of
[0048] or
[0049] , wherein optimizing the adaptive filter (H (z)) is based on the added signal.
[0200]
[0051] A computer program comprising program code causing a computer to perform the method according to anyone of
[0026] to
[0050] , when being carried out on a computer.
[0052] A non-transitory computer-readable recording medium that stores therein a computer program product, which, when executed by a processor, causes the method according to anyone of
[0026] to
[0050] to be performed.
Claims
CLAIMS1. An electronic device comprising circuitry configured to perform noise cancellation on an audio signal based on a reference signal received from a partner device.
2. The electronic device of claim 1, wherein the reference signal is source specific.
3. The electronic device of claim 1, wherein the noise cancellation comprises a source separation of the audio signal.
4. The electronic device of claim 3, wherein the source separation is configured to separate the audio signal into a speech signal and a noise signal.
5. The electronic device of claim 3, wherein the source separation of the audio signal is based on the reference signal received from the partner device.
6. The electronic device of claim 4, wherein the noise cancellation comprises an anti-noise signal generation configured to generate an anti-noise signal based on the noise signal.
7. The electronic device of claim 6, wherein the anti-noise signal generation is configured to generate the anti-noise signal based on only the noise signal, excluding the speech signal obtained from source separation from the noise cancellation.
8. The electronic device of claim 1, wherein the reference signal is a partner ID (21a).
9. The electronic device of claim 1, wherein the reference signal is a reference audio signal received from the partner device.
10. The electronic device of claim 9, wherein the noise cancellation comprises a difference calculation configured to determine a difference signal based on the audio signal and the reference audio signal.
11. The electronic device of claim 10, wherein the noise cancellation comprises a noise cancellation configured to determine an anti-noise signal based on the difference signal.
12. The electronic device of claim 11, wherein the noise cancellation is further configured to determine the anti-noise signal based on an error signal.
13. The electronic device of claim 11, wherein the noise cancellation is based on an adaptive filter which operates on the audio signal and generates the anti-noise signal.
14. The electronic device of claim 12, wherein the noise cancellation comprises a second adaptive filter which is configured to determine, based on the difference signal, a filtered difference signal.
15. The electronic device of claim 14, wherein the second adaptive filter is configured to approximate a transfer function of the transmission path from the partner device to the electronic device.
16. The electronic device of claim 15, wherein the approximated transfer function of the transmission path from the partner device to the electronic device includes a reconstruction of direction of arrival.
17. The electronic device of claim 16, wherein the reconstruction of direction of arrival is based on a geometric configuration of the electronic device and the partner device.
18. The electronic device of claim 17, wherein the reconstruction of direction of arrival is based on a geometric configuration of the electronic device and the partner device relative to each other.
19. The electronic device of claim 18, wherein the geometric configuration of the electronic device and the partner device relative to each other is determined based on localization.
20. The electronic device of claim 14, wherein the second adaptive filter is obtained by recursively subjecting the refence audio signal to the second adaptive filter to obtain a filtered refence audio signal and determine a second difference signal between the audio signal and the filtered reference audio signal.
21. The electronic device of claim 14, wherein the adaptive filter is optimized based on the filtered difference signal and the error signal in a way that the error signal is minimized.
22. The electronic device of claim 21, wherein the noise cancellation is further configured to determine an added signal based on the filtered difference signal and the error signal.
23. The electronic device of claim 22, wherein optimizing the adaptive filter is based on the added signal.
24. A method comprising: performing noise cancellation on an audio signal based on a reference signal received from a partner device.
25. The method of claim 24, wherein the reference signal is source specific.
Citation Information
Patent Citations
Method, apparatus and system
EP3201917A1
Signal processing apparatus, method, and system
US20220335923A1
Electronic device for controlling ambient sound based on audio scene and operating method thereof
US20230112073A1
Active noise cancellation method, device, and system
US20230335101A1