Systems and methods for environmental noise detection, identification, and management
By integrating microphone, speaker and processing unit in headphone equipment to identify and filter specific noise, the problem of indistinguishable and filtering specific noise in the prior art is solved, and personalized noise management is realized to ensure that users do not miss important information in noisy environments.
Patent Information
- Application Number
- CN202080088365.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-12-19
- Filing Date
- 2020-12-14
- Publication Date
- 2025-08-19
- Estimated Expiration
- 2040-12-14
AI Technical Summary
The prior art cannot effectively distinguish and filter ambient noise defined by a particular user, and traditional devices may limit social communication or cannot effectively suppress target noise.
The headphone equipment is equipped with a microphone and speaker, combined with a processing unit, a memory unit and an identification filtering processing unit, and automatically remove, weaken or mask disgusting sounds by identifying feature maps and suppressing models.
Personalized filtering of specific noise is achieved to ensure that users can hear important voice information while reducing the negative impact of environmental noise.
Smart Images

Figure CN114846539B_ABST
Abstract
Description
Technical Field
[0001] The present invention generally relates to systems and methods for detecting, identifying / classifying, and managing ambient aversive sounds, and more particularly to intelligent systems and methods for ambient aversive sound reduction and suppression. Background Art
[0002] There are several types of objectionable sounds, such as industrial and construction noise, which are known to be harmful to the ears, and which can be harmful when they are too loud, even for a short period of time. Every day, many people around the world are exposed to these noises. A noisy environment during a conversation can be irritating, making it difficult for people to focus on what another person is saying and they may miss important information in the speech. In addition, some people have a reduced tolerance for everyday noise (hyperacusis) and react negatively to very specific noises. Sound sensitivity affects the general public and is very common in people with autism spectrum disorder (ASD). Negative reactions to sound can be extremely debilitating, interfering with social interactions and participation in daily activities. Therefore, addressing this issue is an important issue for everyone, but especially for people who are sensitive to general sounds and / or specific sounds.
[0003] Known systems and devices exist that address this issue using signal processing methods and / or mechanical devices. Some of these known systems focus on ear protection by controlling the volume of sound and denoising speech, as taught in U.S. Patent Application 2014 / 0010378A1, or by masking noise by playing music or other specific sounds, as disclosed in U.S. Patent Application 20200312294. Several documents, such as U.S. Patent Applications US9524731B2 and US9648436B2, address this issue by extracting specific features from the user's digitized ambient sound and providing appropriate information to aid the listener's hearing or protect them from dangerous sounds.
[0004] In some known devices, active noise cancellation techniques are implemented to address this issue, such that incoming sounds from the earpiece are detected and a signal out of phase with the unpleasant signal is generated to cancel the unpleasant sound. These active noise cancellation techniques are most effective for lower frequency and continuous sounds between 50 Hz and 1 kHz.
[0005] A major drawback of existing systems and devices is that they are designed to attenuate or remove all ambient noise, regardless of the nature of the sound. For example, a person may want to listen to the sounds of children playing while suppressing unpleasant street noise. Another limitation is that devices that attenuate or eliminate ambient noise may limit social interaction because the user cannot hear speech. Known devices that can denoise speech in noisy environments are designed for industrial settings where workers wear the same device and communicate via a telecommunications transmission link (e.g., US20160126914A1).
[0006] There are also known systems and devices that use multi-microphone technology located at different spatial locations of the device to suppress offensive sounds (as taught in EPO patent application EP3096318A1). However, these systems are not practical because in most cases these multi-microphone technologies cannot successfully suppress the target noise, especially when the microphones capture the same signal from the surrounding environment, or they move and shake when the user is active. In addition, based on the acoustic design of earbuds and headphones, the implementation of multi-channel microphones and intelligent algorithms is quite difficult and expensive. On the other hand, newer methods use single-channel audio to identify and suppress noise (as disclosed in EPO patent application EP3175456B1). Systems and methods based on single-channel audio recognition and noise suppression in practical applications are effective and practical, but such systems have limitations in quality and accuracy.
[0007] Therefore, what is needed is a system and method that can perceive the noise content to filter specific user-defined ambient noise, rather than a system that treats all ambient noise equally. Summary of the Invention
[0008] In one aspect, a system for detecting, identifying, and managing unpleasant ambient sounds is provided. The system includes: a headphone device having a microphone and a speaker, the microphone configured to capture the ambient sounds surrounding a user as samples of small segments of ambient sound; an interface having a hardware processor programmed with executable instructions for obtaining input data, transmitting data, and providing output data; a memory unit storing input information, a library of unpleasant ambient sound signals, unpleasant sound feature maps, a recognition prediction model, unpleasant sound recognition categories, and an unpleasant sound suppression prediction model; and a processing unit in communication with the headphone device, the memory unit, and the interface. The processing unit includes a recognition unit coupled to the headphone device and the memory unit; and a filtering processing unit coupled to the recognition unit and the memory unit. The recognition unit is programmed with executable instructions to identify unpleasant ambient sound signals in a sound segment by extracting at least one feature of the sound in the sound segment and creating a feature map for the sound segment. The recognition unit processes a feature map of the sound segment using a recognition prediction model stored in the memory unit to compare at least one feature in the feature map with a feature map of an unpleasant sound signal stored in the memory unit. When an unpleasant ambient sound is identified, the recognition unit classifies the unpleasant sound signal using a recognition category. The filtering unit is programmed with executable instructions to receive a mixed signal of the ambient sound segment and the unpleasant sound signal, process the mixed signal, calculate the amplitude and phase of the mixed signal to generate a feature map of the mixed signal, compare the feature map with the stored feature maps using at least one unpleasant sound suppression model, and provide a recommended action to manage the unpleasant sound signal. The headphone device also includes a valve for regulating the ambient sound segment transmitted to the user.
[0009] In one aspect, the recommended action is to remove the offensive sound, and the filtering processing unit is further programmed with executable instructions to automatically remove the identified offensive signal from the mixed signal using the generated feature map to obtain a clean sound in the mixed signal reconstructed from the frequency domain to the time domain, and combine the phase of the mixed signal with the amplitude of the clean sound of the final segment to create a clean sound signal for transmission to the speaker.
[0010] In another aspect, the recommended action is to attenuate the offensive sound, and the filter processing unit further includes a bypass having a gain to create the attenuated offensive sound. The filter processing unit is programmed with executable instructions to automatically add the attenuated offensive sound signal to the clean sound signal.
[0011] In yet another aspect, the filtering processing unit is programmed with executable instructions to automatically add pre-recorded sounds onto the mixed signal to mask the objectionable ambient sound signal.
[0012] In one aspect, the system includes a first activation device in communication with the earphone device to manually trigger a valve to suppress or attenuate an objectionable ambient sound signal, and a second activation device in communication with the earphone device and a memory unit to manually access stored pre-recorded sounds and play them over a mixed signal to mask the objectionable ambient sound signal.
[0013] In one aspect, the system includes an alert system that communicates with the interface and / or headset device to generate an alert signal to alert the user of the recommended action, the alert signal being selected from one of a visual signal, a tactile signal, an audible signal, or any combination thereof.
[0014] In another aspect, the system includes at least one physiological sensor in communication with a processing unit, the at least one physiological sensor configured to detect at least one physiological parameter of a user. The processing unit identifies an objectionable ambient sound signal in the ambient sound segment if the detected parameter is outside a predetermined range of at least one of the detected physiological parameters.
[0015] In one aspect, a method for detecting, identifying, and managing unpleasant ambient sounds is provided. The method includes: capturing a stream of ambient sound clips using a microphone in a headphone device; storing input information, a library of unpleasant ambient sound signals, an unpleasant sound feature map, a recognition prediction model, an unpleasant sound recognition category, and an unpleasant sound suppression prediction model in a memory unit; and processing the captured sound clips by a processing unit. The processing steps include: extracting at least one feature of a sound signal in the sound clip, creating a feature map for the sound clip, comparing at least one feature in the feature map with a feature map of unpleasant sound signals in the recognition prediction model stored in the memory unit, identifying unpleasant sound signals in the captured sound clips using the recognition prediction model, and classifying the identified unpleasant sound signals using a recognition category; filtering a mixed signal including the ambient sound clips and the unpleasant ambient sound signals, calculating the amplitude and phase of the mixed signal to generate a feature map of the mixed signal, comparing the feature map with the stored features using at least one unpleasant sound suppression model, and providing recommended actions to manage the unpleasant sound signals.
[0016] In addition to the aspects and embodiments described above, further aspects and embodiments will become apparent by reference to the drawings and by study of the following detailed descriptions. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 An exemplary schematic diagram of a system for detecting, identifying, and managing environmental objectionable sounds is shown.
[0018] Figure 2 A simplified block diagram illustrating an example of a method for detecting, identifying / classifying, and managing environmental objectionable sounds according to one embodiment of the present invention.
[0019] Figure 3 A simplified block diagram showing an example of a recognition / classification method according to one embodiment of the present invention is shown.
[0020] Figure 4 A simplified block diagram illustrating an example of a real-time objectionable sound filtering and processing method according to one embodiment of the present invention.
[0021] Figure 5 An example of a graphical user interface according to one embodiment of a system for detecting, identifying, and managing environmental objectionable sounds is shown. DETAILED DESCRIPTION
[0022] Embodiments of the present invention provide device details, features, systems and methods for ambient objectionable sound detection and identification / classification, as well as systems and methods for ambient objectionable sound isolation and suppression.
[0023] Figure 1A schematic example of a system 100 for detecting, identifying / classifying, and managing objectionable ambient sounds is shown. System 100 includes an earpiece device 101, such as a set of earbuds or earmuffs, with a microphone 104. In some embodiments, system 100 may also be equipped with a speaker 105 and / or a valve 102. Valve 102 can suppress or isolate ambient sounds. Valve 102 can be opened or closed, such that when the valve is in "open mode," ambient sounds can pass through system 100 with some attenuation due to the passive noise isolation provided by the materials in the earbuds or earmuffs, while when valve 102 is in "closed mode," ambient sounds are attenuated or isolated. In one embodiment, valve 102 can be electronically activated / deactivated using a solenoid, voice coil, or the like. For example, an activation button 103 can be used to operate valve 102. Button 103 can be positioned somewhere on earpiece device 101 or on a wearable or handheld device 108 or 110, and an electronic trigger signal can be transmitted to activation device 103 via a wired or wireless mechanism. The activation device 103 may be any known mechanical or electronic activation mechanism. In another embodiment, the device does not include a valve, and the device 101 may include a standard in-ear bud or an earphone, each equipped with a microphone and a speaker, and the system 100 may perform the classification and sound management operations described below.
[0024] The system 100 also includes a processing unit 107. The processing unit 107 may be located in the headset device 101, or may be located external to the headset device 101. For example, the processing unit 107 may be in a smart wearable device 110, or a handheld device 108 (e.g., a mobile phone, a tablet, a laptop), or any other suitable smart device. In one implementation, the processing unit 107 may be in a hardware server or a virtual server (cloud server) that communicates with the headset device 101 using Internet communication. In one embodiment, a portion of the processing unit 107 may be in the headset device 101, and another portion may be in a smart wearable / device, or any other computer, or server. The processing unit 107 may include a memory unit, such as a non-volatile memory, for storing data, identifiers, or classifier models and instructions ( Figure 3 ), and filter processing models and instructions ( Figure 4 ). The system 100 may also include Figure 5 5. The user can provide input data to the system 100 through the interface 500. The memory unit stores the input data, the captured ambient sound data, the library of known ambient sounds, the sound feature map, and the prediction model used by the classifier 203 and the filter processor 205.
[0025] Figure 2 The diagram illustrates the overall steps 200 performed by the system 100 to detect, classify, and manage objectionable ambient sounds. The microphone 104 of the earphone device 101 captures ambient sounds 201 and can play these sounds to the speaker 105. When the device is operating in "normal" mode (the valve 102 is in "open mode," or the speaker plays the undistorted ambient sounds captured by the microphone), the user can hear the ambient sounds. When objectionable ambient sounds are detected, the system 100 is activated and the system 100 can take recommended actions such as (1) suppressing the detected sounds including the objectionable ambient sounds by activating the valve 102 so that the device operates in "off mode" and the valve / plug blocks the objectionable ambient sounds from entering the ear canal, independent of other operations; (2) suppressing the signal by stopping the transmission of ambient sounds from the microphone to the speaker 105; (3) attenuating the ambient sounds by reducing the volume so that the earphones of the headset play sounds at a reduced volume (the user selects the volume or there is a predefined lower volume); (4) removing the objectionable portion of the ambient sounds using the system 100 as described below, or (5) masking the ambient sounds by playing pre-recorded sounds to the speaker 105 as an additional masking feature (e.g., playing music, white noise, or any sounds that the user prefers). By playing music, white noise, or other preferred sounds during the off mode, the system 100 can maximize the masking of ambient sounds beyond what passive isolation (options 1 and 2) can provide. When no longer detecting the offensive sound, the operation of the system will return to "normal" mode. Microphone 104 can be mounted on the outside of the headphone device 101 and can capture ambient sound. Microphone 104 is configured to capture samples of small segments of the sound stream that are considered to be frame selection 202. The frame size is selected so that the human auditory system cannot understand the processing, decision-making, and associated delays of offensive sound suppression for each sound segment / frame. The human ear can tolerate a delay (latency) between 20 and 40ms, and system 100 is configured to have a delay within this range to work smoothly in real time. For example, a frame size of 24ms can be used, and microphone 104 can capture a 24ms sound frame at a sampling rate of 8Khz at a time. In this setting, the segment of signal 201 includes 192 data samples.
[0026] An ambient sound segment is input to processing unit 107, where it is processed to identify whether the segment contains an unpleasant sound. As previously described, this processing can be performed within the headphone device 101, or remotely within a smart wearable device 110, a handheld device 108 (e.g., a mobile phone, tablet, laptop), or any other suitable hardware server or virtual server (cloud server) that communicates with the headphone device 101 using internet communications. Processing unit 107 includes an identifier / classifier 203 having a prediction model that is or will be trained to detect and identify unpleasant sounds, which may be predefined by the user or may be unpleasant sounds that the user has not previously heard. Classifier 203 can use the classification prediction model and a library of unpleasant sounds to determine whether the segment contains an unpleasant ambient sound. If classifier 203 does not identify any unpleasant sounds in the sound segment, such a signal is transmitted to speaker 105 and the user can hear it. If the classifier 203 identifies that the ambient audio segment includes an offensive sound 204, the mixed signal of the ambient audio segment and the offensive sound is processed by a filtering processor 205, where the offensive sound suppression prediction model is used to determine specific features of the mixed signal for offensive sound suppression purposes. In one implementation of the system 100, the processing unit 107 can use the results of the filtering processor 205 to automatically remove or suppress the offensive sound signal from the mixed signal, thereby generating a clean sound 206 as output. The clean sound 206 can then be played on the speaker 105, which is located in the headphone device 101 (e.g., earphone or earpiece).
[0027] In one implementation, the processing unit 107 can provide a recommended action to the user. For example, the processing unit 107 can use the interface 500 to send an alert with a recommended action. The recommended action can be, for example, (1) suppressing the sound signal by closing the valve 102, or (2) suppressing the signal by stopping the transmission of the signal from the microphone to the speaker, or (3) reducing the ambient sound by lowering the volume, or (4) removing the offensive sound signal from the mixed signal, or (5) masking the ambient sound by playing a pre-recorded sound. The user can then use the interface 500, or the activation device 103, or in some implementations, a plug to decide and manually perform the recommended action. In another implementation, the system can automatically trigger the valve or activation device to suppress or reduce the signal, or activate a player that stores a pre-recorded masking sound, or provide instructions to the processing unit 107 to remove the offensive sound from the signal.
[0028] The masking sounds may be pre-recorded and stored in a memory unit, or a separate player storing pre-recorded sounds may be provided in communication with the headphone device 101 which may be triggered manually by the user or automatically by the system 100. In one embodiment, the masking sounds may be stored in an application such as Spotify or similar. TM , and can be accessed manually by the user or automatically by the system 100. In embodiments where the system 100 automatically suppresses, attenuates, removes, or masks objectionable ambient sounds, the user may have the ability to override the system's recommended actions by manually deactivating and activating the system 100 using, for example, the interface 500. The system 100 may continuously monitor ambient sounds and sample new sound segments and process such segments to identify objectionable ambient sounds and suppress, attenuate, remove, or mask the objectionable sounds accordingly, as described above.
[0029] Graphical User Interface 500( Figure 5 ) can be compatible with every operating system running on the smart wearable device or handheld device 108, 110. The interface 500 has a hardware processor that is programmed with executable instructions for acquiring input data, transmitting data, and providing output data. For example, the graphical user interface 500 can be responsible for all communications, settings, customizations, and user-defined operations. The computing programs including artificial intelligence prediction models, model training, data sampling, and data processing are executed by the processor 107, which can be in the headphone device 101 or any device 108, 110, or on a server connected to the system via an Internet network, or any combination thereof. The system 100 can also include a rechargeable battery 106 that provides the required power for the headphone device 101. The battery can be charged using a complimentary charger unit.
[0030] In one implementation, the system 100 may also include a set of physiological sensors to capture body reactions, such as skin conductance, heart rate, electroencephalogram (EEG), electrocardiogram (ECG), electromyogram (EMG), or other signals. These sensors may be embedded in a wearable device 110 or any other smart device 108, or attached to the user's clothing or body. The provided sensor data may be transmitted to the system 100 using a wireless or wired connection 109 to determine body reactions. For example, a heart rate monitor may be used to detect the user's heart rate, and the processing unit 107 may process such signals to determine the user's stress state, such as, for example, an increased heart rate may be a signal that the user's stress level is increasing, or a skin conductance monitor may detect an increased level of physiological arousal in the user. Such body reactions may be the result of an unpleasant environmental sound, so the system 100 may automatically recommend that the correct action be performed accordingly, as described herein.
[0031] In some embodiments, the system 100 is programmed and includes algorithms / models that can distinguish and suppress some offensive sounds. These models can be based on machine learning, deep learning, or other sound detection / recognition and suppression techniques.
[0032] Figure 3 An example of an identifier / classifier 203 and a method 300 performed by the classifier 203 is shown. The classifier receives as input a segmented ambient sound signal 301 obtained from the microphone 104. Signal 301 is then preprocessed in step 302 to flatten or create a signal representation in the frequency domain using a Fast Fourier Transform (FFT) of the signal. For example, the preprocessing step may also include signal processing operations (e.g., normalization, sampling rate conversion, windowing, FFT, flattening) to provide input to a classifier prediction model (e.g., a deep neural network, a support vector machine (SVM), linear regression, a perceptron, etc.) for unpleasant sound detection / recognition. Predetermined features are then calculated / extracted in step 303 using, for example, Mel-Frequency Cepstral Coefficients (MFCCs), Short-Time Energy (STE), etc. For example, the intensity and power (i.e., amplitude and frequency) of signal 301 are extracted, and a feature map 303 is created and input into the classifier prediction model 304. The classification prediction model is a pre-trained model, which in one example is an artificial neural network for learning a program, comprising multiple hidden layers (e.g., dense layers, GRU layers, etc.). The classification prediction model is trained on a comprehensive set of offensive sound categories, and therefore each offensive sound category has a corresponding category number. Feature map 303 is fed to model 304 and output 305 is a category number based on the similarity between signal 301 and the sound pattern in the library. The identified category can be a number that identifies a specific offensive sound. For example, category number 2 can be the sound of an air conditioner, and category number 7 can be, for example, a whistle, etc. In one implementation, the identification category 305 can be a non-written (verbal) text.
[0033] In one embodiment, if the identifier / classifier identifies an objectionable sound in the ambient sound, the user is notified using a graphical interface and can choose to use the system 100 to automatically close the valve, suppress, attenuate, or remove the objectionable sound and play the remaining ambient sound into the earpiece, or play music, white noise, or any preferred sound, or even a combination of music and suppressed ambient sound. The user can customize the device to perform any of the above operations, and these operations can also be changed when an objectionable sound is identified based on the user's determination of the objectionable nature of the sound. The mixed signal including the objectionable ambient sound identified with the identified category 305 is then input into the filtering processor 205.
[0034] Figure 4 An example of a filtering processor 205 and a method 400 performed by the filtering processor 205 is shown. In one embodiment, the filtering processor is activated when a classifier identifies an unpleasant sound in an ambient sound environment. Therefore, the input 301 is a mixed signal of ambient sound, including the unpleasant sound. The filtering processor 205 receives a segment of the mixed signal as an input frame and calculates the signal's amplitude and phase. The amplitude is then preprocessed, as shown in step 401, to construct a feature map 402, which is subsequently used as input to a prediction model for unpleasant sound suppression. In one example, the preprocessing includes obtaining an FFT of the signal, performing spectral analysis, and performing mathematical operations to create a representation of the mixed signal in the frequency domain as a feature map. The preprocessing steps in the filtering process may differ from those in the recognition process depending on the purpose. In one embodiment, the filtering processor 205 may obtain several overlapping, segmented frames of the mixed signal and perform preprocessing to calculate the power and frequency in the mixed signal and construct the feature map. For example, pre-processing may include signal processing operations to provide input to an unpleasant sound suppression model 403 (e.g., a deep neural network, SVM, linear regression, perceptron) for unpleasant sound suppression. The unpleasant sound suppression model 403 may be trained to obtain a flattened (pre-processed) feature map and remove the unpleasant components of the identified input mixed signal 301. This produces a clean sound in the input mixed signal. The clean sound is in the frequency domain and therefore needs to be reconstructed into a clean sound in the time domain by using, for example, an inverse fast Fourier transform (IFFT) 404, and the phase 405 of the mixed signal extracted from the original mixed signal 301 is combined with the amplitude of the clean sound to create a clean signal. In one embodiment, a post-processing tool 406 including powering and flattening may be applied to the clean signal to form a smooth clean signal. For example, in one embodiment, the post-processing tool 406 may use an overlap-add method that can be executed in real time to generate a time domain of clean sound. As previously described, the frame size of the captured ambient signal is selected so that human perception cannot understand the associated delays in processing, decision-making, and unpleasant sound suppression for each sound frame. The human ear can tolerate a delay between 20 and 40 ms, and the system 100 can have a delay within this range to work smoothly in real time. For example, the microphone 104 captures a 24 ms frame at a sampling rate of 8 kHz, so that the segment of the signal includes 192 data samples. Taking into account the 50% overlap of the overlap addition method, the post-processing tool 406 adds a 12 ms overlap of the clean signal, which is played to the speaker 105. The overlap technique provides the advantages of smooth continuous framing and keeps information at the edge of the frame to generate clean, smooth and continuous sound without losing sound quality. Therefore, the output of the post-processing 406 can be an estimated clean signal 407.
[0035] The inherent structure and pattern of offensive sounds can be categorized into three basic types: Figure 3 Unpleasant sounds are considered based on the categories 305 identified in the image processing unit 206. These categories include static noises such as air conditioners, engines, etc.; non-static noises such as trains, wind, etc.; and highly dynamic noises such as dog barking, whistles, and babies crying. In one embodiment, the system and method can perform different pre-processing and post-processing on the mixed signal 410 based on the identified category of unpleasant sounds. For example, digital signal processing filtering assistance such as adaptive filters is applied to non-static category noises. For one embodiment, the filtering processor 205 may include multiple different unpleasant sound suppression models to select between them based on the category / category of the unpleasant sound identified in order to generate clean sound. The model for unpleasant sound suppression is trained on a comprehensive dataset that results in accuracy and high performance, and is trained using a reliable deep neural network model or other machine learning model.
[0036] In one embodiment, the filter processor 205 includes an unpleasant signal bypass 408 that is configured to attenuate unpleasant sounds and add the attenuated unpleasant sounds to the clean sound. For example, a user can use the settings 506 in the interface 500 to select the attenuation level of the unpleasant sounds. In some implementations, a user can use the slider 505 in the interface 500 to manually attenuate the unpleasant sound level. The bypass 408 with gain is considered and multiplied by the estimated unpleasant signal, which is the subtraction between the mixed signal and the estimated clean signal, to create the attenuated unpleasant signal. This signal is then added to the clean signal so that the user can hear the attenuated unpleasant sound with the clean sound through the speaker 105. The attenuation level from zero to maximum attenuation in the bypass gain 408 can be set in the settings or using a designed knob, slider, or button in the user interface 500.
[0037] Figure 5 An example of a graphical user interface 500 of the system 100 is shown. The graphical user interface 500 can be installed on any smart device, including a smartphone, a smart watch, a personal computer on any suitable operating system such as iOS, Windows, Android, etc. The installed user interface can interact with the headset device 101 using a wired or wireless connection. Figure 5In the illustrated example, screen 501 can indicate the system operating mode. For example, the operating mode can be "normal mode," which refers to the absence of unpleasant sounds identified by the system and method, and "unpleasant mode," which refers to unpleasant sounds identified by the system and method. In addition, screen 501 can be used to alert the user of actions recommended by processing unit 107, such as, for example, suppressing, reducing, or masking the sound. Activation slider 502 is configured to manually turn system 100 on and off. The user can use slider 503 to activate the recorder or access a music storage device. The user can play and pause music, white noise, or any other preferred sound based on the surrounding environment or their preferences. In addition, the user can use slider 504 to control the volume of the sound entering the earpiece. When an unpleasant sound is identified and notified, the user can use slider 505 to selectively reduce the unpleasant sound identified. The settings 506 are configured for user customization so that the user can enter their preset preferences, such as, for example, selecting and uploading favorite music / sounds to be stored in the memory unit or a separate recording unit, selecting an alarm setting (e.g., an audible alarm or any alarm signal, such as light, or vibration, or text), selecting an alarm sound, specifying an action to be performed when an offensive sound is identified (e.g., automatic offensive sound management or manual offensive sound management), a list of offensive sounds to be suppressed and / or notified, a button to activate a single learning process. The interface 500 may also include an exit button 507 to close the entire program. For those who do not Figure 5 In the embodiment of the settings in, the customization process can take place using a personal computer, laptop, smartphone, etc. The user can complete the customization process as described above by logging into a website or a dedicated interface provided to them, and after completion, the final settings can be transmitted to the memory of the headset, a cloud server, etc. for implementation by the processing unit.
[0038] In one embodiment, an alert system is used to notify the user of the presence of an unpleasant sound, and the system 100 automatically suppresses, attenuates, or masks the unpleasant sound. When the system 100 identifies an unpleasant sound nearby (e.g., a sound defined by the user as unpleasant), the alert system notifies the user of the unpleasant sound by playing a specific piece of music, white noise, an alert voice, a light alarm (e.g., a colored LED), a text message, a beep in the headset, a vibration, or any combination of these. The user can also be informed of the nature and intensity of the unpleasant sound and a recommended action, such as suppressing, attenuating, or masking, or any combination thereof. The user can veto any recommended action using the interface 500 or a veto button on the headset. The alert system can be incorporated into the interface 500 and can communicate with the user using any smart wearable device, handheld device, or headset. The alert system can be customized by the user. The system can also notify the user of the removal of the unpleasant sound.
[0039] In one embodiment, the user can add a user-defined list of offensive sounds for the system to detect and manage. In such an embodiment, the system 100 can start with a predefined list of offensive sounds, but when the user hears an offensive sound that is not predefined, they can activate the system 100 to record the ambient sounds and situation (obtaining appropriate samples of the sounds and situations), process the sounds to identify the various components (offline or online), communicate the findings to the user, and ask the user to specify which of the identified sounds is offensive, and finally add the offensive sound to the user's customized library of offensive sounds. In one embodiment, the learning component can be based on a one-shot or few-shot learning approach, or other intelligent algorithms such as machine learning and deep learning. For example, the system can record ambient sounds and time-stamp events. The user will then be notified of the recorded sounds in real time or offline and asked to identify the offensive sound by remembering the sound / situation at the time of the event or by listening to a sample. If the user identifies the sound as offensive, it will be added to the offensive sound library.
[0040] In one embodiment, physiological sensors can be used to detect unpleasant situations from the user's physiological signals (e.g., skin conductance, heart rate, EEG, ECG, EMG, or other signals). These signals (independent or fused) can be used to identify the occurrence of unpleasant sounds / situations, and the methods described above can be used to identify unpleasant sound components in mixed signals. When such an unpleasant sound is detected, it can be added to an unpleasant sound library, and the system can take recommended actions to reduce / filter / mask / block such unpleasant sounds. As described above, the user will be notified and he / she can manually overrule the recommended action. The unpleasant component of the sound can be detected and reported to the user for verification, after which it will be added to the library and / or the recommended action will be implemented.
[0041] Although the specific elements, embodiments and applications of the present disclosure have been shown and described, it should be understood that the scope of the present disclosure is not limited thereto, as those skilled in the art can make modifications without departing from the scope of the present disclosure, especially in view of the foregoing teachings. Therefore, for example, in any method or process disclosed herein, the actions or operations constituting the method / process can be performed in any suitable order and are not necessarily limited to any specific disclosed order. In various embodiments, elements and components can be configured or arranged, combined and / or eliminated in different ways. The various features and processes mentioned above can be used independently of each other or can be combined in various ways. All possible combinations and sub-combinations are intended to fall within the scope of the present disclosure. References to "some embodiments", "one embodiment" etc. throughout this disclosure mean that the specific features, structures, steps, processes or characteristics described in conjunction with the embodiments are included in at least one embodiment. Therefore, the appearance of phrases "in some embodiments", "in one embodiment" etc. throughout this disclosure do not necessarily refer to the same embodiment and can refer to one or more of the same or different embodiments.
[0042] Various aspects and advantages of the embodiments have been described where appropriate. It will be appreciated that not all of these aspects or advantages may be achieved in accordance with any particular embodiment. Thus, for example, it will be appreciated that various embodiments may be performed in a manner that achieves or optimizes one advantage or group of advantages as taught herein without necessarily achieving other aspects or advantages as taught or suggested herein.
[0043] Conditional language used herein, such as "can," "might," "would," "may," "for example," etc., unless otherwise expressly stated or understood otherwise in the context in which it is used, is generally intended to convey that certain embodiments include certain features, elements, and / or steps, while other embodiments do not include certain features, elements, and / or steps. Therefore, such conditional language is generally not intended to imply that features, elements, and / or steps are in any way necessary for one or more embodiments, or that one or more embodiments must include a method for deciding whether these features, elements, and / or steps are included or will be performed in any particular embodiment with or without operator input or prompting. No single feature or group of features is required or indispensable for any particular embodiment. The terms "include," "comprise," "have," etc. are synonymous and are used inclusively in an open-ended manner and do not exclude other elements, features, actions, operations, etc. In addition, the term "or" is used in its inclusiveness (rather than exclusivity), for example, when used to connect a list of elements, the term "or" means one, some, or all of the elements in the list.
[0044] The example results and parameters of the embodiments described herein are intended to illustrate, not to limit, the disclosed embodiments.Other embodiments may be configured and / or operate differently than the illustrative examples described herein.
Claims
1. A system for detecting, identifying, and managing environmental objectionable sounds, the system comprising: - a headphone device comprising a microphone and a speaker, the microphone being configured to capture ambient sound surrounding the user, the microphone being configured to capture samples of small segments of the ambient sound; - an interface having a hardware processor programmed with executable instructions for obtaining input data, transmitting data, and providing output data; a memory unit storing: input information, a library of offensive environmental sound signals, offensive sound feature maps, a recognition prediction model, offensive sound recognition categories, and an offensive sound suppression prediction model; as well as a processing unit in communication with the headphone device, the memory unit, and the interface, and comprising: a recognition unit coupled to the headphone device and the memory unit and programmed with executable instructions to recognize an offensive ambient sound signal in a sound segment by extracting at least one feature of the sound in the ambient sound segment and creating a feature map of such a sound segment, processing the feature map of the sound segment using the recognition prediction model stored in the memory unit to compare at least one feature in the feature map with the offensive sound signal feature map in the memory unit, wherein, when an offensive ambient sound is recognized, the recognition unit classifies the offensive sound signal using a recognition category; and o A filtering processing unit coupled to the recognition unit and the memory unit and programmed with executable instructions to receive a mixed signal of an ambient sound segment and an offensive sound signal, and to process the mixed signal, calculate an amplitude and a phase of the mixed signal to generate a feature map of the mixed signal, compare the feature map with stored feature maps using at least one offensive sound suppression model, and provide a recommended action to manage such offensive sound signal.
2. The system according to claim 1, wherein: The headphone device further comprises a valve for suppressing or isolating the ambient sound segment transmitted to the user.
3. The system according to claim 1, wherein: The filtering processing unit is also programmed with executable instructions for: automatically removing the identified offensive signal from the mixed signal using the generated feature map to obtain a clean sound in the mixed signal reconstructed from the frequency domain to the time domain; and combining the phase of the mixed signal with the amplitude of the clean sound of the final segment to create a clean sound signal for transmission to the speaker.
4. The system according to claim 3, wherein: The filtering processing unit is further programmed with executable instructions, and the instructions are used to: post-process the clean sound signal and generate a smooth clean sound signal.
5. The system according to claim 3, wherein: The filtering processing unit further includes a bypass having a gain for creating a reduced objectionable sound, the filtering processing unit being programmed with executable instructions for automatically adding the reduced objectionable sound signal to the clean sound signal.
6. The system according to claim 3, wherein: The recommended action is to remove the identified objectionable ambient sound signal.
7. The system according to claim 5, wherein: The recommended action is to attenuate the identified objectionable ambient sound signal.
8. The system according to claim 1, wherein: The memory unit also stores a static unpleasant sound suppression prediction model, a non-static unpleasant sound suppression prediction model, and a highly dynamic unpleasant sound suppression prediction model, and the filtering processing unit is programmed to access one of the static unpleasant sound suppression prediction model, the non-static unpleasant sound suppression prediction model, or the highly dynamic unpleasant sound suppression prediction model according to the category of the identified unpleasant environmental sound signal.
9. The system according to claim 1, wherein: The processing unit is programmed to record the newly identified offensive environmental sound signal into a library of offensive sound signals.
10. The system according to claim 9, wherein: The library of offensive sound signals includes offensive sound signals recognized by a user.
11. The system according to claim 1 further includes an alarm system, which communicates with the interface and / or the headphone device to generate an alarm signal to alert the user of the recommended action, and the alarm signal is selected from a visual signal, a tactile signal, a sound signal, or any combination thereof.
12. The system according to any one of claims 3 to 7, wherein: The headphone device further comprises a valve for suppressing or isolating the ambient sound segment transmitted to the user, and the system further comprises a first activation device in communication with the headphone device for manually triggering the valve to suppress or attenuate the objectionable ambient sound signal.
13. The system of claim 1, wherein: The memory unit also stores pre-recorded sounds.
14. The system according to claim 13, wherein: The filtering processing unit is programmed with executable instructions for automatically adding the pre-recorded sound to the mixed signal to mask the objectionable ambient sound signal.
15. The system of claim 13 , further comprising a second activation device in communication with the headphone device and the memory unit, the user using the second activation device to access the stored pre-recorded sound to play the pre-recorded sound over the mixed signal to mask the objectionable ambient sound signal.
16. The system of claim 1, wherein: The memory unit is embedded in the headphone device or located remote from the headphone device, and the memory unit communicates with the headphone device and the processing unit by wire, wirelessly, or using an Internet network.
17. The system of claim 1, wherein: The processing unit is embedded in the headphone device or located remote from the headphone device, and the processing unit communicates with the headphone device and the memory unit through a wired, wireless, or using an Internet network.
18. The system of claim 1, wherein: The interface is located remote from the headphone device, and the interface communicates with the headphone device and the processing unit through a wired, wireless or internet network.
19. The system of claim 1 , further comprising at least one physiological sensor in communication with the processing unit, the at least one physiological sensor being configured to detect at least one physiological parameter of the user, the processing unit identifying the offensive ambient sound signal in the ambient sound segment if the detected parameter is outside a predetermined range of at least one of the detected physiological parameters.
20. The system of claim 19, wherein: The identified offensive environmental sound signals are recorded in a library of offensive sound signals.
21. A method for environmental objectionable sound detection, identification, and management, the method comprising: - capturing ambient sound around the user using a microphone in the headset device, the microphone being configured to capture small samples of the ambient sound; - storing input information, a library of offensive environmental sound signals, offensive sound feature maps, recognition prediction models, offensive sound recognition categories, and offensive sound suppression prediction models on a memory unit; as well as - Processing the captured sound clips by a processing unit, the processing steps comprising: extracting at least one feature of a sound signal in the sound clip, creating a feature map of the sound clip, comparing at least one feature in the feature map with a feature map of offensive sound signals in the recognition prediction model stored in the memory unit, identifying offensive sound signals in the captured sound clip using the recognition prediction model, and classifying the identified offensive sound signals using a recognition category; and o Filtering a mixed signal comprising an ambient sound segment and the unpleasant ambient sound signal, calculating the amplitude and phase of the mixed signal to generate a feature map of the mixed signal, comparing the feature map with stored feature maps using at least one unpleasant sound suppression model, and providing recommended actions to manage the unpleasant sound signal.
22. The method according to claim 21, further comprising: Input data is obtained from the user, data is transmitted, and output data is provided using an interface having a hardware processor programmed with executable instructions.
23. The method according to claim 21, wherein The filtering process also includes: removing the identified offensive signal from the mixed signal to obtain a clean sound in the mixed signal, reconstructing the clean sound from the frequency domain to the time domain; combining the phase of the mixed signal with the amplitude of the clean sound to create a clean sound signal and transmitting the clean sound signal to a speaker.
24. The method according to claim 23, wherein The filtering process further includes: performing post-processing on the clean sound signal to generate a smooth clean sound signal.
25. The method according to claim 23, wherein The filtering process also includes creating a reduced offensive sound using a bypass having a gain, and adding the reduced offensive sound signal to the clean sound signal.
26. The method of claim 21, further comprising: The newly identified offensive environmental sound signal is recorded into a library of offensive sound signals.
27. The method of claim 21, further comprising: Pre-recorded sounds are stored on the memory unit.
28. The method according to claim 27, further comprising: The pre-recorded sound is played to mask the objectionable ambient sound signal.
Citation Information
Patent Citations
Noise reduction in multi-microphone systems
EP3096318A1
Noise suppression system and method
EP3175456B1
Advanced communication earpiece device and method
US20140010378A1
Advanced communication earpiece device and method
US20160126914A1
Spectrum matching in noise masking systems
US20200312294A1