Predicting and triggering a future response to a predicted background noise based on a sound sequence

CN116802732BActive Publication Date: 2026-09-25TELEFONAKTIEBOLAGET LM ERICSSON (PUBL)
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202180091908.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-03-18
Publication Date
2026-09-25
Estimated Expiration
2041-03-18

AI Technical Summary

Technical Problem

因此几乎不可能过滤掉噪声——如果语音和噪声重叠,则算法无法区分二者

Benefits of technology

[0012]本文公开的设备的这些和另外的操作可以提供许多潜在的优点。本公开的潜在优点包括:在涉及在线会议时对不想要的声音更快做出响应,因为这些操作预测具有定义的干扰特性的后续声音出现的概率,然后可以通过触发补救动作规则对其进行响应以在后续声音出现之前使后续声音静音或抑制后续声音。在具有定义的干扰特性的后续声音出现之前执行这种补救动作可以避免后续声音的任何部分被发送给在线会议中的远程参与者。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116802732B_ABST
    Figure CN116802732B_ABST
Patent Text Reader

Abstract

A device performs operations including identifying occurrence of a trigger sound in at least one microphone signal. The operations also include predicting a probability of occurrence of a subsequent sound having defined disturbance characteristics in the at least one microphone signal after the occurrence of the trigger sound. The operations also include triggering a remedial action to be performed to mute the at least one microphone signal or suppress the subsequent sound in the at least one microphone signal when the probability of occurrence of the subsequent sound having the defined disturbance characteristics satisfies a remedial action rule.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to electronic devices, for example, that actively eliminate or reduce noise present in microphone signals during online meetings. Background Technology

[0002] Several solutions exist for hearing protection, noise cancellation, and handling unwanted sounds associated with online meetings.

[0003] One example is noise-canceling headphones, which suppress or block out external noise and allow the wearer to focus on a favorite song or ongoing conversation. A technology called Active Noise Control (ANC) works by using microphones to pick up (low-frequency) noise and neutralize it before it reaches the ear. Also known as noise cancellation or Active Noise Reduction (ANR), ANC is a method of reducing unwanted sounds by adding a second sound specifically designed to cancel out the first. The headphones generate a sound signal that is 180 degrees out of phase with the unwanted noise, resulting in two sounds canceling each other out.

[0004] Another example is a hearing protection device (HPD), which reduces the amount of sound reaching the eardrum through a combination of electronic devices and structural components. An HPD is an ear protection device worn in or on the ear when exposed to harmful noise to help prevent noise-induced hearing loss. HPDs reduce (rather than clear) the level of noise entering the ear. HPDs can also prevent other effects of noise exposure, such as tinnitus and hyperacusis. Many different types of HPDs exist for use, including earmuffs, earplugs, electronic hearing protection devices, and semi-inserted devices. Some electronic HPDs, known as hearing enhancement protection systems, provide hearing protection against high-level sounds while allowing other sounds, such as speech, to be transmitted. Some electronic HPDs also have the ability to amplify low-level sounds. This type can be beneficial for users in noisy environments who still need access to lower-level sounds. For example, hunters who rely on detecting and locating the faint sounds of wild animals still want to protect their hearing from gunshot wounds.

[0005] Microsoft has demonstrated real-time noise suppression using artificial intelligence (AI) to detect and suppress distracting background noise during calls. Real-time noise suppression filters out the sounds of someone typing on their keyboard, the rustling of a bag of chips, and a vacuum cleaner running in the background. AI removes background noise in real time, so you only hear the voice on the other end of the call.

[0006] Noise suppression has been present in Microsoft Teams, Skype, and Skype for Business applications for years. Other communication tools and video conferencing applications also have some form of noise suppression. However, this noise suppression covers static noise, such as computer fans or air conditioners running in the background. Traditional noise suppression methods look for pauses in speech, estimate a baseline of noise, assume continuous background noise does not change over time, and filter it out.

[0007] Isolating human speech from unwanted background noise is no easy task, as they can overlap at the same frequencies. In the spectrogram of a speech signal, unwanted noise appears in the gaps between speech words and overlaps with them. Therefore, it's nearly impossible to filter out noise—if speech and noise overlap, the algorithm cannot distinguish between them. Instead, the algorithm may need to be trained beforehand to understand what noise looks like and therefore what speech looks like. Microsoft trained a machine learning model to understand the difference between noise and speech, and then the machine learning model attempted to suppress noise during inference while keeping the speech unaffected.

[0008] Machine learning comprises computer algorithms that automatically improve through experience. It is considered a part of artificial intelligence. Machine learning algorithms build models based on sample data called "training data" to make predictions or decisions without being explicitly programmed to do so. Machine learning algorithms are used in a variety of applications such as email filtering and computer vision, where it is difficult or impossible to develop conventional algorithms to perform the required tasks. Summary of the Invention

[0009] Some embodiments disclosed herein relate to an apparatus comprising: at least one processor configured to receive at least one microphone signal from at least one microphone; and at least one memory storing program code executable by the at least one processor. Operations performed by the at least one processor include: identifying the occurrence of a trigger sound in the at least one microphone signal. Operations further include: predicting, after the occurrence of the trigger sound, the probability of a subsequent sound having defined interference characteristics occurring in the at least one microphone signal. Operations further include: triggering a remedial action to be performed to mute the at least one microphone signal or suppress the subsequent sound in the at least one microphone signal when the probability of the subsequent sound having defined interference characteristics satisfies a remedial action rule.

[0010] Some embodiments relate to a method performed by a device, the method comprising: identifying the occurrence of a trigger sound in at least one microphone signal received from at least one microphone. The method further comprises: predicting, after the occurrence of the trigger sound, the probability of a subsequent sound having defined interference characteristics occurring in the at least one microphone signal. The method further comprises: triggering a remedial action to be performed to mute or suppress the subsequent sound in the at least one microphone signal when the probability of the subsequent sound having defined interference characteristics satisfies a remedial action rule.

[0011] Some embodiments relate to a computer program product including a non-transitory computer-readable medium storing program code executable by at least one processor of a device to perform operations. The operations include: identifying the occurrence of a trigger sound in at least one microphone signal received from at least one microphone. The operations also include: predicting, after the occurrence of the trigger sound, the probability of a subsequent sound having defined interference characteristics occurring in the at least one microphone signal. The operations further include: triggering a remedial action to be performed to mute the at least one microphone signal or suppress the subsequent sound in the at least one microphone signal when the probability of the subsequent sound having defined interference characteristics satisfies a remedial action rule.

[0012] These and other operations of the device disclosed herein can provide numerous potential advantages. These potential advantages include: a faster response to unwanted sounds in online meetings, because these operations predict the probability of subsequent sounds with defined interference characteristics occurring, and can then be responded to by triggering remedial action rules to mute or suppress the subsequent sounds before they occur. Performing such remedial action before subsequent sounds with defined interference characteristics occur can prevent any part of the subsequent sound from being sent to remote participants in the online meeting.

[0013] Other devices, methods, and computer program products according to the embodiments will be apparent to or will become apparent to those skilled in the art after reviewing the following figures and detailed description. All such devices, methods, and computer program products are intended to be included in this specification, within the scope of this disclosure, and protected by the appended claims. Furthermore, it is intended that all embodiments disclosed herein can be practiced individually or in any manner and / or in combination. Attached Figure Description

[0014] Various aspects of this disclosure are illustrated by way of example and are not limited to the accompanying drawings. In the drawings:

[0015] Figure 1 Examples of sound sequences that can be processed by a device to suppress sounds with defined interference characteristics, according to some embodiments of the present disclosure, are shown;

[0016] Figure 2 A system diagram is shown that includes components configured to operate according to some embodiments of the present disclosure;

[0017] Figures 3 to 8 Flowcharts illustrating operations performed by the device according to some embodiments of the present disclosure are shown;

[0018] Figure 9 A computing system for controlling sound playback via a user device, according to some embodiments, is illustrated;

[0019] Figure 10 The component circuitry of another computing system configured according to some embodiments is shown;

[0020] Figure 11 This is a block diagram of component circuitry of a computing server configured to operate according to some embodiments of the present disclosure; and

[0021] Figure 12 This is a block diagram of component circuitry that may include the functionality of an adaptive music system or be communicatively connected to a user equipment according to some embodiments of this disclosure;

[0022] Figure 13 A Markov chain for sound sequences is shown according to some embodiments of this disclosure; and

[0023] Figure 14 The conditional probabilities of subsequent sounds according to some embodiments of this disclosure are shown. Detailed Implementation

[0024] In the following description, the inventive concept will be more fully described with reference to the accompanying drawings, in which examples of embodiments of the inventive concept are illustrated. However, the inventive concept can be embodied in many different forms and should not be construed as being limited to the embodiments set forth herein. Rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the scope of the various inventive concepts to those skilled in the art. It should also be noted that these embodiments are not mutually exclusive. Components from one embodiment may be assumed by default to be present / used in another embodiment.

[0025] Various embodiments of this disclosure describe an apparatus and method that uses artificial intelligence (AI) or other machine learning to suppress or mute unwanted sounds by analyzing sound sequences and their temporal and spatial correlations.

[0026] Existing active noise cancellation techniques are insufficient to suppress background noise being sent to participants in online meetings.

[0027] Hearing protection devices (HPDs) work by attempting to reduce the amount of sound reaching the eardrums of a local listener in the event of a loud sound, such as a gunshot, by trying to quickly suppress noise with high transients. HPD technology is also insufficient to suppress background noise being sent to participants in online meetings.

[0028] Microsoft has demonstrated real-time noise suppression using artificial intelligence to detect and suppress distracting background noise during calls. However, a challenging issue is isolating human speech, as other noises also occur simultaneously. One alternative approach is to train a machine learning model to understand the difference between noise and speech.

[0029] Furthermore, suppressing sudden, high-transient sounds is challenging in terms of both detection and appropriate action. For example, a dog barking or a door slamming sound has a short, high-peak transient that requires immediate suppression or silencing. If silencing or suppression is not fast enough, at least some of the noise will get through. Typically, machine learning models require time to detect or classify sounds, which can lead to a noise remediation response that is too late, and some noise being sent to other devices.

[0030] Figure 1 Examples of sound sequences that can be processed by a device to suppress sounds with defined interfering characteristics according to some embodiments of the present disclosure are shown. The example sound sequences include at least some of the following: a detectable sound 100, such as footsteps on a front porch, that does not trigger a device response; a triggering sound 102, such as a knocking sound; a predicted noise 104, such as a dog barking; and someone shouting “Quiet” 106 after the dog barking. The sound sequence can have a high probability of occurrence after the initial triggering sound occurs; in the illustrated example, the initial triggering sound is footsteps on a front porch 100. For example, a package delivery person walking on the front porch approaching a door could generate sound 100, followed by a knocking sound 102, which would trigger a dog barking 104, and then someone shouting “Quiet” 106.

[0031] Various embodiments of this disclosure relate to analyzing sound sequences (e.g., Figure 1 The system uses the sound sequences shown and their individual temporal and spatial relationships to predict the probability of a subsequent sound or a subsequent sound sequence occurring after the current triggering sound, in order to use AI to suppress unwanted sounds or mute unwanted sounds.

[0032] Figure 2 A system diagram is shown that includes components configured to operate according to some embodiments of the present disclosure. In the system shown, the microphone 202 of the host user device 200 (“device”) detects a sequence of sounds, such as at least some of sounds 100, 102, 104, and 106.

[0033] The host user device 200 may be, for example, a laptop computer, tablet computer, smartphone, extended reality headset, etc. The host user device 200 includes at least one processor configured to receive at least one microphone signal from at least one microphone 202 and may be configured to provide audio signals to speakers. Microphones 202 and any speakers may be physically or wirelessly (e.g., via Bluetooth or Wi-Fi headsets) connected to the host user device 200. Multiple microphones or microphone arrays and / or speakers may be physically or wirelessly connected to the host user device 200.

[0034] The host user device 200 can be communicatively coupled to the virtual conference server 210. The virtual conference server 210 may include a predictive audio correction component 212. The predictive audio correction component 212 may alternatively be located in the host user device 200. The predictive audio correction component 212 can be trained using machine learning algorithms. The virtual conference server 210 is configured to provide audio streams from the host user device 200 to the participant user devices 220 and 222 via wired and / or wireless network connections.

[0035] According to various embodiments disclosed herein, a system via predictive sound remedy component 212 can be configured to predict the probability of subsequent sounds with defined interference characteristics (e.g., knocking 102 and barking 104) occurring in a microphone signal after a triggering sound (e.g., footsteps 100 on the front porch) occurs. When the probability of subsequent sounds with defined interference characteristics occurs satisfies a remedy action rule, the system can trigger a remedy action to be performed to mute the microphone signal or suppress subsequent sounds in the microphone signal.

[0036] Figure 3 Flowcharts illustrating operations performed by devices such as host user device 200 and / or virtual meeting server 210 according to some embodiments of this disclosure are shown. For convenience, Figure 3 The operations are described in the context of being performed by the virtual meeting server 210, although they may be performed additionally or alternatively by the host user device 200 and / or by another component of the system.

[0037] In some embodiments, the virtual meeting server 210 is configured to identify the occurrence of a trigger sound (e.g., "footsteps on the front porch" 100 and / or "knocking on the door" 102) in at least one microphone signal 300. The virtual meeting server 210 is also configured to, after the occurrence of the trigger sound (e.g., "footsteps on the front porch" 100 and / or "knocking on the door" 102), predict 302 the probability of a subsequent sound with defined interference characteristics (e.g., "barking" 104) occurring in at least one microphone signal. The device is further configured to, when the probability of the subsequent sound with defined interference characteristics (e.g., "barking" 104 and "quiet" 106) occurring satisfies a remedial action rule, trigger a remedial action to be performed by 304 to mute at least one microphone signal or suppress the subsequent sound (e.g., "barking" 104) and the shout "quiet" 106 in at least one microphone signal.

[0038] A potential advantage of some embodiments of this disclosure is that they provide a rapid response to triggering sounds and can initiate remedial actions before subsequent sounds with defined interference characteristics (e.g., “dog barking” 104 and “silence” 106) occur. During online meetings, these embodiments can prevent any portion of the subsequent sound in the microphone signal of the host user device 200 from being transmitted in the audio stream to the participant user devices 220 and 222.

[0039] In some embodiments, predicting the probability of a subsequent sound with defined interference characteristics appearing in at least one microphone signal includes predicting the probability of a subsequent sound meeting a defined interference level.

[0040] In some embodiments, predicting the probability of a subsequent sound occurring that satisfies a defined interference level includes determining the probability that at least one of the following conditions is met: the predicted peak decibel level of the subsequent sound exceeds a peak threshold; the predicted duration of the subsequent sound exceeds a duration threshold; the predicted frequency component of the subsequent sound is within a defined frequency band; and the subsequent sound has a predicted sound category that has been defined as unacceptable.

[0041] A machine learning model can be trained to detect a specific sound in a sequence of detected ambient sounds, i.e., a trigger sound. The trigger sound is then used to predict the next sound in the sound sequence and to predict the probability X of the next sound occurring with defined interference characteristics. If the next sound in the sound sequence is predicted using the probability X and the interference characteristics defined by Y, then: if X and Y are less than a threshold (e.g., a threshold rule is not met), the operation predicts the probability of the next sound occurring in the sound sequence after that; and if X and Y are greater than or equal to the threshold (e.g., a threshold rule is met), the operation trigger can include actions such as muting the microphone signal or performing sound suppression.

[0042] Various embodiments can be coupled to online conferencing applications (such as Microsoft Teams or Zoom) and configured to provide the online conferencing application with an indication that a predicted unwanted sound will soon occur (e.g., within milliseconds, seconds, or minutes), triggering the online conferencing application to mute or suppress the sound before it occurs. For example, a countdown signal can be provided to the online conferencing application indicating the predicted amount of time remaining before the expected sound interference to be mute or suppressed occurs. The online conferencing application can be indicated with respect to which of a plurality of microphone signals should be mute or suppressed and the duration of the expected sound interference, allowing the online conferencing application to control the duration of the microphone signal mute or sound suppression accordingly. The triggered action can be to mute all microphones, a selected subset of microphones, a specific microphone, or a hardware port associated with a microphone. The triggered action can also, or alternatively, be to adjust a detection threshold or other probabilistic parameters of the algorithm for detecting and / or classifying sound.

[0043] In some embodiments, determining when the probability of a subsequent sound with defined interference characteristics meeting a remedial action rule includes determining whether at least one of the following contextual parameters meets the remedial action rule: device location data when the sound is triggered; time data when the sound is triggered; date data when the sound is triggered; characteristics of a background noise component in at least one microphone signal; user input data indicating whether a remedial action should be triggered; an indication that a defined sound source type has been identified by a camera; and user demographic characteristics. An example indication that a defined sound source type has been identified by a camera is an indication that a security camera has indicated the presence of a dog.

[0044] The sensitivity of operations used to predict the probability of subsequent sounds, define what constitutes interference characteristics, and / or define remedial action rules can be adjusted based on context parameters. Example context parameters can be set to indicate whether the user is at work or at home. Context parameters can be defined to adapt to certain predefined contexts, such as "home alone," "home with family," or "home with pets." Context parameters can indicate the user's surrounding environment, such as being in an outdoor park, in a first aid station such as a fire station, in a car, on a train, or on an airplane. Context parameters can indicate the user's location, time, and date. Context parameters can indicate how the device is being used, such as for work or personal use.

[0045] The sensitivity of operations used to predict the probability of subsequent sounds, define what constitutes a defined interference characteristic, and / or define what constitutes a remedial action rule can be increased or decreased between defined thresholds or threshold ranges (high-medium-low). For example, the sound level associated with a "high" interference level in the context of "at work" can be associated with a "low" interference level in the context of being home alone during the weekend.

[0046] Another aspect of some embodiments is that other sensors can be used to provide input about contextual parameters to the operation (e.g., to a machine learning model). For example, a video camera can generate contextual parameters based on recognizing the presence of a dog, a microphone can generate contextual parameters based on hearing relatively high background noise in a work environment, contextual parameters can be generated based on determining that a user is logged into a home network, and contextual parameters can be defined based on the time of day to indicate that a user may be in a certain environment, etc.

[0047] In another embodiment, a machine learning model and / or a portion thereof may be trained centrally based on a general sound sequence and / or a sound sequence collected from a demographic group, and then the trained model is provided to the predictive sound remediation component 212.

[0048] Figure 4 and Figure 5 A flowchart illustrating operations performed by a device (e.g., host user device 200 and / or virtual meeting server 210) according to some embodiments of the present disclosure is shown.

[0049] First refer to Figure 4 The operation that triggers 304 to perform a remedial action to mute at least one microphone signal or suppress subsequent sound in at least one microphone signal includes: predicting 400 the duration of the subsequent sound. Then, the operation, within a duration determined based on the predicted duration of the subsequent sound, mutes 402 at least one microphone signal or suppresses 402 subsequent sound in at least one microphone signal.

[0050] Now for reference Figure 5 The operation of triggering the remedial action to be performed at 304 to mute at least one microphone signal or suppress subsequent sound in at least one microphone signal includes: predicting the time delay between the occurrence of the trigger sound at 500 and the start of the subsequent sound. Then, after the trigger sound occurs, the operation triggers the start of the remedial action at 502 based on the expiration of the time delay.

[0051] Automatic muting can be performed by the host's user device 200 participating in the online video conference. A predictive sound remediation component 212 can be part of the online conferencing application and can use a trained machine learning model executed locally and / or in a virtual conferencing server 210. The machine learning model can be configured to identify a large number of sound sequences related to different contexts and how the sounds are correlated with each other (time-dependent occurrence).

[0052] In one example, a user of the host user device 200 participates in an online meeting from home, and the context parameters can be defined to indicate the presence of a wife, children, and a barking dog at home. In this scenario, a person approaching the front door makes footsteps on the front porch 100, then knocks on the door 102, which triggers the dog barking 104, followed by the wife shouting "Quiet!" 106. This scenario is highly disruptive to both the user and other participants in the online meeting.

[0053] Various embodiments of this disclosure address the problem through operations that can be performed using machine learning models. These operations identify triggering sounds (e.g., footsteps 100 on the front porch) and predict the probability of subsequent sound sequences occurring, which can trigger remedial actions. The machine learning model detects the occurrence of a triggering sound (e.g., footsteps 100 on the front porch) and, after the triggering sound occurs, predicts the probability of subsequent sounds (e.g., knocking 102) occurring with defined interference characteristics. In one example, the probability of subsequent sounds occurring is determined to be "low".

[0054] The machine learning model detects subsequent sounds in the sequence (e.g., knocking on the door 102), determines the subsequent sounds in the sequence with a certain probability, and predicts that the probability of the next subsequent sound "dog barking" 104 and a person shouting "quiet" 106 is "high".

[0055] Because the probability of the next following sound, "dog barking" 104, and the person shouting "quiet" 106, which have defined interference characteristics, occurring is "high," the probability of their occurrence satisfies the remedial action rule. This triggers the remedial action to be performed to mute at least one microphone signal or suppress the next following sound in at least one microphone signal. The remedial action can be performed by the predictive sound remediation component 212 and / or by notifying the online conferencing application of the reason for silencing the "sudden high background noise."

[0056] This might start a timer on the first device / application, which will restore the silence to non-silence when it expires. This timer can be set based on the predicted duration of subsequent sounds.

[0057] Figure 9 A computing system for controlling sound playback via a user device, according to some embodiments, is shown.

[0058] refer to Figure 9 The predictive sound correction component includes at least one processing circuitry 912 as part of the predictive sound correction component 910. The predictive sound correction component 910 may be located on or communicatively coupled to the device 900. For ease of illustration of the various functional operations of the processing circuitry 912, in... Figure 9 In one embodiment, the processing circuit 912 is shown to include an analysis circuit 920, a machine learning processing circuit 930, and a remedial action circuit 940. The processing circuit 912 may have a higher... Figure 9 The circuitry shown may include more or fewer circuits. For example, as further explained below, any one or more of the analysis circuitry 920, machine learning processing circuitry 930, and remedial action circuitry 940 may be combined as an integrated circuit or divided into two or more separate circuits. User equipment 900 may be configured to receive microphone signals that may be provided by microphone circuitry within user equipment 900 or connected thereto via a wired or wireless connection. For example, a headset may include a microphone configured to provide digitized microphone signals to user equipment 900.

[0059] Although only for ease of illustration and explanation, the analysis circuit 920, the machine learning processing circuit 930, and the remedial action circuit 940 are described herein. Figure 9 And various other circuits are shown as separate blocks in the diagrams, but any two or more of these circuits may be implemented in a shared circuit, and any one of these circuits may be implemented at least partially in a digital circuit, for example by program code stored in at least one memory circuit, which is executed by at least one processor circuit 912.

[0060] Figure 10 Component circuitry of another computing system configured according to some embodiments is shown. Although the predictive sound remediation component 910 is shown as separate from and communicatively connected to the user equipment 1002 and the database 1000 of pre-recorded sound sequences of various types shown via a network 1010, some or all of the circuitry components of the predictive sound remediation component 910 (e.g., analysis circuitry 920, remediation action circuitry 940, machine learning processing circuitry 930, training circuitry 1042, etc.) may be implemented by circuitry implemented in any one or more user equipment 1002 and / or in the database 1000.

[0061] refer to Figure 10 The training circuit 1042 is configured to train the machine learning model 932 based on a combination of many parameters discussed in various embodiments herein.

[0062] Analysis circuit 920 is configured to analyze inputs from user equipment 1002 and / or database 1000 containing pre-recorded sound sequences for use in training machine learning processing circuit 930.

[0063] Analysis circuitry 920 can characterize sound sensed by the microphone of user equipment 1002 and / or sound obtained from database 1000. For example, characterization may include at least one of characterizing the sound spectrum (e.g., zero-crossing rate, spectral centroid, spectral roll-off, overall shape of the spectral envelope, chromatic frequencies, etc.), sound acoustic fingerprint (based on a time-frequency graph of ambient noise, which may also be referred to as a spectrogram), sound loudness, and sound noise repetition patterns.

[0064] Zero-crossing rate can correspond to the rate of sign change along a signal, i.e., the speed at which a signal changes from positive to negative or from negative to positive. Spectral envelope can correspond to the location of the "centroid" of a sound and can be calculated as a weighted average of the frequencies present in the sound. Spectral roll-off can correspond to a shape measurement of a signal, such as the frequency representing a specified percentage of the total spectral energy below it. Overall shape can correspond to the Mel-frequency cepstral coefficients (MFCC) of a sound, which is a small set of features (typically around 10 to 20) that concisely describes the overall shape of the spectral envelope. Chromatic frequencies can correspond to a representation of a sound where the entire spectrum is divided into a defined number (e.g., 12) of intervals, representing a defined number (e.g., 12) of different semitones (or chromaticities) of the sound's spectral octaves.

[0065] Analysis circuit 920 can characterize the sound sequences that appear in the sounds sensed by the microphone of user equipment 1002 and / or the sounds obtained from database 1000.

[0066] The analysis circuit 920 can characterize remedial actions performed by the predictive sound remediation component 212 and / or remedial actions performed by the user in response to the occurrence of the characterized sound. For example, the analysis circuit 920 can characterize user actions in response to the occurrence of the characterized sound: muting the microphone, increasing the speaker volume, sensing the movement of the host user device 200, pausing audio playback, sensing the closing of a door, sensing the closing of a window, etc.

[0067] Analysis circuit 920 can predict the probability of subsequent sounds with defined interference characteristics based on the characterized sound and the characterized sound sequence.

[0068] The machine learning processing circuit 930 is configured to be trained to predict, after a triggered sound occurs, the probability of a subsequent sound with defined interference characteristics occurring in at least one microphone signal. When the probability of the subsequent sound with defined interference characteristics occurs satisfies a remedial action rule, a remedial action to be performed is triggered to mute or suppress the subsequent sound in at least one microphone signal. The machine learning processing circuit 930 may trigger a remedial action circuit 940 to perform a remedial action to mute or suppress the subsequent sound in at least one microphone signal. The remedial action circuit 940 may reside at least partially within each user equipment 1002.

[0069] The machine learning processing circuit 930 can operate in both run mode and training mode, although these modes are not mutually exclusive and at least some training can be performed during run.

[0070] During operation, the representation data output by the analysis circuit 920 can be adjusted by the data preconditioning circuit 1020 to, for example, normalize and / or filter the representation data values ​​before it reaches the machine learning processing circuit 930 via the running path 1040. The machine learning processing circuit 930 includes a machine learning model 932, which in some embodiments includes a neural network circuit 934. The representation data is processed by the machine learning model 932 to predict the probability of a subsequent sound with defined interference characteristics appearing in at least one microphone signal after the occurrence of a triggered sound, and when the probability of a subsequent sound with defined interference characteristics appears satisfies a remedial action rule, a remedial action to be performed is triggered to mute or suppress the subsequent sound in at least one microphone signal.

[0071] During training, training circuit 1042 adapts machine learning model 932 based on representation data from analysis circuit 920 to predict the probability of sound sequences with defined interference characteristics occurring, wherein the representation data can be adjusted by pre-conditioning circuit 1020. When machine learning model 932 includes neural network circuit 934, training may include adapting the weights of combined nodes in the neural network layers and / or adapting the firing thresholds used by the combined nodes of neural network circuit 934. Training circuit 1042 may train machine learning processing circuit 930 based on historical representation data values ​​obtainable from historical data store 1030. Historical data store 1030 may be populated over time with representation data values ​​output by analysis circuit 920.

[0072] Figure 11 This is a block diagram of the component circuitry of a computing server 910 configured to operate according to some embodiments of this disclosure. The computing server 910 may, for example, correspond to a virtual meeting server 210. Figure 2 See also Figure 11 The computing server 910 includes a wired / wireless network interface circuit 1120, at least one processing circuit 1100 (processing circuit), and at least one storage circuit 1110 (memory), which is also described below as a computer-readable medium. The processing circuit 1100 may correspond to... Figure 9 The processing circuitry 912 is included. Memory 1110 stores program code 1112, which is executed by processing circuitry 1100 to perform the operations disclosed herein for at least one embodiment of the computing server. Program code 1112 may include machine learning model code 932, configured to perform at least some of the operations described herein for machine learning. Processing circuitry 1100 may include one or more data processing circuits, such as general-purpose and / or special-purpose processors (e.g., microprocessors and / or digital signal processors), that may be co-located or distributed across one or more data networks. Computing server 910 may also include a display device 1150 and a user input interface 1160.

[0073] Figure 12 This is a block diagram of component circuitry for a user equipment 900 according to some embodiments of the present disclosure. This component circuitry may include the functionality of a predictive voice remediation component or may be communicatively connected to a computing server. User equipment 900 may, for example, correspond to host user equipment 200 (…). Figure 2 ) or user equipment 1002 ( Figure 10 User equipment 900 may include wireless network interface circuitry 1220, at least one processing circuitry 1200 (processing circuitry), and at least one storage circuitry 1210 (memory), the at least one storage circuitry 1210 being hereinafter also described as a computer-readable medium. The processing circuitry 1200 may correspond to... Figure 9The processing circuitry 912 is included. Memory 1210 stores program code 1212, which is executed by the processing circuitry 1200 to perform the operations disclosed herein for at least one embodiment of the user equipment. Program code 1212 may include machine learning model code 932, configured to perform at least some of the operations described herein for machine learning. Processing circuitry 1200 may include one or more data processing circuits, such as general-purpose and / or dedicated processors (e.g., microprocessors and / or digital signal processors), that may be co-located or distributed across one or more data networks. User equipment 900 may also include location determination circuitry 1270, microphone 1230, display device 1250, and user input interface 1260 (e.g., keyboard or touch-sensitive display). Location determination circuitry 1270 may operate to determine the geographic location of user equipment 900 based on satellite positioning (e.g., GNSS (Global Navigation Satellite System), GPS (Global Positioning System), GLONASS, BeiDou, or Galileo) and / or based on terrestrial network-assisted positioning (e.g., cell tower triangulation based on signaling time-of-flight or Wi-Fi-based positioning). User equipment 900 may include other sensors 1240, such as a camera.

[0074] Some embodiments of this disclosure include a machine learning model for detecting a specific sound, i.e., a trigger sound, in a sequence of detected ambient sounds. The machine learning model is trained based on sound sequences categorized according to: a first trigger sound 102 recorded by at least one microphone and the transient / level and duration of the sound; a next sound 104 in the sequence with timing relative to the first sound 102 and the transient / level and duration of the next sound; a subsequent sound 106 with timing data and the transient / level and duration of the subsequent sound; and contextual data.

[0075] The machine learning model is trained using input from the device's microphone. The training focuses on sound sequences, sound levels, spectral characteristics, and their temporal and spatial relationships with (user) context parameters.

[0076] The machine learning model will infer / predict the sequence of sounds following the triggering sound based on the triggering sound inference. The prediction will determine which sound(s) in the sound sequence the hardware or application should adapt to based on probability and the level of interference (e.g., high-medium-low related to its transients, decibel levels, duration, direction, current context parameters, etc.).

[0077] Figure 6 , Figure 7 and Figure 8A flowchart illustrating operations performed by a device (e.g., host user device 200 or virtual meeting server 210) according to some embodiments of the present disclosure is shown.

[0078] exist Figure 6 In one operational embodiment, the operation includes training a 600 machine learning algorithm to classify sounds received by the device and / or another device, and identifying probabilistic correlations between classified sounds appearing in a sequence. The operation also includes processing a 602 trigger sound using a machine learning algorithm to predict the probability of subsequent sounds with defined interference characteristics appearing in at least one microphone signal.

[0079] In some embodiments, the operation further includes selecting a machine learning algorithm to be trained from a set of machine learning algorithms based on at least one of the following contextual parameters: device location data when one of the sounds to be classified occurs; time data when one of the sounds to be classified occurs; date data when one of the sounds to be classified occurs; characteristics of the background noise component present when one of the sounds to be classified occurs; and sensor data indicating object or environmental parameters of the type sensed.

[0080] In some embodiments, training the machine learning algorithm 600 further includes training the machine learning algorithm based on user feedback indicating whether the sound to be classified has defined interference characteristics.

[0081] refer to Figure 7 The operational example, training the machine learning algorithm 600, further includes: selecting 700 pre-recorded sound sequences from a database of pre-recorded sound sequences, based on the fact that a set of pre-recorded sound sequences is received by a device that satisfies a similarity rule with the device. The operation also includes training a machine learning algorithm 702 based on this set of pre-recorded sound sequences.

[0082] Other embodiments involve repeating operations on a sound sequence, for example, it can be represented as follows: Figure 13 The Markov state model shown is described below.

[0083] refer to Figure 8 The operation example further includes processing 800 subsequent sounds using a machine learning algorithm to predict the next probability of a next subsequent sound with defined interference characteristics occurring in at least one microphone signal, wherein the next subsequent sound is received by the device after the subsequent sound is received. The operation also includes triggering a remedial action 802 to perform, such that at least one microphone signal is muted or the next subsequent sound in at least one microphone signal is suppressed, when the next probability of a next subsequent sound with defined interference characteristics occurs satisfies a remedial action rule.

[0084] In some embodiments, the operation further includes training a machine learning algorithm to indicate when the probability of a subsequent sound with defined interference characteristics occurs satisfies a remedial action rule.

[0085] Training a machine learning algorithm can include training it based on user feedback indicating when a remedial action is triggered by the user.

[0086] Training the machine learning algorithm may include training the algorithm based on user feedback indicating that the user has performed at least one of the following remedial actions: the user mutes at least one microphone; the user increases the speaker volume; the user removes the device from its position when subsequent sound occurs; and the device detects indications that the user has performed an action separate from the device's operation to suppress subsequent sound. This "detection" includes detecting when the user has closed doors, windows, etc.

[0087] In another embodiment, a machine learning model and / or portions thereof are trained centrally based on general sound sequences and / or sound sequences collected from demographic groups, and then the trained model is pushed to a user device for inference.

[0088] Various embodiments of this disclosure describe a way of training a model on a sequence of sounds and training how the sounds are correlated in time and order depending on different contexts. For example, device capabilities, a first user context (at home, at home with family, time of day, etc.), whether the user is in a meeting, etc. Another example includes the correlation between a first (triggered) sound and subsequent sounds after the triggered sound, where the subsequent sounds are associated with certain probabilities, cascading to third and fourth level sounds, each with a reasonable probability of T3% and F4%. For example, "doorbell" will cause "dog barking" with a 91% probability, and "dog barking" falls into the spectral mask of "Daisy, dog" with a 99% probability.

[0089] Another example involves the correlation between the level of interference and a particular sound in the sequence. This can be done using supervised learning, or a combination of unsupervised and supervised learning.

[0090] It is also possible to train sound sequences that are correlated with time (audible time) and frequency samples (e.g., 8 to 16 kHz). Furthermore, the timing between sounds in the sequence is important.

[0091] The establishment of a first sound (i.e., a triggering sound) to cause a sequence of subsequent sounds or actually to cause a second sound that can evolve into a subsequent sound or terminate can be considered as a Markov chain. Figure 13 A Markov chain representation of a sound sequence according to some embodiments of the present disclosure is shown. Figure 13 Includes example probabilities and matrices used to illustrate Markov chains.

[0092] refer to Figure 13 For simplicity, assume three sounds G, M, and A. The assumed Markov state can be represented by T, where Tij is the probability that one sound follows another. A fundamental property of Markov chains is that only the nearest point in the event path (called the trajectory) affects what happens next; this is often referred to as the Markov property. Let {X0, X1, X2, ...} be a sequence of discrete random variables. Then {X0, X1, X2, ...} is a Markov chain that satisfies the Markov property: for all t = 1, 2, 3, ..., and for all states s0, s1, ..., st, s, P(Xt+1 = s | Xt = st, ..., X0 = s0) = P(Xt+1 = s | Xt = st).

[0093] Then, in, for example, a three-state model, assuming that a triggering sound G is detected, the probability of sound G reappearing is 80%, and the probability of sound M appearing is 20%; if then in the “next time instance” we appear at sound M, the probability of sound G appearing next is 80%, and the probability of us returning to sound G is 20%; if then in the “next time instance” we somehow end up at sound A, the probability of another sound A following is 90%, the probability of sound G reappearing is 10%, and the probability of sound M following sound A is zero (i.e., it has not been detected by the machine learning model before).

[0094] Then operations can determine the probability of any particular path; and given the Markov property that only the nearest point in the event path affects what happens next, these operations can calculate the probability of any trajectory by multiplying the starting probability by the probabilities of all subsequent single steps. For example, the calculation of P(X2=5|X0=1) means considering the transition from state 1 at time 0 to state 5 at time 2.

[0095] One approach could include using machine learning models to detect and classify sound sequences, thereby determining the probability factors of a Markov state transition matrix.

[0096] In addition, one approach may include a user context that can be described by different state transition matrices; a “working matrix” and, for example, an “idle time” matrix; alternatively, a full transition state matrix representing all possible states (given the physical existence of the object), but some state-state transitions may be prohibited (i.e., considered non-causal or non-physical), such as “the doorbell generates a dog barking, but the dog is not home”, etc.

[0097] Furthermore, it is assumed that the machine learning model trains (adjusts) the state transition entries in the context of the current object in the user's context.

[0098] Based on the known principle of conditional probability, that is, given a doorbell with a known initial state, according to... Figure 14 What is the conditional probability of triggering a disruptive barking? Figure 14 The conditional probabilities of subsequent sounds following a first sound according to some embodiments of this disclosure are shown.

[0099] Relaxing the requirement of only the previous state, other carefully designed calculation schemes can also be considered.

[0100] The above methods, or conditional probability methods, can also be considered in the context of deep learning networks, decision trees, Bayesian networks, etc., to adjust the transition coefficients between events (states) using machine learning models.

[0101] Federation learning (“FL”) systems can be used in various embodiments of this disclosure. In an FL system, a central server (referred to as the principal or master entity) is responsible for maintaining a global model, which is created by aggregating models / weights that are trained using local data during the iteration process of participating nodes / clients (referred to as workers or worker entities).

[0102] FL depends on workers continuously participating in the iterative process of training the model and transmitting model weights to the principal. The principal can communicate with varying numbers of workers (ranging from tens to millions), and the size of the transmitted model weight updates ranges from kilobytes to tens of megabytes.

[0103] Joint learning (FL) is a method that can be used to train models for use on different systems. However, the lifecycle of models typically used for joint learning can be strict because:

[0104] 1) New workers are not allowed to join the union of activities. New workers can only join during the selection process.

[0105] 2) It does not solve the feature selection problem because different features may not be equally important to all operators.

[0106] 3) The combined product is a joint averaging model, which may not be able to match the characteristics of a single operator when trying to capture the characteristics of all operators.

[0107] 4) Although federated learning can run on any type of device (e.g., mobile devices, base stations, network nodes, etc.), it typically assumes that all devices have the same capabilities and may not consider the data transmission costs that may arise when federated learning occurs. Therefore, federation can only follow a very strict subject-to-worker cycle, which may be limiting in some cases.

[0108] 5) Joint learning generates joint models. While this may be meaningful in the long run, there are cases where isolated training or even centralized training may produce better performance.

[0109] Some of the embodiments described herein address one or more challenges associated with joint learning. Specifically, some of the embodiments described herein address one or more of these challenges via joint feature selection, joint model fine-tuning, and / or dynamic selection of computational resources for the joint model.

[0110] Joint feature selection involves choosing system features to include in a neural network. Model fine-tuning involves using joint information provided by the master entity to tune the local model. Dynamic selection of computational resources for the joint model can involve computation or estimation of memory requirements, processing power (e.g., floating-point operations per second or FLOPS), resource availability, and network data transfer resources to create a computational topology for the FL model for training / inference. Depending on the capabilities / availability of different devices, decisions can be made regarding joint or non-joint, pre-training, no pre-training, fallback to a more specific model, etc.

[0111] In some embodiments, the operation further includes generating visual, auditory, and / or tactile notifications to the user indicating that a remedial action will be triggered and / or has already been triggered.

[0112] In these embodiments, information about unwanted sounds that are about to be suppressed or muted is further relayed to the user as a user interface (“UX”) element, thus enabling the user to selectively override system defaults in certain situations. One example is a mute dog button, which can be provided to the user [in mobile devices, extended reality (“XR”) glasses, etc.] if the system, based on the current context and possible sound sequences, contains a barking sound component inferred to originate from a dog source. In other words, the user will have the ability to turn the suppression or muting of unwanted sound components in multiple component sequences on or off, and each component can be displayed to the user as a representation of its most probable source based on object recognition from the audio input. A camera can be used to further improve object recognition.

[0113] In some embodiments, the occurrence of a triggering event observed by the home agent system is identified, wherein the operation of predicting the probability of a subsequent sound with defined interference characteristics occurring in at least one microphone signal after the occurrence of the triggering sound is further based on the occurrence of the identified triggering event. The home agent may include a camera configured to recognize the presence of a person, a specific person, the presence of an animal, a specific animal, the opening and closing of a door, the opening and closing of a window, etc. The home agent may include a microphone configured to recognize certain types of sounds (e.g., doorbells, telephone rings, fire alarms, etc.).

[0114] In these embodiments, the triggering events observed by the home agent system include at least one of the following: doorbell; fire alarm; notification of an upcoming package or service delivery; and scheduled incoming call.

[0115] In emerging home automation solutions, it may become commonplace to find solutions that manage, in addition to "just" smart home agents that listen to residents' verbal requests for pizza delivery and order pizzas accordingly (like Alexa), also manage things like doorbells and lighting systems.

[0116] The solution for the detection event chain discussed above can / cannot cause some interference later, and the learning of the same machine learning model, etc., can be managed by the home automation system; because the system controlling, for example, the doorbell and the speaker / microphone associated with the user can detect in the first step that the doorbell has been called by someone outside (but has not yet started playing the jingle), in the second step identify the doorbell as the trigger sound for the dog barking later (given the user context, etc.) and determine from it that some selected user speakers may be muted, in the third step call the microphone to mute, and in a later step determine the speakers to be muted according to the selected rules, and then call the doorbell jingle to play.

[0117] In this regard, a home agent system can be designed with a "Do Not Disturb" setting, which can be automatically invoked in the context of a given user.

[0118] Other definitions and examples:

[0119] In the above description of various embodiments of the inventive concept, it is to be understood that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the inventive concept. Unless otherwise defined, all terms used herein (including technical and scientific terms) have the same meaning as commonly understood by one of ordinary skill in the art to which the inventive concept pertains. It will be understood that terms such as those defined in general dictionaries should be interpreted as having a meaning consistent with their meaning in the context of this specification and related art, and not as having an ideal or overly formal meaning, unless otherwise expressly defined herein.

[0120] When an element is referred to as “connected to,” “coupled to,” “responsive to,” or a variation thereof, it may be directly connected to, coupled to, or responsive to another element, or there may be intermediate elements present. Conversely, when an element is referred to as “directly connected to,” “directly coupled to,” “directly responsive to,” or a variation thereof, there are no intermediate elements present. Throughout the text, similar reference numerals are used to denote similar elements. Furthermore, the terms “coupled,” “connected,” “responsive,” or variations thereof as used herein may include wireless coupling, connection, or responsiveness. As used herein, the singular forms “a,” “an,” and “described” are intended to also include the plural forms unless the context clearly indicates otherwise. For the sake of brevity and / or clarity, well-known functions or structures may not be described in detail. The term “and / or” includes any and all combinations of one or more of the listed items.

[0121] It will be understood that although the terms first, second, third, etc., may be used herein to describe various elements / operations, these elements / operations should not be limited by these terms. These terms are used only to distinguish one element / operation from another. Thus, a first element / operation in some embodiments may be referred to as a second element / operation in other embodiments without departing from the teachings of the inventive concept. Throughout the specification, the same reference numerals or reference symbols denote the same or similar elements.

[0122] As used herein, the terms “comprise,” “comprising,” “comprises,” “include,” “including,” “have,” “has,” or variations thereof are open-ended and include one or more of the stated features, integers, elements, steps, components, or functions, but do not preclude the presence or addition of one or more other features, integers, elements, steps, components, functions, or combinations thereof. Furthermore, as used herein, the common abbreviation “eg (e.g.)” (derived from the Latin phrase “exempli gratia”) may be used to introduce or specify a general example of a previously mentioned item and is not intended as a limitation of that item. The common abbreviation “ie (ie)” (derived from the Latin phrase “id Est”) may be used to specify a specific item in a broader sense of reference.

[0123] This document describes exemplary embodiments with reference to block diagrams and / or flowcharts illustrating computer-implemented methods, apparatus (systems and / or devices), and / or computer program products. It should be understood that the blocks in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented by computer program instructions executed by one or more computer circuits. These computer program instructions can be provided to processor circuitry of general-purpose computer circuitry, special-purpose computer circuitry, and / or other programmable data processing circuitry to produce a machine, such that instructions executed by the processor of a computer and / or other programmable data processing apparatus translate and control transistors, values ​​stored in memory locations, and other hardware components within such circuitry to implement the functions / actions specified in the block diagrams and / or flowcharts, and wherein means (functional bodies) and / or structures for implementing the functions / actions specified in the block diagrams and / or flowcharts are created.

[0124] These computer program instructions may also be stored in a tangible computer-readable medium capable of directing a computer or other programmable data processing apparatus to function in a specific manner, such that the instructions stored in the computer-readable medium produce an article of writing, including instructions that implement the functions / actions specified in the blocks of a block diagram and / or flowchart. Therefore, embodiments of the inventive concept can be implemented in hardware and / or software (including firmware, stored software, microcode, etc.) running on a processor such as a digital signal processor, which may be collectively referred to as a "circuit," a "module," or a variation thereof.

[0125] It should also be noted that in some alternative implementations, the functions / actions marked in the boxes may not occur in the order indicated in the flowchart. For example, depending on the functions / actions involved, two boxes shown consecutively may actually be executed substantially simultaneously, or the boxes may sometimes be executed in reverse order. Furthermore, the functionality of a given box in a flowchart and / or block diagram may be divided into multiple boxes, and / or the functionality of two or more boxes in a flowchart and / or block diagram may be at least partially integrated. Finally, without departing from the scope of the inventive concept, other boxes may be added / inserted between the shown boxes, and / or boxes / actions may be omitted. Moreover, although some boxes include arrows indicating the main direction of communication regarding the communication path, it should be understood that communication may occur in the opposite direction to the indicated arrows.

[0126] Many changes and modifications can be made to the embodiments without substantially departing from the inventive concept. All such changes and modifications are intended to be included within the scope of the inventive concept herein. Therefore, the subject matter disclosed above should be understood as exemplary rather than restrictive, and the examples of the appended embodiments are intended to cover all such modifications, improvements, and other embodiments falling within the spirit and scope of the inventive concept. Therefore, to the fullest extent permitted by law, the scope of the inventive concept should be determined by the widest permissible interpretation of this disclosure, including the following examples of embodiments and their equivalents, and should not be limited to or restricted to the specific embodiments described above.

Claims

1. A device (200, 210, 900, 910), comprising: At least one processor (1200, 1100) is configured to receive at least one microphone signal from at least one microphone (1230); as well as At least one memory (1210, 1110) stores program code executable by at least one processor to perform operations including: The trigger sound is detected in the signal from at least one microphone; After the trigger sound occurs, predict the probability of a subsequent sound having defined interference characteristics and satisfying a defined interference level occurring in the at least one microphone signal; When the probability of a subsequent sound with the defined interference characteristics occurring satisfies the remedial action rule, a remedial action to be performed is triggered to mute the at least one microphone signal or suppress the subsequent sound in the at least one microphone signal.

2. The device (200, 210, 900, 910) according to claim 1, wherein, The operation of predicting the probability of subsequent sounds occurring at a defined level of interference includes determining the probability that at least one of the following conditions is met: The predicted peak decibel level of the subsequent sound exceeds the peak threshold; The predicted duration of the subsequent sound exceeds the duration threshold; The predicted frequency components of the subsequent sound are within the defined frequency band; and The subsequent sound has a predicted sound category that has been defined as unacceptable.

3. The device (200, 210, 900, 910) according to any one of claims 1 to 2, wherein, The operation of determining when the probability of a subsequent sound with the defined interference characteristics meets the remedial action rule includes determining whether at least one of the following context parameters meets the remedial action rule: Device location data when the trigger sound occurs; Time data when the trigger sound occurs; Date data when the trigger sound appears; Characteristics of the background noise component in the at least one microphone signal; User input data indicating whether to trigger the remedial action; An indication that the defined sound source type has been recognized by the camera; as well as User demographic characteristics.

4. The device (200, 210, 900, 910) according to any one of claims 1 to 2, wherein, The operation that triggers the remedial action to mute or suppress the subsequent sound in the at least one microphone signal includes: Predict the duration of the subsequent sound; and For a duration determined based on the predicted duration of the subsequent sound, the at least one microphone signal is muted or the subsequent sound in the at least one microphone signal is suppressed.

5. The device (200, 210, 900, 910) according to any one of claims 1 to 2, wherein, The operation that triggers the remedial action to mute or suppress the subsequent sound in the at least one microphone signal includes: Predict the time delay between the occurrence of the trigger sound and the start of the subsequent sound; and After the trigger sound appears, the remedial action is triggered based on the expiration of the time delay.

6. The device (200, 210, 900, 910) according to any one of claims 1 to 2, wherein, The operation also includes: Train a machine learning algorithm to classify sounds received by said device and / or another device, and identify probabilistic correlations between classified sounds appearing in a sequence; and The trigger sound is processed by the machine learning algorithm to predict the probability of a subsequent sound with the defined interference characteristics appearing in the at least one microphone signal.

7. The device (200, 210, 900, 910) according to claim 6, wherein, The operation also includes: Select the machine learning algorithm to be trained from the set of machine learning algorithms based on at least one of the following context parameters: Device location data when one of the sounds to be classified occurs; Time data when one of the sounds to be classified occurs; Date data when one of the sounds to be categorized appears; The characteristics of the background noise component that appears when one of the sounds to be classified appears; and Sensor data indicating the type of object or environmental parameters sensed.

8. The device (200, 210, 900, 910) according to claim 6, wherein, The operation of training the machine learning algorithm further includes training the machine learning algorithm based on user feedback indicating whether the sound to be classified has defined interference characteristics.

9. The device (200, 210, 900, 910) according to claim 6, wherein, The operation of training the machine learning algorithm also includes: Based on a set of pre-recorded sound sequences received by a device that satisfies a similarity rule with the device, the set of pre-recorded sound sequences is selected from a database of pre-recorded sound sequences; and The machine learning algorithm is trained based on the set of pre-recorded sound sequences.

10. The device (200, 210, 900, 910) according to claim 6, wherein the operation further comprises: The subsequent sound is processed using the machine learning algorithm to predict the next probability of a next subsequent sound having the defined interference characteristics occurring in the at least one microphone signal, wherein the next subsequent sound is received by the device after the subsequent sound is received; and When the probability of the next subsequent sound with the defined interference characteristics occurs satisfies the remedial action rule, the remedial action to be performed is triggered to mute the at least one microphone signal or suppress the next subsequent sound in the at least one microphone signal.

11. The device (200, 210, 900, 910) according to any one of claims 1 to 2, wherein, The operation also includes: A machine learning algorithm is trained to indicate when the probability of a subsequent sound with the defined interference characteristics satisfies the remedial action rule.

12. The device (200, 210, 900, 910) according to claim 11, wherein, The operation of training the machine learning algorithm further includes training the machine learning algorithm based on user feedback indicating when the remedial action is triggered by the user.

13. The device (200, 210, 900, 910) according to claim 12, wherein, The operation of training the machine learning algorithm also includes training the machine learning algorithm based on user feedback indicating that the user has performed at least one of the following remedial actions: The user mutes at least one of the microphones; The user increased the speaker volume; When the subsequent sound occurs, the user removes the device from its position; as well as The device detects an indication that the user has performed an action separate from the operation of the device to suppress the subsequent sound.

14. The device (200, 210, 900, 910) according to any one of claims 1 to 2, wherein, The operation also includes: Generate visual, auditory, and / or tactile notifications to the user indicating that the remedial action will be triggered and / or has already been triggered.

15. The device (200, 210, 900, 910) according to any one of claims 1 to 2, wherein, The operation also includes: The operation of identifying the occurrence of a triggering event observed by the home agent system, wherein, after the occurrence of the triggering sound, the operation of predicting the probability of a subsequent sound having the defined interference characteristics occurring in the at least one microphone signal is further based on the occurrence of the identified triggering event.

16. The device (200, 210, 900, 910) according to claim 15, wherein, The triggering events observed by the home agent system include at least one of the following: doorbell; Fire alarm; Notifications of upcoming package or service deliveries; and The scheduled call.

17. A method performed by a device, the method comprising: Identify (300) the occurrence of a trigger sound in at least one microphone signal received from at least one microphone; After the trigger sound appears, predict (302) the probability of a subsequent sound having defined interference characteristics and satisfying a defined interference level appearing in the at least one microphone signal; When the probability of a subsequent sound with the defined interference characteristics occurs satisfies the remedy action rule, the remedy action to be performed (304) is triggered to mute or suppress the subsequent sound in the at least one microphone signal.

18. The method according to claim 17, wherein, The prediction (302) of the probability of subsequent sound occurrence satisfying the defined interference level includes determining the probability that at least one of the following conditions is met: The predicted peak decibel level of the subsequent sound exceeds the peak threshold; The predicted duration of the subsequent sound exceeds the duration threshold; The predicted frequency components of the subsequent sound are within the defined frequency band; and The subsequent sound has a predicted sound category that has been defined as unacceptable.

19. The method according to any one of claims 17 to 18, wherein, Determining when the probability of a subsequent sound with the defined interference characteristics satisfies the remedial action rule includes determining whether at least one of the following context parameters satisfies the remedial action rule: Device location data when the trigger sound occurs; Time data when the trigger sound occurs; Date data when the trigger sound appears; Characteristics of the background noise component in the at least one microphone signal; User input data indicating whether to trigger the remedial action; An indication that the defined sound source type has been recognized by the camera; as well as User demographic characteristics.

20. The method according to any one of claims 17 to 18, wherein, The remedial action to be performed by the trigger (304) to mute or suppress the subsequent sound in the at least one microphone signal includes: Predict the duration of the subsequent sound (400); and For a duration determined based on the predicted duration of the subsequent sound, the at least one microphone signal is muted (402) or the subsequent sound in the at least one microphone signal is suppressed (402).

21. The method according to any one of claims 17 to 18, wherein, The remedial action to be performed by the trigger (304) to mute or suppress the subsequent sound in the at least one microphone signal includes: Predict (500) the time delay between the occurrence of the trigger sound and the start of the subsequent sound; and After the trigger sound appears, the remedial action is triggered (502) based on the expiration of the time delay.

22. The method according to any one of claims 17 to 18, further comprising: Train (600) a machine learning algorithm to classify sounds received by said device and / or another device and identify probabilistic correlations between classified sounds appearing in a sequence; as well as The trigger sound is processed (602) by the machine learning algorithm to predict the probability of a subsequent sound with the defined interference characteristics appearing in the at least one microphone signal.

23. The method of claim 22, further comprising: Select the machine learning algorithm to be trained from the set of machine learning algorithms based on at least one of the following context parameters: Device location data when one of the sounds to be classified occurs; Time data when one of the sounds to be classified occurs; Date data when one of the sounds to be categorized appears; as well as The characteristics of the background noise component that appears when one of the sounds to be classified appears.

24. The method according to claim 22, wherein, The training (600) machine learning algorithm further includes training the machine learning algorithm based on user feedback indicating whether the sound to be classified has defined interference characteristics.

25. The method according to claim 22, wherein, The training (600) machine learning algorithm also includes: Based on a set of pre-recorded sound sequences received by a device that satisfies a similarity rule with the device, (700) the set of pre-recorded sound sequences is selected from a database of pre-recorded sound sequences; and The machine learning algorithm (702) is trained based on the set of pre-recorded sound sequences.

26. The method of claim 22, further comprising: The subsequent sound is processed (800) by the machine learning algorithm to predict the next probability of a next subsequent sound having the defined interference characteristics occurring in the at least one microphone signal, wherein the next subsequent sound is received by the device after the subsequent sound is received; and When the probability of the next subsequent sound with the defined interference characteristics occurs satisfies the remedial action rule, the remedial action to be performed (802) is triggered to mute or suppress the next subsequent sound in the at least one microphone signal.

27. The method according to any one of claims 17 to 18, further comprising: A machine learning algorithm is trained to indicate when the probability of a subsequent sound with the defined interference characteristics satisfies the remedial action rule.

28. The method according to claim 27, wherein, The training of the machine learning algorithm also includes training the machine learning algorithm based on user feedback indicating when the remedial action is triggered by the user.

29. The method according to claim 28, wherein, The training of the machine learning algorithm also includes training the machine learning algorithm based on user feedback indicating that the user has performed at least one of the following remedial actions: The user mutes the at least one microphone; The user increased the speaker volume; When the subsequent sound occurs, the user removes the device from its position; as well as The device detects an instruction from the user that they have performed an action separate from the operation of the device to suppress the subsequent sound.

30. The method according to any one of claims 17 to 18, further comprising: Generate visual, auditory, and / or tactile notifications to the user indicating that the remedial action will be triggered and / or has already been triggered.

31. The method according to any one of claims 17 to 18, further comprising: The operation of identifying the occurrence of a triggering event observed by the home agent system, wherein, after the occurrence of the triggering sound, the operation of predicting the probability of the occurrence of a subsequent sound with defined interference characteristics in the at least one microphone signal is further based on the occurrence of the identified triggering event.

32. The method according to claim 31, wherein, The triggering events observed by the home agent system include at least one of the following: doorbell; Fire alarm; Notifications of upcoming package or service deliveries; and The scheduled call.

33. A computer program product comprising: A non-transitory computer-readable medium storing program code executable by at least one processor of the device to perform operations including: Identify (300) the occurrence of a trigger sound in at least one microphone signal received from at least one microphone; After the trigger sound appears, predict (302) the probability of a subsequent sound having defined interference characteristics and satisfying a defined interference level appearing in the at least one microphone signal; When the probability of a subsequent sound with the defined interference characteristics occurs satisfies the remedy action rule, the remedy action to be performed (304) is triggered to mute or suppress the subsequent sound in the at least one microphone signal.

34. The computer program product according to claim 33, wherein, The program code executable by at least one processor of the device also performs the operation of any of the methods described according to claims 18 to 32.

35. An apparatus (200, 210, 900, 910) adapted to perform the method according to any one of claims 17 to 32.

Citation Information

Patent Citations

  • Event masking

    EP3779810A1