Assistive technology

By designing a computer device including an audio stream acquisition unit, a sound detector and a situation determiner, the problem of determining the situation using non-verbal sound prompts is solved, and efficient auxiliary facility generation and user environment adaptation are achieved.

CN112700765BActive Publication Date: 2025-06-17CTRL-LABS CORP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202011043265.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-10-07
Filing Date
2020-09-28
Publication Date
2025-06-17
Estimated Expiration
2040-09-28

AI Technical Summary

Technical Problem

The prior art is difficult to effectively utilize nonverbal voice cues to determine situations and provide auxiliary facilities.

Method used

A computer device is designed, including an audio stream acquisition unit, a sound detector and a situation determiner. The device determines whether a specific situation is satisfied by detecting a nonverbal sound identifier in the audio sample stream, and generates auxiliary output based on the situation.

Benefits of technology

The ability to determine situations and generate auxiliary responses based on nonverbal sounds and scenes is realized, the intelligence and personalization of auxiliary facilities are improved, and adaptability to the user environment is enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112700765B_ABST
    Figure CN112700765B_ABST
Patent Text Reader

Abstract

Provided is a device or system that is configured to detect one or more sound events and / or scenarios associated with a predetermined situation and provide an auxiliary output when the situation is met.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure generally relates to providing assistive facilities to a user based on a context determined from non-verbal cues. Background Art

[0002] Background information on voice recognition systems and methods can be found in the applicant's PCT application WO2010 / 070314, the entire content of which is incorporated herein by reference.

[0003] The present applicant has recognized the potential of new applications of voice recognition systems. Summary of the Invention

[0004] Techniques have been widely adopted to assist users in their daily lives. It has become common for users to deploy assistive technologies as sources of information or to provide them with cues or reminders for performing certain tasks.

[0005] For example, in a home environment, computer-assisted devices can be deployed that are configured to provide a user with reminders in the form of a display, an audible alarm, a tactile stimulus, or computer-generated speech according to a schedule. Additionally, or alternatively, such devices can provide facilities for automating certain actions. Thus, for example, an assistive device can issue instructions to be implemented by appropriate collaborative devices to turn on or off house lighting, or open or close curtains, or generate a sound output intended to wake a sleeping person. Such actions can be pre-arranged by the user of the device.

[0006] For example, in an automotive environment, it is well known to provide a navigation system that is designed to provide a driver with graphical and audible instructions to reach a destination as effectively as possible. Such instructions can be adapted based on information about road traffic conditions or other criteria.

[0007] Generally, a device or system is provided that is configured to detect one or more sound events and / or scenarios associated with a predetermined context and to provide an assistive output when that context is met.

[0008] Aspects of the present disclosure provide a computer device operable to generate an assistive output based on context determination, the device including: an audio stream acquisition unit for acquiring a stream of audio samples; a sound detector for detecting one or more non-verbal sound identifiers on the audio sample stream, each non-verbal sound identifier identifying a non-verbal sound signature on the audio sample stream; a context determiner for determining that a particular context has been met based on the detection of one or more indicative non-verbal sound identifiers and for generating an assistive output based on the context.

[0009] Aspects of the present disclosure provide a computer device that can determine whether a predetermined situation has been met based on recognizable non-verbal sounds and / or scenarios on an audio input stream and thus generate an assistive response to the situation.

[0010] The determination of whether a situation has been met can be made in a variety of ways. In a simple example, a single instance of a particular sound event may result in the satisfaction of a situation. Combinations of sound events can satisfy a situation. More complex combination methods can be further used to determine the satisfaction of a situation. The satisfaction of a situation can be relative to a situation model. The situation model can include a processing network model, such as a neural network or a decision tree, and a machine model can be developed using machine learning on training data that consists of "valid" combinations of sound events associated with a particular situation. The machine learning may be adaptive in use, and the device may obtain further training from user feedback in response to potential incorrect responses to real data.

[0011] It will be understood that the functionality of the devices described herein can be divided into multiple modules. Alternatively, the functionality can be provided in a single module or processor. The processor or each processor can be implemented with any known suitable hardware, such as a microprocessor, a digital signal processing (DSP) chip, an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), a GPU (graphics processing unit), a TPU (tensor processing unit), or an NPU (neural processing unit), etc. The processor or each processor can include one or more processing cores, each core configured to execute independently. The processor or each processor can have connectivity to a bus to execute instructions and process information stored, for example, in a memory.

[0012] The present invention also provides processor control code for implementing the above-mentioned system and method, for example, on a general-purpose computer system, a digital signal processor (DSP), or a specially designed mathematical acceleration unit (such as a graphics processing unit (GPU) or a tensor processing unit (TPU)). The present invention also provides a carrier carrying the processor control code for implementing any of the above methods during runtime, especially on a non-transitory data carrier, such as a disk, a microprocessor, a CD- or DVD-ROM, a programmable memory, such as a read-only memory (firmware), or a data carrier (such as an optical or electrical signal carrier). The code can be provided on a carrier such as a disk, a microprocessor, a CD- or DVD-ROM, a programmable memory such as a non-volatile memory (e.g., flash memory) or a read-only memory (firmware). The code (and / or data) for implementing an embodiment of the present invention can include source code, object code, or executable code (interpreted or compiled) in a conventional programming language (such as C) or assembly code, code for setting or controlling an ASIC (application-specific integrated circuit) or an FPGA (field-programmable gate array), or code for a hardware description language (such as VerilogTM or VHDL) (hardware description language for high-speed integrated circuits). As those skilled in the art will understand, such code and / or data can be distributed among multiple coupled components communicating with each other. The present invention can include a controller that includes a microprocessor, a working memory, and a program memory coupled to one or more components of the system.

[0013] These and other aspects will become very clear from the embodiments described below. The scope of the present disclosure is neither limited to this overview nor to embodiments that must address any or all of the noted disadvantages. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] To better understand the present disclosure and to illustrate how embodiments work, reference is made to the accompanying drawings, in which:

[0015] Figure 1 A block diagram of an example device in a monitored environment is shown;

[0016] Figure 2 A block diagram of a computing device is shown.

[0017] Figure 3 A block diagram of software implemented on a computing device is shown;

[0018] Figure 4 is a flowchart showing a process for providing an auxiliary output according to an embodiment.

[0019] Figure 5 is a process architecture diagram showing the implementation of an embodiment and indicating the functions and structures of such an implementation. DETAILED DESCRIPTION

[0020] Embodiments are now described by way of example only.

[0021] Figure 1 A computing device 102 in a monitored environment 100 is shown, which can be an indoor space (e.g., a house, a gym, a store, a train station, etc.), an outdoor space, or in a vehicle.

[0022] The network 106 can be a wireless network, a wired network, or can include a combination of wired and wireless connections between devices.

[0023] As described in more detail below, the computing device 102 can perform audio processing to identify (i.e., detect) a target sound in the monitored environment 100. In an alternative embodiment, a sound recognition device 104 external to the computing device 102 can perform audio processing to identify a target sound in the monitored environment 100 and then alert the computing device 102 that the target sound has been detected.

[0024] Figure 2 A block diagram of the computing device 102 is shown. It will be appreciated from the following that Figure 2 is illustrative only, and the computing device 102 of embodiments of the present disclosure may not include Figure 2 all of the components shown in

[0025] The computing device 102 can be a PC, a mobile computing device such as a laptop computer, a smart phone, a tablet PC, a consumer electronic device (e.g., a smart speaker, a TV, headphones, a wearable device, etc.), or other electronic device (e.g., a vehicle-mounted device). The computing device 102 can be a mobile device such that the user 103 can move the computing device 102 around the monitored environment. Alternatively, the computing device 102 can be fixed at a certain position in the monitored environment (e.g., mounted on a panel on a house wall). Alternatively, the user can wear the device by attaching it to a body part or placing it on a body part or by attaching it to a piece of clothing.

[0026] The computing device 102 includes a processor 202 coupled to a memory 204, and the memory 204 stores computer program code of application software 206 that can operate with data elements 208. As Figure 3 shown, a mapping of the memory in use is shown. The sound recognition software 206a is used to identify a target sound by comparing the detected sound with one or more sound models 208a stored in the memory 204. The sound models 208a can be associated with one or more target sounds (which can be, for example, the sound of broken glass, a smoke alarm sound, a baby crying sound, a sound indicating that an action is being performed, etc.).

[0027] The context determination software 206b optionally combines other factors such as geographical location or time of day to determine if a context is met by detecting specific sound events and / or scenarios (such as those mentioned above). The context determination software 206b is enabled by one or more context models 208b and is developed to identify contexts based on one or more relationships between specific sound events and / or scenarios that characterize a particular context.

[0028] The assistance software 206c manages the response to context satisfaction. Thus, in response to a particular context being met, the assistance software responds by generating an assistance output. For example, this can be a signal to the user, such as a display element, an audible output, a tactile stimulus, or a remote alert. On the other hand, or additionally, it can be an electrical signal or other signal for a collaborating device to receive to actuate another device, such as an electrical switch. It can also be a telecommunications, such as the initiation of a message or a telephone communication session.

[0029] The user interface software 206d initiates the generation of a user interface for inviting the user to perform a user input action. Such a user interface can take various forms. Thus, for example, the user interface can include a graphical user interface that provides display elements for inviting the user to input an action, such as selecting a button on the screen or entering information into a designated screen input field. On the other hand, or additionally, the user interface can be audio-based. In this case, the user interface software 206d is capable of receiving and interpreting speech audio and converting it into data input for controlling other aspects of the implementation. In that case, the user interface software 206d can be used to generate computer-synthesized speech output for interacting with the user.

[0030] The user interface software 206d, however it is implemented, is supported by user interface data 208d that stores information from which the user interface can be implemented.

[0031] The computing device 102 can include, for example, one or more input devices. Physical buttons (including a single button, a keypad, or a keyboard) or physical controls (including a knob or a dial, a scroll wheel, or a touch bar) 210 and / or a microphone 212. The computing device 102 can include one or more output devices, for example, a speaker 214 and / or a display 216. It should be understood that the display 216 can be a touch-sensitive display and thus can be used as an input device.

[0032] The computing device 102 can also include a communication interface 218 for communicating with one or more controllable devices 108 and / or sound recognition devices 104. The communication interface 218 can include a wired interface and / or a wireless interface.

[0033] As Figure 3As shown, the computing device 102 can store the sound model locally (in the memory 204), and thus does not need to maintain constant communication with any remote system for identifying the captured sound. Alternatively, the sound model 208a is stored on a remote server coupled to the computing device 102 ( Figure 2 not shown in the figure), and the sound recognition software 206 on the remote server is used to perform processing of the audio received from the computing device 102 to identify that the sound captured by the computing device 102 corresponds to the target sound. This advantageously reduces the processing performed on the computing device 102.

[0034] Recognition of acoustic models and acoustic events and / or scenarios

[0035] Based on the processing of the captured sound corresponding to the sound event and / or scene category, a sound model 208a associated with the sound event and / or scene is generated. Preferably, multiple instances of the same sound are captured multiple times to improve the reliability of the sound model generated by the captured sound event and / or scene category.

[0036] To generate the sound model, the captured sound event and / or scene category is processed, and parameters are generated for the specific captured sound event and / or scene category. The generated sound model includes these generated parameters and other data that can be used to characterize the captured sound event and / or scene category.

[0037] There are various ways to generate a sound model associated with the target sound category. The sound model of the captured sound can be generated using machine learning techniques or predictive modeling techniques, such as: hidden Markov models, neural networks, support vector machines (SVMs), decision tree learning, etc.

[0038] The applicant's PCT application WO2010 / 070314 (incorporated herein by reference in its entirety) details various methods for identifying sounds. Generally speaking, the input sample sound is processed by decomposing it into frequency bands, for example, PCA / ICA can be used for decorrelation, and then this data is compared with one or more Markov models to generate log-likelihood ratio (LLR) data for the input sound to be identified. Then a (hard) confidence threshold can be used to determine whether the sound has been identified. If a "fit" to two or more stored Markov models is detected, the system preferentially selects the most likely model. By effectively comparing the sound to be identified with the expected frequency domain data predicted by the Markov model, the sound can be "fitted" into the model. False positives can be reduced by correcting / updating the mean and variance in the model based on interfering (including background) noise.

[0039] It should be understood that other techniques than those described herein can be employed to create the sound model.

[0040] A voice recognition system can use compressed audio or uncompressed audio. For example, the time-frequency matrix for a 44.1KHz signal might be a 1024-point FFT with 512 overlaps. This is approximately a 20-millisecond window with a 10-millisecond overlap. The resulting 512 frequency bands are then grouped into sub-bands, or quarter octaves between, for example, 62.5 and 8000Hz, giving 30 sub-bands.

[0041] A lookup table can be used to map from the compressed or uncompressed frequency bands to a new sub-band representation of the frequency bands. For a given example of sample rate and STFT size, for each supported sample rate / bin (segment) number pair, the array might include an array of (Binsize÷2) x 6 ((segment size÷2) x 6). These rows correspond to the segment number (center) - STFT size or the number of frequency coefficients. The first two columns determine the lower and higher quarter octave segment index numbers. The next four columns determine the proportion of the segment size that should be placed in the corresponding quarter octave segments starting from the low quarter octave defined in the first column to the high quarter octave defined in the second column. For example, if the segment overlaps two quarter octave ranges, the sum of the proportion values in columns 3 and 4 is 1, while the proportion values in columns 5 and 6 are zero. If the segment overlaps more than one sub-band, more columns will have proportion size values. This example models the critical frequency bands in the human auditory system. This reduced time / frequency representation is then processed by the outlined normalization method. This process is repeated for all frames with a 10-ms hop size incrementally moving the frame position. The overlapping window (hop size not equal to window size) improves the time resolution of the system. This is considered a sufficient representation of the signal frequency and can be used to summarize the perceptual characteristics of the sound. Then, the normalization phase divides each frame in the sub-band decomposition by the square root of the average power in each sub-band. The average is calculated as the total power in all frequency bands divided by the number of frequency bands. This normalized time-frequency matrix is passed to the next part of the system, where a voice recognition model and its parameters can be generated to fully characterize the frequency distribution and time trend of the sound.

[0042] The next stage of voice characterization needs to be further defined.

[0043] Machine learning models are used to define and obtain the trainable parameters required to recognize voices. The definition of such a model is:

[0044] - A set of trainable parameters θ, such as but not limited to, the means, variances, and transitions of a Hidden Markov Model (HMM), the support vectors of a Support Vector Machine (SVM), the weights, biases, and activation functions of a Deep Neural Network (DNN),

[0045] - A dataset with audio observations o and associated sound tags l, e.g., a set of audio recordings that capture a set of target sounds of interest for identification, such as baby cries, dog barks, or smoke alarms, and other background sounds that are not the target sounds to be identified and may be adversely identified as target sounds. This audio observation dataset is associated with a set of tags l that indicate the location of the target sounds of interest, e.g., the time and duration at which a baby cry occurs in the audio observation o.

[0046] Generating model parameters is a problem of defining and minimizing a loss function over the entire set of audio observations where the minimization is performed by a training method, such as but not limited to, the Baum-Welsh algorithm for HMMs, soft margin minimization for SVMs, or stochastic gradient descent for DNNs.

[0047] To classify new sounds, an inference algorithm uses the model and its parameters θ to determine the probability or score P(C|o,θ) that the new incoming audio observation o is associated with one or more sound classes C. Then, the probability or score is converted to a discrete sound class symbol by a decision method such as but not limited to thresholding or dynamic programming.

[0048] These models will operate under many different acoustic conditions, and since it is actually limited to give examples representing all the acoustic conditions the system will encounter, the model will be internally adjusted so that the system can operate under all these different acoustic conditions. Many different methods can be used for this update. For example, the method can include taking the average of the individual subbands, e.g., the one-quarter octave frequency values over the last T seconds. These averages are added to the model values to update the internal model of the sounds in that acoustic environment.

[0049] In an embodiment where a computing device 102 performs audio processing to identify target sounds in a monitored environment 100, the audio processing includes a microphone 212 of the computing device 102 that captures sounds, and a sound recognition 206a that analyzes the captured sounds. In particular, the sound recognition 206a compares the captured sounds with one or more sound models 208a stored in a memory 204. If the captured sound matches the stored sound model, the sound is identified as a target sound.

[0050] The sequence of identified target sounds can thus be passed to a sequence-to-sequence model 206b for processing in the context of controlling the navigation of a document 208c supported by a browser 206c.

[0051] In the present disclosure, the target sounds of interest are non-verbal sounds. Many use cases will be described in due course, but the reader will understand that various non-verbal sounds can trigger navigation actions. The present disclosure and the specific selection of examples employed herein should not be construed as limiting the scope of applicability of the underlying concepts.

[0052] Situation determination

[0053] The resulting sequence of non-verbal sound identifiers generated by the sound recognition 206a is passed to the context determination software 206b to determine whether it characterizes a context defined in the context definition model 208b.

[0054] The context definition model 208b encodes a context as one or more relationships between a set of sound events and / or sounds collected in a scene. Relationships can include, but are not limited to, the order of occurrence of sound events and / or scenes collected in the set under consideration, their co-occurrence within a predefined time window, their distance in time, their probability - occurrence count (n-gram) or any other form of weighted or unweighted graph. These context definitions can be obtained in a variety of ways, such as, but not limited to, through manual programming of an expert system, or through machine learning (e.g., but not limited to) using deep neural networks, decision trees, Gaussian mixture models, or probabilistic n-grams.

[0055] It should be noted that although the sound recognition process 206a converts an audio stream into one or more sound events and / or scenes (possibly with timestamps), context recognition converts a set of (possibly timestamped) sound descriptors into a decision about the context. For example, a context definition model can be defined as "having breakfast" or "leaving the house". Each of these will be stored as a set of sound events and / or scenes and one or more of their relationships. Detecting one or more relationships between sound events and / or scenes in a set of sound events and / or scenes emitted from the sound recognition process will result in the determination that a specific recognized context has been satisfied.

[0056] Auxiliary output

[0057] As a result of satisfying a specific recognized context, an auxiliary output is generated. The auxiliary output can be directly mapped to the satisfied context, and this mapping is stored in the memory.

[0058] The auxiliary output can be (non-exhaustively) a synthetic speech audio output, an audible alarm, a graphical display, electromagnetic communication with another device, wired electrical communication with another device, or any combination of the above.

[0059] Processing

[0060] Figure 4FIG. 400 is a flowchart showing a process 400 for controlling a user interface of a computing device according to a first embodiment. The steps of process 300 are performed by processor 202.

[0061] In step S402, processor 202 identifies one or more sound events and / or scenarios in the monitored environment 100.

[0062] Microphone 212 of computing device 102 is arranged to capture sound in the monitored environment 100. Step S402 may be performed by the processor to convert the captured sound pressure wave into digital audio samples and execute sound recognition software 206 to analyze the digital audio samples (the processor may compress the digital audio samples before performing this analysis). In particular, sound recognition software 206 compares the captured sound with one or more sound models 208 stored in memory 204. If the captured sound matches the stored sound model, the captured sound is identified as the target sound. Alternatively, processor 202 may send the captured sound to a remote server via communication interface 218 for processing to identify whether the sound captured by computing device 102 corresponds to the target sound. That is, processor 202 may identify the target sound in the monitored environment 100 based on a message received from the remote server that the sound captured by computing device 102 corresponds to the target sound.

[0063] Alternatively, the microphone of sound recognition device 104 may be arranged to capture sound in the monitored environment 100 and process the captured sound to identify whether the sound captured by sound recognition device 104 corresponds to the target sound. In this example, sound recognition device 104 is configured to send a message to computing device 102 via network 106 to alert computing device 102 that a target sound has been detected. That is, processor 202 may identify the target sound in the monitored environment 100 based on a message received from sound recognition device 104.

[0064] Regardless of where the processing of the captured sound is performed, the identification of the target sound includes identifying non-verbal sounds that may be generated in the environment of the sound capture device (computing device 102 or sound recognition device 104), such as the sound of breaking glass, a smoke alarm, a baby crying, onomatopoeic words, the sound of a quiet house, or the sound of a train station.

[0065] In step S404, the processor 202 determines the satisfaction of the situation, as defined by the situation model 208b. This can be an ongoing process - the processor may be configured to load a specific situation model 208b into short-term memory and thus focus on the stream of sound events and / or scenes to detect whether the sound relationships indicating that situation are satisfied. This can be pre-established by user input operations. Thus, for example, the user can input to the device a need to be warned about the presence of a specific one or more situations. The user can further configure the device to determine whether an alarm should be set once or each time the situation is encountered.

[0066] The sequence of (possibly timestamped) sound events and / or scene descriptors received by the situation determination process is analyzed as a set, where the set is not necessarily ordered. The situation model is represented, for example, by a graph of sound event and / or scene co-occurrences and can be decoded by, for example, the Viterbi algorithm, but other models can also be used to learn co-occurrence models from data, such as decision trees or deep neural networks.

[0067] Other methods are possible - upon receiving a specific sound event and / or scene, the processor 202 can search the situation model to find candidate situations that are satisfied and then monitor future sound events and / or scenes until one of those conditions is met.

[0068] In step S406, the processor 202 issues an auxiliary output or alarm corresponding to the satisfied situation.

[0069] Use case

[0070] The following are some use cases intended to illustrate the scope of applicability of the above technology. None of these use cases should be construed as a limitation on potential applicability.

[0071] Situations can be defined to monitor the progress and completion of a child's morning routine from waking up until leaving for school. In this case, an audio input stream can be obtained from a bathroom smart speaker. Using that speaker, it can be detected that a specific individual (e.g., a child) has brushed their teeth, used the toilet, and washed their hands before going to school. Thus, a first situation can be defined as an "in progress" situation where the morning routine has started but is not yet complete. In response to detecting this, information can be pushed to a smartphone (e.g., the parent's smartphone) so that the parent can be updated on the progress of the entire morning routine. When it is detected that the morning routine has been completed, a further alarm can be sent to the smartphone, thus entering the "ready for school" situation.

[0072] In another case, a situation can be defined around home security. For example, a home assistant can monitor actions and events associated with the occupants of the house getting ready to go to work. The home assistant can detect sound events and / or scenarios associated with the occupants making final preparations to leave, such as putting on shoes or picking up a bunch of keys. The home assistant can respond to such events and / or scenarios to determine whether the prior sequence of events and / or scenarios matches the expected multiple events and / or scenarios associated with the morning routine. In response to any mismatch, an auxiliary output may be generated. Thus, for example, the home assistant may generate an output in response to any such detected mismatch, such as "Wait, you forgot to fill the dishwasher" or "Wait, you forgot to turn off the kitchen faucet."

[0073] In another case, a system including multiple appropriately configured devices can enable monitoring of a user's exposure to audible noise. The goal of such devices is to track, monitor, and establish better routines around daily noise exposure. Sounds collected from wearable devices or the user's smart headphones, etc., can accurately record the user's state of exposure to sounds, including sound level, time intensity, and sound type. There is a recognized link between exposure to harmful sounds (noise) and mood. Noise can trigger stress in extreme cases. The system can be configured to define a situation associated with exceeding a daily dose of exposure to certain sounds and issue an alert to the user in response to input of that situation.

[0074] In another case, a device can be configured to detect a period of relative quiet indoors (despite the presence of a human user) as a sound event and / or scenario. Thus, the device can define a situation around this quiet period as an opportunity for the user to rest. In this case, an auxiliary output can be generated as an audible synthetic voice output to the user, such as "You've been so busy. It's noon... How about some relaxing music?"

[0075] In another case, a situation can be defined around the user's healthy sleep cycle. Thus, a smart speaker in the bedroom can be deployed to detect the stages of the user's sleep cycle based on the intensity and occurrence of breathing sounds, movement in the bed, and the time of day / night. Based on this, a morning home heating cycle can be initiated using a sound output or a trigger such as triggering a heating system to determine the time that is most suitable or healthiest for waking up the user. Other outputs that may be triggered include sending a message to an automatic shower system to start the water flow so that the user can walk to a preheated shower, a coffee machine starting to brew a pot of coffee, or starting other audiovisual effects such as a TV presentation, an email, a browser, or other appropriate actions on appropriate devices.

[0076] Similar devices can be further configured to determine whether a user has had a poor sleep night based on a detected sequence of sound events and / or scenarios. In response to detecting such a situation, the device can be configured to trigger a corresponding secondary output. Thus, the device can, for example, output an audible synthetic voice to convey information to the user to encourage morning relaxation (e.g., reading, listening to music), or connect to a family member for support. For example, the device can be used by a frail or elderly person, especially one with diminished speech capabilities, to alert a third party of a change in health condition or the need for assistance.

[0077] In a network of suitably collaborating devices, such as in a home, sound events and / or scenarios may be monitored to enable users to share facilities more effectively. Thus, for example, a device may be able to monitor whether the bathroom is in use. The device can be configured to monitor the vacancy of the bathroom and emit a sound output in response thereto. Thus, for example, a user may initiate the monitoring process by issuing a voice command such as "tell me when the bathroom is empty", and one or more devices respond to sound events and / or scenarios associated with the opening of the bathroom door and other sounds that may indicate the bathroom is vacant. In such a case, one or more devices will emit an audible synthetic voice output, such as "the bathroom is empty". Similarly, the sound of placing a breakfast tray on the table may trigger an output of "breakfast is almost ready" to the child's bedroom.

[0078] In one embodiment, further facilities can be provided to enable a user to configure the device to operate in a specific manner. Thus, for example, the device can accept a user input action regarding how the user wishes to receive an alert related to the occurrence of a specific situation, such as a spoken user input action. For example, the situation detection can be turned on or off by the user, or the monitoring of a specific situation can be enabled or disabled. In addition, it can be configured whether the situation occurs once, a predetermined number of times (e.g., "snooze" function) or an alert is issued each time it occurs.

[0079] In a situation model, the relationship between sounds can be enhanced by leveraging the relationship with other information items corresponding to the sounds. For example, the occurrence of sounds in space or time can be recorded as part of a sound event. Using the identification of the sound event and optionally the time or location of occurrence of the sound event, further conclusions can be drawn about the identifiable situations defined in the situation model. Thus, for example, if sounds associated with the preparation or consumption of breakfast occur in the morning or at a specific location in the house related to breakfast (e.g., the kitchen), there may be a stronger correlation with the situation described as "having breakfast". Similarly, if sounds associated with a sound occur at a specific time or at a location associated with that sound (e.g., the dining room), the sounds associated with the sound referred to as "eating" may have a stronger relationship with that sound.

[0080] Overview

[0081] As Figure 5 shown, the overall structure and functionality of a system 500 designed to implement the above use case are presented. In this case, a first digital audio acquisition block 510 receives an audio signal from a microphone 502 and generates a series of waveform samples. These samples are passed to a sound detection block 520, which generates a sound identifier for each sound event and / or scene detectable on the waveform samples. Each sound identifier includes information identifying the sound event and / or scene, namely what the sound is, whether the sound is starting or ending (or in some cases, the duration of the event and / or scene), and the time of the event and / or scene.

[0082] The functionality of the sound detection block 520 is further constituted by data stored in a control sound recognition and alert block 550, which itself is configured by user input actions on a user interface 540. In this embodiment, a typical user input action is: setting an alert for an audio situation. Thus, for example, a user can enter a request such that if a sound associated with breakfast preparation is recognized, an alert will be sent to the user's device 560 (which can be, for example, a smartphone).

[0083] Thus, by being appropriately configured, the sound detection unit 520 is actively monitoring for the occurrence of sound events and / or scenes, and because they are related in a particular way, it identifies the situation of the breakfast being prepared. Then, a decision 530 is made regarding whether the situation has been met. If it has not been met, the sound detection block 520 continues to detect sounds. If it has been met, this decision is relayed back to the control sound recognition and alert block 550, and an alert associated with the situation is sent to the user's device 560.

[0084] Separate computers can be used for the various stages of processing. Thus, for example, the user input can be on a first device, which can be a smartphone. The configuration of the sound detection and situation detection can be performed on another device. In fact, Figure 5 all of the functions shown can be performed on separate computers and can be networked with each other. Alternatively, all of the above functions can be provided on the same computing device.

[0085] Aspects of the embodiments disclosed herein can provide certain advantages to users in terms of the utility of a computing device. For example, the combination of artificial intelligence used in an automatic sound event and / or scene recognition system, in combination with a context detection system, can increase the relevance of an alert to a context. Thus, for example, an alert can be associated with a context rather than a specific time, allowing the system to adapt to the user rather than strictly adhering to a real-time schedule. Embodiments can also relieve people of the attention of monitoring a series of detectable events and / or scenes that can be recognized through sound events and / or scenes. Embodiments can also enhance the ability of humans to monitor a series of sound events and / or scenes and / or scenes indicating contexts that have occurred in many rooms, while sleeping, or on various sound sensors, which is a task that humans cannot perform because they cannot place themselves at multiple monitoring points simultaneously or in a short period of time.

Claims

1. A computer device capable of operating to generate an auxiliary output based on a context determination, the device comprising: An audio stream acquisition unit for acquiring an audio sample stream; A sound detector for detecting a plurality of non-verbal sound events and / or scenarios from the audio sample stream; A sound processor for processing the plurality of non-verbal sound events and / or scenarios to determine, based on the plurality of non-verbal sound events and / or scenarios, a sound event and / or scenario identifier for each of the plurality of non-verbal sound events and / or scenarios, each of the plurality of non-verbal sound event and / or scenario identifiers identifying a non-verbal sound event and / or scenario from the audio sample stream; An activity context determiner for determining that a specific activity context has been satisfied based on a plurality of determined non-verbal sound event and / or scenario identifiers, the activity context being associated with a completion state of an activity including a plurality of associated actions or events, wherein satisfaction of the specific activity context is defined by an activity context model of the specific activity context, and wherein the activity context determiner is configured to: input the plurality of determined non-verbal sound event and / or scenario identifiers into the activity context model of the specific activity context; and receive an indication that the specific activity context has been satisfied from the activity context model, wherein the specific activity context is associated with a recommended user action; and An auxiliary output generator for generating, based on the indication that the specific activity context has been satisfied, an auxiliary output for the specific activity context to the user, the auxiliary output communicating auxiliary information to the user, the auxiliary information being operable to prompt the user to adopt the recommended user action.

2. The computer device according to claim 1, wherein, The activity context determiner is operable to determine satisfaction of the activity context based on detection of non-verbal sound event and / or scenario identifiers related to the activity context.

3. The computer device according to claim 1, wherein, The activity context determiner is operable to determine satisfaction of the activity context based on a time metric, the time metric being a metric relative to real time or an instance of a non-verbal sound event and / or scenario with respect to detection of one or more non-verbal sound events and / or scenarios combined.

4. The computer device according to claim 1, wherein, The activity context determiner is operable to determine satisfaction of the activity context based on a location metric in combination with detection of one or more non-verbal sound events and / or scenarios.

5. The computer device according to claim 1, wherein, The activity context determiner is operable to determine, based on a plurality of activity context definitions, which activity context definition is satisfied in the presence of activity context definitions by detecting one or more non-verbal sound event and / or scenario identifiers.

6. The computer device according to claim 1, wherein, The activity context model is implemented using machine learning.

7. The computer device according to claim 1, wherein, The activity context determiner includes a decision tree.

8. The computer device according to claim 1, wherein, The activity context determiner includes a neural network.

9. The computer device according to claim 1, wherein, The activity context determiner includes a weighted graph model.

10. The computer device according to claim 1, wherein, The activity context determiner includes a hidden Markov model.

11. The computer device according to claim 1, wherein, The auxiliary output generator is operable to output an alarm signal based on the satisfied context.

12. The computer device according to claim 11, wherein, The alarm signal includes at least one of an auditory alarm, a visual alarm, a tactile alarm, and a remote alarm.

13. The computer device according to claim 1, wherein, The auxiliary output generator is operable to output an auxiliary output associated with the satisfied activity context.

14. The computer device according to claim 1, comprising a user interface unit capable of operating to implement a user interface for receiving a signal corresponding to a user input action, and wherein, The activity context determiner responds to a user input action to associate a context with an auxiliary output.

15. The computer device according to claim 1, comprising a user interface unit capable of operating to implement a user interface for receiving a signal corresponding to a user input action, and wherein, The activity context determiner responds to a user input action to associate the fulfillment of an activity context with the detection of one or more non-verbal sound identifiers.

16. A computer-implemented method for generating an auxiliary output based on a context determination, the method comprising: Obtain an audio sample stream; Detect one or more non-verbal sound events and / or scenes from the audio sample stream; Process the plurality of non-verbal sound events and / or scenes to determine, based on the plurality of non-verbal sound events and / or scenes, a sound event and / or scene identifier for each of the plurality of non-verbal sound events and / or scenes, each of the plurality of non-verbal sound event and / or scene identifiers identifying a non-verbal sound event and / or scene from the audio sample stream; Based on the plurality of determined non-verbal sound event and / or scene identifiers, determine that a specific activity context has been satisfied, wherein the satisfaction of the specific activity context is defined by an activity context model of the specific activity context, the activity context being associated with a completion state of an activity including a plurality of associated actions or events, wherein determining that the specific activity context has been satisfied includes: inputting the plurality of determined non-verbal sound event and / or scene identifiers into the activity context model of the specific activity context; and receiving an indication from the activity context model that the specific activity context has been satisfied, wherein the specific activity context is associated with a recommended user action; and Based on the indication that the specific activity context has been satisfied, generate an auxiliary output for the specific activity context to the user, the auxiliary output communicating auxiliary information to the user, the auxiliary information being operable to prompt the user to adopt the recommended user action.

17. A non - transitory computer - readable medium stores computer - executable instructions that, when executed by a general - purpose computer, cause the general - purpose computer to perform the following steps: Obtain an audio sample stream; Detect one or more non - verbal sound events and / or scenarios from the audio sample stream; Process the plurality of non-verbal sound events and / or scenes to determine, based on the plurality of non-verbal sound events and / or scenes, a sound event and / or scene identifier for each of the plurality of non-verbal sound events and / or scenes, each of the plurality of non-verbal sound event and / or scene identifiers identifying a non-verbal sound event and / or scene from the audio sample stream; Based on the plurality of determined non-verbal sound event and / or scene identifiers, determine that a specific activity context has been satisfied, wherein the satisfaction of the specific activity context is defined by an activity context model of the specific activity context, the activity context being associated with a completion state of an activity including a plurality of associated actions or events, wherein determining that the specific activity context has been satisfied includes: inputting the plurality of determined non-verbal sound event and / or scene identifiers into the activity context model of the specific activity context; and receiving an indication from the activity context model that the specific activity context has been satisfied, wherein the specific activity context is associated with a recommended user action; and Based on the indication that the specific activity context has been satisfied, generate an auxiliary output for the specific activity context to the user, the auxiliary output communicating auxiliary information to the user, the auxiliary information being operable to prompt the user to adopt the recommended user action.

Citation Information

Patent Citations

  • Sound identification systems

    WO2010070314A1

  • System and Method for Audio Scene Understanding of Physical Object Sound Sources

    US20170105080A1