System for providing real-time feedback to reduce undesirable speaking patterns and method of use thereof

The system provides real-time feedback on filler speech using portable devices, addressing the limitations of existing solutions by offering discreet and effective correction of speech habits during face-to-face interactions.

JP2026502393APending Publication Date: 2026-01-22CDC PHONE APIPIPIPII 2023 LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025549420
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-10-31
Filing Date
2023-10-30
Publication Date
2026-01-22

AI Technical Summary

Technical Problem

Existing solutions for reducing excessive filler speech during face-to-face conversations are ineffective for real-time feedback and can cause embarrassment, as they often require advance setup and are limited to formal settings.

Method used

A system that detects filler speech in real-time using audio signals, providing unobtrusive tactile, auditory, or visual feedback through portable devices like smartphones or watches, allowing speakers to correct their speech habits discreetly.

Benefits of technology

Enables immediate, unobtrusive feedback to reduce filler speech during any conversation, maintaining privacy and effectiveness without distracting others, thus improving speech behavior.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026502393000001_ABST
    Figure 2026502393000001_ABST
Patent Text Reader

Abstract

A system and method for identifying filler speech and providing real-time feedback to a speaker for correction. The system receives a live audio signal from a user's speech, analyzes the audio signal for filler speech, authenticates the identity of the speaker, and, if filler speech is detected from an authenticated user, provides subtle sensory feedback to make the speaker aware of this behavior so that it can be corrected in real time, with the added benefit of helping the speaker improve their speech behavior by eliminating the use of filler speech.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a system and method for improving speech behavior and reducing the use of filler words such as "like" or "uh, you know" using unobtrusive real-time feedback to speakers during any face-to-face conversation that is not feasible with current alternatives. [Background technology]

[0002] The use of hesitations or "filler speech" is common in informal conversations and professional presentations. Examples of filler words include "like," "you know," and "really," while examples of filler sounds include "uhm," "uh," and forced laughter or giggles. While filler speech may be acceptable in some contexts, excessive use of filler speech can have negative consequences, such as distracting the listener from what the speaker is saying, reducing the speaker's credibility, or creating the undesirable impression that the speaker is inexperienced, nervous, unprepared, and / or unsure of themselves. Therefore, the speech behavior of individuals who use excessive filler speech needs to be modified. However, speakers who frequently use excessive filler speech often do so unconsciously and habitually as a learned behavior and are therefore less well-suited to modifying their own behavior. Identifying and publicly informing the speaker of a speaker's behavior risks social stigma or embarrassment, especially in front of the speaker's peers.

[0003] Current solutions for correcting undesirable speech behavior include the use of personal computers and display screens to provide speakers with visual feedback; recording speech, counting filler words, and displaying statistics after completion of speaking; and counseling by speech coaches. While these traditional approaches can prove useful, they generally require advance planning and setup (e.g., placement of a display screen), are available only at limited times (e.g., during formal speeches or video conferences), and are ineffective for providing real-time feedback during any face-to-face conversation, where feedback is most needed and effective.

[0004] Therefore, there is a need for a means to improve speech during any face-to-face conversation by reducing the use of excessive filler speech and by providing unobtrusive feedback to the speaker without causing embarrassment or discomfort. Summary of the Invention [Means for solving the problem]

[0005] The present invention includes a system and method for reducing undesirable speech behavior by detecting the occurrence of filler speech and providing unobtrusive, real-time feedback to the speaker during any face-to-face conversation for correction of the undesirable behavior.

[0006] A system according to the present invention includes a processor, a memory accessible by the processor, and programmed instructions and data stored in the memory and executable by the processor. The system is configured to receive an audio signal at an input device, preferably a smartphone, based on the speech of an individual within proximity of the system; evaluate the audio signal for the presence of one or more instances of filler speech; authenticate the identity of the speaker and determine whether the audio signal originates from the speech of a target user of the system; and, upon detecting the presence of at least one instance of filler speech and confirming that the audio signal originates from an authenticated target user of the system, output to the target user a subtle sensory signal, preferably a tactile signal using a smartphone, watch, or other portable device, that notifies the target user of the detection of filler speech. The system is configured to output the subtle sensory signal in the form of one or more of a tactile signal, an auditory signal, and a visual signal. The real-time sensory signal subtly alerts the speaker to the use of undesired filler speech, allowing the speaker to practice competing responses, such as pauses, in place of filler words without others noticing the signal. The system may also record a history of detected filler speech for review by the user.

[0007] The system is configured to evaluate the audio signal for detection of filler speech in the form of filler words and filler sounds, and upon detecting the presence of at least one instance of a filler word or filler sound, confirm the presence of at least one instance of filler speech in the audio signal. The audio signal is evaluated for the presence of filler words using a text classification model and for the presence of filler sounds using an acoustic classification model.

[0008] The text classification model converts the audio signal into a text transcript and uses a text search of the text transcript for words that match a predetermined list of filler words, the predetermined list of filler words being stored in memory. Optionally, the system may be further configured, upon detecting a filler word in the text transcript, to identify surrounding words proximate to the detected filler word and determine, based on the surrounding words, whether the detected filler word was used in an appropriate non-filler context. The system may also be configured to allow a user to selectively update the predetermined list of filler words to remove or add words that the system uses to detect filler words.

[0009] The acoustic classification model compares waveforms of the audio signal with waveforms of predetermined filler sound sound files and determines whether filler sounds are present in the audio signal based on one or more matching waveforms. The predetermined filler sound sound files are stored in memory, and the system is configured so that a user can selectively update the predetermined filler sound sound files to remove or add sound files, thereby adding or removing sounds that the system uses to detect filler sounds.

[0010] The system is configured to authenticate the identity of a speaker by comparing the waveform of the audio signal with one or more waveforms from a sound file of a user voice recording, the sound file of the user voice recording being stored in memory.

[0011] Both the foregoing general description and the following detailed description are exemplary and explanatory only and are intended to provide further explanation of the invention as claimed. The accompanying drawings are included to provide a further understanding of the invention, are incorporated in and constitute a part of this specification, illustrate embodiments of the invention, and together with the description, serve to explain the principles of the invention.

[0012] Further features and advantages of the present invention can be seen from the following detailed description, taken in conjunction with the drawings described below. [Brief explanation of the drawings]

[0013] [Figure 1] 1 illustrates an example process for detecting the use of filler speech and providing real-time feedback informing about the use of filler speech. [Figure 2] An example of a system for use in carrying out the process of FIG. 1 is shown. [Figure 3] 2 shows the results of a first example of a text classification model in the process of FIG. 1. [Figure 4] 2 shows the results of a second example of a text classification model in the process of FIG. 1. [Figure 5] 2 shows the results of an example acoustic classification model in the process of FIG. 1. DETAILED DESCRIPTION OF THE INVENTION

[0014] The following disclosure will describe the invention with reference to examples illustrated in the accompanying drawings, but it is not intended to limit the invention to these examples.

[0015] Any examples provided herein, or the use of exemplary language (e.g., "such as"), are intended merely to better illustrate the invention and do not limit the scope of the invention unless otherwise asserted. No language in the specification should be construed as indicating any non-claimed element as essential or otherwise critical to the practice of the invention unless the context makes clear otherwise.

[0016] As used herein, the singular forms "a," "an," and "the" include plural referents unless the context clearly dictates otherwise. The term "or" is to be understood as a non-exclusive "or" unless the context indicates otherwise. Terms such as "first," "second," and "third," used to describe multiple devices or elements, are used only to convey the relative action, positioning, and / or functionality of the separate devices and do not require either a particular order of such devices or elements or a particular quantity or ranking of such devices or elements.

[0017] As used herein, the word "substantially" as used with respect to any characteristic or circumstance refers to a degree of deviation that is small enough so as not to significantly impair the identified characteristic or circumstance. The precise degree of deviation that is permissible in each situation will depend on the specific context, as will be understood by those skilled in the art.

[0018] It will be understood that as used herein, the terms "comprises" and / or "comprising", except as otherwise indicated herein or clearly contradicted by context, specify the presence of stated features, integers, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.

[0019] As used herein, the term "filler speech" will be understood to encompass filler words, filler sounds, hesitations, and other undesirable speaking patterns and behaviors, unless the context makes clear otherwise.

[0020] Unless otherwise indicated, recitations of ranges of values ​​herein serve as shorthand for referring individually to each separate value falling within each range, including the range endpoints, each separate value within the range, and all intermediate ranges subsumed throughout the range, each of which is incorporated herein as if individually set forth herein.

[0021] Unless otherwise indicated or clearly contradicted by context, the methods described herein may be carried out with individual steps performed in any suitable order, including the exact order disclosed, with no intermediate steps or with one or more additional steps intervening between the disclosed steps; the disclosed steps performed in an order other than the exact order disclosed; one or more steps performed simultaneously; and one or more disclosed steps omitted.

[0022] The present invention includes systems and methods for reducing the use of filler speech and providing speakers with subtle, real-time feedback for correction. Undesirable speaking patterns and behaviors may include excessive use of filler words, filler sounds, and / or hesitations, but may also include the use of socially unacceptable words and / or phrases (e.g., abusive language), negative self-talk (e.g., self-degrading words or phrases), and patterns / behaviors that deviate from pre-established norms (e.g., speaking too fast or too slow compared to a predetermined speech rate, e.g., a speech rate based on a predetermined words per minute rate).

[0023] Figure 1 shows an example process 100 for detecting the presence of filler speech and providing real-time feedback to a user to modify their speech behavior. Figure 2 shows an example system 200 for performing process 100.

[0024] Generally, process 100 includes a first step 102 of receiving an audio signal representing an individual's speech. The audio signal is stored in a temporary memory (e.g., an audio buffer) of system 200 while processor 204 performs three separate evaluations of the audio signal, including [1] a filler word evaluation for detecting filler words, [2] a filler sound evaluation for detecting filler sounds, and [3] a speaker authentication evaluation for confirming the authenticated target speaker. In the illustrated example, the three evaluations [1]-[3] are performed in parallel, but in other examples, the three evaluations may instead be performed sequentially, one after the other, and in any desired order. Optionally, system 200 may be configured to allow a user to disable the speaker authentication evaluation so that process 100 requires only the filler word evaluation and the filler sound evaluation.

[0025] In filler word evaluation, in step 104, the audio signal is converted into a text transcript by one or more speech-to-text models. A speech-to-text model has two parts: an acoustic model and a language model. First, the acoustic model takes the audio as input and converts it into probabilities for letters of the alphabet. Next, the language model helps turn these probabilities into words in a logical language. The language model assigns probabilities to words and phrases based on statistics from training data. Suitable speech-to-text algorithms include, but are not limited to, the speech recognition and transcription programs available through Apple, Inc.'s iOS mobile operating system and Google, LLC's Android operating system.

[0026] In step 106, the text transcript is then searched for filler words. A list of filler words may be stored in memory 208 as one or more text files, and processor 204 may communicate with memory 208 to identify filler words to search for in the text transcript. Filler words may be any words that are predetermined to be commonly used to fill pauses or breaks in speech without meaningfully contributing to the content of the speech. Examples of filler words include, but are not limited to, "like," "you know," and "really."

[0027] In some examples, memory 208 may be preloaded with a predetermined list of filler words before they are provided to the user, and the user may subsequently modify the stored list of filler words. For example, an extensive list of words predetermined as suggested filler words may be stored in memory 208, and the user may interact with system 200 via input device 202 to select and / or deselect specific words from the predetermined list that they want system 200 to treat as filler words. In some examples, the user may add words to the list of predetermined words, which may be done via input device 202 in the form of either typed text input or speech-to-text input. By allowing the user to edit the list of filler words, the user can customize system 200 to identify speech behaviors unique to individual users and provide feedback based on the speech behaviors.

[0028] In step 108, the processor 204 determines a word count of the filler words identified in the text transcript of the audio signal. If the text search is completed without identifying any filler words in the text transcript, the word count is set to a zero count (i.e., no filler words are present). If the text search is completed with one or more filler words found in the text transcript, the word count is set to a non-zero count, which may be done by recording the presence of at least one filler word or by recording the exact number of filler words detected in the text transcript. In some examples, the word count may be set to the exact number of filler words detected in the text transcript (e.g., 10 filler words detected) and may further identify a count of each specific filler word detected (e.g., 1 count of "really," 2 counts of "you know," and 7 counts of "like").

[0029] In some examples, there may be an optional step 110 in which processor 204 performs a contextual review of the results from the text search to determine whether any of the filler words (if any) detected by the text search were actually used in an appropriate context (i.e., not as filler words). For example, the word "like" may be the most commonly used filler word, but it also has appropriate non-filler uses, such as when used to convey a similarity between two separate objects or events. Contextual review may be used to identify appropriate uses of such filler words to avoid false positive reports of filler speech.

[0030] If step 110 relating to context review is included, the list of filler words stored in memory 208 may include additional information identifying listed filler words that also have appropriate contextual usage. Once the text search is completed with the discovery of one or more filler words, processor 204 may communicate with memory 208 to determine whether any of the detected filler words are known to have appropriate contextual usage. If one or more detected filler words are identified as having appropriate contextual usage, processor 204 performs a further search of the text transcript for each occurrence of those detected filler words with appropriate contextual usage to identify the number of words preceding and following each occurrence of each such filler word. For example, if one or more occurrences of the word "like" are detected in the text transcript, processor 204 searches the text transcript for each occurrence of the word "like" and further identifies the number of words preceding and following each such occurrence. Contextual review may be performed using identification of any number of preceding and following words, including, for example, three preceding and following words, five preceding and following words, or ten or more preceding and following words.

[0031] Upon identifying surrounding words (preceding and following) for an occurrence of a filler word, processor 204 uses a text classification model to determine whether the surrounding words indicate whether the corresponding occurrence of the filler word was in fact an appropriate contextual usage or a filler usage. If processor 204 determines that one or more occurrences of the detected filler word were appropriate contextual usages, processor 204 updates the results from the text search to decrease the word count of the corresponding detected filler word by the number of appropriate contextual usages identified for that filler word so that the updated word count (step 108) more accurately represents the true count of filler usages for the corresponding filler word.

[0032] For filler sound assessment, in step 114, processor 204 generates a waveform of the audio signal, and in step 116, processor 204 uses an acoustic classification model to determine whether the audio signal contains filler sounds. Suitable acoustic classification models include, but are not limited to, the CoreML model available from Apple, Inc. or the TensorFlow Lite ML model available from Google, LLC. The acoustic classification model is trained on a dataset of common filler sounds such as "ums," "uhs," and giggles or laughter, as well as other speech and background sounds.

[0033] The collection of predetermined filler sounds may be stored in memory 208 as one or more sound files, and processor 204 may communicate with memory 208 to perform a waveform analysis in which one or more waveforms generated from the audio signal are compared with the waveforms of the stored filler sounds to identify the presence of filler sounds in the received audio signal. Positive identification of a filler sound may be conditioned on a predetermined confidence level based, for example, on a minimum matching percentage (e.g., 75%) between the waveforms of the audio signal and the waveforms of the stored filler sounds.

[0034] In some examples, memory 208 may be preloaded with a predetermined list of filler sounds before being provided to the end user, and the user may subsequently modify the list of saved filler sounds. For example, system 200 may be provided with an extensive list of sounds that have been predetermined to be known filler sounds, and the user may interact with system 200 via input device 202 to select and / or deselect particular sounds from the predetermined list that they want system 200 to treat as filler sounds. In some examples, the user may add additional sounds to the list of predetermined filler sounds, which may be done via audio input device 202. By allowing the user to edit the list of filler sounds, the user can customize system 200 to identify speech behaviors that are unique to individual users and provide feedback based on the speech behaviors.

[0035] In step 118, processor 204 determines a sound count for filler sounds identified in the audio signal. If processor 204 does not identify a filler sound in the audio signal, the sound count is set to a zero count (i.e., no filler sounds are present). If processor 204 identifies one or more filler sounds in the audio signal, the sound count is set to a non-zero count, which may be done by recording the presence of at least one filler sound or by recording the exact number of filler sounds detected. In some examples, processor 204 may identify the exact number of filler sounds detected (e.g., five filler sounds detected) and may further identify a count for each particular filler sound detected (e.g., one count for "uhm," two counts for "uh," and three counts for "giggles").

[0036] In the speaker authentication assessment, in step 124, the processor 204 analyzes the received audio signal to generate one or more waveforms of the audio signal. In step 126, the processor 204 compares the waveforms generated from the audio signal with one or more pre-stored waveforms to determine whether the audio signal contains speech from an authenticated speaker who is a target user for speech behavior modification, thereby authenticating the speaker's identity using a voice matching algorithm. The waveforms used for speaker authentication and sound classification may include one or more mel spectrograms. Suitable voice matching algorithms include, but are not limited to, Pytorch or the NeMo Speaker Recognition model available from Nvidia Corporation.

[0037] The voice matching algorithm is trained on a dataset of speech files recorded in system 200 before being provided to the end user. During system initialization and setup, the end user creates one or more unique recorded speech files for comparison purposes. The unique recorded speech files can be updated at any time to more accurately reflect the speaker's hearing environment. For example, a user may record multiple speech files recording their own speech in several different environments and / or situations (e.g., low ambient noise environment, medium ambient noise environment, high ambient noise environment, one-on-one conversation, public presentation, large social gathering, etc.), and system 200 can be adapted to recognize the environment and / or situation and select the corresponding speaker's voice recording for use in speaker authentication assessment. System 200 can also be adapted to proactively allow the user to select the environment and / or situation. Preferably, the recorded speech files contain utterances spoken by the user based on prompts provided by system 200, which is programmed to prompt the user with specific utterances predetermined to be most useful for use in verifying the speaker's identity and identifying common filler speech. For example, the system 200 may ask the user to repeat utterances into the audio input device 202, which may include utterances that mimic common filler words and filler sounds.

[0038] The recorded speech files may be stored in memory 208 as one or more sound files, and processor 204 may communicate with memory 208 to compare one or more waveforms of the audio signal with one or more waveforms from the stored speech files to determine whether the audio signal contains speech originating from a speaker authenticated as a target user of the system for speech behavior modification. Positive identification of an authenticated speaker may be contingent on a predetermined confidence level based, for example, on a minimum matching percentage (e.g., 75%) between the waveforms of the audio signal and the waveforms of the stored speech files. Comparing the audio signal waveforms to the speech file waveforms may include assessing the cosine similarity of the waveforms.

[0039] In some examples, in addition to system 200 prompting the user to train the voice matching algorithm by repeating a predetermined utterance in response to the audio input, the user may provide additional utterances of their choice to further train the voice matching algorithm. In some examples, training of the voice matching algorithm may continue after system initialization and setup, as well as during operation, so that the voice matching algorithm may be continually improved to increase its accuracy in correctly identifying speakers as target users as the system continues to receive additional user audio input during any real-time conversation.

[0040] In step 128, the processor 204 determines whether the audio signal contains speech originating from an authenticated speaker who is a target user of the system 200 using either a positive identification setting (e.g., 1, Yes, True, etc.) or a negative identification setting (e.g., 0, No, False, etc.).

[0041] In step 130, the processor 204 verifies the occurrence of filler speech by the authenticated speaker. If the word count from the text search confirms the presence of one or more filler words and / or the sound count from the audio signal analysis confirms the presence of one or more filler sounds, the processor 204 confirms the presence of filler speech. If the system 200 is adapted to identify specific occurrences of filler words and / or filler sounds, the processor 204 may save this additional information in the memory 208. If the identification setting from the speaker authentication assessment confirms that the audio signal contains speech from an authenticated speaker, the processor 204 confirms that the filler speech originated from the target user of speech behavior modification. If the processor 204 confirms that there was at least one instance of filler speech with positive speaker authentication of the target user, the process proceeds to step 132, where the processor 204 triggers the sensory signal transmitter 206 to provide real-time feedback to the user notifying them of the filler speech by outputting a subtle sensory signal to the target user. If the processor 204 determines that there was no instance of filler speech or that there was a negative speaker identification in the absence of the target user, the process ends at step 134 without sensory signal feedback to the user.

[0042] Optionally, system 200 may be configured to allow a user to selectively override speaker authentication assessment. In such a scenario, process 100 omits the steps for speaker authentication assessment (steps 124, 126, 128), and step 130 only requires confirmation that at least one instance of filler speech was present for process 100 to proceed to step 132, where processor 204 triggers sensory signal transmitter 206 to provide feedback to the user. Without being bound by theory, speaker authentication can be difficult in environments with high levels of background noise (e.g., large crowds, loud machinery, etc.). In such environments, a user may selectively override speaker authentication assessment (which may provide more reliable identification of filler speech in such environments) so that system 200 can more reliably identify occurrences of filler speech without requiring confirmation of the speaker's identity.

[0043] In some examples, system 200 may output a sensory signal to the target user for each detected instance of filler speech by the target user. In other examples, system 200 may output a sensory signal to the target user only after detecting a predetermined number of instances of filler speech (e.g., three occurrences, five occurrences, seven occurrences, etc.) within a predetermined period of time (e.g., 10 seconds, 30 seconds, 60 seconds, etc.). In some examples, system 200 may be adapted to output a predetermined number of sensory signals (e.g., one or more individual signal pulses), and in other examples, processor 204 may instruct sensory signal transmitter 206 to output a sensory signal for a predetermined duration (e.g., repeating pulses continuously for 30 seconds), with the option for the user to interact with system 200 to trigger early termination of the sensory signal (e.g., by triggering an alert termination switch) before the completion of the predetermined duration.

[0044] Optionally, system 200 may be configured to allow a user to selectively set the number of occurrences of filler speech before a sensory signal is output, the period of time within which a predetermined number of occurrences of filler speech must be detected before a sensory signal is output, and / or the duration for which a sensory signal may be repeatedly output. Enabling selective setting of these parameters allows a user to tailor system 200 to provide feedback that is most helpful in correcting a particular speech behavior. For example, if a user's frequency of filler speech is relatively high, it may be undesirable for system 200 to provide sensory feedback for every occurrence of filler speech because a large number of sensory signals may be distracting to the user, whereas a single sensory signal after multiple occurrences of filler speech may be sufficient to alert the user to correct their speech behavior. Without being bound by theory, it is also expected that when system 200 is configured to require multiple occurrences of filler speech within a predetermined period of time, there will be fewer instances of false-positive sensory feedback based on the detection of one or more filler words actually used in the appropriate context (e.g., the word "like" used in a comparison context).

[0045] The system 200 is adapted to provide sensory signal feedback as an unobtrusive sensory signal that is easily perceptible by the target user, but relatively reduced or imperceptible to others. Examples of sensory signals that may be used with the present invention include, but are not limited to, tactile signals (e.g., mechanical vibrations or electrical stimulation), auditory signals, and visual signals (e.g., light emitted from a light source). Examples of sensory signal devices for generating unobtrusive sensory signals include, but are not limited to, mobile phones, wristwatches, earphones, and any other electronic devices that may be worn or carried in close proximity to the user's body.

[0046] In one example, a wristwatch may provide a user with a subtle sensory signal in the form of a tactile signal perceptible only to the user through skin contact with the wristwatch. In another example, electronic earphones may be used to generate a sensory signal to the user in the form of a low-intensity auditory signal perceptible only to the user through the proximity of the earphones to the user's ear canal. In a further example, a light source on the inside surface of eyeglasses may be used to generate a visual sensory signal to the user in the form of a low-intensity light signal perceptible only to the user through the proximity of the eyeglasses to the user's eyes. In yet another example, the screen of a phone, tablet, laptop, or other such consumer electronic device may emit a subtle flash to provide a visual sensory signal to the user. In general, any small electronic device can be adapted to provide one or more subtle sensory signals to the user, provided that it is always in close proximity to the user's body. For example, a mobile phone may be programmed to generate a detectable tactile signal based on the proximity of the mobile phone to the user's body while the phone is being carried; clothing and neck jewelry may be adapted to generate a tactile signal based on body contact; ear jewelry may be adapted to generate a tactile and / or auditory signal based on body contact and / or proximity to the ear canal, etc.

[0047] In some examples, system 200 can be adapted to provide additional functionality beyond simply providing a sensory signal. For example, when using a sensory signal device with a relatively more robust computing system (e.g., a smartphone or smartwatch), the system can be adapted to generate a usage report that identifies specific filler speech identified as the basis for triggering the sensory signal, the number of occurrences of each identified specific filler speech, and / or transcripts of audio signals identified as containing filler speech. In some examples, the system can also provide historical data reporting the number of occurrences of identified filler speech over a long period of time, along with counts of occurrences in individual instances within that long period of time (e.g., a count of the total number of times the word "like" was said in the past 30 days, along with a count of the number of times it was said on each individual day within that 30-day period).

[0048] 2 provides an exemplary block diagram of a system 200 in accordance with the present invention. System 200 may include one or more processors 204 in the form of (CPUs) 204A-204N, input / output circuitry 202, memory 208, and a sensory signal transmitter 206. Optionally, system 200 may further include a network adapter for connecting system 200 to a network, such as the Internet, for transmission of data, which may include downloading software updates and / or uploading user data to remote storage.

[0049] The processor 204 executes program instructions to perform the functions of the present invention and may be implemented as one or more microprocessors, microcontrollers, or processors within a system-on-chip. FIG. 2 illustrates an example in which the system 200 is implemented as a single multiprocessor system, in which multiple processors 204A-204N share system resources such as memory 208, input / output circuitry 202, and sensory signal transmitter 206. In other examples, the system 200 may be implemented as a single processor system. The input / output circuitry 202 provides the ability to input data into or output data from the system 200. For example, the input / output circuitry may include input devices such as a microphone, sensor, keypad, touchscreen, etc.; output devices such as a speaker and display screen; and / or input / output devices providing combined functionality of one or more of the aforementioned input and output devices. The sensory signal transmitter 206 may be any device for transmitting a subtle sensory signal to a target user to notify the user of a detected instance of filler speech, which may include all such devices described herein.

[0050] The memory 208 stores program instructions executed by the processor 204, as well as data used and processed by the processor 204, to perform the functions of the system 200. The memory 208 may include electronic memory devices such as, for example, random access memory (RAM), read-only memory (ROM), programmable read-only memory (PROM), electrically erasable programmable read-only memory (EEPROM), flash memory, etc. The memory 208 may include a sensor data acquisition routine 210, a signal processing routine 212, a data processing routine 214, and stored data such as signal data 216, manifold data 218, classification data 220, and an operating system 222.

[0051] The sensor data capture routines 210 may include routines for receiving and processing sensor input data, such as capture of a user's speech at the input device 202, to form signal data 216. The signal processing routines 212 may include routines for processing the signal data 216, such as text classification models and acoustic classification models, to form aggregate data 218 (e.g., conclusions regarding the detection of filler words and filler sounds, as described above). The data processing routines 214 may include routines for processing the aggregate data 218 for operation of the system 200 (e.g., instructing a sensory signal transmitter based on the detection of filler speech). The classification data 220 may include saved lists of filler words, sound files of filler sounds, and voice recordings of authenticated speakers. The operating system 222 provides overall system functionality.

[0052] An example was created by training two text classification models to detect filler words and an acoustic classification model to detect filler sounds using the CreateML tool from Apple, Inc. The two text classification models were trained to detect occurrences of the filler word "like" and to determine whether each occurrence was the use of the word as filler speech or in an appropriate context.

[0053] The first text classification model used the "Transfer Learning BERT Embeddings" algorithm for Apple iOS (versions 17 and above), and the second text classification model used the "Conditional Random Field" algorithm for Apple iOS (versions prior to 17). These models were trained using a natural language model to remove punctuation and segment the text transcript into words. They were then trained using ChatGPT to generate paragraphs in which the word "something like" was used as a filler word and paragraphs in which the word "something like" was used in its appropriate context. The paragraphs were then converted into an audio signal using TTSMaker, a text-to-speech program, and the text classification model detected all occurrences of the word "something like" and, for each occurrence, captured the three preceding and three following words, forming sentences that were subsequently used for further training. Figure 4 shows the results of the first model, which achieved 95.6% accuracy (training and validation) after 10 iterations. Figure 5 shows the results of the second model, which achieved 95.7% accuracy (training and validation) after three iterations.

[0054] An acoustic classification model for detecting filler sounds was trained from a dataset of collected audio samples of multiple filler sounds (e.g., um, uh, and giggles / laughter) as well as a dataset of background noise and speech. The background noise and speech datasets were collected from a collection of sound samples from the Columbia University Sound Sample Database, which is publicly available for use as "background noise" in composite audio signals to simulate real-world situations, and from a collection of sound samples from Pixabay, a stock media website, which is publicly available for use as ambient sound effects in composite audio signals to simulate real-world situations. All collected sound samples were mono channel, 16 kHz, and 1 second in duration. The datasets were labeled and trained with an acoustic classification model with the following parameters: Feature extractor: Audio feature print, Iterations: 55, Window duration: 0.5, Window overlap: 25%. This model was trained using assemblyai_ to determine the occurrence of filler sounds in audio samples. A timestamp function was used to identify the start and end of individual words in the audio signal, segmenting the audio signal into separate segments, from which individual occurrences of filler sounds such as "Um" and "Uh" were then identified. Figure 3 shows the results of the acoustic classification model, where accuracy (training and validation) quickly exceeded 95% in the early iterations and converged to an accuracy range of 97.9% (validation) to 100% (training) after 54 iterations.

[0055] The systems and methods herein provide significant benefits, including real-time feedback for immediate identification of instances of undesirable speaking behavior; continuous, uninterrupted monitoring and feedback throughout the day, including during any face-to-face conversation; real-time feedback to the target user in an unobtrusive manner that is not embarrassing; feedback that is perceptible to the target user with minimal distraction or diverting the user's attention; and feedback that is not distracting, intrusive, or perceptible by others. By combining these benefits, the present invention is expected to be far more effective in correcting undesirable speech behavior than is possible with conventional approaches that do not possess such advantages.

[0056] While the present invention has been described with reference to particular embodiments, it will be understood by those skilled in the art that the foregoing disclosure deals with exemplary embodiments only; that the scope of the invention is not limited to the disclosed embodiments; and that the scope of the invention may encompass additional embodiments encompassing any combination of all or part of the disclosed embodiments, as well as various changes and modifications to the examples disclosed herein, without departing from the scope of the invention as defined in the appended claims and their equivalents.

[0057] As an example, while the foregoing disclosure and accompanying figures address system 200 in the form of a single device, it will be understood that system 200 may be provided in the form of a multi-device having two or more devices. A two-device system may be provided in which a first device performs a first portion of the functionality (e.g., audio input, processing, analysis, and filler speech detection), a second device performs a second portion of the functionality (e.g., providing a subtle sensory signal), and the two devices are in remote communication with each other (e.g., signal transmitters in both devices for the first device to instruct the second device when to provide the sensory signal). A three-device system may be provided in which a first device performs a first portion of the functionality (e.g., audio input and processing), a second device performs a second portion of the functionality (e.g., signal analysis and filler speech detection), and a third device performs a third portion of the functionality (e.g., providing a subtle sensory signal), and the three devices are in remote communication with each other (e.g., signal transmitters in all three devices to transmit signals to each other). Any other number of multi-device systems having four or more devices may also be provided. Such a multi-device system may be preferable when system 200 is configured to provide unobtrusive sensory signals via particularly small sensory signal devices (e.g., earphones, clothing accessories, etc.), as the division of functional systems may allow for a more compact construction of one or both of the audio input receiving device and / or sensory signal generating device, which may further facilitate positioning of those devices in close proximity to the user's body, while allowing for complex computational work to be performed via a more robust remote server (e.g., a cloud server).

[0058] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. Each block in the flowcharts or block diagrams may represent a module, segment, or portion of instructions, which includes one or more executable instructions for implementing one or more specified logical functions. In some implementations, the functions noted in the blocks may occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently, or the blocks may sometimes be executed in the reverse order, depending on the functionality involved. It should also be noted that each block in the block diagrams and / or flowchart diagrams, and combinations of blocks in the block diagrams and / or flowchart diagrams, may be implemented by a special-purpose hardware-based system that performs the specified functions or acts or executes a combination of special-purpose hardware and computer instructions.

[0059] To the extent necessary to understand or complete the disclosure of the present invention, all publications, patents, and patent applications mentioned in this specification are expressly incorporated by reference herein to the same extent as if each were individually incorporated.

[0060] The present invention is not limited to the exemplary embodiments set forth herein, but is instead characterized by the appended claims, which in no way limit the scope of the present disclosure.

Claims

1. 1. A system for providing speech-related feedback, comprising: a processor; a memory accessible by the processor; and programmed instructions and data stored in the memory and executable by the processor, whereby the system: receiving an audio signal at an input device based on speech of an individual within proximity of the system; evaluating the audio signal for the presence of one or more instances of filler speech; authenticating the identity of the speaker and evaluating the audio signal to determine whether the audio signal originates from the speech of a target user of the system; upon detecting the presence of at least one instance of filler speech and verifying that the audio signal originates from an authenticated target user of the system, outputting a subtle sensory signal to the target user, notifying the target user of the detection of the filler speech; A system configured to:

2. 10. The system of claim 1, wherein the system is configured to evaluate the audio signal for detection of filler speech in the form of filler words and filler sounds, and determine the presence of at least one instance of filler speech upon detecting the presence of at least one instance of either one or more filler words, one or more filler sounds, or one or more hesitant speaking patterns or behaviors.

3. 3. The system of claim 2, wherein the system is configured to evaluate the audio signal for detection of filler words using a text classification model and to evaluate the audio signal for detection of filler sounds using an acoustic classification model.

4. The system of claim 1 , wherein the system is configured to evaluate the audio signal for detection of filler speech in the form of filler words.

5. The system of claim 4 , wherein the system is further configured to determine a context of the detected filler word to determine whether the detected filler word was used in an appropriate non-filler context.

6. The system, for filler word detection, converting the audio signal into a text transcript; performing a text search of the text transcript for words that match a predetermined list of filler words, the predetermined list of filler words being stored in the memory; The system of claim 4 , configured to evaluate the audio signal by:

7. 7. The system of claim 6, wherein the system is further configured, upon detecting a filler word in the text transcript, to identify surrounding words proximate to the detected filler word and determine, based on the surrounding words, whether the detected filler word was used in an appropriate non-filler context.

8. The system of claim 6 , wherein the system is configured to allow a user to selectively update the predetermined list of filler words to remove or add words.

9. The system of claim 1 , wherein the system is configured to evaluate the audio signal for detection of filler speech in the form of filler sounds.

10. 10. The system of claim 9, wherein the system is configured to evaluate the audio signal for filler sound detection by comparing a waveform of the audio signal to waveforms of sound files of predetermined filler sounds, the sound files of the predetermined filler sounds being stored in the memory.

11. The system of claim 10 , wherein the system is configured for a user to selectively update the sound files of the predetermined filler sounds to remove or add sound files.

12. 2. The system of claim 1, wherein the system is configured to authenticate a speaker's identity by comparing one or more waveforms of the audio signal with one or more waveforms from a sound file of a user voice recording, the sound file of the user voice recording being stored in the memory.

13. The system of claim 1 , wherein the system is configured to output a subtle sensory signal in the form of one or more of a tactile signal, an auditory signal, and a visual signal.

14. The system of claim 1 , wherein the system is configured to record a history of detected filler speech for review by a user.