Systems and methods for human speech verification
Patent Information
- Application Number
- PCT/US2024/049333
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-05-09
- Filing Date
- 2024-09-30
- Publication Date
- 2025-05-08
AI Technical Summary
The increasing prevalence of deepfake technology and AI-generated content poses a significant challenge in verifying the authenticity of human speech, leading to concerns about disinformation, fraud, and the inability to distinguish between human and synthetic media.
A system that combines traditional microphone technology with radar sensors to verify human speech by correlating acoustic signals with physiological signals, providing real-time authentication and linking the certificate of authenticity to the media file.
This solution effectively verifies the authenticity of human speech, preventing the spread of misinformation and enhancing security by ensuring that audio content is generated by a human, rather than AI or deepfakes.
Smart Images

Figure US2024049333_08052025_PF_FP_ABST
Abstract
Description
Attorney’s Docket No.: 055743-816976 (M24-042P-US1-b) SYSTEMS AND METHODS FOR HUMAN SPEECH VERIFICATION CROSS REFERENCE TO RELATED APPLICATIONS
[0001] This is a non-provisional application that claims benefit to U.S. provisional application serial number 63 / 541,773 filed on September 29, 2023, U.S. provisional application serial number 63 / 541,778 filed on September 29, 2023, and to U.S. provisional application serial number 63 / 645,081 filed on May 9, 2024; all of which are herein incorporated by reference in their entireties. FIELD
[0002] The present disclosure generally relates to technologies for distinguishing human users from automated systems or bots, and particularly to a microphone-based system that can certify on a recording device that a recorded acoustic signal originates from a human who is speaking into a microphone.
[0003] With the rise of AI and deepfake technology, we find ourselves in an era where distinguishing between real and synthetic media is becoming increasingly more challenging. The year 2023 was an inflection point for deepfakes. Consider the following three incidents:
[0004] In June of 2023, a deepfake of Russian President Vladimir Putin was broadcast on Russian television and radio. In the video, “Putin” declared martial law and stated that Ukraine’s army had invaded Russia. Authorities later said that some broadcast channel streams had been hacked, which allowed the deepfake to be broadcast on the national airwaves. This incident underscores the potential of deepfakes to propagate disinformation, impersonate global leaders, and incite panic or conflict.
[0005] In April of 2023, a song called “Heart on My Sleeve” that uses AI to simulate the music of pop stars Drake and The Weeknd was published and went viral. The song was so convincing that some people even thought it surpassed the real pop stars’ talents1. However, the popularity and revenue-earning potential of AI- generated songs have put music industry gatekeepers on guard. Universal Music Group, the label owner of Drake and The Weeknd, immediately invoked copyright violation to get the platforms to take “Heart on My Sleeve” down1. There have been several other examples of AI-manipulated songs since then. The legal ramifications98974549.21Attorney’s Docket No.: 055743-816976 (M24-042P-US1-b) concerning music and copyright law will continue to dominate the debate on AI covers in the music industry.
[0006] In April of 2023, a mother in Arizona received a scam call that used AI voice cloning technology to mimic her daughter’s voice. The caller claimed to have kidnapped her daughter and demanded ransom. The mother was initially convinced that it was her daughter’s voice on the phone. This incident highlights the dangers of AI deepfakes and voice cloning technology, which can be used to deceive people and commit fraud.
[0007] Deepfakes pose significant societal challenges, including the potential to spread disinformation, incite panic or conflict, and undermine trust in media. They can impersonate global leaders, leading to geopolitical instability, as seen in the Putin incident. In the music industry, AI-generated songs can infringe on copyright laws, disrupt revenue streams, and challenge traditional gatekeepers. Furthermore, voice cloning technology can be used for malicious purposes such as fraud and deception, as demonstrated by the scam call incident in Arizona.
[0008] Unfortunately, these unsettling incidents expose a significant gap in the digital landscape. We lack a reliable way to ensure the authenticity of the voices we hear as human, to verify that there was indeed a human behind the microphone moving their lips that produced the recorded audio. Furthermore, there is a pressing need to make it accessible to everyone. Currently, for any piece of media – whether it’s human-generated, AI-generated, or manipulated media, the moment that source file is saved and uploaded to a distribution center (e.g., YouTube, social media, TV broadcast), it is disconnected from the method of creation. That is, for speech, we don’t know whether there was a human speaking on the other end of the recording device or whether the speech was partially or completely generated via AI.
[0009] Similarly, in both packet-switched and circuit-switched communication channels, the moment a voice packet leaves a talker’s device, it becomes disconnected from the talker. We have no way of knowing whether there was a human talker on the phone or whether the voice is synthetic.
[0010] The pressing question this invention aims to answer is: how can we uphold the authenticity of 'human-made' media in this age of advanced deepfake technology?98974549.22Attorney’s Docket No.: 055743-816976 (M24-042P-US1-b)
[0011] A related, but different, problem is that of distinguishing between human users and automated bots, not necessarily specific to media creation. CAPTCHA, an acronym for "Completely Automated Public Turing test to tell Computers and Humans Apart," is a technology utilized to distinguish between human users and automated systems or bots. CAPTCHAs were conceived to mitigate automated form submissions, unauthorized access, web scraping, spamming, and similar kinds of malicious activities perpetrated by computer programs on the web. This technology generally incorporates a challenge-response test where users must perform a task that, at least traditionally, is difficult for machines to execute. These tasks include identifying distorted text or images, performing simple mathematical calculations, and identifying specific objects within a visual scene or a series of images.
[0012] However, the reliability and effectiveness of existing CAPTCHA methodologies are being seriously challenged by recent advancements in Artificial Intelligence (AI) and Machine Learning (ML). Traditional CAPTCHAs, often dependent on vision-based tasks, are being bypassed by modern AI algorithms that have proven adept at image and text recognition. Optical Character Recognition (OCR) techniques have shown remarkable progress, enabling machines to read distorted texts, an ability that was once exclusive to human cognition. Similarly, Convolutional Neural Networks (CNNs), a class of deep learning neural networks, have shown exceptional performance in visual recognition tasks, thereby diminishing the efficacy of image-based CAPTCHAs.
[0013] Moreover, AI models are learning to mimic human behavior more closely, enhancing their abilities to solve puzzles and challenges that were initially designed for human intellect. AI's progress in natural language processing, predictive modeling, and cognitive computing further diminish the effectiveness of the traditional CAPTCHA system. These developments indicate that traditional CAPTCHA systems are becoming increasingly outdated and inadequate in fulfilling their primary purpose, which is differentiating human users from AI systems. Therefore, it is imperative to develop and implement new methodologies that are more resilient to rapidly evolving AI technologies.
[0014] It is with these observations in mind, among others, that various aspects of the present disclosure were conceived and developed.98974549.23Attorney’s Docket No.: 055743-816976 (M24-042P-US1-b)98974549.24Attorney’s Docket No.: 055743-816976 (M24-042P-US1-b) SUMMARY
[0015] The present disclosure provides a number of examples that describe methods and systems associated with human speech verification / authentication. In the context of the disclosed methods, devices, techniques, apparatus, systems, and so on, the terms “operable to,” “configured to,” and “capable of” used herein are interchangeable.
[0016] In a first set of illustrative examples, the inventive concept includes a method of authenticating human speech during a speech event or utterance of speech, comprising the steps of accessing data associated with utterance of speech from multiple sources, the data including a combination of a plurality of signals comprising at least one physiological signal and at least one acoustic signal; extracting a plurality of features from the plurality of signals including spectral features and temporal features from the one or more signals, the spectral features defining frequency components of the one or more physiological signals and the temporal features representing changes in the one or more physiological signals over time; and identifying an association between feature values calculated from the data over the utterance of speech to verify the acoustic signal as originating from a human and at least one physiological signal as a biosignal from the human.
[0017] In some examples, the method can further include steps of filtering the acoustic signal and physiological signal to a same frequency range, aligning the extracted features from the acoustic and physiological signals over time to accommodate correspondence between features extracted to common time frames, comparing feature variation over time between the acoustic and physiological signals for each corresponding set of extracted features from the acoustic and physiological signals, and confirming human authenticity of the speech when the acoustic and physiological signals demonstrate synchronous feature variation over time.
[0018] The foregoing examples broadly outline various aspects, features, and technical advantages of examples according to the disclosure in order that the detailed description that follows may be better understood. It is further appreciated that the above operations described in the context of the illustrative example method, device, and computer-readable medium are not required and that one or more operations may be excluded and / or other additional operations98974549.25Attorney’s Docket No.: 055743-816976 (M24-042P-US1-b) discussed herein may be included. Additional features and advantages will be described hereinafter. The conception and specific examples illustrated and described herein may be readily utilized as a basis for modifying or designing other structures for carrying out the same purposes of the present disclosure. Such equivalent constructions do not depart from the spirit and scope of the appended claims.98974549.26Attorney’s Docket No.: 055743-816976 (M24-042P-US1-b) BRIEF DESCRIPTION OF THE DRAWINGS
[0019] FIG.1 is a flow chart that shows a process of human authentication via measurement of physiological signals.
[0020] FIG.2 is a flow chart of a process for measuring physiological signals from radar.
[0021] FIG.3 is a flow chart of a process for detection of a human user based on passive monitoring via radar.
[0022] FIG.4 is a flow chart of a process for determining whether an active task is needed at the time of authentication.
[0023] FIG.5 is a flow chart of a process for verification via an active task.
[0024] FIG.6 is an illustration showing a source filter model and inverse in speech processing.
[0025] FIG.7 is a graphical representation of physiological signals by radar.
[0026] FIG.8 is an illustration and related graphical representations of a radar signal measured during human speech production.
[0027] FIG.9A is a time-frequency representation of a radar signal acquired during speaking; and FIG.9B is a time-frequency representation with the low-frequency physiological signals removed.
[0028] FIGS.10A and 10B are graphical representations showing feature values calculated over an utterance for both a microphone and a radar sensor.
[0029] FIG.11 is a graphical representation showing a normalized distribution (z-scored) of correlations when there is a mismatch between the acoustic signal and the radar signal.
[0030] FIG.12 shows radar returns for a human talker versus a loudspeaker at three playback levels.
[0031] FIG.13 is a simplified illustration showing a Data Fusion Engine for audio authentication; and
[0032] FIG.14 is a simplified diagram showing an example computing system for implementation of aspects of the systems and methods outlined herein.98974549.27Attorney’s Docket No.: 055743-816976 (M24-042P-US1-b)
[0033] Corresponding reference characters indicate corresponding elements among the view of the drawings. The headings used in the figures do not limit the scope of the claims.98974549.28Attorney’s Docket No.: 055743-816976 (M24-042P-US1-b) DETAILED DESCRIPTION
[0034] To date we have relied on listeners’ senses to verify the authenticity of media. Until recently, synthetic media sounded and looked synthetic. As the examples highlighted in the previous section demonstrate, this is no longer possible as the quality and quantity of AI-generated media is already good and rapidly improving.
[0035] The disclosed techniques for solving this have been reactive – using AI to detect AI-generated content. While none of these solutions have been productized, the idea is that a pre-trained model that can detect AI-generated media is installed at the distribution center (e.g. in YouTube servers, at the broadcast station) and every time a media file is uploaded, it is evaluated via these models to determine whether it was AI generated. There are several problems with this approach. First, AI to detect deepfakes is a cat-and-mouse game that isn’t likely to yield long-term success according to experts in the field.
[0036] Furthermore, even if the AI model at the distribution center works perfectly, the communication channel is vulnerable to stream hacks. That is, the stream between the distribution center and a user’s device can be hacked and a new audio sample streamed. Under this scenario, the user still receives manipulated media, as in the example with Putin above. To detect such attacks using reactive methods, the AI model would have to be installed on each user's device to process all incoming samples. It’s unlikely that power-limited edge devices can support persistent AI detection for all incoming media.
[0037] Various examples of a core, proactive technical solution to this problem are disclosed herein. This solution includes combination of physiological signals with acoustic signals to verify human speech. In some examples, the concept is based on new microphone technology that can certify on the recording device itself that the recorded acoustic signal originates from a human who is speaking into the microphone. This certificate can be linked to the media file and then accessed on the listener side to provide verification that a human generated the audio. The microphone technology is based on a combination of a traditional MEMS (Micro Electronic Mechanical System) microphone (or other type of microphone), which measures the acoustic signal, and a radar sensor, which measures the physiological vibrations from the human during speech production. These two98974549.29Attorney’s Docket No.: 055743-816976 (M24-042P-US1-b) signals are verified relative to each other at the time of recording to prove the personhood of the speech generator. EXAMPLE 1
[0038] A next-generation CAPTCHA system is disclosed in the present example that relies on direct measurement of physiological signals emitted from the human body to distinguish humans from automated bots.
[0039] The present system uses a radar sensor to evaluate physiological signals while a user is passive or during active task administration when necessary. Some existing websites have a button saying "Click here to prove you are human" or a similar variant, without a subsequent task. They often rely on more subtle methods to distinguish between humans and bots. For example, many operate by monitoring mouse movements or touch events. The path and behavior of a human moving a mouse or interacting with a touchscreen provide evidence of human interaction. So, if the button is clicked and there is a corresponding mouse movement or touch event in a way that looks human, the site can reasonably conclude that the user is indeed human. These methods can be less disruptive to the user experience than a traditional CAPTCHA system, but they might also be less secure, as they can be easier for sophisticated bots to bypass. This system combines the less disruptive user experience of this type of CAPTCHA, while improving security. This is accomplished by directly measuring physiological signals of the person to authenticate their personhood (rather than relying on just mouse or touch movements).
[0040] FIG.1 shows a flow chart, describing the proposed process of human authentication via measurement of physiological signals. Physiological signals of the user are passively collected via a radar sensor (Algorithm 1) as shown in FIG.2. These signals may include respiration, cardiac activity, lip movement, and / or vocal fold vibration. These signals are then used to detect whether a human is using the device (Algorithm 2) as shown in FIG.3. If it cannot be conclusively determined if a human user is present, an active task is administered (Algorithm 3) as shown in FIG.4.
[0041] The active task involves a speaking task where a prompt is provided to the user. The user is asked to respond to the prompt verbally, generating98974549.210Attorney’s Docket No.: 055743-816976 (M24-042P-US1-b) both an acoustic signal and a radar signal. The acoustic signal is captured by the microphone. The radar signal results from vocal fold vibration, articulator movement, as well as heartbeat and respiration. Both signals are then processed to validate the radar signal relative to the acoustic signal (Algorithm 4) as shown in FIG.5. This validation ensures that the task was performed correctly, that both signals originated from the same source, and that the signals were indeed generated by a human user. The final result of this process is a validation that a human generated the signals, confirming the user's authenticity. A description of each of the algorithms is provided in the ensuing sections.
[0042] This system provides a robust and reliable method for human verification, combining passive physiological signal detection with active task administration when necessary. The use of both a microphone and a radar sensor ensures accurate and comprehensive user verification.
[0043] Generating a range-Doppler representation from radar signals suitable for measuring physiological signals from a human involves several pre- processing steps. Some common processing steps include:
[0044] Signal Acquisition: The first step is to acquire the raw radar signals. This involves setting up a radar system and capturing backscattered signals from a human subject.
[0045] Range Compression: This step involves processing the received radar signal to improve the range resolution. This is typically done using a matched filter or pulse compression technique, which maximizes the signal-to-noise ratio (SNR).
[0046] Doppler Processing: The next step is to perform Doppler processing. This involves applying a Fast Fourier Transform (FFT) to the range- compressed signals to separate them based on their Doppler shifts. The Doppler shift is a change in frequency due to the relative motion between the radar system and the human subject.
[0047] Clutter Suppression: Clutter refers to unwanted radar reflections that can interfere with the desired signals from the human subject. Clutter can come from stationary objects in the environment or from parts of the human body that are not of interest. Clutter suppression techniques, such as high-pass filtering or adaptive filtering, are used to minimize these effects.98974549.211Attorney’s Docket No.: 055743-816976 (M24-042P-US1-b)
[0048] Range-Doppler Map Generation: After clutter suppression, the processed signals are used to generate a range-Doppler map. This is a 2D representation that shows the Doppler shift (indicating velocity) on one axis and range (indicating distance) on the other axis. The intensity at each point in the map represents the radar reflectivity at that range and velocity.
[0049] Normalization: The range-Doppler map is often normalized to a standard scale to facilitate further analysis. This can involve subtracting the mean and dividing by the standard deviation or scaling the values to lie between 0 and 1.
[0050] Algorithm 2, as depicted in the flow chart of FIG.3, is designed to detect whether a human user is interacting with the device. This algorithm operates on the data representation output from Algorithm 1, which provides a suitable format for the extraction of physiological signals from an acquired radar signal.
[0051] As shown, the first step in Algorithm 2 is feature extraction. This process involves distilling the data representation into a more compact form by extracting both spectral and temporal features. Spectral features capture the frequency components of the physiological signals, while temporal features represent the changes in these signals over time. This dual extraction allows the algorithm to capture a comprehensive profile of the user's physiological signals.
[0052] Once the features have been extracted, they are evaluated by a pre-trained model. This model will be pre-trained on a human database of the same features, allowing it to effectively distinguish between human and non-human users. The model evaluates the extracted features and determines whether they align with the patterns of a human user.
[0053] This process is not a one-time operation. Instead, Algorithm 2 is designed to repeat this process periodically as a user is working with the device. This ensures continuous monitoring and periodically provides a continuous (or discrete) prediction of whether the user is human.
[0054] Algorithm 3 is designed to determine whether an active task is required for human authentication or whether passive radar-based detection is sufficient. The operation of Algorithm 3 is based on the output of Algorithm 2, which detects whether a human user is interacting with the device. If Algorithm 2 provides a discrete decision that a human was detected, then there is no need for an active98974549.212Attorney’s Docket No.: 055743-816976 (M24-042P-US1-b) task. In this case, the user is authenticated as human based solely on passive radar- based detection. If Algorithm 2 decides that a human was not detected, an active task is activated.
[0055] If the output of Algorithm 2 is continuous, then the decision to activate an active task is based on a threshold. If the continuous output value is below the threshold, indicating uncertainty in human detection, an active task is activated. If the output value is above the threshold, indicating a high confidence in human detection, no active task is needed.
[0056] The process shown in FIG.3 ensures that the active task, which requires user interaction, is only activated when necessary, thereby enhancing the user experience while maintaining robust human authentication.
[0057] Active Task Activation: This is the initial stage where the active task is triggered, typically when the passive detection from Algorithm 3 is uncertain or below a certain threshold.
[0058] Task Prompt for Verbal Output: In this stage, the device prompts the user to perform a task that requires a verbal response. This could be a simple spoken phrase or answering a question.
[0059] Acoustic and Radar Sensor Activation: As the user speaks in response to the prompt, the acoustic sensor (microphone) captures the acoustic speech signal and the radar sensor captures the corresponding physiological signals.
[0060] Time Alignment of Sensors: The signals from the acoustic sensor and the radar sensor are aligned in time. This ensures that the features extracted from both sensors correspond to the same time frames of the speech production process. The sensors maybe co-located or dispersed based on application.
[0061] Feature Extraction: Features related to vocal fold vibration and articulator movement are extracted from both the acoustic signal and the radar signal. These features capture essential characteristics of the speech production mechanism.
[0062] Feature Comparison: The extracted features from the acoustic signal and the radar signal are compared. This comparison aims to ensure that the98974549.213Attorney’s Docket No.: 055743-816976 (M24-042P-US1-b) features represent the same speech production mechanism, validating that the microphone signal and the radar signal originated from the same source.
[0063] Validation of Same Speech Production Source: The final stage is the validation of the source of the speech production. If the comparison of features indicates that both signals likely originated from the same source, the user is validated as a human. If not, the process may be repeated or other measures may be taken. Novel aspects of this invention
[0064] The novel aspects of this system include:
[0065] Use of Radar for CAPTCHA: The use of radar sensors in a CAPTCHA device is unique. The radar sensor passively evaluates physiological signals.
[0066] Passive and Active Human Detection: The device operates in both passive and active modes for human detection. In the passive mode, it uses radar signals to evaluate physiological signals like breathing, heartbeat, lip movement, and vocal fold vibration. If the passive detection is uncertain, the device switches to an active mode where it administers a speaking task to the user. The microphone captures the user's speech during an active task. This combination of radar sensor and microphone during active tasks allows for a more robust and reliable human detection and validation process.
[0067] Use of Algorithms for Signal Processing and Human Detection: The device includes a processor that executes a series of algorithms for pre-processing radar signals, extracting and comparing features from the radar and acoustic signals, and determining whether an active task is required. These algorithms executed by the processor enable the device to accurately determine whether the user is a human and validate the source of the speech production.
[0068] Validation of Radar Signal Relative to Acoustic Signal: The device validates the radar signal measurements of vocal fold vibration and lip movement relative to the speech signal. The process ensures that the radar and microphone data have the same source, further enhancing the reliability of the human detection process.98974549.214Attorney’s Docket No.: 055743-816976 (M24-042P-US1-b)
[0069] Periodic Human Detection: The device periodically evaluates the user's physiological signals as the user is working with the device. This feature allows for continuous monitoring and authentication of the user, which is particularly useful in preventing automated form submissions, unauthorized access, web scraping, spamming, and similar kinds of malicious activities perpetrated by computer programs on the web. Advantages over current technology and impact
[0070] A CAPTCHA device as disclosed herein presents several advantages over current technology and could have a significant impact in various fields:
[0071] Enhanced Security: The combination of passive and active human detection mechanisms significantly enhances security. By using physiological signals and speech, the device can more accurately distinguish between humans and automated systems, reducing the risk of unauthorized access or malicious activities.
[0072] Non-Intrusive Authentication: The passive detection mode of the device is non-intrusive, as it uses radar to evaluate physiological signals. This can improve user experience as it does not require active participation unless necessary.
[0073] Continuous Authentication: The device's ability to periodically evaluate a user's physiological signals allows for continuous authentication. This is a significant advantage over traditional one-time authentication methods, as it can detect if the authenticated user is replaced by an unauthorized user during a session.
[0074] Versatility: The device's use of both acoustic and radar signals makes it versatile and adaptable to various environments and conditions. It can be used in noisy environments where acoustic signals alone may not be reliable, or in situations where physiological signals can provide additional verification.
[0075] Impact on Accessibility: This technology could also have a positive impact on accessibility. For users who may have difficulties with traditional CAPTCHA tasks, such as those with visual impairments, the use of physiological and vocal signals could provide an alternative method of verification.98974549.215Attorney’s Docket No.: 055743-816976 (M24-042P-US1-b)
[0076] Potential for Integration: The technology could potentially be integrated into a wide range of devices and systems, from smartphones and computers to secure entry systems and online platforms, enhancing security across a broad spectrum of applications. EXAMPLE 2
[0077] A core, proactive solution to the voice cloning (deepfake) problem is disclosed in the present example. This solution is based on new microphone technology that can certify on the recording device itself that the recorded acoustic signal originates from a human who is speaking into the microphone. This certificate can be linked to a given media file and then accessed on the listener side to provide verification that a human generated the audio. The microphone technology is based on a combination of a traditional MEMS (Micro Electronic Mechanical System) microphone (or other type of microphone), which measures the acoustic signal, and a radar sensor, which measures the physiological vibrations from the human speech production mechanism. These two signals are verified relative to each other at the time of recording to prove the personhood of the speech generator. Speech Generation and the Link Between Acoustic Signal, Vocal Fold Vibration, and Articulator Motion
[0078] Speech generation is a complex, physiological process involving a finely coordinated interplay of various body parts. A simple physiological model of this process is shown in FIG.6, where sound is created by vocal fold vibrations and further shaped as it passes through the oral and nasal cavities, resulting in speech. The process begins with the production of voiced sound in the larynx, a process known as phonation. Herein, the vocal folds, also known as vocal cords, positioned within the larynx, play a pivotal role. As air from the lungs passes through the tightly closed vocal folds, it causes them to vibrate. These rapid, repetitive openings and closings of the vocal folds modulate the airflow and generate a series of acoustic pressure waves
[0079] These pressure waves form the primary acoustic signal, which carries the fundamental frequency of the talker's voice. This frequency is largely determined by the physical properties of the individual's vocal folds, such as their98974549.216Attorney’s Docket No.: 055743-816976 (M24-042P-US1-b) length and tension, and the rate at which they vibrate. This fundamental frequency typically lies in a range of about 80 to 240 Hz in adults.
[0080] The primary acoustic signal then travels through the vocal tract - the throat, oral cavity, and nasal passages. As it passes through these structures, the signal is shaped and filtered by the resonant properties of the vocal tract, resulting in the rich, complex acoustic signal that we perceive as a person’s voice.
[0081] Thus, the physiological measurements of the talker are intrinsically linked to the generation of the speech signal. As a result, the vibration of the vocal folds and the movement of the articulators is directly related to acoustic data captured by a microphone. The Source-Filter Model and Inverse Filtering in Speech Processing
[0082] In the field of speech processing, the source-filter model serves as a fundamental framework to conceptualize speech production and to process speech signals. According to this model, the generation of speech sounds involves two primary components: the source and the filter. The source component refers to the voiced or unvoiced excitation generated at the larynx, essentially the vocal fold vibration for voiced sounds. The filter represents the vocal tract configuration, which resonates and shapes these source excitations into discernible speech sounds.
[0083] The acoustic signal, as perceived at the talker's mouth, is a convolution of the source excitation and the vocal tract filter response, resulting in the intricate and unique phonetic output we recognize as individual speech sounds. The vocal tract imposes its resonant properties on the source signal, generating a series of formants or spectral peaks in the acoustic output. Mathematically, this can be modeled as the convolution of a source signal (vocal fold vibrations) with a vocal tract filter (movement of articulators) (see top block diagram in FIG.6).
[0084] Inverse filtering is a technique used to deconvolve, or separate, the source and filter components from the observed speech signal. Essentially, it aims to reverse the filtering effect of the vocal tract to obtain the original source excitation, which in the case of voiced sounds, is the vocal fold vibration signal (see bottom block diagram in FIG.6). This can be achieved through linear predictive coding (LPC) or cepstral analysis, among other methods, which estimate the vocal tract filter and subsequently 'remove' its effect from the acoustic signal. By successfully applying inverse filtering, the vocal fold vibration signal and98974549.217Attorney’s Docket No.: 055743-816976 (M24-042P-US1-b) configuration of the articulators can be extracted from the acoustic speech signal, thereby providing valuable information about the physiological aspects of voice production; namely, the vibration of the vocal folds and the movement of the articulators (lips, jaw, soft palate).
[0085] The estimated source signal via inverse filtering is directly related to the vibration of the vocal folds and the estimated filter signal is directly related to the movement of the articulator (lips, tongue, jaw, soft palate). That is, if the vibration of the vocal folds or the movement of the articulators were measured directly, that signal would closely match the source signal estimated from the measured acoustic speech signal. Measuring human physiological signals during speaking via radar
[0086] Radar, which stands for Radio Detection and Ranging, is a technology that uses radio waves to detect and determine the distance, direction, and speed of objects. It works by emitting a radio signal, which then bounces off any object in its path, and the return signal is captured by the radar system. Radar technology precisely measures several key parameters of detected objects. It measures the 'range' or 'distance' to the object by calculating the time it takes for the radio wave to return after bouncing off the object. It determines the 'velocity' or 'speed' of the object using the Doppler Effect, where the frequency shift of the returned signal indicates motion towards or away from the radar.
[0087] Utilizing a radar system to target an individual engaged in speaking, it is feasible to obtain a range of physiological signals by measuring oscillations emitted from the individual’s body, which correspond to various physiological processes. These encompass the oscillation due to heartbeats, driven by the heart's rhythmic contraction and relaxation, respiratory movements from the expansion and contraction of the lungs, articulatory actions including lip and jaw kinetics, as well as vocal fold vibrations integral to phonation.
[0088] FIG.7 shows each of these physiological processes, the unique frequency bands in which they fall, and the magnitudes of their corresponding radar signals. Respiration typically falls within a range of 0.2 to 0.5 Hz, reflecting the slow, regular expansion and contraction of the lungs. Cardiac activity, more specifically heartbeats, are found within a slightly broader and higher frequency band of 0.2 to 298974549.218Attorney’s Docket No.: 055743-816976 (M24-042P-US1-b) Hz. The articulatory movements of the lips during speech occur within the band of 3 to 7 Hz, attributable to the rapid, complex movements involved in speech production.
[0089] Furthermore, vocal fold vibrations occur at a significantly higher frequency band of 80 to 240 Hz. The swift, repetitive opening and closing of the vocal folds during phonation cause these vibrations, which are integral to the generation of voiced sound.
[0090] Although there is a degree of overlap in the frequency bands of respiration and heartbeats, vocal fold vibrations reside in a distinctly separate frequency band. There is no overlap of this band with any other human-generated physiological signal, rendering it uniquely identifiable. This characteristic separation simplifies the task of isolating the vocal fold vibration signal from the comprehensive physiological data gathered by the radar system.
[0091] Respiratory movements typically generate the largest radar signals. This is due to the sizable chest wall displacement and the substantial internal movement of the lungs during breathing. The sheer scale of this action results in a large-reflected radar signal.
[0092] Cardiac activity, though regular, generates smaller radar signals due to the relatively smaller physical displacement involved. The heart is a smaller organ, and its movement during contractions is limited compared to the larger scale breathing motion.
[0093] Lip and jaw movements during speech, while having a relatively higher frequency, produce relatively smaller radar signals. The body parts involved in articulation are smaller in comparison and exhibit limited movement even when articulating complex speech sounds, thus reflecting less radar signal back to the receiver.
[0094] Lastly, vocal fold vibrations, despite their high frequency, yield the smallest magnitude radar signals. Vocal folds are small, their movement during speech is of very small amplitude, and they are located deep within the body. These factors result in a very small radar cross-section and the movement creates only minor changes in the overall reflected radar signal. However, the uniqueness of the frequency band of vocal fold vibrations enables their detection and isolation for analysis, despite their low signal magnitude.98974549.219Attorney’s Docket No.: 055743-816976 (M24-042P-US1-b)
[0095] This is shown in FIG.8, where a radar system is set up in front of a human talker. Here a miniature 77 GHz millimeter wave (mmWave) radar sensor is used for simultaneous human vocal sound and heart sound measurement while a test subject is enunciating simple words or sentences. The talker is located a distance of 0.7 m from the microphone and the radar sensor. The acquired radar signatures in the time domain and in the frequency domain are shown in FIG.8. The resulting spectrum shows a clear separation of the low-frequency physiological signals (breathing, heartbeat, articulator movements) from the high-frequency physiological signals (vocal fold vibration).
[0096] A time-frequency representation of a radar signal acquired during speaking is further shown in FIG.9A. This figure shows the vocal fold vibration associated with the speech signal at the higher frequencies and the low- frequency physiological signals. The low-frequency physiological signals are then removed, resulting in an isolated representation of the speech signal captured by the radar sensor, as shown in FIG.9B. This figure shows changes in the pitch and pitch harmonics over time. This signal can then be converted back to the time domain to generate a time domain estimate of the produced speech signal. If the speech signal is also captured by a microphone, it allows the radar signal shown in FIG.9B to be compared to the acoustic signal acquired via the microphone. Validating the acoustic signal relative to the radar signal
[0097] Synchronizing and validating speech and radar signals to ensure they originate from the same individual involves several processing stages. The data collected from both the radar sensor and the microphone represent a combination of personal physiological and acoustic features. The system then can automatically and algorithmically validate that a measured speech signal was generated by a human by tying these data streams together and validating that they originated from the same source. This includes the following steps:
[0098] Data Acquisition and Time Synchronization: Simultaneous data collection is performed using a radar sensor and a microphone. The two sensors are synchronized via a common clock or timestamping the data from both sensors.
[0099] Speech Signal Processing: The microphone-recorded speech signal is analyzed using various digital signal processing techniques. This could98974549.220Attorney’s Docket No.: 055743-816976 (M24-042P-US1-b) involve an inverse filtering process to isolate the source and the filter, and extracting features, such as fundamental frequency, formants, spectral characteristics, and timing features of the speech signal. This information represents the unique characteristics of an individual's speech.
[0100] Radar Signal Processing: The radar signals representing the physiological movements related to speech production are processed. This involves steps as explained earlier, such as Range-Doppler processing and signal filtering. Specifically, the high-frequency vocal fold vibrations and articulatory movements are of interest, as they are tightly coupled with speech features extracted from the source and filter.
[0101] Cross-modal comparison and validation: Key features from both the radar signal and the speech signal, such as the pattern and frequency of vocal fold vibrations, the timing of articulatory movements, and other speech characteristics, are then compared to each other. By correlating these two datasets, it can be validated that they are indeed originating from the same individual. Advanced signal processing or machine learning algorithms may be used to perform this cross-modal integration. These algorithms may compare and contrast the feature sets from both signals, thereby establishing a relationship between them. High correlation or congruence may suggest that the radar signal and the speech signal originated from the same person.
[0102] Continual Monitoring and Authentication: By continually monitoring and comparing the radar and speech signals, it is possible to certify the persistency of the same talker. Any significant deviation in the comparison metrics may suggest a change in the talker or a change in the acoustic signal.
[0103] In a similar setup to the one used in FIGS.8, 9A and 9B, an utterance from a talker (‘She had your dark suit in greasy wash water all year’) was collected and recorded using a microphone and a radar sensor. The acoustic signal and the radar signal are filtered to the same frequency range (0 to 500 Hz) and then time-aligned by finding a peak in the cross-correlation. From each signal, a series of features was extracted that track harmonicity, temporal variation, and several spectral parameters of both signals. FIG.10A shows how a normalized harmonicity feature varies in a synchronous fashion across both modalities when there is a match between the acoustic signal and the radar signal. FIG.10B shows how the98974549.221Attorney’s Docket No.: 055743-816976 (M24-042P-US1-b) same feature does not vary in a synchronous fashion when there is not a match between the acoustic signal and the radar signal. It is evident from FIG.10A that the features across both modalities track in time, making it possible to perform cross- modal validation. That is, the high correlation between the features indicates that it is possible to validate that the acoustic signal recorded by the microphone originated from the vocal fold vibrations and articulator movements of the talker measured by the radar sensor. The lack of correlation between the two signals in FIG.10B shows that in this instance, such a cross-modal validation can be performed with high confidence.
[0104] To further evaluate the confidence level with which the cross- modal validation can be performed, a distribution of feature correlations with a mismatch between the acoustic signal and the radar signal was generated. We randomly select from an existing corpus of 350 audio samples, not collected at the same time as the one captured by the radar. That is, these speech samples are mismatched from the one collected by the radar and spoken by different talkers speaking different sentences. For each audio sample, we compute the feature values for the acoustic signal and the radar signal and compute the correlation between the two sets of features per speech utterance. The normalized distribution (z-scored) of correlations is shown in blue in FIG.11. In the same figure, the true correlation (z-scored) when there is a match between both signals is shown in red. As the figure shows, even using a single feature, the true correlation value of this feature falls almost 4 standard deviations away from the mean, indicating that it is possible to determine a match between the acoustic signal and the radar signal with very high confidence even on a single utterance. Linking the certificate to the media
[0105] Post-validation, the resulting authentication certificate can be linked to the media to verify its origin and ensure that it has been human generated and not altered or forged. There are several methods that can be utilized to link the certificate to the media. These include, but are not limited to:
[0106] Using Existing Software to Digitally Sign: Some media editing or security software may allow you to digitally sign media files directly. Tools tailored to digital rights management (DRM) or secure media handling might include this functionality.98974549.222Attorney’s Docket No.: 055743-816976 (M24-042P-US1-b)
[0107] Hashing and Signing: Create a hash of the audio file, which is a unique numerical value that represents the file's content. Use a private key to sign this hash, creating a digital signature. The signature and public key can be stored alongside the audio file or embedded within metadata, allowing recipients with the corresponding public key to verify the signature.
[0108] Embedding the Signature in the File's Metadata: Metadata in audio files can often store additional information, such as artist details, album name, etc. A digital signature can also be embedded in this metadata, although care must be taken to ensure that this does not disrupt the playback of the file.
[0109] Embedding the Signature in a Watermark: Apply a digital watermark to the media. This can be a visible or invisible pattern that includes the authentication information.
[0110] Blockchain Technology: Utilize blockchain to create a verifiable and unchangeable record of the media. For example, platforms like Ethereum can be used to write a smart contract, encoding the authenticity information of the media within the blockchain.
[0111] Using Standard File Signing Tools with Wrappers: If the audio file is contained within a wrapper format (like a ZIP file) that can be digitally signed using standard file signing tools, then the whole package can be signed, including the audio content. Recipients can then verify the signature of the package, thereby verifying the authenticity of the audio file within.
[0112] Other Custom Solutions: Depending on the specific requirements, a custom solution could be developed to digitally sign media files. For example, the radar features can be linked to the audio file via existing tools (e.g. watermarking, blockchain, etc.). Then, during authentication, the speech features can be extracted and evaluated against the radar features to ensure that there is a match. Complexity of Vocal Fold Vibration and Challenges in Spoofing Radar Captured Signal
[0113] The process of vocal fold vibration is a highly intricate physiological phenomenon that directly results in the acoustic signal humans produce as they speak. The vocal folds, situated in the larynx, oscillate at high frequencies ranging from approximately 80 to 240 Hz in adults during voiced speech98974549.223Attorney’s Docket No.: 055743-816976 (M24-042P-US1-b) production. However, the vibration is not a simple periodic motion but rather a complex, three-dimensional, mucosal wave action that involves longitudinal, horizontal, and vertical displacements. These vibrations are influenced by numerous factors, including the length, tension, and mass of the vocal folds, subglottal pressure, vocal tract configuration, and individual physiological characteristics.
[0114] A radar signal resulting from vocal fold vibrations carries these complex movement patterns as frequency modulations, along with other physiological features of the talker’s vocal anatomy. When combined with the other physiological signals captured by the radar sensor, such as respiration and cardiac activity, the combination of these characteristics forms a unique “human print”, which is highly challenging to artificially reproduce or spoof.
[0115] To attempt to spoof such a radar signal, it would not only be required to mimic the unique frequency characteristics of the talker's vocal fold vibration but also to replicate their specific physiological features affecting the radar cross-section. Additionally, the intricacy of the three-dimensional, high-frequency mucosal wave action of the vocal folds further amplifies this challenge. Given currently available technology and current understanding of human physiology and radar signal processing, successfully spoofing a radar signal resulting from vocal fold vibrations presents a considerable challenge.
[0116] One simple attempt to spoof the system could be to place a loudspeaker in front of the radar such that the radar measures the vibration of the loudspeaker’s membrane while the microphone measures the acoustic signal. To detect attempts to spoof the system, machine learning algorithms can be implemented to ensure that a human is in front of the microphone (and not a synthetic vibrating membrane). Through machine learning, an algorithm can be trained on a vast array of radar signals from humans speaking, learning to recognize the subtle patterns and characteristics of radar returns that are typical of living human talkers during speech. This model can then be used to evaluate whether an acquired signal is real. In FIG.12 below, we show the radar returns from a human talker and the returns from a vibrating loudspeaker membrane (at three playback levels). This figure shows clear differences between the human returns and the loudspeaker. The first obvious difference is the difference in turns when there is no voiced speech being generated. The background noise levels are much higher for98974549.224Attorney’s Docket No.: 055743-816976 (M24-042P-US1-b) the loudspeaker’s returns. The second obvious difference is that the magnitude of the loudspeaker’s returns are greater (more red in the loudspeaker returns). The third obvious difference is the difference in returns for the higher-frequency harmonics. As the figure shows, the radar returns for higher-frequency harmonics are more pronounced, particularly for the medium and high-volume levels. This figure provides evidence that it should be possible to learn a statistical model via machine learning that captures the radar returns from a human speaker. Novel aspects of invention
[0117] The novel aspects of this invention lie in its unique approach to audio authenticity verification, which combines traditional acoustic signal recording with physiological vibration measurements. Some key innovative elements of this approach include:
[0118] Dual-Sensor Technology: The integration of a traditional MEMS microphone with a radar sensor is a novel approach. While the microphone records the acoustic signal, the radar sensor measures the physiological vibrations from the human speech production mechanism. This dual-sensor setup provides a more comprehensive and accurate way to verify the presence of a human speaker. The two sensors can either be co-located or dispersed based on the application.
[0119] Real-Time Verification: The ability to verify the personhood of the speech generator at the time of recording is a significant advancement. This real- time verification can prevent the damaging effects of deepfakes or other synthetic audio from the outset.
[0120] Certification Linking: The ability to link the certificate of authenticity to the media file itself is another novel aspect. This feature allows listeners to access the verification information directly, providing an immediate assurance of the audio's authenticity.
[0121] Physiological Vibration Measurement: The use of radar technology to measure physiological vibrations associated with human speech production is a unique aspect. This approach provides a biological confirmation of a human talker, which is difficult to replicate artificially.
[0122] On-Device Verification: The fact that the verification process occurs on the computing device as the acoustic signal is being recorded is also98974549.225Attorney’s Docket No.: 055743-816976 (M24-042P-US1-b) innovative. This on-device processing can enhance the security and efficiency of the verification process. Advantages over current technology and impact:
[0123] The proposed speech certification system offers several advantages over current technology and could have a significant impact:
[0124] Authenticity Verification: The primary advantage of the invention is the ability to verify the authenticity of audio content. This technology can provide a certificate of authenticity that proves a human, not an AI or deepfake, generated the audio. This feature is increasingly important in an era where deepfakes and synthetic media are becoming more prevalent.
[0125] Prevention of Misinformation: By verifying the source of an audio signal, this invention can help prevent the spread of misinformation and disinformation. It can be a crucial tool in maintaining the integrity of news broadcasts, political speeches, and other public communications.
[0126] Security and Privacy: The invention can enhance security in communication systems. For instance, it can prevent voice cloning scams, as it ensures the speaker is a real person. It can also be used in secure voice biometric systems, adding an additional layer of security.
[0127] Legal and Forensic Applications: In legal contexts, this invention can help verify the authenticity of audio evidence. It can also be used in forensic investigations to confirm the identity of talkers in audio recordings.
[0128] Music and Entertainment Industry: In the music and entertainment industry, the invention can help protect artists' rights by ensuring that an audio recording was created by a human artist and is not artificially generated.
[0129] The impact of this invention could be far-reaching, potentially changing the way humans trust and interact with audio media. It could restore faith in digital media by providing a reliable way to verify authenticity, which is becoming increasingly important in the digital age. Applications of this technology
[0130] Ensuring safe communication: Scams based on voice cloning have begun to appear. Having a human verification check mark during a conversation ensures that the person on the other end is a human talker and not an98974549.226Attorney’s Docket No.: 055743-816976 (M24-042P-US1-b) AI bot. In addition to consumer applications, the applications in defense for such a technology are considerable.
[0131] Next generation CAPTCHA: With improvements in AI technology, current CAPTCHA technology will soon become obsolete. It has become increasingly easier for AI to detect deformed digits / text or to identify specific objects in a visual scene. We posit that next generation CAPTCHA will require human verification via physiological measurement. A radar-based system allows for efficient verification.
[0132] Authenticating music, podcasts, and other forms of content as human-made: As generative AI improves, AI-generated entertainment content is likely to outnumber human-generated content online. Under this scenario, distinguishing between human-made media and AI-generated media provides opportunities for new business models (e.g. different monetization strategies for human content vs. AI content).
[0133] Celebrities / influencers protecting and monetizing their voices: As AI models improve, synthetic media from celebrities and influencers will be easy to generate. Synthetic voices have already been used to generate celebrity hate speech (e.g. see https: / / www.theverge.com / 2023 / 1 / 31 / 23579289 / ai-voice- clone-deepfake-and style of famous artists (https: / / www.npr.org / 2023 / 04 / 21 / 1171032649 / ai-music-heart- on-my-sleeve-drake-the-weeknd). Validation of human-generated content will allowand their brand. There are also very interesting future markets. For example, some celebrities have developed AI duplicates where they allow fans to interact with the duplicate (e.g. see https: / / www.washingtonpost.com / technology / 2023 / 05 / 13 / caryn-ai-technology-gpt-4 / ).interactions are more expensive than interactions with AI-generated bots.
[0134] Authenticating virtual assistants: Human users can be verified before any command is executed by a virtual assistant, such as Alexa or Siri, providing an added layer of security when communicating with smart devices.
[0135] Proof-of-personhood / Proof-of-humanity: Proof of personhood / Proof-of-humanity is a way to verify that a user is a unique, real human being and not a bot or fake account. There are two main forms of proof of 98974549.227Attorney’s Docket No.: 055743-816976 (M24-042P-US1-b) personhood: social-graph-based and biometric. Social-graph-based proof of personhood relies on some form of vouching, while biometric proof of personhood involves verifying some physical or behavioral trait that distinguishes humans from bots and individual humans from each other. Voice biometrics can play a role in biometric proof of personhood by verifying an enrolled person’s identity. The biometric voice recognition system captures a new speech sample, creates a template from the sample, and compares it against the enrollment template. A strong match between templates indicates that the same person spoke both samples, thus verifying the person’s identity. The proposed technology can be used in the context of such a system to provide additional evidence that the user generating the speech is a human. EXAMPLE 3
[0136] The present example outlines a method and system for authenticating live human speech and protecting against deepfake audio impersonations. It integrates a variety of biometric sensors to capture physiological and acoustic signals that are indicative of natural human speech production.
[0137] The system includes a wearable device capable of authenticating live human speech by detecting a suite of biometric and acoustic signals. Its general form is adaptable and can be embodied as an in-ear device, headphones, or a head-mounted apparatus. The device incorporates a range of sensors, including bone conduction microphones, photoplethysmography sensors, accelerometers, gyroscopes, and temperature sensors. These sensors work in concert to capture the vibrational, physiological, and thermal signatures unique to human speech as produced by an individual. The objective is to provide a robust method for distinguishing between live human speech and non-human sources, with potential applications in security, personal identification, health monitoring, and interactive technology.
[0138] Deepfake technology has necessitated the development of reliable methods to authenticate the veracity of multimedia content, especially human speech. The systems and methods outlined herein are a response to the need for a robust, real-time method that can differentiate authentic human speech from deepfake-generated audio. The method capitalizes on recent advances in sensor technology and algorithmic analysis to assess the biometric and acoustic98974549.228Attorney’s Docket No.: 055743-816976 (M24-042P-US1-b) signals inherent to natural speech, offering a new layer of defense against the burgeoning threat of deepfakes.
[0139] Consistent with the previous example, once human speech is verified, it can be linked to an associated media file to verify a listener that the speech is human generated. Contributions
[0140] The systems and methods outlined herein contribute the following:
[0141] Multi-Sensor Fusion: The method integrates a variety of sensors such as bone conduction microphones, PPG sensors, accelerometers, gyroscopes, and thermal sensors in a wearable format to capture a comprehensive data set of physiological and acoustical signals related to speech production. This fusion of multiple data types for live speech verification is unique.
[0142] Real-Time Deepfake Detection: By analyzing physiological signals that are difficult to synthesize, such as the subtle blood flow changes associated with speech or the unique patterns of jaw movements, the system can detect deepfakes in real-time, which is a significant improvement over existing technologies that typically analyze only audio and lack real-time processing capabilities.
[0143] Wearable Format: The method is designed to be integrated into wearable devices (e.g., hearables, head-mounted microphones, combinations, etc.), providing a discreet and convenient form factor for users.
[0144] Physiological Signal Authentication: Existing technologies primarily focus on the sound wave analysis of speech for deepfake detection. In contrast, the systems outlined herein authenticate speech through underlying physiological processes, offering a novel approach to verification that is inherently more secure against current and future deepfake technologies.
[0145] Enhanced Security: By utilizing physiological data alongside acoustic signals, the systems outlined herein provide a level of security that is inherently more resistant to spoofing attempts, including sophisticated deepfake audio impersonations. Current deepfake detection technologies relying on the acoustic output exhibit variable performance and are being outpaced by the rapid progress in deepfake generation.98974549.229Attorney’s Docket No.: 055743-816976 (M24-042P-US1-b)
[0146] Real-Time Authentication: The method enables real-time processing and authentication of live speech, which is crucial for dynamic security systems and interactive applications, unlike many current technologies that operate on pre-recorded samples and lack the capability for on-the-fly verification.
[0147] Versatility in Application: The wearable nature of the device allows for a wide range of applications, from mobile security to healthcare monitoring, without the need for specialized environments or setups required by some current technologies.
[0148] Non-Invasive and User-Friendly: The system outlined herein is designed to be non-intrusive and user-friendly, integrating seamlessly into everyday devices such as headphones or eyeglasses, offering a frictionless experience for the user.
[0149] Robustness to Environmental Variations: With the ability to capture and analyze physiological signals, the system outlined herein is less susceptible to environmental noise and variations that can affect solely acoustic- based technologies. Core Components and Methodology Sensors Array
[0150] The system employs an array of sensors that can be configured in various wearable forms, including but not limited to in-ear devices, neckbands, headsets, or eyeglasses. These sensors include, but are not limited to:
[0151] Bone Conduction Microphones: To capture vocal fold vibrations transmitted through the skull.
[0152] Photoplethysmography (PPG) Sensors: To detect blood volume changes in the tissues of the ear or skin, which correspond to human skin and may correspond with the effort of speech.
[0153] Accelerometers and Gyroscopes: To monitor the micro- movements associated vibrations during speech production and with jaw and head motions during speech.
[0154] Thermal Sensors: To measure the temperature near the skin's surface, which can fluctuate with speech due to blood flow changes.98974549.230Attorney’s Docket No.: 055743-816976 (M24-042P-US1-b)
[0155] High-Fidelity Microphones: To capture the audio signal within the ear canal or near the mouth, using ambient noise reduction techniques. Signal Processing
[0156] The system employs advanced signal processing algorithms to analyze the collected data from multiple sensing modalities, preparing them for the data fusion engine. This includes extraction of features for the different sensing modalities.
[0157] Temporal Features: Features that measure the timing and sequence of speech-related physiological changes.
[0158] Spectral Features: Features that decompose audio signals into their constituent frequencies, useful in identifying acoustic, vibrational, and other physiological cues associated with speech production. Data Fusion Engine
[0159] FIG.13 illustrates a Data Fusion engine 1300 that can be implemented using a computing device (e.g., computing device 1400 of FIG.14) onboard or otherwise in communication with various wearable devices outlined herein. The Data Fusion Engine ensures the human authenticity of audio. This can be done in two ways. First, the collected signals are fused to create a multi- dimensional profile of a speech event using the features extracted in the previous step. This profile is compared against known profiles of authentic speech to determine the likelihood of the speech being live and human. Next, several of the sensing modalities are evaluated relative to each other to ensure that there is live human speech. For example, the vibrational signature (measured from the bone conduction microphone), the movement of the articulators (measured from the accelerometer), and the acoustic output (measured from the microphone) can be validated relative to each other to ensure that they have the same source (the human speaker). Output and Integration
[0160] The system provides an output that indicates the human authenticity of the speech. This output can be integrated into security systems requiring voice authentication, communication devices to verify the identity of98974549.231Attorney’s Docket No.: 055743-816976 (M24-042P-US1-b) speakers, media platforms to screen for deepfake content, or content authentication platforms for verification of authentic human speech. User Interface
[0161] The system includes a user interface for calibration, monitoring, and alerting users regarding authentication status. It could be a standalone application or integrated into existing software ecosystems. Adaptability and Embodiments
[0162] The device form factor can be adapted for discreet use, comfort, and user-specific applications, ranging from consumer electronics to high-security communication devices.
[0163] The sensor array and processing algorithms can be customized for different levels of security, from basic authentication to highly secure environments where voice spoofing poses significant risks.
[0164] The method is adaptable to different populations and languages, accommodating variations in speech patterns and physiological responses.
[0165] The system can be designed for both real-time live speech authentication and post-analysis of recorded speech to verify its authenticity retrospectively. Relevant features for the signal processing component
[0166] Possible features that could be extracted from each sensing modality to generate the human profile during speech include the following (this is not an exhaustive list): Bone Conduction Microphones:
[0167] Vibrational Energy: Measures the energy of vocal fold vibrations transmitted through the skull, indicative of live speech production.
[0168] Fundamental Frequency (F0): Tracks the pitch of the speaker's voice, which varies naturally in live human speech. Photoplethysmography (PPG) Sensors:98974549.232Attorney’s Docket No.: 055743-816976 (M24-042P-US1-b)
[0169] Blood Volume Pulse (BVP): Detects heart rate variability that may correlate with the effort and emotional state of speech.
[0170] Oxygen Saturation: Monitors changes in blood oxygen levels, which could vary with breathing patterns associated with speech. Accelerometers and Gyroscopes:
[0171] Jaw Movement Patterns: Captures the dynamics of jaw movements during speech, which are unique to individuals and difficult to replicate in synthetic speech.
[0172] Head Motion: Analyzes head movements that accompany natural speech, providing cues to the authenticity of the speech act. Thermal Sensors:
[0173] Temperature Variability: Measures fluctuations in skin temperature near the mouth or throat, which could change with speech due to blood flow variations.
[0174] Breath Temperature: Detects the warmth of exhaled air during speech, which could indicate live speech production. High-Fidelity Microphones:
[0175] Spectral Features: Extracts features such as Mel-frequency cepstral coefficients (MFCCs) that capture the unique timbral qualities of a speaker's voice.
[0176] Temporal Dynamics: Analyzes the timing and duration of phonemes, syllables, and words, which are characteristic of natural speech rhythms. Data Fusion Engine
[0177] The Data Fusion Engine integrates and analyzes data from multiple sensors to authenticate the human origin of speech signals. Steps include: Multi-dimensional Profile Creation:
[0178] The engine first aggregates features extracted from the array of sensors, each providing a unique perspective on the speech event. This includes vibrational data from bone conduction microphones, physiological signals from PPG98974549.233Attorney’s Docket No.: 055743-816976 (M24-042P-US1-b) sensors, movement data from accelerometers and gyroscopes, temperature variations from thermal sensors, and acoustic signals from high-fidelity microphones. These features may include energy, blood volume pulse, jaw movement patterns, temperature variability, and spectral features. These features are indicative of natural human speech production mechanisms.
[0179] These extracted features are fused to create a multi-dimensional profile of the speech event. This profile encapsulates the complex interplay of physiological and acoustic signals that characterize live human speech.98974549.234Attorney’s Docket No.: 055743-816976 (M24-042P-US1-b) Comparison Against Known Profiles:
[0180] The multi-dimensional profile is compared against a database of profiles associated with authentic human speech. This database is built from extensive samples of live speech events, capturing a wide range of natural variations in speech production across different individuals and contexts.
[0181] The comparison process involves machine learning algorithms that capture the subtle nuances that characterize live human speech. An example machine learning model could include a Gaussian Mixture Model that models the distribution of the multi-dimensional profile. Cross-modal Validation:
[0182] In addition to comparing the multi-dimensional profile against known profiles, the engine also performs cross-modal validation. This involves evaluating the consistency and correlation between different sensing modalities.
[0183] For instance, the vibrational signature captured by the bone conduction microphone is checked against the movement data from accelerometers (articulator movements) and the acoustic output from high-fidelity microphones. The engine assesses whether these different data streams corroborate each other, indicating they originate from the same live speech event.
[0184] This step is crucial for detecting deepfake attempts that might convincingly replicate one aspect of speech production (e.g., acoustic output) but fail to accurately mimic the holistic, interconnected nature of live human speech. Output Generation:
[0185] Based on the outcomes of the profile comparison and cross- modal validation, the engine generates an output indicating the likelihood that the speech event is live and human. This output can be a binary decision (authentic or not), a confidence score, or a detailed report highlighting the analysis's findings.
[0186] This output can then be used by various downstream systems requiring voice authentication, such as security systems, communication devices, and content verification platforms, to make informed decisions about the authenticity of the speech.98974549.235Attorney’s Docket No.: 055743-816976 (M24-042P-US1-b)
[0187] The verification signal can also be embedded within the speech signal as a watermark. Or added to the recording as meta-information. Or can be used to cryptographically sign the audio as authentic human speech. Discussion
[0188] The present disclosure outlines systems and methods for authenticating live human speech by analyzing physiological and acoustic signals using an array of sensors integrated into a wearable device. This method addresses the urgent need for reliable verification in the face of advanced deepfake technologies, which pose a growing threat to personal security, information integrity, and privacy. The commercial potential of the invention lies in its application across various sectors including security, personal electronics, health monitoring, and communications. It offers a novel solution for companies looking to enhance the authenticity of digital interactions and voice-controlled systems and holds promise for a wide range of licensing opportunities in fields where verification of human presence is crucial. Computer-implemented System
[0189] FIG.14 is a schematic block diagram of an example computing device 1400 that may be used with one or more embodiments described herein, e.g., as a component of the system including the wearable devices outlined herein and implementing aspects of the methods outlined herein.
[0190] Computing device 1400 comprises one or more network interfaces 1410 (e.g., wired, wireless, PLC, etc.), at least one processor 1420, and a memory 1440 interconnected by a system bus 1450, as well as a power supply 1460 (e.g., battery, plug-in, etc.).
[0191] Computing device 1400 can include or otherwise communicate with a display device 1430 that communicates a user interface and / or outputs of the methods outlined herein to a user.
[0192] Network interface(s) 1410 include the mechanical, electrical, and signaling circuitry for communicating data over the communication links coupled to a communication network. Network interfaces 1410 are configured to transmit and / or receive data using a variety of different communication protocols. As illustrated, the box representing network interfaces 1410 is shown for simplicity, and98974549.236Attorney’s Docket No.: 055743-816976 (M24-042P-US1-b) it is appreciated that such interfaces may represent different types of network connections such as wireless and wired (physical) connections. Network interfaces 1410 are shown separately from power supply 1460, however it is appreciated that the interfaces that support PLC protocols may communicate through power supply 1460 and / or may be an integral component coupled to power supply 1460.
[0193] Memory 1440 includes a plurality of storage locations that are addressable by processor 1420 and network interfaces 1410 for storing software programs and data structures associated with the embodiments described herein. In some embodiments, computing device 1400 may have limited memory or no memory (e.g., no memory for storage other than for programs / processes operating on the device and associated caches). Memory 1440 can include instructions executable by the processor 1420 that, when executed by the processor 1420, cause the processor 1420 to implement aspects of the systems and the methods outlined herein.
[0194] Processor 1420 comprises hardware elements or logic adapted to execute the software programs (e.g., instructions) and manipulate data structures 1445. An operating system 1442, portions of which are typically resident in memory 1440 and executed by the processor, functionally organizes computing device 1400 by, inter alia, invoking operations in support of software processes and / or services executing on the device. These software processes and / or services may include audio authentication processes / services 1490, which can include aspects of the methods and / or implementations of various modules described herein. Note that while audio authentication processes / services 1490 is illustrated in centralized memory 1440, alternative embodiments provide for the process to be operated within the network interfaces 1410, such as a component of a MAC layer, and / or as part of a distributed computing network environment.
[0195] It will be apparent to those skilled in the art that other processor and memory types, including various computer-readable media, may be used to store and execute program instructions pertaining to the techniques described herein. Also, while the description illustrates various processes, it is expressly contemplated that various processes may be embodied as modules or engines configured to operate in accordance with the techniques herein (e.g., according to the functionality of a similar process). In this context, the term module and engine98974549.237Attorney’s Docket No.: 055743-816976 (M24-042P-US1-b) may be interchangeable. In general, the term module or engine refers to model or an organization of interrelated software components / functions. Further, while the audio authentication processes / services 1490 is shown as a standalone process, those skilled in the art will appreciate that this process may be executed as a routine or module within other processes.
[0196] It should be understood from the foregoing that, while particular embodiments have been illustrated and described, various modifications can be made thereto without departing from the spirit and scope of the invention as will be apparent to those skilled in the art. Such changes and modifications are within the scope and teachings of this invention as defined in the claims appended hereto.98974549.238
Claims
Attorney’s Docket No.: 055743-816976 (M24-042P-US1-b) CLAIMS What is claimed is:
1. A method for authenticating human speech during a speech event, comprising: accessing data associated with utterance of speech from multiple sources, the data including a combination of a plurality of signals comprising at least one physiological signal and at least one acoustic signal; extracting a plurality of features from the plurality of signals including spectral features and temporal features from the one or more signals, the spectral features defining frequency components of the one or more physiological signals and the temporal features representing changes in the one or more physiological signals over time; and identifying an association between feature values calculated from the data over the utterance of speech to verify the acoustic signal as originating from a human and at least one physiological signal as a biosignal from the human.
2. The method of claim 1, wherein the plurality of signals is captured by a microphone system comprising a microphone which measures the acoustic signal, and a time of flight (TOF) sensor that measures physiological vibrations from human speech production.
3. The method of claim 1, further comprising: extracting the features including speech features from the microphone and features from the combination of physiological signals captured during the speech event, including vocal fold vibration or articulator movement or heart beats or respiratory signals, capturing defining characteristics of the human speech production mechanism.98974549.239Attorney’s Docket No.: 055743-816976 (M24-042P-US1-b) 4. The method of claim 1, further comprising: extracting the features including extracting spectral features and temporal features from the one or more signals related to speech production, the spectral features defining frequency components of the one or more physiological signals and the temporal features representing changes in the one or more physiological signals over time.
5. The method of claim 1, further comprising: accessing physiological signals from the combination of signals captured via a TOF sensor; and generating a range-Doppler map from the physiological signals, the range-Doppler map comprising a two-dimensional representation defining a Doppler shift indicating velocity on one axis and range indicating distance on another axis, an intensity at each point in the range-Doppler map representing reflectivity of the sensor at that range and velocity informative as to human speech production.
6. The method of claim 1, further comprising: fusing the features as extracted to create a multi-dimensional profile of the speech event, the multi-dimensional profile defining an interplay of physiological and acoustic signals that characterize live human speech; and comparing the multi-dimensional profile with known profiles of authentic speech to assess a likelihood of the speech event being live and human.
7. The method of claim 1, wherein the plurality of signals includes physiological signals including blood volume changes in tissues of an ear, skin of the speaker captured by a PPG sensor, or vibrations or movements of the articulators.98974549.240Attorney’s Docket No.: 055743-816976 (M24-042P-US1-b) 8. The method of claim 1, further comprising: annotating a media file associated with the speech event with an identifier reflecting verification of the plurality of signals as originating from a human.
9. The method of claim 1, further comprising: generating an isolated representation of speech from the data by removing low-frequency physiological signals from the combination of signals.
10. The method of claim 1, wherein the plurality of signals includes: a vibrational signal transmitted through bone conduction or accelerometer during the speech event, and an acoustic signal associated with the speech event measured from a microphone.
11. The method of claim 1, further comprising: filtering the plurality of signals to a same frequency range; aligning the extracted features from the plurality of signals over time; and comparing feature variation over time between the at least one acoustic signal and the at least one physiological signal for the extracted features.
12. The method of claim 11, further comprising: identifying the association between feature values including confirming authenticity of the speech when the at least one acoustic signal and the at least one physiological signal demonstrate synchronous feature variation over time.98974549.241
Citation Information
Patent Citations
Methods and Apparatus for Dynamic Low Frequency Noise Suppression
US20160019910A1
Live event video stream service
US20200260150A1
System for mitigating the problem of deepfake media content using watermarking
US20210233204A1
Voice recognition using accelerometers for sensing bone conduction
US20230045064A1
Method and apparatus for focus-of-attention control
US8913103B1