Systems and methods for human voice verification
Patent Information
- Application Number
- CN202480075259.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-05-09
- Filing Date
- 2024-09-30
- Publication Date
- 2026-09-22
AI Technical Summary
[0020]前述示例大致上概括了根据本公开的示例的各个方面、特征和技术优势,以便可以更好地理解随后的详细描述。还应当领会,不要求在说明性示例方法、设备和计算机可读介质的上下文中描述的上述操作,并且可以排除一个或多个操作和/或可以包括本文讨论的其他额外操作。下文将描述附加的特征和优点。本文示出和描述的概念和具体示例可以容易地用作修改或设计用于实现本公开的相同目的的其他结构的基础。这种等同构造不脱离所附权利要求书的精神和范围。
Smart Images

Figure CN122804261A_ABST
Abstract
Description
[0001] Cross-references to related applications This non-provisional application claims the benefits of U.S. Provisional Application Serial No. 63 / 541773, filed September 29, 2023; U.S. Provisional Application Serial No. 63 / 541778, filed September 29, 2023; and U.S. Provisional Application Serial No. 63 / 645081, filed May 9, 2024; all applications are incorporated herein by reference in their entirety. Technical Field
[0002] This disclosure generally relates to techniques for distinguishing human users from automated systems or robots, and more particularly to microphone-based systems that enable recording devices to demonstrate that recorded acoustic signals originate from a person speaking directly into the microphone.
[0003] With the rise of AI and deepfake technology, we find ourselves in an era where distinguishing between real and synthetic media is becoming increasingly challenging. 2023 was a turning point for deepfakes. Consider the following events:
[0004] In April 2023, a song titled "Heart on My Sleeve" was released and went viral, using AI to simulate the music of pop stars Drake and The Weeknd. The song was so convincing that some even argued it surpassed the talent of real pop stars. However, the popularity and revenue potential of the AI-generated song raised concerns among music industry gatekeepers. Universal Music Group, the owner of Drake and The Weeknd's record label, immediately cited copyright infringement, demanding the platform remove "Heart on My Sleeve." Since then, several other examples of AI-manipulated songs have emerged. Legal disagreements regarding music and copyright law will continue to dominate the debate surrounding AI album covers within the music industry.
[0005] In April 2023, a mother in Arizona received a scam phone call from someone using AI voice cloning technology to mimic her daughter's voice. The caller claimed to have kidnapped her daughter and demanded a ransom. The mother initially believed the voice on the phone was her daughter's. This incident highlights the dangers of AI deepfakes and voice cloning technologies that can be used to deceive and commit fraud.
[0006] Deepfakes pose significant social challenges, including the potential to spread misinformation, incite panic or conflict, and undermine trust in the media. In the music industry, AI-generated songs could violate copyright laws, disrupt revenue streams, and challenge traditional gatekeepers. Furthermore, voice cloning technology could be used for malicious purposes such as fraud and deception, as demonstrated by the Arizona phone scam incident.
[0007] Unfortunately, these disturbing incidents expose a significant gap in the digital realm. We lack a reliable method to ensure the authenticity of the voices we hear, to verify that someone is indeed moving their lips behind the microphone, producing the recorded audio. Furthermore, there is an urgent need to make it available to everyone. Currently, for any piece of media—whether human-generated, AI-generated, or manipulated—it is disconnected from the method of creation the moment the source file is saved and uploaded to a distribution center (such as YouTube, social media, or television broadcasting). That is, with speech, we don't know if someone is speaking at the other end of the recording device, nor whether the speech is partially or entirely generated by AI.
[0008] Similarly, in packet-switched and circuit-switched communication channels, when a voice packet leaves the speaker's device, it becomes disconnected from the speaker. We cannot know whether there is a human speaker on the phone or whether the voice is synthesized.
[0009] The pressing question this invention aims to answer is: In this era of advanced deepfake technology, how can we maintain the authenticity of "man-made" media? A related but distinct problem is distinguishing human users from automated robots, a problem not necessarily unique to media creation. CAPTCHA, an acronym for "Common Turing Test for Fully Automated Systems That Differentiate Between Computers and Humans," is a technique used to differentiate human users from automated systems or robots. CAPTCHA is believed to be aimed at reducing automated form submissions, unauthorized access, web scraping, spam, and similar malicious activities carried out by computer programs on the internet. This technique typically involves a challenge-response test where the user must perform tasks that are, at least traditionally, difficult for machines to perform. These tasks include recognizing distorted text or images, performing simple mathematical calculations, and identifying specific objects within a visual scene or a series of images.
[0010] However, the reliability and effectiveness of existing CAPTCHA methods are being seriously challenged by recent advancements in artificial intelligence (AI) and machine learning (ML). Traditional CAPTCHA, which typically relies on vision-based tasks, is being bypassed by modern AI algorithms that have proven adept at image and text recognition. Optical character recognition (OCR) technology has shown significant progress, enabling machines to read distorted text—a capability once exclusive to human cognition. Similarly, convolutional neural networks (CNNs), a type of deep learning neural network, have demonstrated superior performance in visual recognition tasks, thus diminishing the effectiveness of image-based CAPTCHA.
[0011] Furthermore, AI models are learning to more closely mimic human behavior, enhancing their ability to solve problems and challenges originally designed for human intelligence. Advances in AI in natural language processing, predictive modeling, and cognitive computing further diminish the effectiveness of traditional CAPTCHA systems. These developments indicate that traditional CAPTCHA systems are becoming increasingly obsolete and insufficient to fulfill their primary purpose of distinguishing between human users and AI systems. Therefore, it is imperative to develop and implement new, more resilient approaches to rapidly evolving AI technologies.
[0012] These observations were taken into account when conceiving and developing various aspects of this disclosure. Summary of the Invention
[0013] This disclosure provides several examples of methods and systems associated with human voice verification / authentication. In the context of the disclosed methods, devices, techniques, apparatuses, systems, etc., the terms “operable as,” “configurable as,” and “capable” are used interchangeably.
[0014] In the first set of illustrative examples, the present invention concept includes a method for authenticating human speech, comprising the steps of: capturing data associated with a speech event from multiple sources, the data including a combination of signals related to both physiological and acoustic features of the speech event; extracting features from the captured signals that indicate human-generated speech in the event; analyzing features such as those extracted from the acoustic and physiological signals to authenticate the human-generated speech; and verifying the combination of signals relative to each other to further confirm the authenticity of the speech event.
[0015] In some examples, the above method can be translated into machine-readable instructions executable by at least one processor.
[0016] In a second example, the inventive concept includes a method for human verification via physiological signals emitted from a human body, comprising: accessing physiological data from a human user passively collected via a radar sensor; processing the physiological data to generate one or more physiological signals from the physiological data; extracting multiple features from the one or more signals, including spectral features and temporal features, the spectral features defining frequency components of the one or more physiological signals, and the temporal features representing the changes of the one or more physiological signals over time; and evaluating the multiple features via a model to determine whether the multiple features are aligned with patterns of a human user, the model being pre-trained on a human database of the same features.
[0017] In this example, the method may further include the following steps: accessing a voice signal captured via a microphone, the voice signal being elicited by a human user in response to a prompt, including capturing the voice signal while passively collecting physiological data via a radar sensor; and verifying the human user by a combination of the voice signal and one or more physiological signals.
[0018] In this example, the method may further include the following steps: temporally aligning one or more physiological signals and speech signals such that features extracted from both the radar sensor and the microphone correspond to a common time frame; extracting speech features related to vocal cord vibration and speech organ movement from both the speech signal and the one or more physiological signals, the speech features defining the characteristics of the speech generation mechanism; and determining whether the speech signal and the one or more physiological signals originate from the same source by analyzing whether the speech features from both the speech signal and the one or more physiological signals represent a common speech generation mechanism.
[0019] In some examples, the method may also include the following steps: filtering acoustic and physiological signals to the same frequency range, temporally aligning features extracted from acoustic and physiological signals such that the extracted features correspond to a common time frame, comparing temporal feature changes between acoustic and physiological signals for each corresponding set of extracted features from acoustic and physiological signals, and confirming the human authenticity of the speech when acoustic and physiological signals exhibit synchronous feature changes in time.
[0020] The foregoing examples broadly summarize various aspects, features, and technical advantages of the examples according to this disclosure in order to better understand the following detailed description. It should also be understood that the operations described above in the context of the illustrative example methods, apparatus, and computer-readable media are not required, and one or more operations may be excluded and / or other additional operations discussed herein may be included. Additional features and advantages will be described below. The concepts and specific examples shown and described herein can be readily used as the basis for modifying or designing other structures for achieving the same purpose of this disclosure. Such equivalent constructions do not depart from the spirit and scope of the appended claims. Attached Figure Description
[0021] Figure 1 This is a flowchart illustrating the process of human authentication via physiological signal measurement.
[0022] Figure 2 This is a flowchart of the process used to measure physiological signals from radar.
[0023] Figure 3 This is a flowchart for the process of detecting human users based on passive monitoring via radar.
[0024] Figure 4 This is a flowchart used to determine whether an active task is required during authentication.
[0025] Figure 5 It is a flowchart used for the verification process via active tasks.
[0026] Figure 6 This is a diagram illustrating the source filter model and inverse model in speech processing.
[0027] Figure 7 It is a graphical representation of physiological signals obtained through radar.
[0028] Figure 8 These are illustrations and related graphical representations of radar signals measured during human speech production.
[0029] Figure 9A shows the time-frequency representation of the radar signal acquired during the speech; and Figure 9B It is a time-frequency representation that removes low-frequency physiological signals.
[0030] Figures 10A and 10B are graphical representations showing the characteristic values calculated for the microphone and radar sensor in terms of sound emission.
[0031] Figure 11 This is a graphical representation of the normalized distribution (z-score) of the correlation when there is a mismatch between acoustic and radar signals.
[0032] Figure 12 The image shows the radar echo of a human speaker relative to a loudspeaker at three playback levels.
[0033] Figure 13 This is a simplified diagram illustrating the data fusion engine used for audio authentication; and Figure 14 This is a simplified diagram of an example computational system used to implement aspects of the systems and methods outlined in this paper.
[0034] Corresponding reference characters indicate the corresponding elements in the views of the accompanying drawings. The headings used in the figures do not limit the scope of the claims. Detailed Implementation
[0035] Until now, we have relied on the senses of our audience to verify the authenticity of media. Until recently, synthetic media sounded and looked synthetic. As the examples highlighted in the previous section demonstrate, this is no longer possible, as the quality and quantity of AI-generated media are already quite good and are rapidly improving.
[0036] The publicly available technology for addressing this problem has been passive—using AI to detect AI-generated content. While none of these solutions have been commercialized, the idea is to install pre-trained models at distribution centers (such as YouTube servers or broadcasting stations) that can detect AI-generated media, and then evaluate the media file via these models each time it is uploaded to determine if it is AI-generated. This approach has several problems. First, according to experts in the field, AI detection of deepfakes is a cat-and-mouse game unlikely to be successful in the long run.
[0037] Furthermore, even if the AI model at the distribution center works perfectly, the communication channel is vulnerable to streaming attacks. That is, the stream between the distribution center and the user's device could be compromised, and new audio samples could be streamed. In this scenario, the user still receives manipulated media. To detect such attacks using a passive approach, the AI model would have to be installed on every user's device to process all incoming samples. Power-constrained edge devices are unlikely to support continuous AI detection of all incoming media.
[0038] This article discloses various examples of core, proactive technological solutions to this problem. Such solutions involve a combination of physiological and acoustic signals used to verify human speech. In some examples, the concept is based on novel microphone technology that can certify on the recording device itself that the recorded acoustic signals originated from a person speaking into the microphone. This certificate can be linked to a media file and then accessed on the listener's side to provide verification that the audio was generated by a human. This microphone technology is based on a combination of a conventional MEMS (microelectromechanical systems) microphone (or other type of microphone) that measures acoustic signals and a radar sensor that measures physiological vibrations from the human during speech generation. These two signals are validated against each other during recording to verify the personality of the speech generator.
[0039] Example 1 This example discloses a next-generation CAPTCHA system that relies on direct measurement of physiological signals emitted from the human body to distinguish between humans and automated robots.
[0040] This system uses radar sensors to assess physiological signals when the user is passive or, when necessary, during active task management. Some existing websites have buttons that say "Click here to prove you are human" or similar variations without any subsequent tasks. They often rely on more subtle methods to distinguish between humans and robots. For example, many methods operate by monitoring mouse movement or touch events. The path and behavior of a human moving a mouse or interacting with a touchscreen provide evidence of human interaction. Therefore, if the button is clicked and there is a corresponding mouse movement or touch event that appears human, the website can reasonably conclude that the user is indeed human. Compared to traditional CAPTCHA systems, these methods are less disruptive to the user experience, but they may also be less secure, as sophisticated robots may be able to bypass them more easily. This system combines the less disruptive user experience of this type of CAPTCHA with improved security. This is achieved by directly measuring a person's physiological signals to authenticate their personality (rather than relying solely on mouse or touch movements).
[0041] Figure 1A flowchart illustrating the proposed process for human authentication via physiological signal measurements is shown. Figure 2 As shown, physiological signals of the user are passively collected via a radar sensor (Algorithm 1). These signals may include breathing, heart activity, lip movements, and / or vocal cord vibration. These signals are then used to detect whether the human is using [a specific device / mechanism]. Figure 3 The device shown is Algorithm 2. If it cannot be definitively determined whether a human user exists, then manage the active task (Algorithm 3), such as... Figure 4 As shown in the image.
[0042] The active task involves a speaking task that provides cues to the user. The user is asked to verbally respond to the cues, thereby generating both acoustic and radar signals. The acoustic signal is captured by a microphone. The radar signal originates from vocal cord vibrations, vocal organ movements, and heartbeat and respiration. Both signals are then processed to validate the radar signal relative to the acoustic signal (Algorithm 4), as follows: Figure 5 As shown in the diagram. This verification ensures that the task was executed correctly, that both signals originate from the same source, and that the signals were indeed generated by a human user. The end result of this process is to verify that the signals were generated by a human, thereby confirming the user's authenticity. Descriptions of each algorithm are provided in subsequent chapters.
[0043] This system provides a robust and reliable method for human verification, combining passive physiological signal detection with active task management where necessary. The use of both microphones and radar sensors ensures accurate and comprehensive user verification.
[0044] Generating a range-Doppler representation from radar signals suitable for measuring physiological signals from humans involves several preprocessing steps. Some common processing steps include: Signal Acquisition: The first step is to acquire the raw radar signal. This involves setting up the radar system and capturing the signal backscattered from the human body.
[0045] Range compression: This step involves processing the received radar signal to improve range resolution. This is typically achieved using matched filters or pulse compression techniques that maximize the signal-to-noise ratio (SNR).
[0046] Doppler Processing: The next step is to perform Doppler processing. This involves applying a Fast Fourier Transform (FFT) to the range compression signals to separate them based on their Doppler frequency shifts. The Doppler frequency shifts are frequency variations caused by the relative motion between the radar system and the human subject.
[0047] Clutter suppression: Clutter refers to unwanted radar reflections that can interfere with desired signals from human subjects. Clutter can originate from stationary objects in the environment or uninteresting human body parts. Clutter suppression techniques, such as high-pass filtering or adaptive filtering, are used to minimize these effects.
[0048] Range-Doppler Map Generation: After clutter suppression, the processed signal is used to generate a range-Doppler map. This is a 2D representation showing the Doppler frequency shift (indicating velocity) on one axis and the range (indicating distance) on the other axis. The intensity at each point in the map represents the radar reflectivity at that range and velocity.
[0049] Normalization: Distance-Doppler plots are typically normalized to a standard scale for easier analysis. This may involve subtracting the mean and dividing by the standard deviation, or scaling the values to between 0 and 1.
[0050] like Figure 3 As shown in the flowchart, Algorithm 2 is designed to detect whether a human user is interacting with the device. This algorithm operates on the data representation output from Algorithm 1, which provides a suitable format for extracting physiological signals from the acquired radar signals.
[0051] As shown in the figure, the first step of Algorithm 2 is feature extraction. This process involves refining the data representation into a more compact form by extracting spectral and temporal features. Spectral features capture the frequency components of physiological signals, while temporal features represent the changes of these signals over time. This dual extraction allows the algorithm to capture a comprehensive profile of the user's physiological signals.
[0052] Once the features are extracted, they are evaluated using a pre-trained model. This model is pre-trained on a human database with the same features, allowing it to effectively distinguish between human and non-human users. The model evaluates the extracted features and determines whether they align with patterns observed by human users.
[0053] This process is not a one-time operation. Instead, Algorithm 2 is designed to repeat the process periodically as the user uses the device. This ensures continuous monitoring and provides periodic (or discrete) predictions of whether the user is human.
[0054] Algorithm 3 is designed to determine whether an active task is needed for human authentication, or whether detection based on passive radar is sufficient. Algorithm 3 operates based on the output of Algorithm 2, which detects whether a human user is interacting with the device. If Algorithm 2 provides a discrete decision that a human has been detected, an active task is not needed. In this case, the user is authenticated as human solely based on passive radar detection. If Algorithm 2 determines that no human has been detected, an active task is activated.
[0055] If the output of Algorithm 2 is continuous, the decision to activate the active task is based on a threshold. If the continuous output values are below the threshold, indicating uncertainty in human detection, the active task is activated. If the output values are above the threshold, indicating high confidence in human detection, the active task is not needed.
[0056] Figure 3 The process shown ensures that proactive tasks requiring user interaction are activated only when needed, thereby enhancing the user experience while maintaining robust human authentication.
[0057] Active task activation: This is the initial stage where the active task is typically triggered when the passive detection from Algorithm 3 is uncertain or falls below a certain threshold.
[0058] Verbal task prompts: At this stage, the device prompts the user to perform tasks that require a verbal response. This could be a simple spoken phrase or answering a question.
[0059] Acoustic and radar sensor activation: When the user speaks in response to a prompt, the acoustic sensor (microphone) captures the acoustic speech signal, and the radar sensor captures the corresponding physiological signal.
[0060] Sensor time alignment: Signals from acoustic and radar sensors are aligned in time. This ensures that features extracted from both sensors correspond to the same time frame of the speech generation process. Depending on the application, the sensors can be located collaboratively or separately.
[0061] Feature extraction: Extracting features related to vocal cord vibration and speech organ movement from both acoustic and radar signals. These features capture the essential characteristics of the speech production mechanism.
[0062] Feature Comparison: This section compares the extracted features from the acoustic and radar signals. This comparison aims to ensure that the features represent the same speech generation mechanism, thus verifying that the microphone and radar signals originate from the same source.
[0063] Verification of the same speech source: The final stage is to verify the speech source. If the comparison of features indicates that the two signals may originate from the same source, the user is verified as human. If not, the process can be repeated or other measures can be taken.
[0064] Novel aspects of the invention The novel aspects of this system include: Using radar in CAPTCHA: The use of radar sensors in CAPTCHA devices is unique. Radar sensors passively assess physiological signals.
[0065] Passive and Active Human Detection: The device operates in both passive and active modes for human detection. In passive mode, it uses radar signals to assess physiological signals such as breathing, heartbeat, lip movements, and vocal cord vibrations. If passive detection is inconclusive, the device switches to active mode, where it manages the speaking task for the user. A microphone captures the user's speech during active tasks. This combination of radar sensors and microphones during active tasks allows for a more robust and reliable human detection and verification process.
[0066] The algorithm is used for signal processing and human detection: The device includes a processor that executes a series of algorithms for preprocessing radar signals, extracting and comparing features from radar and acoustic signals, and determining whether an active task is required. These algorithms, executed by the processor, enable the device to accurately determine whether the user is human and verify the source of the speech.
[0067] Verifying radar signals against acoustic signals: This device verifies radar signal measurements of vocal cord vibration and lip movement against speech signals. This process ensures that radar and microphone data have the same source, thereby further enhancing the reliability of human detection processes.
[0068] Periodic Human Detection: While the user is working on the device, the device periodically assesses the user's physiological signals. This feature allows for continuous monitoring and authentication of the user, which is particularly useful in preventing automated form submissions, unauthorized access, web scraping, spam, and similar malicious activities carried out on the network by computer programs.
[0069] Advantages and impacts relative to current technology The CAPTCHA device disclosed in this article exhibits several advantages over current technologies and can have a significant impact across various fields: Enhanced security: The combination of passive and active human detection mechanisms significantly enhances security. By using physiological signals and voice, the device can more accurately distinguish between humans and automated systems, thereby reducing the risk of unauthorized access or malicious activity.
[0070] Non-invasive authentication: The device's passive detection mode is non-invasive because it uses radar to assess physiological signals. This improves the user experience because it does not require active participation unless necessary.
[0071] Continuous authentication: The device's ability to periodically assess a user's physiological signals allows for continuous authentication. This is a significant advantage over traditional one-time authentication methods because it can detect whether an authenticated user has been replaced by an unauthorized user during a session.
[0072] Versatility: The device's ability to use both acoustic and radar signals makes it versatile and adaptable to various environments and conditions. It can be used in noisy environments where acoustic signals alone may be unreliable, or in situations where physiological signals can provide additional verification.
[0073] Impact on accessibility: This technology may also have a positive impact on accessibility. For users who may have difficulty with traditional CAPTCHA tasks, such as those with visual impairments, the use of physiological and auditory signals can provide an alternative verification method.
[0074] Integration potential: This technology has the potential to be integrated into a wide range of devices and systems, from smartphones and computers to secure access systems and online platforms, thereby enhancing security across a broad range of applications.
[0075] Example 2 This example discloses a core, proactive solution to the problem of voice cloning (deepfakes). This solution is based on a novel microphone technology that can prove, on the recording device itself, that the recorded acoustic signal originated from a human speaking into the microphone. This certificate can be linked to a given media file and then accessed on the listener's side to provide verification that the audio was generated by a human. The microphone technology is based on a combination of a conventional MEMS (microelectromechanical systems) microphone (or other type of microphone) that measures acoustic signals and a radar sensor that measures physiological vibrations from the mechanisms of human speech generation. These two signals are verified against each other during recording to prove the personality of the speech generator.
[0076] The relationship between speech generation and acoustic signals, vocal cord vibration, and speech organ movement Speech generation is a complex physiological process involving the finely coordinated interactions of various parts of the body. A simple physiological model of this process is as follows: Figure 6 As shown, sound is produced by the vibration of the vocal cords and further shaped as it passes through the oral and nasal cavities, thus producing speech. This process begins with the production of voiced sounds in the larynx, a process known as phonation. In this text, the vocal cords, also called vocal folds, located within the larynx, play a crucial role. When air from the lungs passes through the tightly closed vocal cords, it causes them to vibrate. These rapid, repetitive opening and closing of the vocal cords regulate airflow and generate a series of sound pressure waves.
[0077] These pressure waves form the primary acoustic signal, which carries the fundamental frequency of the speaker's voice. This frequency is largely determined by the physical properties of an individual's vocal cords, such as their length and tension, and the rate at which they vibrate. In adults, this fundamental frequency is typically in the range of approximately 80 to 240 Hz.
[0078] The primary acoustic signal then travels through the vocal tract—the throat, mouth, and nasal cavity. As it passes through these structures, the signal is shaped and filtered by the resonant properties of the vocal tract, producing a rich and complex acoustic signal that we perceive as a human voice.
[0079] Therefore, the speaker's physiological measurements are intrinsically linked to the generation of the speech signal. Consequently, the vibration of the vocal cords and the movement of the articulatory organs are directly related to the acoustic data captured by the microphone.
[0080] Source-filter model and inverse filtering in speech processing In the field of speech processing, the source-filter model serves as the fundamental framework for conceptualizing speech generation and processing speech signals. According to this model, speech generation involves two main parts: sources and filters. The source components refer to the voiced or unvoiced excitations generated in the larynx, essentially the vocal cord vibrations of voiced sounds. The filters represent the vocal tract configuration that occurs and shapes these source excitations into a discernible speech pattern.
[0081] The acoustic signal perceived by a speaker is a convolution of the source excitation and the response of the vocal tract filters, producing the complex and unique speech output we recognize as personal speech. The vocal tract imposes its resonant properties on the source signal, generating a series of formants or spectral peaks in the acoustic output. Mathematically, this can be modeled as the convolution of the source signal (vocal cord vibration) with the vocal tract filters (movements of the vocal organs) (see [link to relevant documentation]). Figure 6 (Top box diagram in the middle).
[0082] Inverse filtering is a technique used to deconvolve or separate the source and filter components from an observed speech signal. Essentially, it aims to reverse the filtering effect of the vocal tract to obtain the original source excitation, which, in the case of voiced speech, is the vocal cord vibration signal (see [link to original text]). Figure 6 (See the bottom block diagram). This can be achieved through linear predictive coding (LPC) or cepstral analysis, among other methods, which estimate the vocal tract filters and subsequently “eliminate” their effects from the acoustic signal. By successfully applying inverse filtering, vocal cord vibration signals and the configuration of the articulatory organs can be extracted from the acoustic speech signal, providing valuable information about the physiological aspects of speech production; namely, the vibration of the vocal cords and the movement of the articulatory organs (lips, jaw, soft palate).
[0083] The estimated source signal obtained through inverse filtering is directly related to the vibration of the vocal cords, and the estimated filtered signal is directly related to the movement of the articulatory organs (lips, tongue, jaw, soft palate). In other words, if the vibration of the vocal cords or the movement of the articulatory organs is directly measured, the signal will closely match the source signal estimated from the measured acoustic speech signal.
[0084] Human physiological signals during speech measured by radar Radar, representing radio detection and ranging, is a technology that uses radio waves to detect and determine the distance, direction, and velocity of objects. It works by transmitting radio signals that are bounced off any object in their path, and the reflected signals are captured by the radar system. Radar technology precisely measures several key parameters of the detected object. It measures the "range" or "distance" to the object by calculating the time it takes for the radio waves to return after bouncing off the object. It determines the object's "velocity" or "rate" by utilizing the Doppler effect, which indicates motion toward or away from the radar, based on the frequency shift of the returned signal.
[0085] It is feasible to use a radar system to target a speaking individual and obtain a range of physiological signals by measuring oscillations emanating from the individual's body corresponding to various physiological processes. These include oscillations caused by the rhythmic contraction and relaxation of the heart, respiratory movements caused by the expansion and contraction of the lungs, vocalization movements including lip and jaw dynamics, and the vocal cord vibrations essential for phonation.
[0086] Figure 7 This diagram illustrates each of these physiological processes, the unique frequency band they fall into, and the amplitude of their corresponding radar signals. Breathing typically falls within the 0.2 to 0.5 Hz range, reflecting the slow, regular expansion and contraction of the lungs. Cardiac activity, more specifically the heartbeat, occurs in a slightly wider and higher frequency band of 0.2 to 2 Hz. The lip movements during speech occur in the 3 to 7 Hz range, attributed to the rapid and complex movements involved in speech production.
[0087] Furthermore, vocal cord vibrations occur in a significantly higher frequency band, from 80 to 240 Hz. These vibrations are caused by the rapid, repetitive opening and closing of the vocal cords during phonation, and are essential for the generation of voiced sounds.
[0088] While there is some overlap in the frequency bands of respiration and heartbeat, vocal cord vibration lies in a distinctly separate band. This band does not overlap with any other human-generated physiological signals, making it uniquely identifiable. This separation simplifies the task of isolating the vocal cord vibration signal from the comprehensive physiological data collected by radar systems.
[0089] Respiratory movements typically generate the largest radar signals. This is due to the considerable chest wall displacement and extensive internal lung movements during respiration. The absolute scale of this action results in a large radar reflection signal.
[0090] Although cardiac activity is regular, it generates a smaller radar signal because it involves relatively small physical displacements. The heart is a smaller organ, and its movement during contraction is limited compared to the larger-scale respiratory movements.
[0091] While lip and jaw movements during speech have a relatively higher frequency, they produce a relatively smaller radar signal. In contrast, the body parts involved in phonation are smaller and exhibit limited movement, even when producing complex speech, thus reflecting less radar signal back to the receiver.
[0092] Finally, vocal cord vibration, despite its high frequency, produces minimal radar signal amplitude. The vocal cords are small, their movement during speech has a very small amplitude, and they are located deep within the body. These factors result in a very small radar cross-section, and the movement produces only a tiny variation throughout the reflected radar signal. However, the unique frequency band of vocal cord vibration enables their detection and isolation for analysis, despite their low signal amplitude.
[0093] This is Figure 8 As shown, the radar system is positioned in front of the human speaker. A miniature 77 GHz millimeter-wave (mmWave) radar sensor is used here for simultaneous measurement of human vocalization and heart sounds while the test subject is uttering simple words or sentences. The speaker is positioned 0.7 meters away from the microphone and radar sensor. Figure 8 The image shows radar characteristics acquired in the time and frequency domains. The resulting spectrum shows a clear separation between low-frequency physiological signals (respiration, heartbeat, and vocal organ movements) and high-frequency physiological signals (vocal cord vibration).
[0094] Figure 9A further illustrates the time-frequency representation of the radar signal acquired during speech. This figure shows the vocal cord vibrations associated with the speech signal at higher frequencies and the low-frequency physiological signal. The low-frequency physiological signal is then removed, producing an isolated representation of the speech signal captured by the radar sensor, as shown below. Figure 9B As shown in the figure, this diagram illustrates the changes in pitch and pitch harmonics over time. This signal can then be converted back to the time domain to generate a time-domain estimate of the resulting speech signal. If the speech signal is also captured by a microphone, it allows for... Figure 9B The radar signal shown is compared with the acoustic signal acquired via a microphone.
[0095] Acoustic signals compared to radar signals Synchronizing and verifying voice and radar signals to ensure they originate from the same individual involves several processing stages. Data collected from both the radar sensor and microphone represents a combination of an individual's physiological and acoustic characteristics. Then, by combining these data streams and verifying that they originate from the same source, the system can automatically and algorithmically verify that the measured voice signal was generated by a human. This includes the following steps: Data acquisition and time synchronization: Simultaneous data collection is performed using radar sensors and microphones. The two sensors are synchronized via a common clock or by timestamping the data from both sensors.
[0096] Speech signal processing: Analyzing speech signals recorded by a microphone using various digital signal processing techniques. This may involve inverse filtering processes that isolate the source and filters, as well as extracting features of the speech signal, such as fundamental frequency, formants, spectral characteristics, and temporal features. This information represents the unique characteristics of an individual's speech.
[0097] Radar signal processing: Radar signals representing the physiological movements associated with speech production are processed. This involves steps as explained above, such as range-Doppler processing and signal filtering. Specifically, high-frequency vocal cord vibrations and articulation movements are of interest because they are closely coupled to the speech features extracted from the source and filters.
[0098] Cross-modal comparison and verification: Key features from both radar and speech signals, such as vocal cord vibration patterns and frequencies, the timing of articulation movements, and other speech characteristics, are then compared. By correlating these two datasets, it can be verified that they indeed originate from the same person. This cross-modal integration can be performed using advanced signal processing or machine learning algorithms. These algorithms compare and contrast feature sets from two signals, thereby establishing relationships between them. High correlation or consistency may suggest that the radar and speech signals originated from the same person.
[0099] Continuous monitoring and verification: By continuously monitoring and comparing radar and speech signals, it is possible to prove the continuity of the same speaker. Any significant deviations in the comparative metrics may indicate a change in the speaker or a change in the acoustic signal.
[0100] In Figure 8 In setups similar to those used in 9A and 9B, a microphone and radar sensor are used to collect and record the speaker's vocalizations ("She's been putting your dark suit in greasy laundry water all year round"). The acoustic and radar signals are filtered to the same frequency range (0 to 500 Hz) and then time-aligned by finding peaks in the cross-correlation. A series of features are extracted from each signal to track the harmonics, time variations, and several spectral parameters of both signals. Figure 10A illustrates how the normalized harmonic characteristics change synchronously across both modes when there is a match between the acoustic and radar signals. Figure 10B This illustrates how identical features do not change synchronously when there is no match between the acoustic and radar signals. It is clear from Figure 10A that features across both modalities are tracked in time, making cross-modal verification possible. That is, the high correlation between features suggests that it is possible to verify the speaker's vocal cord vibrations and articulatory organ movements, sourced from the acoustic signal recorded by the microphone and measured by the radar sensor. Figure 10B The lack of correlation between the two signals suggests that, in this case, such cross-modal verification can be performed with high confidence.
[0101] To further evaluate the confidence level at which cross-modal verification can be performed, a distribution of feature correlations with mismatches between acoustic and radar signals was generated. We randomly selected from an existing corpus of 350 audio samples that were not collected at the same time as the radar-captured samples. That is, these speech samples do not match samples collected by radar and spoken by different speakers saying different sentences. For each audio sample, we computed feature values for both acoustic and radar signals and calculated the correlation between the two sets of features for each speech utterance. The normalized distribution of the correlation (z-score) is shown in the figure. Figure 11 The values are shown in blue. In the same figure, the true correlation (z-score) when two signals match is shown in red. As this figure shows, even using a single feature, the true correlation value of that feature deviates from the mean by almost 4 standard deviations, indicating that even for a single sound, it is possible to determine a match between an acoustic signal and a radar signal with a very high confidence level.
[0102] Link the certificate to the media After verification, the resulting identity authentication certificate can be linked to media to verify its origin and ensure that it was human-generated and not tampered with or forged. Several methods can be used to link the certificate to media. These methods include, but are not limited to: Digital signing using existing software: Some media editing or security software may allow you to digitally sign media files directly. Tools tailored for digital rights management (DRM) or secure media disposal may include this functionality.
[0103] Perform hashing and signing: Create a hash for the audio file, which is a unique numerical value representing the file's content. Sign this hash using a private key, thus creating a digital signature. The signature and public key can be stored alongside the audio file or embedded in the metadata, allowing recipients with the corresponding public key to verify the signature.
[0104] Embedding a signature in the file's metadata: Metadata in audio files often stores additional information, such as artist details and album names. A digital signature can also be embedded in this metadata, although care must be taken to ensure that this does not interfere with the file's playback.
[0105] Embedding signatures in watermarks: Applying digital watermarks to media. This can be a visible or invisible pattern that includes authentication information.
[0106] Blockchain technology: Utilizing blockchain to create verifiable and immutable media records. For example, platforms like Ethereum can be used to write smart contracts that encode the authenticity of media within the blockchain.
[0107] Using a standard file signing tool with a wrapper: If the audio file is contained in a wrapper format (similar to a ZIP file) that can be digitally signed using a standard file signing tool, the entire package, including the audio content, can be signed. The recipient can then verify the package's signature, thereby verifying the authenticity of the audio file inside.
[0108] Other customized solutions: Custom solutions can be developed to digitally sign media files based on specific requirements. For example, radar signatures can be linked to audio files via existing tools such as watermarking, blockchain, etc. Then, during authentication, voice features can be extracted and evaluated against radar signatures to ensure a match.
[0109] The complexity of vocal cord vibration and the challenge of deceiving radar signal acquisition. The process of vocal cord vibration is a highly complex physiological phenomenon that directly results in the acoustic signals produced when humans speak. In adults, during the production of voiced sounds, the vocal cords, located in the larynx, oscillate at high frequencies ranging from approximately 80 to 240 Hz. However, the vibration is not a simple periodic motion, but a complex three-dimensional mucosal undulation involving longitudinal, horizontal, and vertical displacements. These vibrations are influenced by many factors, including the length, tension, and mass of the vocal cords, subglottic pressure, vocal tract structure, and individual physiological characteristics.
[0110] The radar signals generated by vocal cord vibrations, like frequency modulation, carry these complex motion patterns, as well as other physiological characteristics of the speaker's vocal anatomy. When combined with other physiological signals captured by radar sensors, such as breathing and heart activity, these characteristics form a unique "human imprint," which is extremely challenging to replicate or deceive.
[0111] To attempt to deceive such radar signals, it is necessary not only to mimic the unique frequency characteristics of a speaker's vocal cord vibrations but also to replicate the specific physiological features that affect the radar cross-section. Furthermore, the complexity of the three-dimensional, high-frequency mucosal undulations of the vocal cords further amplifies this challenge. Given the currently available technology and current understanding of human physiology and radar signal processing, successfully deceiving radar signals generated by vocal cord vibrations presents a considerable challenge.
[0112] A simple attempt to deceive a radar system might involve placing a speaker in front of it, causing the radar to measure the vibration of the speaker's diaphragm while the microphone measures the acoustic signal. To detect such attempts, machine learning algorithms can be implemented to ensure a human is in front of the microphone (rather than a synthetically vibrating diaphragm). Through machine learning, the algorithm can be trained on a large amount of radar signals from human speech, learning to recognize subtle patterns and characteristics of radar echoes typical of a living human speaker during speech. This model can then be used to evaluate whether the acquired signal is genuine. (See below...) Figure 12 In the figure, we show radar echoes from a human speaker and echoes from a vibrating speaker diaphragm (at three playback levels). The figure illustrates the significant differences between the human echo and the speaker echo. The first significant difference is the difference in rounds when no voiced speech is generated. For the speaker echo, the background noise level is much higher. The second significant difference is the larger amplitude of the speaker echo (more red in the speaker echo). The third significant difference is the difference in high-frequency harmonic echoes. As shown, the radar echoes of high-frequency harmonics are more pronounced, especially at mid-to-high volume levels. This figure provides evidence that it should be possible to learn statistical models to capture radar echoes from human speakers via machine learning.
[0113] Novel aspects of the invention The novelty of this invention lies in its unique method for verifying audio authenticity, which combines traditional acoustic signal recording with physiological vibration measurements. Some key innovative elements of this method include: Dual-sensor technology: The integration of a traditional MEMS microphone with a radar sensor is a novel approach. While the microphone records acoustic signals, the radar sensor measures physiological vibrations from the mechanisms of human speech production. This dual-sensor setup provides a more comprehensive and accurate way to verify the presence of a human speaker. Depending on the application, the two sensors can be used collaboratively or separately.
[0114] Real-time verification: The ability to verify the personality of a speech generator during recording is a significant advancement. This real-time verification can prevent the destructive effects of deepfakes or other synthetic audio from the outset.
[0115] Certificate Linking: The ability to link authenticity certificates to the media file itself is another novel aspect. This feature allows listeners direct access to verification information, thus providing an immediate guarantee of the audio's authenticity.
[0116] Physiological vibration measurement: Measuring the physiological vibrations associated with human speech production using radar technology is a unique aspect. This method provides biological confirmation of the human speaker, which is difficult to replicate artificially.
[0117] On-device verification: The fact that the verification process occurs on a computing device while the acoustic signal is being recorded is also innovative. This on-device processing enhances the security and efficiency of the verification process.
[0118] Advantages and impacts relative to current technology: The proposed voice authentication system offers several advantages over current technologies and could have a significant impact: Authenticity Verification: A key advantage of this invention is its ability to verify the authenticity of audio content. This technology can provide an authenticity certificate proving that the audio was generated by a human, not AI or a deepfake. This feature is becoming increasingly important in an era where deepfakes and synthetic media are becoming more prevalent.
[0119] Preventing misinformation: By verifying the source of the audio signal, this invention can help prevent the spread of misinformation and false information. It can be a key tool for maintaining the integrity of news broadcasts, political speeches, and other shared communications.
[0120] Security and Privacy: This invention can enhance the security of communication systems. For example, it can prevent voice cloning scams because it ensures the speaker is a real person. It can also be used in secure voice biometric systems, adding an extra layer of security.
[0121] Legal and forensic applications: In a legal context, this invention can help verify the authenticity of audio evidence. It can also be used in forensic investigations to identify the speaker in an audio recording.
[0122] Music and Entertainment Industry: In the music and entertainment industry, this invention can help protect artists' rights by ensuring that audio recordings are created by human artists rather than artificially generated.
[0123] The impact of this invention could be profound, potentially transforming how humans trust and interact with audio media. It could restore trust in digital media by providing a reliable method for verifying authenticity, something that is becoming increasingly important in the digital age.
[0124] Application of this technology Ensuring secure communications: Scams based on voice cloning are emerging. Performing human verification checks during conversations can ensure that the person on the other end is a human speaker, not an AI bot. Beyond consumer applications, this technology also has significant potential applications in the defense sector.
[0125] Next-Generation CAPTCHA: With advancements in AI technology, current CAPTCHA techniques will soon become obsolete. AI is finding it increasingly easy to detect distorted numbers / text or identify specific objects in visual scenes. We hypothesize that next-generation CAPTCHA will require human verification via physiological measurements. Radar-based systems would allow for effective verification.
[0126] Certifying music, podcasts, and other forms of content as human-generated: As generative AI improves, AI-generated entertainment content may surpass human-generated online content. In this scenario, differentiating between human-generated and AI-generated media presents opportunities for new business models (such as different monetization strategies for human and AI content).
[0127] Celebrities / Influencers Protecting and Monetizing Their Voices: As AI models improve, synthetic media from celebrities and influencers will be easier to generate. Synthetic voices have already been used to generate celebrity hate speech (e.g., see https: / / www.theverge.com / 2023 / 1 / 31 / 23579289 / ai-voice-clone-deepfake-abuse-4chan-elevenlabs) and AI-generated songs using the voices and styles of famous artists (https: / / www.npr.org / 2023 / 04 / 21 / 1171032649 / ai-music-heart-on-my-sleeve-drake-the-weeknd). Validation of human-generated content will allow celebrities to protect their voices and their brands. There is also a very promising future market. For example, some celebrities have developed AI clones in which they allow fans to interact with the clones (e.g., see https: / / www.washingtonpost.com / technology / 2023 / 05 / 13 / caryn-ai-technology-gpt-4 / ). Validation of human-generated content allows for tiered pricing, where genuine interactions are more expensive than interactions with AI-generated bots.
[0128] Verify virtual assistants: Human users can be verified before any commands are executed by a virtual assistant (such as Alexa or Siri), providing an extra layer of security when communicating with smart devices.
[0129] Personality verification / humanity verification: Personality verification / humanity verification is a method to verify that a user is a unique, real human being and not a bot or fake account. There are two main forms of personality verification: social graph-based and biometric. Social graph-based personality verification relies on some form of guarantee, while biometric personality verification involves verifying certain physical or behavioral traits that distinguish humans from bots and individual humans from one another. Voice biometrics can play a role in biometric personality verification by verifying the identity of the registrant. A biometric speech recognition system captures new speech samples, creates a template from the samples, and compares it to a registered template. A strong match between the templates indicates that the same person spoke both samples, thus verifying that person's identity. The proposed technique can be used in such a system to provide additional evidence that the user generating the speech is human.
[0130] Example 3 This example outlines methods and systems for authenticating live human speech and preventing deepfake audio imitation. It integrates various biometric sensors to capture physiological and acoustic signals indicative of the production of natural human speech.
[0131] The system includes a wearable device capable of authenticating live human speech by detecting a set of biometric and acoustic signals. It is generally adaptable and can take the form of an in-ear device, headphones, or a head-mounted device. The device incorporates a suite of sensors, including a bone conduction microphone, photoplethysmography (PPG) sensors, an accelerometer, a gyroscope, and a temperature sensor. These sensors work together to capture vibrational, physiological, and thermal signals specific to human speech produced by an individual. The goal is to provide a robust method for distinguishing live human speech from non-human sources, with potential applications in security, personal identification, health monitoring, and interactive technologies.
[0132] Deepfake technology necessitates the development of reliable methods to authenticate multimedia content, particularly the authenticity of human speech. The systems and methods outlined in this paper are a response to the need for robust, real-time methods capable of distinguishing between genuine human speech and deepfake-generated audio. This approach leverages recent advances in sensor technology and algorithmic analysis to assess the inherent biometric and acoustic signals of natural speech, thus providing a new layer of defense against emerging threats to deepfakes.
[0133] Consistent with the previous example, once human speech is verified, it can be linked to an associated media file so that listeners can verify that the speech was generated by a human.
[0134] contribute The system and methodological contributions summarized in this paper are as follows: Multi-sensor fusion: This method integrates various sensors in a wearable format, such as bone conduction microphones, PPG sensors, accelerometers, gyroscopes, and thermal sensors, to capture a comprehensive dataset of physiological and acoustic signals associated with speech generation. This fusion of multiple data types for on-site speech verification is unique.
[0135] Real-time deepfake detection: By analyzing physiological signals that are difficult to synthesize, such as subtle blood flow changes associated with speech or unique patterns of jaw movement, this system can detect deepfakes in real time. This is a major improvement over existing technologies, which typically only analyze audio and lack real-time processing capabilities.
[0136] Wearable format: This method is designed to be integrated into wearable devices (such as hearing devices, headsets, and combinations) to provide users with a discreet and convenient form factor.
[0137] Physiological signal authentication: Existing technologies for deepfake detection primarily focus on the acoustic analysis of speech. In contrast, the system outlined in this paper authenticates speech through underlying physiological processes, providing a novel verification method that is inherently more secure than current and future deepfake technologies.
[0138] Enhanced Security: By leveraging physiological data and acoustic signals, the system outlined in this paper offers an inherently higher level of security against deception attempts, including sophisticated deepfake audio imitations. Current deepfake detection techniques that rely on acoustic output exhibit inconsistent performance and are being surpassed by the rapid advancements in deepfake generation.
[0139] Real-time authentication: Unlike many current technologies that operate on pre-recorded samples and lack the ability to verify them instantly, this method enables real-time processing and authentication of live voice, which is crucial for dynamic security systems and interactive applications.
[0140] Versatility of applications: The wearable nature of this device allows for a wide range of applications, from mobile security to healthcare monitoring, without requiring the special environments or settings that some current technologies require.
[0141] Non-invasive and user-friendly: The system outlined in this article is designed to be non-invasive and user-friendly, seamlessly integrating into everyday devices such as headphones or glasses, thus providing users with a frictionless experience.
[0142] Robustness to environmental changes: By employing the ability to capture and analyze physiological signals, the systems outlined in this paper are not susceptible to environmental noise and changes that may only affect acoustic-based technologies.
[0143] Core components and methods sensor array The system employs a sensor array that can be configured into various wearable forms, including but not limited to in-ear devices, neckband devices, headphones, or glasses. These sensors include, but are not limited to: Bone conduction microphone: Used to capture vocal cord vibrations transmitted through the skull.
[0144] Photoplethysmography (PPG) sensors: used to detect changes in blood volume in tissues of the ear or skin, corresponding to changes in human skin and potentially to the force of speaking.
[0145] Accelerometers and gyroscopes: used to monitor micro-motions associated with vibrations during speech production, as well as micro-motions associated with jaw and head movements during speech.
[0146] Thermal sensor: Used to measure the temperature near the surface of the skin, as the temperature will fluctuate with speech due to changes in blood flow.
[0147] High-fidelity microphone: Used to capture audio signals inside the ear canal or near the mouth using ambient noise reduction technology.
[0148] Signal processing This system employs advanced signal processing algorithms to analyze data collected from multiple sensing modalities, preparing it for a data fusion engine. This includes extracting features from different sensing modalities.
[0149] Temporal characteristics: The temporal and sequential characteristics of speech-related physiological changes.
[0150] Spectral characteristics: Decomposing an audio signal into its constituent frequencies can be used to identify acoustic, vibrational, and other physiological cues associated with speech production.
[0151] Data Fusion Engine Figure 13 The data fusion engine 1300 is shown, which can use on-board computing devices of various wearable devices (e.g., as outlined in this document) Figure 14 This can be achieved using a computing device (1400) or by communicating with various wearable devices in other ways. The data fusion engine ensures the human authenticity of the audio. This can be done in two ways. First, the collected signals are fused using features extracted in the previous step to create a multi-dimensional profile of the speech event. This profile is compared to known profiles of real speech to determine the likelihood that the speech is live and human. Next, several sensing modalities are evaluated relative to each other to ensure the presence of live human speech. For example, vibrational features (measured from a bone conduction microphone), vocal organ movement (measured from an accelerometer), and acoustic output (measured from a microphone) can be verified relative to each other to ensure they share the same source (a human speaker).
[0152] Output and Integration This system provides an output indicating the human authenticity of the voice. This output can be integrated into security systems that require voice authentication, communication devices that verify the speaker's identity, media platforms that filter deepfake content, or content authentication platforms that verify genuine human voice.
[0153] user interface The system includes a user interface for calibrating, monitoring, and alerting users about their certification status. It can be a standalone application or integrated into an existing software ecosystem.
[0154] Adaptability and Implementation Examples The device's form factor is designed to suit discreet, comfortable, and user-specific applications, ranging from consumer electronics to high-security communication devices.
[0155] Sensor arrays and processing algorithms can be customized for different security levels, from basic authentication to highly secure environments where voice spoofing poses a significant risk.
[0156] This method can be adapted to different populations and languages, and can adapt to changes in speech patterns and physiological responses.
[0157] This system can be designed for real-time on-site voice authentication and post-analysis of recorded speech to retrospectively verify its authenticity.
[0158] Relevant characteristics of signal processing components Possible features that can be extracted from each sensing modality during speech to generate a human silhouette include the following (this is not an exhaustive list): Bone conduction microphone: Vibrational energy: The energy of vocal cord vibrations generated by indicative speech transmitted through the skull.
[0159] Fundamental frequency (F0): Tracks the pitch of the speaker's voice, which varies naturally during live human speech.
[0160] Photoplethysmography (PPG) sensor: Blood volume pulse (BVP): measures heart rate variability, which can be correlated with the force of speaking and emotional state.
[0161] Blood oxygen saturation: Monitors changes in blood oxygen levels, which can vary with breathing patterns associated with speaking.
[0162] Accelerometer and gyroscope: Jaw movement patterns: Capture the dynamics of jaw movements during speech, which are unique to the individual and difficult to replicate in synthesized speech.
[0163] Head movements: Analyzing head movements that accompany natural speech can provide clues to the authenticity of speech actions.
[0164] Thermal sensor: Temperature variability: Measures the fluctuations in skin temperature near the mouth or throat, as skin temperature can change as one speaks due to changes in blood flow.
[0165] Breath temperature: This measures the temperature of the air exhaled during speech, which can indicate the level of speech produced on the spot.
[0166] High-fidelity microphone: Spectral features: Extracting features such as Mel-frequency cepstral coefficients (MFCCs) that capture the unique timbre quality of the speaker's voice.
[0167] Temporal dynamics: Analyzing the temporal and duration characteristics of phonemes, syllables, and words in natural speech rhythm.
[0168] Data Fusion Engine The data fusion engine integrates and analyzes data from multiple sensors to authenticate the human origin of the voice signal. The steps include: Multi-dimensional contour creation: The engine first aggregates features extracted from a sensor array, each providing a unique perspective on a speech event. These include vibrational data from bone conduction microphones, physiological signals from PPG sensors, motion data from accelerometers and gyroscopes, temperature changes from thermal sensors, and acoustic signals from high-fidelity microphones. These features can include energy, blood volume pulse, jaw movement patterns, temperature variability, and spectral characteristics. These features indicate the natural mechanisms of human speech production.
[0169] These extracted features are fused to create a multidimensional profile of the speech event. This profile encompasses the complex interaction of physiological and acoustic signals characterizing live human speech.
[0170] Comparison with known contours: The multidimensional contours are compared with a contour database associated with real human speech. This database was built from a large number of live speech event samples, which capture a wide range of natural variations in speech production across different individuals and environments.
[0171] The comparison process involves machine learning algorithms that capture subtle differences characterizing live human speech. Example machine learning models could include Gaussian mixture models that model the distribution of multidimensional contours.
[0172] Cross-modal verification: In addition to comparing the multidimensional contours relative to known contours, the engine also performs cross-modal verification. This involves evaluating the consistency and correlation between different sensing modalities.
[0173] For example, vibration signal characteristics captured by a bone conduction microphone are examined relative to motion data (vocal organ movement) from an accelerometer and acoustic output from a high-fidelity microphone. The engine evaluates whether these different data streams corroborate each other, indicating that they originate from the same live speech event.
[0174] This step is crucial for detecting deepfake attempts that may convincingly replicate one aspect of speech production (such as acoustic output) but fail to accurately mimic the overall and interconnected characteristics of live human speech.
[0175] Output generation: Based on the results of contour comparison and cross-modal verification, the engine generates output indicating the likelihood that a speech event is real or human. This output can be a binary decision (real or unreal), a confidence score, or a detailed report emphasizing the analysis results.
[0176] The output can then be used by various downstream systems that require voice authentication, such as security systems, communication devices, and content verification platforms, to make informed decisions about the authenticity of the voice.
[0177] The verification signal can also be embedded in the speech signal as a watermark, added to the record as metadata, or used to encrypt and sign audio as real human speech.
[0178] discuss This disclosure outlines systems and methods for authenticating live human voice by analyzing physiological and acoustic signals using sensor arrays integrated into wearable devices. This approach addresses the urgent need for reliable verification in the face of advanced deepfake technologies that pose increasing threats to personal safety, information integrity, and privacy. The commercial potential of this invention lies in its applications across various fields, including security, personal electronics, health monitoring, and communications. It provides a novel solution for companies seeking to enhance the authenticity of digital interaction and voice control systems and offers promising prospects for a wide range of licensing opportunities in areas where verifying human presence is crucial.
[0179] Computer-implemented system Figure 14 This is a schematic block diagram of an example computing device 1400 that can be used with one or more embodiments described herein, such as as a component of a system that includes wearable devices outlined herein and implements some aspects of the methods outlined herein.
[0180] The computing device 1400 includes one or more network interfaces 1410 (e.g., wired, wireless, PLC, etc.) interconnected via a system bus 1450, at least one processor 1420 and memory 1440, and also includes a power supply 1460 (e.g., battery, plug-in, etc.).
[0181] The computing device 1400 may include or otherwise communicate with a display device 1430, which communicates user interface and / or the output of the methods outlined herein to a user.
[0182] One or more network interfaces 1410 include mechanical, electrical, and signaling circuitry for transmitting data via a communication link coupled to a communication network. Network interfaces 1410 are configured to transmit and / or receive data using various communication protocols. As shown in the figures, for simplicity, blocks representing network interfaces 1410 are shown, and it should be understood that these interfaces can represent different types of network connections, such as wireless and wired (physical) connections. Network interfaces 1410 are shown separate from power supply 1460; however, it should be understood that interfaces supporting PLC protocols can communicate through power supply 1460 and / or can be integrated components coupled to power supply 1460.
[0183] Memory 1440 includes multiple storage locations addressable by processor 1420 and network interface 1410 for storing software programs and data structures associated with the embodiments described herein. In some embodiments, computing device 1400 may have limited memory or no memory (e.g., no memory for storage other than programs / processes running on the device and associated caches). Memory 1440 may include instructions executable by processor 1420 that, when executed by processor 1420, cause processor 1420 to implement some aspects of the systems and methods outlined herein.
[0184] Processor 1420 includes hardware elements or logic adapted to execute software programs (e.g., instructions) and manipulate data structures 1445. Operating system 1442 (parts of which typically reside in memory 1440 and are executed by the processor) functionally organizes computing device 1400, particularly by invoking operations that support software processes and / or services executing on the device. These software processes and / or services may include an audio authentication process / service 1490, which may include aspects of the methods and / or various modules implemented herein. Note that although the audio authentication process / service 1490 is shown in centralized memory 1440, alternative embodiments provide processes to operate within network interface 1410, such as components of the MAC layer, and / or as part of a distributed computing network environment.
[0185] It will be apparent to those skilled in the art that other processor and memory types, including various computer-readable media, can be used to store and execute program instructions related to the techniques described herein. Furthermore, while various processes are illustrated in the description, it is clearly understood that these processes can be embodied as modules or engines configured to operate according to the techniques described herein (e.g., according to the functionality of similar processes). In this context, the terms module and engine are used interchangeably. Generally, the terms module or engine refer to a model or organization of related software components / functions. Additionally, although the audio authentication process / service 1490 is shown as a standalone process, those skilled in the art will appreciate that it can be executed as a routine or module within other processes.
[0186] It should be understood from the foregoing that, although specific embodiments have been shown and described, it will be apparent to those skilled in the art that various modifications can be made thereto without departing from the spirit and scope of the invention. Such changes and modifications are within the scope and teachings of the invention as defined by the appended claims.
Claims
1. A method for authenticating human speech, comprising: Data associated with speech events is captured from multiple sources, the data including a combination of signals related to the physiological and acoustic characteristics of the speech events; Extract features of human-generated speech from the captured signal; The features extracted from acoustic and physiological signals are analyzed to authenticate human-generated speech; as well as The combination of the verification signals is compared with each other to further confirm the authenticity of the voice event.
2. The method according to claim 1, further comprising: The combined speech signal passively captured via a microphone; as well as Simultaneously, the combined physiological signals are captured via at least one sensor.
3. The method according to claim 1, further comprising: Extracting the features includes extracting speech features related to vocal cord vibration and speech organ movement from both the combined speech signal and one or more physiological signals from the signals, the speech features defining the characteristics of the speech production mechanism.
4. The method according to claim 1, further comprising: Extracting the features includes extracting spectral features and temporal features from one or more signals related to speech production, wherein the spectral features define the frequency components of the one or more physiological signals, and the temporal features represent the temporal variation of the one or more physiological signals.
5. The method of claim 1, further comprising: Access to the physiological signals from the combination of signals captured by the radar sensor; as well as A distance-Doppler map is generated from the physiological signal. The distance-Doppler map includes a two-dimensional representation that defines a Doppler frequency shift indicating velocity on one axis and a distance indicating spacing on another axis. The intensity at each point in the distance-Doppler map represents the reflectivity of the sensor at the distance and velocity, and the reflectivity provides information about speech generation.
6. The method of claim 1, further comprising: The extracted features are fused to create a multidimensional profile of the speech event, which defines the interaction of physiological and acoustic signals characterizing the speech in the field. as well as The multidimensional profile is compared with known profiles of real speech to assess the likelihood that the speech event is real and human.
7. The method of claim 5, further comprising: Feature representations are generated for radar features and acoustic signals accessed from data captured by the radar sensors; as well as Identify the correlation between feature values calculated from the radar features and the acoustic signal data during a segment of speech production to verify that the acoustic signal originates from humans and that the radar signal features are biological signals originating from humans.
8. The method of claim 1, further comprising: Generate a media file associated with the voice event, the media file including a certificate of human origin for verifying the combination of signals.
9. The method of claim 1, further comprising: An isolated representation of the speech event is generated from the data by removing low-frequency physiological signals from the combination of signals.
10. The method according to claim 1, wherein, The combination of signals includes: Vibration signals transmitted via bone conduction or accelerometer during the speech event. Used to verify the presence of other physiological signals from a human speaker at the scene, and Acoustic signals associated with the speech event, measured from the microphone.
11. The method of claim 1, further comprising: The combination of analyzed signals includes The acoustic and physiological signals are filtered to the same frequency range; The extracted features from the acoustic and physiological signals are aligned temporally; as well as Based on the extracted features, the temporal feature changes between the acoustic signal and the physiological signal are compared.
12. The method of claim 11, further comprising: The combination of verification signals includes confirming the authenticity of the speech event when the acoustic and physiological signals exhibit synchronous characteristic changes over time.