System and method for voice analysis
The system and method address the limitations of conventional voice analysis by employing ultrasonic acoustic sensors and digital processing to extract and analyze inaudible vocal features, enhancing biometric identification and medical diagnosis.
Patent Information
- Application Number
- GB2020015934
- Authority / Receiving Office
- GB · GB
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2020-09-30
- Publication Date
- 2026-02-11
AI Technical Summary
Conventional voice analysis systems are limited to the human audible range (20 Hz - 20 kHz) and fail to utilize ultrasonic acoustic features for robust biometric identification, medical diagnosis, and security applications.
A system and method utilizing an acoustic sensor with a measured frequency response in the ultrasonic range (>20 kHz) and digital signal processing to extract and compare ultrasonic vocal acoustic features, enabling identification, presence, and condition assessment of an individual.
Enables robust biometric identification and medical diagnosis by leveraging unique ultrasonic features from vocal signals, providing a more reliable and automated assessment of an individual's identity and condition.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
Technical Field of the Invention The invention relates to the field of voice analysis and more specifically to a system and method for analysing the vocal signals from an individual. Backgrgund........to......the......Invention Voice analysis is the study of speech sounds and is relevant to the fields of medical diagnosis, therapy, criminal justice and intelligence. It is also widely used in connection with security of electronic systems and devices, such as digital assistants and online services, such as banking. It is generally known that an individual’s voice comprises unique acoustic features, which can be captured, recorded and digitally processed using available signal processing techniques. These unique acoustic features can include both anatomical and behavioural components, including accent, syntax, pitch, phonation, loudness, rate and breathing patterns. Analysis of these features can contribute to the diagnosis of problems with the vocal cords or larynx or be indicative of an individual's state of mind or emotion. The presence of these unique acoustic features also enables a stored speech sample to be utilised as a “vocal signature” or biometric identifier to ascertain or verify a speaker’s identity. Many electronic and online communications devices and business services rely on voice biometrics to provide speaker verification. Speaker verification generally requires the user to record a predefined word or phrase to be stored in an authentication database. In order to access the service or device subsequently, the user is requested to input their vocal password, which is then compared to the stored vocal signature for that user. If the input vocal sample matches the stored vocal signature the user’s identity is authenticated and access to the device or service is permitted. Voice analysis is also used for speaker recognition, which aims to identify an unknown speaker by comparing the speaker’s voice biometric to all stored speech samples, whilst speaker diarisation is a method for identifying when the same speaker is speaking by partitioning the audio stream into homogeneous segments according to the speaker identity. The applications above rely on the analysis of unique acoustic features associated with an individual and should not be confused with speech recognition, which is a method of capturing and translating speech into text by computers. In this case it is the content of the speech rather than the identity of the speaker which is critical. All of the above fields rely, to some extent, on the fidelity of voice samples subjected to digital signal processing techniques i.e. input and output through electronic equipment by means of sensors or transducers. It is an aim of the invention to provide an improved system and method for analysing the vocal signals from an individual. Summary of the Invention According to a first aspect of the invention, there is provided a system for analysing vocal signals from an individual comprising: an acoustic sensor configured to have a measured frequency response in the ultrasonic range; means for collecting audio data from the acoustic sensor; processing means configured to analyse collected audio data by extracting a plurality of vocal acoustic features in the ultrasonic range and comparing the extracted features with stored reference data; and output means, such that inaudible vocal acoustic features in the ultrasonic range are used to assess the presence, identity or condition of an individual. According to a second aspect of the invention, there is provided a method for analysing vocal signals from an individual comprising the steps of: collecting audio data from an acoustic sensor configured to have a measured frequency response in the ultrasonic range; processing the collected audio data by extracting a plurality of vocal acoustic features in the ultrasonic range and comparing the extracted features with stored reference data; and outputting data, such that inaudible vocal acoustic features in the ultrasonic range are used to assess the presence, identity or condition of an individual. The human ear can generally hear sound at frequencies between 20 Hz and 20 kHz. This frequency range is referred to as the human audible range or audible sound. Sound having frequencies above 20 kHz is referred to as ultrasound and sound having frequencies below 20 Hz is referred to as infrasound. Sound in the ultrasonic range (>20 kHz) is generally inaudible to humans. Sound propagates as a pressure wave to a receptor, such as an ear or other acoustic sensor, where it is converted into nerve or electrical impulses. For the purposes of electronic digital signal processing a sound wave is converted into a series of discrete impulse samples over time using an analogue-to-digital converter (ADC) after which it may be reconstructed. For human speech it is generally accepted that the useful energy falls within the 100 Hz-4 kHz frequency range and that an audio sampling rate of 8 kHz is adequate. This is the industry standard for telephone and wireless communication systems. For this reason conventional voice sampling and analysis systems are effectively constrained to operating at a sampling rate of 8 kHz (with a considered upper feasible limit of 8 kHz audio, 16 Khz sampling). The inventor has found that vocal signals include unique ultrasonic acoustic features emanating from both verbal and non-verbal content, for example, from breathing and from involuntary, inaudible output from the larynx and vocal cords. By configuring an acoustic sensor to have a measured frequency response in the ultrasonic range and extracting and comparing ultrasonic vocal acoustic features, by sampling at a rate greater than 40 kHz and preferably greater than 80 kHz, it is possible to gather data indicative of the size, shape and condition of an individual's vocal cords, with or without the individual speaking or making any audible noises at all. This data can be used either alone or in conjunction with data from audible features to support identification / presence of an individual, medical diagnosis, therapy, criminal justice, intelligence and security applications. Furthermore, the use of inaudible acoustic features can contribute to a more robust biometric identifier, since the originator is unaware and unable to consciously adjust such features. Ultrasound is, of course, well known and is used in many applications such as imaging, therapy, cleaning, welding, detection etc. but crucially, because of the limitations associated with the human audible range and accepted conventions within the communications industries, ultrasound has not been utilised previously in the field of voice analysis. The acoustic sensor may be any type which is capable of detecting sound, and capable of being configured to have a measured frequency response in the ultrasonic range. Preferably the acoustic sensor is a microphone and more preferably a Digital Microelectromechanical systems (MEMS) microphone. MEMS microphones have the advantage of providing precise and readily matched performance characteristics. Furthermore, they inherently have increased sensitivity at higher frequency and are well suited to operating in the ultrasonic range. In practice, the inventor has discovered that a MEMs microphone can be sampled much faster than its specification suggests and the capability to record ultrasonics has been verified using an anechoic chamber and full audio lab. In some embodiments the system is configured to provide a broadband frequency response, for example in the range from 20 Hz to 100 kHz and the processing means is configured to analyse audio data across the broadband range of the system, so that acoustic features in both the audible and ultrasonic range are analysed and contribute to the output. This can be achieved by using a single sensor having a broadband frequency response or a number of sensors configured to respond to different frequency ranges. To allow for digital signal processing the input sound wave is processed to ensure it is provided in digital form for further analysis. Accordingly the means for collecting audio data in the apparatus is an analogue to digital converter (ADC). The ADC may be a separate component within the apparatus or may be an integral part of the acoustic sensor. For example a MEMS microphone combines an acoustic sensor with an ADC so that a sound wave input is sampled over time to provide a digital output. Conveniently, the processing means may be a standard microcomputer programmed to perform the steps of extracting a plurality of vocal acoustic features in the ultrasonic range and comparing the extracted features with stored reference data. In some embodiments the processing means is configured to operate across a broad band frequency range from 20 Hz to 100 kHz. In particular, the processing means is configured to extract and analyse acoustic features comprising the frequency roll-off of shockwaves in the collected audio data. It has been found that vocal signals include a unique pattern of shockwaves in the ultrasonic range, emanating from both verbal and non-verbal content. For example, ultrasonic signals in speech may be related to vowel sounds from the larynx, starting off at high frequency and dropping off as the vowel continues. The speed and range of the roll-off of these shockwaves can be analysed and used to discriminate an individual’s presence, identity or condition. Output data can be conveyed to a system operator or to another system or device, in a number of known ways. For example a spectrogram, sonograph or voiceprint is a visual representation of frequency against time. An operator may then be responsible for analysing the content of this visual information. Alternatively, output data may be presented as an auditory output. In order to achieve this the processing means is further configured to reduce the frequency of the collected audio data such that ultrasonic acoustic features can be output in the audible range. In other embodiments the output might be a match / no match indication. Digital processing means that the invention has the potential to operate in real time thus providing an automated match / no match discrimination or providing an operator with an objective assessment of inaudible acoustic features, which may be used in conjunction with the operator’s auditory analysis. The system and method is suitable for collecting and processing data obtained from a single acoustic sensor but is equally well suited to collecting and processing audio data which has been collected from a plurality of acoustic sensors arranged in an array, wherein each acoustic sensor is configured to have a measured frequency response in the ultrasonic range. Microphone arrays are well known in the art and it will be well understood by the skilled person that the data obtained from such arrays may additionally be subjected to techniques such as beam forming as is standard in the art to change and / or improve directionality of the sensor array. Brief Description of the Drawings The invention will now be described by way of example only and with reference to the accompanying drawing, in which: Fig. 1 illustrates a spectrogram and wave plot showing frequency distribution with time. The drawing is for illustrative purposes only and is not to scale. DetailedDescription A MEMS microphone was connected to and driven by an FPGA and the output was stored on a micro SD card. This enabled a signal recording of up to 96-150Khz. Fig.1 illustrates a spectrogram and wave plot 10 showing a series of broad band features recorded from an individual’s speech - in this case from a single vowel sound. The larynx causes a series of broad frequency shock waves 14 (only three referenced), present above 20kHz and notably above 24kHz, which are picked up by the microphone. The shock waves 14 can be seen as vertical lines on the spectrogram and as sharp drop-offs on the wave plot. Analysis of these shockwaves, the frequency roll-off and the gaps and lack of gaps in the broad frequency band signal can be used to provide characteristic data relating to an individual’s presence identity or condition. By reducing the frequency of the wave by 80% it is possible for the human ear to hear the inaudible ultrasonic content.
Claims
1. A system for analysing vocal signals from an individual comprising:an acoustic sensor configured to have a measured frequency response in the ultrasonic range;means for collecting audio data from the acoustic sensor;processing means configured to analyse collected audio data by extracting a plurality of vocal acoustic features in the ultrasonic range and comparing the extracted features with stored reference data; andoutput means,such that inaudible vocal acoustic features in the ultrasonic range are used to assess the presence, identity or condition of an individual.
2. A system according to claim 1 wherein the acoustic sensor is configured to operate with a sampling rate greater than 40 kHz.
3. A system according to claim 1 or claim 2 wherein the acoustic sensor is a microphone.
4. A system according to any preceding claim wherein the acoustic sensor comprises a MEMS microphone.
5. A system according to any preceding claim configured to provide a broadband frequency response in the range from 20 Hz to 100 kHz.
6. A system according to any preceding claim wherein the means for collecting audio data is an analogue to digital converter.
7. A system according to any preceding claim wherein the extracted acoustic features comprise the frequency roll-off of shockwaves in the collected audio data.
8. A system according to any preceding claim wherein the processing means is further configured to reduce the frequency of the collected audio data such that ultrasonic acoustic features can be output in the audible range.
9. A system according to any preceding claim comprising a plurality of acousticsensors arranged in an array wherein each acoustic sensor is configured to have a measured frequency response in the ultrasonic range.
10. A method for analysing vocal signals from an individual comprising the steps of:collecting audio data from an acoustic sensor configured to have a measured frequency response in the ultrasonic range;processing the collected audio data by extracting a plurality of vocal acoustic features in the ultrasonic range and comparing the extracted features with stored reference data; andproviding an output,such that inaudible vocal acoustic features in the ultrasonic range are used to assess the presence, identity or condition of an individual.
11. A method according to claim 10 wherein the acoustic sensor is configured to operate with a sampling rate greater than 40 kHz.
12. A method according to claim 10 or 11 wherein the acoustic sensor comprises a microphone.
13. A method according to any of claims 10 to 12 wherein the acoustic sensor comprises a MEMS microphone.
14. A method according to any of claims 10 to 13 wherein the acoustic sensor is configured to have a broadband frequency response in the range from 20 Hz to 100 kHz.
15. A method according to any of claims 10 to 14 wherein the means for collecting audio data is an analogue to digital converter.
16. A method according to any of claims 10 to 15 wherein the extracted acoustic features comprise the frequency roll-off of shockwaves in the collected audio data.
17. A method according to any of claims 10 to 16 wherein the processing means is further configured to reduce the frequency of the collected audio data such that ultrasonic acoustic features can be output in the audible range.
18. A method according to any of claims 10 to 17 comprising a plurality of acoustic sensors arranged in an array wherein each acoustic sensor is configured to have a measured frequency response in the ultrasonic range.c
Citation Information
Patent Citations
Voice identification device
JP2015191076A