Voice biometric documentation for determining speaker impersonation
The method addresses deep fake vulnerabilities in speaker recognition by refining voice prints and implementing a verification pipeline with x-vector extraction and clustering, ensuring secure access and robust voice recognition across diverse speaking styles.
Patent Information
- Application Number
- GB2023017847
- Authority / Receiving Office
- GB · GB
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-11-22
- Publication Date
- 2025-05-28
AI Technical Summary
Speaker recognition systems are vulnerable to deep fakes, posing risks of unauthorized access and privacy breaches due to the increasing sophistication of AI-generated audio mimicking genuine voices, necessitating improved methods to detect and prevent such impersonations.
A computer-implemented method using an enrolment pipeline to refine voice prints and an evaluation pipeline to authenticate speakers, incorporating x-vector extraction, outlier detection, and HDBSCAN clustering to ensure secure access, with a threshold of 75% cosine similarity for verification, and additional security measures like proximity detection and secondary identification.
Enhances speaker recognition systems by accurately distinguishing genuine voices from deep fakes, ensuring secure access and reducing the risk of unauthorized access and privacy breaches, while allowing for flexible voice recognition across various speaking styles.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
Background The present invention pertains to the field of audio processing and analysis, with a specific focus on enhancing speaker recognition and diarization systems. Speaker recognition involves identifying individual speakers within an audio segment by analysing their unique acoustic features. Speaker diarization, on the other hand, is the process of segregating an audio stream into homogeneous segments associated with specific speakers. In recent years, the demand for efficient speaker recognition and diarization systems has surged due to their extensive applications in various domains. Existing methodologies often rely on techniques like Gaussian Mixture Models (GMMs), Hidden Markov Models (HMMs), and clustering algorithms. Speaker recognition technology has witnessed substantial advancements, particularly in the context of smart speakers and mobile phones, contributing to a significant enhancement in security and user experience. This technology uses distinctive vocal patterns and features unique to individuals, making it a powerful tool for authentication and identification. In the realm of smart speakers and mobile phones, speaker recognition is employed for seamless and secure access control. By analysing vocal characteristics during the authentication process, these devices can determine whether the user is authorised to access specific features, apps, or personal information. Users can unlock their devices, authorise transactions, or access sensitive data simply by speaking a predetermined passphrase. Voice-controlled access not only adds an extra layer of security but also enhances the user experience by eliminating the need for traditional PINs or passwords. The unique vocal patterns ensure that only the authorised user can access the device, protecting their personal data from unauthorised usage. Speaker recognition enables a high degree of personalisation in smart speakers and mobile devices. By identifying different users based on their unique voiceprints, these devices can customise preferences, settings, and content according to individual preferences. This level of personalisation enhances user satisfaction and simplifies interactions with the device. For example, a smart speaker can recognise distinct voices within a household, providing personalised music playlists, calendars, or recommendations based on each user's preferences. In mobile phones, personalised user experiences can include tailored app layouts, preferred language settings, and individualised recommendations. In mobile device use, speaker recognition contributes to robust data security and fraud prevention. By employing voice authentication for various sensitive activities, such as mobile banking or authorising transactions, the technology significantly reduces the risk of unauthorised access and fraudulent activities. In cases of fraudulent attempts, speaker recognition can detect anomalies or inconsistencies in vocal patterns, triggering additional security measures or denying access. This ensures that only the legitimate device owner can authorise critical actions, enhancing overall security. Speaker recognition technology has therefore evolved to play a vital role in bolstering security and enhancing user experiences in the domain of smart speakers and mobile phones. Its applications span from access control and user authentication to personalisation and data security. By leveraging unique vocal characteristics, these devices can offer a more secure, efficient, and personalised user experience, paving the way for a future where voice-controlled interactions become an integral part of our daily lives. The above benefits however come at the price of leaving a user's phone open to hacking via an accurate emulation of the user's voice. Speaker recognition technology, though highly effective in enhancing security and user experience, is not without its risks. One significant concern is the potential vulnerability to deep fakes—a form of artificial intelligence-generated audio that mimics a person's voice to a near-perfect extent. Deep fakes have rapidly advanced, making it increasingly challenging to distinguish between genuine and fabricated audio. Hackers and malicious actors can use deep learning algorithms to create sophisticated imitations of an individual's voice based on publicly available recordings or social media clips. As a result, these adversaries can exploit the technology used for speaker recognition, gaining unauthorised access to voice-secured applications and devices. Such fraudulent use of a user's voiceprint can lead to unauthorised access to sensitive information, unauthorised transactions, or even manipulation of personal data. For instance, a malicious actor could impersonate a user's voice to make fraudulent financial transactions or send false instructions to connected devices. Furthermore, the widespread use of smart assistants, which rely heavily on speaker recognition, raises concerns about privacy invasion. Users often interact with these devices in the comfort of their homes, discussing personal and sensitive information. If a hacker gains access to these interactions by mimicking the user's voice, it could result in a severe breach of privacy. To mitigate these risks, it is imperative to continuously enhance speaker recognition systems, incorporating advanced techniques to detect and prevent deep fake attempts. In conclusion, while speaker recognition technology offers remarkable advantages in terms of security and convenience, the rising threat of deep fakes highlights the need for ongoing research and vigilance. There is therefore a need in the art for improved systems and methods of speaker recognition. Summary The present invention in its various aspects is as set out in the appended claims. As such, the present invention provides a computer implemented method for determining speaker impersonation that comprises: providing a secure access portal to the secure area, the access portal enabled by a speaker recognition system comprising means of recording, detecting and processing the voice of the speaker; the system comprising: an enrolment pipeline and an evaluation pipeline. The enrolment pipeline serves to enrol speakers' voices into a database (speaker database) so they can later be recognised in the evaluation pipeline. The enrolment pipeline comprises taking input of one or more enrolment audio data containing audio of a speaker to be enrolled; and then for each of the enrolment audio data: detecting voice activity of the speaker to be enrolled; extracting one or more x-vectors from the detected voice activity; extracting one or more voice prints from the extracted x-vectors. The enrolment pipeline then refines the one or more voice prints and saves the refined voice prints of the speaker in a speaker database alongside a unique speaker ID. The enrolment pipeline facilitates efficient and accurate speaker recognition by automatically processing and extracting critical voice features (x-vectors and voice prints) from the provided audio data. This ensures an effective speaker database containing voice prints associated with a unique speaker ID. The speaker ID may be the speaker's name. Alternatively, the speaker ID may be a number or alphanumeric string. In the case that the speaker's name is not the speaker ID, the database may additionally store the speaker's name alongside the speaker ID and associated voice prints so that the speaker's name can be easily determined from the speaker ID. The unique speaker ID is unique to the identity of the individual speaking. The enrolment pipeline subsequently creates and associates a permission or permissions to access the secure area with the unique speaker ID. This enables there to be users enrolled in the database who have differing permissions. The evaluation pipeline comprises taking input of audio data to be evaluated. From the inputted audio data to be evaluated, detects voice activity in the audio data to be evaluated. The evaluation pipeline extracts one or more x-vectors from the detected voice activity and performs speaker diarazation on the audio data to be evaluated to determine the different speakers in the file. Then, for each determined speaker, the evaluation pipeline performs voice print extraction and searches the speaker database for the one or more extracted voice prints. If an extracted voice print in the evaluation pipeline is matched with the voiceprint of a speaker ID in the database the evaluation pipeline will identify the extracted voiceprint with the speaker ID in the database. A match may be identified if, when compared to a voiceprint in the database using cosine similarity, the extracted voiceprint is above an identification threshold of 60% as measured by cosine similarity. Preferably the identification threshold is 75% cosine similarity, as in experiments this threshold has been shown to identify speakers correctly and discount imitations of the speaker using for example deepfake audio. When an extracted voice print is matched with said refined voice print of a speaker in the database the evaluation pipeline identifies the extracted voiceprint with the speaker in the database using the unique speaker ID and if said permission to access the secure area is present then the speaker is granted access to the secure area. This allows the secure area to be protected by the voice recognition provision of the enrolment and evaluation pipelines ensuring access only by authorised users. When the extracted voice print is not matched with a refined voice print of a speaker in the database; the speaker is denied access. A user may request access through the portal by providing the audio data to be evaluated via a microphone in proximity to the portal. This ensures that the user is present proximate to the portal and reduces the likelihood of effective impersonation. The microphone may be the microphone of the user’s mobile device. In this case proximity may be determined by the strength of a radio signal emitted by the device and carrying the audio data. The radio signal may be a Bluetooth signal. Alternatively, the microphone may be the microphone of the user’s mobile device and proximity is determined by a user scanning a code with their mobile device, the code being proximate to the portal. For example, the user may be prompted to scan a QR code proximate to the portal with the camera of their mobile device, the scanning of the QR code resulting in the user being prompted to provide the audio data to be evaluated via the microphone of their mobile device. To add a further layer of security, the QR code may change after a given time period to ensure that a person must be proximate to the portal to gain access. This prevents a user having a remote copy of the code as a saved copy would become out of date after the given time period. The portal may be a door, and when access to the portal is permitted, it is permitted by means of releasing a lock securing the door. The lock may be an electronic lock. If access through the portal is denied, the evaluation pipeline may trigger an alert in the form of a physical signal in the environment. The physical signal may be one or both of lighting changes and an alarm sounding. Upon denial of access, the portal may be opened and access granted to a further secure, holding, area in which the speaker is asked questions automatically and the responses are recorded and processed using said enrolment pipeline, the subsequent permissions step being used to either grant or deny the speaker further access when released from the secure holding area. In addition, the response to denial of access the speaker may be automatically asked if they wish to enter the further secure holding area. Access to the secure holding area will then be granted or not in accordance with the user’s response. The secure holding area may include a camera for providing one or more of facial recognition, retina scanning and fingerprint scanning as a secondary form of identification. A user’s secondary identification information (facial data, fingerprint data and retina data as appropriate) may be stored in the speaker database alongside the unique speaker ID of a user for the event that voice recognition does not provide a match. An example of a time that voice recognition may not provide a match could be if a person has developed a vocal impediment, for example due to being unwell in a manner that significantly affected their voice. In the case that a secure holding area is used and a user is s, a further unique speaker ID may be created and associated with an existing unique speaker ID and the same permissions granted for adapting the method to a use with a vocal impediment, such as a nasal infection. The evaluation pipeline may save voice prints that have been evaluated, and if matching voiceprints are used within a predetermined time period to request access through the portal and are denied, the evaluation pipeline will trigger an alert. The portal may be a digital portal, for example a digital portal allowing access to a user’s account or an intranet page. In the case that the portal is a digital portal, the permission or permissions may be levels of access to a smart speaker and its functions. For example, in a family with adults and children in the home, the adults may have permissions granted that allow them to make purchases through the smart speaker. The children on the other hand may not be granted this permission, but may have permission to ask the speaker to play music. The permissions may stipulate a time frame within which the permissions are granted. In the case of the portal being a physical door, a user may have access granted only within hours of work for example. In the case of the portal being digital and governing access to a smart speaker and its functions, a user may only have access to certain functions at certain times of the day The enrolment audio data may comprise enrolment audio data and associated metadata, the enrolment audio data comprising an audio recording that includes the voice of the speaker to be enrolled and the metadata includes the name of the speaker being enrolled. The metadata may additionally comprise one or more timestamps that indicate at what time in the audio data the speaker to be enrolled is speaking. This is of particular use in cases where there is more than one speaker in the enrolment data, the method may include automatically associating the names of the different speakers in the speaker database alongside the associated speaker ID and voiceprints. Using the metadata in this way may reduce the amount of time and processing power spent doing voice activity detection. The enrolment pipeline may only perform voice activity detection on regions of the audio enrolment data that are timestamped as including speech from the speaker to be enrolled. In the case that there are a plurality of speakers to be enrolled and the enrolment audio data may therefore comprise the voices of the plurality of speakers. In this case, the metadata preferably comprises the names of the speakers to be enrolled and the timestamps of when in the audio recording each of the plurality of speakers is speaking. Diarization is an algorithm that analyses speech to determine speaker changes. It does not determine the identity of a speaker but detects if a speaker changes. It produces an output like: Speaker 1: Blah Speaker 2: Blah Blah Blah Speaker 1: Blah Blah In the case that the enrolment audio data comprises the voice data of a plurality of speakers to be enrolled, the enrolment pipeline may perform a step of speaker diarization following x-vector extraction to determine the different speakers in the file. Following diarization the enrolment pipeline will then perform voice print extraction for each of the determined speakers separately. The enrolment pipeline would then perform voice print refinement for each speaker separately. The enrolment pipeline may then compare the voice prints of each speaker to voice prints already in the speaker database using cosine similarity to identify if the speakers are new speakers, found speakers or maybe speakers. New speakers preferably have a cosine similarity if less than 70% with all other voice prints in the database. The enrolment pipeline may save the voiceprint(s) of any new speakers in the speaker database alongside a unique speaker ID. For the found and maybe speakers, the cosine similarity may preferably be 76% and above for the found speakers and between 70 and 75 % for the maybe speakers. The enrolment pipeline may check metadata associated with the enrolment audio data for a unique speaker ID associated with the found and maybe speakers. For the found speakers, the enrolment pipeline may compare the unique speaker ID from the metadata to the unique speaker ID associated with the voiceprint in the database that the found speaker was matched to. If the speaker IDs match, then the enrolment pipeline may additionally flag to a user the identity of the found speaker. If the speaker IDs do not match, then the enrolment pipeline may additionally flag to the user that an error in enrolment may have occurred previously. This prevents enrolment of two different speakers with voiceprints that are similar enough that could lead to incorrect identification and also serves as a check of the database to ensure that previous enrolments were correct. For the maybe speakers, the enrolment pipeline may compare the unique speaker ID from the metadata to the unique speaker ID associated with the voiceprint in the database that the maybe speaker was matched to. If the speaker IDs match, then the enrolment pipeline may additionally flag to a user the identity of the maybe speaker. If the speaker IDs do not match, then the enrolment pipeline may additionally flag to the user that an error in enrolment may have occurred previously. For the maybe speakers, if there is no metadata, the enrolment pipeline may alternatively offer the user the option to force enrolment (i.e., to choose to save the speaker in the database alongside a new speaker ID). Following the step of voice print refinement, the enrolment pipeline may further comprise comparing the resulting refined voice print to existing voiceprints in the database and; if the resulting voice print matches an existing voice print in the database; not saving the resulting voiceprint in the database. This prevents redundancy in the database and optimizes storage use. The enrolment pipeline may preferably compare the resulting refined voice print to existing voiceprints in the database using cosine similarity and not save the resulting voiceprint in the database if the cosine similarity between it and any other voiceprint in the database exceeds 75%. This has been shown to be the optimum threshold for preventing redundancy in the database. If the resulting voiceprint is a match with a voiceprint in the database at greater than 75% cosine similarity, the enrolment pipeline may offer a user the option to force enrolment, i.e., to save the voice print in the database regardless of its similarity to an existing voiceprint in the database. Voice print extraction can be considered to be the process of generating a distinct vocal signature from the acoustic features present in a person's speech. Extracting the voice print is preferably done by outlier detection followed by HDBSCAN clustering. The use of an effective outlier detection algorithm in this case can eliminate noisy, distorted and overlapping x-vectors leading to the extraction of high-quality voice prints for improved speaker recognition. HBDSCAN clustering clusters the x-vectors according to speaking styles. This enables the database to include the speaker of interest across a variety of different speaking styles that they may employ in different domains. In outlier detection, preferably all the x-vectors representing a speaker are grouped together and the system calculates a cosine similarity matrix between all the x-vectors and eliminates any noisy x-vectors. Noisy x-vectors are identified based on the cosine similarity measure and the vectors that cannot demonstrate a strong association with any of the major clusters are discarded. The remaining vectors may then be processed with HDBSCAN clustering with an aggressive setting by enabling the ’allow single cluster’ parameter i.e., the x-vectors are re-clustered. The HDBSCAN clustering with the aggressive setting will yield a minimum of one cluster. The number of clusters indicates the distinct speaking styles captured from a speaker’s vocal features enabling the system to capture and identify the speaker of interest across a variety of domains. Finally, the centres of the obtained clusters (simple centroid calculation) are extracted as the voice prints of the speaker. An X-vector is a mathematical representation of a speaker's voice characteristics, often used for speaker verification and speaker identification tasks. x-vectors are typically extracted from a speaker's voice recording using deep neural networks (DNNs) or convolutional neural networks (CNNs) trained on large datasets of speakers. These networks are often referred to as neural speaker embeddings or deep speaker embeddings models. The resulting x-vector is a fixed-dimensional vector that encodes various acoustic and phonetic characteristics of the speaker's voice. More specifically, the output of the affine (linear) components of the layer before the final softmax layer in the CNN is used to generate the x-vector. This is preferably performed using a pretrained ResNet101 CNN model and sample every 144ms of audio with a 120ms overlap to extract audio features and generate the x-vectors. As such, X-vector extraction alone does not provide a precise vector representing an individual, it can be slightly different (unlike a finger print) for the same speaker as there can be slight variability in vocal production, and it can be slightly different due to environmental noise, also people pitch their voices in noisy environment etc. Voice print extraction tries to find the average voice print for a particular voice that accounts for different speaking style, pitching, etc. Voice print extraction preferably uses outlier detection to remove outlier x-vectors leaving the ‘core’ of the persons voice print. The present invention provides the benefit that there may be multiple detected cores. Extracted voice prints for an individual will be slightly different for when the person is shouting, talking or rapping for example. The present invention allows for all these different voice prints to be associated with a single speaker. Allowing this flexibility makes for more accurate recognition. Whilst a single individual may have more than one voice prints, no two individuals may share a voice print. A voice print is defined as a collection of x-vectors that accounts for different speaking styles such as when speaking normally, singing, shouting or other ways people pitch their voices. The strength of our voice biometric system is that we allow for more x-vector representations within what we call a voice print. Voice prints are produced within the core extraction - which involves calculating the mean of extracted x-vectors for a speaker and removing outliers. Refining the voiceprints preferably comprises comparing all of the obtained voiceprints against each other using cosine similarity and discarding the cores with similarity greater than a predetermined threshold. Whilst counterintuitive, this prevents duplication of voiceprints in the database, reducing the amount of database storage required over all. Preferably voiceprints with a similarity greater than 85% are discarded and this has been determined experimentally to be the ideal threshold to prevent unnecessary duplication of voiceprints in the database. The speaker database is preferably stored on the cloud. This allows the database to be accessed remotely, increasing ease of access to the database. It also allows for scalability of the speaker database. Taking input of audio data to be evaluated the evaluation pipeline preferably comprises taking input of a purported speaker ID. The purported speaker ID being speaker ID of a person who is purported to be speaking in the audio data to be evaluated. In this case, if none of the voice prints extracted from the audio data to be evaluated match the voice print of the purported speaker ID in the database, the evaluation pipeline may alert a user to a lack of a match. This is of particular use in determining whether or not a speaker is who they purport to be. The purported speaker ID may be determined by scanning the ID card of a user or by facial or retina recognition. This would have benefit for example in real time access control. An individual may scan an ID card - giving the evaluation pipeline a purported speaker ID - then be prompted to speak into a microphone for voice verification by the evaluation pipeline. Alternatively, the purported speaker ID may be found in metadata of the audio data to be evaluated. This is useful in the case of post processing audio where the speaker themselves may not be present to provide a secondary means of verification. In addition to, or instead of, alerting the user to the lack of a match, the evaluation pipeline may take additional actions such as, requesting additional voice data, restricting access to a device or alerting the authorities. This has applications in voice activated access. For example, if a user were to unlock for example their mobile device using voice recognition, the voice recognition method of the present invention will be able to detect if the voice detected is not the voice of the intended user of the device and restrict access to the device. The audio data to be evaluated may be processed in real time. Again, this has applications in voice activated access where a user requesting access to a device / place / process will speak, into a microphone, their detected voice analysed in real time and access only granted if their voice print is one of an authorised user. Whether or not a user is an authorised user may be stored against the user’s unique speaker ID in the speaker database. The enrolment pipeline may identify the dominant speaker present across the provided files and if a dominant speaker is identified in files above a dominance threshold, then the voiceprint from the dominant speaker may be enrolled as metadata in the database. In an additional aspect, the present invention provides a data processing system comprising means for carrying out any of the methods set out above. The data processing system may be a voice assistant device or a mobile phone. In a further aspect, the present invention comprises a computer program product comprising instructions which, when the program is executed by a computer, cause the computer to carry out the method set out above. In yet another aspect, the present invention may provide a data carrier signal carrying the computer program product set out above. Detailed Description of Figures The present invention will now be described in terms of the following figures: Figure 1 - Flow chart depicting an Enrolment Pipeline Figure 2 - Flow chart depicting an Evaluation Pipeline Although some of the processes in the enrolment and evaluation pipeline are the same, in the following description they are given different numerals to make it clear which pipeline is being referred to. Figure 1 discloses an example evaluation pipeline 10 according to the present invention. The evaluation pipeline takes an input 100 of enrolment audio data. The enrolment audio data comprising audio data of the voice of a speaker to be enrolled and may comprise one or more files containing audio data of the speaker to be enrolled. The enrolment audio data may also include metadata. The metadata may include the names and / or speaker IDs of the at least one speaker to be enrolled. The metadata may also include timestamps indicating when the at least one speaker is speaking. The evaluation pipeline 10 performs voice activity detection 110 on the inputted enrolment audio data. Voice activity detection 110 determines where speech is in the enrolment audio data. One approach to executing voice activity detection may be setting a threshold signal energy. Preferably the voice activity detection comprises analysing the energy across the low-frequency bands where speech occurs (up to 3KHz). Doing so means the enrolment pipeline can tell the difference between noise sources that operate in specific and isolated frequency bands (like for example the humming of a fridge) from multi-band sources like someone humming. Following voice activity detection 110, the evaluation pipeline 10 performs x-vector extraction 120 on the sections of audio identified as having voice activity. Next, the evaluation pipeline 10 performs core extraction 130. Core extraction 130 is also known as voice print extraction and is the method by which the enrolment pipeline extracts voice prints from the enrolment audio data. Core extraction 130 comprises the steps of outlier detection 140 and HBDSCAN clustering 150. The use of an effective outlier detection algorithm in this case can eliminate noisy, distorted and overlapping x-vectors leading to the extraction of high-quality voice prints for improved speaker recognition. HBDSCAN clustering clusters the x-vectors according to speaking styles. This enables the database to include the speaker of interest across a variety of different speaking styles that they may employ in different domains. If there are more than one files of inputted audio enrolment data, the method of the enrolment pipeline will loop through steps 100 to 150 to process 170 all the files containing the speaker of interest (the speaker to be enrolled). After all of the audio enrolment data has been processed for the speaker to be enrolled. The enrolment pipeline 10 comprises performing core refinement 160 on the extracted voice prints. Core refinement may also be referred to as voice print refinement. The enrolment pipeline then enrols 180 the speaker into the database. Enrolling the speaker into the database comprises saving the refined voice prints of the speaker in a speaker database alongside a unique speaker ID. Figure 1 discloses an example where there is a single speaker to be enrolled, however, it is also envisaged that there may be a plurality of speakers to be enrolled. In this case the enrolment audio data comprises the voice data of a plurality of speakers to be enrolled and the enrolment pipeline may perform a step of speaker diarization following x-vector extraction to determine the different speakers in the file. Following diarization the enrolment pipeline will then perform voice print extraction for each of the determined speakers separately. The enrolment pipeline would then perform voice print refinement for each speaker separately before enrolling each of the identified speakers into the database. Figure 2 describes the evaluation pipeline 20. The evaluation pipeline 20 begins with taking in an input 200 of audio data to be evaluated. The evaluation pipeline performs voice activity detection 210 on the audio data to be evaluated to determine where speech is in the audio data to be evaluated. Next, the evaluation pipeline comprises x-vector extraction 220 followed by speaker diarization 230. These are performed on the audio data where voice activity was detected in step 210. Core (voice print) extraction 240 is the next step in the evaluation pipeline and is the method by which the evaluation pipeline extracts voice prints from the audio data to be evaluated. Core extraction is performed separately for each speaker determined by the diarization step. Like in the enrolment pipeline, Core extraction 240 in the evaluation pipeline comprises the steps of outlier detection 250 and HBDSCAN clustering 260. The use of an effective outlier detection algorithm in this case can eliminate noisy, distorted and overlapping x-vectors leading to the extraction of high-quality voice prints for improved speaker recognition. HBDSCAN clustering clusters the x-vectors according to speaking styles. The evaluation pipeline will then search 270 the speaker database for the extracted cores. If a match for the extracted core is found within the speaker database, then the extracted core will be identified with a speaker ID 290. A match may be identified if the extracted voiceprint is above an identification threshold of 60% as measured by cosine similarity to a voiceprint stored in the speaker database. Preferably the identification threshold is 75% cosine similarity, as in experiments this threshold has been shown to identify speakers correctly and discount imitations of the speaker using for example deepfake audio. The evaluation pipeline 20 repeats steps 240 to 270 to repeat the process for all the speakers in the audio data to be evaluated 280.
Claims
1. A method of detecting speaker impersonation for granting access to a secure area, the method comprising providing a secure access portal to the secure area, the access portal enabled by a speaker recognition system including means of recording, detecting and processing the voice of the speaker; the system comprising:an enrolment pipeline and;an evaluation pipeline;wherein the enrolment pipeline comprises:taking input of one or more enrolment audio data containing audio of a speaker to be enrolled;for each of the enrolment audio data:detecting voice activity of the speaker to be enrolled;extracting one or more x-vectors from the detected voice activity;extracting one or more voice prints from the extracted x-vectors,refining the voice prints;saving the refined voice prints the speaker in the speaker database alongside a unique speaker ID;and subsequentlycreating and associating a permission or permissions to access the secure area with the unique speaker ID;wherein the evaluation pipeline comprises:recording audio input thereby creating audio data to be evaluated;detecting voice activity in the audio data to be evaluated;processing by:extracting one or more x-vectors from the detected voice activity;performing speaker diarization on the audio data to be evaluated to determine the different speakers in the file;for each determined speaker:performing voice print extraction;searching the speaker database for a match of the one or more extracted voice print;when the extracted voice print is matched with said refined voice print of a speaker in the database; identifying the extracted voiceprint with the speaker in the database using the unique speaker ID and if said permission to access the secure area is present then the speaker is granted access; andwhen the extracted voice print is not matched with a refined voice print of a speaker in the database; the speaker is denied access.
2. The method of claim 1 wherein a user requests access through the portal by providing the audio data to be evaluated via microphone in proximity to the portal.
3. The method of claim 2 wherein the microphone is the microphone of the user’s mobile device and proximity is determined by evaluating the strength of a radio signal emitted by the device and carrying the audio data.
4. The method of claim 3 wherein the radio signal is a bluetooth signal.
5. The method of any preceding claim wherein the portal is a door and when accessthrough the portal is permitted it is permitted by means of releasing a lock securing the door.
6. The method of claim 5 wherein when access through the portal is permitted, an electronic lock is unlocked.
7. The method of any preceding claim wherein if access through the portal is denied, the evaluation pipeline triggers an alert in the form of a physical signal in the environment.
8. The method of any of claims 1 to 6 wherein upon denial of access, the portal is opened and access granted to a further secure, holding, area in which the speaker is asked questions automatically and the responses are recorded and processed using said enrollment pipeline, the subsequent permissions step being used to either grant or deny the speaker further access when released from the secure holding area.
9. The method of step 8 wherein the speaker is automatically asked if they wish to enter the further secure holding area and access granted accordingly.
10. The method of step 8 or step 9 wherein the further unique speaker ID created is associated with an existing unique speaker ID and the same permissions granted for adapting the method to a user with a vocal impediment, such as a nasal infection.
11. The method of any of claims 1-10 wherein the evaluation pipeline saves voice prints that have been evaluated, and if matching voiceprints are used within a predetermined time period to request access through the portal and are denied, the evaluation pipeline triggers an alert.
12. The computer implemented method of any preceding claim wherein extracting the voice print is done by outlier detection followed by HDBSCAN clustering.
13. The computer implemented method of any preceding claim wherein the database is stored in the cloud.
14. A data processing system comprising means for carrying out the method of claim any preceding claim.
15. The data processing system of claim 11 wherein the data processing system is a voice assistant device.
16. A computer program product comprising instructions which, when the program is executed by a computer, cause the computer to carry out the method of any of claims 1-13.
17. A data carrier signal carrying the computer program product of claim 16.18
Citation Information
Patent Citations
Modern authentication
US20200334344A1
Limiting identity space for voice biometric authentication
US20220392453A1