Artificial intelligence (AI)-based system for voice biometric authentication and method thereof

The AI-based voice biometric system addresses spoofing and linguistic diversity by processing vocal characteristics, generating secure voiceprints, and adapting to user changes, improving authentication accuracy and security.

WO2026099798A1PCT designated stage Publication Date: 2026-05-15VOXMIND LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
VOXMIND LTD
Filing Date
2025-11-07
Publication Date
2026-05-15

AI Technical Summary

Technical Problem

Existing voice biometric authentication systems are susceptible to spoofing attacks, struggle with diverse linguistic environments, and lack robust anti-spoofing mechanisms, leading to inaccurate authentication and privacy concerns.

Method used

An AI-based system with a phoneme extraction subsystem, multi-dimensional feature analysis, voiceprint-generating subsystem, and liveness-enhanced anti-spoofing authentication to process vocal characteristics, generate secure voiceprints, and adapt to user changes over time.

Benefits of technology

Enhances authentication accuracy across diverse languages and accents, detects spoofing attempts, and ensures secure, reliable voice biometric authentication with continuous learning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IB2025061379_15052026_PF_FP_ABST
    Figure IB2025061379_15052026_PF_FP_ABST
Patent Text Reader

Abstract

An artificial intelligence (AI)-based system (102) for voice biometric authentication and method thereof are disclosed. The AI-based system (102) comprises a voice-capturing unit (106), a phoneme extraction subsystem (206), a multi-dimensional feature analysis subsystem (208), a voiceprint- generating subsystem (210), a voiceprint data encryption subsystem (212), a match-score generating subsystem (214), a real-time authentication subsystem (216), a liveness-enhanced anti-spoofing authentication subsystem (218), and a continuous learning subsystem (220). The AI-based system (102) receives an audio signal from one or more users to identify one or more phonemes. The AI- based system (102) analyse multi-dimensional features in the audio signal to generate an anti- spoofing voice biometric profile. The AI-based system (102) create distinctive voiceprint data as an encrypted mathematical model for generating a match score. The AI-based system (102) authenticate an identity of the one or more users with voice biometric authentication based on the generated match score.
Need to check novelty before this filing date? Find Prior Art

Description

ARTIFICIAL INTELLIGENCE (AI)-BASED SYSTEM FOR VOICE BIOMETRIC AUTHENTICATION AND METHOD THEREOFTECHNICAL FIELD

[0001] Embodiments of the present disclosure relate to biometric authentication systems, and more particularly relate to an artificial intelligence (Al)-based system and method for voice biometric authentication with anti-spoofing and continuous learning competencies.BACKGROUND

[0002] Voice biometric authentication systems have emerged as a promising technology for secure user identification and access control across various applications. The voice biometric authentication systems leverage distinct features of the user’s voice, such as pitch, tone, and cadence, to create voice data for secure access to the various applications. The voice biometric authentication systems provide a convenient and non-intrusive alternative to traditional authentication methods like passwords, Personal Identification Numbers (PINs), and even fingerprints, which are prone to forgotten, stolen, compromised, theft, hacking, or spoofing attempts.

[0003] The voice biometric authentication systems offer several advantages, including non-invasive data collection and the ability to verify identity remotely. However, the existing voice biometric authentication systems face significant challenges. One of the primary concerns with existing voice biometric authentication systems is their susceptibility to spoofing attacks, where an attacker uses recorded or synthesised voice samples to deceive the voice biometric authentication systems. Many current voice biometric authentication systems rely heavily on the static analysis of voiceprints, making them vulnerable to sophisticated spoofing techniques.

[0004] The voice biometric authentication systems struggle to maintain accuracy when dealing with users who speak in different accents and languages. This limitation stems from the fact that the many voice biometric authentication systems are trained on a limited dataset, which may not adequately represent a full range of phonetic variations present in global populations. Asa result, these voice biometric authentication systems generate higher false rejection rates (FRR) or false acceptance rates (FAR) for non-native speakers or those with strong regional accents.

[0005] In another aspect, real-world environments are often noisy, which significantly degrades the performance of the voice biometric authentication systems. Background noise, overlapping speech, and other auditory interferences distort the user’s voice, leading to inaccurate voice data analysis and potential authentication failures. Further, current voice biometric authentication systems typically lack advanced anti-spoofing mechanisms that effectively distinguish between a live human voice and a recorded or synthetically generated one. This shortfall leaves the voice biometric authentication systems vulnerable to a variety of attack vectors, particularly in high- security applications.

[0006] Many existing voice biometric authentication systems do not scale well in dynamic environments where user populations are large and diverse. Moreover, these voice biometric authentication systems often lack the ability to adapt over time to changes in a user’s voice due to ageing, illness, or other factors, which lead to authentication failures and decreased user satisfaction. The storage and transmission of the voice data present significant privacy and security challenges. Many current voice biometric authentication systems do not implement sufficiently robust encryption techniques, leaving sensitive biometric data exposed to potential breaches and unauthorized access.

[0007] In the existing technology, an adversarial robust voice biometrics, secure recognition, and identification system are disclosed. The system primarily focuses on detecting fraudulent attempts using replay attacks or deep fake emulations. The system authenticates the user in real time based on the user’s unique vocal characteristics. While this system addresses some security concerns, it may not fully account for the challenges of accurately identifying users across diverse languages and accents. The system is configured with Al algorithms to process the user’s voice input based on distinct factors that are unique to the user’s vocal tract. However, the system does not adequately address the issue of adapting to natural changes in a user's voice over time, which is highlighted as a limitation of conventional systems. While the system incorporates some multi-factor authentication elements, it does not fully address the seamless integration of voice authentication with other security measures, which is noted as a limitation in existing systems.

[0008] There are various technical problems with the voice biometric authentication systems in the prior art. In the existing technology, the voice biometric authentication systems have significant challenges in accurately identifying and authenticating users, particularly in diverse linguistic environments where accents, dialects, and languages vary widely. Many current voice biometric authentication systems rely on generating embeddings from the voice data and computing voice similarity scores, which leads to increased error rates when processing the voice data. Additionally, the voice biometric authentication systems are often susceptible to spoofing attacks, where attackers use pre-recorded or synthesized voice samples to falsely authenticate as legitimate users. The lack of robust anti-spoofing mechanisms and inadequate handling of background noise further compromise the reliability of the voice biometric authentication systems, making them prone to false acceptances and rejections. Moreover, existing voice biometric authentication systems typically do not incorporate sophisticated encryption techniques to secure voiceprint data, leaving it vulnerable to unauthorized access and potential data breaches. These deficiencies highlight the need for more advanced voice biometric authentication solutions that can overcome these limitations and provide secure, accurate, and reliable authentication across a wide range of user scenarios.

[0009] Therefore, there is a need for a system to address the aforementioned issues by providing advanced voice biometric authentication solutions that overcome these limitations and provide secure, accurate, and reliable authentication across a wide range of user scenarios.SUMMARY

[0010] This summary is provided to introduce a selection of concepts, in a simple manner, which is further described in the detailed description of the disclosure. This summary is neither intended to identify key or essential inventive concepts of the subject matter nor to determine the scope of the disclosure.

[0011] In accordance with an embodiment of the present disclosure, an artificial intelligence (AI)- based system for voice biometric authentication is disclosed. The Al-based system comprises a voice-capturing unit, one or more hardware processors, and a memory unit. The voicecapturing unit is configured to receive an audio signal from each user of the one or more users for the voice biometric authentication analysis. The voice-capturing unit is further configuredto pre-process the received audio signal by performing at least one of: noise reduction, echo cancellation, and normalization of audio levels for the voice biometric authentication analysis.

[0012] According to other aspects of the present disclosure, the memory unit is operatively connected to the one or more hardware processors. The memory unit comprises a set of computer- readable instructions in form of a plurality of subsystems. The plurality of subsystems is configured to be executed by the one or more hardware processors. The plurality of subsystems comprises a phoneme extraction subsystem, a multi-dimensional feature analysis subsystem, a voiceprint-generating subsystem, a voiceprint data encryption subsystem, a match-score generating subsystem, a real-time authentication subsystem, a liveness-enhanced anti-spoofing authentication subsystem, and a continuous learning subsystem.

[0013] In an embodiment, the phoneme extraction subsystem is configured to break down the received audio signal into a plurality of constituent phonemes for identifying one or more phonemes across at least one of: diverse languages and diverse accents. The phoneme extraction subsystem is coupled to one or more machine learning (ML) models. The one or more ML models comprise a neural phonetic aligner configured to align one or more phonetic sequences with the received audio signal within a time-domain waveform. The neural phonetic aligner is trained to accurately detect one or more boundaries of each phoneme of the one or more phonemes in the real-time audio signal comprises at least one of: background noise and overlapping of the one or more phonemes. The one or more ML models comprise a deep convolutional neural network (CNN) configured to analyse pitch data of the received audio signal in a real-time to optimise the identification of the one or more phonemes. The phoneme extraction subsystem is configured to map the identified one or more phonemes to the associated pitch data for generating a dataset with the one or more phonemes and the pitch data of the received audio signal.

[0014] The phoneme extraction subsystem is configured to extract the pre-defined one or more phonemes during the initial configuration of the voice biometric authentication of each user of the one or more users. The pre-defined one or more phonemes are stored in the one or more databases for performing comparative analysis between the identified one or more phonemes against the pre-defined one or more phonemes to generate a phoneme match score. The phoneme match score comprises a defined range. If the phoneme match score is one of: within a threshold phoneme score of the defined range and equal to the threshold phoneme score ofthe defined range, the Al-based system determines that the identified one or more phonemes in the received audio signal is a negative match with the pre-defined one or more phonemes and restrict for the voice biometric authentication. If the phoneme match score exceeds the threshold phoneme score of the defined range, the Al-based system determines that the identified one or more phonemes in the received audio signal is a positive match with the predefined one or more phonemes and allows for the voice biometric authentication.

[0015] In another embodiment, the multi-dimensional feature analysis subsystem is configured to analyse at least one of: one or more vocal features in the received audio signal, frequencies of the plurality of constituent phonemes, and transitions in the plurality of constituent phonemes for generating an anti-spoofing voice biometric profile for each user of the one or more users to store in one or more databases. The one or more vocal features comprise at least one of: pitch patterns, intonation patterns, voice timbre, voice resonance, speech rhythm, speech pacing, and articulatory patterns, to optimise the anti-spoofing voice biometric profile.

[0016] In yet another embodiment, the voiceprint-generating subsystem is coupled to the one or more ML models. The voiceprint-generating subsystem is configured to create distinctive voiceprint data based on the stored anti-spoofing voice biometric profile. The voiceprintgenerating subsystem is configured to generate a voice similarity score by performing a comparative analysis between the real-time voiceprint data against the stored distinctive voiceprint data by using the deep CNN. If the voice similarity score is one of: within a threshold voice similarity score of the defined frequency range and equal to the threshold voice similarity score of the defined frequency range, the Al-based system determines that the real-time voiceprint data is benign for the voice biometric authentication. If the voice similarity score exceeds the threshold voice similarity score of the defined frequency range, the Al-based system determines that the real-time voiceprint data is malicious for voice biometric authentication.

[0017] In an embodiment, the voiceprint data encryption subsystem is configured with one or more cryptographic procedures to encrypt the distinctive voiceprint data for storing the distinctive voiceprint data as an encrypted mathematical model in the one or more databases. The one or more cryptographic procedures is selected from a group comprises at least one of: an Advanced Encryption Standard (AES) with a 256-bit key, asymmetric encryption models,symmetric encryption models, and lattice-based cryptography models, for encrypting the distinctive voiceprint data.

[0018] In another embodiment, the match-score generating subsystem is configured to generate a match score by analysing at least one of: a) real-time voiceprint data generated from a realtime audio signal of the one or more users against the stored distinctive voiceprint data, and b) the identified one or more phonemes against a pre-defined one or more phonemes registered at an initial configuration of the voice biometric authentication. The match score is generated based on a weighted combination of the phoneme match score and the voice similarity score.

[0019] In yet another embodiment, the real-time authentication subsystem is configured to authenticate an identity of the one or more users based on the generated match score of each user of the one or more users for the voice biometric authentication. The match score comprises a defined match range. If the match score is one of: within a threshold match score of the defined match range and equal to the threshold match score of the defined match range, the Al-based system determines that the received audio signal from the one or more users is benign for the voice biometric authentication. If the match score exceeds the threshold match score of the defined match range, the Al-based system determines that the received audio signal from the one or more users is malicious for the voice biometric authentication.

[0020] In an embodiment, the liveness-enhanced anti-spoofing authentication subsystem is configured to detect a spoofing audio signal from the one or more users based on analysing micro-patterns across the one or more phonemes of the received audio signal. The continuous learning subsystem is configured to update at least one of the: stored distinctive voiceprint data and pre-defined one or more phonemes at a pre-defined time period for optimising the voice biometric authentication of the one or more users.

[0021] In accordance with an embodiment of the present disclosure, an Al-based method for voice biometric authentication is disclosed. In the first step, the Al-based method includes receiving, by a voice-capturing unit, an audio signal from each user of the one or more users for the voice biometric authentication analysis. In the next step, the Al-based method includes breaking down, by the one or more hardware processors through the phoneme extraction subsystem, the received audio signal into the plurality of constituent phonemes to identify the one or more phonemes across at least one of the: diverse languages and diverse accents.

[0022] In the next step, the Al-based method includes analysing, by the one or more hardware processors through the multi-dimensional feature analysis subsystem, at least one of the: one or more vocal features in the received audio signal, frequencies of the plurality of constituent phonemes, and transitions in the plurality of constituent phonemes to generate the antispoofing voice biometric profile for each user of the one or more users to store in the one or more databases. In the next step, the Al-based method includes creating, by the one or more hardware processors through the voiceprint-generating subsystem coupled to the one or more ML models, the distinctive voiceprint data based on the stored anti-spoofing voice biometric profile.

[0023] In the next step, the Al-based method includes encrypting, by the one or more hardware processors through the voiceprint data encryption subsystem configured with the one or more cryptographic procedures, the distinctive voiceprint data to store the distinctive voiceprint data as the encrypted mathematical model in the one or more databases. In the next step, the Al-based method includes generating, by the one or more hardware processors through the match-score generating subsystem, the match score by analysing at least one of: a: real-time voiceprint data generated from a real-time audio signal of the one or more users against the stored distinctive voiceprint data, and b) the identified one or more phonemes against a predefined one or more phonemes registered at an initial configuration of the voice biometric authentication.

[0024] In the next step, the Al-based method includes authenticating, by the one or more hardware processors through the real-time authentication subsystem, the identity of the one or more users based on the generated match score of each user of the one or more users for the voice biometric authentication. In the next step, the Al-based method includes detecting, by the one or more hardware processors through the liveness-enhanced anti-spoofing authentication subsystem, the spoofing audio signal from the one or more users based on analysing micropatterns across the one or more phonemes of the receive audio signal. In the next step, the AI- based method includes updating, by the one or more hardware processors through the continuous learning subsystem, at least one of the: stored distinctive voiceprint data and predefined one or more phonemes at the pre-defined time period for optimising the voice biometric authentication of the one or more users.

[0025] In accordance with an embodiment of the present disclosure, a non-transitory computer- readable storage medium storing computer-executable instructions that, when executed by the one or more hardware processors, cause the one or more hardware processors to perform operations for the voice biometric authentication, the operations comprising: a) breaking down the received audio signal through the voice-capturing unit from one or more users into the plurality of constituent phonemes to identify the one or more phonemes across at least one of the: diverse languages and diverse accents, b) analysing at least one of the: one or more vocal features in the received audio signal, frequencies of the plurality of constituent phonemes, and transitions in the plurality of constituent phonemes to generate the antispoofing voice biometric profile for each user of the one or more users to store in the one or more databases, c) creating the distinctive voiceprint data based on the stored anti-spoofing voice biometric profile by using the one or more ML models d) encrypting the distinctive voiceprint data to store the distinctive voiceprint data as the encrypted mathematical model in the one or more databases using one or more cryptographic procedures, e) generating a match score by analysing at least one of: real-time voiceprint data generated from a real-time audio signal of the one or more users against the stored distinctive voiceprint data, and the identified one or more phonemes against a pre-defined one or more phonemes registered at an initial configuration of the voice biometric authentication, f) authenticating the identity of the one or more users based on the generated match score of each user of the one or more users for the voice biometric authentication.

[0026] To further clarify the advantages and features of the present disclosure, a more particular description of the disclosure will follow by reference to specific embodiments thereof, which are illustrated in the appended figures. It is to be appreciated that these figures depict only typical embodiments of the disclosure and are therefore not to be considered limiting in scope. The disclosure will be described and explained with additional specificity and detail with the appended figures.BRIEF DESCRIPTION OF DRAWINGS

[0027] The disclosure will be described and explained with additional specificity and detail with the accompanying figures in which:

[0028] Figure 1 illustrates an exemplary block diagram representation of a network architecture depicting an artificial intelligence (Al)-based system for voice biometric authentication, in accordance with an embodiment of the present disclosure;

[0029] Figure 2 illustrates an exemplary block diagram representation of the Al-based system as shown in Figure 1 for voice biometric authentication, in accordance with an embodiment of the present disclosure;

[0030] Figure 3 illustrates an exemplary flow chart of an Al-based method for voice biometric authentication, in accordance with an embodiment of the present disclosure;

[0031] Figure 4 illustrates an exemplary block diagram representation of a server platform for implementation of the disclosed Al-based system, in accordance with an embodiment of the present disclosure; and

[0032] Figure 5 illustrates an exemplary cloud-based architecture for implementation of the disclosed Al-based system, in accordance with an embodiment of the present disclosure.

[0033] Further, those skilled in the art will appreciate that elements in the figures are illustrated for simplicity and may not have necessarily been drawn to scale. Furthermore, in terms of the construction of the device, one or more components of the device may have been represented in the figures by conventional symbols, and the figures may show only those specific details that are pertinent to understanding the embodiments of the present disclosure so as not to obscure the figures with details that will be readily apparent to those skilled in the art having the benefit of the description herein.DETAILED DESCRIPTION OF THE DISCLOSURE

[0034] For the purpose of promoting an understanding of the principles of the disclosure, reference will now be made to the embodiment illustrated in the figures and specific language will be used to describe them. It will nevertheless be understood that no limitation of the scope of the disclosure is thereby intended. Such alterations and further modifications in the illustrated system, and such further applications of the principles of the disclosure as would normally occur to those skilled in the art are to be construed as being within the scope of the present disclosure. It will be understood by those skilled in the art that the foregoing generaldescription and the following detailed description are exemplary and explanatory of the disclosure and are not intended to be restrictive thereof.

[0035] In the present document, the word “exemplary” is used herein to mean “serving as an example, instance, or illustration.” Any embodiment or implementation of the present subject matter described herein as “exemplary” is not necessarily to be construed as preferred or advantageous over other embodiments.

[0036] The terms “comprise”, “comprising”, or any other variations thereof, are intended to cover a non-exclusive inclusion, such that one or more devices or sub-systems or elements or structures or components preceded by “comprises... a" does not, without more constraints, preclude the existence of other devices, sub-systems, additional sub-modules. Appearances of the phrase "in an embodiment”, "in another embodiment" and similar language throughout this specification may, but not necessarily do, all refer to the same embodiment.

[0037] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this disclosure belongs. The system, methods, and examples provided herein are only illustrative and not intended to be limiting.

[0038] A computer system (standalone, client or server computer system) configured by an application may constitute a “module” (or “subsystem”) that is configured and operated to perform certain operations. In one embodiment, the “module” or “subsystem” may be implemented mechanically or electronically, so a module includes dedicated circuitry or logic that is permanently configured (within a special-purpose processor) to perform certain operations. In another embodiment, a “module” or “subsystem” may also comprise programmable logic or circuitry (as encompassed within a general-purpose processor or other programmable processor) that is temporarily configured by software to perform certain operations.

[0039] Accordingly, the term “module” or “subsystem” should be understood to encompass a tangible entity, be that an entity that is physically constructed permanently configured (hardwired) or temporarily configured (programmed) to operate in a certain manner and / or to perform certain operations described herein.

[0040] Referring now to the drawings, and more particularly to Figure 1 through Figure 4, where similar reference characters denote corresponding features consistently throughout the figures, there are shown preferred embodiments, and these embodiments are described in the context of the following exemplary system and / or method.

[0041] Figure 1 illustrates an exemplary block diagram representation of a network architecture 100 depicting an artificial intelligence (Al)-based system 102 for voice biometric authentication, in accordance with an embodiment of the present disclosure.

[0042] According to an exemplary embodiment of the present disclosure, the network architecture 100 may include the Al-based system 102, one or more databases 116, and one or more communication devices 104. The Al-based system 102 may be communicatively coupled to the one or more databases 116, the one or more communication devices 104 via one or more communication networks 114. The one or more communication networks 114 may be, but not limited to, a wired communication network and / or a wireless communication network. The wired communication network may comprise, but not limited to, at least one of: Ethernet connections, Fiber Optics, Power Line Communications (PLCs), Serial Communications, Coaxial Cables, Quantum Communication, Advanced Fiber Optics, Hybrid Networks, and the like. The wireless communication network may comprise, but not limited to, at least one of: wireless fidelity (wi-fi), cellular networks (including fourth generation (4G) technologies and fifth generation (5G) technologies), Bluetooth, ZigBee, long-range wide area network (LoRaWAN), satellite communication, radio frequency identification (RFID), 6G (sixth generation) networks, advanced loT protocols, mesh networks, non-terrestrial networks (NTNs), near field communication (NFC), and the like.

[0043] In an exemplary embodiment, the one or more communication devices 104 may represent various network endpoints, such as, but not limited to, user devices, mobile devices, smartphones, Personal Digital Assistants (PDAs), tablet computers, phablet computers, wearable computing devices, Virtual Reality / Augmented Reality (VR / AR) devices, laptops, desktops, and the like. The one or more communication devices 104 are configured to function as an intermediate unit between one or more users and the Al-based system 102. The one or more communication devices 104 are configured with at least one of a: voice-capturing unit 106 and one or more user interfaces, and the like, which allow the one or more users to interact with the Al-based system 102.

[0044] In an exemplary embodiment, the voice-capturing unit 106 is configured to receive an audio signal from each user of the one or more users for the voice biometric authentication analysis. The voice-capturing unit 106 is further configured to pre-process the received audio signal by performing at least one of: noise reduction, echo cancellation, and normalization of audio levels for the voice biometric authentication analysis. The voice-capturing unit 106 is configured with advanced noise reduction techniques to filter out ambient and background noise that may interfere with the clarity of the received audio signal. This is crucial in real- world environments where background noise is often present, such as in public spaces, offices, or when using mobile devices. The noise reduction process ensures that the captured audio signal retains the essential vocal characteristics needed for accurate one or more phoneme extraction and voiceprint data generation. The voice-capturing unit 106 is also configured with echo cancellation capabilities. Echoes may occur in environments where sound waves bounce off surfaces and return to the voice-capturing unit 106, potentially distorting the audio signal. Echo cancellation models are applied to detect and remove these unwanted echoes from the received audio signal, ensuring that the audio signal remains clear and undistorted for the voice biometric authentication analysis. The voice-capturing unit 106 normalizes the audio levels of the received audio signal to ensure consistency in volume and clarity, variations in the voice-capturing unit 106 sensitivity, each user distance from the voicecapturing unit 106, and speaking volume may lead to inconsistent audio signals. By normalizing the audio levels, the voice-capturing unit 106 ensures that the audio signal is within an optimal range for analysis, reducing the likelihood of errors during the subsequent stages of the voice biometric authentication analysis. The pre-processing steps collectively enhance the quality and consistency of the audio signal received by the voice-capturing unit 106.

[0045] In another exemplary embodiment, the voice-capturing unit 106 may comprise at least one of: one or more microphones, one or more noise filters and one or more pop filters. The one or more noise filters are configured to reduce the impact of wind interference when the Al-based system 102 is used outdoors, while one or more pop filters are configured to minimize the distortion caused by plosive sounds (such as "p" and "b" sounds) that may occur when the one or more users speaks directly into the one or more microphones.

[0046] In an exemplary embodiment, the one or more user interfaces may include, but not limited to, at least one of: graphical displays, touchscreens, voice recognition, and other input / output mechanisms that facilitate easy access to the Al-based system 102 and control functions from the one or more communication devices 104. The one or more user interfaces provide one or more clickable prompts to each user of the one or more users to generate a user profile. The user profile creation process may involve the user providing specific information, such as personal identifiers, voice samples, and security preferences, which are essential for setting up the voice biometric authentication.

[0047] The user interface may guide each user through the process of recording their voice by displaying or announcing specific phrases or words that each user must repeat. These prompts are configured to capture a broad spectrum of the user's vocal characteristics, including one or more phonemes, intonations, and speech patterns, which are necessary for building a robust and unique voiceprint data for each user of the one or more users. The voice samples collected during initial configuration is analysed in real-time by the Al-based system 102, which then generates distinctive voiceprint data and pre-defined one or more phonemes for each user.

[0048] Additionally, the user interface may allow each user to review and approve the generated voiceprint data and pre-defined one or more phonemes to store in an anti-spoofing voice biometric profile for each user. If necessary, each user may re-record their voice samples or adjust settings to improve the accuracy and reliability of the voiceprint data. The user interface may also provide options for setting up multi-factor authentication, such as combining the voice biometric authentication with a password protection or other biometric methods like , but not limited to, at least one of a: fingerprint and facial recognition, to enhance security.

[0049] The anti-spoofing voice biometric profile generated through the user interface is securely stored within the one or more databases 116 and it serves as a reference for future voice authentication attempts. The user interface may also offer features for managing the antispoofing voice biometric profile, such as updating voice samples, adjusting security settings, or viewing authentication logs, providing the one or more users with comprehensive control over their authentication experience.

[0050] In an exemplary embodiment, the one or more databases 116 may configured to store, and manage data related to various aspects of the Al-based system 102. The one or more databases116 may store at least one of, but not limited to, the anti-spoofing voice biometric profile, the distinctive voiceprint data, the pre-defined one or more phonemes, and other relevant data for the voice biometric authentication. The one or more databases 116 may include different types of databases such as, but not limited to, at least one of: relational databases, non- Structured Query Language (NoSQL) databases, distributed databases, time-series databases, an OpenSearch databases, in-memory databases, and the like, depending on the specific requirements of the Al-based system 102. The relational databases may be utilized for structured data storage, where information is organized into tables with predefined relationships between them. This structure is particularly useful for storing anti-spoofing voice biometric profile, authentication logs, and other data that require consistency and integrity in data relationships. The relational databases may employ SQL (Structured Query Language) for querying and managing the stored data. The NoSQL databases may be employed for storing unstructured or semi- structured data, such as the voiceprint data and the one or more phoneme sequences, which may not fit neatly into the tabular format of relational databases. The NoSQL databases are designed to handle large volumes of data and may provide flexible schemas that accommodate the diverse and complex data types generated by the Al-based system 102. The distributed databases may be implemented to ensure data availability and redundancy across multiple locations. The distributed databases are configured to distribute the storage and processing load across several servers, enhancing the Al-based system’s 102 fault tolerance and scalability. The distributed databases are particularly beneficial for ensuring that the voiceprint data is accessible from different geographical regions, improving the responsiveness and reliability of the Al-based system 102.

[0051] Furthermore, the one or more databases 116 may incorporate advanced encryption mechanisms to protect sensitive data, including the distinctive voiceprint data and antispoofing voice biometric profiles, from unauthorized access or breaches. The one or more databases 116 may also include features for regular backups, data replication, and disaster recovery to ensure data integrity and availability in case of system failures or cyber- attacks.

[0052] In an exemplary embodiment, the Al-based system 102 may be implemented by way of a single device or a combination of multiple devices that may be operatively connected or networked together. The Al-based system 102 may be implemented in hardware or a suitablecombination of hardware and software. The Al-based system 102 includes one or more hardware processors 108 and a memory unit 110. The “one or more hardware processors 108” may comprise a combination of discrete components, an integrated circuit, an applicationspecific integrated circuit, a field-programmable gate array, a digital signal processor, or other suitable hardware. The “software” may comprise one or more objects, agents, threads, lines of code, subroutines, separate software applications, two or more lines of code, or other suitable software structures operating in one or more software applications or one or more processors.

[0053] The one or more hardware processors 108 execute a set of computer-readable instructions for dynamically recommending the course of action sequences for the voice biometric authentication. The one or more hardware processors 108 are high-performance processors capable of handling large volumes of data and complex computations. The one or more hardware processors 108 may be, but not limited to, at least one of: multi-core central processing units (CPU), graphics processing units (GPUs), and specialized Artificial Intelligence (Al) accelerators that enhance an ability of the Al-based system 102 to process real-time data from one or more sources simultaneously.

[0054] The one or more hardware processors 108 is responsible for executing one or more artificial intelligence (Al) models that analyse at least one of the: one or more phonemes, one or more vocal features in the received audio signal, frequencies of a plurality of constituent phonemes, and transitions in the plurality of constituent phonemes, the distinctive voiceprint data, match score and the like for the voice biometric authentication. The one or more hardware processors 108 may also include, for example, microprocessors, microcomputers, microcontrollers, digital signal processors, central processing units, state machines, logic circuits, and / or any devices that manipulate data or signals based on operational instructions. Among other capabilities, the one or more hardware processors 108 may fetch and execute the set of computer-readable instructions in the memory unit 110 operationally coupled with the AI- based system 102 for performing tasks such as data processing, input / output processing, and / or any other functions. Any reference to a task in the present disclosure may refer to an operation being or that may be performed on data.

[0055] The memory unit 110, which works in conjunction with the one or more hardware processors108. The memory unit 110 comprises the set of computer-readable instructions in form of aplurality of subsystems 112. The memory unit 110 is composed of at least one of: a non- transitory volatile memory and a non-volatile memory, ensuring that the data is readily accessible for processing while also being securely stored for long-term analysis and historical reference. The plurality of subsystems 112 is configured to be executed by the one or more hardware processors 108. The plurality of subsystems 112 is operatively connected to one or more one or more machine learning (ML) models 118 for enhancing the performance and accuracy of the voice biometric authentication. The ML models 118 are configured to process and analyse at least one of the: one or more phonemes and distinctive voiceprint data to generate the match score and enable the Al-based system 102 to learn from patterns and improve over time.

[0056] Though few components and the plurality of subsystems 112 are disclosed in Figure 1, there may be additional components and subsystems which is not shown, such as, but not limited to, ports, routers, repeaters, firewall devices, network devices, the one or more databases 116, network attached storage devices, assets, machinery, instruments, facility equipment, emergency management devices, image capturing devices, any other devices, and combination thereof. The person skilled in the art should not be limiting the components / subsystems shown in Figure 1. Although Figure 1 illustrates the Al-based system 102, and the one or more communication devices 104 connected to the one or more databases 116, one skilled in the art can envision that the Al-based system 102, and the one or more communication devices 104 may be connected to several user devices located at various locations and several databases via the one or more communication networks 114.

[0057] Those of ordinary skilled in the art will appreciate that the hardware depicted in Figures 1 may vary for particular implementations. For example, other peripheral devices such as an optical disk drive and the like, local area network (LAN), wide area network (WAN), wireless (e.g., wireless-fidelity (Wi-Fi)) adapter, graphics adapter, disk controller, input / output (VO) adapter also may be used in addition or place of the hardware depicted. The depicted example is provided for explanation only and is not meant to imply architectural limitations concerning the present disclosure.

[0058] Those skilled in the art will recognize that, for simplicity and clarity, the full structure and operation of all data processing systems suitable for use with the present disclosure are not being depicted or described herein. Instead, only so much of the Al-based system 102 as isunique to the present disclosure or necessary for an understanding of the present disclosure is depicted and described. The remainder of the construction and operation of the Al-based system 102 may conform to any of the various current implementations and practices that were known in the art.

[0059] Figure 2 illustrates an exemplary block diagram representation 200 of the Al-based system 102 as shown in Figure 1 for voice biometric authentication, in accordance with an embodiment of the present disclosure.

[0060] In an exemplary embodiment, the Al-based system 102 (hereinafter referred to as the system 102). The system 102 comprises the one or more hardware processors 108, the memory unit 110, and a storage unit 204. The one or more hardware processors 108, the memory unit 110, and the storage unit 204 are communicatively coupled through a system bus 202 or any similar mechanism. The system bus 202 functions as the central conduit for data transfer and communication between the one or more hardware processors 110, the memory unit 110, and the storage unit 204. The system bus 202 facilitates the efficient exchange of information and instructions, enabling the coordinated operation of the system 102. The system bus 202 may be implemented using various technologies, including but not limited to, parallel buses, serial buses, or high-speed data transfer interfaces such as, but not limited to, at least one of a: universal serial bus (USB), peripheral component interconnect express (PCIe), and similar standards.

[0061] In an exemplary embodiment, the memory unit 110 is operatively connected to the one or more hardware processors 108. The memory unit 110 comprises the plurality of subsystems 112 in the form of programmable instructions executable by the one or more hardware processors 108. The plurality of subsystems 112 comprises a phoneme extraction subsystem 206, a multi-dimensional feature analysis subsystem 208, a voiceprint-generating subsystem 210, a voiceprint data encryption subsystem 212, a match-score generating subsystem 214, a real-time authentication subsystem 216, a liveness-enhanced anti-spoofing authentication subsystem 218, and a continuous learning subsystem 220.

[0062] The one or more hardware processors 108, as used herein, means any type of computational circuit, such as, but not limited to, the microprocessor unit, microcontroller, complex instruction set computing microprocessor unit, reduced instruction set computingmicroprocessor unit, very long instruction word microprocessor unit, explicitly parallel instruction computing microprocessor unit, graphics processing unit, digital signal processing unit, or any other type of processing circuit. The one or more hardware processors 108 may also include embedded controllers, such as generic or programmable logic devices or arrays, application-specific integrated circuits, single-chip computers, and the like.

[0063] The memory unit 110 may be the non-transitory volatile memory and the non-volatile memory. The memory unit 110 may be coupled to communicate with the one or more hardware processors 108, such as being a computer-readable storage medium. The one or more hardware processors 108 may execute machine-readable instructions and / or source code stored in the memory unit 110. A variety of machine -readable instructions may be stored in and accessed from the memory unit 110. The memory unit 110 may include any suitable elements for storing data and machine-readable instructions, such as read-only memory, random access memory, erasable programmable read-only memory, electrically erasable programmable read-only memory, a hard drive, a removable media drive for handling compact disks, digital video disks, diskettes, magnetic tape cartridges, memory cards, and the like. In the present embodiment, the memory unit 110 includes the plurality of subsystems 112 stored in the form of machine-readable instructions on any of the above-mentioned storage media and may be in communication with and executed by the one or more hardware processors 108.

[0064] The storage unit 204 may be a cloud storage or the one or more databases 116 such as those shown in Figure 1. The storage unit 204 may store, but not limited to, recommended course of action sequences dynamically generated by the system 102. These action sequences are based on at least one of: phoneme extraction, multi-dimensional feature analysis, generating the distinctive voiceprint data, encrypting the distinctive voiceprint data, generating the match score, and the like. The storage unit 204 ensures that the action sequences are readily accessible for analysis and implementation. By storing this information, the system 102 provides data related to the voice biometric authentication. The storage unit 204 may also store historical data related to user profiles, previous match scores, logs of network activities, and any prior alerts or notifications. Additionally, the storage unit 204 may retain configuration settings, user preferences, and security policies, ensuring that the system 102 operates consistently and in alignment with organizational security protocols. The storage unit204 may be any kind of database such as, but not limited to, relational databases, dedicated databases, dynamic databases, monetized databases, scalable databases, cloud databases, distributed databases, any other databases, and a combination thereof.

[0065] In an exemplary embodiment, the phoneme extraction subsystem 206 is configured to break down the received audio signal into a plurality of constituent phonemes for identifying one or more phonemes across at least one of: diverse languages and diverse accents. This is essential for the voice biometric authentication, as each user of the one or more users may have different speech patterns, accents, and speak in various languages. The phoneme extraction subsystem 206 ensures that the system 102 may accurately process and authenticate users' voices, regardless of linguistic or accent-based variations.

[0066] The phoneme extraction subsystem 206 is operatively coupled to the one or more ML models 118. The one or more ML models 118 comprise a neural phonetic aligner configured to align one or more phonetic sequences with the received audio signal within a time-domain waveform. The neural phonetic aligner is trained to accurately detect one or more boundaries of each phoneme of the one or more phonemes in the real-time audio signal even in challenging conditions comprising at least one of: background noise and overlapping of the one or more phonemes. The ability to recognize the one or more phonemes in noisy environments or when the one or more phonemes blend into each other is essential for accurate voice authentication, ensuring that external disturbances do not compromise the process.

[0067] The one or more ML models 118 comprise a deep convolutional neural network (CNN) configured to analyse pitch data of the received audio signal in real-time, optimising the identification of the one or more phonemes by correlating the pitch data with the extracted one or more phonemes. The deep CNN provides sophisticated pitch analysis that enables the system to precisely identify phonetic variations and differences between one or more users, which contributes to the accuracy of the voiceprint data creation and matching process. Once the one or more phonemes are identified, the phoneme extraction subsystem 206 maps the identified one or more phonemes to the corresponding pitch data to generate a dataset. This dataset contains the one or more phonemes and the pitch data of the received audio signal and is utilized for further analysis, including comparison with pre-defined phonemes stored in the one or more databases 116.

[0068] During the initial configuration of the voice biometric authentication for each user, the phoneme extraction subsystem 206 is configured to extract the pre-defined one or more phonemes from the user's voice. These pre-defined one or more phonemes serve as the user's unique voiceprint and are stored in the one or more databases 116. The system 102 uses these pre-defined one or more phonemes as a reference for future comparative analysis, ensuring that subsequent voice inputs from the user may be accurately verified against this stored predefined one or more phonemes.

[0069] When a new audio signal is received for the voice biometric authentication, the phoneme extraction subsystem 206 performs a comparative analysis between the identified one or more phonemes and the pre-defined one or more phonemes, generating a phoneme match score. This phoneme match score falls within a defined range and is used to determine the authenticity of the voice input. The defined range may fall within, but not limited to, one of: 0 and 1, 0 and 100, 0 and 1000, -1 and 1, 0 and 10, 0 and 255, 1 and 10, 0.0 and 1.0, 0 and 500, 0 and 1024, 1 and 1000, 0.01 and 1, and the like.

[0070] The phoneme match score is evaluated based on a threshold phoneme score. If the phoneme match score is within the threshold phoneme score of the defined range or equal to the threshold phoneme score of the defined range, the system 102 determines that the identified one or more phonemes in the received audio signal is a negative match with the pre-defined one or more phonemes and restrict for the voice biometric authentication. If the phoneme match score exceeds the threshold phoneme score of the defined range, the system 102 determines that the identified one or more phonemes in the received audio signal is a positive match with the pre-defined one or more phonemes and allows for the voice biometric authentication. For instance, the phoneme match score is computed using the number of phonemes that have been statistically matched between the identified one or more phonemes in the received audio signal in real-time against the pre-defined one or more phonemes registered at the initial configuration. The higher the one or more phonemes that match indicates the positive match if the defined range is between 1 and 0. If the system 102 identifies that 90% of the phonemes in the received audio signal match the pre-defined one or more phonemes, the phoneme match score would be 0.9, indicating a strong positive match. This means the user is authenticated, as most phonemes have matched. Conversely, if onlyand potentially a negative authentication result. If none of the phonemes match, the score would be 0, indicating a complete mismatch and the system 102 may restrict access or deny the voice biometric authentication.

[0071] In an exemplary embodiment, the multi-dimensional feature analysis subsystem 208 is configured to analyse at least one of: one or more vocal features in the received audio signal, frequencies of the plurality of constituent phonemes, and transitions in the plurality of constituent phonemes for generating the anti-spoofing voice biometric profile for each user of the one or more users to store in one or more databases 116. The one or more vocal features comprise at least one of: pitch patterns, intonation patterns, voice timbre, voice resonance, speech rhythm, speech pacing, and articulatory patterns, to optimise the anti-spoofing voice biometric profile. The pitch patterns are the frequencies at which the user's voice operates, which varies between individuals and provides a distinctive marker for each user. The intonation patterns are variations in the pitch as each user speaks, allowing the system 102 to distinguish between different speech tones and emotional expressions. The voice timbre is a quality and texture of each user's voice, which may include elements such as roughness, warmth, or sharpness. The voice resonance is a depth in the user's voice, which is influenced by their vocal tract and other physical characteristics. The speech rhythm is a timing and flow of speech, including how quickly or slowly a user speaks, as well as their use of pauses. The speech pacing is speed at which each user articulates individual words or phrases, which may vary significantly from user to user. The articulatory patterns are unique ways in which each user forms and pronounces specific phonemes, contributing to the overall distinctiveness of each user's speech. By analysing the multi-dimensional vocal features, the system 102 ensures that the anti-spoofing voice biometric profile is highly accurate and resistant to common spoofing techniques, such as voice synthesis or voice recordings.

[0072] In an exemplary embodiment, the voiceprint-generating subsystem 210 is operatively connected to the one or more ML models 118, facilitating the creation of the distinctive voiceprint data based on the stored anti-spoofing voice biometric profile. The voiceprintgenerating subsystem 210 is essential in forming a unique voiceprint for each user, leveraging the one or more ML models 118 to ensure accuracy and precision in identifying each user's audio signal. The voiceprint-generating subsystem 210 is configured to perform real-time comparative analysis between the live or real-time voiceprint data captured from each userduring an authentication attempt and the stored distinctive voiceprint data during the initial configuration. The analysis is conducted using a deep CNN, which processes complex voice patterns and identifies intricate similarities and differences between the real-time voice data and the stored distinctive voiceprint data.

[0073] The system 102 generates a voice similarity score based on this comparative analysis. The voice similarity score measures how closely the real-time voiceprint data matches the stored distinctive voiceprint data. This voice similarity score is compared against a defined frequency range. If the voice similarity score falls within a threshold voice similarity score or is equal to the threshold voice similarity score of the defined frequency range, the system 102 determines that the real-time voiceprint data is benign. This means that the current voiceprint data is considered to be a legitimate match with the previously registered user, allowing the authentication process to proceed. Conversely, if the voice similarity score exceeds the threshold voice similarity score of the defined frequency range, the system 102 flags the realtime voiceprint data as malicious, indicating that the voiceprint data does not match the stored distinctive voiceprint data and is a fraudulent or spoofed attempt. In such cases, the system will restrict or deny voice biometric authentication.

[0074] For instance, the voice similarity score is computed using an encoder after both the registered voice sample (captured during the initial configuration) and the verification voice sample (provided during the real-time authentication process) are processed. The encoder is a key component of the system 102 that analyses the two voice samples, extracting distinctive voice features and comparing them to assess how similar or dissimilar the two voiceprints are. If the voice similarity score value ranges between 0 and 1. The voice similarity score closer to 0 indicates a high degree of similarity between the registered and verification voice samples. This means the two voice samples are identical, reflecting that the user providing the voice sample during verification is the same as the one who initially registered their voice in the system 102. The voice similarity score closer to 1 indicates a high degree of dissimilarity, meaning the verification voice sample is significantly different from the registered voice sample. This suggests the voice may belong to a different individual, or that there is a substantial variance from the registered voice, triggering potential rejection or suspicion of spoofing. By using this voice similarity score, the system 102 efficiently determines whetherthe voice in real-time matches the stored distinctive voiceprint data, ensuring a robust, secure, and accurate voice biometric authentication.

[0075] In an exemplary embodiment, the voiceprint data encryption subsystem 212 is configured with one or more cryptographic procedures to ensure the secure storage of the distinctive voiceprint data by encrypting it before storing it in the one or more databases 116. The encryption process converts the distinctive voiceprint data into an encrypted mathematical model that may only be accessed or decrypted by authorized systems or users. The one or more cryptographic procedures is selected from a group comprises at least one of: an Advanced Encryption Standard (AES) with a 256-bit key, asymmetric encryption models, symmetric encryption models, and lattice-based cryptography models, for encrypting the distinctive voiceprint data. The AES with the 256-bit key provides a prominent level of security and is considered impenetrable to brute-force attacks.

[0076] In an exemplary embodiment, the match-score generating subsystem 214 is configured to generate a match score that evaluates the authenticity of the received audio signal from the one or more users during the voice biometric authentication. The match score is derived by analysing multiple aspects of the real-time voiceprint data and the one or more phonemes. Specifically, the match-score generating subsystem 214 performs a two-pronged analysis that includes a) the match-score generating subsystem 214 is configured to compare the real-time voiceprint data, which is extracted from the real-time audio signal of the one or more users, against the stored distinctive voiceprint data in the one or more databases 116. This comparison generates the voice similarity score, which indicates how closely the real-time voice matches the previously registered voiceprint data, b) the match-score generating subsystem 214 is configured to compare the identified one or more phonemes from the realtime audio signal against the pre-defined one or more phonemes registered at the initial configuration of the voice biometric authentication. This process results in the generation of the phoneme match score.

[0077] The match score is generated based on a weighted combination of the phoneme match score and the voice similarity score. The weighted combination is determined based on the relative importance of phoneme accuracy and voice similarity in ensuring robust and secure voice biometric authentication. For example, in certain cases where voice similarity may vary slightly due to environmental factors or voice conditions, a higher weight may be assigned tothe phoneme match score to ensure the overall security of the system 102. The match-score generating subsystem 214 is optimized to work in real time, ensuring that the authentication process is both fast and accurate. If the match score falls within a pre-defined threshold, the system confirms the authenticity of the user and allows access. Conversely, if the match score falls below a threshold match score, it signals a potential mismatch and can either trigger additional authentication steps or deny access to protect against unauthorised entry.

[0078] In an exemplary embodiment, the real-time authentication subsystem 216 is configured to authenticate an identity of the one or more users based on the generated match score of each user of the one or more users for the voice biometric authentication. The real-time authentication subsystem 216 operates in real-time, processing the received audio signals from each user and performing authentication checks based on the analysis of the audio signal. The match score is a numerical value that represents how closely the user's real-time voiceprint data and the one or more phonemes match the stored distinctive voiceprint data and pre-defined one or more phonemes. The match score is compared against a defined match range, which establishes the threshold for determining whether the user's voice is benign (legitimate) or malicious (fraudulent).

[0079] The system 102 operates according to the following rules for voice biometric authentication: If the match score falls within the threshold match score or is equal to the threshold match score of the defined match range, the system 102 determines that the received audio signal is benign, meaning that the user's voice matches the previously registered voiceprint data and the one or more phonemes. In this case, the system 102 allows the user to proceed with the authentication process and grants access to the requested resources or services. If the match score exceeds the threshold match score of the defined match range, the system 102 flags the audio signal as potentially malicious. This indicates that the user's real-time voice data and the one or more phonemes do not sufficiently match the stored voiceprint data or phoneme profile, suggesting an attempt at spoofing or unauthorized access. In such a case, the system 102 may deny authentication, trigger an alert, or prompt for additional security checks, depending on the specific implementation and security policies.

[0080] The defined match range may be adjusted based on security requirements and user profiles.For instance, a more stringent threshold may be applied in high- security environments where the risk of voice spoofing is greater, whereas a more lenient threshold may be acceptable inlower-risk scenarios. This flexibility allows the real-time authentication subsystem 216 to balance between user convenience and security.

[0081] For instance, if the match score is scaled between 0 and 1, the system 102 can interpret the score as follows: The match score closer to 0 indicates a higher likelihood that the real-time voice data matches the stored distinctive voiceprint data and the one or more phonemes. Conversely, a match score closer to 1 suggests a significant deviation, indicating the possibility of a spoofing attempt or incorrect user data.

[0082] In an exemplary embodiment, the liveness-enhanced anti-spoofing authentication subsystem 218 is configured to detect a spoofing audio signal from the one or more users based on analysing micro-patterns across the one or more phonemes of the received audio signal. The liveness-enhanced anti-spoofing authentication subsystem 218 uses advanced signal processing techniques and the one or more ML models to identify minute variations in the one or more phonemes and acoustic characteristics of the received audio signal that are indicative of real human speech versus synthesized or replayed audio. The micro-patterns may include subtle shifts in pitch, frequency, energy, and timing variations across the one or more phonemes, which are difficult to replicate in a spoofed or synthetic voice sample. These micro-patterns are compared against pre-defined thresholds to assess whether the audio signal originates from a legitimate live user or an artificial source, such as a recording or a synthesized voice.

[0083] The liveness-enhanced anti-spoofing authentication subsystem 218 continuously learns and adapts to new forms of attack by employing a combination of neural networks, statistical models, and deep learning models. The learning and adapting may also account for changing environmental conditions, such as background noise, by filtering and normalizing the received audio signals before analysing the micro-patterns. This real-time analysis facilitates that the system 102 may detect spoofing attempts with high accuracy and prevent unauthorized access through forged or manipulated audio signals.

[0084] In an exemplary embodiment, the continuous learning subsystem 220 is configured to update at least one of the: stored distinctive voiceprint data and pre-defined one or more phonemes at a pre-defined time period for optimising the voice biometric authentication of the one or more users. The continuous learning subsystem 220 uses the one or more ML models to refineand enhance the stored distinctive voiceprint data and the pre-defined one or more phonemes periodically, based on the real-time authentication results and user interactions. The continuous learning subsystem 220 may perform this update process at regular intervals or in response to significant deviations between the real-time audio signals and the stored distinctive voiceprint data and the pre-defined one or more phonemes. The continuous learning process also assists in optimizing the voice biometric authentication process by improving the accuracy of the anti-spoofing and phoneme matching functions. For instance, as the system collects more data on each user’s voice over time, it may more effectively distinguish between legitimate variations in a user’s speech and potential spoofing attempts. The updated distinctive voiceprint data and pre-defined phoneme profiles are securely stored in the one or more databases 116, ensuring that the system 102 remains robust against evolving threats and continues to authenticate the one or more users with precision. This updating process may be automated, requiring minimal manual intervention, while ensuring that the voice biometric authentication remains current and reliable.

[0085] Figure 3 illustrates an exemplary flow chart of an Al-based method 300 for voice biometric authentication, in accordance with an embodiment of the present disclosure.

[0086] In according to an exemplary embodiment of the present disclosure, the Al-based method 300 for voice biometric authentication is disclosed. At step 302, the Al-based method 300 includes receiving, by the voice-capturing unit, an audio signal from each user of the one or more users for the voice biometric authentication analysis. In the next step, the Al-based method 300 includes pre-processing, by the voice-capturing unit, the received audio signal by performing at least one of the: noise reduction, echo cancellation, and normalization of audio levels for the voice biometric authentication analysis.

[0087] At step 304, the Al-based method 300 includes breaking down, by the one or more hardware processors through the phoneme extraction subsystem, the received audio signal into the plurality of constituent phonemes to identify the one or more phonemes across at least one of the: diverse languages and diverse accents. The breaking down of the received audio signal comprises: a) coupling the phoneme extraction subsystem to the one or more ML models, the one or more ML models comprises the neural phonetic aligner and the deep CNN; b) aligning, by the neural phonetic aligner, one or more phonetic sequences with the received audio signal within the time-domain waveform; c) detecting, by the neural phonetic aligner, the one ormore boundaries of each phoneme of the one or more phonemes in the real-time audio signal comprises at least one of the: background noise and overlapping of the one or more phonemes; d) analysing, by the deep CNN, the pitch data of the received audio signal in the real-time to optimize the identification of the one or more phonemes; and d) mapping the identified one or more phonemes to the associated pitch data to generate the dataset with the one or more phonemes and the pitch data of the received audio signal.

[0088] In step 304, the Al-based method 300 includes extracting, by the phoneme extraction subsystem, the pre-defined one or more phonemes during the initial configuration of the voice biometric authentication of each user of the one or more users. In step 304, the Al-based method 300 includes performing, by the phoneme extraction subsystem, a comparative analysis between the identified one or more phonemes against the pre-defined one or more phonemes to generate the phoneme match score. In step 304, the Al-based method 300 includes determining, by the phoneme extraction subsystem, the identified one or more phonemes in the received audio signal is the negative match with the pre-defined one or more phonemes, and restricting the voice biometric authentication, if the phoneme match score is one of: within the threshold phoneme score of the defined range and equal to the threshold phoneme score of the defined range. In step 304, the Al-based method 300 includes determining, by the phoneme extraction subsystem, whether the identified one or more phonemes in the received audio signal is the positive match with the pre-defined one or more phonemes and allowing the voice biometric authentication, if the phoneme match score exceeds the threshold phoneme score of the defined range.

[0089] At step 306, the Al-based method 300 includes analysing, by the one or more hardware processors through the multi-dimensional feature analysis subsystem, at least one of the: one or more vocal features in the received audio signal, frequencies of the plurality of constituent phonemes, and transitions in the plurality of constituent phonemes to generate the antispoofing voice biometric profile for each user of the one or more users to store in the one or more databases. The one or more vocal features comprise at least one of the: pitch patterns, intonation patterns, voice timbre, voice resonance, speech rhythm, speech pacing, and articulatory patterns, to optimise the anti-spoofing voice biometric profde.

[0090] At step 308, the Al-based method 300 includes creating, by the one or more hardware processors through the voiceprint-generating subsystem coupled to the one or more MLmodels, the distinctive voiceprint data based on the stored anti-spoofing voice biometric profile. In step 308, the Al-based method 300 includes generating, by the voiceprintgenerating subsystem, the voice similarity score by performing the comparative analysis between the real-time voiceprint data against the stored distinctive voiceprint data using the deep CNN. In step 308, the Al-based method 300 includes determining, by the voiceprintgenerating subsystem, whether the real-time voiceprint data is benign for the voice biometric authentication if the voice similarity score is one of: within the threshold voice similarity score of the defined frequency range and equal to the threshold voice similarity score of the defined frequency range. In step 308, the Al-based method 300 includes determining, by the voiceprint-generating subsystem, that the real-time voiceprint data is malicious for the voice biometric authentication if the voice similarity score exceeds the threshold voice similarity score of the defined frequency range.

[0091] At step 310, the Al-based method 300 includes encrypting, by the one or more hardware processors through the voiceprint data encryption subsystem configured with the one or more cryptographic procedures, the distinctive voiceprint data to store the distinctive voiceprint data as the encrypted mathematical model in the one or more databases. The one or more cryptographic procedures are selected from a group comprises at least one of: the AES with a 256-bit key, asymmetric encryption models, symmetric encryption models, and lattice-based cryptography models, for encrypting the distinctive voiceprint data.

[0092] At step 312, the Al-based method 300 includes generating, by the one or more hardware processors through the match-score generating subsystem, the match score by analysing the real-time voiceprint data generated from the real-time audio signal of the one or more users against the stored distinctive voiceprint data. At step 312, the Al-based method 300 includes generating, by the one or more hardware processors through the match-score generating subsystem, the match score by analysing the identified one or more phonemes against the predefined one or more phonemes. Generating the match score based on the weighted combination of the phoneme match score and the voice similarity score. The match score comprises the defined match range.

[0093] At step 314, the Al-based method 300 includes authenticating, by the one or more hardware processors through the real-time authentication subsystem, the identity of the one or more users based on the generated match score of each user of the one or more users for the voicebiometric authentication. At step 314, the Al-based method 300 includes determining, by the real-time authentication subsystem, whether the received audio signal from the one or more users is benign for the voice biometric authentication if the match score is one of: within the threshold match score of the defined match range and equal to the threshold match score of the defined match range. At step 314, the Al-based method 300 includes determining, by the real-time authentication subsystem, whether the received audio signal from the one or more users is malicious for the voice biometric authentication if the match score exceeds the threshold match score of the defined match range.

[0094] Additionally, the Al-based method 300 includes detecting, by the one or more hardware processors through the liveness-enhanced anti-spoofing authentication subsystem, the spoofing audio signal from the one or more users based on analysing micro-patterns across the one or more phonemes of the receive audio signal. Furthermore, the Al-based method 300 includes updating, by the one or more hardware processors through the continuous learning subsystem, at least one of the: stored distinctive voiceprint data and pre-defined one or more phonemes at the pre-defined time period for optimising the voice biometric authentication of the one or more users.

[0095] Figure 4 illustrates an exemplary block diagram representation of one or more server platforms 400 for implementation of the disclosed Al-based system, in accordance with an embodiment of the present disclosure.

[0096] In an exemplary embodiment, for the sake of brevity, the construction, and operational features of the system 102 which are explained in detail above are not explained in detail herein. Particularly, computing machines such as but not limited to internal / external server clusters, quantum computers, desktops, laptops, smartphones, tablets, and wearables may be used to execute the system 102 or may include the structure of the one or more server platforms 400. As illustrated, the one or more server platforms 400 may include additional components not shown, and some of the components described may be removed and / or modified. For example, a computer system with the multiple graphics processing units (GPUs) may be located on at least one of: internal printed circuit boards (PCBs) and external-cloud platforms including Amazon® Web Services (AWS), internal corporate cloud computing clusters, or organizational computing resources.

[0097] The one or more server platforms 400 may be a computer system such as the system 102 that may be used with the embodiments described herein. The computer system may represent a computational platform that includes components that may be in the one or more servers or another computer system. The computer system may be executed by one or more hardware processors 108 (e.g., single, or multiple processors) or other hardware processing circuits, the methods, functions, and other processes described herein. These methods, functions, and other processes may be embodied as machine -readable instructions stored on a computer-readable medium, which may be non-transitory, such as hardware storage devices (e.g., RAM (random access memory), ROM (read-only memory), EPROM (erasable, programmable ROM), EEPROM (electrically erasable, programmable ROM), hard drives, and flash memory). The computer system may include the one or more hardware processors 108 that execute software instructions or code stored on a non-transitory computer-readable storage medium 402 to perform methods of the present disclosure. The software code includes, for example, instructions to gather data and analyse the network environment data. For example, the plurality of subsystems 112 includes the phoneme extraction subsystem 206, the multidimensional feature analysis subsystem 208, the voiceprint-generating subsystem 210, the voiceprint data encryption subsystem 212, the match-score generating subsystem 214, the real-time authentication subsystem 216, the liveness-enhanced anti-spoofing authentication subsystem 218, and the continuous learning subsystem 220.

[0098] The instructions on the computer-readable storage medium 402 are read and stored the instructions in the storage unit or random-access memory (RAM) 404. The storage unit 204 may provide a space for keeping static data where at least some instructions could be stored for later execution. The stored instructions may be further compiled to generate other representations of the instructions and dynamically stored in the RAM 404. The one or more hardware processors 108 may read instructions from the RAM 404 and perform actions as instructed.

[0099] The computer system may further include an output device 406 to provide at least some of the results of the execution as output including, but not limited to, visual information to the one or more users. The output device 406 may include a display on computing devices and virtual reality glasses. For example, the display may be a mobile phone screen or a laptop screen. GUIs and / or text may be presented as an output on the display screen. The computer systemmay further include an input device 408 to provide the one or more users or another device with mechanisms for entering data and / or otherwise interacting with the computer system. The input device 408 may include, for example, a keyboard, a keypad, a mouse, or a touchscreen. Each of these output devices 406 and input device 408 may be joined by one or more additional peripherals.

[0100] A network communicator 410 may be provided to connect the computer system to a network and in turn to other devices connected to the network including other entities, servers, data stores, and interfaces. The network communicator 410 may include, for example, a network adapter such as a LAN adapter or a wireless adapter. The computer system may include a data sources interface 412 to access a data source 414. The data source 414 may be an information resource about voice biometric data. As an example, the one or more databases 116 of exceptions and rules may be provided as the data source 414. Moreover, knowledge repositories and curated data may be other examples of the data source 414. The data source 414 may include libraries containing, but not limited to, datasets related to cryptographic procedures, device configurations, historical cryptographic procedures, cryptographic keys, and other essential information for the voice biometric authentication. Moreover, the data sources interface 412 enables the system 102 to dynamically access and update these data repositories as latest information is collected, analysed, and utilized.

[0101] Figure 5 illustrates an exemplary cloud-based architecture 500 for implementation of the disclosed Al-based system 102, in accordance with an embodiment of the present disclosure.

[0102] In an exemplary embodiment, the one or more users accessing the system 102 via the one or more communication devices 104. The one or more users interact with the system 102 through the one or more user interfaces to enroll, verify, and authenticate. In the illustrative embodiment, the server platform 400 is the AWS. The system 102 is deployed within a specific AWS region, here EU-WEST-2, indicating that the infrastructure is regionalized for data sovereignty, performance, or compliance. The system 102 is deployed within a secure and isolated Virtual Private Cloud (VPC) i.e.,VoxmindMVP-VPC, which is further divided into one or more subnets based on functionality. The one or more subnets comprises a presentation subnet 502 and an application subnet 504. The presentation subnet 502 is publicfacing services accessible to the one or more users and one or more service providers. The backoffice and one or more application programming interface (API) endpoints facilitateinteractions, while a streams service may handle real-time audio streaming from the one or more service providers.

[0103] The application subnet 504 is configured with web user interface (UI) 506 to provide access to core application interfaces and management consoles. The web user UI 506 operatively connected to core services. The core services comprises a verification service 508, an enrolment service 510, an authentication service 512, and a management service 514. The verification service 508 is configured to process the real-time voiceprint data against stored profiles. The enrolment service 510 is configured to manage new one or more user enrolments and capturing their voice biometric data. The authentication service 512 is configured to handle a main authentication logic, verifying users' identities using voice biometrics. The management service 514 is configured to overseeing the system’s operational parameters and configurations.

[0104] The core services is operatively connected to one or more internal services 516. The one or more internal services 516 comprises an anti-spoofing service 518, a core application programming interface (API) service 520, and task processing service 522. The anti-spoofing service 518 is configured to employ the one or more machine learning (ML) models to detect spoofing attempts based on real-time audio analysis. The core API service 520 and the task processing service 522 facilitate backend processing and orchestration of tasks within the system 102. The core services and the one or more internal services 516 are connected to a storage application programming interface (API) 524. The storage API 524 is configured to manage the data flow between the one or more databases 116 and the core services.

[0105] In an exemplary embodiment, an Amazon® Elastic Container Service (ECS) 526 is configured to manage the deployment of containers for various micro services, ensuring scalability and efficient resource utilization. A data subnet 528 is operatively connected to the storage API 524. The data subnet 528 comprises an elasticache 530 and an Amazon® aurora 532. The elasticache 530 is configured to provide in-memory caching to optimize performance, likely caching frequently accessed data like recent voiceprints or match scores. The Amazon® aurora 532 is the relational database used to store persistent data, such as user profiles, voiceprint data, and phoneme-related information. Network services 536 are configured used for storing unstructured data, which could include encrypted voiceprint data, phoneme data, and audio recordings used for training and verification.

[0106] A message (MQ) broker 534 is operatively connected to the task processing service 522. The MQ broker 534 is configured to handle asynchronous message processing, allowing different components to communicate efficiently. The handle asynchronous message processing may be critical for processing voice data streams and managing task distribution across services. One or more audit & logging services 538 are configured to collectively provide monitoring, logging, and tracing. The one or more audit & logging services 538 are essential for real-time monitoring of system health, security auditing, and tracking requests to ensure compliance and detect anomalies. In an illustrative embodiment, the one or more audit & logging services 538 comprises Amazon® CloudWatch, Elasticsearch Service, and AWS X-Ray.

[0107] In an illustrative embodiment, one or more security services 540 are configured with at least one of an: Amazon® Web Services (AWS) Web Application Firewall (WAF), AWS certificate manager, AWS Key Management Service (KMS), AWS shield, and AWS Identity and access management. The AWS WAF is configured to protect against common web exploits. The AWS certificate manager is configured to manage Secure Sockets Eayer (SSE) and Transport Eayer Security (TES) certificates for secure communication. The KMS is configured to encrypt sensitive data like voiceprint data, ensuring it is stored securely in the one or more databases 116. The AWS shield protects the application from Distributed Denial-of-Service (DDoS) attacks. The AWS Identity and access management is configured to manage identity and access controls, ensuring that only authorized one or more users and one or more services are able to access sensitive components. Further, the cloud-based architecture 500 comprises an Amazon® elastic container registry (ECR) 542. The Amazon® ECR 542 is configured to manage container images for the system’s microservices, providing a repository for secure and scalable container storage. The cloud-based architecture 500 demonstrates a secure, scalable, and efficient deployment of the Al-based system 102 for voice biometric authentication. The robust configuration supports continuous learning and updates, enabling the Al-based system 102 to adapt to diverse languages, accents, and potential security threats over time.

[0108] In another exemplary embodiment, the system integrated with one or more artificial intelligence (Al) models to enhance the accuracy, efficiency, and adaptability of the voice biometric authentication. The integration of the one or more Al models allows the system 102 to dynamically learn from user interactions, detect complex patterns in the audio signal, andmake more accurate authentication decisions in real time. The phoneme extraction subsystem 206 may use the one or more Al models to break down the received audio signal into the plurality of constituent phonemes for identifying the one or more phonemes across at least one of the: diverse languages and diverse accents. The voiceprint-generating subsystem 210 may also use the one or more Al models to create the distinctive voiceprint data based on the stored anti-spoofing voice biometric profile. The one or more Al models continuously refine the identification and matching of the one or more phonemes against the pre-defined one or more phonemes, improving the accuracy of the voice authentication by handling dialectal variations and reducing false positives or negatives.

[0109] Numerous advantages of the present disclosure may be apparent from the discussion above. In accordance with the present disclosure, the Al-based system for the voice biometric authentication provides several technical advantages that significantly improve the security, accuracy, and efficiency of voice-based identity verification. One key advantage is the multidimensional feature analysis, which enhances the system's ability to detect and prevent spoofing attacks by analysing a wide range of vocal features, such as pitch patterns, voice timbre, and phoneme transitions. The use of one or more ML models, including deep CNN, further optimizes the identification and comparison of voiceprint data, improving the overall accuracy and speed of authentication, even in noisy environments or with overlapping speech.

[0110] Another advantage is the system's real-time authentication capabilities, which enable it to process and analyse the received audio signals, providing instantaneous feedback on whether the user’s identity is legitimate. The integration of liveness detection ensures that the system can distinguish between live human voices and recorded or synthesized audio, thereby reducing the risk of fraudulent access attempts through spoofing. The system incorporates a continuous learning subsystem, which allows it to evolve and adapt to changes in a user's voice over time. This feature ensures that the system remains accurate and effective, even as users experience natural variations in their speech due to factors like ageing, illness, or environmental changes. Additionally, the secure encryption of voiceprint data through robust cryptographic techniques, such as AES-256, provides the prominent level of data protection, ensuring that sensitive biometric data remains secure and inaccessible to unauthorized parties. Overall, the system provides a scalable, adaptive, and secure solution for voice biometric authentication, addressing key challenges in existing technologies by combining real-timeanalysis, advanced machine learning models, and multi-dimensional feature extraction to deliver accurate and secure authentication results across a wide range of applications.

[0111] A description of an embodiment with several components in communication with each other does not imply that all such components are required. On the contrary, a variety of optional components are described to illustrate the wide variety of possible embodiments of the invention. When a single device or article is described herein, it will be apparent that more than one device / article (whether or not they cooperate) may be used in place of a single device / article. Similarly, where more than one device or article is described herein (whether or not they cooperate), it will be apparent that a single device / article may be used in place of more than one device or article, or a different number of devices / articles may be used instead of the shown number of devices or programs. The functionality and / or the features of a device may be alternatively embodied by one or more other devices which are not explicitly described as having such functionality / features. Thus, other embodiments of the invention need not include the device itself.

[0112] The illustrated steps are set out to explain the exemplary embodiments shown, and it should be anticipated that ongoing technological development will change the manner in which particular functions are performed. These examples are presented herein for purposes of illustration, and not limitation. Further, the boundaries of the functional building blocks have been arbitrarily defined herein for the convenience of the description. Alternative boundaries can be defined so long as the specified functions and relationships thereof are appropriately performed. Alternatives (including equivalents, extensions, variations, deviations, etc., of those described herein) will be apparent to persons skilled in the relevant art(s) based on the teachings contained herein. Such alternatives fall within the scope and spirit of the disclosed embodiments. Also, the words "comprising," "having," "containing," and "including," and other similar forms are intended to be equivalent in meaning and be open-ended in that an item or items following any one of these words is not meant to be an exhaustive listing of such item or items or meant to be limited to only the listed item or items. It must also be noted that as used herein and in the appended claims, the singular forms “a,” “an,” and “the” include plural references unless the context clearly dictates otherwise.

[0113] Finally, the language used in the specification has been principally selected for readability and instructional purposes, and it may not have been selected to delineate or circumscribe theinventive subject matter. It is therefore intended that the scope of the invention be limited not by this detailed description, but rather by any claims that issue on an application based here on. Accordingly, the embodiments of the present invention are intended to be illustrative, but not limited, of the scope of the invention, which is outlined in the following claims.

Claims

CLAIMS1. An artificial intelligence (Al)-based system for voice biometric authentication, comprising: a voice-capturing unit configured to receive an audio signal from each user of one or more users for the voice biometric authentication analysis; one or more hardware processors; and a memory unit operatively connected to the one or more hardware processors, wherein the memory unit comprises a set of computer-readable instructions in form of a plurality of subsystems, configured to be executed by the one or more hardware processors, wherein the plurality of subsystems comprises: a phoneme extraction subsystem configured to break down the received audio signal into a plurality of constituent phonemes for identifying one or more phonemes across at least one of: diverse languages and diverse accents; a multi-dimensional feature analysis subsystem configured to analyse at least one of: one or more vocal features in the received audio signal, frequencies of the plurality of constituent phonemes, and transitions in the plurality of constituent phonemes for generating an anti- spoofing voice biometric profile for each user of the one or more users to store in one or more databases; a voiceprint-generating subsystem coupled to one or more machine learning (ML) models, configured to create distinctive voiceprint data based on the stored anti-spoofing voice biometric profile; a voiceprint data encryption subsystem configured with one or more cryptographic procedures to encrypt the distinctive voiceprint data for storing the distinctive voiceprint data as an encrypted mathematical model in the one or more databases; a match-score generating subsystem configured to generate a match score by analysing at least one of: real-time voiceprint data generated from a real-time audio signal of the one or more users against the stored distinctive voiceprint data; andthe identified one or more phonemes against a pre-defined one or more phonemes registered at an initial configuration of the voice biometric authentication; and a real-time authentication subsystem configured to authenticate an identity of the one or more users based on the generated match score of each user of the one or more users for the voice biometric authentication.

2. The artificial intelligence (Al)-based system according to claim 1, wherein the voice-capturing unit is further configured to pre-process the received audio signal by performing at least one of noise reduction, echo cancellation, and normalization of audio levels for the voice biometric authenti cati on analy si s .

3. The artificial intelligence (Al)-based system according to claim 1, wherein the phoneme extraction subsystem is coupled to the one or more machine learning (ML) models, the one or more machine learning (ML) models comprise a neural phonetic aligner configured to align one or more phonetic sequences with the received audio signal within a time-domain waveform, the neural phonetic aligner is trained to accurately detect one or more boundaries of each phoneme of the one or more phonemes in the real-time audio signal comprises at least one of background noise and overlapping of the one or more phonemes; the one or more machine learning (ML) models comprise a deep convolutional neural network (CNN) configured to analyse pitch data of the received audio signal in a real-time to optimise the identification of the one or more phonemes, and the phoneme extraction subsystem is configured to map the identified one or more phonemes to the associated pitch data for generating a dataset with the one or more phonemes and the pitch data of the received audio signal.

4. The artificial intelligence (Al)-based system according to claim 1, wherein the phoneme extraction subsystem is configured to extract the pre-defined one or more phonemes during the initial configuration of the voice biometric authentication of each user of the one or more users, the pre-defined one or more phonemes stored in the one or more databases for performing comparative analysis between the identified one or more phonemes against the pre-defined one or more phonemes to generate a phoneme match score.

5. The artificial intelligence (Al)-based system according to claim 4, wherein the phoneme match score comprises a defined range, wherein if the phoneme match score is one of within a threshold phoneme score of the defined range and equal to the threshold phoneme score of the defined range, the artificial intelligence (Al)-based system determines that the identified one or more phonemes in the received audio signal is a negative match with the pre-defined one or more phonemes, and restrict for the voice biometric authentication; and wherein if the phoneme match score exceeds the threshold phoneme score of the defined range, the artificial intelligence (Al)-based system determines that the identified one or more phonemes in the received audio signal is a positive match with the pre-defined one or more phonemes and allows for the voice biometric authentication.

6. The artificial intelligence (Al)-based system according to claim 1, wherein the one or more vocal features comprise at least one of pitch patterns, intonation patterns, voice timbre, voice resonance, speech rhythm, speech pacing, and articulatory patterns, to optimise the anti-spoofing voice biometric profile.

7. The artificial intelligence (Al)-based system according to claim 1, wherein the voiceprintgenerating subsystem is configured to generate a voice similarity score by performing a comparative analysis between the real-time voiceprint data against the stored distinctive voiceprint data by using the deep convolutional neural network (CNN).

8. The artificial intelligence (Al)-based system according to claim 7, wherein the voice similarity score comprises a defined frequency range, wherein if the voice similarity score is one of within a threshold voice similarity score of the defined frequency range and equal to the threshold voice similarity score of the defined frequency range, the artificial intelligence (Al)-based system determines that the real-time voiceprint data is benign for the voice biometric authentication; and wherein if the voice similarity score exceeds the threshold voice similarity score of the defined frequency range, the artificial intelligence (Al)-based system determines that the real-time voiceprint data is malicious for voice biometric authentication.

9. The artificial intelligence (Al)-based system according to claim 1, wherein the one or more cryptographic procedures are selected from a group comprises at least one of an Advanced Encryption Standard (AES) with a 256-bit key, asymmetric encryption models, symmetric encryption models, and lattice-based cryptography models, for encrypting the distinctive voiceprint data.

10. The artificial intelligence (Al)-based system according to claim 1, wherein the match score is generated based on a weighted combination of the phoneme match score and the voice similarity score.

11. The artificial intelligence (Al)-based system according to claim 1, wherein the real-time authentication subsystem is configured to authenticate the identity of the one or more users based on the generated match score, the match score comprises a defined match range, wherein if the match score is one of within a threshold match score of the defined match range and equal to the threshold match score of the defined match range, the artificial intelligence (AI)- based system determines that the received audio signal from the one or more users is benign for the voice biometric authentication; and wherein if the match score exceeds the threshold match score of the defined match range, the artificial intelligence (Al)-based system determines that the received audio signal from the one or more users is malicious for the voice biometric authentication.

12. The artificial intelligence (Al)-based system according to claim 1, wherein the plurality of subsystems comprises a liveness-enhanced anti-spoofing authentication subsystem and a continuous learning subsystem, the liveness-enhanced anti-spoofing authentication subsystem configured to detect a spoofing audio signal from the one or more users based on analysing micro-patterns across the one or more phonemes of the received audio signal; and the continuous learning subsystem configured to update at least one of the: stored distinctive voiceprint data and pre-defined one or more phonemes at a pre-defined time period for optimising the voice biometric authentication of the one or more users.

13. An artificial intelligence (Al)-based method for voice biometric authentication, comprising:receiving, by a voice-capturing unit, an audio signal from each user of one or more users for the voice biometric authentication analysis; breaking down, by one or more hardware processors through a phoneme extraction subsystem, the received audio signal into a plurality of constituent phonemes to identify one or more phonemes across at least one of: diverse languages and diverse accents; analysing, by the one or more hardware processors through a multi-dimensional feature analysis subsystem, at least one of: one or more vocal features in the received audio signal, frequencies of the plurality of constituent phonemes, and transitions in the plurality of constituent phonemes to generate an anti-spoofing voice biometric profile for each user of the one or more users to store in one or more databases; creating, by the one or more hardware processors through a voiceprint-generating subsystem coupled to one or more machine learning (ML) models, distinctive voiceprint data based on the stored anti-spoofing voice biometric profile; encrypting, by the one or more hardware processors through a voiceprint data encryption subsystem configured with one or more cryptographic procedures, the distinctive voiceprint data to store the distinctive voiceprint data as an encrypted mathematical model in the one or more databases; generating, by the one or more hardware processors through a match-score generating subsystem, a match score by analysing at least one of: real-time voiceprint data generated from a real-time audio signal of the one or more users against the stored distinctive voiceprint data; and the identified one or more phonemes against a pre-defined one or more phonemes registered at an initial configuration of the voice biometric authentication; and authenticating, by the one or more hardware processors through a real-time authentication subsystem, an identity of the one or more users based on the generated match score of each user of the one or more users for the voice biometric authentication.

14. The artificial intelligence (Al)-based method according to claim 13, comprising:pre-processing, by the voice-capturing unit, the received audio signal by performing at least one of: noise reduction, echo cancellation, and normalization of audio levels for the voice biometric authentication analysis.

15. The artificial intelligence (Al)-based method according to claim 13, wherein breaking down the received audio signal comprises: coupling the phoneme extraction subsystem to the one or more machine learning (ML) models, the one or more machine learning (ML) models comprise a neural phonetic aligner and a deep convolutional neural network (CNN), aligning, by the neural phonetic aligner, one or more phonetic sequences with the received audio signal within a time-domain waveform; detecting, by the neural phonetic aligner, one or more boundaries of each phoneme of the one or more phonemes in the real-time audio signal comprises at least one of: background noise and overlapping of the one or more phonemes; analysing, by the deep convolutional neural network (CNN), pitch data of the received audio signal in a real-time to optimize the identification of the one or more phonemes; and mapping the identified one or more phonemes to the associated pitch data to generate a dataset with the one or more phonemes and the pitch data of the received audio signal.

16. The artificial intelligence (Al)-based method according to claim 13, comprising: extracting, by the phoneme extraction subsystem, the pre-defined one or more phonemes during the initial configuration of the voice biometric authentication of each user of the one or more users; and performing, by the phoneme extraction subsystem, comparative analysis between the identified one or more phonemes against the pre-defined one or more phonemes to generate a phoneme match score.

17. The artificial intelligence (Al)-based method according to claim 16, wherein the phoneme match score comprises a defined range, and the method further comprises: determining, by the phoneme extraction subsystem, the identified one or more phonemes in the received audio signal is a negative match with the pre-defined one or more phonemes, andrestricting the voice biometric authentication, if the phoneme match score is one of: within a threshold phoneme score of the defined range and equal to the threshold phoneme score of the defined range; and determining, by the phoneme extraction subsystem, the identified one or more phonemes in the received audio signal is a positive match with the pre-defined one or more phonemes and allowing the voice biometric authentication, if the phoneme match score exceeds the threshold phoneme score of the defined range.

18. The artificial intelligence (Al)-based method according to claim 13, wherein the one or more vocal features comprise at least one of: pitch patterns, intonation patterns, voice timbre, voice resonance, speech rhythm, speech pacing, and articulatory patterns, to optimise the anti-spoofing voice biometric profile.

19. The artificial intelligence (Al)-based method according to claim 13, comprising: generating, by the voiceprint-generating subsystem, a voice similarity score by performing a comparative analysis between the real-time voiceprint data against the stored distinctive voiceprint data using the deep convolutional neural network (CNN).

20. The artificial intelligence (Al)-based method according to claim 19, wherein the voice similarity score comprises a defined frequency range, determining, by the voiceprint-generating subsystem, the real-time voiceprint data is benign for the voice biometric authentication if the voice similarity score is one of: within a threshold voice similarity score of the defined frequency range and equal to the threshold voice similarity score of the defined frequency range; and determining, by the voiceprint-generating subsystem, the real-time voiceprint data is malicious for the voice biometric authentication if the voice similarity score exceeds the threshold voice similarity score of the defined frequency range.

21. The artificial intelligence (Al)-based method according to claim 13, wherein the one or more cryptographic procedures are selected from a group comprises at least one of: an Advanced Encryption Standard (AES) with a 256-bit key, asymmetric encryption models, symmetric encryption models, and lattice-based cryptography models, for encrypting the distinctive voiceprint data.

22. The artificial intelligence (Al)-based method according to claim 13, wherein generating the match score based on a weighted combination of the phoneme match score and the voice similarity score.

23. The artificial intelligence (Al)-based method according to claim 13, authenticating the identity of the one or more users based on the generated match score comprises; the match score comprises a defined match range, determining, by the real-time authentication subsystem, the received audio signal from the one or more users is benign for the voice biometric authentication if the match score is one of within the threshold match score of the defined match range and equal to the threshold match score of the defined match range; and determining, by the real-time authentication subsystem, the received audio signal from the one or more users is malicious for the voice biometric authentication if the match score exceeds the threshold match score of the defined match range.

24. The artificial intelligence (Al)-based method according to claim 13, comprising: detecting, by the one or more hardware processors through a liveness-enhanced anti-spoofing authentication subsystem, a spoofing audio signal from the one or more users based on analysing micro-patterns across the one or more phonemes of the received audio signal; and updating, by the one or more hardware processors through a continuous learning subsystem, at least one of the: stored distinctive voiceprint data, and pre-defined one or more phonemes at a predefined time period for optimising the voice biometric authentication of the one or more users.

25. A non-transitory computer-readable storage medium storing computer-executable instructions that, when executed by one or more hardware processors, cause the one or more hardware processors to perform operations for voice biometric authentication, the operations comprising: breaking down the received audio signal through a voice-capturing unit from one or more users into a plurality of constituent phonemes to identify one or more phonemes across at least one of: diverse languages and diverse accents; analysing at least one of: one or more vocal features in the received audio signal, frequencies of the plurality of constituent phonemes, and transitions in the plurality of constituent phonemes to generate an anti-spoofing voice biometric profile for each user of the one or more users to store in one or more databases;creating distinctive voiceprint data based on the stored anti-spoofing voice biometric profile by using one or more machine learning (ML) models; encrypting the distinctive voiceprint data to store the distinctive voiceprint data as an encrypted mathematical model in the one or more databases using one or more cryptographic procedures; generating a match score by analysing at least one of: real-time voiceprint data generated from a real-time audio signal of the one or more users against the stored distinctive voiceprint data; and the identified one or more phonemes against a pre-defined one or more phonemes registered at an initial configuration of the voice biometric authentication; and authenticating an identity of the one or more users based on the generated match score of each user of the one or more users for the voice biometric authentication.