Privacy protection type biological characteristic identity authentication method and system based on dynamic voice

By combining voiceprint and speech features, a two-factor encryption authentication method based on dynamic speech is adopted, which solves the problems of complexity and insufficient liveness detection in existing multi-factor schemes, and realizes contactless, efficient and secure biometric authentication.

CN121508958APending Publication Date: 2026-02-10NORTHEASTERN UNIV CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511654566.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-12
Publication Date
2026-02-10

AI Technical Summary

Technical Problem

Existing privacy-preserving biometric authentication technologies suffer from problems such as complex multi-factor schemes, high costs, cumbersome operation, difficulty in supporting dynamic behavioral features and liveness detection, and vulnerability to replay attacks.

Method used

A two-factor encryption authentication method based on dynamic voice is adopted, which combines voiceprint and voice features, realizes liveness detection through time-synchronized dynamic password mechanism, and uses dynamic time warping algorithm and position-sensitive integer hashing algorithm for encryption matching to build a contactless multi-factor authentication system.

Benefits of technology

Simplify system architecture, improve user experience, enhance security and privacy protection capabilities, effectively resist replay attacks, and achieve liveness detection and efficient authentication.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121508958A_ABST
    Figure CN121508958A_ABST
Patent Text Reader

Abstract

The invention provides a privacy protection type biological characteristic identity authentication method and system based on dynamic voice, and relates to the technical field of identity authentication. In the initialization stage, the client and the server respectively initialize system parameters for encryption template generation and secure communication; the client extracts standard voice features and generates an encrypted voice template based on the isolated word set of the dynamic password space, and uploads the encrypted voice template to the server to establish an encrypted voice template database; a client side collects voice data of a user, voiceprint features are extracted through a voiceprint recognition model irrelevant to a text, encryption processing is executed to generate an encrypted voiceprint template, and an encrypted voiceprint template database is constructed; in the identity verification stage, the client generates a real-time dynamic password and generates a two-factor encryption authentication credential according to the real-time dynamic password; and the client performs two-factor encryption authentication on the encrypted voice credential and the encrypted voiceprint credential uploaded by the client based on the encrypted voice template database and the encrypted voiceprint template database.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of identity authentication technology, and in particular to a privacy-preserving biometric identity authentication method and system based on dynamic voice. Background Technology

[0002] In recent years, with the rapid development of artificial intelligence technologies such as intelligent sensing and pattern recognition, a new generation of remote identity authentication systems based on biometrics as the core credential has gradually become mainstream. Traditional password- and token-based authentication methods, due to their susceptibility to loss and leakage, are no longer sufficient to meet the dual demands of security and convenience in the network environment. In contrast, authentication schemes based on biometrics such as face, fingerprint, voiceprint, and iris recognition, with their uniqueness and difficulty in forgery, are widely used in mobile payments, smart terminals, IoT devices, and other scenarios. However, with the large-scale application of biometric information in open network environments, user privacy and data security issues are becoming increasingly prominent. Once a biometric template is stolen, its irreplaceable and irreversible nature will permanently invalidate the authentication system, causing irreversible security risks. In recent years, attack methods targeting biometric templates have continuously evolved, with new threats emerging one after another, including template reverse reconstruction, cross-matching attacks, and feature replay attacks.

[0003] Against this backdrop, privacy-preserving biometric authentication has become an important research direction in this field. The formal implementation of the Data Security Law of the People's Republic of China and the Personal Information Protection Law of the People's Republic of China has further promoted the development of technology and the enhancement of compliance requirements in this field.

[0004] From the perspective of biometric selection, current mainstream research mainly focuses on static physiological features, aiming to achieve encrypted storage and matching calculation of biometric templates while maintaining authentication accuracy. For example, single-factor encrypted authentication schemes for static features such as faces, voiceprints, and fingerprints can already balance privacy protection and matching accuracy to a certain extent. However, dynamic behavioral features (such as gestures, gait, keystroke rhythm, and operation trajectory) exhibit unique security enhancement potential in the field of identity authentication due to their characteristics of being difficult to replicate, continuously verifiable, and randomly variable. Nevertheless, existing encrypted authentication mechanisms generally rely on similarity measurement methods between static templates and input features, making it difficult to effectively adapt to dynamic behavioral features that have temporal correlation and content variability. Therefore, there is still a lack of encrypted authentication schemes specifically for dynamic behavioral features.

[0005] Furthermore, two-factor or multi-factor encryption authentication technologies, by integrating multiple biometric information (such as fingerprints + face, voiceprints + iris, etc.), require simultaneous verification of multiple features to pass authentication, offering significant advantages in security and reliability compared to single-factor schemes. This type of scheme has become the mainstream development direction for biometric encryption authentication. However, existing multi-factor schemes generally rely on independent acquisition by multiple sensors, which not only leads to complex system structures and high costs but also increases the cumbersomeness of the operation process, affecting user experience and authentication efficiency. An ideal solution should achieve contactless single-sensor acquisition, capable of simultaneously acquiring multiple biometric features on the same sensor, thereby improving system convenience and usability while maintaining high security.

[0006] From a security authentication strength perspective, liveness detection is a crucial step in biometric authentication systems to ensure that credentials originate from genuine users. However, existing static template-based encryption authentication schemes typically lack temporal correlation information, enabling only static matching and failing to build dynamic liveness detection mechanisms. This makes them ill-suited to effectively defend against demonstration-type threats such as replay attacks and synthetic attacks. This not only weakens the authentication system's resistance to attacks but may also lead to a fragile binding relationship between authentication credentials and the real biometric entity. The root of the problem lies in the fact that traditional encryption templates use a fixed storage method, failing to reflect the time-varying characteristics and liveness attributes of biometrics. Therefore, how to implement liveness detection within a cryptographic framework to confirm that encrypted credentials indeed originate from real-time collection of a living individual is a significant challenge in current research on privacy-preserving biometric authentication.

[0007] In summary, current privacy-preserving biometric authentication technologies still have the following main problems:

[0008] (1) Multi-factor authentication schemes generally rely on independent data collection from multiple sensors, resulting in complex system structures, high deployment costs, cumbersome user operation processes, and poor overall experience.

[0009] (2) Existing encryption matching mechanisms are difficult to effectively support dynamic behavioral characteristics, lack time correlation and liveness detection capabilities, and the system is vulnerable to demonstration-type threats such as replay attacks and synthetic attacks. Summary of the Invention

[0010] To address the shortcomings of existing technologies, this invention provides a privacy-preserving biometric authentication method and system based on dynamic voice. Using voice data, which possesses both static physiological characteristics and dynamic behavioral characteristics, as the core authentication factor, a dynamic password mechanism is introduced within a cryptographic framework. This comprehensively utilizes the physiological uniqueness of voiceprints and the time-varying characteristics of voice content to achieve privacy-preserving liveness-based encrypted authentication under dynamic voice conditions, constructing a secure and reliable encrypted authentication system that integrates liveness detection. This invention aims to solve the problem of the lack of a liveness attribute reflection mechanism in the encrypted authentication process of existing technologies, achieving a secure binding between authentication credentials and the authentication entity, thereby significantly improving the security and usability of biometric encrypted authentication.

[0011] On the one hand, a privacy-preserving biometric authentication system based on dynamic voice includes a client and a server.

[0012] The client is used for collecting biometric data, extracting features, encrypting data, and generating authentication credentials.

[0013] Specifically: The client participates in system parameter initialization, including dynamic password parameters, voiceprint feature encryption parameters, and speech feature encryption parameters; the client uses a speech content feature extractor to extract speech features from the standard speech signals of all isolated words in the dynamic password space, completes encryption processing locally, and finally generates an encrypted speech template and uploads it to the server; the client collects the user's real-time registration voice through a voice sensor, extracts voiceprint features through a voiceprint feature extractor, completes encryption processing locally to obtain an encrypted voiceprint template, and uploads it to the server; the user submits real-time authentication voice according to the dynamic password given by the system, the client extracts features from the real-time authentication voice to obtain dynamic behavioral features, i.e., speech content, and static physiological features, i.e., voiceprint features, and generates encrypted voiceprint features and encrypted speech features as authentication features, which are then uploaded to the server;

[0014] The server is used for storing and managing the encrypted template database, performing encrypted matching calculations, and determining authentication results.

[0015] Specifically: The server participates in the initialization of system parameters, including dynamic password parameters and voiceprint feature encryption parameters; the server receives encrypted voice templates and encrypted voiceprint templates, and constructs encrypted voice template databases and encrypted voiceprint template databases; after receiving authentication credentials uploaded by the client, the server performs comparison and verification based on the encrypted voiceprint template database and the encrypted voice template database, and completes remote two-factor encrypted identity authentication in conjunction with the dynamic password verification mechanism.

[0016] The two-factor encrypted identity authentication specifically includes an encrypted voiceprint authentication mechanism and an encrypted voice recognition mechanism.

[0017] In the encrypted voiceprint authentication mechanism, the client collects user voice data, extracts voiceprint features that can represent the user's identity using a text-independent VPR model, and generates an encrypted voiceprint template which is then uploaded to the server for storage. The user submits real-time authentication voice according to a dynamic password given by the system. The client extracts voiceprint features from the real-time authentication voice and generates encrypted voiceprint features as authentication credentials, which are then sent to the server for verification. The server performs encrypted voiceprint matching based on the encrypted voiceprint template database. When the similarity between the encrypted voiceprint features and the preset encrypted voiceprint template is within a preset threshold range, it is determined that the subject corresponding to the voiceprint is consistent with the registered identity, i.e., the voiceprint identity authentication is successful.

[0018] The encrypted speech recognition mechanism introduces a Dynamic Time Warping (DTW) algorithm to achieve optimal matching between feature sequences of inconsistent lengths. Combined with a Position-Sensitive Integer Hash (PSH) algorithm, it constructs an efficient and secure encrypted speech recognition mechanism. Specifically, it performs speech segmentation using endpoint detection technology, dividing the continuous speech signal into multiple independent speech segments corresponding to each character in the dynamic password. For a given speech template (i.e., a preset isolated word standard speech pattern) and real-time authentication speech features, the system determines whether they represent the same password character through template matching. For any isolated word in the dynamic password space, the client preprocesses the speech signal using an ASR model, converting the preprocessed speech signal into a feature vector usable for template matching and generating a corresponding encrypted speech template, thus constructing an encrypted speech template database. The user submits real-time authentication speech within a valid time window based on the dynamic password generated by the client. The client extracts the speech features, generates encrypted speech features as authentication credentials, and sends them to the server. The server performs encrypted speech recognition based on the encrypted speech template database. When the recognition result is completely consistent with the current dynamic password on the server, it is determined that the authentication credentials originate from the user's real-time operation.

[0019] A privacy-preserving biometric authentication method based on dynamic voice, implemented through the aforementioned privacy-preserving biometric authentication system based on dynamic voice, includes the following steps:

[0020] Step 1: Initialization phase, the client and server initialize system parameters respectively for encrypted template generation and secure communication;

[0021] Step 1.1: Initialize the Time Synchronization Dynamic Password Module (TOTP) on both the client and server sides;

[0022] The Time Synchronization Dynamic Password (TOTP) module allows the client and server to independently calculate identical dynamic passwords based on the same key and current time, thus achieving authentication without real-time network interaction. The workflow of the TOTP module is as follows:

[0023] Step 1.1.1: Initial Key Synchronization: The client and server need to share a unique seed key in advance; the seed key is sent from the server to the client when the client first enters the system; the seed key is only transmitted once in the initial stage and will not be transmitted again afterward.

[0024] Step 1.1.2: Time Window Alignment: Set a fixed time window to divide the current time into multiple consecutive time segments; the client and server will determine the current time window based on the current machine time and use it as the calculation input to ensure that the same time base is used;

[0025] Step 1.1.3: Dynamic Password Generation: The client and server independently calculate the password and then verify the identity through the verification process; the client calculates a hash value based on the locally stored seed key and the current time window using a hash algorithm, then truncates and encodes the hash value to generate a 6-8 digit dynamic password, and sends it to the server.

[0026] Step 1.1.4: Verification and Matching: After receiving the dynamic password entered by the user, the server will call the locally stored seed key and the current time window to perform the same hash operation and encoding process as the client to generate a dynamic password on the server side. The dynamic password entered by the user is compared with the password generated by the server side. If they match, the authentication is successful; otherwise, access is denied. Only when it is within the valid time window and the seed keys used by the server and the client are exactly the same can both ends calculate the same dynamic password, thus ultimately confirming the legitimate identity.

[0027] Step 1.2: Initialize voice feature encryption parameters on both the client and server sides. ,in These are random numbers generated based on the system time. and Indicates the scale of the disruption. Indicates feature selection parameters;

[0028] Step 1.3: Initialize voiceprint feature encryption parameters on both the client and server sides. ,in, The coordinate axis parameters represent the overall safety profile. Indicates unit distance, This represents the unit distance contained in each interval. To indicate the number of intervals on the axis, then on the axis... The interval starting from is denoted as , The center point is used as the interval representative value, that is... ; win represents the sliding window parameters, f is the random projection function, each random projection is defined by the normal vector ω and the bias term θ, and the random projection equation is y=ω T x+θ, where x is the data point within the sliding window; for each sliding window, if... If the value is 1, then the encoding bit is 1; otherwise, the encoding bit is 0.

[0029] Step 2: The client extracts standard speech features and generates encrypted speech templates based on the isolated word set of the dynamic password space, and uploads the encrypted speech templates to the server to establish an encrypted speech template database;

[0030] Step 2.1: For any isolated word w in the dynamic password space, extract speech sequence features based on the standard speech samples of w. ,in Indicates the first Feature vectors of a frame of speech; for any Based on random seed Generate a perturbation vector, where, Represents the Hadema product operation:

[0031] ;

[0032] ;

[0033] Statistical vectors Medium-small index The number of elements pointed to is denoted as . This value is the voice frame. Integer hash encoding; by performing integer hash encoding on all speech frame features, the encrypted speech template of the isolated word w is obtained. ;

[0034] Step 2.2: For each isolated word w, the client sends the data item <isolated word w, encrypted voice template T> to the server for storage, thus building an encrypted voice template database;

[0035] Step 3: The client collects the user's voice data, extracts voiceprint features through a text-independent voiceprint recognition model, performs encryption processing to generate encrypted voiceprint templates, and builds an encrypted voiceprint template database.

[0036] Step 3.1: The client collects the user's voice signal within a fixed duration using sensors, and obtains the voiceprint template using a text-independent voiceprint feature extractor. , template for 3D feature vector;

[0037] Step 3.2: For the given voiceprint template The sliding window parameters win and the random projection function f are generated. 3D binary encoded vector And obtain a preprocessed voiceprint template. , express for 3D eigenvectors; calculate the eigenvalue of each feature point on the axis. The interval above, generate dimensional region vector This yields the voiceprint template. of 5D Encrypted Voiceprint Template s n express for 3D feature vectors.

[0038] Step 3.3: The client will send <User ID, Region Vector> Encrypted voiceprint template The registration information is sent to the server for storage, and an encrypted voiceprint template database is built.

[0039] Step 4: During the authentication phase, the client generates a real-time dynamic password based on the Time Synchronization Dynamic Password (TOTP) algorithm, and generates a two-factor encrypted authentication credential accordingly.

[0040] Step 4.1: The client calculates the real-time dynamic password, which is a random sequence of fixed length within the dynamic password space, and displays it to the user in real time as an authentication command;

[0041] Step 4.2: The user reads aloud the displayed dynamic password, and the client collects the user's real-time voice signal through sensors;

[0042] Step 4.3: The client extracts speech features from the collected speech signal and calculates the encrypted speech credential; specifically, endpoint detection technology is used to segment the continuous speech signal into independent speech segments with the same length as the dynamic password; the sequence features of each speech segment are extracted and denoted as follows. and using encrypted parameters Encryption processing is performed to generate corresponding encrypted voice features. The set of encrypted voice features of all independent voice segments constitutes the encrypted voice authentication credential for this authentication.

[0043] The client simultaneously extracts voiceprint features from the collected speech and calculates encrypted voiceprint credentials; specifically, it extracts voiceprint features using a text-independent voiceprint feature extractor. Using encrypted parameters Encode to obtain a binary encoded vector and encrypted voiceprint features This is the encrypted voiceprint authentication credential used in this authentication.

[0044] Step 4.5: The client sends the generated encrypted voice authentication credential and encrypted voiceprint authentication credential together as a two-factor encrypted authentication credential <User ID, encrypted voice credential, encrypted voiceprint credential> to the server.

[0045] Step 5: The client performs two-factor encryption authentication on the encrypted voice credentials and encrypted voiceprint credentials uploaded by the client based on the encrypted voice template database and the encrypted voiceprint template database.

[0046] Step 5.1: The server uses the encrypted voice credentials to match against the encrypted voice template database. For each encrypted voice feature in the encrypted voice credentials... The server compares the data with all data items <isolated word w, encrypted speech template T> in the database in turn, and uses the Dynamic Time Warping (DTW) algorithm to calculate the similarity.

[0047] Specifically, for a given <isolated word w, encrypted speech template T> and encrypted speech features... The server first constructs a distance matrix, where each element represents the Euclidean distance between corresponding elements in two encrypted sequences. Starting from the bottom left corner of the matrix, the cumulative distance of each point is updated sequentially using dynamic programming, and the path with the smallest cumulative distance is selected as the optimal matching path. The final cumulative distance is the encrypted voice template. and encrypted voice features The DTW distance; the smaller the DTW distance, the higher the similarity between the two; if a encrypted voice template exists in the encrypted voice template database. This makes it consistent with encrypted voice features. If the DTW distance is minimized, then the encrypted voice feature is considered to be the one that minimizes the DTW distance. The corresponding audio segment is identical to the isolated word w represented by the template <isolated word w, encrypted audio template T>.

[0048] Step 5.2: The server performs matching calculations on all encrypted voice features in the encrypted voice authentication credential to reconstruct the complete dynamic password content. If the parsed dynamic password matches the server's current time-synchronized dynamic password, the encrypted voice credential is determined to originate from the user's current real-time reading operation, thus passing the detection; otherwise, it is considered an abnormal encrypted credential.

[0049] Step 5.3: The server uses the encrypted voiceprint credentials to retrieve the corresponding data item <User ID, Region Vector> from the encrypted voiceprint template database based on the user ID. Encrypted voiceprint template >. Encrypted voiceprint features for users claiming the same user ID. Calculate the displacement vector And accordingly, recalculate the corresponding region vector. ;like Then it is considered that the encrypted voiceprint features With encrypted voiceprint template If the voiceprint matches the same entity, the voiceprint identity is confirmed to be successfully matched.

[0050] Step 5.4: The system determines that the two-factor encryption authentication is successful only if the verification results of Step 5.2 and Step 5.3 both meet the preset conditions, and confirms that the operation originated from a legitimate live user. At this time, the server can authorize the corresponding operation permissions or allow the execution of subsequent security processes. Otherwise, access will be denied.

[0051] The beneficial effects of adopting the above technical solution are as follows:

[0052] This invention provides a privacy-preserving biometric authentication method and system based on dynamic voice. It uses voice data to simultaneously carry voiceprint and voice content features, achieving contactless multi-factor authentication, simplifying the system structure and improving user experience. By introducing a Time-Synchronized Dynamic Password (TOTP) mechanism, encrypted credentials possess time-related and dynamic characteristics, thereby enabling liveness detection, resisting replay and forgery attacks, and significantly enhancing the security and privacy protection capabilities of the authentication process. Attached Figure Description

[0053] Figure 1 This is an architecture diagram of the privacy-preserving biometric identity authentication system based on dynamic voice according to the present invention.

[0054] Figure 2 This is an example diagram illustrating the encrypted voice template calculation process in a specific embodiment of the present invention.

[0055] Figure 3 This is an example diagram illustrating the encrypted voice template matching calculation process in a specific embodiment of the present invention. Detailed Implementation

[0056] The specific implementation methods of this application will be further described in detail below with reference to the accompanying drawings and embodiments.

[0057] Example 1:

[0058] On the one hand, a privacy-preserving biometric authentication system based on dynamic voice, such as Figure 1 As shown, this includes both the client and server sides;

[0059] The client is used for collecting biometric data, extracting features, encrypting data, and generating authentication credentials.

[0060] Specifically: The client participates in system parameter initialization, including dynamic password parameters, voiceprint feature encryption parameters, and speech feature encryption parameters; the client uses a speech content feature extractor to extract speech features from the standard speech signals of all isolated words in the dynamic password space, completes encryption processing locally, and finally generates an encrypted speech template and uploads it to the server; the client collects the user's real-time registration voice through a voice sensor, extracts voiceprint features through a voiceprint feature extractor, completes encryption processing locally to obtain an encrypted voiceprint template, and uploads it to the server; the user submits real-time authentication voice according to the dynamic password given by the system, the client extracts features from the real-time authentication voice to obtain dynamic behavioral features, i.e., speech content, and static physiological features, i.e., voiceprint features, and generates encrypted voiceprint features and encrypted speech features as authentication features, which are then uploaded to the server;

[0061] The server is used for storing and managing the encrypted template database, performing encrypted matching calculations, and determining authentication results.

[0062] Specifically: The server participates in the initialization of system parameters, including dynamic password parameters and voiceprint feature encryption parameters; the server receives encrypted voice templates and encrypted voiceprint templates, and constructs encrypted voice template databases and encrypted voiceprint template databases; after receiving authentication credentials uploaded by the client, the server performs comparison and verification based on the encrypted voiceprint template database and the encrypted voice template database, and completes remote two-factor encrypted identity authentication in conjunction with the dynamic password verification mechanism.

[0063] The core of this invention lies in a two-factor encryption authentication mechanism based on dynamic voice, which includes two key processes: encrypted voiceprint authentication and encrypted speech recognition. Only when the similarity between the encrypted voiceprint features and the preset encrypted voiceprint template reaches a predetermined threshold, and the encrypted speech recognition result accurately matches the expected content, can the two-factor encryption authentication be considered successful, confirming that the operation originates from a legitimate, live user, thereby granting corresponding permissions or allowing subsequent operations to be performed.

[0064] The two-factor encrypted identity authentication specifically includes an encrypted voiceprint authentication mechanism and an encrypted voice recognition mechanism.

[0065] In the encrypted voiceprint authentication mechanism, the client collects user voice data, extracts voiceprint features that can represent the user's identity using a text-independent VPR (Voice Print Recognition) model, and generates an encrypted voiceprint template which is uploaded to the server for storage. The user submits real-time authentication voice according to the dynamic password given by the system. The client extracts voiceprint features from the real-time authentication voice and generates encrypted voiceprint features as authentication credentials, which are sent to the server for verification. The server performs encrypted voiceprint matching based on the encrypted voiceprint template database. When the similarity between the encrypted voiceprint features and the preset encrypted voiceprint template is within a preset threshold range, it is determined that the subject corresponding to the voiceprint is consistent with the registered identity, that is, the voiceprint identity authentication is passed.

[0066] In the encrypted voice recognition mechanism, such as Figure 2 , Figure 3As shown, a Dynamic Time Warping (DTW) algorithm is introduced to achieve optimal matching between feature sequences of inconsistent lengths. Combined with a position-sensitive integer hashing (PSH) algorithm, an efficient and secure encrypted speech recognition mechanism is constructed. Specifically, since the speech signals of users reading dynamic passwords generally contain obvious endpoint information, endpoint detection technology can be used for speech segmentation. This divides the continuous speech signal into multiple independent speech segments corresponding to each character in the dynamic password, transforming the dynamic voice password recognition problem into a problem of recognizing multiple isolated words. Given a speech template (i.e., a preset isolated word standard speech pattern) and real-time authentication speech features, the system determines whether the two represent the same password character through template matching. In the isolated word speech recognition process based on template matching, due to the randomness and non-uniformity of the speech signal, even for the same content, the length of the speech signal from different users or the same user at different times is often inconsistent. To address this issue, this invention introduces the Dynamic Time Warping (DTW) algorithm to achieve optimal matching between feature sequences of inconsistent lengths. Combined with a position-sensitive integer hashing (PSH) algorithm, for any isolated word in the dynamic password space, the client preprocesses the speech signal using an Automatic Speech Recognition (ASR) model. The preprocessed speech signal is converted into a feature vector (such as Mel-frequency cepstral coefficients, MFCC) suitable for template matching, generating a corresponding encrypted speech template and constructing an encrypted speech template database. This encrypted speech template does not require specialized training for specific speakers or large-scale adjustments for incremental users, exhibiting good scalability and universality. Users submit real-time authentication speech within a valid time window based on the dynamic password generated by the client. The client extracts the speech features, generates encrypted speech features as authentication credentials, and sends them to the server. The server performs encrypted speech recognition based on the encrypted speech template database. When the recognition result is completely consistent with the current dynamic password on the server, it is determined that the authentication credentials originate from the user's real-time operation, confirming their liveness and achieving dynamic speech two-factor encrypted authentication with liveness detection.

[0067] A privacy-preserving biometric authentication method based on dynamic voice, implemented through the aforementioned privacy-preserving biometric authentication system based on dynamic voice, includes the following steps:

[0068] Step 1: Initialization phase, the client and server initialize system parameters respectively for encrypted template generation and secure communication;

[0069] Step 1.1: The client and server initialize the Time Synchronization Dynamic Password (TOTP) module to ensure the synchronized generation of subsequent dynamic passwords. The core working principle of TOTP is that the client and server independently calculate completely identical dynamic passwords based on the same key and current time, thus achieving authentication without real-time network interaction. Its workflow can be broken down into three key steps to ensure consistency in the calculation results between the two ends.

[0070] (1) Initial key synchronization: The client and the server need to share a unique "seed key" in advance. This key is usually sent from the server to the client when the client first enters the system. The key is only transmitted once in the initial stage and will not be transmitted again afterward.

[0071] (2) Time window alignment: TOTP relies on time as the calculation benchmark, so the times at both ends must be synchronized. The system sets a fixed "time window" (commonly 30 or 60 seconds) to divide the current time into multiple consecutive time segments. The client and server determine the current "time window" based on the current machine time and use it as the calculation input to ensure that the same time benchmark is used.

[0072] (3) Dynamic password generation: The client and server calculate the password independently, and then confirm the identity through the verification process. The client calculates a hash value based on the locally stored "seed key" and the current "time window" using hash algorithms such as HMAC-SHA1, and then truncates and encodes the value to generate a 6-8 digit dynamic password, which is then sent to the server.

[0073] (4) Verification and Matching: After receiving the dynamic password input by the user, the server will call the locally stored "seed key" and the current "time window" to perform the same hash operation and encoding process as the client, generating a dynamic password on the server side. The dynamic password input by the user is compared with the password generated by the server side. If they match, the authentication is successful; otherwise, access is denied. Only when it is within the valid time window and the seed key used by the server and the client is exactly the same can both ends calculate the same dynamic password, thus ultimately confirming the legitimate identity.

[0074] In this invention, dynamic identity authentication is achieved using the TOTP dynamic password mechanism as the basic framework: the client and server generate a dynamic password synchronously through the TOTP algorithm based on a consistent seed key and time window; during authentication, the user verbally reads the dynamic password presented by the client, the client collects the voice signal, extracts the voiceprint features and voice content and encrypts it; in the encrypted state, the server verifies the legitimacy of the identity by comparing the voiceprint features and ensures the real-time nature of authentication by verifying the consistency between the voice content and the dynamic password, ultimately achieving liveness authentication that combines identity authenticity and real-time operation.

[0075] Step 1.2: Initialize voice feature encryption parameters on both the client and server sides. ,in These are random numbers generated based on the system time. and Indicates the scale of the disruption. The parameter represents the feature selection parameter; the voice feature encryption parameter is used for the encryption and authentication of voice features; through unified hash coding rules and constraints, it is ensured that the distance sensitivity between features can be maintained when the standard voice template and the authenticated voice feature use the same encryption parameters, so as to achieve feature comparability under encryption conditions.

[0076] Step 1.3: Initialize voiceprint feature encryption parameters on both the client and server sides. Voiceprint feature encryption parameters are used for voiceprint feature encryption and authentication. Among them, The coordinate axis parameters represent the overall safety profile. Indicates unit distance, This represents the unit distance contained in each interval. To indicate the number of intervals on the axis, then on the axis... The interval starting from is denoted as , The center point is used as the interval representative value, that is... ;win represents the sliding window parameter, used for local smoothing and feature stabilization of voiceprint features. f is the random projection function, used for binary encoding of voiceprint features. Each random projection is defined by the normal vector ω and the bias term θ, and the random projection equation is y=ω T x+θ, where x is the data point within the sliding window; for each sliding window, if... If the value is 1, the encoding bit is 1; otherwise, the encoding bit is 0. Through the above projection and encoding method, the encoding bits of all sliding windows are combined into a binary encoding sequence, thereby effectively avoiding the influence of feature outliers and improving system fault tolerance.

[0077] Step 2: The client extracts standard speech features and generates encrypted speech templates based on the isolated word set of the dynamic password space, and uploads the encrypted speech templates to the server to establish an encrypted speech template database;

[0078] Step 2.1: For any isolated word w in the dynamic password space, extract speech sequence features based on the standard speech samples of w. ,in Indicates the first Feature vectors of a frame of speech; for any Based on random seed Generate a perturbation vector, where, Represents the Hadema product operation:

[0079] ;

[0080] ;

[0081] Statistical vectors Medium-small index The number of elements pointed to is denoted as . This value is the voice frame. Integer hash encoding; by performing integer hash encoding on all speech frame features, the encrypted speech template of the isolated word w is obtained. ;

[0082] Step 2.2: For each isolated word w, the client sends the data item <isolated word w, encrypted voice template T> to the server for storage, thus building an encrypted voice template database;

[0083] Step 3: The client collects the user's voice data, extracts voiceprint features (static information representing the user's physiological identity features) through a text-independent voiceprint recognition model, and performs encryption processing to generate encrypted voiceprint templates, building an encrypted voiceprint template database; the server only saves the encrypted feature data and does not store any plaintext features, thereby preventing the leakage of user privacy.

[0084] Step 3.1: The client collects the user's voice signal within a fixed duration using sensors, and obtains the voiceprint template using a text-independent voiceprint feature extractor. , template for 3D feature vector;

[0085] Step 3.2: For the given voiceprint template The sliding window parameters win and the random projection function f are generated. 3D binary encoded vector And obtain a preprocessed voiceprint template. , express for 3D eigenvectors; calculate the eigenvalue of each feature point on the axis. The interval above, generate dimensional region vector This yields the voiceprint template. of 5D Encrypted Voiceprint Template s n express for 3D feature vectors.

[0086] Step 3.3: The client will send <User ID, Region Vector> Encrypted voiceprint template The registration information is sent to the server for storage, and an encrypted voiceprint template database is built.

[0087] Step 4: During the authentication phase, the client generates a real-time dynamic password based on the Time Synchronization Dynamic Password (TOTP) algorithm, and generates a two-factor encrypted authentication credential accordingly.

[0088] Step 4.1: The client calculates the real-time dynamic password, which is a random sequence of fixed length within the dynamic password space, and displays it to the user in real time as an authentication command;

[0089] Step 4.2: The user reads aloud the displayed dynamic password, and the client collects the user's real-time voice signal through sensors;

[0090] Step 4.3: The client extracts speech features from the collected speech signal and calculates the encrypted speech credential; specifically, endpoint detection technology is used to segment the continuous speech signal into independent speech segments with the same length as the dynamic password; the sequence features of each speech segment are extracted and denoted as follows. and using encrypted parameters Encryption processing is performed to generate corresponding encrypted voice features. The set of encrypted voice features of all independent voice segments constitutes the encrypted voice authentication credential for this authentication.

[0091] The client simultaneously extracts voiceprint features from the collected speech and calculates encrypted voiceprint credentials; specifically, it extracts voiceprint features using a text-independent voiceprint feature extractor. Using encrypted parameters Encode the data to obtain a binary encoded vector. and encrypted voiceprint features This is the encrypted voiceprint authentication credential used in this authentication.

[0092] Step 4.5: The client sends the generated encrypted voice authentication credential and encrypted voiceprint authentication credential together as two-factor encrypted authentication credentials <User ID, encrypted voice credential, encrypted voiceprint credential> to the server for subsequent encrypted verification and matching processes.

[0093] Step 5: The client performs two-factor encryption authentication on the encrypted voice credentials and encrypted voiceprint credentials uploaded by the client based on the encrypted voice template database and the encrypted voiceprint template database.

[0094] Step 5.1: The server uses the encrypted voice credentials to match against the encrypted voice template database. For each encrypted voice feature in the encrypted voice credentials... The server compares the data with all data items <isolated word w, encrypted speech template T> in the database in turn, and uses the Dynamic Time Warping (DTW) algorithm to calculate the similarity.

[0095] Specifically, for a given <isolated word w, encrypted speech template T> and encrypted speech features... The server first constructs a distance matrix, where each element represents the Euclidean distance between corresponding elements in two encrypted sequences. Starting from the bottom left corner of the matrix, the cumulative distance of each point is updated sequentially using dynamic programming, and the path with the smallest cumulative distance is selected as the optimal matching path. The final cumulative distance is the encrypted voice template. and encrypted voice features The DTW distance; the smaller the DTW distance, the higher the similarity between the two; if a encrypted voice template exists in the encrypted voice template database. This makes it consistent with encrypted voice features. If the DTW distance is minimized, then the encrypted voice feature is considered to be the one that minimizes the DTW distance. The corresponding audio segment is identical to the isolated word w represented by the template <isolated word w, encrypted audio template T>.

[0096] Step 5.2: The server performs matching calculations on all encrypted voice features in the encrypted voice authentication credential to reconstruct the complete dynamic password content. If the parsed dynamic password matches the server's current time-synchronized dynamic password, it is determined that the encrypted voice credential originates from the user's current real-time reading operation, i.e., it passes the liveness detection; otherwise, it is considered an abnormal encrypted credential.

[0097] Step 5.3: The server uses the encrypted voiceprint credentials to retrieve the corresponding data item <User ID, Region Vector> from the encrypted voiceprint template database based on the user ID. Encrypted voiceprint template >. Encrypted voiceprint features for users claiming the same user ID. Calculate the displacement vector And accordingly, recalculate the corresponding region vector. ;like Then it is considered that the encrypted voiceprint features With encrypted voiceprint template If the voiceprint matches the same entity, the voiceprint identity is confirmed to be successfully matched.

[0098] Step 5.4: The system determines that the two-factor encryption authentication is successful only if the verification results of Step 5.2 and Step 5.3 both meet the preset conditions, and confirms that the operation originated from a legitimate live user. At this time, the server can authorize the corresponding operation permissions or allow the execution of subsequent security processes. Otherwise, access will be denied.

[0099] This invention can be applied to remote identity verification scenarios such as online government services or financial transaction platforms. When a user initiates a remote operation request, the client generates a real-time dynamic password and displays it to the user, who then reads the password aloud using a built-in microphone. After collecting the voice signal, the client extracts voiceprint and voice content features and generates encrypted credentials, which are sent to the server. The server verifies the user's identity using an encrypted voiceprint template and simultaneously performs dynamic password recognition on the encrypted voice content, achieving liveness detection and identity confirmation. Only when both voiceprint verification and voice content matching pass, the system determines that the operation originates from a legitimate user, thus allowing the execution of sensitive operations. This solution requires no additional hardware, ensuring the security of remote identity verification while also protecting user privacy and ease of operation, making it suitable for high-security remote service scenarios.

[0100] Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or part of the technical solution, can be embodied in the form of a computer program product.

[0101] The various embodiments in this application are described in a progressive manner. The same or similar parts between the various embodiments can be referred to each other. Each embodiment focuses on describing the differences from other embodiments.

[0102] The scope of protection of this application is not limited to the embodiments described above. Obviously, those skilled in the art can make various modifications and variations to this disclosure without departing from the scope and spirit of this disclosure. If such modifications and variations fall within the scope of the disclosed solution and its equivalents, then the intent of this disclosure also includes these modifications and variations.

Claims

1. A privacy-preserving biometric authentication system based on dynamic voice, characterized in that, Including both client and server sides; The client is used for collecting biometric data, extracting features, encrypting data, and generating authentication credentials. Specifically, the client participates in the initialization of system parameters, including dynamic password parameters, voiceprint feature encryption parameters, and voice feature encryption parameters; The client uses a speech content feature extractor to extract speech features from the standard speech signals of all isolated words in the dynamic password space. After local encryption, it generates an encrypted speech template and uploads it to the server. The client also collects the user's real-time registration speech through a speech sensor, extracts voiceprint features through a voiceprint feature extractor, encrypts the voiceprint template locally, and uploads it to the server. The user submits real-time authentication speech based on the dynamic password given by the system. The client extracts features from the real-time authentication speech to obtain dynamic behavioral features (speech content) and static physiological features (voiceprint features), and generates encrypted voiceprint features and encrypted speech features as authentication features, which are then uploaded to the server. The server is used for storing and managing the encrypted template database, performing encrypted matching calculations, and determining authentication results. Specifically, the server participates in the initialization of system parameters, including dynamic password parameters and voiceprint feature encryption parameters; The server receives encrypted voice templates and encrypted voiceprint templates, and constructs encrypted voice template databases and encrypted voiceprint template databases. After receiving authentication credentials uploaded by the client, the server performs comparison and verification based on the encrypted voiceprint template database and the encrypted voice template database, and completes remote two-factor encrypted identity authentication in conjunction with a dynamic password verification mechanism.

2. The privacy-preserving biometric authentication system based on dynamic voice according to claim 1, characterized in that, The two-factor encrypted identity authentication specifically includes an encrypted voiceprint authentication mechanism and an encrypted voice recognition mechanism. In the encrypted voiceprint authentication mechanism, the client collects user voice data, extracts voiceprint features that can represent the user's identity using a text-independent VPR model, and generates an encrypted voiceprint template which is then uploaded to the server for storage. The user submits real-time authentication voice according to the dynamic password given by the system. The client extracts voiceprint features from the real-time authentication voice and generates encrypted voiceprint features as authentication credentials, which are then sent to the server for verification. The server performs encrypted voiceprint matching based on the encrypted voiceprint template database. When the similarity between the encrypted voiceprint features and the preset encrypted voiceprint template is within a preset threshold range, it is determined that the subject corresponding to the voiceprint is consistent with the registered identity, i.e., the voiceprint identity authentication is passed. In the encrypted speech recognition mechanism, the Dynamic Time Warping (DTW) algorithm is introduced to achieve optimal matching between feature sequences with inconsistent lengths. Combined with the position-sensitive integer hashing (PSH) algorithm, an efficient and secure encrypted speech recognition mechanism is constructed. Specifically, speech segmentation is performed using endpoint detection technology to divide the continuous speech signal into multiple independent speech segments corresponding to each character in the dynamic password. For a given speech template, namely the preset isolated word standard speech pattern and the real-time authentication speech features, the system determines whether the two represent the same password character through template matching. For any isolated word in the dynamic password space, the client preprocesses the speech signal using an ASR model, converts the preprocessed speech signal into a feature vector that can be used for template matching, and generates a corresponding encrypted speech template, thus constructing an encrypted speech template database. The user submits real-time authentication speech within a valid time window based on the dynamic password generated by the client. The client extracts the speech features, generates encrypted speech features as authentication credentials, and sends them to the server. The server performs encrypted speech recognition based on the encrypted speech template database. When the recognition result is completely consistent with the current dynamic password on the server, it is determined that the authentication credentials originate from the user's real-time operation.

3. A privacy-preserving biometric authentication method based on dynamic voice, implemented through the privacy-preserving biometric authentication system based on dynamic voice as described in claim 1, characterized in that... Includes the following steps: Step 1: Initialization phase, the client and server initialize system parameters respectively for encrypted template generation and secure communication; Step 2: The client extracts standard speech features and generates encrypted speech templates based on the isolated word set of the dynamic password space, and uploads the encrypted speech templates to the server to establish an encrypted speech template database; Step 3: The client collects the user's voice data, extracts voiceprint features through a text-independent voiceprint recognition model, performs encryption processing to generate encrypted voiceprint templates, and builds an encrypted voiceprint template database. Step 4: During the authentication phase, the client generates a real-time dynamic password based on the Time Synchronization Dynamic Password (TOTP) algorithm, and generates a two-factor encrypted authentication credential accordingly. Step 5: The client performs two-factor encryption authentication on the encrypted voice credentials and encrypted voiceprint credentials uploaded by the client, based on the encrypted voice template database and the encrypted voiceprint template database.

4. The privacy-preserving biometric authentication method based on dynamic voice according to claim 3, characterized in that, Step 1 specifically includes the following steps: Step 1.1: Initialize the Time Synchronization Dynamic Password Module (TOTP) on both the client and server sides; The Time Synchronization Dynamic Password (TOTP) module allows the client and server to independently calculate identical dynamic passwords based on the same key and current time, thus achieving authentication without real-time network interaction. The workflow of the TOTP module is as follows: Step 1.2: Initialize voice feature encryption parameters on both the client and server sides. ,in These are random numbers generated based on the system time. and Indicates the scale of the disruption. Indicates feature selection parameters; Step 1.3: Initialize voiceprint feature encryption parameters on both the client and server sides. ,in, The coordinate axis parameters represent the overall safety profile. Indicates unit distance, This represents the unit distance contained in each interval. To indicate the number of intervals on the axis, then on the axis... The interval starting from is denoted as , The center point is used as the interval representative value, that is... ; win represents the sliding window parameters, f is the random projection function, each random projection is defined by the normal vector ω and the bias term θ, and the random projection equation is y=ω T x+θ, where x is the data point within the sliding window; for each sliding window, if... If the value is 1, then the encoding bit is 1; otherwise, the encoding bit is 0.

5. The privacy-preserving biometric authentication method based on dynamic voice according to claim 4, characterized in that, Step 1.1 specifically includes the following steps: Step 1.1.1: Initial Key Synchronization: The client and server need to share a unique seed key in advance; the seed key is sent from the server to the client when the client first enters the system; the seed key is only transmitted once in the initial stage and will not be transmitted again afterward. Step 1.1.2: Time Window Alignment: Set a fixed time window to divide the current time into multiple consecutive time segments; the client and server will determine the current time window based on the current machine time and use it as the calculation input to ensure that the same time base is used; Step 1.1.3: Dynamic Password Generation: The client and server independently calculate the password and then verify the identity through the verification process; the client calculates a hash value based on the locally stored seed key and the current time window using a hash algorithm, then truncates and encodes the hash value to generate a 6-8 digit dynamic password, and sends it to the server. Step 1.1.4: Verification and Matching: After receiving the dynamic password entered by the user, the server will call the locally stored seed key and the current time window to perform the same hash operation and encoding process as the client to generate a dynamic password on the server side. The dynamic password entered by the user is compared with the password generated by the server side. If they match, the authentication is successful; otherwise, access is denied. Only when it is within the valid time window and the seed keys used by the server and the client are exactly the same can both ends calculate the same dynamic password, thus ultimately confirming the legitimate identity.

6. The privacy-preserving biometric authentication method based on dynamic voice according to claim 3, characterized in that, Step 2 specifically includes the following steps: Step 2.1: For any isolated word w in the dynamic password space, extract speech sequence features based on the standard speech samples of w. ,in Indicates the first Feature vectors of a frame of speech; for any Based on random seed Generate a perturbation vector, where, Represents the Hadema product operation: ; ; Statistical vectors Medium-small index The number of elements pointed to is denoted as . This value is the voice frame. Integer hash encoding; by performing integer hash encoding on all speech frame features, the encrypted speech template of the isolated word w is obtained. ; Step 2.2: For each isolated word w, the client sends the data item <isolated word w, encrypted voice template T> to the server for storage, thus building an encrypted voice template database.

7. The privacy-preserving biometric authentication method based on dynamic voice according to claim 3, characterized in that, Step 3 specifically includes the following steps: Step 3.1: The client collects the user's voice signal within a fixed duration using sensors, and obtains the voiceprint template using a text-independent voiceprint feature extractor. , template for 3D feature vector; Step 3.2: For the given voiceprint template The sliding window parameters win and the random projection function f are generated. 3D binary encoded vector And obtain a preprocessed voiceprint template. , express for 3D eigenvectors; calculate the eigenvalue of each feature point on the axis. The interval above, generate dimensional region vector This yields the voiceprint template. of 5D encrypted voiceprint template s n express for 3D feature vector; Step 3.3: The client will send <User ID, Region Vector> Encrypted voiceprint template The registration information is sent to the server for storage, and an encrypted voiceprint template database is built.

8. The privacy-preserving biometric authentication method based on dynamic voice according to claim 3, characterized in that, Step 4 specifically includes the following steps: Step 4.1: The client calculates the real-time dynamic password, which is a random sequence of fixed length within the dynamic password space, and displays it to the user in real time as an authentication command; Step 4.2: The user reads aloud the displayed dynamic password, and the client collects the user's real-time voice signal through sensors; Step 4.3: The client extracts speech features from the collected speech signal and calculates the encrypted speech credential; specifically, endpoint detection technology is used to segment the continuous speech signal into independent speech segments with the same length as the dynamic password; the sequence features of each speech segment are extracted and denoted as follows. and using encrypted parameters Encryption processing is performed to generate corresponding encrypted voice features. The set of encrypted voice features of all independent voice segments constitutes the encrypted voice authentication credential for this authentication. The client simultaneously extracts voiceprint features from the collected speech and calculates encrypted voiceprint credentials; specifically, it extracts voiceprint features using a text-independent voiceprint feature extractor. Using encrypted parameters Encode the data to obtain a binary encoded vector. and encrypted voiceprint features This is the encrypted voiceprint authentication credential used in this authentication. Step 4.5: The client sends the generated encrypted voice authentication credential and encrypted voiceprint authentication credential together as a two-factor encrypted authentication credential <User ID, encrypted voice credential, encrypted voiceprint credential> to the server.

9. A privacy-preserving biometric authentication method based on dynamic voice according to claim 3, characterized in that, Step 5 specifically includes the following steps: Step 5.1: The server uses the encrypted voice credentials to match them in the encrypted voice template database; for each encrypted voice feature in the encrypted voice credentials... The server compares the data with all data items <isolated word w, encrypted speech template T> in the database in turn, and uses the Dynamic Time Warping (DTW) algorithm to calculate the similarity. Step 5.2: The server performs matching calculations on all encrypted voice features in the encrypted voice authentication credential to restore the complete dynamic password content; if the parsed dynamic password is consistent with the dynamic password synchronized with the current time on the server, it is determined that the encrypted voice credential originates from the user's current real-time reading operation, i.e., it passes the detection; otherwise, it is considered that the encrypted credential is abnormal. Step 5.3: The server uses the encrypted voiceprint credentials to retrieve the corresponding data item <User ID, Region Vector> from the encrypted voiceprint template database based on the user ID. Encrypted voiceprint template >; For encrypted voiceprint features that declare the same user ID Calculate the displacement vector And accordingly, recalculate the corresponding region vector. ;like Then it is considered that the encrypted voiceprint features With encrypted voiceprint template If the voiceprint matches the same entity, the voiceprint identity is confirmed to be successfully matched. Step 5.4: The system determines that the two-factor encryption authentication is successful only if the verification results of Step 5.2 and Step 5.3 both meet the preset conditions, and confirms that the operation originated from a legitimate live user. At this time, the server can authorize the corresponding operation permissions or allow the execution of subsequent security processes. Otherwise, access will be denied.

10. A privacy-preserving biometric authentication method based on dynamic voice according to claim 7, characterized in that, Step 5.1 specifically involves: for a given <isolated word w, encrypted speech template T> and encrypted speech features... The server first constructs a distance matrix, where each element represents the Euclidean distance between corresponding elements in two encrypted sequences. Starting from the bottom left corner of the matrix, the cumulative distance of each point is updated sequentially using dynamic programming, and the path with the smallest cumulative distance is selected as the optimal matching path. The final cumulative distance is the encrypted voice template. and encrypted voice features The DTW distance; the smaller the DTW distance, the higher the similarity between the two; if a encrypted voice template exists in the encrypted voice template database. This makes it consistent with encrypted voice features. If the DTW distance is minimized, then the encrypted voice feature is considered to be the one that minimizes the DTW distance. The corresponding audio segment is identical to the isolated word w represented by the template <isolated word w, encrypted audio template T>.