System and method for enhancing session security based on voiceprint recognition
By adopting a voiceprint recognition-based system in real-time communication applications, using voiceprint collection, preprocessing, extraction, matching and early warning modules, the problems of identity disguise and deception in real-time communication applications are solved, and the security and response speed of the conversation are significantly improved.
Patent Information
- Application Number
- CN202510264431.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-06
- Publication Date
- 2025-06-13
AI Technical Summary
In real-time communication applications, there are third parties who participate in chat sessions by pretending their identities and synthesizing voices, committing fraud and fraud, resulting in the threat of the security of the session.
A system based on voiceprint recognition is adopted to collect audio signals through a microphone and a sound card, and combine wearable devices, noise reduction devices and isolation devices to monitor and record user's voice in real time. Voiceprint features are extracted through the Mel frequency cepspectral coefficient and convolutional neural network algorithm, and match and warning are performed through dynamic time regularization and deep neural network algorithm.
It significantly improves the security of the session. Through real-time voiceprint matching and early warning mechanisms, it can promptly identify and warn of potential identity disguises and deceptions, improving the accuracy and response speed of user identity verification.
Smart Images

Figure CN120151007A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to technologies such as conversation systems, voiceprint recognition, identity authentication, and security alerts, and particularly relates to a system and method for enhancing conversation security based on voiceprint recognition. Background Art
[0002] In the era of mobile Internet, instant messaging applications based on mobile phone terminals are everywhere and are widely used in life. In instant messaging application APPs, the content and methods of chat conversations, including text, voice, real-time audio, real-time video, etc., with the advent of the 5G era, the proportion of media-like voice, real-time audio, real-time video, etc. in chat conversations is increasing.
[0003] In the field of life, the authentication of the identity of the login user in application APPs generally uses the traditional methods of username and password, as well as biometric methods of fingerprint, voiceprint, face recognition, etc. of the human body to achieve. However, during specific chat conversations within the application, real-time identity authentication is not performed. Therefore, there are third parties and fraudsters who participate in chat conversations by disguising their identities and synthesizing voices, implementing deception and fraud to achieve ulterior motives.
[0004] In special fields, through information interaction of the instant messaging system, commands are issued by superiors and tasks are cooperatively executed. In this case, there are also identity disguises and frauds, and serious consequences will be produced. Summary of the Invention
[0005] The purpose of the present invention is to propose a system and method for enhancing conversation security based on voiceprint recognition.
[0006] The technical solution for achieving the purpose of the present invention is: A system for enhancing conversation security based on voiceprint recognition, comprising:
[0007] Device layer, including a microphone and a sound card, wearable devices, noise reduction devices, and isolation devices, for collecting audio information, wherein the microphone is used to capture audio analog signals, and the sound card converts the audio analog signals into digital signals; wearable devices, with portability, for real-time monitoring and recording of the user's voice; noise reduction devices for reducing background noise in the environment and improving the quality of audio signals; isolation devices providing sound isolation to ensure the purity of the recording;
[0008] The storage layer is deployed on both mobile terminals and servers simultaneously and is used to store information, including personnel information, session information, alarm information, voiceprint data, configuration information, and log information. Among them, personnel information is used to identify and verify user identities, and the parameters include name, ID number, and voiceprint feature vector; session information is used to record the interaction process between users and the system, for auditing and analyzing user behaviors, and the parameters include session start time, end time, participants, session duration, and session content summary; alarm information is used for the system to generate alarms when abnormal behaviors or security threats are detected, and the parameters include alarm time, alarm type, alarm level, and alarm description; configuration information is used to customize system behaviors, and the parameters include system parameters and user preference settings; voiceprint data is the core of voiceprint recognition, used to verify user identities and trigger security warnings, and the parameters include sampling rate, bit depth, frame length, and frame shift; log information records system operations and events, for troubleshooting and security monitoring, and the parameters include log time, operation type, operation result, operating user, and affected resources;
[0009] The processing layer includes a voiceprint acquisition module, a voiceprint storage module, a voiceprint preprocessing module, a voiceprint extraction module, a voiceprint matching module, and a voiceprint warning module. Among them, the voiceprint acquisition module is used for free speech acquisition and emotional speech acquisition; the voiceprint storage module is used to store user audio digital signals, audio signals after voiceprint preprocessing, fixed-length feature vectors after voiceprint extraction processing, preset standard user audio feature vectors, convolutional neural network models, and deep neural network models; the voiceprint preprocessing module is used for endpoint detection and speech enhancement; the voiceprint extraction module uses Mel Frequency Cepstral Coefficients and convolutional neural network algorithms to learn local and global features of audio data; the voiceprint matching module uses Dynamic Time Warping and deep neural network algorithms to determine whether two voice samples are from the same speaker; the voiceprint warning module uses methods based on semantic analysis and acoustic features to determine whether there is a warning situation;
[0010] The display layer provides the UI page of the APP, including user login, personnel list, message session, search filter, and permission management.
[0011] Furthermore, the voiceprint acquisition module captures the user's voice signal through a high-precision audio acquisition device and converts it into a digital format, samples and quantizes the analog signal according to the set sampling rate and bit depth, to ensure that the acquired audio data is clear, accurate, and can effectively reflect the unique voiceprint characteristics of the speaker.
[0012] Further, the voiceprint preprocessing module performs voice activity detection on the audio digital signal obtained by the voiceprint acquisition module to form an audio signal marked with voice activity regions, which is represented as a binary marking sequence indicating whether there is voice activity in each time period; the audio signal marked with voice activity regions is enhanced through voice enhancement, and the combination of voice activity detection and voice enhancement is achieved through the following three methods:
[0013] 1) Sequential processing: First, perform voice activity detection to determine the valid part of the voice, and then apply voice enhancement technology to these parts;
[0014] 2) Iterative optimization: Voice activity detection and voice enhancement are performed iteratively, that is, the preliminary voice activity detection results are used to guide the preliminary voice enhancement, and the enhanced voice signal is fed back to the voice activity detection for more accurate voice activity detection;
[0015] 3) Parameter dynamic adjustment: According to the characteristics of the real-time signal and the changes in the noise environment, dynamically adjust the threshold of voice activity detection and the parameters of the voice enhancement algorithm to achieve the best preprocessing effect.
[0016] Further, the voiceprint extraction module has the following processing steps:
[0017] 1) Segment the audio signal obtained by the voiceprint preprocessing module into multiple time windows to obtain windowed audio data, and each window contains audio data of dozens of milliseconds;
[0018] 2) Perform FFT on each windowed audio data to convert the time-domain signal into a frequency-domain signal;
[0019] 3) Apply a Mel filter bank to filter the FFT result, map the frequency-domain signal to the Mel frequency scale, simulate the human ear's perception of different frequencies, and obtain Mel spectrum data;
[0020] 4) Perform a logarithmic operation on the Mel spectrum data and calculate the energy of each filter to obtain a logarithmic energy spectrum;
[0021] 5) Apply DCT to convert the logarithmic energy spectrum into MFCC coefficients, and the MFCC coefficients capture the pitch, timbre, and dynamic change characteristics of the audio signal;
[0022] 6) Extract the first N coefficients from the MFCC coefficients to form an MFCC feature vector;
[0023] 7) Stack the MFCC feature vectors of several consecutive frames together to form a three-dimensional data structure, where the first dimension is the time series, the second dimension is the frequency dimension, and the third dimension is the MFCC coefficients;
[0024] 8) Construct a CNN model, which includes multiple convolutional layers, pooling layers, and fully connected layers. The convolutional layers are used to extract local features and extract multiple two-dimensional feature maps from the three-dimensional data structure. The pooling layers are used to reduce the spatial dimension of the features and convert the feature maps output by the convolutional layers into downsampled feature maps. The fully connected layers are used to integrate the features output by the pooling layers and perform final classification or regression, and finally obtain a fixed-length feature vector for voiceprint matching processing, where
[0025] The mathematical expression of the CNN model is as follows:
[0026]
[0027] where,
[0028] "X [l-1] " represents the activation output of the previous layer, that is, the three-dimensional data structure of the MFCC feature vector;
[0029] "W [l] " represents the convolutional kernel of the l-th layer;
[0030] "b [l] " represents the bias of the l-th layer;
[0031] "*" represents the convolution operation;
[0032] "W [l] *X [l-1] +b [l] " represents the convolution operation of the l-th layer, applying the convolutional kernel to the input data and then adding the bias to obtain the original output of the convolutional layer, that is, the two-dimensional feature map;
[0033] "ReLU" represents the activation function, which is used to increase the non-linearity of the network, enable the network to learn more complex features, and convert the output of the convolutional layer into non-negative values;
[0034] "MaxPool" represents the max pooling operation, which converts the feature maps output by the convolutional layer into downsampled feature maps;
[0035] "W [L] " represents the connection weights of the fully connected layer;
[0036] represents the output of the CNN model, which is a fixed-length feature vector for comparison and recognition.
[0037] Furthermore, the voiceprint matching processing module has the following processing steps:
[0038] 1) Use DTW to align the feature sequences of the two voice samples, namely the feature vector obtained from voiceprint extraction processing and the pre-set standard user audio feature vector, to solve the time offset and variation between them;
[0039] 2) Use the aligned feature sequence as input and extract the speaker feature representation through a pre-trained DNN model;
[0040] 3) Use the speaker feature representation extracted by the DNN to calculate the similarity between speakers. Use the Euclidean distance and cosine similarity methods to measure the distance or similarity between features and obtain the similarity score;
[0041] 4) Use the similarity value output in step 3 to determine whether two voice samples are from the same speaker and obtain the matching result. The mathematical expression is as follows:
[0042]
[0043] where,
[0044] "X', Y'" represent the feature vector obtained by speaker feature extraction processing and the preset standard user audio feature vector;
[0045] "D" represents the deep learning model, which is used to extract a more advanced feature representation from the feature vector;
[0046] "D(X'), D(Y')" represent the speaker feature representations obtained after the two feature vectors undergo feature extraction (DNN);
[0047] "T" represents a threshold, which is used for speaker similarity matching judgment;
[0048] "||……||" represents the Euclidean norm of the vector;
[0049] "·" represents the dot product of the vectors.
[0050] Furthermore, the speaker warning processing module has the following processing steps:
[0051] 1) Use automatic speech recognition technology to convert the audio signal into text based on semantic analysis, analyze the semantic meaning of the speech content, understand the actual meaning and context information in the statement, and identify possible threatening language, inappropriate remarks or keywords for the warning system to trigger a security response;
[0052] 2) The method based on acoustic features is responsible for analyzing the acoustic attributes of the speech signal, including pitch, volume, speech rate, rhythm, etc., and detecting abnormal speech patterns or behaviors;
[0053] 3) Evaluate the speech information from different dimensions through a weighted fusion method, understand both the actual content of the speech and analyze the voice characteristics of the speaker, so as to more comprehensively judge whether there is a warning situation.
[0054] Furthermore, the storage layer, the processing layer, and the display layer are all software parts, deployed on the server and the terminal device, and the device layer is the hardware part, which is selectively connected to the server and the terminal according to the scenario requirements.
[0055] A method for enhancing session security based on voiceprint recognition, using the system for enhancing session security based on voiceprint recognition as described above, to achieve enhanced session security based on voiceprint recognition. The specific steps are as follows:
[0056] (1) Collect audio data through the terminal microphone, collect voiceprint information, associate and store the logged-in user information. This step is implemented by the voiceprint collection and voiceprint storage modules.
[0057] (2) During the chat session, the user communicates through voice files or real-time voice, obtains audio information in real time, and performs preprocessing of the audio data. This step is implemented by the voiceprint preprocessing module.
[0058] (3) The preprocessed audio data is subjected to main voice extraction and voiceprint calculation to obtain the user's real-time voiceprint information. This step is implemented by the voiceprint extraction module.
[0059] (4) After obtaining the voiceprint information, calculate the matching degree between the real-time voiceprint and the standard voiceprint by matching the standard voiceprint information of the user stored in the database. This step is implemented by the voiceprint matching module.
[0060] (5) Compare the matching degree value with a preset threshold. If it is greater than the threshold, the user's identity is considered reliable; otherwise, send a warning to the recipient. This step is implemented by the voiceprint warning module. The voiceprint warning module also combines the methods based on semantic analysis and acoustic features to identify the languages and audio metrics that need to be warned about and outputs a warning signal.
[0061] (6) The recipient ends the session or conducts further identity confirmation in a timely manner according to the warning information.
[0062] Compared with the prior art, the significant advantages of the present invention are as follows: 1) It not only focuses on the physical characteristics of the voice, but also extracts the feature vectors related to emotions by analyzing the intonation, volume, speech rate, etc. in the voice, increasing the depth and meticulousness of voiceprint recognition; 2) Combining the traditional MFCC algorithm with CNN can automatically learn and extract the local and global features suitable for voiceprint recognition from the original audio data, improving the robustness and accuracy of voiceprint recognition; 3) Automatically adjust the voiceprint model according to the changes in the user's voice or the environment. This adaptive ability enables the system to maintain high performance under changing conditions, improving the stability and accuracy of voiceprint recognition; 4) Design a real-time voiceprint matching and warning mechanism. When it is detected that the voiceprint matching degree is lower than the preset threshold, a warning can be sent in a timely manner, significantly improving the security and response speed of the session. Description of the Drawings
[0063] Figure 1 It is a system architecture diagram for enhancing session security based on voiceprint recognition.
[0064] Figure 2 It is a diagram of system devices and connection relationships for enhancing session security based on voiceprint recognition.
[0065] Figure 3 It is a system interaction diagram for enhancing session security based on voiceprint recognition.
[0066] Figure 4 It is a real-time voiceprint processing flow chart. Detailed implementation manners
[0067] In order to make the objectives, technical solutions and advantages of this application clearer, the following further elaborates on this application in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely used to explain this application and are not used to limit this application.
[0068] As Figure 1 shown, a system for enhancing session security based on voiceprint recognition according to the present invention is divided into a device layer, a storage layer, a processing layer and a display layer.
[0069] The device layer includes a microphone and a sound card, wearable devices, noise reduction devices, isolation devices and other external acquisition devices for collecting audio information. Among them, the microphone is used to capture audio analog signals, and the sound card converts these analog signals into digital signals for system processing; wearable devices have portability, can monitor and record the user's voice in real time, and are closer to the user's body, which helps to improve the accuracy of voiceprint recognition; noise reduction devices are used to reduce background noise in the environment and improve the quality of audio signals; isolation devices provide a high degree of sound isolation to ensure the purity of recordings. The coordinated action of these devices significantly improves the accuracy and security of voiceprint recognition, while optimizing the user experience and adapting to diverse application scenarios.
[0070] The storage layer is deployed on both the mobile terminal and the server. The stored information mainly includes personnel information, session information, alarm information, voiceprint data, configuration information, log information, etc. Among them, personnel information is used to identify and verify the user's identity, and the parameters include name, ID number, voiceprint feature vector, etc.; session information is used to record the interaction process between the user and the system, for auditing and analyzing user behavior, and the parameters include session start time, end time, participants, session duration, session content summary, etc.; alarm information is used for the system to generate alarms when abnormal behaviors or security threats are detected, and the parameters include alarm time, alarm type (such as identity forgery, illegal access, etc.), alarm level, alarm description, etc.; configuration information is used to customize the system behavior, and the parameters include system parameters (such as noise reduction level, recognition threshold, etc.), user preference settings, etc.; voiceprint data is the core of voiceprint recognition, used to verify the user's identity and trigger security warnings, and the parameters include sampling rate, bit depth, frame length, frame shift, etc.; log information records system operations and events, for troubleshooting and security monitoring, and the parameters include log time, operation type, operation result, operating user, affected resources, etc.
[0071] The processing layer includes processing modules such as voiceprint acquisition, voiceprint storage, voiceprint preprocessing, voiceprint extraction, voiceprint matching, and voiceprint warning. Among them, the voiceprint acquisition module uses free speech acquisition and emotional speech acquisition; the voiceprint storage module uses feature vector storage and voice model storage. The voiceprint preprocessing module uses endpoint detection and speech enhancement, the voiceprint extraction module uses Mel Frequency Cepstral Coefficients (MFCC) and Convolutional Neural Network Algorithm (CNN), the voiceprint matching module uses Dynamic Time Warping (DTW) and Deep Neural Network Algorithm (DNN), and the voiceprint warning module uses methods based on semantic analysis and acoustic features.
[0072] The display layer mainly refers to the UI page of the APP, including user login, personnel list, message session, search and filtering, permission management, etc.
[0073] The detailed design of the functional modules in the processing layer is as follows:
[0074] (1) In the voiceprint acquisition process, a high-precision audio acquisition device is used to capture the user's voice signal and convert it into a digital format. Using analog-to-digital conversion technology, the analog signal is sampled and quantized according to the set sampling rate and bit depth to ensure that the collected audio data is clear, accurate, and can effectively reflect the unique voiceprint characteristics of the speaker. This module is implemented at the device layer and outputs the user's audio digital signal.
[0075] (2) In voiceprint preprocessing, two methods of voice activity detection and voice enhancement are combined. Voice activity detection is responsible for determining the start and end points of the voice signal, that is, identifying when to start speaking and when to stop speaking in the audio, so as to process or analyze the effective voice part; voice enhancement is to improve the quality of the voice signal, so that the system can process and recognize the voice signal more accurately. At the same time, the system dynamically adjusts the threshold and parameters of voice activity detection in a timely manner according to the characteristics and changes of the audio signal and the noise environment, and dynamically adjusts the sensitivity and accuracy of voice activity detection according to the characteristics of the real-time signal.
[0076] The audio digital signal obtained by the voiceprint acquisition module forms an audio signal marked with the voice activity area after voice activity detection. Usually, it is a binary marking sequence indicating whether there is voice activity in each time period; the audio signal marked with the voice activity area forms an enhanced voice signal after voice enhancement, improving the clarity and intelligibility of the voice. The combination of the two methods of voice activity detection and voice enhancement is achieved in the following 3 ways:
[0077] 1) Sequential processing: First, perform voice activity detection to determine the effective part of the voice, and then apply voice enhancement technology to these parts.
[0078] 2) Iterative optimization: Voice activity detection and voice enhancement can be performed iteratively, that is, the initial voice activity detection results are used to guide the initial voice enhancement, and the enhanced voice signal is fed back to voice activity detection for more accurate voice activity detection.
[0079] 3) Parameter dynamic adjustment: According to the characteristics of the real-time signal and the changes in the noise environment, dynamically adjust the threshold of voice activity detection and the parameters of the voice enhancement algorithm to achieve the best preprocessing effect.
[0080] (3) In voiceprint extraction processing, Mel Frequency Cepstral Coefficients (MFCC) are combined with the Convolutional Neural Network (CNN) algorithm, enabling the CNN to learn both local and global features of the original audio data simultaneously, thereby improving the robustness and distinguishability of voiceprint features.
[0081] The main steps of voiceprint extraction processing are as follows:
[0082] 1) Divide the audio signal obtained from voiceprint preprocessing into multiple time windows to obtain windowed audio data. Each window usually contains audio data of dozens of milliseconds.
[0083] 2) Perform FFT on each windowed audio data to convert the time-domain signal into a frequency-domain signal.
[0084] 3) Apply a Mel filter bank to filter the FFT result, map the frequency-domain signal to the Mel frequency scale, simulate the human ear's perception of different frequencies, and obtain Mel spectrum data.
[0085] 4) Perform a logarithmic operation on the Memel spectral data and calculate the energy of each filter to obtain the logarithmic energy spectrum;
[0086] 5) Apply DCT to convert the logarithmic energy spectrum into MFCC coefficients, which capture key features such as pitch, timbre, and dynamic changes of the audio signal;
[0087] 6) Extract the first N coefficients from the MFCC coefficients, and these coefficients form the MFCC feature vector;
[0088] 7) Stack the MFCC feature vectors of consecutive frames together to form a three-dimensional data structure, where the first dimension is the time series, the second dimension is the frequency dimension, and the third dimension is the MFCC coefficients;
[0089] 8) Design a CNN model, which includes multiple convolutional layers, pooling layers, and fully connected layers. The convolutional layers are used to extract local features and extract multiple two-dimensional feature maps from the three-dimensional data structure; the pooling layers are used to reduce the spatial dimension of the features and convert the feature maps output by the convolutional layers into downsampled feature maps. The pooling operation reduces the spatial size of the feature maps but retains the most important information; the fully connected layers are used to integrate features and perform final classification or regression, and finally obtain a fixed-length feature vector for speaker verification processing.
[0090] Through the multi-layer structure of the CNN, the model learns the mapping from MFCC coefficients to higher-level features. These higher-level features are non-linear combinations of the original MFCC coefficients and no longer directly correspond to any specific attributes of the audio signal. Instead, they capture deeper patterns and structures and can express more complex speech attributes, such as the emotional state of the speaker and the change in speech rate. These attributes are not easily directly observable in the original MFCC coefficients. These higher-level features have stronger expressive power and discriminability.
[0091] The mathematical expressions applied by the multi-layer structure of the CNN are as follows:
[0092]
[0093] Among them,
[0094] "X [l-1] " represents the activation output of the previous layer, that is, the three-dimensional data structure of the MFCC feature vector;
[0095] "W [l] " represents the convolutional kernel of the l-th layer;
[0096] "b [l] " represents the bias of the l-th layer;
[0097] "*" represents the convolution operation;
[0098] “W [l] *X [l-1] +b [l] ” represents the convolutional operation of the l-th layer. This operation applies the convolutional kernel to the input data and then adds the bias to obtain the original output of the convolutional layer, i.e., the two-dimensional feature map;
[0099] “ReLU” represents the activation function, which is used to increase the non-linear characteristics of the network so that the network can learn more complex features. It converts the output of the convolutional layer into non-negative values;
[0100] “MaxPool” represents the max pooling operation, which converts the feature map output by the convolutional layer into a feature map with reduced dimensions;
[0101] “W [L] ” represents the connection weights of the fully connected layer;
[0102] represents the output of the CNN model. The CNN finally generates a fixed-length feature vector that can be used for comparison and recognition.
[0103] (4) In the voiceprint storage and processing, the feature vectors of the user's voiceprint and the speech model are effectively stored, managed, and updated to support the requirements of the voiceprint recognition system. It includes the user's audio digital signal, the audio signal after voiceprint preprocessing, the fixed-length feature vector after voiceprint extraction and processing, the trained CNN model, the trained DNN model, and the preset standard user audio feature vector. These data can be retrieved and used in the subsequent voiceprint matching and verification processes.
[0104] (5) In the voiceprint matching process, dynamic time warping is combined with the deep neural network algorithm to overcome the limitations of feature design in traditional methods and give full play to the advantages of deep learning in voiceprint feature extraction, so as to make full use of the flexibility and accuracy of DTW for time information.
[0105] The main steps of the voiceprint matching process are as follows:
[0106] 1) Use DTW to align the feature sequences of the two voice samples, namely the feature vector obtained from voiceprint extraction and the preset standard user audio feature vector, to solve the time offset and variation between them. This step will generate an aligned feature sequence, making their corresponding positions on the time axis more consistent;
[0107] 2) Take the aligned feature sequence as the input and extract the voiceprint feature representation through the pre-trained DNN model;
[0108] 3) Calculate the similarity between speakers using the voiceprint feature representation extracted by DNN. Use the Euclidean distance and cosine similarity methods to measure the distance or similarity between features, and obtain a similarity score;
[0109] 4) Use the similarity value output in step 3 to determine whether two voice samples are from the same speaker, and obtain a matching result.
[0110] The mathematical expressions applied in the above steps are as follows:
[0111]
[0112] Among them,
[0113] "X', Y'" represent the feature vectors obtained from voiceprint extraction processing and the preset standard user audio feature vectors;
[0114] "D" represents a deep learning model used to extract higher-level feature representations from feature vectors;
[0115] "D(X'), D(Y')" represent the voiceprint feature representations obtained after two feature vectors undergo feature extraction (DNN);
[0116] "T" represents a threshold used for voiceprint similarity matching judgment;
[0117] "||……||" represents the Euclidean norm (L2 norm) of a vector;
[0118] "·" represents the dot product of vectors.
[0119] (6) In voiceprint early warning processing, combine the method based on semantic analysis and the method based on acoustic features. The input is the audio signal after voiceprint preprocessing, and the output is an early warning signal indicating that safety response measures need to be taken immediately.
[0120] 1) The method based on semantic analysis uses automatic speech recognition technology to convert the audio signal into text, analyzes the semantic meaning of the speech content, understands the actual meaning and context information in the statement, and can identify possible threatening languages, inappropriate remarks or keywords for the early warning system to trigger a safety response.
[0121] 2) The method based on acoustic features is responsible for analyzing the acoustic attributes of the speech signal, such as pitch, volume, speech rate, rhythm, etc., and can detect abnormal speech patterns or behaviors, such as nervousness, anxiety or aggressive tones, which may be indicators of potential risks
[0122] 3) By combining these two methods, the system evaluates voice information from different dimensions through a weighted fusion method, not only understanding the actual content of the voice but also analyzing the voice characteristics of the speaker, so as to more comprehensively judge whether there is a warning situation and improve the accuracy and reliability of the warning system.
[0123] The hardware devices of the system and their connection relationships are as Figure 2 shown. The system includes a server, mobile terminal devices, and external audio collection devices. Figure 1 The storage layer, processing layer, and display layer described above are all software parts, deployed on the server and terminal devices. The device layer is the hardware part, which can be selectively connected to the server and the terminal according to the scenario requirements. The server serves as the backend storage of the entire system, responsible for storing system information and processing the system session interaction logic; the terminal device faces the user, responsible for human-computer interaction operations, and at the same time, the internally integrated microphone device can also collect audio; external audio collection devices, such as sound cards, wearable devices, noise reduction devices, isolation devices, etc., are responsible for collecting audio data.
[0124] The processing flow of the entire information system: First, collect the user's voice information, extract the voiceprint for preservation, and associate the user account information; then, in the chat session, process the voice information in real time, reduce the interference of noise and background sound through audio preprocessing, and obtain the characteristic information of the voice through voiceprint extraction; judge the matching degree between the voice of the actual speaker and the application login person through voiceprint matching; through voiceprint warning, send a warning message to the voice receiver, and perform further operations such as ending the session and confirming the identity in a timely manner, so as to enhance the security of the system in the entire chat session. The specific steps are as follows:
[0125] (1) Collect audio data through the terminal microphone, collect voiceprint information, associate the logged-in user information, and store it. This step is implemented by the voiceprint collection and voiceprint storage modules.
[0126] (2) When the user is in a chat session, communicate through voice files or real-time voice. The APP obtains the audio information in real time and preprocesses the audio data through algorithms such as noise reduction, echo cancellation, and background sound separation. This step is implemented by the voiceprint preprocessing module.
[0127] (3) The preprocessed audio data is subjected to main voice extraction and voiceprint calculation to obtain the user's real-time voiceprint information. This step is implemented by the voiceprint extraction module.
[0128] (4) After obtaining the voiceprint information, calculate the matching degree between the real-time voiceprint and the standard voiceprint in the database by matching the standard voiceprint information of the user stored in the database. This step is implemented by the voiceprint matching module.
[0129] (5) Compare the matching degree value with a preset threshold. If it is greater than the threshold, the user identity is considered reliable; otherwise, a warning is sent to the recipient. This step is implemented by the voiceprint warning module.
[0130] (6) The recipient ends the session or conducts further identity verification in a timely manner based on the warning information.
[0131] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.
[0132] The above-described embodiments merely represent several implementation manners of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the appended claims.
Claims
1. A system for enhancing conversation security based on voiceprint recognition, characterized in that: include: The device layer includes microphones and sound cards, wearable devices, noise reduction devices, and isolation devices, which are used to collect audio information. The microphone is used to capture audio analog signals, and the sound card converts audio analog signals into digital signals. Wearable devices are portable and can monitor and record the user's voice in real time. Noise reduction equipment is used to reduce background noise in the environment and improve the quality of audio signals; isolation equipment provides sound isolation to ensure the purity of recordings; The storage layer is deployed on both the mobile terminal and the server and is used to store information, including personnel information, session information, alarm information, voiceprint data, configuration information and log information. Personnel information is used to identify and verify the user's identity, and its parameters include name, ID number and voiceprint feature vector; session information is used to record the interaction process between the user and the system and is used to audit and analyze user behavior. Its parameters include session start time, end time, participants, session duration and session content summary; alarm information is used to generate alarms when abnormal behavior or security threats are detected. Its parameters include alarm time, alarm type, alarm level and alarm description; configuration information is used to customize system behavior. Its parameters include system parameters and user preference settings; voiceprint data is the core of voiceprint recognition and is used to verify the user's identity and trigger security warnings. Its parameters include sampling rate, bit depth, frame length and frame shift; log information records system operations and events for troubleshooting and security monitoring. Its parameters include log time, operation type, operation result, operation user and affected resources; The processing layer includes a voiceprint collection module, a voiceprint storage module, a voiceprint preprocessing module, a voiceprint extraction module, a voiceprint matching module and a voiceprint warning module, wherein the voiceprint collection module is used for free speech collection and emotional speech collection; the voiceprint storage module is used to store user audio digital signals, audio signals after voiceprint preprocessing, fixed-length feature vectors after voiceprint extraction processing, preset standard user audio feature vectors, convolutional neural network models and deep neural network models; the voiceprint preprocessing module is used for endpoint detection and speech enhancement; the voiceprint extraction module adopts Mel-frequency cepstral coefficients and convolutional neural network algorithms to learn local and global features of audio data; the voiceprint matching module adopts dynamic time warping and deep neural network algorithms to determine whether two sound samples come from the same speaker; the voiceprint warning module adopts methods based on semantic analysis and acoustic features to determine whether there is a warning situation; The display layer provides the UI page of the APP, including user login, personnel list, message conversation, search filtering and permission management.
2. The system for enhancing conversation security based on voiceprint recognition according to claim 1, characterized in that: The voiceprint acquisition module captures the user's voice signal through a high-precision audio acquisition device and converts it into a digital format. It samples and quantizes the analog signal according to the set sampling rate and bit depth to ensure that the collected audio data is clear and accurate, and can effectively reflect the speaker's unique voiceprint characteristics.
3. The system for enhancing conversation security based on voiceprint recognition according to claim 1, characterized in that: The voiceprint preprocessing module performs voice endpoint detection on the audio digital signal obtained by the voiceprint acquisition module to form an audio signal with the voice activity area marked, which is represented as a binary mark sequence, indicating whether there is voice activity in each time period; the audio signal with the voice activity area marked is subjected to voice enhancement to form an enhanced voice signal. The combination of voice endpoint detection and voice enhancement is achieved through the following three methods: 1) Sequential processing: First, speech endpoint detection is performed to determine the valid parts of the speech, and then speech enhancement techniques are applied to these parts; 2) Iterative optimization: Speech endpoint detection and speech enhancement are performed iteratively, that is, the preliminary speech endpoint detection results are used to guide the preliminary speech enhancement, and the enhanced speech signal is fed back to the speech endpoint detection for more accurate speech activity detection; 3) Dynamic parameter adjustment: According to the characteristics of the real-time signal and the changes in the noise environment, the threshold of the speech endpoint detection and the parameters of the speech enhancement algorithm are dynamically adjusted to achieve the best preprocessing effect.
4. The system for enhancing conversation security based on voiceprint recognition according to claim 1, characterized in that: Voiceprint extraction module, the processing steps are as follows: 1) Divide the audio signal obtained by the voiceprint preprocessing module into multiple time windows to obtain windowed audio data, where each window contains audio data of several tens of milliseconds; 2) Perform FFT on each windowed audio data to convert the time domain signal into a frequency domain signal; 3) Apply the Mel filter bank to filter the FFT result, map the frequency domain signal to the Mel frequency scale, simulate the human ear's perception of different frequencies, and obtain the Mel spectrum data; 4) Perform logarithmic operation on the Mel spectrum data and calculate the energy of each filter to obtain the logarithmic energy spectrum; 5) Apply DCT to convert the logarithmic energy spectrum into MFCC coefficients, which capture the pitch, timbre, and dynamic change characteristics of the audio signal; 6) Extract the first N coefficients from the MFCC coefficients to form the MFCC feature vector; 7) Stack the MFCC feature vectors of several consecutive frames together to form a three-dimensional data structure, in which the first dimension is the time series, the second dimension is the frequency dimension, and the third dimension is the MFCC coefficient; 8) Construct a CNN model, which includes multiple convolutional layers, pooling layers and fully connected layers. The convolutional layer is used to extract local features and extract three-dimensional data structures to obtain multiple two-dimensional feature maps; the pooling layer is used to reduce the spatial dimension of the features and convert the feature maps output by the convolutional layer into feature maps after dimensionality reduction; the fully connected layer is used to integrate the features output by the pooling layer and perform the final classification or regression, and finally obtain a fixed-length feature vector for voiceprint matching processing, where The mathematical expression of the CNN model is as follows: in, "X [l-1] " represents the activation output of the previous layer, that is, the three-dimensional data structure of the MFCC feature vector; "W [l] " represents the convolution kernel of the lth layer; "b [l] ” represents the bias of the lth layer; "*" indicates convolution operation; "W [l] *X [l-1] +b [l] " represents the convolution operation of the lth layer, which applies the convolution kernel to the input data and then adds the bias to obtain the original output of the convolution layer, that is, the two-dimensional feature map; "ReLU" represents the activation function, which is used to increase the nonlinear characteristics of the network, enabling the network to learn more complex features and convert the output of the convolutional layer into a non-negative value; "MaxPool" represents the maximum pooling operation, which converts the feature map output by the convolutional layer into a feature map after dimensionality reduction; "W [L] ” represents the connection weight of the fully connected layer; Represents the output of the CNN model, which is a fixed-length feature vector used for comparison and recognition.
5. The system for enhancing conversation security based on voiceprint recognition according to claim 1, characterized in that: Voiceprint matching processing module, the processing steps are as follows: 1) Use DTW to align the feature sequences of the two sound samples, the feature vector obtained by voiceprint extraction and the preset standard user audio feature vector, to resolve the time offset and change between them; 2) Taking the aligned feature sequence as input, the voiceprint feature representation is extracted through the pre-trained DNN model; 3) Use the voiceprint feature representation extracted by DNN to calculate the similarity between speakers, use the Euclidean distance and cosine similarity methods to measure the distance or similarity between features, and obtain the similarity score; 4) Use the similarity value output from step 3 to determine whether the two sound samples come from the same speaker and obtain the matching result. The mathematical expression is as follows: in, "X', Y'" represent the feature vector obtained by voiceprint extraction and the preset standard user audio feature vector; "D" represents a deep learning model, which is used to extract higher-level feature representations from feature vectors; "D(X'), D(Y')" represents the voiceprint feature representation obtained after two feature vectors are extracted (DNN); "T" represents a threshold value, which is used for voiceprint similarity matching judgment; "||……||" represents the Euclidean norm of a vector; "·" represents the dot product of vectors.
6. The system for enhancing conversation security based on voiceprint recognition according to claim 1, characterized in that: Voiceprint warning processing module, the processing steps are as follows: 1) The semantic analysis-based method uses automatic speech recognition technology to convert audio signals into text, analyze the semantic meaning of the speech content, understand the actual meaning and contextual information in the sentence, and identify possible threatening language, inappropriate remarks or keywords for use in the early warning system to trigger a security response; 2) Acoustic feature-based methods are responsible for analyzing the acoustic properties of speech signals, including pitch, volume, speaking rate, rhythm, etc., to detect abnormal speech patterns or behaviors; 3) Evaluate speech information from different dimensions through weighted fusion methods, understand the actual content of the speech, and analyze the speaker's voice characteristics, so as to more comprehensively determine whether there is a warning situation.
7. The system for enhancing conversation security based on voiceprint recognition according to claim 1, characterized in that: The storage layer, processing layer, and display layer are all software parts, deployed on servers and terminal devices. The device layer is the hardware part, which is selectively connected to servers and terminals according to scenario requirements.
8. A method for enhancing conversation security based on voiceprint recognition, characterized in that: The system for enhancing session security based on voiceprint recognition according to any one of claims 1 to 7 is used to implement session security enhancement based on voiceprint recognition, and the specific steps are as follows: (1) Collecting audio data through the terminal microphone, collecting voiceprint information, associating the login user information and storing it. This step is implemented by the voiceprint collection and voiceprint storage module; (2) In a chat session, the user communicates through voice files or real-time voice, obtains audio information in real time, and pre-processes the audio data. This step is implemented by the voiceprint pre-processing module; (3) The pre-processed audio data is extracted and the voiceprint is calculated to obtain the user's real-time voiceprint information. This step is implemented by the voiceprint extraction module; (4) After obtaining the voiceprint information, the matching degree between the real-time voiceprint and the standard voiceprint is calculated by matching the user's standard voiceprint information stored in the database. This step is implemented by the voiceprint matching module; (5) Compare the matching value with a preset threshold. If the matching value is greater than the threshold, the user identity is considered reliable. Otherwise, an early warning is issued to the recipient. This step is implemented by the voiceprint early warning module. The voiceprint early warning module also combines the semantic analysis-based and acoustic feature-based methods to identify the language and audio indicators that need to be warned and output a warning signal. (6) The recipient promptly ends the session or performs further identity confirmation based on the warning information.
Citation Information
Cited By
Interview data encryption protection method and system and interview method for examination tutoring
CN120710804A
Anti-fraud method and device based on voiceprint recognition, equipment and medium
CN121053998A