Television terminal multi-mode identity authentication voice verification code generation method and system
By using multimodal identity authentication technology, combined with voiceprint and spatial location features to generate personalized voice verification codes, the problems of recording replay attacks and environmental adaptability in TV-based identity authentication are solved, thereby improving security and recognition accuracy.
Patent Information
- Application Number
- CN202511477090.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-16
- Publication Date
- 2025-11-11
- Estimated Expiration
- 2045-10-16
AI Technical Summary
Existing TV-based identity authentication technologies suffer from several problems: voice verification codes are vulnerable to recording and playback attacks; voice clarity is insufficient in living room environments; recognition accuracy is low; and they cannot effectively distinguish between authorized and unauthorized users.
By integrating multimodal features such as user voiceprint and spatial location, a personalized voice verification code is generated. This includes acoustic compensation processing, voiceprint feature extraction, spatial positioning, multimodal identity authentication, personalized speech synthesis, and spectral coding, forming a user-exclusive voice verification code that is bound to the user's identity feature code.
It improves the security and accuracy of identity authentication on TV, prevents recording replay attacks, adapts to the living room environment, enables personalized verification code generation, and enhances user experience.
Smart Images

Figure CN120935408A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of speech recognition technology, and in particular to a method and system for generating voice verification codes for multimodal identity authentication on a television. Background Technology
[0002] With the rapid development of smart TVs and home entertainment systems, user authentication on TVs has become a crucial technological requirement for protecting family privacy and providing personalized services. Existing TV authentication technologies primarily rely on traditional single-modal authentication methods such as remote control button input, simple voice recognition, or basic facial recognition. Regarding voice verification code technology, current solutions typically use standardized text-to-speech synthesis technology to generate uniformly formatted voice verification codes. All users receive verification codes that are essentially identical in terms of voice characteristics, timbre, and playback parameters, differing only in the content of the verification code.
[0003] However, existing technologies have significant shortcomings in security and applicability. First, the uniform format of voice verification codes is vulnerable to replay attacks, allowing attackers to bypass authentication by recording and replaying the voice verification code. Second, existing voice verification codes do not consider the acoustic characteristics of the living room environment on a TV, resulting in insufficient voice clarity and low recognition accuracy in noisy environments at a distance. Third, single-content voice verification cannot effectively distinguish between authorized and unauthorized users, lacking a personalized feature binding mechanism for user identity. Summary of the Invention
[0004] This application provides a method and system for generating voice verification codes for multimodal identity authentication on a TV, which solves the problems of existing voice verification codes being monotonous, vulnerable to attacks, and unable to be personalized based on the user's multimodal identity features. By integrating multimodal features such as user voiceprint and spatial location, it realizes the dynamic generation of personalized voice verification codes, thereby improving the security and user experience of TV-based identity authentication.
[0005] In a first aspect, this application provides a method for generating a voice verification code for multimodal identity authentication on a television terminal, the method comprising: Step S1: Perform acoustic compensation processing on the acoustic signals of the living room environment collected by the TV to obtain acoustic correction parameters; Step S2: Extract voiceprint features from the user's voice to obtain the user's voiceprint template, and simultaneously perform spatial localization processing on the user's voice location to obtain the spatial location feature vector. Step S3: Perform multimodal identity authentication processing on the user voiceprint template and the spatial location feature vector to obtain the user identity feature code; Step S4: Perform personalized speech synthesis processing on the verification code sequence based on the user identity feature code and the acoustic correction parameters to obtain the user-exclusive voice verification code audio; Step S5: Perform spectrum encoding on the user-specific voice verification code audio to obtain a multimodal voice verification code for the TV. Step S6: Associate and bind the TV-side multimodal voice verification code with the user identity feature code to obtain the TV-side identity authentication credential.
[0006] Secondly, this application provides a voice verification code generation system for multimodal identity authentication on a television terminal, the voice verification code generation system for multimodal identity authentication on a television terminal includes: The compensation module is used to perform acoustic compensation processing on the acoustic signals of the living room environment collected by the TV to obtain acoustic correction parameters; The extraction module is used to extract voiceprint features from the user's voice to obtain the user's voiceprint template, and at the same time, to perform spatial positioning processing on the user's voice position to obtain the spatial position feature vector. The authentication module is used to perform multimodal identity authentication processing on the user voiceprint template and the spatial location feature vector to obtain the user identity feature code; The synthesis module is used to perform personalized speech synthesis processing on the verification code sequence based on the user identity feature code and the acoustic correction parameters to obtain user-exclusive voice verification code audio. The encoding module is used to perform spectrum encoding processing on the user-specific voice verification code audio to obtain a multimodal voice verification code for the TV. The binding module is used to associate and bind the TV terminal multimodal voice verification code with the user identity feature code to obtain the TV terminal identity authentication credential.
[0007] Thirdly, a voice verification code generation device for multimodal identity authentication on a television is provided, comprising: a memory and at least one processor, wherein the memory stores instructions; the at least one processor invokes the instructions in the memory to cause the voice verification code generation device for multimodal identity authentication on a television to execute the aforementioned voice verification code generation method for multimodal identity authentication on a television.
[0008] Fourthly, a computer-readable storage medium is provided, wherein instructions are stored therein, which, when executed on a computer, cause the computer to perform the aforementioned method for generating voice verification codes for multimodal identity authentication on a television terminal.
[0009] The technical solution provided in this application obtains acoustic correction parameters by performing acoustic compensation processing on the acoustic signals of the living room environment collected from the TV terminal. This effectively solves the problem of voice verification code distortion in the living room environment in existing technologies, ensuring that the voice verification code can adaptively adjust according to different living room acoustic environments. Simultaneously, voiceprint feature extraction processing is performed on the user's voice to obtain a user voiceprint template, and spatial positioning processing is performed on the user's voice location to obtain a spatial location feature vector. This achieves multi-dimensional feature capture of user identity, significantly enhancing the accuracy and anti-counterfeiting capabilities of identity authentication compared to existing single-modal authentication methods that rely solely on voice content verification. The technical features of obtaining the user identity feature code through multi-modal identity authentication processing of the user voiceprint template and spatial location feature vector, by fusing acoustic and spatial features, achieve comprehensive user identity discrimination, avoiding the security risks of easy forgery of single features in existing technologies. Personalized speech synthesis processing is performed on the verification code sequence based on the user's identity feature code and acoustic school positive parameters to obtain the user's exclusive voice verification code audio. This breaks through the technical limitation of the existing technology that all users use the same format voice verification code, realizes personalized verification code generation based on user identity features, and fundamentally solves the security threat of recording replay attacks.
[0010] The technical features of this application include: spectral encoding of user-specific voice verification code audio to obtain a multimodal voice verification code for television; and associating and binding the multimodal voice verification code for television with the user's identity feature code to obtain a television identity authentication credential. This achieves end-to-end personalized processing of the voice verification code from content generation to final authentication. Particularly in the specific application area of voice verification code generation for multimodal identity authentication on television, the radial basis function acoustic compensation algorithm used in this application can accurately model the acoustic propagation characteristics of the living room environment; the 8-direction chain code quantization algorithm effectively extracts the personalized feature patterns of the user's voiceprint; and the attention weight algorithm realizes the dynamic weight allocation of voiceprint features and spatial location features. The synergistic effect of these algorithmic features ensures that the generated voice verification code not only possesses a unique identifier for the user but is also adapted to the television usage environment. Compared with existing standardized speech synthesis technologies, it significantly improves in terms of anti-attack capability, personalization, and environmental adaptability. Attached Figure Description
[0011] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0012] Figure 1This is a schematic diagram of an embodiment of the voice verification code generation method for multimodal identity authentication on a television terminal in this application. Figure 2 This is a schematic diagram of an embodiment of the voice verification code generation system for multimodal identity authentication on a television terminal in this application. Figure 3 This is a schematic block diagram of the structure of the voice verification code generation device for multimodal identity authentication on the TV side in an embodiment of the present invention. Detailed Implementation
[0013] This application provides a method and system for generating voice verification codes for multimodal identity authentication on a television. The terms "first," "second," "third," "fourth," etc. (if present) in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments described herein can be implemented in a sequence other than that illustrated or described herein. Furthermore, the terms "comprising" or "having" and any variations thereof are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or device that includes a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or devices.
[0014] For ease of understanding, the specific process of the embodiments of this application is described below. Please refer to [link / reference]. Figure 1 One embodiment of the voice verification code generation method for multimodal identity authentication on a television terminal in this application includes: Step S1: Perform acoustic compensation processing on the acoustic signals of the living room environment collected by the TV to obtain acoustic correction parameters; Step S2: Extract voiceprint features from the user's voice to obtain the user's voiceprint template, and simultaneously perform spatial localization processing on the user's voice location to obtain the spatial location feature vector. Step S3: Perform multimodal identity authentication processing on the user's voiceprint template and spatial location feature vector to obtain the user's identity feature code; Step S4: Perform personalized speech synthesis processing on the verification code sequence based on the user's identity feature code and acoustic correction parameters to obtain the user's exclusive voice verification code audio; Step S5: Perform spectrum encoding on the user-specific voice verification code audio to obtain the TV-side multimodal voice verification code; Step S6: Associate and bind the TV-side multimodal voice verification code with the user's identity feature code to obtain the TV-side identity authentication credential.
[0015] It is understood that the executing entity of this application can be a voice verification code generation system for multimodal identity authentication on a TV, or it can be a terminal or a server; the specific implementation is not limited here. This application's embodiment uses a server as an example for illustration.
[0016] Specifically, the TV's built-in microphone array first collects raw acoustic data of the living room environment, including background noise, echo, and reverberation information. The radial basis function (RBF) is a function with distance as its independent variable. It establishes an acoustic distance model by calculating the Euclidean distance between the TV microphone array position and the user's voice position. Specifically, the acoustic signals detected by the microphone array include a reverberation time parameter, which reflects the reflection attenuation characteristics of sound waves within the living room space, and a sound propagation delay parameter, representing the time difference between the sound wave's arrival at each microphone from the voice source. The RBF calculates the spatial distance between the sound source and the receiving point based on these acoustic characteristic parameters, and then, combined with the frequency response characteristics of the living room, calculates compensation coefficients for different frequency bands. The acoustic correction parameter is obtained by multiplying the compensation coefficients by a preset reference gain value; this parameter is used for environmental adaptation in the subsequent speech synthesis process.
[0017] Simultaneously, voiceprint feature extraction and spatial positioning processing are performed. The fundamental frequency extraction process analyzes the periodic characteristics of the user's speech signal to identify the time-series trajectory of the fundamental frequency. First-order difference calculates the difference in fundamental frequency values between adjacent time points, reflecting the trend of fundamental frequency change; second-order difference calculates the difference in first-order difference, reflecting the acceleration characteristics of fundamental frequency change. Eight-directional chain code quantization maps the direction of fundamental frequency change to eight discrete directions, corresponding to different fundamental frequency change patterns such as rising, falling, and stationary, forming a unique voiceprint template for the user. Spatial positioning processing uses a time difference of arrival algorithm, calculating the time difference of arrival of the same sound source signal at different microphones, and using triangulation principles to determine the user's three-dimensional spatial coordinates. The horizontal angle represents the user's left-right offset angle relative to the front of the television, the vertical angle represents the user's vertical position deviation, and distance layering divides the distance between the user and the television into three levels: near, medium, and far, forming a spatial position feature vector.
[0018] Multimodal feature fusion authentication is achieved. Data standardization converts user voiceprint templates into a standardized format with zero mean and one variance, eliminating dimensional differences in voiceprint data from different users. A fully connected network maps spatial location feature vectors through a multi-layer neural network, performing a non-linear transformation to generate a high-dimensional spatial location code. An attention weight algorithm calculates the importance weight of each feature component, and a softmax function normalizes the weight values to ensure that the sum of all weights is one. The fused feature vector combines the standardized voiceprint features and spatial location code into a unified feature representation through weighted summation. The identity classifier uses a multi-layer perceptron structure to map the fused feature vector into a unique identifier for the user.
[0019] Personalized voice verification codes are generated based on user identity feature codes. Personalized parameter mapping uses the user identity feature code as a seed, and a hash function is used to generate a verification code sequence bound to that user, ensuring different users receive different verification code content. The speech synthesis network adopts an adversarial network architecture of encoder-generator-discriminator. The encoder converts the verification code text into a text feature vector, and the compression encoding mechanism uses the Lempel-Ziv compression algorithm to ensure the uniqueness and unpredictability of the generated verification codes. The generator network uses a deconvolutional layer structure to progressively upsample the text feature vector into a speech spectrum representation. The discriminator network judges the authenticity of the generated speech through multi-scale analysis. Living room environment adaptation processing combines the initial voice verification code with acoustic correction parameters to adjust the volume, pitch, and spectral characteristics of the speech to adapt to the acoustic environment of a living room. Audio fingerprint embedding embeds the user identity feature code as a watermark into the high-frequency components of the speech signal, forming a user-specific voice verification code audio.
[0020] The time-domain to frequency-domain transformation uses Fast Fourier Transform (FFT) to convert the speech time-domain signal into a frequency-domain representation, obtaining the spectral distribution information of the voice verification code. Frequency-domain adaptive filtering divides the spectrum into 12 key frequency bands, each corresponding to the main frequency range of human voice. The amplitude and phase characteristics of each frequency band are extracted using adaptive filters. Spectral envelope modeling calculates the envelope curve for each frequency band, reflecting the energy distribution pattern of the speech signal in that band. A spectral coding matrix organizes the envelope parameters of the 12 frequency bands into a matrix form, forming the digital representation of the multimodal voice verification code on the television terminal.
[0021] The association and binding process combines the spectral encoding matrix of the multimodal voice verification code on the TV with the user's identity feature code using an encrypted hash algorithm to generate a composite credential containing user identity information and verification code content. This credential simultaneously contains the spectral characteristics of the voice verification code and the user's multimodal identity characteristics, enabling personalized voice verification code generation based on individual user characteristics.
[0022] In one specific embodiment, step S1 includes: Acoustic signals from the living room environment are collected and processed using the TV's built-in microphone array to obtain raw environmental acoustic data. The reverberation time and sound propagation delay were extracted from the raw environmental acoustic data to obtain the acoustic characteristic parameters of the living room. The acoustic distance of the sound source is calculated by performing acoustic distance calculation on the acoustic characteristic parameters of the living room based on the radial basis function. The acoustic compensation coefficient is obtained by calculating the compensation coefficient of the frequency response function based on the spatial distance value of the sound source. The acoustic correction parameters are obtained by multiplying the acoustic compensation coefficients with the reference gain parameters.
[0023] Specifically, the acoustic signal acquisition and processing of the TV's built-in microphone array involves simultaneously recording sound wave signals from the living room environment using multiple microphones. Each microphone converts the mechanical vibrations of the sound waves into voltage signals, which are then converted into digital sampling sequences via an analog-to-digital converter. The microphone array typically uses an L-shaped or circular configuration, containing 4 to 8 independent omnidirectional microphone units, with the microphones spaced at a fixed distance of centimeters. The raw environmental acoustic data includes mixed signals from all sound sources in the living room, including user voices, TV playback sound, air conditioner noise, footsteps, and other acoustic components. During data acquisition, each microphone independently records the time series of sound pressure changes, forming a multi-channel acoustic data matrix. The sampling frequency is typically set to 16kHz or 48kHz to cover the main frequency range of human voices. Reverberation time and sound propagation delay extraction and processing analyze two key characteristic parameters of the living room acoustic environment. Reverberation time represents the time required for sound to attenuate to one-thousandth of its original intensity after reflection within the living room. It is calculated by analyzing the attenuation envelope curve of the sound signal, specifically by measuring the duration of the residual echo in the room after the sound stops. Sound propagation delay is obtained by calculating the time difference of sound waves traveling from the sound source to different microphone positions. Utilizing the physical constant that the speed of sound in air is approximately 343 meters per second, and combining this with the time difference of each microphone receiving the same sound signal, the distance difference between the sound source and each microphone is calculated. Living room acoustic characteristic parameters integrate reverberation time and sound propagation delay into a parameter vector reflecting the acoustic environment characteristics of the living room. This vector describes the propagation and attenuation patterns of sound waves within a specific space.
[0024] The radial basis function (RBF) acoustic distance calculation process uses a nonlinear function with distance as the variable to establish a mapping relationship between the sound source location and acoustic characteristics. The core principle of the RBF is to calculate the Euclidean distance between the input point and a preset center point, and then substitute this distance value into an exponential decay function for transformation. The acoustic distance calculation first determines the geometric center of the television microphone array as the reference origin, and then calculates the three-dimensional coordinates of the user's vocal position using triangulation principles based on the time difference of the sound signals received by each microphone. Specifically, by comparing the arrival times of the same sound pulse received by different microphones and combining this with the known distances between the microphones, a system of equations regarding the sound source location is established. The spatial distance value of the sound source represents the straight-line distance between the user's vocal position and the center point of the television; this distance value directly affects the intensity attenuation and spectral characteristics of the sound signal.
[0025] The frequency response function describes the transmission characteristics of sound waves of different frequencies in a living room environment. Low-frequency sound waves attenuate less during spatial propagation but are prone to standing waves, while high-frequency sound waves attenuate faster but are more directional. The compensation coefficient calculation establishes a compensation relationship between distance and frequency response by analyzing the impact of sound source distance on the signal strength of each frequency band. During the calculation, the audible frequency range is divided into multiple sub-bands, and each band is assigned a corresponding compensation weight based on its attenuation characteristics at different distances. The acoustic compensation coefficient reflects the amount of gain adjustment required for each frequency band under the current sound source distance conditions; the farther the sound source, the greater the compensation gain required in the high-frequency range.
[0026] The product operation of acoustic compensation coefficients and reference gain parameters combines frequency-dependent compensation coefficients with the device's inherent reference gain value. The reference gain parameter is the standard amplification factor of a television sound system, reflecting the device's signal amplification capability under ideal conditions. The product operation multiplies each frequency band component of the compensation coefficient with the corresponding reference gain value to obtain the final correction parameter for the current acoustic environment. The acoustic correction parameter incorporates a comprehensive adjustment based on living room environmental characteristics, sound source location information, and device characteristics; this parameter is directly used for audio modulation in the subsequent voice verification code generation process.
[0027] In one specific embodiment, step S2 includes: The user's spoken voice is processed by fundamental frequency extraction to obtain a fundamental frequency trajectory sequence; First-order and second-order difference calculations are performed based on the fundamental frequency trajectory sequence to obtain fundamental frequency change direction data; The fundamental frequency variation direction data is quantized in 8 directions to obtain the user voiceprint template. The time difference of arrival of the user's voice is calculated by using a TV microphone array to obtain three-dimensional spatial coordinate data. Based on the three-dimensional spatial coordinate data, horizontal angle, vertical angle and distance are layered processing to obtain spatial location feature vectors.
[0028] Specifically, the user's voice fundamental frequency extraction process identifies the basic frequency of vocal cord vibration by analyzing the periodic characteristics of the speech signal. The fundamental frequency represents the number of vibrations per second of the vocal cords and is a crucial feature parameter for distinguishing different speakers. The fundamental frequency extraction process first preprocesses the user's speech signal, including pre-emphasis filtering and windowing / framing, dividing the continuous speech signal into several overlapping short frames, each typically 20 to 30 milliseconds long. Then, the periodicity of each frame is detected using the autocorrelation function method. The fundamental frequency period is determined by finding the first peak position of the autocorrelation function, and the reciprocal of the period is converted into the fundamental frequency value. The fundamental frequency trajectory sequence records the trajectory of the fundamental frequency change over time throughout the user's vocalization process, forming time-series data reflecting the user's vocal cord vibration characteristics. The fundamental frequency trajectory sequence not only contains the absolute value of the fundamental frequency but also retains the temporal information of the fundamental frequency changes, laying a data foundation for subsequent voiceprint feature analysis.
[0029] The fundamental frequency trajectory sequence is processed using first-order and second-order difference calculations to extract the velocity and acceleration information of fundamental frequency changes. First-order difference calculation subtracts the fundamental frequency values at adjacent time points in the fundamental frequency trajectory sequence to obtain the change in fundamental frequency within each time interval, reflecting the trend and magnitude of fundamental frequency rise or fall. Second-order difference calculation performs adjacent difference operations again on the first-order difference result to obtain the rate of change of fundamental frequency velocity, reflecting the acceleration characteristics of fundamental frequency changes. Fundamental frequency change direction data is obtained by performing sign determination and amplitude analysis on the first-order and second-order difference results, converting continuous fundamental frequency changes into discrete directional indicators. Positive values indicate a rising fundamental frequency, negative values indicate a falling fundamental frequency, and zero values indicate a stable fundamental frequency; the magnitude of the change reflects the severity of the fundamental frequency change. The fundamental frequency change direction data captures the fluctuation pattern of the user's tone of voice, forming a personalized fundamental frequency change fingerprint.
[0030] The 8-directional chain code quantization process maps continuous fundamental frequency variation information into coded sequences of eight discrete directions. The 8-directional chain code employs a coding method similar to contour tracking in image processing, mapping the combination of fundamental frequency variation direction and amplitude into eight basic directions: sharp rise, slow rise, stable, slow fall, sharp fall, fluctuating rise, fluctuating fall, and complex variation. Quantization processing sets thresholds based on the numerical ranges of the first and second-order differences, classifying fundamental frequency variations into corresponding directional categories. The user voiceprint template statistically analyzes the 8-directional chain code distribution characteristics of each user under different vocal content, forming a template vector reflecting the individual acoustic characteristics of the user. The voiceprint template not only records the statistical distribution of the user's fundamental frequency variation but also retains the temporal correlation information of the variation patterns, constituting the user's identity identifier in voice interaction on the television.
[0031] The TV microphone array's user voice location time difference calculation process utilizes the physical characteristics of sound wave propagation to locate the user's spatial position. The time difference of arrival algorithm is based on the time difference of the same sound signal arriving at different microphones, combined with the speed of sound in air, to calculate the spatial coordinates of the sound source. During the calculation, firstly, characteristic pulses or speech start points in the speech signals received by each microphone are identified, and then the precise timestamps of these characteristic points arriving at each microphone are calculated. By comparing the time differences between different microphone pairs, an overdetermined system of equations regarding the user's position is established, and the least squares method is used to solve for the user's three-dimensional spatial coordinates. The three-dimensional spatial coordinate data is used to establish a Cartesian coordinate system with the center of the TV screen as the origin, recording the user's front-to-back distance, left-to-right offset, and vertical position relative to the TV.
[0032] Layered processing of 3D spatial coordinate data (horizontal angle, vertical angle, and distance) transforms continuous spatial location information into discrete feature representations. The horizontal angle is calculated by determining the user's left-right deflection relative to the front of the TV, using the arctangent of the user's x and z coordinates. The vertical angle is calculated by determining the user's up-down deflection relative to the horizontal centerline of the TV, using the user's y and z coordinates. Layered distance processing divides the straight-line distance between the user and the TV into three levels: near, medium, and far, each corresponding to different voice interaction scenarios and acoustic characteristics. The spatial location feature vector encodes the horizontal, vertical, and distance levels into fixed-dimensional numerical vectors, forming a standardized feature representation describing the user's spatial location.
[0033] In one specific embodiment, step S3 includes: The user's voiceprint template is subjected to data standardization processing to obtain standardized voiceprint features; Spatial location feature vectors are mapped using a fully connected network to obtain spatial location codes; The standardized voiceprint features and spatial location encoding are weighted using an attention weighting algorithm to obtain a fused feature vector. The fused feature vectors are processed by an identity classifier to obtain the user's identity feature code.
[0034] Specifically, the standardization process for user voiceprint template data eliminates the differences in dimensions and numerical ranges between different users' voiceprint data. Data standardization is a preprocessing technique in machine learning, which transforms raw data into a standard format with uniform distribution characteristics through mathematical transformations. Voiceprint templates contain frequency statistics of 8-directional chain codes and amplitude information of fundamental frequency variations. Voiceprint data from different users exhibit significant differences in numerical range and distribution characteristics. The standardization process first calculates the mean and standard deviation of each feature dimension in the voiceprint template, then subtracts the mean from each feature value and divides by the standard deviation, resulting in standardized data with a mean of zero and a variance of one. Standardized voiceprint features eliminate the influence of individual differences on numerical magnitude, highlighting the relative variation characteristics of voiceprint patterns, and ensuring that voiceprint data from different users have equal weight and a basis for comparison. The standardization process maintains the relative relationships of the original voiceprint features, ensuring that the uniqueness of individual user acoustic characteristics is preserved.
[0035] Spatial location feature vector mapping using fully connected networks transforms discrete spatial location information into high-dimensional continuous feature representations. Fully connected networks are a fundamental structure of neural networks, consisting of multiple fully connected layers, each containing several neurons, with all nodes in adjacent layers fully connected. Spatial location feature vectors contain discrete codes for horizontal angles, vertical angles, and distances. Fully connected networks map these discrete features into continuous high-dimensional vectors through multiple layers of nonlinear transformations. In the mapping process, the first layer of the fully connected network receives the spatial location feature vector as input, performs linear transformations through weight matrix multiplication and bias addition, and then applies nonlinear activation functions such as ReLU or Sigmoid for nonlinear mapping. The cascading of multiple fully connected networks allows spatial location encoding to capture complex nonlinear relationships between location features, forming richer spatial representations. Spatial location encoding not only preserves the original location information but also acquires deep semantic representations of location features through network learning.
[0036] The attention weighting algorithm addresses the importance balance problem in multimodal feature fusion by standardizing the weights of voiceprint features and spatial location codes. Attention mechanisms, a technique in deep learning for dynamically allocating feature weights, adjust the importance weights of features by calculating their contribution to the final result. The weight allocation process first calculates the attention scores of standardized voiceprint features and spatial location codes by performing a dot product operation between the feature vector and the learnable query vector. Then, a softmax function is applied to normalize the original scores, ensuring that the sum of the weights of all features equals one, resulting in standardized attention weights. The weight allocation dynamically adjusts the importance ratio of voiceprint features and spatial location features based on different users and usage scenarios. Spatial location features receive higher weights when the user is far away, and voiceprint features receive higher weights when there is significant environmental noise. The fused feature vector combines the standardized voiceprint features and spatial location codes according to the attention weights through a weighted summation, forming a unified representation that comprehensively reflects the user's multimodal identity features.
[0037] The fusion feature vector identity classifier transforms multimodal fusion features into a unique identifier for each user. The identity classifier employs a multilayer perceptron architecture, a fully connected network architecture comprising an input layer, hidden layers, and an output layer. The input layer receives the fusion feature vector, the hidden layers extract a high-level abstract representation of the features through multiple nonlinear transformations, and the output layer generates a probability distribution of user identity categories. In the recognition process, the fusion feature vector is first linearly transformed through the weight matrix of the hidden layers, then a nonlinear activation function is applied for feature transformation. The cascading of multiple hidden layers enables the classifier to learn complex identity discrimination patterns. The output layer uses a softmax activation function to convert the output of the last hidden layer into a probability distribution of user identity categories, with the category with the highest probability corresponding to the user's identity identifier. The user identity feature code is generated by encoding the user category corresponding to the highest probability with its confidence level, forming a composite identifier code containing user identity information and authentication credibility.
[0038] In one specific embodiment, step S4 includes: Personalized parameter mapping is performed on the random verification code sequence based on the user's identity feature code to obtain a user-specific verification code sequence; The user-specific verification code sequence is input into a speech synthesis network for text-to-speech processing to obtain the initial voice verification code audio. The initial voice verification code audio is adapted to the living room environment based on the acoustic positive parameter to obtain the acoustically compensated voice verification code audio. The acoustically compensated voice verification code audio is embedded with the user's identity feature code to obtain the user's unique voice verification code audio.
[0039] Specifically, the personalized parameter mapping process for the random CAPTCHA sequence using user identity feature codes converts user identity information into unique CAPTCHA generation parameters through a hash algorithm. Personalized parameter mapping is a key technology ensuring different users receive different CAPTCHA content. By inputting the user identity feature code as a seed value into a pseudo-random number generator, a CAPTCHA sequence bound to that user is generated. The mapping process first hashes the user identity feature code to obtain a fixed-length numerical sequence, which is then used as the initial seed for the random number generator. This ensures that the CAPTCHA sequences generated by the same user are consistent each time, while the CAPTCHA sequences generated by different users are completely different. The random CAPTCHA sequence typically contains a combination of numbers and letters, with a length of 4 to 6 characters. The mapping algorithm determines the specific content and character combination pattern of the CAPTCHA based on the hash value of the user identity feature code. The user-specific CAPTCHA sequence not only contains the character content of the CAPTCHA but also includes a generation timestamp and sequence number associated with the user's identity, forming a CAPTCHA data structure with user uniqueness and timeliness.
[0040] The user-specific CAPTCHA sequence speech synthesis network uses an encoder-generator-discriminator adversarial network architecture to convert text CAPTCHAs into speech signals. The speech synthesis network is a deep learning-based text-to-speech model comprising three main modules: text encoding, acoustic modeling, and speech generation. The text-to-speech process first inputs the CAPTCHA character sequence into a text encoder. The encoder uses a recurrent neural network or Transformer structure to convert discrete text symbols into continuous text feature vectors, with each character corresponding to a high-dimensional semantic representation vector. The acoustic modeling module converts the text feature vectors into acoustic parameters for speech, including fundamental frequency, spectral envelope, and duration information. It learns the mapping relationship between text and acoustic features through multi-layer neural networks. The speech generation module uses a vocoder to convert the acoustic parameters into the final speech waveform. The vocoder reconstructs the time-domain audio signal from the frequency-domain acoustic parameters through an inverse transform. The initial CAPTCHA audio contains standard pronunciation features and rhythmic characteristics, possessing clear speech intelligibility and an appropriate speech length.
[0041] The initial voice verification code audio is adapted for the living room environment based on acoustic calibration parameters. This adaptation adjusts the spectrum and amplitude of the voice signal according to the characteristics of the living room's acoustic environment. Living room environment adaptation is a specialized processing technology for TV applications, considering the impact of acoustic characteristics, user distance, and background noise on voice propagation. The adaptation process uses previously obtained acoustic calibration parameters, including compensation coefficients and gain adjustments for different frequency bands, to readjust the spectral distribution of the initial voice verification code. During spectrum adjustment, the low-frequency portion is attenuated or enhanced according to the reverberation characteristics of the living room, the mid-frequency portion is optimized for the main frequency range of human voices, and the high-frequency portion is compensated and amplified based on distance attenuation characteristics. Amplitude adjustment adaptively adjusts the overall volume of the voice verification code based on the distance between the user and the TV and the gain information in the acoustic calibration parameters. After acoustic compensation, the voice verification code audio possesses acoustic characteristics adapted to the living room environment, maintaining clear intelligibility and an appropriate volume level at the user's receiving location.
[0042] Acoustic compensation followed by audio fingerprint embedding of user identity feature codes in voice verification codes hides user identity information as a digital watermark within the speech signal. Audio fingerprint embedding is a digital audio anti-counterfeiting technology that achieves information hiding and verification by embedding imperceptible digital identifiers in the frequency or time domain of the audio signal. The embedding process first converts the user identity feature code into a binary sequence, then selects high-frequency components or speech gaps in the speech signal as embedding carriers. High-frequency embedding methods modulate identity information into high-frequency bands insensitive to the human ear, carrying digital information by fine-tuning the spectral amplitude, with the embedding strength controlled within a range that does not affect speech quality. Time-domain embedding methods embed identity information into subtle changes in the speech waveform, carrying binary data by adjusting the amplitude values of sampling points. The embedding algorithm ensures a tight binding between the user identity feature code and the voice verification code content, forming a composite audio signal that contains both the verification code speech content and the user identity identifier. The user-specific voice verification code audio integrates personalized verification code content, acoustic characteristics adapted to the living room environment, and the user's digital fingerprint, realizing personalized voice verification code generation based on multimodal identity features.
[0043] In one specific embodiment, the process of inputting the user-specific verification code sequence into the speech synthesis network for text-to-speech processing can specifically include the following steps: The user-specific verification code sequence is input into the encoder network for text feature encoding to obtain the verification code text feature vector. The uniqueness constraint processing of the CAPTCHA text feature vector is performed based on the compression coding mechanism to obtain the coding constraint features; The encoded constraint features are input into the generator network for adversarial generative processing to obtain the speech spectrum representation; Based on a discriminator network, multi-scale time-frequency domain discrimination processing is performed on the speech spectrum representation to obtain the authenticity verification results; Based on the authenticity verification results, the speech spectrum representation is subjected to inverse Mel spectrum transform to obtain the initial speech verification code audio. The initial voice verification code audio is processed using the far-field speech characteristics of the TV to obtain the voice verification code audio.
[0044] Specifically, the user-specific CAPTCHA sequence encoder network's text feature encoding process uses a deep neural network to convert discrete CAPTCHA characters into continuous high-dimensional feature representations. The encoder network is the core component of the sequence-to-sequence model, employing a multi-layer recurrent neural network or Transformer structure to process text sequence data. Text feature encoding first converts each character in the CAPTCHA into a one-hot encoded vector, with numbers and letters corresponding to different encoding positions. Then, an embedding layer maps the sparse one-hot vectors into dense word vector representations. The encoder's recurrent layers process the characters in the CAPTCHA sequence one by one. At each time step, it receives the word vector of the current character and the hidden state from the previous time step, updates the hidden state through a gating mechanism, and generates the contextual representation of the current character. The multi-layer encoding structure enables the network to capture complex dependencies and semantic information between characters. Finally, the output layer generates a text feature vector containing the semantic information of the entire CAPTCHA sequence. The CAPTCHA text feature vector not only encodes the literal meaning of the characters but also contains deeper information such as character order, phonetic characteristics, and semantic associations, providing a rich feature foundation for subsequent speech synthesis processing.
[0045] The compression coding mechanism for CAPTCHA text feature vector uniqueness constraints employs data compression principles from information theory to ensure the uniqueness and unpredictability of generated CAPTCHAs. Based on the core idea of the Lempel-Ziv algorithm, the compression coding mechanism achieves efficient compression and uniqueness guarantees by analyzing repetitive patterns and redundant information in the data. The uniqueness constraint processing first performs pattern analysis on the CAPTCHA text feature vectors, identifying possible repetitive or similar patterns. Then, a compression algorithm is applied to encode and deduplicate these patterns. During compression, the algorithm constructs a dynamic dictionary to record existing feature patterns. Newly appearing patterns are assigned new encoding identifiers, while existing codes are referenced for repetitive patterns, thus achieving feature compression and unique identification. After compression, the encoded constraint features retain the core information of the original text features while adding unique identification and anti-duplicate mechanisms, ensuring that the feature vector corresponding to each CAPTCHA is unique within the entire CAPTCHA space.
[0046] The Generative Adversarial Network (GAN) for encoding constraint feature generation converts text features into a speech spectral representation using a GAN framework. The generator network, responsible for data generation, progressively upsamples low-dimensional feature vectors into high-dimensional spectral data through multiple deconvolutional layers. In the GAN process, the generator receives encoded constraint features as input, maps the feature vectors to the initial dimension for spectrum generation through linear transformation layers, and then progressively increases the temporal and frequency resolution of the spectrum through multiple deconvolutional operations. Each deconvolutional layer uses different kernel sizes and strides to control the details of spectrum generation, and the Leaky ReLU activation function is used to maintain gradient flow and nonlinear expressiveness. The generator's output layer uses the Tanh activation function to normalize the spectral values to a standard range, forming a two-dimensional spectral matrix containing information on both the time and frequency axes. The speech spectral representation contains complete frequency domain information of the CAPTCHA speech, with each time frame corresponding to a frequency distribution vector. The entire matrix describes the spectral characteristics of the speech signal over time.
[0047] The discriminator network performs multi-scale time-frequency domain discrimination processing for speech spectral representation, evaluating the authenticity and quality of the generated spectrum through an adversarial training mechanism. The discriminator network, responsible for authenticity judgment in the adversarial generative network, employs a multi-scale convolutional neural network structure to analyze the authenticity features of the spectral data. Multi-scale time-frequency domain discrimination includes two dimensions: time-scale discrimination and frequency-scale discrimination. Time-scale discrimination analyzes the continuity and naturalness of the spectrum in the time dimension using convolutional kernels of different window lengths, while frequency-scale discrimination detects the reasonableness and authenticity of the spectrum in the frequency domain using convolutional kernels of different frequency bandwidths. The discriminator's convolutional layers employ multiple parallel convolutional branches, each specifically processing spectral features at a particular scale. Pooling layers and fully connected layers fuse the multi-scale features into the final discrimination result. The authenticity verification result outputs a probability value between 0 and 1 through a sigmoid activation function, representing the confidence level that the generated spectrum is judged as real speech. This result guides the optimization direction of the generator and the evaluation of spectral quality.
[0048] The authenticity verification result's speech spectrum representation undergoes Inverse Mel-Syllable Transform (IMT) processing, which converts the frequency domain spectral data back into a time-domain audio waveform signal. IMT is a classic algorithm in speech signal processing, restoring the frequency domain representation to time-domain audio through an inverse spectral transformation. The inverse transformation process first performs an inverse mapping of the speech spectrum representation using a Mel filter bank, converting the Mel frequency scale back to a linear frequency scale and restoring the original frequency distribution characteristics of the spectrum. Then, an inverse discrete cosine transform is applied to convert the spectral coefficients into power spectral density, followed by an inverse fast Fourier transform to convert the power spectrum in the frequency domain back into a time-domain audio waveform. Phase reconstruction is required during the inverse transformation process. The phase information corresponding to the spectrum is estimated using the Griffin-Lin algorithm or a neural vocoder to ensure the temporal continuity and auditory quality of the reconstructed audio. The initial speech verification code audio, after inverse transformation, obtains complete time-domain waveform data, including the pronunciation features, pitch variations, and duration information of the verification code characters.
[0049] The initial voice verification code audio for TV use undergoes far-field voice characteristic processing, optimizing the audio signal to meet the specific requirements of TV usage scenarios. This processing includes distance attenuation compensation, directionality enhancement, and background noise suppression techniques, specifically addressing the issue of users listening to voice verification codes from a distance in a living room environment. Far-field voice characteristic processing first analyzes the spectral distribution and dynamic range of the initial audio, then adaptively adjusts it based on the characteristics of the TV sound system and the acoustic environment of the living room. Distance compensation enhances high-frequency components to offset high-frequency attenuation during air propagation; directionality enhancement improves the spatial localization of speech by adjusting stereo balance; and background noise suppression reduces the impact of environmental noise on speech clarity through spectral subtraction or Wiener filtering. After far-field characteristic processing, the voice verification code audio possesses acoustic characteristics suitable for TV playback and long-distance reception by users.
[0050] In one specific embodiment, step S5 includes: The audio of the user-specific voice verification code is transformed from the time domain to the frequency domain to obtain the voice verification code spectrum data; Based on the frequency domain adaptive filtering algorithm, the spectral data of the voice verification code is analyzed and processed in 12 key frequency bands to obtain spectral contour features. The spectral contour features are processed by spectral envelope modeling to obtain the spectral envelope parameters of the voice verification code; The multimodal voice verification code for television is obtained by generating a spectral coding matrix based on the spectral envelope parameters of the voice verification code.
[0051] Specifically, the user-specific voice verification code audio time-domain to frequency-domain transformation process uses Fast Fourier Transform (FFT) to convert the time-series audio waveform into frequency-domain spectral distribution data. Time-domain to frequency-domain transformation is a fundamental technology in digital signal processing, converting the amplitude variation of an audio signal in the time dimension into an energy distribution in the frequency dimension. The transformation process first preprocesses the voice verification code audio, including pre-emphasis filtering and windowing / framing, dividing the continuous audio signal into several overlapping short frames. Each frame is typically set to a length of 25 milliseconds, with a frame shift of 10 milliseconds to ensure a balance between the time and frequency resolutions of the spectral analysis. The FFT algorithm performs frequency-domain transformation on each audio frame, converting the time-domain sampling sequence into complex spectral coefficients. The intensity and phase information of the frequency components are obtained by calculating the amplitude and phase spectra. The voice verification code spectral data is organized in a two-dimensional matrix, with rows corresponding to time frames and columns corresponding to frequency components. Each element in the matrix represents the energy intensity at a specific time and frequency, forming a data structure that reflects the complete spectral characteristics of the voice verification code.
[0052] The frequency domain adaptive filtering algorithm analyzes and processes the spectral data of the voice verification code using 12 key frequency bands. Based on the frequency distribution characteristics of human voice, the spectrum is divided into targeted analysis intervals. Frequency domain adaptive filtering is a spectrum analysis technique specifically designed for speech signals. It extracts speech features from different frequency bands by constructing multiple bandpass filters. The division of the 12 key frequency bands is based on phonetics and acoustic principles, covering the main frequency components of speech, such as the fundamental frequency and its harmonics, formant frequencies, and high-frequency fricatives. Specifically, it includes 12 frequency bands: the fundamental frequency interval, the first formant interval, the second formant interval, the third formant interval, the low-frequency energy interval, the mid-frequency unvoiced interval, the high-frequency unvoiced interval, and the ultra-high frequency interval. The adaptive filtering algorithm designs corresponding filter parameters based on the characteristics of each frequency band, including the center frequency, bandwidth, and gain coefficient. The frequency response curve of the filter is specifically optimized for the pronunciation characteristics of numbers and letters in the voice verification code. The analysis and processing calculate parameters such as energy distribution, spectral peak position, bandwidth characteristics, and spectral tilt within each frequency band, forming a feature vector describing the performance of the voice verification code in each frequency band. The spectral profile features are generated by statistically analyzing the main characteristic parameters of each frequency band, including the frequency domain distribution pattern, energy concentration trend, and spectral variation law of the voice verification code.
[0053] Spectral envelope modeling, a technique in speech signal processing, transforms discrete spectral features into a continuous envelope function representation using mathematical modeling methods. Spectral envelope modeling is used to describe the overall shape and trend of the spectrum in speech signal processing, approximating the envelope curve of the spectrum by fitting a mathematical function. The modeling process employs polynomial fitting or spline interpolation, using feature parameters from 12 key frequency bands as control points to construct a continuous function describing the entire spectral contour. Polynomial fitting uses the least squares method to determine the polynomial coefficients, minimizing the error between the fitted curve and the spectral contour feature points. Spline interpolation uses piecewise cubic polynomials to ensure the smoothness and continuity of the envelope curve. Spectral envelope modeling considers both static and dynamic characteristics of the spectrum. Static characteristics describe the distribution shape of the spectrum at a specific moment, while dynamic characteristics describe the trend and rate of change of the spectrum over time. The spectral envelope parameters of the speech CAPTCHA include mathematical feature parameters such as the coefficients of the envelope function, peak positions, valley positions, slope, and symmetry. These parameters comprehensively describe the spectral envelope characteristics of the speech CAPTCHA.
[0054] The spectrum encoding matrix generation process for voice verification codes converts the mathematical description of the spectrum envelope into a digital encoding format suitable for storage and transmission. Spectrum encoding matrix generation is an application of data compression and encoding techniques in speech processing, achieving efficient data representation by organizing and quantizing spectrum envelope parameters in matrix form. The generation process first normalizes and quantizes the spectrum envelope parameters, mapping continuous parameter values to discrete digital codes. The quantization precision is balanced according to speech quality requirements and storage space limitations. Rows in the encoding matrix correspond to different envelope feature parameter types, columns correspond to time series or frequency series, and the values of the matrix elements represent the quantized encoded values of the corresponding parameters. The encoding process employs a compression strategy combining differential coding and entropy coding. Differential coding utilizes the correlation between adjacent parameter values to reduce data redundancy, while entropy coding performs variable-length encoding based on the statistical distribution of parameter values. The multimodal voice verification code for television uses a spectrum encoding matrix to achieve a unified encoded representation of the voice verification code content, user identity features, and acoustic environment information. The encoding matrix simultaneously contains the spectral features of the voice verification code, the user's multimodal identity identifier, and the television's environmental adaptation information.
[0055] The above describes the voice verification code generation method for multimodal authentication on a TV terminal in this application embodiment. The following describes the voice verification code generation system for multimodal authentication on a TV terminal in this application embodiment. Please refer to [link to relevant documentation]. Figure 2 One embodiment of the voice verification code generation system for multimodal identity authentication on a television terminal in this application includes: The compensation module is used to perform acoustic compensation processing on the acoustic signals of the living room environment collected by the TV to obtain acoustic correction parameters; The extraction module is used to extract voiceprint features from the user's voice to obtain the user's voiceprint template, and at the same time, to perform spatial positioning processing on the user's voice position to obtain the spatial position feature vector. The authentication module is used to perform multimodal identity authentication processing on the user voiceprint template and the spatial location feature vector to obtain the user identity feature code; The synthesis module is used to perform personalized speech synthesis processing on the verification code sequence based on the user identity feature code and the acoustic correction parameters to obtain user-exclusive voice verification code audio. The encoding module is used to perform spectrum encoding processing on the user-specific voice verification code audio to obtain a multimodal voice verification code for the TV. The binding module is used to associate and bind the TV terminal multimodal voice verification code with the user identity feature code to obtain the TV terminal identity authentication credential.
[0056] above Figure 2 The voice verification code generation system for multimodal identity authentication on the TV terminal in this embodiment of the invention is described in detail from the perspective of modular functional entities. The voice verification code generation device for multimodal identity authentication on the TV terminal in this embodiment of the invention is described in detail from the perspective of hardware processing.
[0057] Reference Figure 3 This invention also provides a voice verification code generation device for multimodal identity authentication on a television. This device can be a server, and its internal structure can be as follows: Figure 3 As shown, the voice verification code generation device for multimodal identity authentication on a television terminal includes a processor, memory, display screen, input device, network interface, and database connected via a system bus. The processor, designed as a computer, provides computing and control capabilities. The memory of the voice verification code generation device for multimodal identity authentication on a television terminal includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores the operating system, computer programs, and database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The database of the voice verification code generation device for multimodal identity authentication on a television terminal stores the data corresponding to this embodiment. The network interface of the voice verification code generation device for multimodal identity authentication on a television terminal is used for communication with external terminals via a network connection. When the computer program is executed by the processor, it implements the above-described method.
[0058] Those skilled in the art will understand that Figure 3 The structure shown is merely a block diagram of a portion of the structure related to the present invention and does not constitute a limitation on the voice verification code generation device for multimodal identity authentication on television terminals to which the present invention is applied.
[0059] The present invention also provides a computer-readable storage medium, which can be a non-volatile computer-readable storage medium or a volatile computer-readable storage medium, wherein the computer-readable storage medium stores instructions that, when the instructions are executed on a computer, cause the computer to perform the steps of the voice verification code generation method for multimodal identity authentication on a television terminal.
[0060] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0061] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a multimodal identity authentication voice verification code generation device on a television (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0062] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for generating voice verification codes for multimodal identity authentication on a television terminal, characterized in that, The method includes: Step S1: Perform acoustic compensation processing on the acoustic signals of the living room environment collected by the TV to obtain acoustic correction parameters; Step S2: Extract voiceprint features from the user's voice to obtain the user's voiceprint template, and simultaneously perform spatial localization processing on the user's voice location to obtain the spatial location feature vector. Step S3: Perform multimodal identity authentication processing on the user voiceprint template and the spatial location feature vector to obtain the user identity feature code; Step S4: Perform personalized speech synthesis processing on the verification code sequence based on the user identity feature code and the acoustic correction parameters to obtain the user-exclusive voice verification code audio; Step S5: Perform spectrum encoding on the user-specific voice verification code audio to obtain a multimodal voice verification code for the TV. Step S6: Associate and bind the TV-side multimodal voice verification code with the user identity feature code to obtain the TV-side identity authentication credential.
2. The voice verification code generation method for multimodal identity authentication on a television terminal according to claim 1, characterized in that, Step S1 includes: Acoustic signals from the living room environment are collected and processed using the TV's built-in microphone array to obtain raw environmental acoustic data. The original environmental acoustic data is processed to extract reverberation time and sound propagation delay to obtain the acoustic characteristic parameters of the living room; The acoustic distance of the living room acoustic feature parameters is calculated based on the radial basis function to obtain the spatial distance value of the sound source. The acoustic compensation coefficient is obtained by calculating the compensation coefficient of the frequency response function based on the spatial distance value of the sound source. The acoustic compensation coefficient is multiplied by the reference gain parameter to obtain the acoustic correction parameter.
3. The voice verification code generation method for multimodal identity authentication on a television terminal according to claim 1, characterized in that, Step S2 includes: The user's spoken voice is processed by fundamental frequency extraction to obtain a fundamental frequency trajectory sequence; Based on the fundamental frequency trajectory sequence, first-order and second-order difference calculations are performed to obtain fundamental frequency change direction data; The fundamental frequency change direction data is subjected to 8-direction chain code quantization to obtain the user voiceprint template; The time difference of arrival of the user's voice is calculated by using a TV microphone array to obtain three-dimensional spatial coordinate data. Based on the three-dimensional spatial coordinate data, the horizontal angle, vertical angle, and distance are layered to obtain the spatial location feature vector.
4. The method for generating voice verification codes for multimodal identity authentication on a television terminal according to claim 1, characterized in that, Step S3 includes: The user voiceprint template is subjected to data standardization processing to obtain standardized voiceprint features; The spatial location feature vector is mapped using a fully connected network to obtain the spatial location code; The standardized voiceprint features and the spatial location encoding are weighted using an attention weighting algorithm to obtain a fused feature vector. The fused feature vector is processed by an identity classifier to obtain the user identity feature code.
5. The method for generating voice verification codes for multimodal identity authentication on a television terminal according to claim 1, characterized in that, Step S4 includes: Personalized parameter mapping processing is performed on the random verification code sequence based on the user identity feature code to obtain a user-specific verification code sequence. The user-specific verification code sequence is input into a speech synthesis network for text-to-speech processing to obtain the initial voice verification code audio. Based on the acoustic correction parameters, the initial voice verification code audio is adapted to the living room environment to obtain the acoustically compensated voice verification code audio. The acoustically compensated voice verification code audio is combined with the user identity feature code to perform audio fingerprint embedding processing to obtain the user-specific voice verification code audio.
6. The method for generating voice verification codes for multimodal identity authentication on a television terminal according to claim 5, characterized in that, The step of inputting the user-specific verification code sequence into a speech synthesis network for text-to-speech processing to obtain the initial voice verification code audio includes: The user-specific verification code sequence is input into the encoder network for text feature encoding to obtain the verification code text feature vector. The uniqueness constraint processing of the CAPTCHA text feature vector is performed based on the compression coding mechanism to obtain the coding constraint features; The encoded constraint features are input into the generator network for adversarial generation processing to obtain a speech spectrum representation. The speech spectrum representation is subjected to multi-scale time-frequency domain discrimination processing based on a discriminator network to obtain the authenticity verification result. Based on the authenticity verification result, the speech spectrum representation is subjected to inverse Mel-spectrum transform to obtain the initial speech verification code audio. The initial voice verification code audio is processed using the far-field voice characteristics of the TV terminal to obtain the voice verification code audio.
7. The method for generating voice verification codes for multimodal identity authentication on a television terminal according to claim 1, characterized in that, Step S5 includes: The user-specific voice verification code audio is transformed from the time domain to the frequency domain to obtain the voice verification code spectrum data. Based on the frequency domain adaptive filtering algorithm, the spectral data of the voice verification code is analyzed and processed in 12 key frequency bands to obtain spectral contour features. The spectral contour features are subjected to spectral envelope modeling to obtain the spectral envelope parameters of the voice verification code. The multimodal voice verification code for the TV terminal is obtained by generating a spectral coding matrix based on the spectral envelope parameters of the voice verification code.
8. A voice verification code generation system for multimodal identity authentication on a television terminal, characterized in that, The voice verification code generation method for implementing multimodal identity authentication on a television terminal as described in any one of claims 1-7, wherein the voice verification code generation system for multimodal identity authentication on a television terminal comprises: The compensation module is used to perform acoustic compensation processing on the acoustic signals of the living room environment collected by the TV to obtain acoustic correction parameters; The extraction module is used to extract voiceprint features from the user's voice to obtain the user's voiceprint template, and at the same time, to perform spatial positioning processing on the user's voice position to obtain the spatial position feature vector. The authentication module is used to perform multimodal identity authentication processing on the user voiceprint template and the spatial location feature vector to obtain the user identity feature code; The synthesis module is used to perform personalized speech synthesis processing on the verification code sequence based on the user identity feature code and the acoustic correction parameters to obtain user-exclusive voice verification code audio. The encoding module is used to perform spectrum encoding processing on the user-specific voice verification code audio to obtain a multimodal voice verification code for the TV. The binding module is used to associate and bind the TV terminal multimodal voice verification code with the user identity feature code to obtain the TV terminal identity authentication credential.
9. A voice verification code generation device for multimodal identity authentication on a television terminal, characterized in that, The device includes a memory and a processor, wherein the memory stores a computer program that can run on the processor, and the processor executes the computer program to implement the voice verification code generation method for multimodal identity authentication on a television terminal as described in any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is run by the processor, it causes the processor to execute the voice verification code generation method for multimodal identity authentication on a television terminal as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Vocal print secret key generating method and device and logging-in method and system based on vocal print secret key
CN103973453A
Verification code implementation method and device based on semantic voiceprint interaction mode
CN115294993A
Identity verification method and system based on machine learning
CN115424608A
Verification code intelligent identification and interaction method based on multi-modal large model
CN120470578A
Voiceprint password key system
JP2007034031A
Cited By
Government affair form submission system and method based on multi-modal verification
CN121811888A