Voice code generation method and system for multi-modal identity authentication of television end
By generating personalized voice verification codes through multimodal identity authentication technology, the security and accuracy issues of identity authentication on TVs are solved, enabling efficient user authentication in the living room environment.
Patent Information
- Application Number
- CN202511477090.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-16
- Publication Date
- 2025-12-23
- Estimated Expiration
- 2045-10-16
AI Technical Summary
Existing TV-based identity authentication technologies suffer from several problems: voice verification codes are vulnerable to recording and playback attacks; voice clarity is insufficient in living room environments; recognition accuracy is low; and they cannot effectively distinguish between authorized and unauthorized users.
By integrating multimodal features such as user voiceprint and spatial location, a personalized voice verification code is generated, including acoustic compensation processing, voiceprint feature extraction, spatial positioning, multimodal identity authentication, personalized speech synthesis, and spectral coding, forming a user-exclusive multimodal voice verification code.
It improves the security and accuracy of identity authentication on TV, prevents recording replay attacks, adapts to the living room environment, and realizes personalized feature binding and authentication of user identity.
Smart Images

Figure CN120935408B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of speech recognition, and in particular to a television terminal multi-modal identity authentication speech verification code generation method and system. BACKGROUND
[0002] With the rapid development of smart TVs and home entertainment systems, television terminal user identity authentication has become an important technical requirement to protect home privacy and personalized services. Existing television terminal identity authentication technologies mainly rely on traditional remote control key input, simple speech recognition or basic face recognition and other single modal authentication methods. In the aspect of speech verification code technology, existing solutions usually use standardized text-to-speech synthesis technology to generate uniform format speech verification codes. All users receive verification codes that are basically the same in terms of speech characteristics, tone and playback parameters, and only differ in verification code content.
[0003] However, the existing technology has significant security and applicability deficiencies. First, uniform format speech verification codes are vulnerable to recording and replay attacks, and attackers can bypass identity authentication by recording and replaying speech verification codes. Second, existing speech verification codes do not take into account the acoustic characteristics of the living room environment, and there are problems of insufficient speech clarity and low recognition accuracy in the living room environment at a distance and with multiple noises. Third, single speech content verification cannot effectively distinguish between authorized users and unauthorized users, and lacks a user identity personalized feature binding mechanism. SUMMARY
[0004] The present application provides a television terminal multi-modal identity authentication speech verification code generation method and system, which solves the problem of uniform speech verification codes in existing technologies that are vulnerable to attacks and cannot be personalized according to user multi-modal identity features. By fusing user voiceprints, spatial positions and other multi-modal features, the present application realizes dynamic generation of personalized speech verification codes and improves the security and user experience of television terminal identity authentication.
[0005] In a first aspect, the present application provides a television terminal multi-modal identity authentication speech verification code generation method, which includes:
[0006] Step S1: Perform acoustic compensation processing on the living room environment acoustic signals collected by the television terminal to obtain acoustic calibration parameters;
[0007] Step S2: Perform voiceprint feature extraction processing on the user's vocal speech to obtain a user voiceprint template, and perform spatial positioning processing on the user's vocal position to obtain a spatial position feature vector;
[0008] Step S3: performing multi-modal identity authentication processing on the user voiceprint template and the spatial position feature vector to obtain a user identity feature code;
[0009] Step S4: performing personalized speech synthesis processing on a verification code sequence according to the user identity feature code and the acoustic correction parameter to obtain a user-specific speech verification code audio;
[0010] Step S5: performing spectrum coding processing on the user-specific speech verification code audio to obtain a television-side multi-modal speech verification code;
[0011] Step S6: performing association and binding processing on the television-side multi-modal speech verification code and the user identity feature code to obtain a television-side identity authentication credential.
[0012] In a second aspect, the present application provides a television-side multi-modal identity authentication speech verification code generation system, which comprises:
[0013] A compensation module is configured to perform acoustic compensation processing on a living room environment acoustic signal collected by the television side to obtain an acoustic correction parameter;
[0014] An extraction module is configured to perform voiceprint feature extraction processing on user voice to obtain a user voiceprint template, and perform spatial positioning processing on a user voice position to obtain a spatial position feature vector;
[0015] An authentication module is configured to perform multi-modal identity authentication processing on the user voiceprint template and the spatial position feature vector to obtain a user identity feature code;
[0016] A synthesis module is configured to perform personalized speech synthesis processing on a verification code sequence according to the user identity feature code and the acoustic correction parameter to obtain a user-specific speech verification code audio;
[0017] An encoding module is configured to perform spectrum coding processing on the user-specific speech verification code audio to obtain a television-side multi-modal speech verification code;
[0018] A binding module is configured to perform association and binding processing on the television-side multi-modal speech verification code and the user identity feature code to obtain a television-side identity authentication credential.
[0019] In a third aspect, a television-side multi-modal identity authentication speech verification code generation device is provided, which comprises a memory and at least one processor, and the memory stores instructions; the at least one processor invokes the instructions in the memory, so that the television-side multi-modal identity authentication speech verification code generation device performs the television-side multi-modal identity authentication speech verification code generation method described above.
[0020] In a fourth aspect, a computer readable storage medium is provided, and the computer readable storage medium stores instructions which, when executed on a computer, cause the computer to perform the method for generating a voice verification code in a multi-modal identity authentication of a television end described above.
[0021] In the technical scheme provided in the present application, the acoustic compensation processing is performed on the living room environment acoustic signal collected by the television end to obtain the acoustic calibration parameter, effectively solving the problem of transmission distortion of the voice verification code in the living room environment in the prior art, and ensuring that the voice verification code can be adaptively adjusted according to different living room acoustic environments. At the same time, the voiceprint feature extraction processing is performed on the user's voice to obtain the user's voiceprint template, and the spatial positioning processing is performed on the user's voice position to obtain the spatial position feature vector, realizing multi-dimensional feature capture of the user's identity. Compared with the single modal authentication mode of the prior art which only relies on voice content verification, the accuracy and anti-fake ability of identity authentication are significantly enhanced. The technical features of the user's voiceprint template and the spatial position feature vector are processed in a multi-modal identity authentication to obtain the user's identity feature code. By fusing acoustic features and spatial features, comprehensive discrimination of the user's identity is realized, avoiding the security risks of single features being easily forged in the prior art. The personalized voice synthesis processing is performed on the verification code sequence according to the user's identity feature code and the acoustic calibration parameter to obtain the user's exclusive voice verification code audio, breaking through the technical limitation of the prior art that all users use the same format voice verification code, realizing personalized verification code generation based on the user's identity feature, and fundamentally solving the security threat of the recording and replay attack.
[0022] The technical features of the user's exclusive voice verification code audio are processed by spectrum coding to obtain the multi-modal voice verification code of the television end, and the multi-modal voice verification code of the television end and the user's identity feature code are associated and bound to obtain the identity authentication credential of the television end, realizing the whole-link personalized processing of the voice verification code from content generation to final authentication. In particular, in the specific application field of the method for generating a voice verification code in a multi-modal identity authentication of a television end, the radial basis function acoustic compensation algorithm used in the present application can accurately model the acoustic propagation characteristics of the living room environment, the 8-direction chain code quantization algorithm effectively extracts the personalized feature mode of the user's voiceprint, and the attention weight algorithm realizes the dynamic weight distribution of the voiceprint features and the spatial position features. The synergistic effect of these algorithm features makes the generated voice verification code not only have the unique identity of the user's identity, but also adapt to the use environment of the television end. Compared with the existing standardized voice synthesis technology, the voice verification code has significant improvements in attack resistance, personalization degree and environmental adaptability. BRIEF DESCRIPTION OF DRAWINGS
[0023] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings needed to be used in the embodiments description. Obviously, the drawings in the following description are some embodiments of the present application, and other drawings can be obtained by those skilled in the art without any creative effort based on these drawings.
[0024] Figure 1 An embodiment schematic diagram of a voice verification code generation method for television end multi-modal identity authentication in the present application;
[0025] Figure 2 An embodiment schematic diagram of a voice verification code generation system for television end multi-modal identity authentication in the present application;
[0026] Figure 3 An embodiment schematic diagram of a voice verification code generation device for television end multi-modal identity authentication in the present application. DETAILED DESCRIPTION
[0027] The present application provides a voice verification code generation method and system for television end multi-modal identity authentication. The terms "first", "second", "third", "fourth" and the like (if any) in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily mean a specific order or sequence. It should be understood that the data thus used can be interchanged under appropriate circumstances, so that the embodiments described herein can be implemented in an order other than that illustrated or described herein. In addition, the term "comprising" or "having" and any variation thereof is intended to cover non-exclusive inclusion, for example, a process, method, system, product or device comprising a series of steps or units does not necessarily limit to those steps or units clearly listed, but can include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0028] For the convenience of understanding, the specific process of the embodiments of the present application will be described below. Please refer to Figure 1 An embodiment of the voice verification code generation method for television end multi-modal identity authentication in the present application includes:
[0029] Step S1: acoustic compensation processing is performed on the living room environment acoustic signal collected by the television end to obtain an acoustic calibration parameter;
[0030] Step S2: voiceprint feature extraction processing is performed on the user's voice to obtain a user voiceprint template, and spatial positioning processing is performed on the user's voice position to obtain a spatial position feature vector;
[0031] Step S3: multi-modal identity authentication processing is performed on the user voiceprint template and the spatial position feature vector to obtain a user identity feature code;
[0032] Step S4: personalized speech synthesis processing is performed on the verification code sequence according to the user identity feature code and the acoustic calibration parameter to obtain a user-specific voice verification code audio;
[0033] Step S5: spectrum coding processing is performed on the user-specific voice verification code audio to obtain a television-side multi-modal voice verification code;
[0034] Step S6: association and binding processing is performed on the television-side multi-modal voice verification code and the user identity feature code to obtain a television-side identity authentication credential.
[0035] It can be understood that the execution subject of the present application can be a television-side multi-modal identity authentication voice verification code generation system, and can also be a terminal or a server, and the specific implementation is not limited herein. The server is taken as an example for description in the embodiments of the present application.
[0036] Specifically, the television built-in microphone array first collects original acoustic data of the living room environment, including background noise, echo and reverberation information. The radial basis function is a function with distance as the independent variable. By calculating the Euclidean distance between the positions of the television microphone array and the user's sound emission position, an acoustic distance model is established. Specifically, the acoustic signal detected by the microphone array contains a reverberation time parameter, which reflects the reflection attenuation characteristics of sound waves in the living room space. At the same time, the sound propagation delay parameter represents the time difference of sound waves from the sound emission point to each microphone. The radial basis function calculates the spatial distance between the sound source and the receiving point according to these acoustic characteristic parameters, and then combines the frequency response characteristics of the living room to calculate the compensation coefficient for different frequency bands. The acoustic calibration parameter is obtained by multiplying the compensation coefficient with a preset reference gain value, and is used for environmental adaptation in the subsequent speech synthesis process.
[0037] The voiceprint feature extraction and spatial positioning processing are performed simultaneously. The fundamental frequency extraction process identifies the time sequence trajectory of the basic frequency by analyzing the periodic characteristics of the user's voice signal. The first-order difference calculates the difference between the fundamental frequency values of adjacent time points, reflecting the trend of the fundamental frequency change, and the second-order difference calculates the difference between the first-order differences, reflecting the acceleration characteristics of the fundamental frequency change. The 8-direction chain code quantization maps the fundamental frequency change direction to 8 discrete directions, corresponding to different fundamental frequency change modes such as rising, falling, and stable, forming a unique voiceprint template for the user. The spatial positioning processing uses the time difference of arrival algorithm to calculate the time difference of the same sound source signal arriving at different microphones, and determines the three-dimensional spatial coordinates of the user using the triangulation principle. The horizontal angle represents the left-right deviation angle of the user relative to the front of the TV, the vertical angle represents the height position deviation of the user, and the distance layer divides the distance between the user and the TV into three levels: near, medium, and far, forming a spatial position feature vector.
[0038] Multi-modal feature fusion authentication is implemented. Data standardization processing converts the user's voiceprint template into a standardized format with a mean of zero and a variance of one, eliminating the dimensional differences of different user voiceprint data. The full connection network mapping performs nonlinear transformation on the spatial position feature vector through multiple layers of neural networks to generate high-dimensional spatial position encoding. The attention weight algorithm calculates the importance weight of each feature component, and normalizes the weight value through the softmax function to ensure that the sum of all weights is one. The fusion feature vector combines the standardized voiceprint features and spatial position encoding into a unified feature representation through weighted summation. The identity classifier recognition process uses a multi-layer perceptron structure to map the fusion feature vector to a unique identification code of the user's identity.
[0039] Generate personalized voice verification codes based on user identity feature codes. Personalized parameter mapping uses the user identity feature code as a seed to generate a verification code sequence bound to the user through a hash function, ensuring that different users obtain different verification code content. The speech synthesis network uses an encoder-generator-discriminator adversarial network architecture. The encoder converts the verification code text into a text feature vector, and the compression encoding mechanism ensures the uniqueness and unpredictability of the generated verification code through the Lempel-Ziv compression algorithm. The generator network uses an inverse convolution layer structure to gradually upsample the text feature vector into a speech spectrum representation, and the discriminator network judges the authenticity of the generated speech through multi-scale analysis. The living room environment adaptation process combines the initial voice verification code with the acoustic calibration parameters to adjust the volume, pitch, and spectral characteristics of the speech to adapt to the living room acoustic environment. The audio fingerprint embedding embeds the user identity feature code in the form of a watermark into the high-frequency components of the speech signal, forming a user-specific voice verification code audio.
[0040] The time-to-frequency domain transformation converts the speech time domain signal into a frequency domain representation using a fast Fourier transform to obtain the frequency spectrum distribution information of the speech verification code. The frequency domain adaptive filtering divides the frequency spectrum into 12 key frequency bands, each corresponding to the main frequency range of human voice, and extracts the amplitude and phase characteristics of each frequency band through an adaptive filter. The spectral envelope modeling calculates the envelope curve of each frequency band, which reflects the energy distribution law of the speech signal in that frequency band. The spectral coding matrix organizes the envelope parameters of the 12 frequency bands into a matrix form, forming a digital representation of the television end multi-modal speech verification code.
[0041] The association binding process combines the spectral coding matrix of the television end multi-modal speech verification code with the user identity feature code through an encryption hash algorithm to generate a composite credential containing user identity information and verification code content. This credential contains both the spectral features of the speech verification code and the multi-modal identity features of the user, enabling personalized speech verification code generation based on individual user characteristics.
[0042] In a specific embodiment, step S1 includes:
[0043] The original environmental acoustic data is obtained by collecting and processing the acoustic signals of the living room environment through the built-in microphone array of the television.
[0044] The living room acoustic characteristic parameters are obtained by performing reverberation time and sound propagation delay extraction processing on the original environmental acoustic data.
[0045] The acoustic distance value is obtained by performing acoustic distance calculation processing on the living room acoustic characteristic parameters based on the radial basis function.
[0046] The acoustic compensation coefficient is obtained by performing compensation coefficient calculation processing on the frequency response function according to the acoustic source spatial distance value.
[0047] The acoustic compensation coefficient is obtained by performing compensation coefficient calculation processing on the frequency response function according to the acoustic source spatial distance value.
[0048] Specifically, the TV built-in microphone array acoustic signal acquisition and processing records the sound wave signals in the living room environment simultaneously through multiple microphones, each of which converts the mechanical vibration of the sound wave into a voltage signal, and then converts the analog signal into a digital sampling sequence through an analog-to-digital converter. The microphone array usually adopts an L-shaped or circular configuration and contains 4 to 8 independent omnidirectional microphone units, each of which is kept at a fixed distance of centimeters. The original environmental acoustic data contains mixed signals generated by all sound sources in the living room, including user voice, TV playing sound, air conditioner running noise, footstep sound, etc. During the data acquisition process, each microphone independently records the time sequence of sound pressure changes to form a multi-channel acoustic data matrix, and the sampling frequency is usually set to 16 kHz or 48 kHz to cover the main frequency range of human voice. Reverberation time and sound propagation delay extraction processing are respectively analyzed for two key characteristic parameters of the living room acoustic environment. The reverberation time represents the time required for the sound to attenuate to one thousandth of the original intensity in the living room, which is calculated by analyzing the decay envelope curve of the sound signal. The specific method is to measure the duration of the residual echo in the room after the sound stops. The sound propagation delay is obtained by calculating the time difference of the sound wave propagating from the sound source to different microphone positions. Using the physical constant that the speed of sound in air is about 343 meters per second, combined with the time difference of the same sound signal received by each microphone, the distance difference between the sound source and each microphone is calculated. The living room acoustic characteristic parameters integrate the reverberation time and the sound propagation delay into a parameter vector reflecting the characteristics of the living room acoustic environment. This vector describes the propagation and attenuation of sound waves in a specific space.
[0049] The radial basis function acoustic distance calculation processing uses a nonlinear function with distance as a variable to establish the mapping relationship between the sound source position and the acoustic characteristics. The core principle of the radial basis function is to calculate the Euclidean distance between the input point and the preset center point, and then substitute the distance value into the exponential decay function for transformation. The acoustic distance calculation first determines the geometric center of the TV microphone array as the reference origin, and then calculates the three-dimensional coordinates of the user's sound position using the principle of triangulation according to the time difference of the sound signals received by each microphone. In the specific calculation process, by comparing the arrival time of the same sound pulse received by different microphones, and combining the known distance between the microphones, an equation set about the sound source position is established. The sound source spatial distance value represents the straight-line distance between the user's sound position and the center point of the TV, which directly affects the intensity attenuation and spectral characteristic change of the sound signal.
[0050] The frequency response function describes the transmission characteristics of the living room environment to sound waves of different frequencies. Low-frequency sound waves are less attenuated in space propagation but are prone to standing waves, while high-frequency sound waves are quickly attenuated but have stronger directivity. The compensation coefficient calculation analyzes the influence of the distance of the sound source on the signal intensity of each frequency band and establishes a compensation relationship between the distance and the frequency response. In the calculation process, the audible frequency range is divided into multiple sub-frequency bands, and each frequency band is assigned a corresponding compensation weight according to its attenuation characteristics at different distances. The acoustic compensation coefficient reflects the gain adjustment amount required for each frequency band signal under the current sound source distance condition. The farther the sound source, the greater the compensation gain required in the high-frequency part.
[0051] The product operation of the acoustic compensation coefficient and the reference gain parameter combines the frequency-dependent compensation coefficient with the device's inherent reference gain value. The reference gain parameter is the standard amplification factor of the television sound system, reflecting the signal amplification capability of the device under ideal conditions. The product operation multiplies each frequency band component of the compensation coefficient with the corresponding reference gain value to obtain the final correction parameter for the current acoustic environment. The acoustic correction parameter contains the comprehensive adjustment amount of the living room environment characteristics, sound source position information, and device characteristics, and this parameter is directly used for audio modulation in the subsequent voice verification code generation process.
[0052] In a specific embodiment, step S2 includes:
[0053] The user's voiced speech is subjected to a fundamental frequency extraction process to obtain a fundamental frequency trajectory sequence.
[0054] Based on the fundamental frequency trajectory sequence, first-order difference and second-order difference calculation processing is performed to obtain fundamental frequency change direction data.
[0055] The fundamental frequency change direction data is subjected to 8-direction chain code quantization processing to obtain a user voiceprint template.
[0056] The user's voiced position is subjected to time difference of arrival calculation processing by the television microphone array to obtain three-dimensional spatial coordinate data.
[0057] Based on the three-dimensional spatial coordinate data, horizontal angle, vertical angle, and distance layering processing are performed to obtain a spatial position feature vector.
[0058] Specifically, the user voice fundamental frequency extraction process identifies the basic frequency of vocal cord vibration by analyzing the periodicity characteristics of the voice signal. The fundamental frequency represents the number of times the vocal cords vibrate per second, and is an important feature parameter for distinguishing different speakers. The fundamental frequency extraction process first pre-processes the user voice signal, including pre-emphasis filtering and windowing and framing, which divides the continuous voice signal into a plurality of overlapping short-time frames, each frame usually having a length of 20 to 30 milliseconds. Then, the autocorrelation function method is used to detect the periodicity of each frame of voice, and the first peak position of the autocorrelation function is found to determine the fundamental frequency period, and then the period is converted into a fundamental frequency value by taking the reciprocal. The fundamental frequency trajectory sequence records the trajectory of the fundamental frequency changing over time during the entire user voice process, forming time series data reflecting the vocal cord vibration characteristics of the user. The fundamental frequency trajectory sequence not only contains the absolute value of the fundamental frequency, but also retains the timing information of the fundamental frequency change, laying a data foundation for subsequent voiceprint feature analysis.
[0059] The first-order difference and second-order difference calculation processes of the fundamental frequency trajectory sequence extract the speed and acceleration information of the fundamental frequency change. The first-order difference calculation subtracts the fundamental frequency values at adjacent time points in the fundamental frequency trajectory sequence to obtain the change amount of the fundamental frequency in each time interval, reflecting the trend and amplitude of the fundamental frequency rising or falling. The second-order difference calculation further performs a difference operation on the first-order difference results to obtain the change rate of the fundamental frequency change speed, reflecting the acceleration characteristics of the fundamental frequency change. The fundamental frequency change direction data is obtained by symbol judgment and amplitude analysis of the first-order difference and second-order difference results, converting the continuous fundamental frequency change into discrete direction identifiers. A positive value indicates that the fundamental frequency is rising, a negative value indicates that the fundamental frequency is falling, and a zero value indicates that the fundamental frequency is stable, and the magnitude of the change amplitude reflects the degree of change in the fundamental frequency. The fundamental frequency change direction data captures the fluctuation pattern of the user's voice tone, forming a fundamental frequency change fingerprint with personal characteristics.
[0060] The 8-direction chain code quantization process of the fundamental frequency change direction data maps the continuous fundamental frequency change information into an encoded sequence of 8 discrete directions. The 8-direction chain code uses an encoding method similar to contour tracking in image processing, combining the direction and amplitude of the fundamental frequency change into 8 basic directions: sharp rise, slow rise, stable, slow fall, sharp fall, fluctuating rise, fluctuating fall, and complex change. The quantization process sets threshold values according to the numerical range of the first-order difference and second-order difference, and divides the fundamental frequency change into corresponding direction categories. The user voiceprint template is formed by statistically analyzing the 8-direction chain code distribution characteristics of each user under different voice content, forming a template vector reflecting the individual acoustic characteristics of the user. The voiceprint template not only records the statistical distribution of the user's fundamental frequency change, but also retains the timing correlation information of the change pattern, constituting the identity of the user in the voice interaction on the television end.
[0061] The television microphone array user vocalization position time difference of arrival calculation process locates the spatial position of the user using the physical properties of sound wave propagation. The time difference of arrival algorithm calculates the spatial coordinates of the sound source based on the time difference of the same sound signal arriving at different microphones, combined with the propagation speed of sound waves in the air. In the calculation process, first, the characteristic pulses or speech starting points in the speech signals received by each microphone are identified, and then the accurate time stamps of these characteristic points arriving at each microphone are calculated. By comparing the time difference between different microphone pairs, an over-determined equation set about the user's position is established, and the least squares method is used to solve the three-dimensional spatial coordinates of the user. The three-dimensional spatial coordinate data takes the center of the television screen as the origin to establish a rectangular coordinate system, recording the front-back distance, left-right offset, and up-down height position information of the user relative to the television.
[0062] The horizontal angle, vertical angle, and distance layering processing of the three-dimensional spatial coordinate data converts continuous spatial position information into discrete feature representation. The horizontal angle calculates the left-right deflection angle of the user relative to the front of the television, and the inverse tangent value is obtained by calculating the x-coordinate and z-coordinate of the user's position. The vertical angle calculates the up-down deflection angle of the user relative to the horizontal center line of the television, and the vertical deflection angle is calculated by the y-coordinate and z-coordinate of the user's position. The distance layering processing divides the straight-line distance between the user and the television into three levels: close, medium, and far, each level corresponding to different speech interaction scenarios and acoustic characteristics. The spatial position feature vector encodes the horizontal angle, vertical angle, and distance level into a fixed-dimensional numerical vector, forming a standardized feature representation describing the user's spatial position.
[0063] In a specific embodiment, step S3 comprises:
[0064] The user voiceprint template is subjected to data standardization processing to obtain a standardized voiceprint feature;
[0065] The spatial position feature vector is subjected to full connection network mapping processing to obtain a spatial position code;
[0066] The standardized voiceprint feature and the spatial position code are subjected to weight distribution processing based on an attention weight algorithm to obtain a fusion feature vector;
[0067] The fusion feature vector is subjected to identity classifier recognition processing to obtain a user identity feature code.
[0068] Specifically, the user voiceprint template data standardization processing eliminates the dimensional difference and numerical range inconsistency between different user voiceprint data. Data standardization is a preprocessing technique in machine learning, which converts the original data into a standard format with uniform distribution characteristics through mathematical transformation. The voiceprint template contains the frequency statistics of 8-direction chain code and the amplitude information of fundamental frequency variation. The voiceprint data of different users has significant differences in numerical range and distribution characteristics. The standardization processing first calculates the mean and standard deviation of each feature dimension in the voiceprint template, then subtracts the mean and divides by the standard deviation for each feature value, to obtain standardized data with a mean of zero and a variance of one. Standardized voiceprint features eliminate the influence of individual differences on numerical size, highlight the relative change characteristics of voiceprint patterns, and make the voiceprint data of different users have equal weight and comparison basis. The relative relationship of the original voiceprint features is maintained during the standardization process, ensuring that the uniqueness of the user's individual acoustic characteristics is preserved.
[0069] The spatial position feature vector full connection network mapping processing converts discrete spatial position information into high-dimensional continuous feature representation. The full connection network is the basic structure of neural network, composed of multiple full connection layers, each layer contains a number of neuron nodes, and all nodes between adjacent layers are fully connected. The spatial position feature vector contains discrete encoding of horizontal angle, vertical angle and distance hierarchy, and the full connection network maps these discrete features into continuous high-dimensional vectors through multiple nonlinear transformations. In the mapping process, the first layer of full connection network receives the spatial position feature vector as input, performs linear transformation through weight matrix multiplication and bias addition, and then applies nonlinear activation functions such as ReLU or Sigmoid for nonlinear mapping. The cascade of multiple layers of full connection network enables the spatial position code to capture the complex nonlinear relationship between position features, forming a more rich spatial representation. The spatial position code not only preserves the original position information, but also obtains a deep semantic representation of the position features through network learning.
[0070] The attention weighting algorithm addresses the importance balance problem in multimodal feature fusion by standardizing the weights of voiceprint features and spatial location codes. Attention mechanisms, a technique in deep learning for dynamically allocating feature weights, adjust the importance weights of features by calculating their contribution to the final result. The weight allocation process first calculates the attention scores of standardized voiceprint features and spatial location codes by performing a dot product operation between the feature vector and the learnable query vector. Then, a softmax function is applied to normalize the original scores, ensuring that the sum of the weights of all features equals one, resulting in standardized attention weights. The weight allocation dynamically adjusts the importance ratio of voiceprint features and spatial location features based on different users and usage scenarios. Spatial location features receive higher weights when the user is far away, and voiceprint features receive higher weights when there is significant environmental noise. The fused feature vector combines the standardized voiceprint features and spatial location codes according to the attention weights through a weighted summation, forming a unified representation that comprehensively reflects the user's multimodal identity features.
[0071] The fusion feature vector identity classifier transforms multimodal fusion features into a unique identifier for each user. The identity classifier employs a multilayer perceptron architecture, a fully connected network architecture comprising an input layer, hidden layers, and an output layer. The input layer receives the fusion feature vector, the hidden layers extract a high-level abstract representation of the features through multiple nonlinear transformations, and the output layer generates a probability distribution of user identity categories. In the recognition process, the fusion feature vector is first linearly transformed through the weight matrix of the hidden layers, then a nonlinear activation function is applied for feature transformation. The cascading of multiple hidden layers enables the classifier to learn complex identity discrimination patterns. The output layer uses a softmax activation function to convert the output of the last hidden layer into a probability distribution of user identity categories, with the category with the highest probability corresponding to the user's identity identifier. The user identity feature code is generated by encoding the user category corresponding to the highest probability with its confidence level, forming a composite identifier code containing user identity information and authentication credibility.
[0072] In one specific embodiment, step S4 includes:
[0073] Personalized parameter mapping is performed on the random verification code sequence based on the user's identity feature code to obtain a user-specific verification code sequence;
[0074] The user-specific verification code sequence is input into a speech synthesis network for text-to-speech processing to obtain the initial voice verification code audio.
[0075] The initial voice verification code audio is adapted to the living room environment based on the acoustic positive parameter to obtain the acoustically compensated voice verification code audio.
[0076] The audio fingerprint embedding processing is performed on the acoustic compensation voice verification code audio and the user identity feature code to obtain a user-specific voice verification code audio.
[0077] Specifically, the user identity feature code random verification code sequence personalized parameter mapping processing converts the user identity information into a specific verification code generation parameter through a hash algorithm. The personalized parameter mapping is a key technology to ensure that different users obtain different verification code contents. By inputting the user identity feature code as a seed value into a pseudo-random number generator, a verification code sequence bound to the user is generated. The mapping processing first performs a hash operation on the user identity feature code to obtain a fixed-length numerical sequence, and then uses the sequence as the initial seed of the random number generator to ensure that the verification code sequence generated by the same user each time is consistent, while the verification code sequences of different users are completely different. The random verification code sequence usually contains a combination of numbers and letters, with a length of 4 to 6 bits. The mapping algorithm determines the specific content and character combination rule of the verification code according to the hash value of the user identity feature code. The user-specific verification code sequence not only contains the character content of the verification code, but also contains the generation timestamp and sequence number associated with the user identity, forming a verification code data structure with user uniqueness and time effectiveness.
[0078] The user-specific verification code sequence speech synthesis network text-to-speech processing uses an encoder-generator-discriminator adversarial network architecture to convert text verification codes into speech signals. The speech synthesis network is a deep learning-based text-to-speech conversion model, which includes three main modules: text encoding, acoustic modeling, and speech generation. The text-to-speech processing first inputs the verification code character sequence into the text encoder. The encoder uses a recurrent neural network or a Transformer structure to convert discrete text symbols into continuous text feature vectors, with each character corresponding to a high-dimensional semantic representation vector. The acoustic modeling module converts the text feature vector into acoustic parameters of the speech, including the fundamental frequency, spectral envelope, and duration information, and learns the mapping relationship between the text and the acoustic features through a multi-layer neural network. The speech generation module uses a vocoder to convert the acoustic parameters into the final speech waveform. The vocoder reconstructs the time-domain audio signal from the frequency-domain acoustic parameters through inverse transformation. The initial voice verification code audio contains standard pronunciation features and beat characteristics, with clear speech intelligibility and appropriate speech length.
[0079] The initial voice verification code audio is processed according to the characteristics of the living room acoustic environment to adjust the spectrum and amplitude of the voice signal. The living room environment adaptation is a special processing technology for the television application scenario, which considers the influence of the acoustic characteristics of the living room space, the distance of the user and the background noise and other factors on the voice propagation. The adaptation processing utilizes the pre-obtained acoustic calibration parameters, including compensation coefficients and gain adjustment amounts of different frequency bands, to readjust the spectrum distribution of the initial voice verification code. In the spectrum adjustment process, the low-frequency part is attenuated or enhanced according to the reverberation characteristics of the living room, the mid-frequency part is optimized for the main frequency range of the human voice, and the high-frequency part is compensated and amplified according to the distance attenuation characteristics. The amplitude adjustment adaptively adjusts the overall volume of the voice verification code according to the distance relationship between the user and the television and the gain information in the acoustic calibration parameters. The voice verification code audio after acoustic compensation has the acoustic characteristics suitable for the propagation in the living room environment, and can maintain clear intelligibility and appropriate volume level at the user receiving position.
[0080] The user identity feature code audio fingerprint embedding processing hides the user identity information in the form of digital watermark in the voice signal. The audio fingerprint embedding is a digital audio anti-counterfeiting technology, which realizes information hiding and verification by embedding imperceptible digital identifiers in the frequency domain or time domain of the audio signal. The embedding processing first converts the user identity feature code into a binary sequence, and then selects the high-frequency components or the voice gap part of the voice signal as the embedding carrier. The high-frequency embedding method modulates the identity information into the high-frequency band which is not sensitive to human ears, and carries the digital information by fine-tuning the spectral amplitude, and the embedding strength is controlled within the range that does not affect the voice quality. The time domain embedding method embeds the identity information in the subtle changes of the voice waveform, and carries the binary data by adjusting the amplitude value of the sampling point. The embedding algorithm ensures that the user identity feature code is closely bound with the voice verification code content, forming a composite audio signal containing both the verification code voice content and the user identity identifier. The user-specific voice verification code audio integrates personalized verification code content, acoustic characteristics suitable for the living room environment and digital fingerprints of the user identity, and realizes the generation of personalized voice verification code based on multi-modal identity features.
[0081] In a specific embodiment, the process of inputting the user-specific verification code sequence into the speech synthesis network for text-to-speech processing can specifically include the following steps:
[0082] Inputting the user-specific verification code sequence into the encoder network for text feature encoding processing to obtain a verification code text feature vector;
[0083] Conducting unique constraint processing on the verification code text feature vector based on a compression encoding mechanism to obtain an encoding constraint feature;
[0084] The coded constraint feature is input into a generator network for adversarial generation processing to obtain a speech spectrum representation;
[0085] The speech spectrum representation is subjected to multi-scale time-frequency domain discrimination processing based on a discriminator network to obtain a reality verification result;
[0086] The speech spectrum representation is subjected to Mel spectrum inverse transform processing according to the reality verification result to obtain an initial voice verification code audio;
[0087] The initial voice verification code audio is subjected to television-side far-field speech characteristic processing to obtain a voice verification code audio.
[0088] Specifically, the user-specific verification code sequence encoding network text feature encoding processing converts discrete verification code characters into continuous high-dimensional feature representations through a deep neural network. The encoder network is the core component of the sequence-to-sequence model, which uses a multi-layer recurrent neural network or a Transformer structure to process text sequence data. The text feature encoding first converts each character in the verification code into a one-hot encoding vector, with numbers and letters corresponding to different encoding positions, and then maps the sparse one-hot vector to a dense word vector representation through an embedding layer. The recurrent layer of the encoder processes the characters in the verification code sequence one by one, receives the word vector of the current character and the hidden state of the previous time step at each time step, updates the hidden state through a gating mechanism and generates the context representation of the current character. The multi-layer encoding structure enables the network to capture complex dependency relationships and semantic information between characters, and the final output layer generates a text feature vector that contains the semantic information of the entire verification code sequence. The verification code text feature vector not only encodes the literal meaning of the characters, but also contains deep information such as character order, speech characteristics, and semantic associations, providing a rich feature basis for subsequent speech synthesis processing.
[0089] The compression encoding mechanism verification code text feature vector uniqueness constraint processing uses the data compression principle in information theory to ensure the uniqueness and unpredictability of the generated verification code. The compression encoding mechanism is based on the core idea of the Lempel-Ziv algorithm, which analyzes the repetition patterns and redundant information in the data to achieve efficient compression and uniqueness guarantee. The uniqueness constraint processing first analyzes the patterns in the verification code text feature vector to identify possible repeated or similar patterns in the feature vector, and then applies a compression algorithm to encode and remove these patterns. During compression, the algorithm constructs a dynamic dictionary to record the feature patterns that have appeared, and assigns a new encoding identifier to new patterns and references the existing encoding for repeated patterns, thereby achieving feature compression and uniqueness identification. The encoded constraint feature retains the core information of the original text feature after compression, while adding a uniqueness identifier and a duplicate prevention mechanism, ensuring that each verification code corresponds to a unique feature vector in the entire verification code space.
[0090] The encoding constraint feature generator network adopts the framework of a generative adversarial network to convert text features into speech spectral representations. The generator network is a component in the adversarial generation network responsible for data generation, which gradually upsamples low-dimensional feature vectors into high-dimensional spectral data through a multi-layer deconvolution network. In the adversarial generation process, the generator receives the encoding constraint features as input, maps the feature vector to the initial dimension of spectral generation through a linear transformation layer, and then gradually increases the time and frequency resolution of the spectrum through multi-layer deconvolution operations. Each layer of deconvolution uses different convolution kernel sizes and steps to control the generation details of the spectrum, and the activation function uses Leaky ReLU to maintain gradient flow and non-linear expression capability. The output layer of the generator uses a Tanh activation function to normalize the spectral values within a standard range, forming a two-dimensional spectral matrix containing time axis and frequency axis information. The speech spectral representation contains the complete frequency domain information of the verification code speech, with each time frame corresponding to a frequency distribution vector, and the entire matrix describing the spectral characteristics of the speech signal over time.
[0091] The discriminator network speech spectral representation multi-scale time-frequency domain discrimination process evaluates the authenticity and quality of the generated spectrum through an adversarial training mechanism. The discriminator network is a component in the adversarial generation network responsible for authenticity discrimination, which uses a multi-scale convolutional neural network structure to analyze the authenticity features of the spectral data. Multi-scale time-frequency domain discrimination includes time scale discrimination and frequency scale discrimination in two dimensions. Time scale discrimination analyzes the continuity and naturalness of the spectrum in the time dimension through convolution kernels of different window lengths, and frequency scale discrimination detects the reasonableness and authenticity of the spectrum in the frequency domain through convolution kernels of different frequency bandwidths. The convolution layer of the discriminator uses multiple parallel convolution branches, each branch processing specific scale spectral features, and the multi-scale features are fused into the final discrimination result through pooling layers and fully connected layers. The authenticity verification result outputs a probability value between 0 and 1 through a Sigmoid activation function, representing the confidence of the generated spectrum being judged as real speech. This result is used to guide the optimization direction of the generator and evaluate the quality of the spectrum.
[0092] The authenticity verification result voice spectrum representation mel-spectrum inverse transform processing converts the spectral data in the frequency domain back to the time-domain audio waveform signal. The mel-spectrum inverse transform is a classic algorithm in speech signal processing, which restores the frequency-domain representation to the time-domain audio through inverse spectral transform. The inverse transform processing first performs inverse mapping of the mel filter bank on the speech spectrum representation, converting the mel frequency scale back to the linear frequency scale to restore the original frequency distribution characteristics of the spectrum. Then, the inverse discrete cosine transform is applied to convert the spectral coefficients into power spectral density, and the inverse fast Fourier transform is used to convert the power spectrum in the frequency domain to the time-domain audio waveform. Phase reconstruction is required during the inverse transform process, and the Griffin-Lim algorithm or neural vocoder is used to estimate the phase information corresponding to the spectrum to ensure the time-domain continuity and auditory quality of the reconstructed audio. The initial voice verification code audio obtains complete time-domain waveform data through inverse transform, which contains the pronunciation characteristics, tone changes and duration information of the verification code characters.
[0093] The television end far-field voice characteristic processing optimizes and adjusts the audio signal according to the special requirements of the television use scenario. The television end far-field voice characteristics include distance attenuation compensation, directivity enhancement and background noise suppression processing techniques, which are specifically designed to solve the problem of listening to the voice verification code at a long distance in the living room environment. The far-field voice characteristic processing first analyzes the spectral distribution and dynamic range of the initial audio, and then performs adaptive adjustment according to the characteristics of the television sound system and the living room acoustic environment. Distance compensation compensates for high-frequency attenuation in air propagation by enhancing high-frequency components, directivity enhancement improves the spatial positioning of the voice by adjusting the stereo balance, and background noise suppression reduces the impact of environmental noise on voice clarity through spectral subtraction or Wiener filtering. The voice verification code audio has acoustic characteristics suitable for television end playback and user long-distance reception after far-field characteristic processing.
[0094] In a specific embodiment, step S5 comprises:
[0095] The time-domain to frequency-domain transform processing is performed on the user-specific voice verification code audio to obtain voice verification code spectral data.
[0096] The frequency-domain adaptive filtering algorithm is used to analyze and process the voice verification code spectral data in 12 key frequency bands to obtain spectral contour features.
[0097] The spectral envelope modeling processing is performed on the spectral contour features to obtain voice verification code spectral envelope parameters.
[0098] The frequency spectrum coding matrix generation processing is performed according to the voice verification code spectral envelope parameters to obtain the television end multi-modal voice verification code.
[0099] Specifically, the user-specific voice verification audio time-to-frequency domain transformation process converts the time series audio waveform into frequency domain spectral distribution data through Fast Fourier Transform. Time-to-frequency domain transformation is a basic technology of digital signal processing, which converts the amplitude change of audio signal from time dimension to energy distribution from frequency dimension. The transformation process first preprocesses the voice verification audio, including pre-emphasis filtering and windowing framing, which divides the continuous audio signal into several overlapping short-time frames, with a frame length of 25 milliseconds and a frame shift of 10 milliseconds, ensuring the balance of time resolution and frequency resolution of spectral analysis. The Fast Fourier Transform algorithm converts the time domain sample point sequence into complex spectral coefficients through frequency domain conversion for each audio frame, and obtains the intensity and phase information of frequency components through amplitude spectrum and phase spectrum. The voice verification spectral data is organized in a two-dimensional matrix form, with rows corresponding to time frames and columns corresponding to frequency components. Each element in the matrix represents the energy intensity at a specific time and frequency, forming a data structure that reflects the complete spectral characteristics of the voice verification.
[0100] The frequency domain adaptive filtering algorithm voice verification spectral data 12 key frequency band analysis process divides the spectrum into targeted analysis intervals based on the frequency distribution characteristics of human voice. Frequency domain adaptive filtering is a spectral analysis technique specifically designed for speech signals, which extracts speech features in different frequency bands by constructing multiple bandpass filters. The division of 12 key frequency bands is based on phonetics and acoustics principles, covering the main frequency components of speech such as fundamental frequency and its harmonics, formant frequencies, high-frequency fricative sounds, etc. Specifically, it includes 12 frequency bands such as fundamental frequency interval, first formant interval, second formant interval, third formant interval, low-frequency energy interval, mid-frequency unvoiced sound interval, high-frequency unvoiced sound interval, and ultrahigh-frequency interval. The adaptive filtering algorithm designs corresponding filter parameters for each frequency band based on its characteristics, including center frequency, bandwidth, and gain coefficient. The frequency response curve of the filter is specifically optimized for the pronunciation characteristics of numbers and letters in the voice verification. The analysis process calculates the energy distribution, spectral peak position, bandwidth characteristics, and spectral inclination in each frequency band, forming a feature vector that describes the performance of the voice verification in each frequency band. The spectral profile features are generated by statistical analysis of the main feature parameters of each frequency band, including the frequency domain distribution pattern, energy concentration trend, and spectral change law of the voice verification.
[0101] The spectral contour feature spectrum envelope modeling process converts discrete spectral features into continuous envelope function representations through mathematical modeling methods. Spectral envelope modeling is a technique used in speech signal processing to describe the overall shape and variation trend of the spectrum, which approximates the envelope curve of the spectrum by fitting mathematical functions. The modeling process uses polynomial fitting or spline interpolation methods to construct a continuous function that describes the entire spectral contour, with the feature parameters of 12 key frequency bands as control points. Polynomial fitting determines the polynomial coefficients through the least squares method to minimize the error between the fitted curve and the spectral contour feature points, and spline interpolation ensures the smoothness and continuity of the envelope curve through piecewise cubic polynomials. Spectral envelope modeling considers both the static features and dynamic features of the spectrum, with static features describing the distribution shape of the spectrum at a specific time, and dynamic features describing the variation trend and speed of the spectrum over time. The voice verification code spectral envelope parameters include coefficient, peak position, valley position, slope, symmetry, and other mathematical feature parameters, which completely describe the spectral envelope characteristics of the voice verification code.
[0102] The spectral envelope parameter spectrum encoding matrix generation process converts the mathematical description of the spectral envelope into a digital encoding format suitable for storage and transmission. Spectrum encoding matrix generation is the application of data compression and encoding techniques in speech processing, which organizes and quantizes spectral envelope parameters in matrix form to achieve efficient data representation. The generation process first normalizes and quantizes the spectral envelope parameters, mapping continuous parameter values to discrete digital codes, with the quantization precision balanced according to the voice quality requirements and storage space limitations. The rows of the encoding matrix correspond to different envelope feature parameter types, and the columns correspond to time series or frequency series. The numerical value of the matrix element represents the quantized encoding value of the parameter. The encoding process uses a combination of difference encoding and entropy encoding compression strategies. Difference encoding reduces data redundancy by utilizing the correlation between adjacent parameter values, and entropy encoding performs variable-length encoding based on the statistical distribution of parameter values. The television end multi-modal voice verification code realizes unified encoding representation of voice verification code content, user identity features, and acoustic environment information in the form of a spectral encoding matrix, which contains spectral features of the voice verification code, multi-modal identity of the user, and environment adaptation information of the television end.
[0103] The above describes the voice verification code generation method for television end multi-modal identity authentication in the embodiments of the present application. The voice verification code generation system for television end multi-modal identity authentication in the embodiments of the present application is described below. Please refer to Figure 2 An embodiment of the voice verification code generation system for television end multi-modal identity authentication in the embodiments of the present application includes:
[0104] A compensation module for performing acoustic compensation processing on the living room environment acoustic signal collected by the television end to obtain acoustic calibration parameters.
[0105] The extraction module is configured to perform voiceprint feature extraction processing on the user voice to obtain a user voiceprint template, and perform spatial positioning processing on a user voice position to obtain a spatial position feature vector.
[0106] The authentication module is configured to perform multi-modal identity authentication processing on the user voiceprint template and the spatial position feature vector to obtain a user identity feature code.
[0107] The synthesis module is configured to perform personalized speech synthesis processing on a verification code sequence according to the user identity feature code and the voice school correction parameter to obtain a user-specific speech verification code audio.
[0108] The encoding module is configured to perform spectral encoding processing on the user-specific speech verification code audio to obtain a television-side multi-modal speech verification code.
[0109] The binding module is configured to perform association and binding processing on the television-side multi-modal speech verification code and the user identity feature code to obtain a television-side identity authentication credential.
[0110] The above Figure 2 The television-side multi-modal identity authentication speech verification code generation system in the embodiment of the application is described in detail from the perspective of a modular functional entity, and the television-side multi-modal identity authentication speech verification code generation device in the embodiment of the application is described in detail from the perspective of hardware processing.
[0111] Referring to Figure 3 In the embodiment of the application, a television-side multi-modal identity authentication speech verification code generation device is also provided, which can be a server, and the internal structure thereof can be as shown in Figure 3 The television-side multi-modal identity authentication speech verification code generation device includes a processor, a memory, a display screen, an input device, a network interface and a database connected through a system bus. The processor of the computer is configured to provide computing and control capabilities. The memory of the television-side multi-modal identity authentication speech verification code generation device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operating system and the computer program in the non-volatile storage medium to run. The database of the television-side multi-modal identity authentication speech verification code generation device is configured to store corresponding data in the embodiment. The network interface of the television-side multi-modal identity authentication speech verification code generation device is configured to communicate with an external terminal through a network connection. The computer program is executed by the processor to implement the above method.
[0112] Those skilled in the art can understand Figure 3The structure shown in the figure is only a block diagram of part of the structure related to the scheme of the present application, and does not constitute a limitation on the television terminal multi-modal identity authentication voice verification code generation device to which the scheme of the present application is applied.
[0113] The present application also provides a computer readable storage medium, which can be a non-volatile computer readable storage medium, and can also be a volatile computer readable storage medium, and the computer readable storage medium stores instructions, and when the instructions run on a computer, the computer executes the steps of the television terminal multi-modal identity authentication voice verification code generation method.
[0114] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described system, system and unit can refer to the corresponding processes in the foregoing method embodiments, which will not be described here.
[0115] The integrated unit, if realized in the form of a software function unit and sold or used as an independent product, can be stored in a computer readable storage medium. Based on such understanding, the technical scheme of the present application or the whole or part of the technical scheme that essentially contributes to the prior art can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a plurality of instructions for causing a television terminal multi-modal identity authentication voice verification code generation device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The foregoing storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and various program code storage media.
[0116] The above embodiments are only used to illustrate the technical scheme of the present application, rather than limit it; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that they can still modify the technical scheme recorded in the foregoing embodiments, or make equivalent replacement for part of the technical features; and these modifications or replacements do not make the essence of the corresponding technical scheme deviate from the spirit and scope of the technical scheme of each embodiment of the present application.
Claims
1. A method for generating voice verification codes for multimodal identity authentication on a television terminal, characterized in that, The method includes: Step S1: Perform acoustic compensation processing on the acoustic signals of the living room environment collected by the TV to obtain acoustic correction parameters; Step S2: Extract voiceprint features from the user's voice to obtain the user's voiceprint template, and simultaneously perform spatial localization processing on the user's voice location to obtain the spatial location feature vector. Step S3: Perform multimodal identity authentication processing on the user voiceprint template and the spatial location feature vector to obtain the user identity feature code, including: performing data standardization processing on the user voiceprint template to obtain standardized voiceprint features; performing fully connected network mapping processing on the spatial location feature vector to obtain spatial location encoding; performing weight allocation processing on the standardized voiceprint features and the spatial location encoding based on the attention weight algorithm to obtain a fused feature vector; and performing identity classifier recognition processing on the fused feature vector to obtain the user identity feature code. The user identity feature code is generated by encoding the user category corresponding to the highest probability with its confidence level. Step S4: Perform personalized speech synthesis processing on the user-specific verification code sequence according to the user identity feature code and the acoustic correction parameters to obtain the user-specific voice verification code audio. In this step, perform personalized parameter mapping processing on the random verification code sequence according to the user identity feature code to obtain the user-specific verification code sequence. Step S5: Perform spectrum encoding on the user-specific voice verification code audio to obtain a multimodal voice verification code for the TV. Step S6: Associate and bind the TV-side multimodal voice verification code with the user identity feature code to obtain the TV-side identity authentication credential.
2. The voice verification code generation method for multimodal identity authentication on a television terminal according to claim 1, characterized in that, Step S1 includes: Acoustic signals from the living room environment are collected and processed using the TV's built-in microphone array to obtain raw environmental acoustic data. The original environmental acoustic data is processed to extract reverberation time and sound propagation delay to obtain the acoustic characteristic parameters of the living room; The acoustic distance of the living room acoustic feature parameters is calculated based on the radial basis function to obtain the spatial distance value of the sound source. The acoustic compensation coefficient is obtained by calculating the compensation coefficient of the frequency response function based on the spatial distance value of the sound source. The acoustic compensation coefficient is multiplied by the reference gain parameter to obtain the acoustic correction parameter.
3. The method for generating voice verification codes for multimodal identity authentication on a television terminal according to claim 1, characterized in that, Step S2 includes: The user's spoken voice is processed by fundamental frequency extraction to obtain a fundamental frequency trajectory sequence; Based on the fundamental frequency trajectory sequence, first-order and second-order difference calculations are performed to obtain fundamental frequency change direction data; The fundamental frequency change direction data is subjected to 8-direction chain code quantization to obtain the user voiceprint template; The time difference of arrival of the user's voice is calculated by using a TV microphone array to obtain three-dimensional spatial coordinate data. Based on the three-dimensional spatial coordinate data, the horizontal angle, vertical angle, and distance are layered to obtain the spatial location feature vector.
4. The method for generating voice verification codes for multimodal identity authentication on a television terminal according to claim 1, characterized in that, Step S4 includes: The user-specific verification code sequence is input into a speech synthesis network for text-to-speech processing to obtain the initial voice verification code audio. Based on the acoustic correction parameters, the initial voice verification code audio is adapted to the living room environment to obtain the acoustically compensated voice verification code audio. The acoustically compensated voice verification code audio is combined with the user identity feature code to perform audio fingerprint embedding processing to obtain the user-specific voice verification code audio.
5. The method for generating voice verification codes for multimodal identity authentication on a television terminal according to claim 4, characterized in that, The step of inputting the user-specific verification code sequence into a speech synthesis network for text-to-speech processing to obtain the initial voice verification code audio includes: The user-specific verification code sequence is input into the encoder network for text feature encoding to obtain the verification code text feature vector. The uniqueness constraint processing of the CAPTCHA text feature vector is performed based on the compression coding mechanism to obtain the coding constraint features; The encoded constraint features are input into the generator network for adversarial generation processing to obtain a speech spectrum representation. The speech spectrum representation is subjected to multi-scale time-frequency domain discrimination processing based on a discriminator network to obtain the authenticity verification result. Based on the authenticity verification result, the speech spectrum representation is subjected to inverse Mel-spectrum transform to obtain the initial speech verification code audio. The initial voice verification code audio is processed using the far-field voice characteristics of the TV terminal to obtain the voice verification code audio.
6. The method for generating voice verification codes for multimodal identity authentication on a television terminal according to claim 1, characterized in that, Step S5 includes: The user-specific voice verification code audio is transformed from the time domain to the frequency domain to obtain the voice verification code spectrum data. Based on the frequency domain adaptive filtering algorithm, the spectral data of the voice verification code is analyzed and processed in 12 key frequency bands to obtain spectral contour features. The spectral contour features are subjected to spectral envelope modeling to obtain the spectral envelope parameters of the voice verification code. The multimodal voice verification code for the TV terminal is obtained by generating a spectral coding matrix based on the spectral envelope parameters of the voice verification code.
7. A voice verification code generation system for multimodal identity authentication on a television terminal, characterized in that, The voice verification code generation method for implementing multimodal identity authentication on a television terminal as described in any one of claims 1-6, wherein the voice verification code generation system for multimodal identity authentication on a television terminal comprises: The compensation module is used to perform acoustic compensation processing on the acoustic signals of the living room environment collected by the TV to obtain acoustic correction parameters; The extraction module is used to extract voiceprint features from the user's voice to obtain the user's voiceprint template, and at the same time, to perform spatial positioning processing on the user's voice position to obtain the spatial position feature vector. The authentication module is used to perform multimodal identity authentication processing on the user voiceprint template and the spatial location feature vector to obtain the user identity feature code; The synthesis module is used to perform personalized speech synthesis processing on the verification code sequence based on the user identity feature code and the acoustic correction parameters to obtain user-exclusive voice verification code audio. The encoding module is used to perform spectrum encoding processing on the user-specific voice verification code audio to obtain a multimodal voice verification code for the TV. The binding module is used to associate and bind the TV terminal multimodal voice verification code with the user identity feature code to obtain the TV terminal identity authentication credential.
8. A voice verification code generation device for multimodal identity authentication on a television terminal, characterized in that, The device includes a memory and a processor, wherein the memory stores a computer program that can run on the processor, and the processor executes the computer program to implement the voice verification code generation method for multimodal identity authentication on a television terminal as described in any one of claims 1 to 6.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is run by the processor, it causes the processor to execute the voice verification code generation method for multimodal identity authentication on a television terminal as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Vocal print secret key generating method and device and logging-in method and system based on vocal print secret key
CN103973453A
Verification code implementation method and device based on semantic voiceprint interaction mode
CN115294993A