Flexible throat voice interaction method, system, device and medium for anti-motion interference
Patent Information
- Application Number
- CN202610878687.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-17
- Publication Date
- 2026-09-04
AI Technical Summary
[0004]然而,上述基于材料形变原理的传统接触式传感方法,在用户处于行走、奔跑或头部转动等自然活动状态时,皮肤因肢体运动产生的大幅度拉伸形变会与微弱的声带振动信号在传感器输出中发生混叠,形成高振幅的运动伪影
[0059] The aforementioned flexible laryngeal voice interaction method, system, device, and medium for resisting motion interference acquire the user's original triaxial acceleration data. This data characterizes the three-dimensional mechanical vibration of the user's laryngeal skin during phonation. Noise reduction is applied to the original triaxial acceleration data to obtain clean triaxial vocal cord vibration data. Time-frequency domain features are extracted from the clean triaxial vocal cord vibration data to obtain multi-channel Mel-spectrum data. This multi-channel Mel-spectrum data is then input into an identity command verification convolutional neural network model to obtain voice command recognition results and user authentication results. The identity command verification convolutional neural network model has a dual-branch output structure. This achieves the technical effect of accurately reconstructing the vocal content from mixed interference signals and simultaneously identifying the speaker's identity in a dynamic user motion environment, thereby significantly improving the reliability and accuracy of flexible laryngeal voice interaction technology in complex scenarios.
Smart Images

Figure CN122696005A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of voice interaction, and in particular relates to a flexible laryngeal voice interaction method, system, device and medium for resisting motion interference. Background Technology
[0002] With the deep development of wearable devices and artificial intelligence technology, a laryngeal voice interaction technology with a contact-type flexible skin sensor as its core has emerged. This technology directly senses the vibration of sound by attaching to the skin of the throat, effectively avoiding airborne acoustic noise interference and achieving private and noise-resistant voice signal acquisition.
[0003] Traditionally, contact-type laryngeal speech sensors primarily rely on material deformation principles such as resistivity, capacitance, or triboelectricity for signal acquisition. When a user speaks, the subtle vibrations of the laryngeal skin surface cause stretching or compression of the sensor's sensitive material, resulting in changes in electrical parameters such as resistance, capacitance, or triboelectric charge. By continuously recording the waveforms of these electrical parameters over time, a sensing signal characterizing the vocal content can be obtained. Based on this, time-frequency domain features are extracted from the signal and input into a pre-trained speech recognition model to map laryngeal skin vibrations to speech commands.
[0004] However, traditional contact sensing methods based on material deformation principles suffer from high-amplitude motion artifacts in the sensor output when users are in natural activities such as walking, running, or head turning. These artifacts, with amplitudes much larger than the effective speech signal in the time domain, completely obscure speech features, causing speech recognition models to fail in dynamic scenarios. This results in insufficient reliability and accuracy in the application of flexible laryngeal voice interaction technology. Summary of the Invention
[0005] Therefore, it is necessary to provide a flexible laryngeal voice interaction method, system, device, and medium that can improve the reliability and accuracy of flexible laryngeal voice interaction technology and resist motion interference, in order to address the above-mentioned technical problems.
[0006] In a first aspect, this application provides a flexible laryngeal speech interaction method for resisting motion interference, comprising:
[0007] Obtain the user's raw triaxial acceleration data; the raw triaxial acceleration data is used to characterize the three-dimensional mechanical vibration of the user's laryngeal skin during vocalization;
[0008] The raw triaxial acceleration data was denoised to obtain clean triaxial vocal cord vibration data.
[0009] Time-frequency domain features were extracted from pure triaxial vocal cord vibration data to obtain multi-channel Mel spectrogram data;
[0010] Multi-channel Mel-spectrum data is input into the identity command verification convolutional neural network model to obtain the voice command recognition result and the user identity authentication result; the identity command verification convolutional neural network model has a two-branch output structure.
[0011] Furthermore, the raw triaxial acceleration data includes raw horizontal axis acceleration data, raw vertical axis acceleration data, and raw vertical axis acceleration data. Noise reduction processing is performed on the raw triaxial acceleration data to obtain clean triaxial vocal cord vibration data, including:
[0012] Perform a Fourier transform on the original horizontal axis acceleration data to obtain the frequency domain horizontal axis acceleration data; perform a Fourier transform on the original vertical axis acceleration data to obtain the frequency domain vertical axis acceleration data; perform a Fourier transform on the original vertical axis acceleration data to obtain the frequency domain vertical axis acceleration data.
[0013] The denoised frequency domain horizontal axis acceleration data and vocal cord vibration frequency domain window function are multiplied to obtain the denoised frequency domain horizontal axis data; the denoised frequency domain vertical axis acceleration data and vocal cord vibration frequency domain window function are multiplied to obtain the denoised frequency domain vertical axis data; the denoised frequency domain vertical axis acceleration data and vocal cord vibration frequency domain window function are multiplied to obtain the denoised frequency domain vertical axis data.
[0014] Perform an inverse Fourier transform on the denoised frequency domain horizontal axis data to obtain clean horizontal axis vibration data; perform an inverse Fourier transform on the denoised frequency domain vertical axis data to obtain clean vertical axis vibration data; perform an inverse Fourier transform on the denoised frequency domain vertical axis data to obtain clean vertical axis vibration data.
[0015] Pure triaxial vocal cord vibration data is constructed based on pure horizontal axis vibration data, pure vertical axis vibration data, and pure vertical axis vibration data.
[0016] Furthermore, time-frequency domain feature extraction was performed on the pure triaxial vocal tract vibration data to obtain multi-channel Mel spectrogram data, including:
[0017] Based on the preset segmentation interval parameters, the pure horizontal axis vibration data is segmented to obtain the horizontal axis vibration segmented dataset; based on the segmentation interval parameters, the pure vertical axis vibration data is segmented to obtain the vertical axis vibration segmented dataset; based on the segmentation interval parameters, the pure vertical axis vibration data is segmented to obtain the vertical axis vibration segmented dataset.
[0018] Short-time Fourier transforms are performed on the horizontal axis vibration segment data in the horizontal axis vibration segment dataset to obtain the horizontal axis spectrum sequence; short-time Fourier transforms are performed on the vertical axis vibration segment data in the vertical axis vibration segment dataset to obtain the vertical axis spectrum sequence; short-time Fourier transforms are performed on the vertical axis vibration segment data in the vertical axis vibration segment dataset to obtain the vertical axis spectrum sequence.
[0019] Based on the preset Mel filter, Mel mapping is performed on each horizontal axis spectrum sequence to obtain the horizontal axis Mel spectrum; based on the Mel filter, Mel mapping is performed on each vertical axis spectrum sequence to obtain the vertical axis Mel spectrum; based on the Mel filter, Mel mapping is performed on each vertical axis spectrum sequence to obtain the vertical axis Mel spectrum.
[0020] The horizontal axis Mel spectrum, vertical axis Mel spectrum, and vertical axis Mel spectrum are stacked to obtain multi-channel Mel spectrum data.
[0021] Furthermore, the identity instruction verification convolutional neural network model includes an input layer, a feature extraction layer, a global pooling layer, a fully connected layer for instruction recognition, a fully connected layer for identity authentication, and a result output layer, wherein:
[0022] The input layer receives multi-channel Mel spectrogram data and transmits it to the feature extraction layer.
[0023] The feature extraction layer is used to perform two-dimensional convolution on the multi-channel Mel spectrogram data, extracting local time-frequency texture features and global acoustic semantic features from the multi-channel Mel spectrogram data step by step to obtain compressed feature maps, and then transmitting the compressed feature maps to the global pooling layer;
[0024] The global pooling layer is used to perform global average pooling on the compressed feature map to obtain acoustic feature vectors, and then transmits the acoustic feature vectors to the instruction recognition fully connected layer and the identity authentication fully connected layer.
[0025] The instruction recognition fully connected layer is used to map acoustic feature vectors to obtain instruction category probability data, and then transmit the instruction category probability data to the result output layer;
[0026] The fully connected layer for identity authentication is used to map acoustic feature vectors to obtain the probability data of legitimate users and the probability data of illegitimate users, and then transmit the probability data of legitimate users and the probability data of illegitimate users to the result output layer;
[0027] The result output layer is used to obtain the voice command recognition result based on the command category probability data, and to obtain the user identity authentication result based on the legitimate user probability data and the illegitimate user probability data.
[0028] Furthermore, the identity instruction verification convolutional neural network model is obtained through the following steps:
[0029] Obtain the preset historical dataset; the historical dataset includes historical multi-channel Mel spectrogram data and historical instruction category labels and historical identity labels corresponding to the historical multi-channel Mel spectrogram data;
[0030] The input layer, feature extraction layer, global pooling layer, instruction recognition fully connected layer, identity authentication fully connected layer, and result output layer are connected to obtain the basic neural network framework. The basic neural network framework is then initialized to obtain the initial neural network.
[0031] The historical multi-channel Mel spectrogram data from the historical dataset are input into the initial neural network to obtain the predicted instruction category data and predicted identity data of the historical multi-channel Mel spectrogram data;
[0032] Based on the predicted instruction category data and historical instruction category labels, the instruction recognition loss value is calculated; based on the predicted identity data and historical identity labels, the identity authentication loss value is calculated.
[0033] Based on the instruction recognition loss value and the identity authentication loss value, the joint loss value is calculated, and the expression for the joint loss value is:
[0034]
[0035] In the formula, It is an index to any historical multi-channel Mel spectrogram data. It is the first The joint loss value of historical multi-channel Mel spectrogram data, It is the instruction loss weight. It is the first Command recognition loss value of historical multi-channel Mel spectrogram data. It is the weight of identity loss. It is the first The authentication loss value of historical multi-channel Mel spectrogram data;
[0036] With the goal of minimizing the joint loss, the initial neural network is iteratively trained until the preset convergence condition is met, thus obtaining the identity instruction verification convolutional neural network model.
[0037] Furthermore, the expression for the frequency domain window function of vocal tract vibration is:
[0038]
[0039] In the formula, It is any frequency value. It is a frequency domain window function for vocal cord vibration. It is the minimum frequency of vocal cord vibration. It is the maximum frequency of vocal cord vibration;
[0040] The minimum frequency of vocal cord vibration is obtained through the following steps:
[0041] Acquire the user's silent motion triaxial acceleration data and silent stationary triaxial acceleration data;
[0042] Fourier transform is performed on the three-axis acceleration data of silent motion to obtain silent motion spectrum data; Fourier transform is performed on the three-axis acceleration data of silent stationary motion to obtain silent stationary spectrum data; the silent motion spectrum data includes silent motion horizontal axis spectrum data, silent motion vertical axis spectrum data, and silent motion vertical axis spectrum data; the silent stationary spectrum data includes silent stationary horizontal axis spectrum data, silent stationary vertical axis spectrum data, and silent stationary vertical axis spectrum data.
[0043] The silent motion energy data is obtained by superimposing the horizontal, vertical, and lateral axis spectral data. The expression for the silent motion energy data is as follows:
[0044]
[0045] In the formula, It is any frequency value. It is the frequency value in silent motion energy data. The value of kinetic energy intensity. The frequency values in the horizontal axis spectrum data of silent motion The transverse frequency intensity, The frequency values in the vertical axis spectrum data of silent motion The vertical axis frequency intensity, The frequency values in the vertical axis spectrum data of silent motion The vertical axis frequency intensity;
[0046] The silent static horizontal axis spectrum data, silent static vertical axis spectrum data, and silent static vertical axis spectrum data are superimposed to obtain the silent static energy data. The expression for the silent static energy data is as follows:
[0047]
[0048] In the formula, It is any frequency value. It is the frequency value in silent static energy data. The static energy intensity value, It is the frequency value in the silent, stationary horizontal axis spectrum data. The transverse frequency intensity, It is the frequency value in the silent, stationary vertical axis spectrum data. The vertical axis frequency intensity, It is the frequency value in the silent, stationary vertical axis spectrum data. The vertical axis frequency intensity;
[0049] The static energy intensity value with the largest value in the silent static energy data is determined as the noise threshold, and the motion energy intensity value of each frequency value in the silent motion energy data is compared with the noise threshold to obtain the noise comparison result of each frequency value.
[0050] The frequency values whose motion energy intensity value is greater than the noise threshold are selected from the noise comparison results to form a noise range set, and the frequency value with the largest value in the noise range set is determined as the minimum value of vocal cord vibration frequency.
[0051] Furthermore, the sampling frequency of the original triaxial acceleration data is 1000Hz, and the original triaxial acceleration data is acquired by a flexible sensor. The flexible sensor is an island-bridge heterogeneous integrated architecture, in which the bridging part is made of liquid metal ink doped with silicon dioxide, and the mass fraction of silicon dioxide is 6wt%. The maximum value of the vocal cord vibration frequency is 255Hz, and the minimum value of the vocal cord vibration frequency can also be directly set to 80Hz.
[0052] Secondly, this application also provides a flexible laryngeal voice interaction system for resisting motion interference, comprising:
[0053] The data acquisition module is used to acquire the user's raw triaxial acceleration data; the raw triaxial acceleration data is used to characterize the three-dimensional mechanical vibration of the user's laryngeal skin during vocalization;
[0054] The noise reduction module is used to denoise the raw triaxial acceleration data to obtain clean triaxial vocal cord vibration data.
[0055] The feature extraction module is used to extract time-frequency domain features from pure triaxial vocal cord vibration data to obtain multi-channel Mel spectrogram data;
[0056] The results analysis module is used to input multi-channel Mel-spectrum data into the identity command verification convolutional neural network model to obtain the voice command recognition results and user identity authentication results; the identity command verification convolutional neural network model has a dual-branch output structure.
[0057] Thirdly, this application also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement any of the flexible laryngeal speech interaction methods for resisting motion interference described in the first aspect of this application.
[0058] Fourthly, this application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements any of the flexible laryngeal speech interaction methods for resisting motion interference described in the first aspect of this application.
[0059] The aforementioned flexible laryngeal voice interaction method, system, device, and medium for resisting motion interference acquire the user's original triaxial acceleration data. This data characterizes the three-dimensional mechanical vibration of the user's laryngeal skin during phonation. Noise reduction is applied to the original triaxial acceleration data to obtain clean triaxial vocal cord vibration data. Time-frequency domain features are extracted from the clean triaxial vocal cord vibration data to obtain multi-channel Mel-spectrum data. This multi-channel Mel-spectrum data is then input into an identity command verification convolutional neural network model to obtain voice command recognition results and user authentication results. The identity command verification convolutional neural network model has a dual-branch output structure. This achieves the technical effect of accurately reconstructing the vocal content from mixed interference signals and simultaneously identifying the speaker's identity in a dynamic user motion environment, thereby significantly improving the reliability and accuracy of flexible laryngeal voice interaction technology in complex scenarios. Attached Figure Description
[0060] To more clearly illustrate the technical solutions in the embodiments or related technologies of this application, the accompanying drawings used in the description of the embodiments or related technologies will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0061] Figure 1 A flowchart illustrating a flexible laryngeal voice interaction method for resisting motion interference, provided as an embodiment of this application;
[0062] Figure 2 A schematic diagram of the structure of a flexible laryngeal voice interaction system for resisting motion interference, provided as an embodiment of this application;
[0063] Figure 3 This is a schematic diagram of the structure of a computer device for a flexible laryngeal voice interaction method for resisting motion interference, provided as an embodiment of this application. Detailed Implementation
[0064] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0065] In one embodiment, such as Figure 1As shown, a flexible laryngeal voice interaction method for resisting motion interference is provided. This embodiment illustrates the application of this method to an interactive terminal. It is understood that this method can also be applied to a server, and also to a system including an interactive terminal and a server, and implemented through the interaction between the interactive terminal and the server. In this embodiment, the method includes the following S101-S104, wherein:
[0066] S101, acquire the user's raw triaxial acceleration data; the raw triaxial acceleration data is used to characterize the three-dimensional mechanical vibration of the user's throat skin during vocalization.
[0067] Specifically, the interactive terminal acquires raw triaxial acceleration data of the user's throat. The mathematical form of this raw triaxial acceleration data can be three time-varying one-dimensional time-series vector signals, which can be represented as follows: .in, , and These are the original horizontal axis acceleration data, the original vertical axis acceleration data, and the original vertical axis acceleration data. The expression can be , For discrete sampling timestamps, In the sampling timestamp The user's throat horizontal axis acceleration value, in m / s², is displayed. The original vertical axis acceleration data and the original vertical axis acceleration data are similar in format to the original horizontal axis acceleration data. , , It is a sampling timestamp User's throat longitudinal axis acceleration value It is a sampling timestamp The user's throat vertical axis acceleration value. The raw triaxial acceleration data is a mixture of weak high-frequency vibration information caused by the user's vocalization and large-amplitude low-frequency acceleration impacts caused by the user's body movements (such as walking, running, or turning the head).
[0068] S102 performs noise reduction processing on the original triaxial acceleration data to obtain clean triaxial vocal cord vibration data.
[0069] Specifically, the interactive terminal performs denoising processing on the raw triaxial acceleration data. The purpose of denoising is to utilize the difference in energy distribution between vocalization and limb movement in the frequency domain. Vocal cord vibration energy caused by vocalization is relatively concentrated in the higher frequency range, while motion artifact energy caused by walking, head turning, and other limb movements is mainly concentrated in the lower frequency range. By selectively filtering the signal in the frequency domain, low-frequency motion artifact energy and high-frequency electronic noise energy are removed from the mixed signal, retaining only the effective frequency band components representing vocal cord vibration, resulting in clean triaxial vocal cord vibration data. The clean triaxial vocal cord vibration data can be represented as follows: ,in, It is pure horizontal axis vibration data. It is pure longitudinal axis vibration data. These are pure vertical axis vibration data. The expression can be , For discrete sampling timestamps, In the sampling timestamp The user's throat acceleration value after noise reduction processing, the unit can be m / s². and Form and Similar, that is , , and These are the vertical and longitudinal acceleration values after noise reduction. This pure triaxial vocal cord vibration data can retain the effective vibration information representing the vocal content to the greatest extent, while significantly suppressing the interference of user motion on the signal.
[0070] S103, time-frequency domain feature extraction is performed on pure triaxial vocal cord vibration data to obtain multi-channel Mel spectrum data.
[0071] Specifically, the interactive terminal performs time-frequency domain feature extraction on the pure triaxial vocal cord vibration data. The purpose of this process is to convert the original one-dimensional time-series vibration signal into a two-dimensional feature representation that simultaneously contains information on temporal dynamic changes and frequency-domain energy distribution. This representation is then further mapped onto a Mel scale that conforms to the characteristics of human auditory perception, thereby extracting acoustic features that are highly correlated with the vocal content and robust to noise. During processing, the interactive terminal can independently execute the same feature extraction pipeline on the pure horizontal axis vibration data, pure vertical axis vibration data, and pure vertical shaft axis vibration data, generating three single-channel Mel spectrograms, each reflecting the vibration characteristics of that axis. Finally, these three single-channel Mel spectrograms are stacked and fused along the channel dimension to generate a three-dimensional feature data with a multi-channel structure, i.e., multi-channel Mel spectrogram data. The mathematical form of this multi-channel Mel spectrogram data can be represented as a three-dimensional tensor. ,in It is a time frame dimension, and its value is determined by the sampling timestamp and the preset segmentation parameters. Each frame corresponds to a spectrum snapshot at a specific moment. The Mel frequency dimension reflects the energy intensity at different perceived frequency scales; The channel dimension corresponds to vibration information in the horizontal, vertical, and axial directions, respectively. Multi-channel Mel-spectrum data comprehensively expresses the vibration patterns of the laryngeal skin in three-dimensional space during vocalization.
[0072] S104, input the multi-channel Mel spectrogram data into the identity command verification convolutional neural network model to obtain the voice command recognition result and the user identity authentication result; the identity command verification convolutional neural network model is a two-branch output structure.
[0073] Specifically, the identity command verification convolutional neural network model can be pre-trained using historical multi-channel Mel spectrogram data, along with historical command category labels and historical identity labels corresponding to each segment of historical multi-channel Mel spectrogram data. The value range of the historical command category label is determined by a preset command set, where each command is assigned a unique command category label, such as "I need water" or "Please help me turn over." The historical identity label identifies which legitimate or illegitimate user the segment of historical multi-channel Mel spectrogram data originated from. The interactive terminal inputs the multi-channel Mel spectrogram data into the pre-trained identity command verification convolutional neural network model to obtain the voice command recognition result and the user identity authentication result. The identity command verification convolutional neural network model has a dual-branch output structure. The voice command recognition result is a command category label determination corresponding to the current input, indicating which command in the predefined command set the current user's voice content belongs to. The user identity authentication result is an identity determination corresponding to the current input, indicating whether the current speaker is a user from the preset set of legitimate users or an unauthorized illegitimate user.
[0074] This embodiment provides a flexible laryngeal voice interaction method for resisting motion interference. It acquires raw acceleration data characterizing the three-dimensional mechanical vibration of the laryngeal skin, denoises it using the difference in energy distribution between human movement and vocalization in the frequency domain, obtaining clean three-axis vocal cord vibration data. Multi-channel Mel-spectrum data is then generated through time-frequency domain feature extraction, and finally jointly analyzed by an identity command verification convolutional neural network model to obtain voice command recognition results and user identity authentication results. This method enables accurate reconstruction of vocal content from mixed interference signals and simultaneous speaker identification in dynamic user environments, significantly improving the reliability and accuracy of flexible laryngeal voice interaction technology in complex scenarios.
[0075] In one embodiment, the raw triaxial acceleration data includes raw horizontal axis acceleration data, raw vertical axis acceleration data, and raw vertical axis acceleration data. The raw triaxial acceleration data is then denoised to obtain clean triaxial vocal cord vibration data, including:
[0076] S201. Perform a Fourier transform on the original horizontal axis acceleration data to obtain the frequency domain horizontal axis acceleration data; perform a Fourier transform on the original vertical axis acceleration data to obtain the frequency domain vertical axis acceleration data; perform a Fourier transform on the original vertical axis acceleration data to obtain the frequency domain vertical axis acceleration data.
[0077] Specifically, the interactive terminal performs a Fourier transform on each axis of the original triaxial acceleration data. The Fourier transform is a mathematical transformation method that converts a time-domain signal to a frequency-domain representation. It is used to decompose a time-varying acceleration waveform into a series of sinusoidal wave components of different frequencies, obtaining the amplitude and phase information of each frequency component. Taking the original horizontal axis acceleration data as an example, after performing a Fourier transform on this sequence, the acceleration values, originally with the sampling timestamp as the independent variable, are converted into a spectral representation with the frequency value as the independent variable, resulting in the frequency-domain horizontal axis acceleration data, which can be represented as follows: ,in It is a frequency value, in Hz. modulus Characterized by the frequency value The energy intensity of the original horizontal axis acceleration data. Similarly, the energy intensity of the original vertical axis acceleration data. and the original vertical axis acceleration data Perform Fourier transforms on each component to obtain the corresponding frequency domain vertical axis acceleration data. and frequency domain vertical axis acceleration data .
[0078] S202: Multiply the frequency domain horizontal axis acceleration data and the vocal cord vibration frequency domain window function to obtain the denoised frequency domain horizontal axis data; multiply the frequency domain vertical axis acceleration data and the vocal cord vibration frequency domain window function to obtain the denoised frequency domain vertical axis data; multiply the frequency domain vertical axis acceleration data and the vocal cord vibration frequency domain window function to obtain the denoised frequency domain vertical axis data.
[0079] Specifically, the interactive terminal multiplies the axial frequency domain acceleration data obtained in step S201 with a preset vocal tract vibration frequency domain window function. The vocal tract vibration frequency domain window function is a pre-defined real-valued function defined on the frequency axis; its values at different frequencies determine whether the corresponding frequency components are preserved or suppressed. The vocal tract vibration frequency domain window function can use the minimum and maximum vocal tract vibration frequencies as passband boundaries. Signal components with frequencies within this passband are fully preserved, while signal components with frequencies outside this passband are filtered out. Through the above multiplication operation, motion artifact energy in the low-frequency band and electronic noise energy in the high-frequency band are suppressed or eliminated in the frequency domain, while effective frequency components closely related to vocal tract vibration within the passband are fully preserved. Taking the horizontal axis as an example, the frequency domain horizontal axis acceleration data... With the frequency domain window function of vocal cord vibration At each frequency value Multiply the above to obtain the denoised frequency domain horizontal axis data. Perform the same operation on both the vertical and directional axes to obtain the denoised frequency domain data for the vertical axis. and denoised frequency domain vertical axis data .
[0080] S203. Perform an inverse Fourier transform on the denoised frequency domain horizontal axis data to obtain clean horizontal axis vibration data; perform an inverse Fourier transform on the denoised frequency domain vertical axis data to obtain clean vertical axis vibration data; perform an inverse Fourier transform on the denoised frequency domain vertical axis data to obtain clean vertical axis vibration data.
[0081] Specifically, the interactive terminal performs an inverse Fourier transform on each axial denoised frequency domain data obtained in step S202. The inverse Fourier transform is a mathematical transformation method that is the inverse of the Fourier transform; its function is to reassemble the denoised frequency domain representation into a time-domain waveform signal. Taking the horizontal axis as an example, the denoised frequency domain horizontal axis data... Performing an inverse Fourier transform reconstructs the time-domain acceleration waveform by superimposing the sinusoidal components of each frequency within the retained sound frequency band, yielding pure transverse axis vibration data. The mathematical form of pure transverse axis vibration data can be expressed as: , For discrete sampling timestamps, For sampling timestamp The user's laryngeal acceleration values after denoising. Compared to the original laryngeal acceleration data, this sequence no longer contains large-amplitude low-frequency fluctuations caused by limb movement; the waveform only reflects changes in skin acceleration caused by vocal cord vibration. Similarly, performing inverse Fourier transforms on the denoised frequency domain vertical axis data and the denoised frequency domain vertical axis data respectively yields the clean vertical axis vibration data. and pure vertical axis vibration data .
[0082] S204 is composed of pure triaxial vocal cord vibration data based on pure horizontal axis vibration data, pure vertical axis vibration data, and pure vertical axis vibration data.
[0083] Specifically, the interactive terminal integrates the pure horizontal axis vibration data, pure vertical axis vibration data, and pure vertical shaft axis vibration data obtained in step S203 to obtain pure triaxial vocal cord vibration data. The mathematical form of this pure triaxial vocal cord vibration data can be expressed as follows: It contains noise-reduced laryngeal skin vibration acceleration information in three spatial orthogonal directions, and fully records the three-dimensional mechanical response of vocal cord vibration on the skin surface during the phonation process.
[0084] This embodiment provides a flexible laryngeal voice interaction method for resisting motion interference. It transforms the time-domain signals of each axis in the original triaxial acceleration data to the frequency domain, uses a preset vocal cord vibration frequency-domain window function to suppress the frequency components corresponding to motion artifacts and electronic noise in the frequency domain, and then inversely transforms the purified frequency-domain data back to the time domain, finally combining them to obtain pure triaxial vocal cord vibration data. Based on the difference in frequency domain distribution between human movement and vocalization, this method removes motion interference from the mixed signal while preserving complete vocalization information, providing high-quality input data for subsequent time-frequency domain feature extraction and deep learning recognition, effectively improving the reliability of laryngeal voice interaction in dynamic scenes.
[0085] In one embodiment, time-frequency domain feature extraction is performed on the pure triaxial vocal tract vibration data to obtain multi-channel Mel-spectrum data, including:
[0086] S301, based on the preset segmentation interval parameter, the pure horizontal axis vibration data is segmented to obtain the horizontal axis vibration segmented dataset; based on the segmentation interval parameter, the pure vertical axis vibration data is segmented to obtain the vertical axis vibration segmented dataset; based on the segmentation interval parameter, the pure vertical axis vibration data is segmented to obtain the vertical axis vibration segmented dataset.
[0087] Specifically, the preset segmentation interval parameter controls the configuration parameters of the segmentation operation. This parameter includes at least two sub-parameters: segment length and segmentation step size. The segment length defines the number of sampling points contained in each data segment, which can be denoted as... The segment step size defines the interval distance between the starting sampling points of two adjacent segments, which can be denoted as: ,and To ensure that there are repeated sampling points between adjacent segments, the preset segmentation interval parameter can be set according to the actual work. Based on the preset segmentation interval parameter, the interactive terminal segments the pure horizontal axis vibration data, pure vertical axis vibration data, and pure vertical axis vibration data into segments respectively. Taking the pure horizontal axis vibration data as an example, its mathematical form is: , The sampling timestamp is used. Based on the segment length and segment step parameters in the preset segmentation interval parameters, the pure horizontal axis vibration data is slide-segmented to obtain several horizontal axis vibration segment data. These segment data constitute the horizontal axis vibration segment dataset. Each horizontal axis vibration segment data can be represented as... ,in This is the starting sampling timestamp for the horizontal axis vibration segment data. Similarly, for the pure vertical axis vibration data... and pure vertical axis vibration data Perform the same segmentation operation to obtain the longitudinal axis vibration segmentation dataset and the vertical axis vibration segmentation dataset, respectively.
[0088] S302, perform short-time Fourier transform on each horizontal axis vibration segment data in the horizontal axis vibration segment dataset to obtain the horizontal axis spectrum sequence; perform short-time Fourier transform on each vertical axis vibration segment data in the vertical axis vibration segment dataset to obtain the vertical axis spectrum sequence; perform short-time Fourier transform on each vertical axis vibration segment data in the vertical axis vibration segment dataset to obtain the vertical axis spectrum sequence.
[0089] Specifically, the interactive terminal performs a short-time Fourier transform on each vibration segment data in the vibration segment dataset for each axis obtained in step S301. The short-time Fourier transform is an improved form of the traditional Fourier transform, used to independently calculate the spectral energy distribution for each short time interval, thus preserving the dynamic information of the signal over time while obtaining frequency domain information. Taking the horizontal axis as an example, for each horizontal vibration segment data in the horizontal axis vibration segment dataset, a short-time Fourier transform is performed once to calculate the energy intensity distribution of the horizontal axis vibration segment data at various frequency values, resulting in the corresponding horizontal axis spectrum sequence. Each element in this horizontal axis spectrum sequence represents the energy intensity value of the original pure vibration signal at a specific sampling timestamp and a specific frequency value. Similarly, the same short-time Fourier transform processing is performed on each vertical vibration segment data in the vertical axis vibration segment dataset and each vertical vibration segment data in the vertical axis vibration segment dataset to obtain the corresponding vertical axis spectrum sequence and vertical axis spectrum sequence.
[0090] S303, based on a preset Mel filter, performs Mel mapping on each horizontal axis spectrum sequence to obtain a horizontal axis Mel spectrum; based on the Mel filter, performs Mel mapping on each vertical axis spectrum sequence to obtain a vertical axis Mel spectrum; based on the Mel filter, performs Mel mapping on each vertical axis spectrum sequence to obtain a vertical axis Mel spectrum.
[0091] Specifically, the preset Mel filter is a pre-designed set of triangular bandpass filters, in which the triangular bandpass filters are evenly spaced on the Mel frequency scale. After mapping back to the linear frequency scale, it exhibits a distribution characteristic of dense filters in the low-frequency band and sparse filters in the high-frequency band, to simulate the nonlinear characteristics of the human ear's sensitivity to sound in different frequency bands. This Mel filter set contains a total of A triangular filter, The value is a preset positive integer, which determines the number of frequency dimensions in the Mel spectrogram. The preset Mel filter can be set according to the actual work. The interactive terminal first combines and arranges the spectrum sequences of each axis obtained in step S302 along the time dimension to generate the time spectrogram corresponding to each axis. The time spectrogram is a two-dimensional data representation, with its horizontal axis corresponding to the dimension of sampling timestamps and its vertical axis corresponding to the dimension of frequency values; any coordinate point in the figure... The value at that location represents the sampling timestamp. Frequency value The vibration energy intensity at a given location. Taking the horizontal axis as an example, the horizontal axis spectrum sequences are arranged sequentially along the time dimension according to the starting sampling timestamps of their corresponding horizontal axis vibration segments, resulting in a complete data matrix, which is the horizontal axis time spectrum. The mathematical form of the horizontal axis time spectrum can be represented as a two-dimensional matrix. ,in The total number of time frames is determined by the number of horizontal axis vibration segment data, where each time frame corresponds to the starting sampling timestamp of the horizontal axis vibration segment data. This represents the total number of frequency values, the value of which is determined by the frequency resolution of the short-time Fourier transform. Similarly, performing the same combination processing on the vertical axis spectrum sequence and the vertical axis spectrum sequence yields the corresponding vertical axis time spectrum. Spectrum diagram along the vertical axis The interactive terminal performs Mel-mapping processing on the time-spectrum graphs of each axis based on a preset Mel filter. Taking the horizontal axis as an example, the horizontal axis time-spectrum graph... For each horizontal axis spectral sequence, calculate its pass through all The energy output of a Mel filter is obtained to get a The Mel energy vectors of all horizontal axis spectrum sequences are arranged along the time dimension to obtain the horizontal axis Mel spectrum, which can be mathematically represented as a two-dimensional matrix. Similarly, the same Mel mapping process is performed on the vertical axis time spectrum and the vertical axis time spectrum to obtain the corresponding vertical axis Mel spectrum. and vertical axis Mel spectrum .
[0092] S304 stacks the horizontal axis Mel spectrum, vertical axis Mel spectrum, and vertical axis Mel spectrum to obtain multi-channel Mel spectrum data.
[0093] Specifically, the horizontal, vertical, and tandem Mel spectrograms of the interactive terminal are stacked and fused along the channel dimension. The stacking operation aligns and combines the three axial two-dimensional Mel spectrograms along the channel dimension to form a three-dimensional data block with three channels, i.e., multi-channel Mel spectrogram data. The mathematical form of multi-channel Mel spectrogram data can be represented as a three-dimensional tensor. ,in This represents the total number of time frames, and its value is determined by the number of horizontal axis vibration segment data. The magnitude of the Mel frequency dimension is determined by the number of filters in the preset Mel filter bank; The number of channels corresponds to the vibration information along the horizontal, vertical, and axial directions, respectively. This three-dimensional tensor fully represents the multi-channel pattern of the vibrational energy of the throat skin in three-dimensional space as a function of time and frequency during vocalization.
[0094] This embodiment provides a flexible laryngeal speech interaction method for resisting motion interference. It sequentially processes clean triaxial vocal cord vibration data through segmentation, short-time Fourier transform, time-spectrum graph combination, Mel mapping, and channel stacking. This transforms the one-dimensional temporal vibration signal into multi-channel Mel spectrogram data with a three-dimensional time-frequency-channel structure. The entire process simulates human auditory perception characteristics for feature extraction and fusion, fully preserving the vibration pattern information of the laryngeal skin in three spatial directions during phonation. This provides highly discriminative and noise-robust structured input data for subsequent parallel recognition by deep learning models, effectively ensuring the accuracy of voice command recognition and user authentication in dynamic scenarios.
[0095] In one embodiment, the identity instruction verification convolutional neural network model includes an input layer, a feature extraction layer, a global pooling layer, a fully connected layer for instruction recognition, a fully connected layer for identity authentication, and a result output layer, wherein:
[0096] S401, Input Layer: The input layer receives multi-channel Mel spectrogram data and transmits the multi-channel Mel spectrogram data to the feature extraction layer.
[0097] Specifically, the input layer in the identity instruction verification convolutional neural network model serves as the data entry point for the model, receiving multi-channel Mel spectrogram data and transmitting it to the feature extraction layer.
[0098] S402, the feature extraction layer, is used to perform two-dimensional convolution on the multi-channel Mel spectrogram data, extracting local time-frequency texture features and global acoustic semantic features from the multi-channel Mel spectrogram data step by step to obtain a compressed feature map, and then transmitting the compressed feature map to the global pooling layer.
[0099] Specifically, the feature extraction layer receives multi-channel Mel-ray spectrogram data from the input layer and performs progressively layered two-dimensional convolution processing on it. Two-dimensional convolution is an operation that involves sliding a convolution kernel across a two-dimensional data plane to perform local weighted summation; the weight parameters of the convolution kernel are determined through pre-training of the model. The feature extraction layer may contain multiple sequentially connected two-dimensional convolution modules. Each two-dimensional convolution module performs two-dimensional convolution operations, batch normalization, non-linear activation, and pooling on the input data. Earlier convolution modules use kernels with smaller receptive fields to extract local time-frequency texture features such as edges and corners from the input 3D tensor. Later convolution modules use kernels with progressively larger equivalent receptive fields, further extracting global acoustic semantic features across time and frequency based on the feature map output from the previous layer. As the data is passed deeper, the spatial, temporal, and frequency dimensions of the feature map are gradually compressed, ultimately resulting in a compressed feature map with a smaller spatial size and a larger number of channels. This compressed feature map is then passed to the global pooling layer. This compressed feature map preserves high-level abstract features from the input multi-channel Mel spectrogram data while discarding redundant spatial details. For example, taking the first two-dimensional convolutional module as an example, its input is multi-channel Mel spectrogram data. First, through a set of learnable two-dimensional convolutional kernels and bias Perform a two-dimensional convolution operation on it, where The spatial size of the convolution kernel. This represents the number of output channels for this module. The two-dimensional convolution operation can be expressed as: In the formula, It is multi-channel Mel spectrogram data The first along the time frame direction The first time frame and the first One frequency value, It is a new feature map channel generated by the 2D convolution module after performing convolution operations, and its value is... , The preset number of output channels for this module It is the input channel index, and its value range is... These correspond to the horizontal axis Mel-ray spectrogram, the vertical axis Mel-ray spectrogram, and the center axis Mel-ray spectrogram, respectively. Each output channel The feature map is generated by weighting and summing the spatial dimensions of the input Mel spectrum maps along the three axes using a set of independent two-dimensional convolutional kernels. During batch normalization of the convolutional output, the mean and variance of the values at all spatial locations in each output channel are calculated. The mean is subtracted from the value at each location, and then the result is divided by the standard deviation to maintain a stable statistical distribution for that channel. When applying the nonlinear activation function, the value of each element in the batch-normalized feature map is calculated individually. Positive values are retained, and negative values are set to zero. During max pooling, within a preset space window (e.g., a size of...),... The maximum value of all values within a given window is selected as the representative value of that window. The window is then slid across the time and frequency dimensions with a preset step size to perform spatial downsampling of the feature map. The final output is a window of size [size missing]. The feature map, where and These represent the temporal frame dimension and frequency value dimension, respectively, after pooling downsampling. This feature map serves as the input to the next two-dimensional convolutional module, and the last two-dimensional convolutional module outputs a compressed feature map. ,in and These represent the size of the time dimension and the frequency dimension after multiple downsampling processes, respectively. This represents the number of output channels for the last module.
[0100] S403, the global pooling layer, is used to perform global average pooling on the compressed feature map to obtain acoustic feature vectors, and then transmits the acoustic feature vectors to the instruction recognition fully connected layer and the identity authentication fully connected layer.
[0101] Specifically, the global pooling layer receives the compressed feature map output by the feature extraction layer. Global average pooling is then performed on it. Global average pooling is an operation that aggregates the entire spatial plane into a single statistic. Specifically, it is calculated as follows: for each channel in the compressed feature map... The size corresponding to this channel is Summing the values at all spatial locations in the two-dimensional feature matrix and dividing by the total number of spatial locations. The global average value of this channel is calculated. This process can be expressed as: In the formula, For the first The global average value of each channel. Spatial location in compressed feature map ,aisle The eigenvalues at that location. For all After calculating the global average value for each channel, the global average values of each channel are arranged in channel order to obtain a value of length [value missing]. a one-dimensional vector This is the acoustic feature vector. This acoustic feature vector performs a global aggregation of all time and frequency information in the input multi-channel Mel spectrogram data, uniformly mapping input signals of arbitrary length into a fixed-dimensional feature representation, thus eliminating the influence of changes in the time length of the input signal on subsequent classification layers.
[0102] S404, the fully connected layer for instruction recognition, is used to map acoustic feature vectors to obtain instruction category probability data and transmit the instruction category probability data to the result output layer.
[0103] Specifically, the instruction recognition fully connected layer is used for instruction classification based on acoustic feature vectors. A fully connected layer is a basic neural network structure where each output neuron is connected to all elements of the input vector. It linearly maps the input using a learnable weight matrix and bias vector, and can further apply non-linear activation functions. The instruction recognition fully connected layer receives the acoustic feature vector output from the global pooling layer, transforms it into a vector of length equal to the preset number of instruction categories through linear mapping, and then processes it using a normalized exponential function to obtain instruction category probability data. This instruction category probability data is then transmitted to the output layer. The mathematical form of the instruction category probability data can be expressed as: ,in The total number of instructions in the preset instruction set, and the number of instructions in the vector. The value of the element is the value of the current input multi-channel Mel spectrogram data belonging to the first element. The probability value of class instructions, all The sum of the probability values is 1.
[0104] S405, the fully connected layer for identity authentication, is used to map acoustic feature vectors to obtain probabilities of legitimate and illegitimate users, and then transmits these probabilities to the result output layer.
[0105] Specifically, the instruction recognition fully connected layer receives the acoustic feature vector output from the global pooling layer, transforms it into a vector of length equal to the preset total number of instruction categories through a linear mapping, and then processes it through a normalized exponential function to obtain legitimate user probability data and illegitimate user probability data. These two data are then transmitted to the result output layer. The legitimate user probability data can be denoted as... The probability data of illegal users can be denoted as: The sum of the probabilities of legitimate users and illegitimate users is 1.
[0106] S406, Result Output Layer: The result output layer is used to obtain the voice command recognition result based on the command category probability data, and to obtain the user authentication result based on the legitimate user probability data and the illegitimate user probability data.
[0107] Specifically, the output layer receives the command category probability data output by the command recognition fully connected layer, and the legitimate user probability data and illegitimate user probability data output by the authentication fully connected layer, and generates the final voice command recognition result and user authentication result, respectively. For the command category probability data, the output layer... The system searches for the element with the largest probability value and identifies the command corresponding to that value as the voice command recognition result. This result indicates which specific command in the preset command set the user's voice content belongs to. For both legitimate and illegitimate user probability data, the output layer compares the legitimate user probability data. and illegal user probability data The size, if If the user is legitimate, the authentication result will be used; otherwise, the authentication result will be used.
[0108] This embodiment provides a flexible laryngeal speech interaction method to resist motion interference. It inputs multi-channel Mel-spectrum data into an identity command verification convolutional neural network model comprising an input layer, a feature extraction layer, a global pooling layer, a fully connected layer for command recognition, a fully connected layer for identity authentication, and a result output layer. The multi-channel Mel-spectrum data undergoes sequential two-dimensional convolutional feature extraction, global average pooling compression, and dual-path fully connected mapping processing, ultimately outputting the speech command recognition result and the user identity authentication result in parallel. This method reduces the resource overhead of model deployment and inference while maintaining recognition accuracy, making it suitable for dynamic laryngeal speech interaction scenarios with high real-time requirements.
[0109] In one embodiment, the identity command verification convolutional neural network model is obtained through the following steps:
[0110] S501, Obtain the preset historical dataset; the historical dataset includes historical multi-channel Mel spectrogram data and historical instruction category labels and historical identity labels corresponding to the historical multi-channel Mel spectrogram data.
[0111] Specifically, the interactive terminal acquires a preset historical dataset for subsequent model training. This preset historical dataset is pre-collected and labeled according to actual application needs, including each set of historical multi-channel Mel spectrogram data, the corresponding historical instruction category label, and the historical identity label. The historical multi-channel Mel spectrogram data is a three-dimensional tensor generated by processing the user's vocalizations according to a preset set of instructions while wearing the throat sensor, using the same method as S103 to obtain the multi-channel Mel spectrogram data. Its mathematical form can be... The structure is completely consistent with the multi-channel Mel spectrogram data used in real-time interaction. Historical instruction category labels are used to identify the specific instruction content corresponding to that segment of historical multi-channel Mel spectrogram data; their format can be a set of instructions with a preset length containing the total number of instructions. The one-hot encoded vector is a vector in which only the element corresponding to the correct instruction category is 1, and all other positions are 0. The historical identity tag is used to identify the identity from which the data segment originated. It can be a binary scalar value or a two-dimensional one-hot encoded vector, indicating whether the sample belongs to a legitimate user or an illegitimate user, respectively.
[0112] S502 connects the input layer, feature extraction layer, global pooling layer, instruction recognition fully connected layer, identity authentication fully connected layer and result output layer to obtain the basic neural network framework, and initializes the basic neural network framework to obtain the initial neural network.
[0113] Specifically, the interactive terminal sequentially connects and parallel branches the functional layers constituting the identity command verification convolutional neural network model according to a preset network structure. The output of the input layer is connected to the input of the feature extraction layer, which in turn is connected to the input of the global pooling layer. The output of the global pooling layer is simultaneously connected in parallel to the inputs of both the command recognition fully connected layer and the identity authentication fully connected layer, forming a dual-path branch. The outputs of the command recognition fully connected layer and the identity authentication fully connected layer are respectively connected to the corresponding inputs of the result output layer. The specific structural parameters within each layer, such as the number of two-dimensional convolutional modules in the feature extraction layer, the kernel size of each convolutional module, the number of output channels, and the pooling window size, are all set according to the preset model configuration. After completing the inter-layer connections, the learnable parameters in the entire network are initialized to obtain the initial neural network. Optionally, the initialization method can be the weight matrix of the two-dimensional convolutional kernels. and bias vector Random initialization with variance scaling can be used to effectively propagate gradients in the early stages of training; scaling parameters in batch normalization layers can also be adjusted. Initialize to 1, translation parameter Initialize to 0. The weights and biases of the fully connected layers can be similarly initialized randomly. After the above connections and initialization, we obtain an initial neural network that has not yet been trained, providing a starting point for subsequent iterative training.
[0114] S503, input the historical multi-channel Mel spectrogram data from the historical dataset into the initial neural network to obtain the predicted instruction category data and predicted identity data of the historical multi-channel Mel spectrogram data.
[0115] Specifically, for each historical multi-channel Mel spectrogram data point in the historical dataset, the interactive terminal inputs it into an initial neural network. The initial neural network outputs the predicted instruction category data and predicted identity data for that historical multi-channel Mel spectrogram data. The mathematical form of the predicted instruction category data can be expressed as follows: ,in The total number of instructions in the preset instruction set, and the number of instructions in the vector. The value of the element is the historical multi-channel Mel spectrogram data belonging to the first element. The probability value of the instruction type. The predicted identity data includes the predicted probability data of legitimate users and the predicted probability data of illegitimate users, and the sum of the predicted probability data of legitimate users and the predicted probability data of illegitimate users is 1.
[0116] S504: Based on the predicted instruction category data and historical instruction category labels, the instruction recognition loss value is calculated; based on the predicted identity data and historical identity labels, the identity authentication loss value is calculated.
[0117] Specifically, for each historical multi-channel Mel-ray spectrogram data point, the interactive terminal quantifies the loss values for both tasks based on the deviation between the predicted value and the true label. For the instruction recognition task, a multi-class cross-entropy loss function can be used to calculate the instruction recognition loss value between the predicted instruction category data and the historical instruction category labels. Instructions identify loss values. It can be determined in the following ways: ,in For the historical instruction category label The value of each element, For predicting the first instruction category data The probability value of the instruction class. This represents the total number of instructions in the preset instruction set. For identity authentication tasks, a binary cross-entropy loss function can be used to calculate the identity authentication loss value between the predicted identity data and historical identity tags. Identity authentication loss value This can be achieved by combining historical identity tags and predicted identity data. Calculated, i.e. ,in is probability data predicting a legitimate user, is probability data predicting an illegal user, is a historical identity label.
[0118] S505, calculating a joint loss value based on the instruction recognition loss value and the identity authentication loss value, wherein the expression of the joint loss value is:
[0119]
[0120] in the formula, is an index of any historical multi-channel Mel spectrogram data, is the -th joint loss value of the historical multi-channel Mel spectrogram data, is an instruction loss weight, is the -th instruction recognition loss value of the historical multi-channel Mel spectrogram data, is an identity loss weight, is the -th identity authentication loss value of the historical multi-channel Mel spectrogram data.
[0121] Specifically, the interactive terminal fuses the loss values of two subtasks into a single joint loss value by means of weighted summation. Wherein, and are respectively preset instruction loss weight and identity loss weight, which are used to adjust the relative importance of the instruction recognition task and the identity authentication task in the joint optimization process, and can be set according to actual work. The instruction recognition loss value of any -th historical multi-channel Mel spectrogram data and the identity authentication loss value are obtained in step S504, and the joint loss value of the sample can be obtained by substituting into the above formula .
[0122] S506, performing iterative training on an initial neural network with the goal of minimizing the joint loss value until a preset convergence condition is satisfied, so as to obtain an identity instruction verification convolutional neural network model.
[0123] Specifically, the interactive terminal uses an optimization algorithm to update the model parameters. In each iteration, the terminal selects a batch of samples from the historical dataset, calculates the average joint loss value of that batch according to the methods described in steps S503 to S505, and then uses a gradient descent-based optimizer (such as stochastic gradient descent or the Adam optimizer) to calculate the gradient of the loss function with respect to each model parameter, and updates the model parameters along the negative gradient direction. This process is repeated, performing multiple cycles on all samples in the historical dataset. Preset convergence conditions may include: the joint loss value no longer significantly decreases within a certain number of consecutive iterations, or the preset maximum number of training cycles is reached, or the accuracy of the model's instruction recognition and identity authentication on an independent validation set meets preset performance metrics. When the preset convergence conditions are met, the iterative training terminates, and the model parameters are saved, thus obtaining the trained identity instruction verification convolutional neural network model.
[0124] This embodiment provides a flexible laryngeal speech interaction method for resisting motion interference. It acquires a historical dataset consisting of historical multi-channel Mel spectrogram data and its associated commands and identity tags. A neural network containing a feature extraction layer and a dual-path classification layer is constructed and initialized. Forward propagation is used to generate prediction results, and command recognition loss and identity authentication loss are calculated separately. These two losses are then weighted and fused into a joint loss value as the optimization objective. This process is iteratively trained on the initial neural network until convergence, resulting in an identity command verification convolutional neural network model that can output commands and identity results in parallel. This allows the model to simultaneously learn discrimination patterns related to vocal content and individual identity based on shared acoustic features, reducing the computational and storage overhead of training multiple models individually. Furthermore, multi-objective collaborative optimization enhances the model's generalization ability, providing reliable model support for accurately achieving speech command recognition and user authentication in dynamic environments.
[0125] In one embodiment, the expression for the frequency domain window function of vocal cord vibration is:
[0126]
[0127] In the formula, It is any frequency value. It is a frequency domain window function for vocal cord vibration. It is the minimum frequency of vocal cord vibration. It is the maximum frequency of vocal cord vibration.
[0128] Specifically, the vocal fold vibration frequency domain window function is a binary function defined on the frequency axis, used to selectively filter frequency domain acceleration data along each axis in product operations. The value of this window function on the frequency axis is determined solely by the minimum frequency of vocal fold vibration. Maximum frequency of vocal cord vibration Two boundary parameters determine: when the frequency value Located in a closed interval When the frequency component is within a certain range, the window function value is 1, indicating that the frequency component is completely preserved; when the frequency value is within a certain range... When the frequency component is outside this range, the window function value is 0, indicating that the frequency component is completely filtered out. Taking the horizontal axis as an example, for frequency domain horizontal axis acceleration data... Each frequency value in The interactive terminal compares its value with... Multiply to obtain the denoised frequency domain horizontal axis data. .when exist When within range, The result of multiplying is the original frequency intensity value. In itself, this frequency component is completely preserved; when When outside the range, The product of these two values is 0, meaning that this frequency component has been completely filtered out. The preset maximum vocal cord vibration frequency can be set based on the upper limit of the physiological frequency of larynx vocalization or the effective working bandwidth of the sensor.
[0129] The minimum frequency of vocal cord vibration is obtained through the following steps:
[0130] S601, acquires the user's silent motion triaxial acceleration data and silent stationary triaxial acceleration data.
[0131] Specifically, the interactive terminal acquires the user's silent motion triaxial acceleration data and silent stationary triaxial acceleration data to calibrate the minimum vocal cord vibration frequency. The silent motion triaxial acceleration data is historical triaxial acceleration data collected by a flexible laryngeal sensor and transmitted to the interactive terminal when the user performs a preset limb movement in a silent state. The preset limb movement could be walking continuously at a natural speed on flat ground for a preset duration, during which the user remains silent and does not make any sound, so that the sensor only captures the inertial acceleration signal caused by the limb movement, excluding vocal cord vibration components. The silent stationary triaxial acceleration data is historical triaxial acceleration data collected when the user remains stationary in a silent state, during which the user maintains a sitting or standing posture without making any sound, and the sensor only captures baseline vibration signals caused by electronic noise and minor physiological activities (such as heartbeat or breathing). The mathematical form of the silent motion triaxial acceleration data can be represented by three time-series vectors. , , The mathematical form of silent stationary triaxial acceleration data can be expressed as: , , superscript Indicates state of motion, Indicates a static state. These are discrete sampling timestamps.
[0132] S602, perform Fourier transform on the silent motion triaxial acceleration data to obtain silent motion spectrum data; perform Fourier transform on the silent stationary triaxial acceleration data to obtain silent stationary spectrum data; the silent motion spectrum data includes silent motion horizontal axis spectrum data, silent motion vertical axis spectrum data and silent motion vertical axis spectrum data; the silent stationary spectrum data includes silent stationary horizontal axis spectrum data, silent stationary vertical axis spectrum data and silent stationary vertical axis spectrum data.
[0133] Specifically, the interactive terminal performs Fourier transforms on the acceleration sequences of each axis in the two types of triaxial acceleration data obtained in step S601, converting them from the time domain to the frequency domain. (Using the horizontal axis acceleration data during silent motion...) For example, performing a Fourier transform on it yields the horizontal axis spectrum data of silent motion. ,in This is the frequency value, in Hz. modulus Characterized by the frequency value The energy intensity of the horizontal axis acceleration signal during silent motion. The vertical axis acceleration data during silent motion. Vertical axis acceleration data during silent motion Silent stationary horizontal axis acceleration data Silent stationary longitudinal acceleration data and silent stationary vertical axis acceleration data Perform Fourier transforms on each component to obtain the vertical spectrum data of the silent motion. Silent motion vertical axis spectrum data Silent static horizontal axis spectrum data Silent stationary vertical axis spectrum data and silent static vertical axis spectrum data .
[0134] S603, the silent motion horizontal axis spectrum data, silent motion vertical axis spectrum data, and silent motion vertical axis spectrum data are superimposed to obtain silent motion energy data. The expression for the silent motion energy data is as follows:
[0135]
[0136] In the formula, It is any frequency value. It is the frequency value in silent motion energy data. The value of kinetic energy intensity. The frequency values in the horizontal axis spectrum data of silent motion The transverse frequency intensity, The frequency values in the vertical axis spectrum data of silent motion The vertical axis frequency intensity, The frequency values in the vertical axis spectrum data of silent motion The vertical axis frequency intensity.
[0137] Specifically, the interactive terminal performs frequency-point aligned energy superposition on the three axial spectra in the silent motion spectrum data. For each frequency value... From the horizontal axis spectrum data of silent motion respectively , silent motion longitudinal spectrum data Vertical axis spectrum data of silent motion Extract the frequency intensity value corresponding to the given frequency value, take the square of its respective modulus (i.e., energy value), and then sum the energy values along the three axes to obtain the frequency value. Motion energy intensity value at the location After traversing all frequency values, a line of silent motion energy data, independent of direction, is formed. .
[0138] S604, the silent stationary horizontal axis spectrum data, silent stationary vertical axis spectrum data, and silent stationary vertical axis spectrum data are superimposed to obtain silent stationary energy data. The expression for silent stationary energy data is as follows:
[0139]
[0140] In the formula, It is any frequency value. It is the frequency value in silent static energy data. The static energy intensity value, It is the frequency value in the silent, stationary horizontal axis spectrum data. The transverse frequency intensity, It is the frequency value in the silent, stationary vertical axis spectrum data. The vertical axis frequency intensity, It is the frequency value in the silent, stationary vertical axis spectrum data. The vertical axis frequency intensity.
[0141] Specifically, the interactive terminal performs energy superposition on the three axial spectra in the silent, stationary spectrum data in the same manner as in step S603. For each frequency value Take silent, stationary horizontal axis spectrum data respectively Silent stationary vertical axis spectrum data and silent static vertical axis spectrum data The frequency value is obtained by squaring the modulus and summing the results. The static energy intensity value at the location After traversing the entire frequency range, silent and static energy data is constructed. .
[0142] S605, the static energy intensity value with the largest value in the silent static energy data is determined as the noise threshold, and the motion energy intensity value of each frequency value in the silent motion energy data is compared with the noise threshold to obtain the noise comparison result of each frequency value.
[0143] Specifically, the interactive terminal uses silent static energy data to determine a noise threshold and, based on this, identifies the effective motion energy frequency range within the silent motion energy data. First, it iterates through the silent static energy data. The static energy intensity values corresponding to all frequencies are taken as the largest value as the noise threshold. This noise threshold represents the maximum energy level that non-motion factors can cause across all frequencies, subsequently applied to silent motion energy data. Each frequency value in The kinetic energy intensity value corresponding to this frequency value With noise threshold Numerical comparisons were performed to obtain noise comparison results for each frequency value.
[0144] S606, select the frequency values in the noise comparison result where the motion energy intensity value is greater than the noise threshold, form a noise range set, and determine the frequency value with the largest value in the noise range set as the minimum value of vocal cord vibration frequency.
[0145] Specifically, based on the noise comparison results of each frequency value obtained in step S605, the interactive terminal filters out all frequency values whose motion energy intensity values are greater than the noise threshold. The frequency values are arranged into a noise range set. This set contains all frequency components effectively driven by the user's limb movements. The frequency value with the highest value in this set is taken as the upper limit frequency of the user's motion energy distribution. This frequency is the highest frequency position that the user's limb motion energy can reach in the current state, determined through data-driven analysis. Optionally, to allow a safe frequency interval for noise reduction processing in actual voice interaction and to prevent spectral leakage components of motion artifacts from mixing into the reserved frequency band, a preset protection factor can be applied based on the upper limit frequency of motion energy distribution. (For example That is, to extend the frequency margin by 30%, and combine it with Multiply them, and determine the product as the minimum frequency of vocal cord vibration. The minimum vocal cord vibration frequency is objectively calibrated based on the user's personal motion data, without relying on general physiological frequency assumptions, thus achieving personalized customization of the cutoff frequency under the frequency domain window function of vocal cord vibration.
[0146] This embodiment provides a flexible laryngeal voice interaction method for resisting motion interference. It collects laryngeal acceleration data from users in both non-vocalized motion and non-vocalized static states, performs Fourier transforms on each axis to obtain the spectrum, and then fuses them into a total energy distribution curve. The extreme value of static energy is used as a noise threshold to filter out the upper limit of effective motion frequencies from the motion energy, ultimately determining the minimum vocal cord vibration frequency. The entire process transforms the setting of filter parameters from relying on general physiological experience to adaptive calibration based on the user's individual motion data. This allows the vocal cord vibration frequency domain window function used in the denoising process to accurately match the motion noise spectrum characteristics of a specific user, effectively filtering out motion artifacts while maximizing the preservation of effective vocal frequency bands, thus significantly improving the personalization and robustness of voice interaction in dynamic scenarios.
[0147] In one embodiment, the sampling frequency of the original triaxial acceleration data is 1000Hz, and the original triaxial acceleration data is acquired by a flexible sensor. The flexible sensor is an island-bridge heterogeneous integrated architecture, wherein the bridging part is made of liquid metal ink doped with silicon dioxide, and the mass fraction of silicon dioxide is 6wt%. The maximum value of the vocal cord vibration frequency is 255Hz, and the minimum value of the vocal cord vibration frequency can also be directly set to 80Hz.
[0148] Specifically, the raw triaxial acceleration data can be acquired using a flexible sensor deployed on the user's throat. This flexible sensor is a wearable device with an island-bridge heterogeneous integrated architecture, where the islands integrate locally rigid microelectromechanical systems (MEMS) triaxial accelerometers, and the islands are connected by bridges. The bridges are printed from gallium-indium alloy liquid metal ink doped with silicon dioxide. This liquid metal ink, prepared by mixing gallium-indium alloy with 6 wt% silicon dioxide, possesses thixotropic properties and can withstand tensile strain without electrical failure. The addition of 6 wt% silicon dioxide nanoparticles is mainly used to reduce the surface tension of the liquid metal and adjust its rheological properties, enabling it to be formed into a preset circuit pattern with high resolution through screen printing. This ensures that the printing integrity and electrical conductivity stability of the bridge structure are maintained under the dynamic environment of frequent stretching and torsion in the throat. The flexible sensor is attached to the skin surface at the user's Adam's apple location through a low-modulus elastomer substrate and encapsulation layer to directly sense the combined inertial impacts caused by vocal cord vibration and laryngeal muscle movement on the skin surface during speech. When a user vocalizes or silently recites, the vibrations of the laryngeal skin in space are recorded by the triaxial accelerometers in the island section and synthesized into raw triaxial acceleration data, which is then transmitted to the interactive terminal. The sampling frequency of the raw triaxial acceleration data can be set to 1000Hz, meaning the time interval between each sampling timestamp is 1 millisecond. This allows for the full capture of the rapid vibration details of the laryngeal skin during vocalization, providing sufficient time resolution and effective bandwidth for subsequent frequency domain analysis. (Maximum vocal cord vibration frequency...) It can be set to 255Hz. This value is determined based on the upper limit of the physiological frequency of human vocalization, which can cover the complete distribution range of the fundamental frequency of vocal vibration in the throat of a healthy adult. While retaining all effective information of vocal cord vibration, it filters out electronic noise interference in higher frequency bands. Minimum vocal cord vibration frequency. It can be directly set to 80Hz. This value is determined based on the general laws of human movement physiology. In daily activities such as walking and running, the acceleration energy caused by the limb movement of healthy adults is usually concentrated in the low frequency range below 20Hz. As the lower limit frequency of the passband, 80Hz can provide a sufficient safe frequency interval between the upper limit frequency of motion energy distribution and the lowest fundamental frequency of vocal cord vibration, ensuring that motion artifacts can be effectively filtered out in most daily exercise scenarios, while fully preserving the fundamental frequency and main harmonic components of vocal cord vibration.
[0149] In the aforementioned flexible laryngeal voice interaction method for resisting motion interference, the user's original triaxial acceleration data is acquired. This data characterizes the three-dimensional mechanical vibration of the user's laryngeal skin during phonation. Noise reduction is applied to the original triaxial acceleration data to obtain clean triaxial vocal cord vibration data. Time-frequency domain feature extraction is performed on the clean triaxial vocal cord vibration data to obtain multi-channel Mel-frequency spectrogram data. This multi-channel Mel-frequency spectrogram data is then input into a convolutional neural network model for identity command verification to obtain voice command recognition results and user authentication results. This method achieves the technical effect of accurately reconstructing the vocal content from mixed interference signals and simultaneously identifying the speaker's identity in a dynamic user environment, thus significantly improving the reliability and accuracy of flexible laryngeal voice interaction technology in complex scenarios.
[0150] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.
[0151] Based on the same inventive concept, this application also provides a flexible laryngeal voice interaction system for implementing the aforementioned flexible laryngeal voice interaction method for resisting motion interference. The solution provided by this system is similar to the implementation described in the above method. Therefore, the specific limitations of one or more embodiments of a flexible laryngeal voice interaction system for resisting motion interference provided below can be found in the limitations of the flexible laryngeal voice interaction method for resisting motion interference described above, and will not be repeated here.
[0152] In one exemplary embodiment, such as Figure 2 As shown, a flexible laryngeal voice interaction system 200 for resisting motion interference is provided, comprising:
[0153] The data acquisition module 201 is used to acquire the user's raw triaxial acceleration data; the raw triaxial acceleration data is used to characterize the three-dimensional mechanical vibration of the user's laryngeal skin during vocalization;
[0154] The noise reduction module 202 is used to perform noise reduction processing on the original triaxial acceleration data to obtain clean triaxial vocal cord vibration data;
[0155] Feature extraction module 203 is used to extract time-frequency domain features from pure triaxial vocal cord vibration data to obtain multi-channel Mel spectrogram data;
[0156] The results analysis module 204 is used to input multi-channel Mel spectrogram data into the identity command verification convolutional neural network model to obtain the voice command recognition result and the user identity authentication result; the identity command verification convolutional neural network model has a dual-branch output structure.
[0157] Furthermore, the raw triaxial acceleration data includes raw horizontal axis acceleration data, raw vertical axis acceleration data, and raw vertical axis acceleration data. The noise reduction module can also be used for:
[0158] Perform a Fourier transform on the original horizontal axis acceleration data to obtain the frequency domain horizontal axis acceleration data; perform a Fourier transform on the original vertical axis acceleration data to obtain the frequency domain vertical axis acceleration data; perform a Fourier transform on the original vertical axis acceleration data to obtain the frequency domain vertical axis acceleration data.
[0159] The denoised frequency domain horizontal axis acceleration data and vocal cord vibration frequency domain window function are multiplied to obtain the denoised frequency domain horizontal axis data; the denoised frequency domain vertical axis acceleration data and vocal cord vibration frequency domain window function are multiplied to obtain the denoised frequency domain vertical axis data; the denoised frequency domain vertical axis acceleration data and vocal cord vibration frequency domain window function are multiplied to obtain the denoised frequency domain vertical axis data.
[0160] Perform an inverse Fourier transform on the denoised frequency domain horizontal axis data to obtain clean horizontal axis vibration data; perform an inverse Fourier transform on the denoised frequency domain vertical axis data to obtain clean vertical axis vibration data; perform an inverse Fourier transform on the denoised frequency domain vertical axis data to obtain clean vertical axis vibration data.
[0161] Pure triaxial vocal cord vibration data is constructed based on pure horizontal axis vibration data, pure vertical axis vibration data, and pure vertical axis vibration data.
[0162] Furthermore, the feature extraction module can also be used for:
[0163] Based on the preset segmentation interval parameters, the pure horizontal axis vibration data is segmented to obtain the horizontal axis vibration segmented dataset; based on the segmentation interval parameters, the pure vertical axis vibration data is segmented to obtain the vertical axis vibration segmented dataset; based on the segmentation interval parameters, the pure vertical axis vibration data is segmented to obtain the vertical axis vibration segmented dataset.
[0164] Short-time Fourier transforms are performed on the horizontal axis vibration segment data in the horizontal axis vibration segment dataset to obtain the horizontal axis spectrum sequence; short-time Fourier transforms are performed on the vertical axis vibration segment data in the vertical axis vibration segment dataset to obtain the vertical axis spectrum sequence; short-time Fourier transforms are performed on the vertical axis vibration segment data in the vertical axis vibration segment dataset to obtain the vertical axis spectrum sequence.
[0165] Based on the preset Mel filter, Mel mapping is performed on each horizontal axis spectrum sequence to obtain the horizontal axis Mel spectrum; based on the Mel filter, Mel mapping is performed on each vertical axis spectrum sequence to obtain the vertical axis Mel spectrum; based on the Mel filter, Mel mapping is performed on each vertical axis spectrum sequence to obtain the vertical axis Mel spectrum.
[0166] The horizontal axis Mel spectrum, vertical axis Mel spectrum, and vertical axis Mel spectrum are stacked to obtain multi-channel Mel spectrum data.
[0167] Furthermore, the identity instruction verification convolutional neural network model includes an input layer, a feature extraction layer, a global pooling layer, a fully connected layer for instruction recognition, a fully connected layer for identity authentication, and a result output layer, wherein:
[0168] The input layer receives multi-channel Mel spectrogram data and transmits it to the feature extraction layer.
[0169] The feature extraction layer is used to perform two-dimensional convolution on the multi-channel Mel spectrogram data, extracting local time-frequency texture features and global acoustic semantic features from the multi-channel Mel spectrogram data step by step to obtain compressed feature maps, and then transmitting the compressed feature maps to the global pooling layer;
[0170] The global pooling layer is used to perform global average pooling on the compressed feature map to obtain acoustic feature vectors, and then transmits the acoustic feature vectors to the instruction recognition fully connected layer and the identity authentication fully connected layer.
[0171] The instruction recognition fully connected layer is used to map acoustic feature vectors to obtain instruction category probability data, and then transmit the instruction category probability data to the result output layer;
[0172] The fully connected layer for identity authentication is used to map acoustic feature vectors to obtain the probability data of legitimate users and the probability data of illegitimate users, and then transmit the probability data of legitimate users and the probability data of illegitimate users to the result output layer;
[0173] The result output layer is used to obtain the voice command recognition result based on the command category probability data, and to obtain the user identity authentication result based on the legitimate user probability data and the illegitimate user probability data.
[0174] Furthermore, the system may also include a model building module, which can be used for:
[0175] Obtain the preset historical dataset; the historical dataset includes historical multi-channel Mel spectrogram data and historical instruction category labels and historical identity labels corresponding to the historical multi-channel Mel spectrogram data;
[0176] The input layer, feature extraction layer, global pooling layer, instruction recognition fully connected layer, identity authentication fully connected layer, and result output layer are connected to obtain the basic neural network framework. The basic neural network framework is then initialized to obtain the initial neural network.
[0177] The historical multi-channel Mel spectrogram data from the historical dataset are input into the initial neural network to obtain the predicted instruction category data and predicted identity data of the historical multi-channel Mel spectrogram data;
[0178] Based on the predicted instruction category data and historical instruction category labels, the instruction recognition loss value is calculated; based on the predicted identity data and historical identity labels, the identity authentication loss value is calculated.
[0179] Based on the instruction recognition loss value and the identity authentication loss value, the joint loss value is calculated, and the expression for the joint loss value is:
[0180]
[0181] In the formula, It is an index to any historical multi-channel Mel spectrogram data. It is the first The joint loss value of historical multi-channel Mel spectrogram data, It is the instruction loss weight. It is the first Command recognition loss value of historical multi-channel Mel spectrogram data. It is the weight of identity loss. It is the first The authentication loss value of historical multi-channel Mel spectrogram data;
[0182] With the goal of minimizing the joint loss, the initial neural network is iteratively trained until the preset convergence condition is met, thus obtaining the identity instruction verification convolutional neural network model.
[0183] Furthermore, the expression for the frequency domain window function of vocal tract vibration is:
[0184]
[0185] In the formula, It is any frequency value. It is a frequency domain window function for vocal cord vibration. It is the minimum frequency of vocal cord vibration. It is the maximum frequency of vocal cord vibration;
[0186] The minimum frequency of vocal cord vibration is obtained through the following steps:
[0187] Acquire the user's silent motion triaxial acceleration data and silent stationary triaxial acceleration data;
[0188] Fourier transform is performed on the three-axis acceleration data of silent motion to obtain silent motion spectrum data; Fourier transform is performed on the three-axis acceleration data of silent stationary motion to obtain silent stationary spectrum data; the silent motion spectrum data includes silent motion horizontal axis spectrum data, silent motion vertical axis spectrum data, and silent motion vertical axis spectrum data; the silent stationary spectrum data includes silent stationary horizontal axis spectrum data, silent stationary vertical axis spectrum data, and silent stationary vertical axis spectrum data.
[0189] The silent motion energy data is obtained by superimposing the horizontal, vertical, and lateral axis spectral data. The expression for the silent motion energy data is as follows:
[0190]
[0191] In the formula, It is any frequency value. It is the frequency value in silent motion energy data. The value of kinetic energy intensity. The frequency values in the horizontal axis spectrum data of silent motion The transverse frequency intensity, The frequency values in the vertical axis spectrum data of silent motion The vertical axis frequency intensity, The frequency values in the vertical axis spectrum data of silent motion The vertical axis frequency intensity;
[0192] The silent static horizontal axis spectrum data, silent static vertical axis spectrum data, and silent static vertical axis spectrum data are superimposed to obtain the silent static energy data. The expression for the silent static energy data is as follows:
[0193]
[0194] In the formula, It is any frequency value. It is the frequency value in silent static energy data. The static energy intensity value, It is the frequency value in the silent, stationary horizontal axis spectrum data. The transverse frequency intensity, It is the frequency value in the silent, stationary vertical axis spectrum data. The vertical axis frequency intensity, It is the frequency value in the silent, stationary vertical axis spectrum data. The vertical axis frequency intensity;
[0195] The static energy intensity value with the largest value in the silent static energy data is determined as the noise threshold, and the motion energy intensity value of each frequency value in the silent motion energy data is compared with the noise threshold to obtain the noise comparison result of each frequency value.
[0196] The frequency values whose motion energy intensity value is greater than the noise threshold are selected from the noise comparison results to form a noise range set, and the frequency value with the largest value in the noise range set is determined as the minimum value of vocal cord vibration frequency.
[0197] Furthermore, the sampling frequency of the original triaxial acceleration data is 1000Hz, and the original triaxial acceleration data is acquired by a flexible sensor. The flexible sensor is an island-bridge heterogeneous integrated architecture, in which the bridging part is made of liquid metal ink doped with silicon dioxide, and the mass fraction of silicon dioxide is 6wt%. The maximum value of the vocal cord vibration frequency is 255Hz, and the minimum value of the vocal cord vibration frequency can also be directly set to 80Hz.
[0198] In one embodiment, such as Figure 3 A computer device is provided, comprising:
[0199] At least one processor 301, and a memory 302 communicatively connected to at least one of the processors 301: the memory stores application code executable by at least one of the processors, the application code being executed by at least one of the processors to enable at least one of the processors to perform a flexible laryngeal speech interaction method for resisting motion interference as described above.
[0200] Computer equipment may also include: sensor 303.
[0201] The processor 301, memory 302 and sensor 303 can be connected via a bus or other means, with the bus being an example in the figure.
[0202] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the steps in the above method embodiments.
[0203] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to in the description of the method embodiments. The device embodiments described above are merely illustrative. The components described as separate parts may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this disclosure according to actual needs. Those skilled in the art can understand and implement this without creative effort.
[0204] The above-described embodiments are merely illustrative of several implementation methods of the embodiments of this application, and their descriptions are relatively specific and detailed. However, they should not be construed as limiting the scope of the patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the embodiments of this application, and these modifications and improvements all fall within the protection scope of the embodiments of this application.
Claims
1. A flexible laryngeal voice interaction method for resisting motion interference, characterized in that, The method includes: Acquire the user's raw triaxial acceleration data; the raw triaxial acceleration data is used to characterize the three-dimensional mechanical vibration of the user's laryngeal skin during vocalization; The original triaxial acceleration data is denoised to obtain clean triaxial vocal cord vibration data; Time-frequency domain feature extraction was performed on the pure triaxial vocal cord vibration data to obtain multi-channel Mel spectrum data; The multi-channel Mel-frequency spectrogram data is input into the identity command verification convolutional neural network model to obtain the voice command recognition result and the user identity authentication result; the identity command verification convolutional neural network model has a dual-branch output structure.
2. The method according to claim 1, characterized in that, The original triaxial acceleration data includes original horizontal axis acceleration data, original vertical axis acceleration data, and original vertical axis acceleration data. The denoising process performed on the original triaxial acceleration data to obtain clean triaxial vocal cord vibration data includes: Perform a Fourier transform on the original horizontal axis acceleration data to obtain frequency domain horizontal axis acceleration data; perform a Fourier transform on the original vertical axis acceleration data to obtain frequency domain vertical axis acceleration data; perform a Fourier transform on the original vertical axis acceleration data to obtain frequency domain vertical axis acceleration data. The frequency domain horizontal axis acceleration data and the vocal tract vibration frequency domain window function are multiplied to obtain the denoised frequency domain horizontal axis data; the frequency domain vertical axis acceleration data and the vocal tract vibration frequency domain window function are multiplied to obtain the denoised frequency domain vertical axis data; the frequency domain vertical axis acceleration data and the vocal tract vibration frequency domain window function are multiplied to obtain the denoised frequency domain vertical axis data. Perform an inverse Fourier transform on the denoised frequency domain horizontal axis data to obtain clean horizontal axis vibration data; perform an inverse Fourier transform on the denoised frequency domain vertical axis data to obtain clean vertical axis vibration data; perform an inverse Fourier transform on the denoised frequency domain vertical axis data to obtain clean vertical axis vibration data. The pure triaxial vocal cord vibration data is composed of the pure horizontal axis vibration data, the pure vertical axis vibration data, and the pure vertical axis vibration data.
3. The method according to claim 2, characterized in that, The time-frequency domain feature extraction of the pure triaxial vocal tract vibration data yields multi-channel Mel-ray spectrogram data, including: Based on the preset segmentation interval parameter, the pure horizontal axis vibration data is segmented to obtain a horizontal axis vibration segmented dataset; based on the segmentation interval parameter, the pure vertical axis vibration data is segmented to obtain a vertical axis vibration segmented dataset; based on the segmentation interval parameter, the pure vertical axis vibration data is segmented to obtain a vertical axis vibration segmented dataset. A short-time Fourier transform is performed on each of the horizontal axis vibration segment data in the horizontal axis vibration segment dataset to obtain a horizontal axis spectrum sequence; a short-time Fourier transform is performed on each of the vertical axis vibration segment data in the vertical axis vibration segment dataset to obtain a vertical axis spectrum sequence; a short-time Fourier transform is performed on each of the vertical axis vibration segment data in the vertical axis vibration segment dataset to obtain a vertical axis spectrum sequence. Based on a preset Mel filter, Mel mapping is performed on each of the horizontal axis spectrum sequences to obtain a horizontal axis Mel spectrum; based on the Mel filter, Mel mapping is performed on each of the vertical axis spectrum sequences to obtain a vertical axis Mel spectrum; based on the Mel filter, Mel mapping is performed on each of the vertical axis spectrum sequences to obtain a vertical axis Mel spectrum. The horizontal axis Mel spectrum, the vertical axis Mel spectrum, and the vertical axis Mel spectrum are stacked to obtain the multi-channel Mel spectrum data.
4. The method according to claim 3, characterized in that, The identity instruction verification convolutional neural network model includes an input layer, a feature extraction layer, a global pooling layer, a fully connected layer for instruction recognition, a fully connected layer for identity authentication, and a result output layer, wherein: The input layer is used to receive the multi-channel Mel spectrogram data and transmit the multi-channel Mel spectrogram data to the feature extraction layer; The feature extraction layer is used to perform two-dimensional convolution on the multi-channel Mel spectrogram data, extract local time-frequency texture features and global acoustic semantic features from the multi-channel Mel spectrogram data step by step to obtain a compressed feature map, and transmit the compressed feature map to the global pooling layer; The global pooling layer is used to perform global average pooling on the compressed feature map to obtain an acoustic feature vector, and then transmits the acoustic feature vector to the instruction recognition fully connected layer and the identity authentication fully connected layer. The instruction recognition fully connected layer is used to map the acoustic feature vector to obtain instruction category probability data, and transmit the instruction category probability data to the result output layer; The fully connected layer for identity authentication is used to map the acoustic feature vector to obtain the probability data of legitimate users and the probability data of illegitimate users, and to transmit the probability data of legitimate users and the probability data of illegitimate users to the result output layer. The result output layer is used to obtain the voice command recognition result based on the command category probability data, and to obtain the user identity authentication result based on the legitimate user probability data and the illegitimate user probability data.
5. The method according to claim 4, characterized in that, The identity command verification convolutional neural network model is obtained through the following steps: Obtain a preset historical dataset; the historical dataset includes historical multi-channel Mel spectrogram data and historical instruction category labels and historical identity labels corresponding to the historical multi-channel Mel spectrogram data; The input layer, the feature extraction layer, the global pooling layer, the instruction recognition fully connected layer, the identity authentication fully connected layer, and the result output layer are connected to obtain a basic neural network framework, and the basic neural network framework is initialized to obtain an initial neural network. The historical multi-channel Mel spectrogram data from the historical dataset are input into the initial neural network to obtain the predicted instruction category data and predicted identity data of the historical multi-channel Mel spectrogram data; Based on the predicted instruction category data and the historical instruction category labels, the instruction recognition loss value is calculated; based on the predicted identity data and the historical identity labels, the identity authentication loss value is calculated. Based on the instruction recognition loss value and the identity authentication loss value, a joint loss value is calculated, and the expression for the joint loss value is: In the formula, It is an index to any historical multi-channel Mel spectrogram data. It is the first The joint loss value of historical multi-channel Mel spectrogram data, It is the instruction loss weight. It is the first Command recognition loss value of historical multi-channel Mel spectrogram data. It is the weight of identity loss. It is the first The authentication loss value of historical multi-channel Mel spectrogram data; With the goal of minimizing the joint loss value, the initial neural network is iteratively trained until a preset convergence condition is met, thereby obtaining the identity instruction verification convolutional neural network model.
6. The method according to claim 2, characterized in that, The expression for the frequency domain window function of the vocal tract vibration is: In the formula, It is any frequency value. It is a frequency domain window function for vocal cord vibration. It is the minimum frequency of vocal cord vibration. It is the maximum frequency of vocal cord vibration; The minimum value of the vocal cord vibration frequency is obtained through the following steps: Obtain the user's silent motion triaxial acceleration data and silent stationary triaxial acceleration data; Perform a Fourier transform on the silent motion triaxial acceleration data to obtain silent motion spectrum data; perform a Fourier transform on the silent stationary triaxial acceleration data to obtain silent stationary spectrum data; the silent motion spectrum data includes silent motion horizontal axis spectrum data, silent motion vertical axis spectrum data, and silent motion vertical axis spectrum data; the silent stationary spectrum data includes silent stationary horizontal axis spectrum data, silent stationary vertical axis spectrum data, and silent stationary vertical axis spectrum data. The silent motion horizontal axis spectrum data, the silent motion vertical axis spectrum data, and the silent motion vertical axis spectrum data are superimposed to obtain silent motion energy data. The expression for the silent motion energy data is as follows: In the formula, It is any frequency value. It is the frequency value in silent motion energy data. The value of kinetic energy intensity. The frequency values in the horizontal axis spectrum data of silent motion The transverse frequency intensity, The frequency values in the vertical axis spectrum data of silent motion The vertical axis frequency intensity, It is the frequency value in the vertical axis spectrum data of silent motion. The vertical axis frequency intensity; The silent static horizontal axis spectrum data, the silent static vertical axis spectrum data, and the silent static vertical axis spectrum data are superimposed to obtain silent static energy data. The expression for the silent static energy data is as follows: In the formula, It is any frequency value. It is the frequency value in silent static energy data. The static energy intensity value, It is the frequency value in the silent, stationary horizontal axis spectrum data. The transverse frequency intensity, It is the frequency value in the silent, stationary vertical axis spectrum data. The vertical axis frequency intensity, It is the frequency value in the silent, stationary vertical axis spectrum data. The vertical axis frequency intensity; The static energy intensity value with the largest value in the silent static energy data is determined as the noise threshold, and the motion energy intensity value of each frequency value in the silent motion energy data is compared with the noise threshold to obtain the noise comparison result of each frequency value. The frequency values in which the motion energy intensity value is greater than the noise threshold are selected as the noise range set, and the frequency value with the largest value in the noise range set is determined as the minimum value of the vocal cord vibration frequency.
7. The method according to claim 6, characterized in that, The sampling frequency of the original triaxial acceleration data is 1000Hz, and the original triaxial acceleration data is acquired by a flexible sensor. The flexible sensor is an island-bridge heterogeneous integrated architecture, wherein the bridging part is made of liquid metal ink doped with silicon dioxide, and the mass fraction of silicon dioxide is 6wt%. The maximum value of the vocal cord vibration frequency is 255Hz, and the minimum value of the vocal cord vibration frequency can also be directly set to 80Hz.
8. A flexible laryngeal voice interaction system for resisting motion interference, characterized in that, The system includes: The data acquisition module is used to acquire the user's raw triaxial acceleration data; the raw triaxial acceleration data is used to characterize the three-dimensional mechanical vibration of the user's laryngeal skin during vocalization; The noise reduction module is used to perform noise reduction processing on the original triaxial acceleration data to obtain clean triaxial vocal cord vibration data; The feature extraction module is used to extract time-frequency domain features from the pure triaxial vocal cord vibration data to obtain multi-channel Mel spectrogram data; The results analysis module is used to input the multi-channel Mel spectrogram data into the identity command verification convolutional neural network model to obtain the voice command recognition result and the user identity authentication result; the identity command verification convolutional neural network model has a dual-branch output structure.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 6.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 6.