Voice interaction method and system based on intelligent doll
By collecting and processing neck tissue vibration signals and ambient audio signals, combining biomechanical modeling and environmental noise compensation, high-precision speech signals are generated, which solves the problems of inaccurate speech generation modeling and insufficient adaptability, and achieves personalized and comfortable voice interaction.
Patent Information
- Application Number
- CN202510800152.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-16
- Publication Date
- 2025-08-12
- Estimated Expiration
- 2045-06-16
AI Technical Summary
In the prior art, the speech generation modeling is not accurate enough and the speech generation adaptability is insufficient, the coupling relationship between the biomechanical characteristics of neck tissue and the vibration mode of the vocal cords is not fully considered, and the dynamic compensation mechanism for environmental noise is lacked.
By collecting vibration displacement signals and ambient audio signals of the user's neck tissue, the vibration characteristic matrix and acoustic characteristic matrix are generated, and the modal parameter set of vocal cord vibration is generated in combination with biomechanical modeling, and environmental noise compensation is performed. The piezoelectric speaker array is used to form a directional focused sound field, record user physiological feedback, and update the vibration-acoustic model through reinforcement learning strategies.
It improves the accuracy and noise immunity of voice signal acquisition, enhances the personalization and comfort of voice interaction, and adapts to complex and changeable usage scenarios.
Smart Images

Figure CN120472901A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of intelligent acoustic interaction technology, and in particular to a voice interaction method and system based on an intelligent doll. Background Art
[0002] With the rapid development of artificial intelligence and human-computer interaction technologies, voice, as one of the most natural and direct forms of communication, has gradually become a core interactive tool for smart devices. In particular, in areas such as children's education, emotional companionship, and elderly care, voice-based interactive smart doll systems have attracted widespread attention due to their anthropomorphic features and high immersiveness. Traditional speech recognition and synthesis technologies primarily rely on microphones to capture ambient audio signals and apply deep neural networks and other methods for speech enhancement and semantic understanding.
[0003] Although the signal-to-noise ratio of speech acquisition has been improved to a certain extent, there are still several key technical bottlenecks: First, existing technologies generally do not fully consider the coupling relationship between the biomechanical properties of the neck tissue and the vibration mode of the vocal cords, resulting in inaccurate modeling of the speech generation mechanism; second, there is a lack of dynamic compensation mechanism for environmental noise during the speech generation process, which makes it difficult to adapt to complex and changeable actual usage scenarios. Summary of the Invention
[0004] In view of the above existing problems, the present invention is proposed.
[0005] Therefore, the present invention provides a voice interaction method based on an intelligent doll to solve the problems of inaccurate voice generation modeling and insufficient adaptability of voice generation in the prior art.
[0006] In order to solve the above technical problems, the present invention provides the following technical solutions:
[0007] In a first aspect, the present invention provides a voice interaction method based on a smart doll, which includes collecting vibration displacement signals and environmental audio signals of the user's neck tissue, and preprocessing them to generate a vibration feature matrix and an acoustic feature matrix; performing time-frequency analysis on the vibration feature matrix, combining the density characteristics of the human neck tissue, and generating a modal parameter set of vocal cord vibration through biomechanical modeling; inputting the modal parameter set of vocal cord vibration into a vibration-acoustic model to generate a sound source excitation field, and using the acoustic feature matrix to compensate for environmental noise in the sound source excitation field, and generating a time-domain pure voice signal through the sound wave equation; collecting the user's real-time position coordinates, calculating and driving the piezoelectric speaker array of the smart doll to adjust the phase according to the spectral characteristics of the time-domain pure voice signal, forming a directional focused sound field and recording the user's physiological feedback to generate a user physiological feedback matrix; generating a standardized feedback vector based on the auricle temperature change rate and voice clarity index in the user's physiological feedback matrix, and updating the vibration-acoustic model through a reinforcement learning strategy.
[0008] As a preferred solution of the voice interaction method based on the smart doll of the present invention, the steps of generating the vibration feature matrix and the acoustic feature matrix are as follows:
[0009] The Doppler effect is used to capture the vibration displacement signal of the user's neck tissue, synchronously collect the ambient audio signal, and eliminate the interference of blood vessels and muscle fluctuations to generate pure vocal cord vibration waves;
[0010] Perform multi-scale dynamic decomposition on pure vocal cord vibration waves and generate time-varying feature weights based on the conduction characteristics of biological tissues;
[0011] The time-varying feature weights are compensated for acoustic impedance matching with the ambient audio signal to generate a vibration feature matrix and an acoustic feature matrix.
[0012] As a preferred solution of the voice interaction method based on the smart doll of the present invention, the specific steps of generating the modal parameter set of vocal cord vibration are as follows:
[0013] Based on the vibration characteristic matrix, a time-varying window is dynamically selected, and the local wave energy density function is generated through nonlinear integration and mapped into a wave energy topology matrix that characterizes the tissue vibration curvature.
[0014] The density characteristics of human neck tissue are converted into biomechanical constraints of mass conservation. Combined with the wave energy topology matrix, the displacement mode function vector is iteratively solved through the augmented Lagrangian algorithm.
[0015] The displacement modal function vector is used to guide the neural network to compensate for the residual, and the modal parameter set of vocal fold vibration is analytically generated after fusion.
[0016] As a preferred solution of the voice interaction method based on the smart doll described in the present invention, wherein: the modal parameter set of the vocal cord vibration is input into the vibration-acoustic model to generate the sound source excitation field, and the acoustic characteristic matrix is used to compensate the sound source excitation field for environmental noise, and the time domain pure voice signal is generated by the acoustic wave equation. The specific steps are as follows:
[0017] The modal parameter set of vocal cord vibration is input into the vibro-acoustic model to generate the sound source excitation field;
[0018] The sound source excitation field and the acoustic characteristic matrix are input into the vibration-acoustic model. The acoustic wave equation is solved and the ambient noise field is compensated synchronously through the deep coupling operator to generate the sound pressure field of the sound channel.
[0019] The sound pressure field of the vocal tract is input into the radiation reconstruction process of the vibration-acoustic model, and the time-domain pure speech signal is generated by integrating the acoustic impedance compensation function on the lip surface.
[0020] As a preferred solution of the voice interaction method based on the smart doll of the present invention, the steps of collecting the user's real-time position coordinates, calculating and driving the piezoelectric speaker array of the smart doll to adjust the phase according to the spectrum characteristics of the pure voice signal in the time domain are as follows:
[0021] Perform time-frequency transformation on the time-domain pure speech signal to generate a time-frequency distribution matrix;
[0022] The user's real-time location coordinates are collected, combined with the time-frequency distribution matrix, and the sound wave phase adjustment amount is calculated through quantum optimization processing. The piezoelectric speaker array is driven to adjust the phase according to the sound wave phase adjustment amount.
[0023] As a preferred solution of the voice interaction method based on the smart doll of the present invention, wherein: the forming of the directional focused sound field and recording the user's physiological feedback to generate the user's physiological feedback matrix, the specific steps are as follows:
[0024] Generate a directional focused sound field based on a phase-adjusted piezoelectric speaker array;
[0025] Based on the directional focused sound field, the user's auricle temperature field change rate is collected while receiving the actual voice signal;
[0026] The user physiological feedback matrix of the acoustic-thermal coupling effect is calculated based on the auricle temperature field change rate and the actual speech signal.
[0027] As a preferred solution of the voice interaction method based on the smart doll of the present invention, wherein: based on the auricle temperature change rate and the speech clarity index in the user's physiological feedback matrix, a standardized feedback vector is generated, and the vibration-acoustic model is updated through a reinforcement learning strategy. The specific steps are as follows:
[0028] Normalize the user's physiological feedback matrix to generate a standardized feedback vector;
[0029] Based on the standardized feedback vector and the vibration-acoustic model, a Markov decision process state space is constructed and input into the TD3 reinforcement learning strategy network to generate vibration-acoustic model optimization instructions;
[0030] Physiological safety constraints are verified for the vibration-acoustic model optimization instructions, and the vibration-acoustic model is iteratively updated.
[0031] In a second aspect, the present invention provides a voice interaction system based on an intelligent doll, comprising a signal acquisition module, a modal analysis module, a voice reconstruction module, a sound field focusing module, and a model optimization module. The signal acquisition module is used to collect the vibration displacement signal of the user's neck tissue and the ambient audio signal, and perform preprocessing to generate a vibration feature matrix and an acoustic feature matrix; the modal analysis module is used to perform time-frequency analysis on the vibration feature matrix, combine the density characteristics of the human neck tissue, and generate a modal parameter set of the vocal cord vibration through biomechanical modeling; the voice generation module is used to input the modal parameter set of the vocal cord vibration into the vibration- The acoustic model generates a sound source excitation field, uses the acoustic feature matrix to compensate for ambient noise in the sound source excitation field, and generates a time-domain pure speech signal through the sound wave equation; the sound field focusing module is used to collect the user's real-time position coordinates, calculate and drive the piezoelectric speaker array of the smart doll to adjust the phase according to the spectral characteristics of the time-domain pure speech signal, form a directionally focused sound field, record the user's physiological feedback, and generate a user physiological feedback matrix; the model optimization module is used to generate a standardized feedback vector based on the auricle temperature change rate and speech clarity index in the user's physiological feedback matrix, and update the vibration-acoustic model through a reinforcement learning strategy.
[0032] In a third aspect, the present invention provides a computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: when the computer program is executed by the processor, any step of the voice interaction method based on the smart doll as described in the first aspect of the present invention is implemented.
[0033] In a fourth aspect, the present invention provides a computer-readable storage medium having a computer program stored thereon, wherein: when the computer program is executed by a processor, any step of the voice interaction method based on the smart doll as described in the first aspect of the present invention is implemented.
[0034] The beneficial effects of the present invention are as follows: by solving the vocal cord displacement modal function vector through nonlinear integration and augmented Lagrangian algorithm, physical modeling of the vocal cord vibration state is realized, thereby improving the acquisition accuracy and noise resistance of the voice signal at the source; further, by collecting the user's auricle temperature change rate and speech clarity index to construct a user physiological feedback matrix, adaptive optimization can be performed according to the user's physiological response and speech perception quality, further improving the personalization and comfort of voice interaction. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0036] Figure 1 The figure is a flow chart of the voice interaction method based on the smart doll.
[0037] Figure 2 Flowchart for generating the vocal fold vibration modal parameter set.
[0038] Figure 3 Flowchart for generating time-domain clean speech signal.
[0039] Figure 4 Flowchart for closed-loop optimization of a vibro-acoustic model driven by physiological feedback. DETAILED DESCRIPTION
[0040] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the specific embodiments of the present invention are described in detail below with reference to the accompanying drawings.
[0041] In the following description, many specific details are set forth to facilitate a full understanding of the present invention. However, the present invention may also be implemented in other ways different from those described herein. Those skilled in the art may make similar generalizations without violating the connotation of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below.
[0042] Secondly, the term "one embodiment" or "embodiment" herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The phrase "in one embodiment" appearing in various places throughout this specification does not necessarily refer to the same embodiment, nor does it refer to a separate or selective embodiment that is mutually exclusive of other embodiments.
[0043] Reference Figures 1 to 4 , is an embodiment of the present invention, which provides a voice interaction method based on an intelligent doll, comprising the following steps:
[0044] S1: Collect the vibration displacement signal of the user's neck tissue and the ambient audio signal, and preprocess them to generate the vibration feature matrix and the acoustic feature matrix.
[0045] S1.1: The Doppler effect is used to capture the vibration displacement signal of the user's neck tissue, synchronously collect the ambient audio signal, and eliminate the interference of blood vessel and muscle fluctuations to generate pure vocal cord vibration waves.
[0046] The specific process includes: when the Doppler effect captures the vibration displacement signal of the user's neck tissue, it emits ultrasonic waves of a specific frequency and receives the reflected wave, and analyzes the vibration velocity and displacement amplitude of the tissue surface by analyzing the frequency offset of the reflected wave. The ambient audio signal is synchronously collected through a high-sensitivity microphone array, and the adaptive filtering adopts the normalized least mean square algorithm. The ambient audio signal is used as the reference signal to eliminate the low-frequency mechanical vibration interference caused by blood vessel pulsation and muscle contraction from the mixed signal, and the 20Hz to 1000Hz frequency band signal that matches the vocal cord vibration frequency is retained through band-pass filtering, and finally a pure vocal cord vibration wave is output.
[0047] S1.2: Perform multi-scale dynamic decomposition of pure vocal cord vibration waves and generate time-varying feature weights based on the conduction characteristics of biological tissues.
[0048] The specific process involves decomposing the pure vocal cord vibration wave through wavelet packets into sub-signal components of different frequency bands. The neck tissue density characteristics and acoustic impedance parameters are used to analyze the propagation characteristics of sound waves in biological tissue and determine the attenuation degree of the signals in each frequency band. The sub-signals of each frequency band are dynamically weighted according to the attenuation degree, with the high-frequency component receiving a lower weight due to tissue absorption and the low-frequency component receiving a higher weight. This ultimately generates time-varying feature weights that reflect the conduction characteristics of biological tissue. The time-varying feature weights act directly on the vibration feature matrix. The optimized multi-scale vibration features retain key information about vocal cord vibration.
[0049] The conduction characteristics of biological tissue refer to the physical properties such as frequency-dependent attenuation law, acoustic impedance distribution and energy absorption efficiency exhibited by sound waves when propagating in the human neck tissue. They are specifically determined by tissue density, elastic modulus and viscosity coefficient.
[0050] S1.3: Perform acoustic impedance matching compensation on the time-varying feature weights and the ambient audio signal to generate a vibration feature matrix and an acoustic feature matrix.
[0051] The specific process includes processing the time-varying feature weights and ambient audio through acoustic impedance matching compensation. The time-varying feature weights reflect the propagation characteristics of sound waves in different frequency bands in the neck tissue. The ambient audio contains spatial acoustic information. The two are matched using sound field reconstruction technology. The time-varying feature weights guide the redistribution of energy in each frequency band of the ambient audio and compensate for the frequency response distortion caused by tissue conduction. The ambient audio is subjected to Mel-frequency cepstrum analysis to extract spectral features. The time-varying feature weights are multiplied by the spectral features frequency-by-frequency point to achieve acoustic impedance matching. The vibration feature matrix is generated by weighted fusion of the multi-scale components of the pure vocal cord vibration wave through the time-varying feature weights. The acoustic feature matrix is generated by the compensated ambient audio features. The vibration feature matrix and the acoustic feature matrix maintain time synchronization.
[0052] S2: Perform time-frequency analysis on the vibration characteristic matrix, combine it with the density characteristics of human neck tissue, and generate the modal parameter set of vocal cord vibration through biomechanical modeling.
[0053] S2.1: Based on the vibration characteristic matrix, a time-varying window is dynamically selected, and the local wave energy density function is generated through nonlinear integration and mapped into a wave energy topology matrix that characterizes the tissue vibration curvature.
[0054] The specific process includes: determining the active period of the signal through short-time energy analysis of the vibration characteristic matrix, adaptively selecting the Hanning window as the time-varying window within the active period, and the width of the time-varying window is inversely proportional to the instantaneous frequency of the vibration characteristic matrix. A narrower window is used in the high-frequency region to improve the time resolution, and a wider window is used in the low-frequency region to ensure the frequency resolution. The signal component after window selection is extracted through Hilbert transform, and the instantaneous amplitude is squared and integrated along the time axis to generate a local wave energy density function. The local wave energy density function is Gaussian smoothed to eliminate sudden interference. The smoothed wave energy density distribution is processed by the curvature operator. The curvature operator uses second-order difference to approximate the curvature value of each time-frequency point. Positive curvature corresponds to the wave energy convergence area, and negative curvature corresponds to the wave energy divergence area. Finally, a wave energy topology matrix reflecting the tissue vibration energy distribution characteristics is generated.
[0055] S2.2: The density characteristics of human neck tissue are converted into biomechanical constraints of mass conservation. Combined with the wave energy topology matrix, the displacement modal function vector is iteratively solved by the augmented Lagrangian algorithm. The expression is:
[0056]
[0057] Among them, u (k+1) represents the displacement mode function vector solved by the k+1th iteration, Indicates finding the minimum value of the objective function for the displacement function vector u, u represents the displacement function vector, k represents the sequence number of the current iteration step, represents the transpose of the displacement function vector, K represents the stiffness matrix, T represents the wave energy topology matrix, μ (k) represents the penalty coefficient of the kth iteration, A represents the mass conservation constraint matrix, represents the transpose of the mass conservation constraint matrix, g represents the external force vector, and λ (k) represents the Lagrange multiplier vector of the kth iteration, and b represents the constant term vector of the mass conservation constraint.
[0058] The specific process includes: the density characteristics of human neck tissue are converted into a mass distribution matrix through finite element discretization processing; the mass distribution matrix is jointly constructed with the continuous medium mechanics equation to construct the mass conservation constraint condition; the mass conservation constraint condition is expressed in the form of a matrix equation; the wave energy topology matrix is extracted through eigenvalue decomposition to extract the dominant vibration mode; the dominant vibration mode and the mass conservation constraint condition together constitute the objective function of the optimization problem; the augmented Lagrangian algorithm adds the mass conservation constraint condition as a penalty term to the objective function; the Lagrangian multiplier vector and the penalty coefficient are updated in each iteration; the linear equation group is solved by the conjugate gradient method to obtain the displacement modal function vector; the coefficient matrix of the linear equation group consists of the stiffness matrix, the wave energy topology matrix and the penalty term; the right-hand term contains the external force vector and the constraint compensation term; the iteration termination condition is set as the Euclidean distance between two adjacent displacement modal function vectors is less than the convergence tolerance factor; the final output displacement modal function vector satisfies the mass conservation constraint and reflects the wave energy distribution characteristics.
[0059] The convergence tolerance factor is set according to the physical magnitude of the displacement modal function vector and the calculation accuracy requirements. The specific value is 1% of the initial residual norm and is determined by grid independence verification.
[0060] S2.3: Use the displacement modal function vector to guide the neural network to compensate for the residual, and then analytically generate the modal parameter set of vocal fold vibration after adaptive fusion.
[0061] The specific process includes: the displacement modal function vector is input as the input feature into the pre-trained three-layer feedforward neural network; the three-layer feedforward neural network outputs the residual compensation coefficient through the sigmoid activation function; the residual compensation coefficient is multiplied element-by-element with the displacement modal function vector to realize residual compensation; the compensated displacement modal function vector is extracted from the principal component through singular value decomposition; the principal component is adaptively weighted fused with the vocal cord anatomical structure characteristics; the weighting coefficient is dynamically determined by the energy proportion of the displacement modal function vector; the fused feature vector is converted into the vocal cord vibration modal parameters through the nonlinear mapping layer; the vocal cord vibration modal parameters include key features such as fundamental frequency, vibration amplitude, and phase difference; the final generated vocal cord vibration modal parameter set is consistent with the biomechanical characteristics.
[0062] Furthermore, the specific training process of the pre-trained three-layer feedforward neural network is as follows: first, the network weights and biases are initialized to ensure that the number of nodes in the input layer, hidden layer and output layer matches the task requirements; then, the displacement modal function vector is input into the network, and the weighted sum of the hidden layer and the output layer and the Sigmoid activation function are analyzed layer by layer to obtain the predicted value of the residual compensation coefficient; then, the mean square error (MSE) loss function is used to compare the predicted value with the true label, analyze the error and start back propagation: starting from the output layer, the chain rule is used to obtain the gradient of the weights and biases layer by layer; finally, the weights and biases are updated by the gradient descent method, and the forward propagation, error analysis and back propagation processes are repeated until the loss converges or the preset number of iterations is reached. During this period, early stopping or regularization may be used to prevent overfitting, and the generalization performance is evaluated through the validation set.
[0063] S3: The modal parameter set of vocal cord vibration is input into the vibration-acoustic model to generate the sound source excitation field. The acoustic characteristic matrix is used to compensate the sound source excitation field for ambient noise, and a time-domain pure speech signal is generated through the acoustic wave equation.
[0064] S3.1: Input the modal parameter set of vocal fold vibration into the vibro-acoustic model to generate the sound source excitation field.
[0065] The specific process includes: the modal parameter set of vocal cord vibration is converted into sound source boundary conditions through parameter mapping relationship. The sound source boundary conditions include vibration amplitude, vibration frequency and phase information. The vibration-acoustic model establishes the law of sound wave propagation based on the linearized Euler equation. The sound source boundary conditions are applied to the glottis position as the excitation source. The sound wave equation is solved by the finite difference method to obtain the sound pressure distribution. The sound pressure distribution forms a sound source excitation field on the discrete grid nodes in three-dimensional space. The time domain change of the sound source excitation field reflects the vibration characteristics of the vocal cords, and the spatial distribution conforms to the anatomical structure characteristics of the vocal tract. The final output sound source excitation field contains complete spatiotemporal modulation information.
[0066] Furthermore, the vibro-acoustic model training process first collects a large amount of vocal fold vibration data and corresponding sound field measurement data as training samples. The vocal fold vibration data is acquired synchronously using high-speed video and laser vibrometers, while the sound field measurement data is recorded in an anechoic chamber using a piezoelectric microphone array. After preprocessing, the training samples are divided into a vocal fold vibration feature set and a sound field feature set. The vocal fold vibration feature set includes parameters such as fundamental frequency, vibration amplitude, and phase difference, while the sound field feature set includes sound pressure level distribution and spectral characteristics. The vibro-acoustic model adopts a deep neural network architecture. The input layer receives the vocal fold vibration features, the hidden layer includes three-dimensional convolutional layers and long short-term memory layers to capture spatial and temporal features, and the output layer predicts the sound field distribution. During the training process, mean squared error is used as the loss function. The network weights are optimized through the back-propagation algorithm, and the learning rate is dynamically adjusted using the Adam optimizer. The number of training iterations is determined based on the convergence of the error on the validation set. The resulting vibro-acoustic model can accurately predict the sound field distribution under given vocal fold vibration parameters.
[0067] Preprocessing includes data denoising (such as wavelet denoising, Kalman filtering), signal segmentation and alignment (frame windowing), feature extraction (fundamental frequency, spectral energy), data normalization (amplitude / time standardization) and quality control (residual analysis).
[0068] S3.2: The sound source excitation field and the acoustic characteristic matrix are input into the vibration-acoustic model. The acoustic wave equation solution and the ambient noise field compensation are synchronously performed through the deep coupling operator to generate the sound channel sound pressure field.
[0069] The specific process includes: the sound source excitation field and the acoustic characteristic matrix are spliced together to form a joint input feature through tensor splicing; the deep coupling operator uses the alternating direction multiplication method to synchronize the solution of the acoustic wave equation with the ambient noise field compensation process; the acoustic wave equation is established based on the linearized Euler equation and is numerically solved using the pseudo-spectral method; the ambient noise field compensation constructs the noise covariance matrix through the acoustic characteristic matrix and applies it to the source term of the acoustic wave equation; the sound pressure field and the noise compensation coefficient are updated alternately during the internal iteration of the deep coupling operator; each iteration first fixes the noise compensation coefficient to solve the acoustic wave equation to obtain the transient sound pressure field distribution, and then updates the noise compensation coefficient based on the residual between the transient sound pressure field distribution and the acoustic characteristic matrix; the iterative termination condition is that the relative error between two adjacent sound pressure field updates is less than the sound field convergence criterion; the final output sound channel sound pressure field meets both the physical laws of sound wave propagation and the requirements of ambient noise suppression; the spatial resolution of the sound channel sound pressure field is consistent with the input sound source excitation field, and the time sampling rate is synchronized with the acoustic characteristic matrix.
[0070] The acoustic field convergence criterion is set according to the numerical stability requirements of the sound pressure field solution. The specific value is 0.1% of the initial sound pressure field energy and is determined through acoustic simulation error analysis.
[0071] S3.3: Input the vocal tract sound pressure field into the radiation reconstruction process of the vibration-acoustic model, and generate a time-domain pure speech signal by integrating the acoustic impedance compensation function over the lip surface.
[0072] The specific process includes: the vocal tract sound pressure field is discretized through the lip surface grid, the node coordinates of the lip surface grid are aligned with the spatial sampling points of the vocal tract sound pressure field, and an acoustic impedance compensation function is established based on the acoustic characteristics of the lip tissue to reflect the change in radiation impedance of the sound wave from the vocal tract to the free field. The acoustic impedance compensation function is expressed as a complex transfer function in the frequency domain. The real part of the complex transfer function corresponds to the acoustic resistance component, and the imaginary part corresponds to the acoustic reactance component. The frequency domain distribution of the vocal tract sound pressure field is multiplied point by point by the acoustic impedance compensation function to achieve frequency domain compensation. The compensated frequency domain sound pressure is converted into a time domain signal through an inverse Fourier transform. The sampling rate of the time domain signal is consistent with the original vocal cord vibration signal. The generated time domain pure speech signal is smoothed by a Hanning window to eliminate the frequency domain leakage effect. The final output time domain pure speech signal retains the essential characteristics of the vocal cord vibration and eliminates the distortion caused by radiation impedance.
[0073] S4: Collect the user's real-time location coordinates, calculate and drive the smart doll's piezoelectric speaker array to adjust the phase based on the spectral characteristics of the pure voice signal in the time domain, form a directional focused sound field, record the user's physiological feedback, and generate a user physiological feedback matrix.
[0074] S4.1: Perform time-frequency transformation on the time-domain clean speech signal to generate a time-frequency distribution matrix.
[0075] The specific process includes: the doll's visual positioning device captures the coordinates of the user's facial feature points through a binocular stereo vision algorithm; the user's real-time position is converted into relative position coordinates centered on the doll through coordinate transformation; the time-domain pure voice signal is processed by a windowed short-time Fourier transform; the window length of the Hanning window is dynamically adjusted according to the fundamental frequency of the voice signal; the window function sliding step is set to one-quarter of the window length to ensure time-frequency continuity; the short-time Fourier transform converts the time-domain pure voice signal into a time-frequency distribution matrix in complex form; the rows of the time-frequency distribution matrix correspond to time frames, the columns correspond to frequency points, the amplitude of the matrix elements represents the signal energy, and the phase represents the signal phase information; the time resolution of the time-frequency distribution matrix is synchronized with the refresh rate of the visual positioning device, and the frequency resolution is determined by the window function length and the sampling rate; the generated time-frequency distribution matrix keeps the timestamp aligned with the user's real-time position data.
[0076] S4.2: Collect the user's real-time location coordinates, combine them with the time-frequency distribution matrix, calculate the sound wave phase adjustment through quantum optimization processing, and drive the piezoelectric speaker array to adjust the phase based on the sound wave phase adjustment. The expression is:
[0077]
[0078] Where ΔΦm represents the acoustic wave phase adjustment of the mth piezoelectric speaker array, α represents the wavelength of the acoustic wave, ||r|| represents the Euclidean distance from the user to the geometric center of the piezoelectric speaker array, m represents the index number of the piezoelectric speaker array, δ m represents the calibration distance of the mth piezoelectric speaker array reference point, β represents the frequency response balance coefficient (0.2≤β≤1.0), f1 represents the lower limit frequency of the effective frequency band of the speech signal, f2 represents the upper limit frequency of the effective frequency band of the speech signal, f represents the frequency component of the effective frequency band of the speech signal, W(f) represents the perceptual weighted value of the frequency component f of the effective frequency band of the speech signal, sinc represents the normalized sine integral function, f3 represents the dynamic center frequency of the current processing frame, Δf represents the frequency resolution, t represents the time frame number of the current processing, and θ(t,f) represents the instantaneous phase angle of the time-frequency matrix at time frame t and the frequency component f of the effective frequency band of the speech signal.
[0079] The specific process includes mapping the user's real-time position to a spatial coordinate system with the center of the piezoelectric speaker array as the origin through polar coordinate transformation, and performing parallel optimization calculations on the time-frequency distribution matrix through a quantum annealing processor. The quantum annealing processor models the acoustic phase adjustment problem as a quadratic unconstrained binary optimization problem. The problem Hamiltonian contains a geometric path difference term and a spectral coherence term. The geometric path difference term is determined by the three-dimensional Euclidean distance difference between the user's real-time position and each piezoelectric speaker array unit. The spectral coherence term is constructed using the phase gradient information of the time-frequency distribution matrix. The complex elements of the time-frequency distribution matrix are used to calculate cross-band phase correlation. The normalized sine integral function constrains the optimization process only in the effective frequency band of the voice signal. The frequency response balance coefficient is adaptively adjusted according to the energy center of gravity position of the time-frequency distribution matrix. The dynamic center frequency is determined by the spectral flatness detection of the time-frequency distribution matrix. The optimal solution output by the quantum annealing processor is converted into an acoustic phase adjustment value through a decoder. The digital signal processor of the piezoelectric speaker array receives the acoustic phase adjustment value and generates a corresponding delay control signal. The delay control signal drives the piezoelectric transducer unit to adjust the phase of the transmitted waveform.
[0080] S4.3: Generate a directionally focused sound field based on a phase-adjusted piezoelectric speaker array.
[0081] The specific process includes: the phase-adjusted piezoelectric speaker array generates a directional focused sound field through a beamforming algorithm. The beamforming algorithm phase-synthesizes the transmission signals of each array unit based on the principle of acoustic wave interference. Each unit of the piezoelectric speaker array accurately controls the delay of the transmitted waveform according to the acoustic wave phase adjustment amount. The phase difference between the piezoelectric speaker arrays produces constructive interference at a specific position in space. The position of the constructive interference area is determined by the user's real-time position. The three-dimensional coordinates of the sound field focal point are converted into the phase compensation amount of each array unit through coordinate transformation. The phase compensation amount is converted into the corresponding digital delay line parameters through a lookup table. The digital delay line parameters control the relative delay time of the transmission signal of each array unit. The accuracy of the relative delay time reaches the nanosecond level to ensure the acoustic wave interference effect. The acoustic waves emitted by the array unit at the user's position interfere to form a high-energy sound spot, and its focusing process is achieved by the superposition of nanosecond delay signals controlled by the phase compensation amount. The size of the high-energy sound spot is determined by the array aperture and the wavelength of the sound wave. The sound field energy distribution is spatially shaped through the array radiation pattern. The main lobe of the array radiation pattern points to the user position while suppressing sidelobe interference. The resulting directionally focused sound field reaches the maximum sound pressure level at the user position and maintains the spectral integrity of the voice signal.
[0082] S4.4: Based on the directionally focused sound field, collect the temperature change rate of the user's auricle and receive the actual voice signal at the same time.
[0083] The specific process includes: when the directional focused sound field acts on the user's auricle area, the dual-band infrared thermal imager captures the surface temperature distribution of the auricle at a fixed sampling interval. The temperature distribution data is eliminated by Gaussian filtering to eliminate environmental thermal noise, and the auricle temperature field change rate is generated by differential operation of the temperature distribution of two adjacent frames. The spatial gradient of the temperature field change rate reflects the sound energy absorption distribution characteristics. At the same time, the high-sensitivity microphone array receives the actual voice signal at the user's position. The actual voice signal is processed by pre-emphasis filtering and adaptive gain control. The pre-emphasis filtering compensates for high-frequency attenuation, and the adaptive gain control keeps the signal amplitude stable. The actual voice signal is aligned in the time domain with the reference signal of the directional focused sound field. The propagation delay difference of the aligned actual voice signal and the reference signal of the directional focused sound field is analyzed by cross-correlation. The auricle temperature field change rate and the timestamp of the actual voice signal are strictly synchronized.
[0084] S4.5: Calculate the user physiological feedback matrix of the acoustic-thermal coupling effect based on the auricle temperature field change rate and the actual speech signal. The expression is:
[0085]
[0086] Wherein, B(i) represents the user physiological feedback matrix of the acoustic-thermal coupling effect at the monitoring time point i, i is the current monitoring time point, ξ represents the tissue thermal conductivity coefficient (0.3≤ξ≤0.4), Q represents the average surface area of the auricle, Ω represents the spatial area of the auricle, x represents the three-dimensional spatial coordinates of the auricle surface, D(x,i) represents the auricle surface temperature field measured at the three-dimensional spatial coordinates x and the monitoring time point i, F represents the fast Fourier transform operator, s1 represents the pure reference speech signal, ⊙ represents the element-by-element multiplication of the matrix, H represents the frequency domain perception weighting function, s2 represents the actual received speech signal, γ represents the acoustic-thermal conversion efficiency factor, erf represents the Gaussian error function, represents the spatial gradient of the temperature field, Represents the spatial gradient of the acoustic pressure field.
[0087] The specific process includes: the auricle temperature field change rate is processed through spatial integration and time domain differentiation to obtain a standardized temperature response index; the actual speech signal is converted into frequency domain features through fast Fourier transform and multiplied element-by-element with the frequency domain perception weighting function; the tissue thermal conductivity coefficient adjusts the sensitivity of the temperature response index; the average surface area of the auricle is used to normalize the temperature field integration result; the spatial gradient of the temperature field is calculated by the central difference method to calculate the temperature change of adjacent measurement points on the auricle surface; the spatial gradient of the sound pressure field is obtained by near-field acoustic holography reconstruction of the directionally focused sound field; the frequency domain perception weighting function highlights the effective frequency band of 300-3400Hz in the speech signal; the Gaussian error function maps the spatial correlation of the temperature field gradient and the sound pressure field gradient to the [-1,1] interval; the acoustic-thermal conversion efficiency factor balances the contribution weights of temperature response and speech clarity. The final generated user physiological feedback matrix of the acoustic-thermal coupling effect includes quantitative indicators of three dimensions: temperature change rate, weighted speech signal-to-noise ratio, and gradient correlation. The time resolution of these three indicators is aligned with the monitoring time point, and the spatial resolution is consistent with the distribution of measurement points on the auricle surface.
[0088] S5: Based on the auricle temperature change rate and speech clarity index in the user's physiological feedback matrix, a standardized feedback vector is generated, and the vibration-acoustic model is updated through a reinforcement learning strategy.
[0089] S5.1: Normalize the user's physiological feedback matrix to generate a standardized feedback vector.
[0090] The specific process includes extracting the historical mean and standard deviation of the user physiological feedback matrix through sliding window statistical processing, subtracting the historical temperature mean from the temperature change rate component and dividing it by the historical temperature standard deviation to achieve zero mean normalization, subtracting the historical signal-to-noise ratio from the speech signal-to-noise ratio component and dividing it by four times the historical signal-to-noise ratio standard deviation to compress the value to the range of [-0.25, 0.25], and nonlinearly mapping the gradient correlation component to the range of [-1, 1] through the hyperbolic tangent function. The normalized temperature change rate component reflects the degree of abnormality relative to the historical benchmark, the normalized speech signal-to-noise ratio item represents the deviation from the typical communication quality, and the normalized gradient correlation item maintains the original physical meaning but eliminates the dimensional difference. The three normalized components are spliced column by column to form a standardized feedback vector. The dimension of the standardized feedback vector is consistent with the user physiological feedback matrix. The value range of the elements of the standardized feedback vector is uniformly adjusted to the range of [-1, 1] for subsequent processing. The time resolution of the standardized feedback vector is completely aligned with the sampling time of the user physiological feedback matrix.
[0091] S5.2: Based on the standardized feedback vector and the vibration-acoustic model, a Markov decision process state space is constructed and input into the TD3 reinforcement learning policy network to generate vibration-acoustic model optimization instructions.
[0092] The specific process includes tensor splicing of the standardized feedback vector and the parameter vector of the current vibration-acoustic model to form a joint feature. The state space of the Markov decision process is composed of the joint feature and the historical parameter update record. The dimension of the Markov decision process state space is reduced to a fixed length through principal component analysis. The dual critic networks of the TD3 reinforcement learning strategy network respectively evaluate the long-term cumulative rewards of the state-action pairs. The critic network adopts a three-layer fully connected architecture with layer normalization processing. The actor network outputs the update direction of the vibration-acoustic model parameters. The action space is constrained within the hypersphere of the parameter space to prevent mutations. During the training of the TD3 reinforcement learning strategy network, the target network smoothing technology is used to stabilize the learning process. The exploration noise is generated by truncated normal distribution to ensure the rationality of the action. The final output vibration-acoustic model optimization instruction contains parameter update direction and step size information. The format of the optimization instruction strictly matches the parameter structure of the vibration-acoustic model, and the execution order of the optimization instruction is consistent with the timing relationship of the state space.
[0093] Furthermore, the specific architecture of the TD3 reinforcement learning policy network includes an Actor network (outputting deterministic actions) and two independent Critic networks (used to suppress overestimation of Q values), each equipped with a corresponding target network. Stable training is achieved by delaying policy updates (such as updating the Critic twice and then the Actor once) and smoothing the target policy (adding truncated noise).
[0094] S5.3: Verify the physiological safety constraints of the vibro-acoustic model optimization instructions and iteratively update the vibro-acoustic model.
[0095] The specific process includes: the vibration-acoustic model optimization instruction is verified for safety through temperature change rate threshold detection; when the temperature change rate predicted by the optimization instruction exceeds the temperature change rate threshold, the protection mechanism is triggered, and the protection mechanism proportionally reduces the parameter update amount to a safe range; the verified vibration-acoustic model optimization instruction acts on the vibration-acoustic model through the parameter update rule; the parameter update is implemented using the momentum-accelerated gradient descent method; the momentum coefficient is dynamically adjusted according to the consistency of the historical update direction; the weight matrix of the vibration-acoustic model is iteratively updated according to the direction and step size given by the optimization instruction; the updated vibration-acoustic model is immediately put into the sound field generation process for effect verification; the verification results are fed back to the experience replay buffer of the TD3 reinforcement learning strategy network; the entire update process keeps the stability boundary of the vibration-acoustic model unchanged, ensuring that the acoustic output is always within the physiological safety range, and the frequency of iterative updates is synchronized with the sound field refresh rate.
[0096] The temperature change rate threshold is set according to medical safety standards. The specific value is determined by counting the extreme change rates in historical temperature data and combining it with the safe operating range of the equipment.
[0097] The safety range refers to the limit value that does not cause harm or danger in a certain activity or equipment operation. It is usually set based on industry standards, medical safety data or equipment performance parameters, and the specific value is determined through experimental verification or historical statistics.
[0098] The momentum coefficient is a hyperparameter used in the optimization algorithm to control the impact of historical gradients on the current parameter update. Its value range is [0, 1). It is generally set empirically (such as 0.9) or dynamically adjusted to balance the convergence speed and stability.
[0099] This embodiment also provides a voice interaction system based on an intelligent doll, comprising: a signal acquisition module, a modal analysis module, a voice reconstruction module, a sound field focusing module, and a model optimization module. The signal acquisition module is used to collect the vibration displacement signal of the user's neck tissue and the environmental audio signal, and perform preprocessing to generate a vibration feature matrix and an acoustic feature matrix; the modal analysis module is used to perform time-frequency analysis on the vibration feature matrix, combine the density characteristics of the human neck tissue, and generate a modal parameter set of the vocal cord vibration through biomechanical modeling; the voice generation module is used to input the modal parameter set of the vocal cord vibration into the vibration-acoustic The model generates a sound source excitation field, uses the acoustic feature matrix to compensate for the ambient noise in the sound source excitation field, and generates a time-domain pure speech signal through the acoustic wave equation; the sound field focusing module is used to collect the user's real-time position coordinates, calculate and drive the piezoelectric speaker array of the smart doll to adjust the phase according to the spectral characteristics of the time-domain pure speech signal, form a directional focused sound field, record the user's physiological feedback, and generate a user physiological feedback matrix; the model optimization module is used to generate a standardized feedback vector based on the auricle temperature change rate and speech clarity index in the user's physiological feedback matrix, and update the vibration-acoustic model through a reinforcement learning strategy.
[0100] This embodiment also provides a computer device suitable for the voice interaction method based on a smart doll, including: a memory and a processor; the memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions to implement the voice interaction method based on the smart doll proposed in the above embodiment.
[0101] The computer device may be a terminal, comprising a processor, a memory, a communication interface, a display screen and an input device connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device comprises a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The communication interface of the computer device is used to communicate with an external terminal in a wired or wireless manner, and the wireless manner may be achieved through WIFI, an operator network, NFC (near field communication) or other technologies. The display screen of the computer device may be a liquid crystal display or an electronic ink display screen, and the input device of the computer device may be a touch layer covering the display screen, or a button, trackball or touchpad provided on the housing of the computer device, or an external keyboard, touchpad or mouse.
[0102] This embodiment also provides a storage medium having a computer program stored thereon, which, when executed by a processor, implements the voice interaction method based on the smart doll as proposed in the above embodiment; the storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk.
[0103] In summary, the present invention realizes physical modeling of the vocal cord vibration state by: solving the vocal cord displacement modal function vector through nonlinear integration and augmented Lagrangian algorithm, thereby improving the acquisition accuracy and noise resistance of the voice signal at the source; further, by collecting the user's auricle temperature change rate and speech clarity index to construct a user physiological feedback matrix, it can perform adaptive optimization according to the user's physiological response and speech perception quality, further improving the personalization and comfort of voice interaction.
[0104] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present invention, which should all be included in the scope of the claims of the present invention.
Claims
1. A voice interaction method based on an intelligent doll, characterized by: include, Collect the vibration displacement signal of the user's neck tissue and the ambient audio signal, perform preprocessing on them, and generate the vibration feature matrix and the acoustic feature matrix; Perform time-frequency analysis on the vibration characteristic matrix, combine it with the density characteristics of human neck tissue, and generate the modal parameter set of vocal cord vibration through biomechanical modeling; The modal parameter set of vocal cord vibration is input into the vibration-acoustic model to generate the sound source excitation field, and the acoustic characteristic matrix is used to compensate the ambient noise of the sound source excitation field, and the time domain pure speech signal is generated through the acoustic wave equation; The user's real-time location coordinates are collected, and the spectral characteristics of the pure voice signal in the time domain are calculated and driven to adjust the phase of the smart doll's piezoelectric speaker array to form a directional focused sound field and record the user's physiological feedback to generate a user physiological feedback matrix. Based on the auricle temperature change rate and speech clarity index in the user's physiological feedback matrix, a standardized feedback vector is generated, and the vibration-acoustic model is updated through a reinforcement learning strategy.
2. The voice interaction method based on the smart doll according to claim 1, characterized in that: The specific steps of generating the vibration feature matrix and the acoustic feature matrix are as follows: The Doppler effect is used to capture the vibration displacement signal of the user's neck tissue, synchronously collect the ambient audio signal, and eliminate the interference of blood vessels and muscle fluctuations to generate pure vocal cord vibration waves; Perform multi-scale dynamic decomposition on pure vocal cord vibration waves and generate time-varying feature weights based on the conduction characteristics of biological tissues; The time-varying feature weights are compensated for acoustic impedance matching with the ambient audio signal to generate a vibration feature matrix and an acoustic feature matrix.
3. The voice interaction method based on the smart doll according to claim 2, characterized in that: The specific steps of generating the modal parameter set of vocal cord vibration are as follows: Based on the vibration characteristic matrix, a time-varying window is dynamically selected, and the local wave energy density function is generated through nonlinear integration and mapped into a wave energy topology matrix that characterizes the tissue vibration curvature. The density characteristics of human neck tissue are converted into biomechanical constraints of mass conservation. Combined with the wave energy topology matrix, the displacement mode function vector is iteratively solved through the augmented Lagrangian algorithm. The displacement modal function vector is used to guide the neural network to compensate for the residual, and the modal parameter set of vocal fold vibration is analytically generated after fusion.
4. The voice interaction method based on the smart doll according to claim 3, characterized in that: The modal parameter set of vocal cord vibration is input into the vibration-acoustic model to generate a sound source excitation field, and the acoustic characteristic matrix is used to compensate the sound source excitation field for environmental noise, and a time domain pure speech signal is generated through the acoustic wave equation. The specific steps are as follows: The modal parameter set of vocal cord vibration is input into the vibro-acoustic model to generate the sound source excitation field; The sound source excitation field and the acoustic characteristic matrix are input into the vibration-acoustic model. The acoustic wave equation is solved and the ambient noise field is compensated synchronously through the deep coupling operator to generate the sound pressure field of the sound channel. The sound pressure field of the vocal tract is input into the radiation reconstruction process of the vibration-acoustic model, and the time-domain pure speech signal is generated by integrating the acoustic impedance compensation function on the lip surface.
5. The voice interaction method based on the smart doll according to claim 4, characterized in that: The steps of collecting the user's real-time location coordinates, calculating and driving the piezoelectric speaker array of the smart doll to adjust the phase according to the spectrum characteristics of the pure voice signal in the time domain are as follows: Perform time-frequency transformation on the time-domain pure speech signal to generate a time-frequency distribution matrix; The user's real-time location coordinates are collected, combined with the time-frequency distribution matrix, and the sound wave phase adjustment amount is calculated through quantum optimization processing. The piezoelectric speaker array is driven to adjust the phase according to the sound wave phase adjustment amount.
6. The voice interaction method based on the smart doll according to claim 5, characterized in that: The specific steps of forming a directional focused sound field and recording user physiological feedback to generate a user physiological feedback matrix are as follows: Generate a directional focused sound field based on a phase-adjusted piezoelectric speaker array; Based on the directional focused sound field, the user's auricle temperature field change rate is collected while receiving the actual voice signal; The user physiological feedback matrix of the acoustic-thermal coupling effect is calculated based on the auricle temperature field change rate and the actual speech signal.
7. The voice interaction method based on the smart doll according to claim 6, characterized in that: The method generates a standardized feedback vector based on the auricle temperature change rate and the speech clarity index in the user's physiological feedback matrix, and updates the vibration-acoustic model through a reinforcement learning strategy. The specific steps are as follows: Normalize the user's physiological feedback matrix to generate a standardized feedback vector; Based on the standardized feedback vector and the vibration-acoustic model, a Markov decision process state space is constructed and input into the TD3 reinforcement learning strategy network to generate vibration-acoustic model optimization instructions; Physiological safety constraints are verified for the vibration-acoustic model optimization instructions, and the vibration-acoustic model is iteratively updated.
8. A voice interaction system based on an intelligent doll, based on the voice interaction method based on an intelligent doll according to any one of claims 1 to 7, characterized in that: Including signal acquisition module, modal analysis module, speech reconstruction module, sound field focusing module and model optimization module, The signal acquisition module is used to collect the vibration displacement signal of the user's neck tissue and the ambient audio signal, and perform preprocessing to generate the vibration feature matrix and the acoustic feature matrix; The modal analysis module is used to perform time-frequency analysis on the vibration characteristic matrix, combine the density characteristics of human neck tissue, and generate the modal parameter set of vocal cord vibration through biomechanical modeling; The speech generation module is used to input the modal parameter set of vocal cord vibration into the vibration-acoustic model to generate the sound source excitation field, and use the acoustic characteristic matrix to compensate the sound source excitation field for environmental noise, and generate a time-domain pure speech signal through the acoustic wave equation; The sound field focusing module is used to collect the user's real-time position coordinates, calculate and drive the smart doll's piezoelectric speaker array to adjust the phase based on the spectral characteristics of the pure voice signal in the time domain, form a directional focused sound field, and record the user's physiological feedback to generate a user physiological feedback matrix; The model optimization module is used to generate a standardized feedback vector based on the auricle temperature change rate and speech clarity index in the user's physiological feedback matrix, and update the vibration-acoustic model through a reinforcement learning strategy.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the smart doll-based voice interaction method according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the voice interaction method based on the smart doll according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Voice recognition classification method based on vocal cord characteristic parameters
CN112562650A
Bone transmission microphone speech enhancement method and device, equipment and storage medium
CN115862656A
Method for voice-excluding muscle motion control system
CN117222967A
Interactive foreign language speech training auxiliary system for hearing-impaired children
CN119889135A
Sound coordination method and system based on tinnitus condition
CN119889582A
Cited By
Health record personalized recommendation method and system based on artificial intelligence
CN121171636A