Intelligent doll-based voice interaction method and system

By collecting and analyzing vibration displacement signals of the user's neck tissue and environmental audio signals, and combining biomechanical modeling and acoustic feature matrix for noise compensation, high-quality speech signals are generated. This solves the problems of inaccurate speech generation modeling and insufficient adaptability, and achieves higher precision and more comfortable voice interaction.

CN120472901BActive Publication Date: 2025-11-18XINGFUQUAN (BEIJING) INTERNATIONAL ARTIFICIAL INTELLIGENCE TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510800152.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-16
Publication Date
2025-11-18
Estimated Expiration
2045-06-16

AI Technical Summary

Technical Problem

Existing technologies suffer from inaccurate speech generation modeling and insufficient adaptability. They fail to fully consider the coupling relationship between the biomechanical characteristics of neck tissue and vocal cord vibration modes, and lack a dynamic compensation mechanism for environmental noise, making it difficult to adapt to complex and ever-changing real-world application scenarios.

Method used

The system collects vibration displacement signals of the user's neck tissue and ambient audio signals. It generates a set of modal parameters for vocal cord vibration through Doppler effect and biomechanical modeling, uses acoustic feature matrix for environmental noise compensation, combines piezoelectric loudspeaker array to adjust phase to form a directional focused sound field, and optimizes the model through reinforcement learning strategy.

Benefits of technology

It improves the accuracy and noise resistance of voice signal acquisition, enhances the personalization and comfort of voice interaction, and achieves adaptability to complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120472901B_ABST
    Figure CN120472901B_ABST
Patent Text Reader

Abstract

The application discloses a voice interaction method and system based on an intelligent doll, relates to the technical field of intelligent acoustic interaction, and comprises the following steps: performing time-frequency analysis and processing on a vibration feature matrix, combining with the density characteristics of human neck tissues, and generating a modal parameter set of vocal cord vibration through biomechanical modeling; inputting the modal parameter set of vocal cord vibration into a vibration-acoustic model to generate a sound source excitation field, compensating the sound source excitation field for environmental noise by using an acoustic feature matrix, and generating a time-domain pure speech signal through a sound wave equation; collecting real-time position coordinates of a user, calculating and driving a piezoelectric loudspeaker array of the intelligent doll to adjust a phase according to the spectral characteristics of the time-domain pure speech signal, forming a directional focused sound field, recording physiological feedback of the user, and generating a user physiological feedback matrix. The application solves a vocal cord displacement modal function vector by using a nonlinear integral and an augmented Lagrange algorithm, so that the collection accuracy and the noise resistance of a speech signal are improved at the source.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent acoustic interaction technology, and in particular to a voice interaction method and system based on intelligent dolls. Background Technology

[0002] With the rapid development of artificial intelligence and human-computer interaction technologies, voice, as one of the most natural and direct modes of communication, has gradually become the core interaction method for smart devices. Especially in areas such as children's education, emotional companionship, and elderly care, voice-interactive smart doll systems have attracted widespread attention due to their anthropomorphic features and high level of immersion. Traditional speech recognition and synthesis technologies mainly rely on microphones to collect environmental audio signals and use methods such as deep neural networks for speech enhancement and semantic understanding.

[0003] Although the signal-to-noise ratio of voice acquisition has been improved to some extent, there are still several key technical bottlenecks: First, existing technologies generally do not fully consider the coupling relationship between the biomechanical characteristics of neck tissue and vocal cord vibration modes, resulting in insufficient accuracy in modeling the voice generation mechanism; Second, there is a lack of dynamic compensation mechanism for environmental noise during the voice generation process, making it difficult to adapt to complex and ever-changing actual use scenarios. Summary of the Invention

[0004] In view of the aforementioned existing problems, the present invention is proposed.

[0005] Therefore, this invention provides a voice interaction method based on intelligent dolls to solve the problems of inaccurate voice generation modeling and insufficient adaptability of voice generation in the prior art.

[0006] To solve the above-mentioned technical problems, the present invention provides the following technical solution:

[0007] In a first aspect, the present invention provides a voice interaction method based on a smart doll, comprising: acquiring vibration displacement signals of the user's neck tissue and environmental audio signals, and preprocessing them to generate a vibration feature matrix and an acoustic feature matrix; performing time-frequency analysis on the vibration feature matrix, combining it with the density characteristics of human neck tissue, and generating a set of modal parameters for vocal cord vibration through biomechanical modeling; inputting the set of modal parameters for vocal cord vibration into a vibration-acoustic model to generate a sound source excitation field, and using the acoustic feature matrix to compensate for environmental noise in the sound source excitation field, and generating a time-domain clean speech signal through the sound wave equation; acquiring the user's real-time position coordinates, calculating and driving the piezoelectric speaker array of the smart doll to adjust its phase according to the spectral characteristics of the time-domain clean speech signal, forming a directional focusing sound field and recording the user's physiological feedback to generate a user physiological feedback matrix; generating a standardized feedback vector based on the auricle temperature change rate and speech intelligibility index in the user physiological feedback matrix, and updating the vibration-acoustic model through a reinforcement learning strategy.

[0008] As a preferred embodiment of the voice interaction method based on intelligent dolls described in this invention, the specific steps for generating the vibration feature matrix and acoustic feature matrix are as follows:

[0009] By capturing the vibration displacement signal of the user's neck tissue through the Doppler effect, simultaneously acquiring the ambient audio signal, and eliminating the interference of blood vessel and muscle fluctuations, pure vocal cord vibration waves are generated.

[0010] The pure vocal cord vibration wave is dynamically decomposed at multiple scales, and time-varying feature weights are generated through the conduction properties of biological tissues.

[0011] The time-varying feature weights are matched and compensated with the ambient audio signal to generate a vibration feature matrix and an acoustic feature matrix.

[0012] As a preferred embodiment of the voice interaction method based on intelligent dolls described in this invention, the specific steps for generating the modal parameter set of vocal cord vibration are as follows:

[0013] Based on the dynamic selection of time-varying windows using the vibration feature matrix, a local wave energy density function is generated through nonlinear integration and mapped to a wave energy topology matrix characterizing the tissue vibration curvature.

[0014] The density characteristics of human neck tissue are transformed into biomechanical constraints of mass conservation. Combined with the wave energy topology matrix, the displacement mode function vector is iteratively solved using the augmented Lagrange algorithm.

[0015] The displacement modal function vector is used to guide the neural network to compensate for the residuals, and the resulting data is then fused to generate a set of modal parameters for vocal cord vibration.

[0016] As a preferred embodiment of the voice interaction method based on intelligent dolls described in this invention, the steps include: inputting the modal parameter set of vocal cord vibration into a vibration-acoustic model to generate a sound source excitation field, using an acoustic feature matrix to compensate for environmental noise in the sound source excitation field, and generating a time-domain clean speech signal through the sound wave equation. The specific steps are as follows:

[0017] The modal parameter set of vocal cord vibration is input into the vibration-acoustic model to generate the sound source excitation field;

[0018] The sound source excitation field and the acoustic feature matrix are input into the vibration-acoustic model. The sound wave equation is solved and the environmental noise field is compensated simultaneously by the deep coupling operator to generate the sound pressure field of the sound channel.

[0019] The sound pressure field of the vocal tract is input into the radiation reconstruction process of the vibration-acoustic model, and a time-domain clean speech signal is generated by the surface integral acoustic impedance compensation function of the lip curvature.

[0020] As a preferred embodiment of the voice interaction method based on a smart doll described in this invention, the specific steps of acquiring the user's real-time location coordinates and calculating and driving the piezoelectric speaker array of the smart doll to adjust its phase based on the spectral characteristics of the time-domain pure voice signal are as follows:

[0021] Perform time-frequency transformation on the clean speech signal in the time domain to generate a time-frequency distribution matrix;

[0022] The system collects the user's real-time location coordinates, combines them with the time-frequency distribution matrix, calculates the acoustic wave phase adjustment through quantum optimization, and drives the piezoelectric speaker array to adjust the phase based on the acoustic wave phase adjustment.

[0023] As a preferred embodiment of the voice interaction method based on intelligent dolls described in this invention, the specific steps for forming a directional focusing sound field, recording user physiological feedback, and generating a user physiological feedback matrix are as follows:

[0024] A directional focused sound field is generated based on a phase-adjusted piezoelectric loudspeaker array.

[0025] Based on the directional focusing sound field, the rate of change of the user's auricle temperature field is collected, while the actual voice signal is received.

[0026] The user physiological feedback matrix of the acoustic-thermal coupling effect is calculated based on the rate of change of the auricular temperature field and the actual speech signal.

[0027] As a preferred embodiment of the voice interaction method based on intelligent dolls described in this invention, the steps of generating a standardized feedback vector based on the auricle temperature change rate and speech intelligibility index in the user's physiological feedback matrix, and updating the vibration-acoustic model through a reinforcement learning strategy, are as follows:

[0028] The user's physiological feedback matrix is ​​normalized to generate a standardized feedback vector;

[0029] Based on the standardized feedback vector and the vibration-acoustic model, a Markov decision process state space is constructed and input into the TD3 reinforcement learning policy network to generate vibration-acoustic model optimization instructions.

[0030] Physiological safety constraints were verified on the optimization instructions for the vibration-acoustic model, and the vibration-acoustic model was iteratively updated.

[0031] Secondly, this invention provides a voice interaction system based on an intelligent doll, comprising a signal acquisition module, a modal analysis module, a speech reconstruction module, a sound field focusing module, and a model optimization module. The signal acquisition module is used to acquire vibration displacement signals of the user's neck tissue and environmental audio signals, and perform preprocessing to generate a vibration feature matrix and an acoustic feature matrix. The modal analysis module is used to perform time-frequency analysis processing on the vibration feature matrix, combine it with the density characteristics of human neck tissue, and generate a set of modal parameters for vocal cord vibration through biomechanical modeling. The speech generation module is used to input the set of modal parameters for vocal cord vibration into the vibration-... The acoustic model generates a sound source excitation field and uses an acoustic feature matrix to compensate for environmental noise in the sound source excitation field. It also generates a time-domain clean speech signal through the sound wave equation. The sound field focusing module is used to collect the user's real-time position coordinates, calculate and drive the piezoelectric speaker array of the smart doll to adjust the phase based on the spectral characteristics of the time-domain clean speech signal, form a directional focused sound field, record the user's physiological feedback, and generate a user physiological feedback matrix. The model optimization module is used to generate a standardized feedback vector based on the auricular temperature change rate and speech intelligibility index in the user physiological feedback matrix, and update the vibration-acoustic model through a reinforcement learning strategy.

[0032] Thirdly, the present invention provides a computer device including a memory and a processor, wherein the memory stores a computer program, wherein: when the computer program is executed by the processor, it implements any step of the voice interaction method based on an intelligent doll as described in the first aspect of the present invention.

[0033] Fourthly, the present invention provides a computer-readable storage medium having a computer program stored thereon, wherein: when the computer program is executed by a processor, it implements any step of the voice interaction method based on an intelligent doll as described in the first aspect of the present invention.

[0034] The beneficial effects of this invention are as follows: by solving the vocal tract displacement mode function vector through nonlinear integral and augmented Lagrange algorithm, a physical model of the vocal tract vibration state is realized, thereby improving the acquisition accuracy and noise resistance of speech signals at the source; furthermore, by collecting the user's auricle temperature change rate and speech intelligibility index to construct the user's physiological feedback matrix, adaptive optimization can be performed according to the user's physiological response and speech perception quality, further enhancing the personalization and comfort of voice interaction. Attached Figure Description

[0035] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0036] Figure 1 This is a flowchart of a voice interaction method based on intelligent dolls.

[0037] Figure 2 A flowchart for generating the vocal tract vibration mode parameter set.

[0038] Figure 3 A flowchart for generating a clean speech signal in the time domain.

[0039] Figure 4 The flowchart for closed-loop optimization of a vibration-acoustic model driven by physiological feedback. Detailed Implementation

[0040] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings.

[0041] Many specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the spirit of the invention. Therefore, the invention is not limited to the specific embodiments disclosed below.

[0042] Secondly, the term "one embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The phrase "in one embodiment" appearing in different places in this specification does not necessarily refer to the same embodiment, nor is it a single or selective embodiment that is mutually exclusive with other embodiments.

[0043] Reference Figures 1-4 This is one embodiment of the present invention, which provides a voice interaction method based on a smart doll, including the following steps:

[0044] S1: Collect vibration displacement signals of the user's neck tissue and ambient audio signals, and preprocess them to generate vibration feature matrices and acoustic feature matrices.

[0045] S1.1: Captures vibration displacement signals of the user's neck tissue through the Doppler effect, synchronously acquires ambient audio signals, and eliminates interference from blood vessel and muscle fluctuations to generate pure vocal cord vibration waves.

[0046] The specific process includes: when the Doppler effect captures the vibration displacement signal of the user's neck tissue, ultrasound waves of a specific frequency are emitted and reflected waves are received; the vibration velocity and displacement amplitude of the tissue surface are analyzed by analyzing the frequency offset of the reflected waves; the ambient audio signal is synchronously acquired through a high-sensitivity microphone array; the adaptive filtering adopts the normalized least mean square algorithm, using the ambient audio signal as a reference signal to eliminate low-frequency mechanical vibration interference caused by vascular pulsation and muscle contraction from the mixed signal; the bandpass filter retains the 20Hz to 1000Hz frequency band signal that matches the vocal cord vibration frequency; and finally, a pure vocal cord vibration wave is output.

[0047] S1.2: Perform multi-scale dynamic decomposition on pure vocal cord vibration waves and generate time-varying feature weights through the conduction characteristics of biological tissues.

[0048] The specific process involves wavelet packet decomposition of the pure vocal cord vibration wave into sub-signal components of different frequency bands. Neck tissue density characteristics and acoustic impedance parameters are used to analyze the propagation characteristics of sound waves in biological tissues and determine the attenuation level of each frequency band. Based on the attenuation level, each frequency band sub-signal is dynamically weighted, with high-frequency components receiving lower weights due to tissue absorption and low-frequency components receiving higher weights. This ultimately generates time-varying feature weights reflecting the conduction characteristics of biological tissues. These weights directly affect the vibration feature matrix, and the optimized multi-scale vibration features retain key information about vocal cord vibration.

[0049] Biological tissue conduction characteristics refer to the physical properties exhibited by sound waves when they propagate in human neck tissues, such as frequency attenuation patterns, acoustic impedance distribution, and energy absorption efficiency. These properties are specifically determined by tissue density, elastic modulus, and viscosity coefficient.

[0050] S1.3: Perform acoustic impedance matching compensation between the time-varying feature weights and the environmental audio signal to generate the vibration feature matrix and the acoustic feature matrix.

[0051] The specific process includes: time-varying feature weights and ambient audio are processed through acoustic impedance matching compensation. The time-varying feature weights reflect the propagation characteristics of sound waves in different frequency bands in the neck tissue, and the ambient audio contains spatial acoustic information. The two are matched using sound field reconstruction technology. The time-varying feature weights guide the redistribution of energy in each frequency band of the ambient audio and compensate for frequency response distortion caused by tissue conduction. The ambient audio is subjected to Mel frequency cepstral analysis to extract spectral features. The time-varying feature weights and spectral features are multiplied at frequency points to achieve acoustic impedance matching. The vibration feature matrix is ​​generated by weighted fusion of the multi-scale components of the pure vocal cord vibration wave through time-varying feature weights. The acoustic feature matrix is ​​generated from the compensated ambient audio features. The vibration feature matrix and the acoustic feature matrix are kept in time synchronization.

[0052] S2: Perform time-frequency analysis on the vibration feature matrix, combine it with the density characteristics of human neck tissue, and generate a set of modal parameters for vocal cord vibration through biomechanical modeling.

[0053] S2.1: Based on the vibration feature matrix, a time-varying window is dynamically selected, and a local wave energy density function is generated through nonlinear integration, which is then mapped to a wave energy topology matrix that characterizes the vibration curvature of the tissue.

[0054] The specific process includes: determining the active signal period through short-time energy analysis of the vibration feature matrix; adaptively selecting a Hanning window as a time-varying window within the active period; the width of the time-varying window being inversely proportional to the instantaneous frequency of the vibration feature matrix; using a narrower window in the high-frequency region to improve time resolution and a wider window in the low-frequency region to ensure frequency resolution; extracting the instantaneous amplitude of the signal components after window selection using Hilbert transform; squaring the instantaneous amplitude and integrating along the time axis to generate a local wave energy density function; smoothing the local wave energy density function using Gaussian smoothing to eliminate abrupt interference; processing the smoothed wave energy density distribution using a curvature operator; and using a second-order difference approximation to analyze the curvature value at each time frequency point. Positive curvature corresponds to wave energy convergence regions, and negative curvature corresponds to wave energy divergence regions, ultimately generating a wave energy topology matrix reflecting the characteristics of tissue vibration energy distribution.

[0055] S2.2: The density characteristics of human neck tissue are transformed into biomechanical constraints of mass conservation. Combined with the wave energy topology matrix, the displacement mode function vector is iteratively solved using the augmented Lagrange algorithm. The expression is:

[0056]

[0057] Among them, u (k+1) This represents the displacement mode function vector obtained in the (k+1)th iteration. This indicates finding the minimum of the objective function over a displacement function vector u, where u represents the displacement function vector and k represents the index of the current iteration step. Let K denote the transpose of the displacement function vector, K denote the stiffness matrix, T denote the wave energy topology matrix, and μ denote the wave energy topology matrix. (k) Let A denote the penalty coefficient for the k-th iteration, and let A denote the mass conservation constraint matrix. Let g denote the transpose of the mass conservation constraint matrix, g denote the external force vector, and λ denote the mass conservation constraint matrix. (k) Let represent the Lagrange multiplier vector for the k-th iteration, and b represent the constant term vector of the mass conservation constraint.

[0058] The specific process includes: converting the density characteristics of human neck tissue into a mass distribution matrix through finite element discretization; constructing mass conservation constraints by combining the mass distribution matrix with the continuous medium mechanics equations; expressing the mass conservation constraints as matrix equations; extracting dominant vibration modes from the wave energy topology matrix through eigenvalue decomposition; and using the dominant vibration modes and mass conservation constraints together to form the objective function of the optimization problem. The augmented Lagrange algorithm incorporates the mass conservation constraints as a penalty term into the objective function, updating the Lagrange multiplier vector and penalty coefficient in each iteration. The linear equation system is solved using the conjugate gradient method to obtain the displacement mode function vector. The coefficient matrix of the linear equation system consists of a stiffness matrix, a wave energy topology matrix, and a penalty term. The right-hand side includes an external force vector and a constraint compensation term. The iteration termination condition is set as the Euclidean distance between two adjacent displacement mode function vectors being less than the convergence tolerance factor. The final output displacement mode function vector satisfies the mass conservation constraints and reflects the wave energy distribution characteristics.

[0059] The convergence tolerance factor is set according to the physical magnitude and computational accuracy requirements of the displacement mode function vector. The specific value is 1% of the initial residual norm and is determined through mesh independence verification.

[0060] S2.3: The displacement mode function vector guides the neural network to compensate for the residuals, and after adaptive fusion, the modal parameter set of vocal cord vibration is generated analytically.

[0061] The specific process includes: the displacement mode function vector is input as the input feature into a pre-trained three-layer feedforward neural network; the three-layer feedforward neural network outputs residual compensation coefficients through a sigmoid activation function; the residual compensation coefficients are multiplied element-wise with the displacement mode function vector to achieve residual compensation; the compensated displacement mode function vector is used to extract principal components through singular value decomposition; the principal components are adaptively weighted and fused with the anatomical features of the vocal tract; the weighting coefficients are dynamically determined by the energy proportion of the displacement mode function vector; the fused feature vector is converted into vocal tract vibration mode parameters through a nonlinear mapping layer; the vocal tract vibration mode parameters include key features such as fundamental frequency, vibration amplitude, and phase difference; and the final generated set of vocal tract vibration mode parameters is consistent with biomechanical characteristics.

[0062] Furthermore, the specific training process of the pre-trained three-layer feedforward neural network is as follows: First, the network weights and biases are initialized to ensure that the number of nodes in the input layer, hidden layer, and output layer matches the task requirements; then, the displacement mode function vector is input into the network, and the predicted values ​​of the residual compensation coefficients are obtained by weighted summation of the hidden layer and output layer and analysis layer by layer using the Sigmoid activation function; subsequently, the mean squared error (MSE) loss function is used to compare the predicted values ​​with the true labels, analyze the error, and start backpropagation: starting from the output layer, the gradients of the weights and biases are obtained layer by layer using the chain rule; finally, the weights and biases are updated by gradient descent, and the forward propagation, error analysis, and backpropagation process are repeated iteratively until the loss converges or the preset number of iterations is reached. During this process, early stopping or regularization may be used to prevent overfitting, and the generalization performance is evaluated through a validation set.

[0063] S3: Input the modal parameter set of vocal cord vibration into the vibration-acoustic model to generate the sound source excitation field, and use the acoustic feature matrix to compensate for environmental noise in the sound source excitation field, and generate a time-domain clean speech signal through the sound wave equation.

[0064] S3.1: Input the modal parameter set of vocal cord vibration into the vibration-acoustic model to generate the sound source excitation field.

[0065] The specific process includes: the modal parameter set of vocal cord vibration is converted into sound source boundary conditions through parameter mapping relationship; the sound source boundary conditions include vibration amplitude, vibration frequency and phase information; the vibration-acoustic model establishes the sound wave propagation law based on the linearized Euler equation; the sound source boundary conditions are applied as the excitation source at the glottis position; the sound pressure distribution is obtained by solving the sound wave equation through the finite difference method; the sound pressure distribution forms the sound source excitation field on the discrete grid nodes in three-dimensional space; the temporal variation of the sound source excitation field reflects the vocal cord vibration characteristics; the spatial distribution conforms to the anatomical structure characteristics of the vocal tract; and the final output sound source excitation field contains complete temporal control information.

[0066] Furthermore, the training process of the vibration-acoustic model first collects a large amount of vocal cord vibration data and corresponding sound field measurement data as training samples. Vocal cord vibration data is acquired simultaneously using a high-speed camera and a laser vibrometer, while sound field measurement data is recorded by a piezoelectric microphone array in an anechoic chamber. After preprocessing, the training samples are divided into a vocal cord vibration feature set and a sound field feature set. The vocal cord vibration feature set includes parameters such as fundamental frequency, vibration amplitude, and phase difference, while the sound field feature set includes sound pressure level distribution and spectral characteristics. The vibration-acoustic model employs a deep neural network architecture. The input layer receives vocal cord vibration features, the hidden layers contain three-dimensional convolutional layers and long short-term memory layers to capture spatial and temporal features, and the output layer predicts the sound field distribution. During training, mean squared error is used as the loss function, the network weights are optimized using the backpropagation algorithm, the learning rate is dynamically adjusted using the Adam optimizer, and the number of training iterations is determined based on the convergence of the validation set error. The final vibration-acoustic model can accurately predict the sound field distribution under given vocal cord vibration parameters.

[0067] Preprocessing includes data denoising (such as wavelet denoising and Kalman filtering), signal segmentation and alignment (frame-by-frame windowing), feature extraction (fundamental frequency and spectral energy), data normalization (amplitude / time normalization), and quality control (residual analysis).

[0068] S3.2: Input the sound source excitation field and the acoustic feature matrix into the vibration-acoustic model, and simultaneously execute the sound wave equation solution and environmental noise field compensation through the deep coupling operator to generate the sound pressure field of the sound channel.

[0069] The specific process includes: the sound source excitation field and the acoustic feature matrix are spliced ​​together by tensors to form a joint input feature; the deep coupling operator uses the alternating direction multiplier method to simultaneously process the solution of the acoustic wave equation and the environmental noise field compensation process; the acoustic wave equation is established based on the linearized Euler equation and numerically solved using the pseudospectral method; the environmental noise field compensation constructs a noise covariance matrix through the acoustic feature matrix and applies it to the source term of the acoustic wave equation; the sound pressure field and noise compensation coefficients are alternately updated during the internal iteration of the deep coupling operator; each iteration first fixes the noise compensation coefficients to solve the acoustic wave equation to obtain the transient sound pressure field distribution, and then updates the noise compensation coefficients based on the residual between the transient sound pressure field distribution and the acoustic feature matrix; the iteration terminates when the relative error between two adjacent sound pressure field updates is less than the sound field convergence criterion; the final output channel sound pressure field simultaneously satisfies the physical laws of sound wave propagation and the requirements for environmental noise suppression; the spatial resolution of the channel sound pressure field is consistent with the input sound source excitation field, and the time sampling rate is synchronized with the acoustic feature matrix.

[0070] The sound field convergence criterion is set according to the numerical stability requirements of the sound pressure field solution. The specific value is 0.1% of the initial sound pressure field energy and is determined through acoustic simulation error analysis.

[0071] S3.3: Input the acoustic pressure field of the vocal tract into the radiation reconstruction processing of the vibration-acoustic model, and generate a clean speech signal in the time domain by using the surface integral acoustic impedance compensation function of the lip curvature.

[0072] The specific process includes: discretizing the vocal tract sound pressure field through a lip surface grid; aligning the node coordinates of the lip surface grid with the spatial sampling points of the vocal tract sound pressure field; establishing an acoustic impedance compensation function based on the acoustic characteristics of the lip tissue to reflect the change in radiation impedance of the sound wave from the vocal tract to the free field; representing the acoustic impedance compensation function as a complex transfer function in the frequency domain, with the real part corresponding to the acoustic impedance component and the imaginary part corresponding to the acoustic impedance component; multiplying the frequency domain distribution of the vocal tract sound pressure field point by point with the acoustic impedance compensation function to achieve frequency domain compensation; converting the compensated frequency domain sound pressure into a time domain signal through an inverse Fourier transform; maintaining the sampling rate of the time domain signal consistent with the original vocal cord vibration signal; smoothing the generated time-domain pure speech signal using a Hanning window to eliminate frequency domain leakage effects; and finally outputting a time-domain pure speech signal that retains the essential characteristics of vocal cord vibration and eliminates distortion caused by radiation impedance.

[0073] S4: Collect the user's real-time location coordinates, calculate and drive the piezoelectric speaker array of the smart doll to adjust the phase based on the spectral characteristics of the time-domain pure speech signal, form a directional focused sound field, record the user's physiological feedback, and generate the user's physiological feedback matrix.

[0074] S4.1: Perform time-frequency transformation on the clean speech signal in the time domain to generate a time-frequency distribution matrix.

[0075] The specific process includes: the doll's visual positioning device captures the coordinates of the user's facial feature points using a binocular stereo vision algorithm; the user's real-time position is converted into relative position coordinates centered on the doll through coordinate transformation; the time-domain clean speech signal is processed by windowed short-time Fourier transform; the window length of the Hanning window is dynamically adjusted according to the fundamental frequency of the speech signal; the sliding step size of the window function is set to one-quarter of the window length to ensure time-frequency continuity; the short-time Fourier transform converts the time-domain clean speech signal into a complex form time-frequency distribution matrix; the rows of the time-frequency distribution matrix correspond to time frames, the columns correspond to frequency points, the amplitude of the matrix elements represents the signal energy, and the phase represents the signal phase information; the time resolution of the time-frequency distribution matrix is ​​synchronized with the refresh rate of the visual positioning device, and the frequency resolution is determined by the window function length and the sampling rate; the generated time-frequency distribution matrix is ​​timestamped with the user's real-time position data.

[0076] S4.2: Collect the user's real-time location coordinates, combine them with the time-frequency distribution matrix, calculate the acoustic wave phase adjustment amount through quantum optimization processing, and drive the piezoelectric speaker array to adjust the phase according to the acoustic wave phase adjustment amount. The expression is:

[0077]

[0078] Where, ΔΦm δ represents the acoustic wave phase adjustment of the m-th piezoelectric speaker array, α represents the acoustic wave wavelength, ||r|| represents the Euclidean distance from the user to the geometric center of the piezoelectric speaker array, m represents the index number of the piezoelectric speaker array, and δ m Let f(t) represent the calibration distance of the m-th piezoelectric loudspeaker array reference point, β represent the frequency response balance coefficient (0.2≤β≤1.0), f1 represent the lower limit frequency of the effective frequency band of the speech signal, f2 represent the upper limit frequency of the effective frequency band of the speech signal, f represent the frequency component of the effective frequency band of the speech signal, W(f) represent the perceptual weighting value of the frequency component f of the effective frequency band of the speech signal, sinc represents the normalized sine integral function, f3 represent the dynamic center frequency of the current processing frame, Δf represents the frequency resolution, t represents the current processing time frame number, and θ(t,f) represent the instantaneous phase angle of the time-frequency matrix at time frame t and the frequency component f of the effective frequency band of the speech signal.

[0079] The specific process includes: the user's real-time position is mapped to a spatial coordinate system with the center of the piezoelectric speaker array as the origin through polar coordinate transformation; the time-frequency distribution matrix is ​​optimized in parallel by a quantum annealing processor; the quantum annealing processor models the acoustic phase adjustment problem as a quadratic unconstrained binary optimization problem; the problem Hamiltonian includes a geometric path difference term and a spectral coherence term; the geometric path difference term is determined by the three-dimensional Euclidean distance difference between the user's real-time position and each piezoelectric speaker array unit; the spectral coherence term is constructed using the phase gradient information of the time-frequency distribution matrix; the complex elements of the time-frequency distribution matrix are used to calculate the cross-frequency band phase correlation; the normalized sine integral function constraint optimization process only applies to the effective frequency band of the speech signal; the frequency response balance coefficient is adaptively adjusted according to the energy centroid position of the time-frequency distribution matrix; the dynamic center frequency is determined by the spectral flatness detection of the time-frequency distribution matrix; the optimal solution output by the quantum annealing processor is converted into an acoustic phase adjustment quantity by a decoder; the digital signal processor of the piezoelectric speaker array receives the acoustic phase adjustment quantity and generates a corresponding time delay control signal; the time delay control signal drives the piezoelectric transducer unit to adjust the phase of the transmitted waveform.

[0080] S4.3: Generate a directional focused sound field based on the phase-adjusted piezoelectric loudspeaker array.

[0081] The specific process includes: the phase-adjusted piezoelectric speaker array generates a directional focused sound field through a beamforming algorithm; the beamforming algorithm synthesizes the emitted signals of each array unit based on the principle of acoustic wave interference; each unit of the piezoelectric speaker array precisely controls the time delay of the emitted waveform according to the acoustic wave phase adjustment amount; the phase difference between the piezoelectric speaker arrays generates constructive interference at a specific location in space; the location of the constructive interference region is determined by the user's real-time position; the three-dimensional coordinates of the sound field focus point are converted into the phase compensation amount of each array unit through coordinate transformation; the phase compensation amount is converted into the corresponding digital delay line parameter through a lookup table; the digital delay line parameter controls the relative delay time of the emitted signals of each array unit; the accuracy of the relative delay time reaches the nanosecond level to ensure the acoustic wave interference effect; the acoustic wave interference emitted by the array units at the user's position forms a high-energy sound spot; its focusing process is achieved by superimposing the nanosecond-level delay signals controlled by the phase compensation amount. The size of the high-energy sound spot is determined by the array aperture and the wavelength of the sound wave. The sound field energy distribution is spatially shaped by the array radiation mode. The main lobe of the array radiation mode points to the user's position while suppressing side lobe interference. The final generated directional focused sound field reaches the maximum sound pressure level at the user's position and maintains the spectral integrity of the speech signal.

[0082] S4.4: Based on the directional focusing sound field, it collects the rate of change of the user's auricle temperature field and simultaneously receives the actual voice signal.

[0083] The specific process includes: when the directional focused sound field acts on the user's auricle area, a dual-band infrared thermal imager captures the temperature distribution of the auricle surface at a fixed sampling interval; the temperature distribution data is filtered by Gaussian to eliminate environmental thermal noise; the auricle temperature field change rate is generated by differential calculation of the temperature distribution of two adjacent frames; the spatial gradient of the temperature field change rate reflects the sound energy absorption distribution characteristics; at the same time, a high-sensitivity microphone array receives the actual speech signal of the user's position; the actual speech signal is processed by pre-emphasis filtering and adaptive gain control; the pre-emphasis filtering compensates for high-frequency attenuation; the adaptive gain control maintains the stability of the signal amplitude; the actual speech signal is time-domain aligned with the reference signal of the directional focused sound field; the propagation delay difference between the aligned actual speech signal and the reference signal of the directional focused sound field is analyzed by cross-correlation analysis; and the timestamp of the auricle temperature field change rate and the actual speech signal are strictly synchronized.

[0084] S4.5: Based on the rate of change of the auricle temperature field and the actual speech signal, calculate the user physiological feedback matrix of the acoustic-thermal coupling effect. The expression is as follows:

[0085]

[0086] Where B(i) represents the user physiological feedback matrix of the acoustic-thermal coupling effect at monitoring time point i, i is the current monitoring time point, ξ represents the tissue thermal conductivity coefficient (0.3≤ξ≤0.4), Q represents the average surface area of ​​the auricle, Ω represents the auricle spatial region, x represents the three-dimensional spatial coordinates of the auricle surface, D(x,i) represents the three-dimensional spatial coordinates x of the auricle surface and the auricle surface temperature field measured at monitoring time point i, F represents the fast Fourier transform operator, s1 represents the clean reference speech signal, ⊙ represents the element-wise multiplication of the matrix, H represents the frequency domain sensing weighting function, s2 represents the actual received speech signal, γ represents the acoustic-thermal conversion efficiency factor, and erf represents the Gaussian error function. Represents the spatial gradient of the temperature field. This represents the spatial gradient of the sound pressure field.

[0087] The specific process includes: the rate of change of the auricular temperature field is processed by spatial integration and temporal differentiation to obtain a standardized temperature response index; the actual speech signal is converted into frequency domain features by fast Fourier transform and multiplied element-wise by a frequency domain perceptual weighting function; the thermal conductivity coefficient is used to adjust the sensitivity of the temperature response index; the average surface area of ​​the auricle is used to normalize the temperature field integral result; the spatial gradient of the temperature field is calculated by the central difference method to calculate the temperature change of adjacent measurement points on the auricular surface; the spatial gradient of the sound pressure field is obtained by near-field acoustic holographic reconstruction of the directional focusing sound field; the frequency domain perceptual weighting function highlights the effective frequency band of 300-3400Hz in the speech signal; the Gaussian error function maps the spatial correlation between the temperature field gradient and the sound pressure field gradient to the [-1,1] interval; the acoustic-thermal conversion efficiency factor balances the contribution weights of temperature response and speech intelligibility; and the finally generated user physiological feedback matrix of acoustic-thermal coupling effect includes three dimensions of quantitative indicators: rate of change of temperature, weighted speech signal-to-noise ratio, and gradient correlation. The temporal resolution of these three indicators is aligned with the monitoring time point, and the spatial resolution is consistent with the distribution of measurement points on the auricular surface.

[0088] S5: Based on the auricle temperature change rate and speech intelligibility index in the user's physiological feedback matrix, a standardized feedback vector is generated, and the vibration-acoustic model is updated through a reinforcement learning strategy.

[0089] S5.1: Normalize the user's physiological feedback matrix to generate a standardized feedback vector.

[0090] The specific process includes: extracting the historical mean and standard deviation of the user physiological feedback matrix through sliding window statistical processing; normalizing the temperature change rate component by subtracting the historical temperature mean and dividing by the historical temperature standard deviation to achieve zero mean; compressing the speech signal-to-noise ratio component to the range of [-0.25, 0.25] by subtracting the historical signal-to-noise ratio mean and dividing by four times the historical signal-to-noise ratio standard deviation; mapping the gradient correlation component nonlinearly to the range of [-1, 1] through the hyperbolic tangent function; the normalized temperature change rate component reflecting the degree of abnormality relative to the historical benchmark; the normalized speech signal-to-noise ratio component representing the deviation from typical communication quality; the normalized gradient correlation component maintaining the original physical meaning but eliminating dimensional differences; concatenating the three normalized components column-wise to form a standardized feedback vector; the dimension of the standardized feedback vector being consistent with the user physiological feedback matrix; the element value range of the standardized feedback vector being uniformly adjusted to the range of [-1, 1] for subsequent processing; and the time resolution of the standardized feedback vector being perfectly aligned with the sampling time of the user physiological feedback matrix.

[0091] S5.2: Based on the standardized feedback vector and the vibration-acoustic model, construct the state space of the Markov decision process and input it into the TD3 reinforcement learning policy network to generate vibration-acoustic model optimization instructions.

[0092] The specific process includes: tensor concatenation of the standardized feedback vector and the parameter vector of the current vibration-acoustic model to form joint features; the state space of the Markov decision process is composed of joint features and historical parameter update records; the dimensionality of the Markov decision process state space is reduced to a fixed length through principal component analysis; the dual Critic networks of the TD3 reinforcement learning policy network evaluate the long-term cumulative reward of the state-action pair respectively; the Critic network adopts a three-layer fully connected architecture with layer normalization processing; the Actor network outputs the update direction of the vibration-acoustic model parameters; the action space is constrained within the hypersphere of the parameter space to prevent abrupt changes; the TD3 reinforcement learning policy network uses target network smoothing technology to stabilize the learning process during training; exploration noise is generated by truncating the normal distribution to ensure the rationality of actions; the final output vibration-acoustic model optimization instructions include parameter update direction and step size information; the format of the optimization instructions strictly matches the parameter structure of the vibration-acoustic model; and the execution order of the optimization instructions is consistent with the temporal relationship of the state space.

[0093] Furthermore, the specific architecture of the TD3 reinforcement learning policy network includes an Actor network (outputting deterministic actions) and two independent Critic networks (used to suppress Q-value overestimation), each equipped with a corresponding target network. Stable training is achieved through delayed policy updates (e.g., Critic updates twice, then Actor updates once) and target policy smoothing (adding truncation noise).

[0094] S5.3: Verify the physiological safety constraints of the vibration-acoustic model optimization instructions and iteratively update the vibration-acoustic model.

[0095] The specific process includes: the vibration-acoustic model optimization command is verified for safety through temperature change rate threshold detection; when the temperature change rate predicted by the optimization command exceeds the temperature change rate threshold, a protection mechanism is triggered, which reduces the parameter update amount proportionally to a safe range; the verified vibration-acoustic model optimization command is applied to the vibration-acoustic model through parameter update rules, and the parameter update is achieved using the momentum-accelerated gradient descent method, with the momentum coefficient dynamically adjusted according to the consistency of the historical update direction; the weight matrix of the vibration-acoustic model is iteratively updated according to the direction and step size given by the optimization command; the updated vibration-acoustic model is immediately put into the sound field generation process for effect verification; the verification results are fed back to the experience playback buffer of the TD3 reinforcement learning policy network; the stability boundary of the vibration-acoustic model remains unchanged throughout the update process, ensuring that the acoustic output is always within the physiological safety range; and the frequency of iterative updates is synchronized with the sound field refresh rate.

[0096] The temperature change rate threshold is set according to medical safety standards. The specific value is determined by statistically analyzing the extreme change rates in historical temperature data and combining them with the safe operating range of the equipment.

[0097] The safe range refers to the limit value that will not cause harm or danger during the operation of an activity or equipment. It is usually set based on industry standards, medical safety data or equipment performance parameters, and the specific value is determined through experimental verification or historical statistics.

[0098] The momentum coefficient is a hyperparameter used in optimization algorithms to control the influence of historical gradients on the current parameter update. Its value range is [0, 1). It is generally set empirically (e.g., 0.9) or dynamically adjusted to balance convergence speed and stability.

[0099] This embodiment also provides a voice interaction system based on a smart doll, including: a signal acquisition module, a modal analysis module, a speech reconstruction module, a sound field focusing module, and a model optimization module. The signal acquisition module is used to acquire vibration displacement signals of the user's neck tissue and environmental audio signals, and perform preprocessing to generate a vibration feature matrix and an acoustic feature matrix. The modal analysis module is used to perform time-frequency analysis processing on the vibration feature matrix, combine it with the density characteristics of human neck tissue, and generate a set of modal parameters for vocal cord vibration through biomechanical modeling. The speech generation module is used to input the modal parameter set of vocal cord vibration into the vibration-acoustic model. The model generates a sound source excitation field and uses an acoustic feature matrix to compensate for environmental noise in the sound source excitation field. It also generates a time-domain clean speech signal through the sound wave equation. The sound field focusing module is used to collect the user's real-time position coordinates, calculate and drive the piezoelectric speaker array of the smart doll to adjust the phase based on the spectral characteristics of the time-domain clean speech signal, form a directional focused sound field, record the user's physiological feedback, and generate a user physiological feedback matrix. The model optimization module is used to generate a standardized feedback vector based on the auricular temperature change rate and speech intelligibility index in the user physiological feedback matrix, and update the vibratory-acoustic model through a reinforcement learning strategy.

[0100] This embodiment also provides a computer device applicable to the voice interaction method based on a smart doll, comprising: a memory and a processor; the memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions to implement the voice interaction method based on a smart doll as proposed in the above embodiment.

[0101] The computer device can be a terminal, comprising a processor, memory, communication interface, display screen, and input devices connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, carrier networks, NFC (Near Field Communication), or other technologies. The display screen can be an LCD screen or an e-ink screen. The input devices can be a touch layer covering the display screen, buttons, a trackball, or a touchpad on the computer device's casing, or an external keyboard, touchpad, or mouse.

[0102] This embodiment also provides a storage medium storing a computer program, which, when executed by a processor, implements the voice interaction method based on an intelligent doll as proposed in the above embodiments. The storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read Only Memory (EPROM), Programmable Red-Only Memory (PROM), Read-Only Memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.

[0103] In summary, this invention achieves physical modeling of vocal cord vibration state by solving the vocal cord displacement mode function vector using nonlinear integrals and augmented Lagrange algorithms, thereby improving the acquisition accuracy and noise resistance of speech signals at the source. Furthermore, by collecting the user's auricular temperature change rate and speech intelligibility index to construct a user physiological feedback matrix, adaptive optimization can be performed based on the user's physiological response and speech perception quality, further enhancing the personalization and comfort of voice interaction.

[0104] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.

Claims

1. A voice interaction method based on a smart doll, characterized in that: include, The vibration displacement signal of the user's neck tissue and the ambient audio signal are collected and preprocessed to generate a vibration feature matrix and an acoustic feature matrix. The vibration feature matrix is ​​processed by time-frequency analysis, combined with the density characteristics of human neck tissue, and generated by biomechanical modeling to produce a set of modal parameters for vocal cord vibration. The modal parameter set of vocal cord vibration is input into the vibration-acoustic model to generate the sound source excitation field. The acoustic feature matrix is ​​used to compensate for environmental noise in the sound source excitation field, and a time-domain clean speech signal is generated through the sound wave equation. The system collects the user's real-time location coordinates, calculates and drives the piezoelectric speaker array of the smart doll to adjust its phase based on the spectral characteristics of the time-domain pure speech signal, forms a directional focused sound field, records the user's physiological feedback, and generates a user physiological feedback matrix. Based on the auricular temperature change rate and speech intelligibility index in the user's physiological feedback matrix, a standardized feedback vector is generated, and the vibration-acoustic model is updated through a reinforcement learning strategy.

2. The voice interaction method based on a smart doll as described in claim 1, characterized in that: The specific steps for generating the vibration feature matrix and acoustic feature matrix are as follows: By capturing the vibration displacement signal of the user's neck tissue through the Doppler effect, simultaneously acquiring the ambient audio signal, and eliminating the interference of blood vessel and muscle fluctuations, pure vocal cord vibration waves are generated. The pure vocal cord vibration wave is dynamically decomposed at multiple scales, and time-varying feature weights are generated through the conduction properties of biological tissues. The time-varying feature weights are matched and compensated with the ambient audio signal to generate a vibration feature matrix and an acoustic feature matrix.

3. The voice interaction method based on intelligent dolls as described in claim 2, characterized in that: The specific steps for generating the modal parameter set for vocal cord vibration are as follows: Based on the dynamic selection of time-varying windows using the vibration feature matrix, a local wave energy density function is generated through nonlinear integration and mapped to a wave energy topology matrix characterizing the tissue vibration curvature. The density characteristics of human neck tissue are transformed into biomechanical constraints of mass conservation. Combined with the wave energy topology matrix, the displacement mode function vector is iteratively solved using the augmented Lagrange algorithm. The displacement modal function vector is used to guide the neural network to compensate for the residuals, and the resulting data is then fused to generate a set of modal parameters for vocal cord vibration.

4. The voice interaction method based on intelligent dolls as described in claim 3, characterized in that: The process involves inputting the modal parameter set of vocal cord vibration into a vibration-acoustic model to generate a sound source excitation field, compensating for environmental noise in the sound source excitation field using an acoustic feature matrix, and generating a time-domain clean speech signal through the sound wave equation. The specific steps are as follows: The modal parameter set of vocal cord vibration is input into the vibration-acoustic model to generate the sound source excitation field; The sound source excitation field and the acoustic feature matrix are input into the vibration-acoustic model. The sound wave equation is solved and the environmental noise field is compensated simultaneously by the deep coupling operator to generate the sound pressure field of the sound channel. The sound pressure field of the vocal tract is input into the radiation reconstruction process of the vibration-acoustic model, and a time-domain clean speech signal is generated by the surface integral acoustic impedance compensation function of the lip curvature.

5. The voice interaction method based on a smart doll as described in claim 4, characterized in that: The process of collecting the user's real-time location coordinates, calculating and adjusting the phase of the smart doll's piezoelectric speaker array based on the spectral characteristics of the time-domain pure speech signal, is as follows: Perform time-frequency transformation on the clean speech signal in the time domain to generate a time-frequency distribution matrix; The system collects the user's real-time location coordinates, combines them with the time-frequency distribution matrix, calculates the acoustic wave phase adjustment through quantum optimization, and drives the piezoelectric speaker array to adjust the phase based on the acoustic wave phase adjustment.

6. The voice interaction method based on a smart doll as described in claim 5, characterized in that: The specific steps for forming a directional focused sound field, recording user physiological feedback, and generating a user physiological feedback matrix are as follows. A directional focused sound field is generated based on a phase-adjusted piezoelectric loudspeaker array. Based on the directional focusing sound field, the rate of change of the user's auricle temperature field is collected, while the actual voice signal is received. The user physiological feedback matrix of the acoustic-thermal coupling effect is calculated based on the rate of change of the auricular temperature field and the actual speech signal.

7. The voice interaction method based on intelligent dolls as described in claim 6, characterized in that: The process involves generating a standardized feedback vector based on the auricle temperature change rate and speech intelligibility index from the user's physiological feedback matrix, and updating the vibration-acoustic model using a reinforcement learning strategy. The specific steps are as follows: The user's physiological feedback matrix is ​​normalized to generate a standardized feedback vector; Based on the standardized feedback vector and the vibration-acoustic model, a Markov decision process state space is constructed and input into the TD3 reinforcement learning policy network to generate vibration-acoustic model optimization instructions. Physiological safety constraints were verified on the optimization instructions for the vibration-acoustic model, and the vibration-acoustic model was iteratively updated.

8. A voice interaction system based on a smart doll, based on the voice interaction method based on a smart doll as described in any one of claims 1 to 7, characterized in that: It includes a signal acquisition module, a modal analysis module, a speech reconstruction module, a sound field focusing module, and a model optimization module. The signal acquisition module is used to acquire vibration displacement signals of the user's neck tissue and environmental audio signals, and to preprocess them to generate vibration feature matrices and acoustic feature matrices. The modal analysis module is used to perform time-frequency analysis on the vibration feature matrix, combine the density characteristics of human neck tissue, and generate a set of modal parameters for vocal cord vibration through biomechanical modeling. The speech generation module is used to input the modal parameter set of vocal cord vibration into the vibration-acoustic model, generate the sound source excitation field, use the acoustic feature matrix to compensate for environmental noise in the sound source excitation field, and generate a clean speech signal in the time domain through the sound wave equation. The sound field focusing module is used to collect the user's real-time position coordinates, calculate and drive the piezoelectric speaker array of the smart doll to adjust the phase based on the spectral characteristics of the time-domain pure speech signal, form a directional focused sound field, record the user's physiological feedback, and generate the user's physiological feedback matrix. The model optimization module is used to generate standardized feedback vectors based on the auricle temperature change rate and speech intelligibility index in the user's physiological feedback matrix, and to update the vibratory-acoustic model through a reinforcement learning strategy.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that: When the processor executes the computer program, it implements the steps of the voice interaction method based on intelligent dolls as described in any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by the processor, it implements the steps of the voice interaction method based on intelligent dolls as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Interactive foreign language speech training auxiliary system for hearing-impaired children

    CN119889135A

  • Sound coordination method and system based on tinnitus condition

    CN119889582A