AI toy voiceprint recognition interaction method, device and equipment

Through the combination of sound field estimation parameters and background noise characteristics, deep neural networks and credibility judgment models are used to solve the problems of low recognition accuracy and low user distinction in complex environments, and more efficient user identity recognition and interactive behavior understanding are achieved.

CN120496540AInactive Publication Date: 2025-08-15SHENZHEN PEMI TECHNOLOGY CO LTD
View PDF 0 Cites 9 Cited by

Patent Information

Application Number
CN202510841205.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-23
Publication Date
2025-08-15
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

In the case where multiple users live together, significant environmental noise interference or complex semantic emotional interactions, existing AI toys have low recognition accuracy and low user distinction.

Method used

By obtaining the original audio data set of the target voice signal, acoustic field estimation parameters and background noise characteristics are generated, voiceprint feature extraction is performed, combined with deep neural network and credibility determination model, identity verification and semantic intent analysis are performed, and multimodal response instructions are generated.

Benefits of technology

It improves the recognition accuracy and user distinction in complex environments, achieves more accurate user identity recognition and interactive behavior understanding, and improves the intelligence level and responsiveness of AI toys.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120496540A_ABST
    Figure CN120496540A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of AI toys, and provides an AI toy voiceprint recognition interaction method, device and equipment, and the method comprises the steps: obtaining a to-be-recognized target voice signal, extracting an original audio data set to generate a sound field estimation parameter and a background noise feature, carrying out the voiceprint feature extraction of the target voice signal, obtaining a voiceprint feature vector, and obtaining the voiceprint feature vector; and performing feature clustering on the voiceprint feature vector and a preset child voiceprint vector set to obtain user identity information, a behavior tag and an emotion tag so as to generate a corresponding multi-modal response instruction, and inputting the multi-modal response instruction into a control module of the AI toy. Robustness of a target voice signal in a complex environment is improved by combining sound field estimation parameters and background noise features, and through voiceprint feature vector extraction and clustering and similarity calculation, the target voice signal is improved under the condition that environmental noise interference is remarkable or semantic emotion interaction is complex. The problems of low identification accuracy and low user discrimination degree exist in the prior art.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of AI toys, and in particular to methods, devices, and equipment for interactive voiceprint recognition of AI toys. Background Art

[0002] With the continuous integration of artificial intelligence technology and smart children's toys, AI toys with voice recognition and voice interaction capabilities have gradually become a focus of market attention. By integrating voice interaction modules, these toys can realize functions such as language recognition, semantic understanding, and personalized feedback during the interaction with children, thereby enhancing children's learning interest and interactive experience. In the home setting, voice, as a natural and convenient input method, has become the primary method for AI toys to obtain information and interactive control.

[0003] Among the related technical means, a voice recognition mechanism based on a fixed voice wake-up word is adopted, combined with a shallow voiceprint comparison model for preliminary identity screening. Existing AI toys mostly rely on end-side voice processing chips or connect to cloud services to perform basic voice enhancement processing and keyword extraction on the received voice signals. After that, the template in the voiceprint library is compared with the current voice features for similarity to identify the user's identity and trigger the preset response strategy. This method can relatively accurately complete identity judgment and basic interactive intention response in a static environment, and has been widely used in some commercial children's educational toys and voice companion devices.

[0004] Regarding the above technical solution, although basic voice recognition interaction functions can be achieved through voiceprint comparison and fixed command matching, in actual applications, especially when multiple users coexist, environmental noise interference is significant, or semantic and emotional interactions are complex, there are still problems of low recognition accuracy and low user differentiation. Summary of the Invention

[0005] In order to improve the problems of low recognition accuracy and low user differentiation when there is significant environmental noise interference or complex semantic emotional interaction, the present application provides an AI toy voiceprint recognition interaction method, device and equipment.

[0006] The present invention provides an AI toy voiceprint recognition interaction method, comprising: obtaining a target voice signal to be recognized, extracting an original audio data set corresponding to the target voice signal, and generating sound field estimation parameters and background noise characteristics according to the original audio data set; performing voiceprint feature extraction on the target voice signal using the sound field estimation parameters and the background noise characteristics to obtain a voiceprint feature vector, performing feature clustering on the voiceprint feature vector and a preset children's voiceprint vector set to obtain an identity clustering result and a similarity index; inputting the identity clustering result and the similarity index into a preset credibility determination model to obtain user identity information and a credibility level, and verifying the user identity information based on the credibility level to obtain a verification result; performing identity-weighted screening on the target voice signal using the verification result to obtain semantic intent information and contextual state information, constructing a behavior label based on the semantic intent information, and constructing an emotion label based on the contextual state information; generating corresponding multimodal response instructions based on the user identity information, behavior label, and emotion label, and inputting the multimodal response instructions into a control module of the AI toy.

[0007] As a preferred solution, the steps of obtaining a target speech signal to be recognized, extracting an original audio data set corresponding to the target speech signal, and generating sound field estimation parameters and background noise characteristics based on the original audio data set include: using multiple microphone arrays to synchronously collect the target speech signal to obtain original audio data sets of different spatial channels, performing time-frequency analysis and beam reconstruction on the original audio data set, and extracting frequency domain power spectrum distribution and time domain phase offset characteristics; calculating the sound source direction estimation value and the echo path length based on the frequency domain power spectrum distribution and the time domain phase offset characteristics, and constructing sound field estimation parameters using the sound source direction estimation value and the echo path length; performing background energy statistics and non-speech component detection on the original audio data set, extracting the spectral envelope and stability parameters of the background noise, and generating a background noise feature vector based on the spectral envelope and the stability parameters.

[0008] As a preferred solution, the target speech signal is subjected to voiceprint feature extraction through the sound field estimation parameters and the background noise characteristics to obtain a voiceprint feature vector, and the voiceprint feature vector is subjected to feature clustering with a preset children's voiceprint vector set to obtain an identity clustering result and a similarity index, including the following steps: spatially filtering the sound field estimation parameters to obtain filtered sound field estimation parameters, using the filtered sound field estimation parameters to perform main sound source enhancement on the target speech signal to obtain an enhanced target speech signal; using the background noise characteristics to perform noise adaptive elimination on the enhanced target speech signal to obtain a pure A speech signal is obtained by performing short-time Fourier transform and perceptual spectrum mapping on the pure speech signal to extract multi-scale voiceprint candidate features; the multi-scale voiceprint candidate features are fused with sound source direction information to construct a spatial voiceprint feature vector; wherein the sound source direction information refers to the three-dimensional positioning information of the target sound source constructed according to the sound source direction estimation value in combination with the spatial distribution model of the microphone array; the spatial voiceprint feature vector is subjected to spectral clustering and local density estimation with a preset children's voiceprint vector set to obtain cluster centers and a similarity matrix, identity clustering results are generated based on the cluster centers, and a similarity index is generated using the similarity matrix.

[0009] As a preferred embodiment, the steps of inputting the identity clustering results and the similarity index into a preset credibility determination model to obtain user identity information and a credibility level, and verifying the user identity information based on the credibility level to obtain a verification result include: extracting the feature center and distribution radius of the category according to the identity clustering results, and generating an identity confidence envelope using the feature center and the distribution radius; correlating and matching the similarity index with historical user interaction records to obtain a matching result, and calculating a behavior consistency score and a temporal stability score based on the matching result; wherein the historical user interaction records refer to a historical set of user interaction conversations recorded by the AI toy for the same spatial voiceprint feature vector over multiple time periods, including behavior trajectory data, semantic response labels, tone and emotional patterns, command execution feedback, and voice environment information; inputting the behavior consistency score, the temporal stability score, and the identity confidence envelope into a preset credibility determination model to output user identity information and a credibility level, hierarchically screening the user identity information according to the credibility level to construct an identity credibility zone boundary; and matching and verifying the identity credibility zone boundary with a preset permission configuration in the AI toy, and outputting a verification result.

[0010] As a preferred solution, the steps of using the verification result to perform identity-weighted screening on the target speech signal to obtain semantic intent information and contextual state information, constructing a behavior label based on the semantic intent information, and constructing an emotion label based on the contextual state information include: extracting semantic segment boundaries and temporal structures based on the target speech signal, constructing a preliminary semantic candidate graph based on the semantic segment boundaries, and constructing a context path structure based on the temporal structure; using the verification result to perform identity-weighted screening on the preliminary semantic candidate graph to obtain a semantic intent candidate vector, and generating contextual state information associated with the context path structure based on the semantic intent candidate vector; jointly embedding the semantic intent candidate vector and the user identity information into a preset dual-channel attention gating network to generate an emotion label. The method comprises the following steps: first, forming semantic intention information, constructing a context consistency map based on the context state information, and performing context reasoning on the context consistency map in combination with the semantic segment position and the turn order to obtain a context reasoning result; wherein, the semantic segment position refers to the start and end time index of the semantic unit in the target speech signal in the overall speech and the position of its syntactic structure; the turn order refers to the turn numbering information constructed according to the speech timing and response relationship during the alternating dialogue between the user and the AI toy; analyzing the intonation emotion distribution and the emotion tendency factor through the context reasoning result, the semantic segment boundary and the temporal structure; jointly matching the semantic intention information with the intonation emotion distribution to obtain a behavior label, and fusing the context reasoning result with the emotion tendency factor to generate an emotion label.

[0011] As a preferred solution, the step of analyzing the intonation emotion distribution and emotional tendency factor through the contextual reasoning results, the semantic segment boundaries and the temporal structure includes: constructing semantic temporal alignment data according to the semantic segment boundaries and the temporal structure, and extracting the intonation change envelope and the silent gap distribution of the speech waveform based on the semantic temporal alignment data; extracting the tone change features using the intonation change envelope, and extracting the pause features based on the silent gap distribution; fusing the tone change features and the pause features through the contextual reasoning results to extract the intonation emotion distribution and emotional tendency factor.

[0012] As a preferred solution, the step of generating corresponding multimodal response instructions based on the user identity information, behavior tags and emotion tags, and inputting the multimodal response instructions into the control module of the AI toy includes: inputting the behavior tags into a preset intention classification library for matching to obtain corresponding candidate semantic units, inputting the emotion tags into a preset expression and action library for matching to obtain corresponding visual feedback resources, and extracting candidate image elements and action trajectories based on the visual feedback resources; performing consistent fusion on the candidate semantic units, the candidate image elements and the action trajectories to generate a multimodal candidate response instruction set; obtaining the current interaction permission level according to the user identity information, screening and prioritizing the multimodal candidate response instruction set based on the interaction permission level to obtain multimodal response instructions, and inputting the multimodal response instructions into the control module of the AI toy.

[0013] The present application also provides an AI toy voiceprint recognition interactive device, comprising: an acquisition module, configured to acquire a target voice signal to be recognized, extract an original audio data set corresponding to the target voice signal, and generate sound field estimation parameters and background noise characteristics based on the original audio data set; An extraction module is used to extract voiceprint features of the target voice signal through the sound field estimation parameters and the background noise characteristics to obtain a voiceprint feature vector, and perform feature clustering on the voiceprint feature vector and a preset child voiceprint vector set to obtain an identity clustering result and a similarity index; a verification module is used to input the identity clustering result and the similarity index into a preset credibility judgment model to obtain user identity information and a credibility level, and verify the user identity information based on the credibility level to obtain a verification result; a decoding module is used to perform identity-weighted screening on the target voice signal using the verification result to obtain semantic intent information and contextual state information, construct a behavior label based on the semantic intent information, and construct an emotion label based on the contextual state information; a generation module is used to generate corresponding multimodal response instructions based on the user identity information, behavior label and emotion label, and input the multimodal response instructions into the control module of the AI toy.

[0014] The present application also provides an electronic device, including a memory and a processor, wherein the memory stores a computer program that can be run on the processor, and when the processor executes the computer program, it implements any one of the above-mentioned AI toy voiceprint recognition interaction methods.

[0015] Compared with existing technologies, this application has the following advantages: high recognition accuracy and strong user differentiation. By combining sound field estimation parameters with background noise characteristics, the robustness of the target speech signal in complex environments is improved. The voiceprint feature vector is extracted through a deep neural network, achieving accurate modeling of the child user's voiceprint. Based on clustering and similarity calculation, a credibility judgment model is further introduced to improve the problems of low recognition accuracy and poor user differentiation in situations with significant environmental noise interference or complex semantic and emotional interactions. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0017] The structures, proportions, sizes, etc. depicted in the drawings of this specification are only used to match the contents disclosed in the specification so as to facilitate understanding and reading by persons familiar with this technology. They are not intended to limit the conditions under which the present invention can be implemented and therefore have no substantive technical significance. Any structural modifications, changes in proportional relationships, or adjustments in size should still fall within the scope of the technical contents disclosed in the present invention without affecting the effects and objectives that can be achieved by the present invention.

[0018] Figure 1 1 is a flow chart of an AI toy voiceprint recognition interaction method provided by an embodiment of the present invention; Figure 2 This is a schematic block diagram of the structure of an AI toy voiceprint recognition interaction device provided by an embodiment of the present invention; Figure 3 It is a schematic block diagram of the structure of an electronic device provided by an embodiment of the present invention.

[0019] Description of reference numerals: 10. AI toy voiceprint recognition interactive device; 11. Acquisition module; 12. Extraction module; 13. Verification module; 14. Decoding module; 15. Generation module; 20. Electronic device; 21. Memory; 22. Processor. DETAILED DESCRIPTION

[0020] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0021] The flowcharts shown in the accompanying drawings are for illustrative purposes only and do not necessarily include all contents and operations / steps, nor must they be executed in the order described. For example, some operations / steps may be decomposed, combined, or partially merged, so the actual execution order may vary depending on the actual situation.

[0022] It should also be understood that the terms used in this specification are for the purpose of describing specific embodiments only and are not intended to limit the present application. As used in this specification and the appended claims, the singular forms "a," "an," and "the" are intended to include the plural forms unless the context clearly indicates otherwise.

[0023] It should be further understood that the term "and / or" used in this specification and the appended claims refers to and includes any and all possible combinations of one or more of the associated listed items.

[0024] The technical solution of the present invention will be further described below with reference to the accompanying drawings and through specific implementation methods.

[0025] Example 1: like Figure 1 As shown, the present application provides an AI toy voiceprint recognition interaction method, including steps S100 to S500.

[0026] Step S100: Acquire a target speech signal to be recognized, extract an original audio data set corresponding to the target speech signal, and generate sound field estimation parameters and background noise features according to the original audio data set.

[0027] In this step, the built-in sound pickup module of the AI toy is used to collect external input voice data to obtain the target voice signal to be recognized; then the target voice signal is preprocessed to extract the corresponding original audio data set, and the preprocessing includes frame segmentation, windowing, endpoint detection, pre-emphasis and other steps; specifically, by performing spatial acoustic modeling on the original audio data set, the spatial convolution method is used to extract the propagation characteristics of the target voice signal in three-dimensional space, and combined with the position information of the microphone array, a sound field estimation model is constructed to obtain the sound field estimation parameters; at the same time, a method based on Short-Time Fourier Transform (STFT) is used to perform spectral analysis on the audio frame, and background noise features are extracted from the silent segments and non-speech segments, including statistical features such as noise energy spectral density and frequency band distribution.

[0028] For example, when extracting sound field estimation parameters, a method based on spatial correlation matrix calculation is used to estimate the direction of the sound source. By calculating the cross-correlation function between different microphone pairs, the dominant propagation direction of the target speech signal is determined to form a spatial feature vector. When extracting background noise features, a Wiener filter can be used to model the spectrum of silent segments to generate background noise estimates, further improving the accuracy of subsequent voiceprint extraction.

[0029] Step S200: extract voiceprint features from the target speech signal using sound field estimation parameters and background noise characteristics to obtain a voiceprint feature vector, perform feature clustering on the voiceprint feature vector and a preset children's voiceprint vector set to obtain an identity clustering result and a similarity index.

[0030] In this step, the sound field estimation parameters and background noise characteristics are used to perform noise suppression and spatial filtering on the target speech signal to enhance the dominant components of the target speech; then a voiceprint feature extraction model jointly constructed by a convolutional neural network (CNN) and a long short-term memory network (LSTM) is used to extract deep voiceprint features in the target speech signal and generate a voiceprint feature vector; specifically, during the voiceprint feature extraction process, the audio frame is input into the voiceprint recognition network, and a voiceprint feature vector of a fixed dimension is output; then, a feature clustering operation is performed with the preset children's voiceprint vector set in the system, and the K-means++ algorithm is used to classify all voiceprint vectors to form an identity clustering result, and the similarity index is calculated in combination with the Euclidean distance or cosine similarity.

[0031] For example, in the voiceprint feature extraction model, the input layer is Mel-Frequency Cepstral Coefficients (MFCCs) or Log-Mel spectrograms. After three convolutional layers and two bidirectional LSTM layers are stacked, the final output is a 128-dimensional voiceprint feature vector. When using K-means++, the initial cluster centers are initialized using the maximum distance difference, which improves convergence speed and clustering effect.

[0032] Step S300: Input the identity clustering result and the similarity index into a preset credibility judgment model to obtain user identity information and credibility level, and verify the user identity information based on the credibility level to obtain a verification result.

[0033] In this step, the identity clustering results and similarity indicators are input into the trained credibility judgment model for judgment; specifically, the credibility judgment model is a two-level model based on the fusion of Gradient Boosting Decision Tree (GBDT) and Support Vector Machine (SVM), and the model input features include voiceprint vector stability, similarity threshold fluctuation, clustering distance, etc.; the model output results include user identity information and credibility level, and the credibility level is an interval value (such as high, medium, and low); during the verification process, the current user identity information is compared with the historical authentication records. If the credibility level is high and the identity matches, the verification success result is output; otherwise, the verification failure result is output.

[0034] For example, in actual deployment, the GBDT model first performs a preliminary classification based on multiple statistical features, and then the SVM model determines boundary conditions to improve robustness under small sample sizes. If the similarity of a certain input voiceprint vector is 0.86, the corresponding cluster center number is C3, and the model determines the credibility level as "high", the verification result is "passed".

[0035] Step S400: Use the verification result to perform identity-weighted screening on the target speech signal to obtain semantic intent information and contextual state information, construct a behavior label based on the semantic intent information, and construct an emotion label based on the contextual state information.

[0036] In this step, the voice content is identity-weighted according to the verification results to enhance the semantic understanding priority of trusted users and weaken the impact of voice segments with unclear identities or low credibility. Specifically, the verified target voice signal is semantically parsed, and the semantic intent recognition model of the Transformer architecture is used to extract semantic intent information. At the same time, the emotion recognition model is used to model the pitch, speaking rate, loudness, etc. to extract contextual state information. Subsequently, the semantic intent information is mapped to behavior labels, such as "request toy operation", "ask questions", etc.; the contextual state information is mapped to emotion labels, such as "happy", "doubtful", "anxious", etc.

[0037] For example, the semantic intent recognition model uses BERT-Encoder as the underlying semantic understanding module, using a multi-head attention mechanism to identify keywords and contextual meanings, and output action labels such as "I want to tell a story." The emotion recognition model uses the GRU+Attention mechanism to extract time series audio features and identify the current emotional state as "happy."

[0038] Step S500: Generate corresponding multimodal response instructions based on user identity information, behavior tags, and emotion tags, and input the multimodal response instructions into the control module of the AI toy.

[0039] In this step, specific multimodal response instructions such as voice, image, and action are generated in combination with user identity information, behavior tags, and emotion tags. Specifically, a response decision map is constructed, identity and behavior are mapped to response priorities, and the response resource module is scheduled to generate an instruction set. Finally, the generated multimodal response instructions (such as voice synthesis instructions, expression and lighting control instructions, and action execution instructions) are synchronously input into the control module of the AI toy to control the AI toy to make a situational response.

[0040] For example, when the user identity is identified as "main user", the behavior label is "request to dance", and the emotion label is "happy", the system generates a multimodal instruction including cheerful music playing, flashing lights, and body rhythms, and the control module executes the dance response.

[0041] In this embodiment, a target voice signal to be recognized is obtained, and an original audio data set corresponding to the target voice signal is extracted, and sound field estimation parameters and background noise characteristics are generated according to the original audio data set; then, voiceprint features are extracted from the target voice signal using the sound field estimation parameters and the background noise characteristics to obtain a voiceprint feature vector, and the voiceprint feature vector is feature clustered with a preset child voiceprint vector set to obtain an identity clustering result and a similarity index; then, the identity clustering result and the similarity index are input into a preset credibility judgment model to obtain user identity information and a credibility level, and the user identity information is verified based on the credibility level to obtain a verification result; further, the target voice signal is identity-weightedly screened using the verification result to obtain semantic intent information and contextual state information, and a behavior label is constructed based on the semantic intent information, and an emotion label is constructed based on the contextual state information; finally, a corresponding multimodal response instruction is generated based on the user identity information, the behavior label, and the emotion label, and the multimodal response instruction is input into a control module of the AI toy. It can effectively extract voiceprint feature information from the target voice signal in a complex sound field environment, and optimize the processing based on the sound source spatial information and background noise characteristics, thereby improving the accuracy and robustness of user identity recognition; through the credibility analysis of voiceprint clustering results and similarity indicators, it realizes dynamic verification and hierarchical screening of user identity information, and enhances the security and credibility of interactive behavior; combining semantic intention information and contextual state information to further construct behavioral labels and emotional labels, which helps AI toys to more accurately understand children’s true intentions and emotional states, and achieve a more natural and personalized interactive experience; by generating multimodal response instructions that match the user’s identity and behavioral state, it realizes the fusion control of voice, action, image and other feedback, comprehensively improving the intelligence level and response ability of AI toys in the process of human-computer interaction, and improving the problems of low recognition accuracy and low user differentiation when the environmental noise interference is significant or the semantic emotional interaction is complex.

[0042] Example 2: In step S100, a target speech signal to be recognized is obtained, and an original audio data set corresponding to the target speech signal is extracted. The steps of generating sound field estimation parameters and background noise features based on the original audio data set specifically include: Multiple microphone arrays are used to synchronously collect the target speech signal to obtain the original audio data set of different spatial channels. Time-frequency analysis and beam reconstruction are performed on the original audio data set to extract the frequency domain power spectrum distribution and time domain phase offset features.

[0043] By performing synchronous signal integration operations on the audio signals collected by a multi-channel microphone array, the signal vector corresponding to each channel is extracted. Specifically, the target scene is first modeled based on a preset microphone array layout model. A time synchronization protocol (such as the PTP protocol) is used to calibrate the sampling time of all microphone channels. While ensuring phase consistency, the short-time Fourier transform (STFT) is used to transform the audio signals of each channel into the frequency domain and extract the frequency domain power spectrum (PSD) distribution. At the same time, the time domain phase difference between the relative time delays of each channel is calculated to obtain the time domain phase offset characteristics. During the beam reconstruction process, a beamforming algorithm based on the minimum variance distortionless response (MVDR) is applied to the target sound source signal for directional enhancement processing to minimize the impact of lateral interference sources.

[0044] For example, when the system is configured with a 4×4 array microphone arrangement, the target speech signals are collected in four directions respectively. After the signal is processed by STFT, the power spectrum vector in each window is extracted. For the time delay Δ𝑡 between two adjacent channels, the system calculates the corresponding phase offset 𝜙=2𝜋𝑓Δ𝑡, where 𝑓 is the frequency component. By processing the phase difference matrix of all channels, a high-resolution spatial spectrum is further formed for subsequent sound source direction estimation.

[0045] The sound source direction estimation value and echo path length are calculated based on the frequency domain power spectrum distribution and time domain phase offset characteristics, and the sound field estimation parameters are constructed using the sound source direction estimation value and echo path length.

[0046] By fusing the power spectrum distribution and phase offset information from multiple channels, and inverting the spatial sound pressure gradient distribution and the spatial cross-correlation peak difference, the direction of arrival (DOA) and reflection path of the sound source are estimated. Specifically, the generalized cross-correlation-phase transform (GCC-PHAT) method is used to obtain the time delay estimate between the channel pairs. Combined with the spatial geometric configuration of the microphones, the MUSIC (Multiple Signal Classification) or SRP-PHAT (Steered Response Power with Phase Transform) algorithm is used to accurately calculate the sound source direction angle. At the same time, based on the time difference between the arrival time of the first reflection path and the main path, the echo path length is estimated in combination with the room impulse response (RIR) model, and a sound field estimation parameter set including the sound source direction, reflection path length, microphone layout and spatial sound pressure parameters is further constructed.

[0047] For example, after a voice signal is collected indoors, GCC-PHAT processing determines that the peak delay between the main channel and the reference channel is 0.0003 seconds. Combined with the microphone spacing of 0.04 meters, the sound source azimuth is approximately 45 degrees. By modeling the room dimensions and using a first-order reflection time difference of 0.001 seconds, the reflection path is estimated to be 0.34 meters. The resulting sound field estimation parameters include: azimuth 45°, echo path 0.34m, reflection gain 0.87, and direct sound energy ratio 4.2dB.

[0048] Background energy statistics and non-speech component detection are performed on the original audio dataset, the spectral envelope and stability parameters of the background noise are extracted, and the background noise feature vector is generated based on the spectral envelope and stability parameters.

[0049] By performing energy envelope extraction and modulation frequency analysis on non-speech segments within a long time window, the spectral shape and time-varying characteristics of the background noise are obtained. Specifically, a GMM-HMM acoustic model based on VAD (voice activity detection) is used to divide the speech and non-speech intervals, and Mel spectrum analysis is performed on the audio signal in the non-speech interval to extract the spectral envelope curve. The spectral envelope is then statistically normalized to extract its average spectral slope, low-frequency jitter rate, and modulation frequency response energy as stability parameters. The spectral envelope and stability parameters are combined and encoded into a background noise feature vector, which serves as the basis for adaptive noise suppression in the subsequent voiceprint feature extraction process.

[0050] For example, when a certain segment of original audio is detected to contain a silence segment of about 3 seconds, the system determines it as a non-speech component based on the GMM-HMM model, and then uses a 24-dimensional Mel filter bank to perform spectral analysis on it to obtain the noise spectrum envelope curve. By comparing the power differences between frequency bands, it is concluded that the low-frequency noise slope is -3.1dB / Hz, the high-frequency modulation frequency bandwidth is 120Hz, and the spectral stability index is 0.72. Finally, it is encoded into a set of 32-dimensional background noise feature vectors.

[0051] In step S200, the voiceprint feature extraction of the target speech signal is performed using the sound field estimation parameters and the background noise characteristics to obtain a voiceprint feature vector, and the voiceprint feature vector is subjected to feature clustering with a preset children's voiceprint vector set to obtain an identity clustering result and a similarity index. Specifically, the steps include: The sound field estimation parameters are spatially filtered to obtain filtered sound field estimation parameters, and the target speech signal is enhanced by using the filtered sound field estimation parameters to obtain an enhanced target speech signal.

[0052] By constructing a spatial filter model based on sound source direction constraints, abnormal directional data and non-target reflection information in the sound field estimation parameters are filtered out, thereby improving the directional resolution and response accuracy of the sound field estimation; specifically, a multi-channel spatial filtering technology (such as a linear filter based on MVDR or GEVD) is used to construct a sound source direction selective enhancer, which retains the parameters within the range of ±Δθ near the estimated sound source direction, while imposing a zero response constraint on the directional components corresponding to directions far away from the target direction (such as highly reflective walls or noise sources); then, the filtered directional parameters are input into the beamformer, and the target speech signal is re-weighted and summed to enhance the sound source signal in the target direction and suppress other interfering directions.

[0053] For example, when the system identifies that the main sound source is located at an azimuth angle of 42° and an elevation angle of 15° based on the sound field estimation parameters, it constructs a ±5° directional window and sets the spurious directional responses above 60° to zero gain in the MVDR filter. Ultimately, the signal-to-noise ratio of the enhanced speech signal is improved by approximately 7.4dB, and non-target speech interference is reduced to 32% of the original value.

[0054] The background noise characteristics are used to adaptively eliminate the noise of the enhanced target speech signal to obtain a pure speech signal. The pure speech signal is subjected to short-time Fourier transform and perceptual spectrum mapping to extract multi-scale voiceprint candidate features.

[0055] By constructing a spectral subtraction model or a deep autoencoder denoising network driven by background noise features, the residual noise components in the enhanced target speech signal are adaptively removed. Specifically, the background noise spectral envelope and stability parameters extracted in step S100 are used to generate a noise estimation spectrum and construct an adaptive gated filter, which is applied to the spectral representation of the enhanced signal to achieve dynamic suppression of local frequency band noise. Subsequently, a short-time Fourier transform (STFT) is performed on the pure speech signal to extract the spectrogram, which is then mapped to the perceptual frequency domain (such as a Mel filter bank or a Gammatone filter bank), and its high-frequency transient features, low-frequency resonance peak structure, and energy envelope differences are extracted using multi-scale windows to form a multi-scale voiceprint candidate feature set.

[0056] For example, after a certain enhanced speech signal is processed by the spectral subtraction algorithm, the residual background noise energy is reduced to 18% of the original level. The spectrogram is then calculated using STFT (window length 25ms, overlap rate 60%), and a 40-dimensional Mel filter bank is used to map the perceptual spectrum to form a short-time perceptual spectrogram. Voiceprint feature vectors are extracted at three time scales (20ms, 50ms, and 100ms) and encoded as early resonance vectors, high-frequency change vectors, and energy envelope curves, respectively, which together constitute multi-scale voiceprint candidate features.

[0057] The multi-scale voiceprint candidate features are fused with the sound source direction information to construct a spatial voiceprint feature vector; the sound source direction information refers to the three-dimensional positioning information of the target sound source constructed based on the sound source direction estimation value combined with the microphone array spatial distribution model, which is used to describe the incident direction, sound angle and sound distance of the speech signal in space, and is used to enhance the spatial separability and individual differences of the voiceprint features.

[0058] By performing feature splicing or attention-guided fusion of the voiceprint candidate features and the three-dimensional spatial position information constructed by the sound source direction, the sensitivity of the voiceprint features to the spatial dimension and the ability to capture individual behavioral differences are enhanced. Specifically, first, based on the sound source estimation angle (azimuth, pitch angle) and the microphone array coordinate model, the three-dimensional spatial position (x, y, z) of the target sound source is reconstructed using triangulation and maximum likelihood estimation methods. The direction cosine vector and distance information of the sound path are further calculated in combination with the room response model to form a directional feature vector containing dimensions such as sound direction, distance, and spatial phase consistency. Then, the attention fusion mechanism is used to take the directional features as weight input, and the multi-scale voiceprint features are subjected to direction-sensitive weighted processing in the fusion module, and finally the fused spatial voiceprint feature vector is output.

[0059] For example, the three-dimensional positioning result of a child's voice source is (1.2m, 0.8m, 1.0m), which is about 1.6 meters away from the center of the microphone, and the direction cosine is (0.58, 0.38, 0.65). The system concatenates this vector with the voiceprint feature vector of each frame and inputs it into the Transformer attention module. It uses spatial weights to adjust the channel activation and finally forms a spatial voiceprint feature representation with a dimension of 128. Compared with traditional voiceprint feature models, the similarity error recognition rate is reduced by about 12% in cross-microphone and non-face-to-face speech scenarios.

[0060] The spatial voiceprint feature vector is spectrally clustered and local density estimated with the preset children's voiceprint vector set to obtain the cluster center and similarity matrix. The identity clustering result is generated based on the cluster center, and the similarity index is generated using the similarity matrix.

[0061] By constructing a joint clustering framework based on spectral clustering and local density estimation, similarity graph construction and spectral decomposition are performed on spatial voiceprint feature vectors to achieve separation and identity aggregation between multiple sound sources. Specifically, the Euclidean distance matrix or cosine similarity matrix between the vector to be identified and the children's voiceprint vector set is first calculated to construct a weighted adjacency graph. Then, the Laplace matrix is used for spectral decomposition, and the first K eigenvectors are extracted and clustered using the K-means algorithm. At the same time, the local density and relative density difference of each spatial voiceprint vector in the feature space are calculated to estimate its possibility of being a cluster center, and finally the cluster center is determined. Identity clusters are generated based on the clustering results, and an identity similarity index is constructed by combining the average similarity inside and outside the cluster.

[0062] For example, the system constructs a 64×64 similarity matrix for the currently collected target voiceprint features and 50 preset child voiceprint vectors, uses spectral clustering to extract the first three feature components, and forms three clusters through K-means. The cosine similarity between one cluster center and the target feature reaches 0.94, and the local density peak is located at the center vector. The corresponding identity is identified as a "high-credibility child identity" with a similarity index of 0.91, which exceeds the credibility threshold set by the system (0.85).

[0063] In step S300, the identity clustering result and the similarity index are input into a preset credibility determination model to obtain user identity information and credibility level, and the user identity information is verified based on the credibility level to obtain the verification result. The steps specifically include: According to the identity clustering results, the feature center and distribution radius of the category are extracted, and the identity confidence envelope is generated using the feature center and distribution radius.

[0064] By extracting the center point vector and the maximum boundary distance of the corresponding category from the clustering results of the spatial voiceprint feature vector, a confidence envelope boundary area for identity determination is formed; specifically, the center vector is first calculated for each cluster. , and then calculate the Euclidean distance from all vectors in the cluster to the center , with the maximum distance Or the statistical radius within the 95% confidence interval As the envelope radius; Defined as the identity confidence envelope, it represents the valid identity area that can be covered by the current cluster center.

[0065] For example, for the cluster to which the target voiceprint belongs, the cluster center vector It is a 128-dimensional vector, the maximum Euclidean distance of all samples in it is 2.6, and the distance distribution radius of 95% of the samples is 2.1. The system selects 2.1 as the confidence radius of the identity and constructs a spherical confidence envelope for subsequent credibility judgment.

[0066] The similarity index is associated and matched with the historical user interaction records to obtain the matching results, and the behavioral consistency score and the temporal stability score are calculated based on the matching results; among them, the historical user interaction records refer to the user interaction session history set recorded by the AI toy for the same spatial voiceprint feature vector in multiple time periods, including behavioral trajectory data, semantic response labels, tone and emotional patterns, command execution feedback and voice environment information, which are used to analyze the consistency and deviation between the current interaction behavior and the historical behavior.

[0067] By constructing a user's historical interaction profile, the current interaction features are matched with historical behavior patterns in semantic, emotional and time series dimensions to extract behavioral consistency and temporal stability indicators. Specifically, the historical record sample set that matches the current spatial voiceprint features is first retrieved. , extract key behavior sequence features (such as semantic intent sequence, intonation change pattern, emotion labeling, etc.), perform multi-dimensional vector matching with the behavior labels extracted from the current voice interaction to calculate Jaccard similarity, dynamic time warping (DTW) distance and semantic embedding similarity, and calculate the behavior consistency score SactS_{act}; at the same time, count the time distribution characteristics of historical interactions (such as active time period, interaction cycle), compare them with the current time characteristics, and calculate the time stability score StimeS_{time}.

[0068] For example, the current user's interaction request "Can I listen to a bedtime story?" is judged to be "bedtime entertainment" semantic label. The system finds that the spatial voiceprint vector has appeared in the history of multiple similar requests between 19:00 and 21:00, including "Tell a fairy tale" and "I want to listen to Winnie the Pooh". The semantic matching score is 0.93, the emotional tone consistency is 0.89, and the behavioral consistency score is 0.91; the current time is 20:30, which belongs to the historical high-frequency interaction segment, and the time stability score is 0.95.

[0069] The behavioral consistency score, time stability score and identity confidence envelope are input into the preset credibility judgment model, and the user identity information and credibility level are output. The user identity information is graded and screened according to the credibility level to construct the identity trust area boundary.

[0070] By inputting the spatial voiceprint confidence envelope matching degree (such as the ratio of the distance from the voiceprint vector to the center to the radius), the behavioral consistency score and the temporal stability score as features into the credibility scoring model, the user's identity credibility level label is output; specifically, a credibility judgment model that integrates multi-source features is constructed, such as a weighted logistic regression model, a hierarchical decision tree or a lightweight neural network, and the input features are standardized into a unified vector. The model outputs the credibility level 𝐿={high credibility, medium credibility, low credibility, untrustworthy and identity ID; the system sets the boundary threshold according to the credibility level (such as high credibility requires voiceprint distance <0.75r, behavioral consistency >0.85, and temporal stability >0.8), constructs the identity credibility area boundary and screens out abnormal or suspicious identities.

[0071] For example, the normalized distance between the current user's voiceprint vector and the cluster center is 0.63 (i.e., 1.33 / 2.1), the behavioral consistency is 0.91, and the temporal stability is 0.95. The system inputs the features [0.63, 0.91, 0.95] to the three-layer neural network model and outputs the identity label "User_KidA" and the credibility level "high credibility"; if the value is [0.92, 0.61, 0.40], it is marked as "low credibility" or "pending verification".

[0072] The boundaries of the identity's trusted area are matched and verified with the preset permission configuration in the AI toy to determine whether the identity is within the authorized scope and output the verification result.

[0073] By querying the permission list configured in the AI toy system based on the current user's identity tag and its corresponding credibility level, the permission determination rules are executed and the final authentication result is generated. Specifically, the system maintains a mapping table or rule configuration table, which defines the permission ranges corresponding to different identity tags and credibility levels (such as "High Confidence - Child": voice games, story playback, and voice diary can be started; "Low Confidence - Unknown Identity": only background music is allowed to play or warnings are triggered). When the identity credibility boundary is determined to meet the authorization conditions, the system outputs "Verification Passed"; otherwise, it outputs "Verification Failed" and triggers a downgraded interaction or parental authorization request.

[0074] For example, when the current user is identified as "User_KidA" and the credibility is "high credibility", the system finds that his permissions allow him to access the voice diary module and the personalized voice question and answer module, and the current request instruction is "Open today's voice diary for me", so the verification result is "passed"; if the current user is judged to be "unknown identity" and the credibility is "low", his request "I want to listen to yesterday's recording" will be blocked due to exceeding the authority, and the system outputs "verification failed" and prompts "please get authorization from the guardian".

[0075] In step S400, the target speech signal is subjected to identity-weighted screening using the verification result to obtain semantic intent information and contextual state information, and the steps of constructing a behavior label based on the semantic intent information and constructing an emotion label based on the contextual state information specifically include: Semantic segment boundaries and temporal structures are extracted based on the target speech signal, a preliminary semantic candidate graph is constructed using the semantic segment boundaries, and a context path structure is constructed based on the temporal structure.

[0076] By performing speech recognition and semantic segmentation on the target speech signal, the start and end boundary information of the semantic units in the continuous speech is extracted, and the segments are divided into segments based on features such as syllable stress, pause duration, grammatical transitions, and keyword recognition to construct a semantic candidate graph. Specifically, automatic speech recognition (ASR) is first performed on the speech signal to extract the word transcription results and their timestamps. Then, through syntactic analysis and sentence boundary detection algorithms (such as the pause-attention mechanism based on bidirectional GRU), the boundaries of the semantic segments are identified and a set of semantic segment nodes are generated. Then, the syntactic dependency relationship is used to construct the edges between the segments to preliminarily form a semantic candidate graph. At the same time, the pronunciation order, dialogue turns and pause duration in the speech are extracted, and the arrangement of the semantic segments on the timeline and the front-to-back dependency order are established to form a context path structure for subsequent contextual reasoning.

[0077] For example, the speech "I want to hear a bedtime story. Can you tell me Peppa Pig?" is transcribed into seven phrases by the ASR module. Boundary detection identifies four semantic segments: "I want to hear | bedtime story | can you tell me | can I tell you Peppa Pig?", with time intervals of (0.0s–0.5s), (0.5s–1.1s), (1.1s–1.6s), and (1.6s–2.2s), respectively. "I want to hear" and "bedtime story" form a syntactic dependency pair, and "can you tell me" and "can I tell you Peppa Pig" form a question structure. The system constructs a semantic candidate graph with weighted edges and establishes a context path structure based on the duration of each segment and the speech gap in the speech: P1→P2→P3→P4.

[0078] The verification results are used to perform identity-weighted screening on the preliminary semantic candidate graph to obtain the semantic intent candidate vector, and context state information associated with the context path structure is generated based on the semantic intent candidate vector.

[0079] Through the obtained authentication results, including user ID and credibility level, the nodes in the semantic candidate graph are weighted and screened, giving the semantic nodes under the trusted identity a higher selection weight; specifically, the identity-weighted GAT is adopted to introduce user credibility 𝑤 in the node aggregation stage. uAs a normalization coefficient, the importance ranking of semantic nodes is adjusted, and the Top-K highly correlated semantic intent vectors are extracted from the candidate graph as semantic intent candidate vectors; then, using the context path structure, combined with the situational temporal position of the semantic intent candidate and the logical continuity of the upper and lower nodes, a multidimensional feature vector describing the current context state (such as the most recent turn-based interaction intention, repeated request frequency, content continuity label) is constructed as context state information.

[0080] For example, if the authentication result is "User_KidA" and the credibility is 0.91, the system assigns higher weights to the two segments "Bedtime Story" and "Telling Peppa Pig" in the semantic candidate graph, obtains the semantic intent candidate vector "request_bedtime_story" through graph attention layer aggregation, and recognizes that this semantic has content continuity with "I like Peppa Pig stories" in the previous interaction; combined with the fact that the turn-to-turn time interval is only 4 minutes, the context state information generated by the system is: ["Repeat Request" = 1, "Temporal Coherence" = 1, "Topic Consistency" = High].

[0081] The semantic intent candidate vector and user identity information are jointly embedded in a preset dual-channel attention gating network to generate semantic intent information. A context consistency map is constructed based on the context state information. The context consistency map is combined with the semantic segment position and turn order to perform contextual reasoning to obtain the contextual reasoning result. Among them, the semantic segment position refers to the start and end time index of the semantic unit in the target speech signal in the overall speech and its syntactic structure position, which is used to mark the relative order and hierarchical structure information of the semantic unit in the context; the turn order refers to the turn number information constructed according to the speech timing and response relationship during the alternating dialogue between the user and the AI toy, indicating the interactive turn position of the current speech in the multi-round dialogue structure, and is used to assist in judging the semantic intent and contextual logical consistency.

[0082] By constructing a dual-channel gated attention network, the semantic intent candidate vector and identity information are encoded and input into two attention channels in parallel, which are used for behavioral intention focusing and identity preference adjustment respectively, and finally fused to output accurate semantic intent information; specifically, the semantic candidate vector is embedded together with the identity ID and credibility level, mapped to a unified vector dimension through the shared word vector space, and then input into the attention network, using a gating mechanism to control the information flow between semantic and identity features; and the identity context memory module is used in the second channel to adjust the attention weight to retain preferred vocabulary.

[0083] At the same time, a context consistency graph is constructed, connecting the context dependencies, semantic categories, time positions and historical context states of the current semantic fragment into a heterogeneous graph structure, introducing the semantic fragment position (such as "starting sentence", "ending phrase", "subject structure") and turn sequence number (such as the third round of response) as node meta-information, performing context path traversal and reasoning on the graph neural network, and obtaining the context logical consistency judgment result.

[0084] For example, the system jointly embeds the semantic intent "request_bedtime_story" and the identity "User_KidA@0.91" into the input model. The first channel focuses on the phrases "before bed, story, telling", and the second channel adjusts the preference weight so that "Page" related content is given a high score; the position of the semantic fragment is judged as "sentence-end subject complement structure", which is the main sentence of the current request, and the turn order is the 4th turn (the previous turn is "I want to listen to a story"); the system constructs a graph and infers that the degree of coherence between the request and the historical dialogue is "high", and the context logic is "topic continuation + repeated request".

[0085] The intonation emotion distribution and emotional tendency factors are analyzed through contextual reasoning results, semantic segment boundaries and temporal structure.

[0086] Based on the results of contextual logical consistency judgment and the syntactic structure position of the semantic segment, the multimodal fusion emotion recognition model is implemented to extract emotional trends by combining the pitch, energy distribution and speaking rhythm of the target speech. Specifically, the F0 fundamental frequency contour, energy change gradient and speaking rate of the speech signal within the boundary of the semantic segment are extracted, and combined with contextual reasoning labels, such as the "repeat request + last sentence reinforcement" label, they are input into the emotion recognition module. The module uses the emotion attention network to guide the emotion tendency recognition model to infer the speaker's emotional state, and outputs emotional tendency factor scores such as "excitement, disappointment, happiness, irritability" and their intonation emotion distribution curves.

[0087] For example, the current semantic segment "Can you talk about Peppa Pig?" shows an upward F0 change (the average fundamental frequency rises from 210Hz to 260Hz), the energy is temporarily increased by 13%, and the speaking speed is accelerated by 8%. The system combines the context to infer the results of "repeated request" and "strong expectation" structures, and comprehensively judges the emotional tendency as "expectation". The emotional factor score is: [expectation=0.87, anxiety=0.14], forming a tone emotion distribution curve for subsequent fusion analysis.

[0088] The semantic intent information is jointly matched with the intonation emotion distribution to obtain the behavior label, and the contextual reasoning result is fused with the emotional tendency factor to generate the emotion label.

[0089] By constructing a behavioral emotion label decision system based on labeling rules and semantic-emotion fusion mapping tables, semantic intentions and emotion distributions are jointly matched to obtain specific behavior type identifications, and contextual reasoning structures and emotional factors are jointly used to generate stable and interpretable emotional state labels. Specifically, a mapping template from semantic intentions to behavior types is set, such as "before-bedtime request type" → "passive entertainment", "active request type" → "target behavior"; at the same time, "expectation + repeated request + rising tone" are jointly judged as "active request type behavior"; the emotion label is derived from "high context coherence + significant positive emotion factors" to "pleasure expectation type emotion".

[0090] For example, the current semantic intent is "request_bedtime_story," and the tone expresses happiness and a request. After matching the rule library, the behavior label "active request - bedtime entertainment" is generated. The corresponding emotional tendency factor combination is "expectation = 0.87, pleasure = 0.64," and the contextual logic judgment is "structural coherence." The final emotional label "pleasure expectation" is generated to guide the AI toy to give a positive and enthusiastic response tone and content selection in the next round of dialogue.

[0091] The steps of analyzing the intonation emotion distribution and emotional tendency factors through contextual reasoning results, semantic segment boundaries, and temporal structure include: Semantic temporal alignment data is constructed according to the semantic segment boundaries and temporal structure, and the intonation change envelope and silence interval distribution of the speech waveform are extracted based on the semantic temporal alignment data.

[0092] By aligning the start and end timestamps of the semantic segments with the entire speech signal, a mapping relationship between semantic content and acoustic features is established, and then the analysis window for intonation changes and silence gaps is delineated. Specifically, based on the semantic segment boundaries and context path structure extracted in the previous step, the time interval of each segment in the original speech waveform is calibrated, and the speech signal within this interval is processed by short-time Fourier transform (STFT) and fundamental frequency extraction to obtain a fundamental frequency (F0) trajectory curve reflecting the fluctuations of intonation. The intonation change envelope is further calculated using a sliding window algorithm. At the same time, the silence gaps (zero energy frames) between speech segments are identified, and the duration, position, and density of the silence distribution are statistically analyzed to construct a "semantic segment-time-intonation / silence" ternary alignment structure.

[0093] For example, the target speech "Can you tell me a story... tell me an interesting one" is divided into two semantic segments with boundary times of (0.0s–1.8s) and (2.0s–3.5s), respectively. The temporal structure is a continuous request type. The system extracts the envelope of the fundamental frequency dropping from 230Hz to 190Hz within the interval (0.0s–1.8s) and identifies two silent intervals of 0.3s and 0.5s, respectively. The alignment data structure is constructed as follows: Semantic segment P1 → F0 envelope: [230Hz→210Hz→190Hz].

[0094] Silence interval: position = [1.2s, 1.9s], length = [0.3s, 0.5s].

[0095] The intonation variation envelope is used to extract the tone variation features, and the pause features are extracted based on the distribution of silent intervals.

[0096] By modeling the local change rate and overall trend of the intonation change envelope, the tone change characteristics representing the intensity and change direction of emotional expression are extracted; at the same time, the distribution of silent gaps is statistically analyzed to explore potential pause characteristics such as speech rhythm, hesitation behavior or anxiety tendency; specifically, the first-order derivative and second-order derivative of the F0 envelope are first calculated as indicators of the rate of increase and degree of fluctuation of intonation, and the feature vectors such as "tone excitement" and "emotional fluctuation value" are generated by combining the average F0 and standard deviation; secondly, the number, average length, maximum length and distribution position of silent gaps are extracted to determine whether they are in the middle of a sentence, at the end of a sentence or between fragments, so as to identify emotion-related behavior patterns such as "inter-sentence contemplation", "hesitation at the end of a sentence" and "interruption waiting".

[0097] For example, the intonation envelope shows that the target speech F0 rises rapidly during the period of "tell me an interesting story" (jumping from 200Hz to 270Hz, with an average derivative of +45Hz / s), corresponding to high "emotional excitement"; silence distribution analysis shows that there is a 0.6-second silence between "story?" and "tell me an interesting story", which is located between sentences and is inferred to be a "contemplative pause"; finally, the tone change feature vector extracted is as follows: [rise rate = 45Hz / s, standard deviation = 30Hz, emotional fluctuation value = 0.76], and the pause characteristics are: [total silence = 2, maximum length = 0.6s, type = inter-sentence contemplation].

[0098] The tone change features and pause features are fused through the contextual reasoning results to extract the intonation emotion distribution and emotional tendency factors.

[0099] By taking the results of contextual reasoning as the guiding conditions for emotional tendency, guiding the fusion strategy of tone and pause features in the emotional model, and using the multi-dimensional attention mechanism to assign differentiated weights to different features, the intonation emotional distribution and emotional tendency factors are generated; specifically, based on indicators such as "interaction intention strength", "context logical coherence", and "speech turn weight" in the aforementioned contextual reasoning results, an emotional guidance vector is constructed; the guidance vector is input together with the tone change features and pause features into a multimodal emotional reasoning network (such as the emotion-aware Bi-LSTM+attention pooling structure), and the emotional state labels of different segments on the timeline (such as "pleasure, happiness, anxiety, doubt", etc.) and their tendency factor scores (ranging from 0 to 1) are output to form a complete intonation emotional distribution curve.

[0100] For example, if the contextual reasoning result is "main sentence + continuous request + historical repetition," the system assigns a higher weight to the tone change feature (significant rise) and a supplementary reference weight to the pause feature (moderate). After fusion analysis, the system determines that the main trend of the tone is "active request" and assigns the following emotional tendency factors: Pleasure = 0.72, Anticipation = 0.81, Anxiety = 0.15. A speech intonation and emotion distribution map is also generated, marking the points where different emotions dominate on the timeline. For example, the interval [1.0s–1.8s] is "anticipation-dominated," and the interval [2.1s–3.5s] is "pleasure-dominated." This serves as a reference for the subsequent response strategy generation module.

[0101] In step S500, the steps of generating corresponding multimodal response instructions based on user identity information, behavior tags, and emotion tags, and inputting the multimodal response instructions into the control module of the AI toy specifically include: The behavior label is input into the preset intention classification library for matching to obtain the corresponding candidate semantic unit. The emotion label is input into the preset expression and action library for matching to obtain the corresponding visual feedback resource. Based on the visual feedback resource, the candidate image elements and action trajectories are extracted.

[0102] By inputting behavior labels and emotion labels into a predefined response mapping library, an association query between semantic content and multimodal feedback resources is performed to construct a set of candidate responses for language and visual / motion channels. Specifically, the behavior labels are input into the intent classification library (Intent Repository), and the corresponding semantic response units are selected through exact matching or semantic embedding distance retrieval. The semantic units contain preset answer text templates or semantic structures (such as greetings, replies, and guidance). The emotion labels are then input into the expression and motion library (Emotion-Motion Mapping Set), and the expression image elements (such as smiles, blinks, etc.) and motion trajectories (such as nodding, waving, etc.) that best match the emotion values are matched. The three-dimensional key point paths, expression displacement parameters, and duration control factors are extracted as candidate images and motion feedback elements.

[0103] For example, the behavior label is "active request - before-bed entertainment", and the emotion label is "joyful anticipation"; the system selects the semantic unit from the intent classification library: "Of course, now I will tell you an interesting story"; at the same time, the system retrieves the visual feedback resources from the expression and action library as two matching units: "open mouth smiling, eyes raised" and "upper body leaning forward, waving hands to welcome", and extracts the mouth opening angle α=15° and eyelid lifting angle β=10° respectively, and the action trajectory is "right arm pulled forward 12cm along the Z axis, left arm swung horizontally 8cm", which lasts about 2.5 seconds.

[0104] The candidate semantic units, candidate image elements and action trajectories are consistently fused to generate a multimodal candidate response instruction set.

[0105] By constructing a multimodal consistency fusion model, semantic responses, visual elements, and motion trajectories are jointly adjusted and encoded in terms of timing, semantic tone, and emotion-driven parameters, generating a multimodal candidate response instruction set with consistent structure and unified emotion. Specifically, a multimodal consistency composer (MCC) network is used to align the three types of elements, with semantic units serving as the dominant axis and their textual emotion embeddings serving as the emotional benchmark to drive the movement amplitude of visual expressions and the execution speed of motion trajectories. The fusion model uses behavior-emotion labels as joint conditional input and outputs a JSON-structured response instruction set, including sub-instruction blocks such as speech synthesis instructions (text + emotion tone label), expression rendering parameters (image path, muscle control vector), and motion control trajectory (joint coordinate sequence, timestamp control amount).

[0106] After a user says, "Can you tell me a more interesting story?" the system identifies the behavior label as "active request - fun-enhancing" and the emotion label as "high anticipation and curiosity." The candidate semantic unit matched from the intent classification library is "Of course, I've prepared a wonderful story. Would you like to hear it?" From the expression and action library, the image elements matching "anticipation and joy" are "a smile with upturned corners of the mouth" and "an expression of surprise with eyes wide open." The action trajectory matching this emotion is "head tilted forward 10 degrees, hands outstretched." In the multimodal consistency fusion model, the system first uses the emotion embedding of the semantic unit as a guide, identifying the intent as "positive guidance" and expressing a "confident and friendly" tone. Based on this, the model fine-tunes the image elements, slightly increasing the smile and slightly widening the angle of the eyes to convey the emotional effect of "high confidence and willingness to interact." At the same time, to maintain consistency between movement and tone, the system slightly slows down the rhythm of the original "open hands" movement to align it with the end of the sentence "I've prepared a wonderful story" in the voice intonation, achieving synchronization in semantics and movement time. The multimodal response constructed from the fusion results includes: the voice content has a smiling tone, a gentle rhythm, and a slightly rising tone at the end; the visual presentation is a slight open-mouth smile and wide-open eyes; the body movements are synchronized with the voice playback process, first the head is slightly tilted forward to show attention, and then the arms are naturally opened forward to express a gesture of welcome and expectation. The overall response maintains a high degree of consistency in semantic expression, visual emotion, and movement rhythm, enhancing the emotional appeal and immersion of the AI toy in the process of interacting with children.

[0107] The current interaction permission level is obtained according to the user identity information, and the multimodal candidate response instruction set is screened and prioritized based on the interaction permission level to obtain the multimodal response instruction, which is then input into the control module of the AI toy.

[0108] By retrieving the permission level and preference profile from the user identity information, the complexity, security and response scale of the response content are set according to the permission level, and the candidate response instruction sets are screened and sorted accordingly; specifically, the system maintains a user identity permission mapping table, divides user identities into levels such as "child user, advanced user, restricted user", and defines parameters such as the maximum speaking speed, maximum movement amplitude, and allowed semantic types for each level; during the response screening process, if the current instruction contains content that exceeds the permission (such as searching the Internet, jumping actions, etc.), it will be filtered or replaced with a low-level response; the execution priority of the retained instructions is scored, and sorted according to the user's historical preference records and label weights, and finally the multi-modal response instruction set with the highest ranking is output; the response instructions are sent to the AI toy control module through the ROS protocol or BLE serial interface, which respectively drives the speech synthesis chip, expression motor module and multi-degree-of-freedom servo execution module to complete the synchronous response.

[0109] For example, the current user identity is "User_KidA" and the permission level is "Child Restricted", which prohibits the AI toy from making jumping movements and using anthropomorphic intonations; the system removes the "wink-smirk" visual element containing anthropomorphic expressions and the "jumping_intro" motion trajectory from the candidate response instruction set, retains the semantically complete welcome instructions, and enhances the clarity of the voice and intonation according to user preferences (the speaking speed is adjusted from 1.1 to 1.0); the final selected response instruction is transmitted to the AI toy's underlying control system through the serial port, and the speech synthesis module is executed in sequence to play the voice, the expression motor controls the corners of the mouth to rise, and the servo module drives the arm in sequence to perform the welcome gesture to complete the coordinated response.

[0110] In this embodiment, by constructing a phased recognition process centered on sound source localization, background modeling, voiceprint clustering, and semantic intent recognition, combined with spatial voiceprint modeling and contextual consistency analysis, the system achieves accurate identification and trustworthy judgment of child user identities, and further combines behavioral labels and emotion labels to complete multimodal response generation. Specifically, the system first collects voice signals through a multi-microphone array, performs time-frequency analysis and phase decoding, extracts sound field estimation parameters and background noise characteristics, and accurately obtains the target sound source direction and environmental information. Subsequently, spatial voiceprint features are extracted through main sound source enhancement and adaptive denoising, and identity recognition and similarity analysis are performed by combining spectral clustering algorithms and density estimation techniques. Furthermore, an identity trust region is constructed based on historical user interaction records and identity confidence envelopes to achieve dynamic verification of the current user's identity. In the semantic intent analysis stage, the system integrates semantic segment structure, context path, and dual-channel attention mechanism to complete semantic intent and contextual state extraction. Finally, the system combines intonation emotional factors with contextual reasoning results to generate behavioral labels and emotion labels, respectively, and outputs consistent responses at the speech, expression, and action levels through a multimodal fusion network. This solution improves the personalized experience, recognition credibility and emotional adaptation capabilities during the interaction between AI toys and child users.

[0111] Example 3: like Figure 2 As shown, the present application also provides an AI toy voiceprint recognition interaction device 10, including an acquisition module 11, an extraction module 12, a verification module 13, a decoding module 14 and a generation module 15.

[0112] The acquisition module 11 is mainly used to acquire the target speech signal to be recognized, extract the original audio data set corresponding to the target speech signal, and generate sound field estimation parameters and background noise characteristics according to the original audio data set.

[0113] The extraction module 12 is mainly used to extract the voiceprint features of the target speech signal through the sound field estimation parameters and background noise characteristics, obtain the voiceprint feature vector, perform feature clustering on the voiceprint feature vector and the preset children's voiceprint vector set, and obtain the identity clustering result and similarity index.

[0114] The verification module 13 is mainly used to input the identity clustering results and similarity indicators into a preset credibility judgment model to obtain user identity information and credibility level, verify the user identity information based on the credibility level, and obtain a verification result.

[0115] The decoding module 14 is mainly used to use the verification results to perform identity-weighted screening on the target speech signal, obtain semantic intent information and contextual state information, construct behavior labels based on the semantic intent information, and construct emotion labels based on the contextual state information.

[0116] The generation module 15 is mainly used to generate corresponding multimodal response instructions based on user identity information, behavior tags and emotion tags, and input the multimodal response instructions into the control module of the AI toy.

[0117] In this embodiment, the acquisition module 11 collects the target speech signal of the user in the environment in real time. Combining sound field estimation parameters and background noise characteristics, it accurately locates and extracts effective voiceprint feature vectors, overcoming noise interference in complex acoustic environments. Specifically, the extraction module 12 uses a multi-channel sound field estimation algorithm integrated with spatial filtering technology to improve the signal-to-noise ratio of the target speech and extract high-dimensional voiceprint feature vectors. Subsequently, by density clustering the voiceprint vectors with a preset set of children's voiceprint vectors, the cluster center with the highest identity similarity is identified and the identity clustering results and corresponding similarity indicators are output. The verification module 13, based on a credibility determination model fused with multi-dimensional features, performs a weighted analysis of the identity clustering results and dynamically adjusts the credibility level based on historical interaction traces to complete user identity verification and trust assessment. The decoding module 14 uses the verification results to perform identity-weighted screening on the original speech signal, extracting precise semantic intent information and contextual state information, and constructing fine-grained behavioral and emotional labels to ensure the accuracy of subsequent semantic understanding and emotional adaptation. The generation module 15 combines the user identity, behavior tags and emotion tags, calls the preset multimodal response instruction library, generates structured speech synthesis instructions, expression control parameters and motion trajectory instructions, and realizes the adaptive adjustment of the human-computer interaction behavior of the AI toy.

[0118] It should be noted that those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described device and each module can refer to the corresponding processes in the aforementioned embodiment 1 and will not be repeated here.

[0119] Example 4: like Figure 3 As shown, the present application also provides an electronic device 20, including a memory 21 and a processor 22, the memory 21 stores a computer program that can be run on the processor 22, and when the processor 22 executes the computer program, the AI toy voiceprint recognition interaction method of Example 1 is implemented.

[0120] In this embodiment, the AI toy voiceprint recognition and interaction method described in Example 1 is implemented through a computer program pre-installed in memory 21. Combined with the high-speed computing power of processor 22, this method enables rapid processing and real-time feedback of complex voiceprint data. Specifically, when processor 22 executes the program, it first calls acquisition module 11 to collect the target voice signal and generate sound field estimation parameters. It then calls extraction module 12 to extract voiceprint features and cluster identities, and uses verification module 13 to complete identity credibility assessment. Subsequently, decoding module 14 performs weighted identity screening based on the verification results, constructing behavior and emotion labels. Finally, generation module 15 invokes a multimodal response instruction set based on the label information, implementing a coordinated response for speech synthesis, expression rendering, and motion control of the AI toy. Through the execution of this computer program, electronic device 20 can continuously track the user's identity and emotional changes during multiple rounds of interaction, achieving efficient, accurate, and personalized voiceprint recognition and interaction capabilities, improving user experience and system security.

[0121] The structures, proportions, sizes, etc. depicted in the drawings of this specification are only used to match the contents disclosed in the specification so as to facilitate understanding and reading by persons familiar with this technology. They are not intended to limit the conditions under which the present invention can be implemented and therefore have no substantive technical significance. Any structural modifications, changes in proportional relationships, or adjustments in size should still fall within the scope of the technical contents disclosed in the present invention without affecting the effects and objectives that can be achieved by the present invention.

[0122] As described above, the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that the technical solutions described in the above embodiments can still be modified, or some of the technical features thereof can be replaced by equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. An AI toy voiceprint recognition interaction method, characterized in that: include: Acquire a target speech signal to be recognized, extract an original audio data set corresponding to the target speech signal, and generate sound field estimation parameters and background noise characteristics based on the original audio data set; Extracting voiceprint features from the target speech signal using the sound field estimation parameters and the background noise characteristics to obtain a voiceprint feature vector, performing feature clustering on the voiceprint feature vector and a preset children's voiceprint vector set to obtain an identity clustering result and a similarity index; Inputting the identity clustering result and the similarity index into a preset credibility determination model to obtain user identity information and a credibility level, and verifying the user identity information based on the credibility level to obtain a verification result; Performing identity-weighted screening on the target speech signal using the verification result to obtain semantic intent information and contextual state information, constructing a behavior label based on the semantic intent information, and constructing an emotion label based on the contextual state information; A corresponding multimodal response instruction is generated based on the user identity information, behavior tag, and emotion tag, and the multimodal response instruction is input into a control module of the AI toy.

2. The AI toy voiceprint recognition interaction method according to claim 1, characterized in that: The steps of obtaining a target speech signal to be recognized, extracting an original audio data set corresponding to the target speech signal, and generating sound field estimation parameters and background noise characteristics based on the original audio data set include: Using multiple microphone arrays to synchronously collect target speech signals to obtain raw audio data sets of different spatial channels, performing time-frequency analysis and beamforming on the raw audio data sets to extract frequency domain power spectrum distribution and time domain phase shift features; Calculating a sound source direction estimation value and an echo path length based on the frequency domain power spectrum distribution and the time domain phase offset characteristics, and constructing a sound field estimation parameter using the sound source direction estimation value and the echo path length; Background energy statistics and non-speech component detection are performed on the original audio data set, a spectral envelope and stability parameters of the background noise are extracted, and a background noise feature vector is generated based on the spectral envelope and the stability parameters.

3. The AI toy voiceprint recognition interaction method according to claim 2, characterized in that: The step of extracting voiceprint features from the target speech signal using the sound field estimation parameters and the background noise features to obtain a voiceprint feature vector, performing feature clustering on the voiceprint feature vector and a preset children's voiceprint vector set to obtain an identity clustering result and a similarity index includes: performing spatial filtering on the sound field estimation parameters to obtain filtered sound field estimation parameters, and performing main sound source enhancement on the target speech signal using the filtered sound field estimation parameters to obtain an enhanced target speech signal; Adaptively removing noise from the enhanced target speech signal using the background noise feature to obtain a clean speech signal, performing short-time Fourier transform and perceptual spectrum mapping on the clean speech signal to extract multi-scale voiceprint candidate features; The multi-scale voiceprint candidate features are combined with the sound source direction information to construct a spatial voiceprint feature vector; wherein the sound source direction information refers to the three-dimensional positioning information of the target sound source constructed based on the sound source direction estimate value and the spatial distribution model of the microphone array; The spatial voiceprint feature vector and a preset children's voiceprint vector set are subjected to spectral clustering and local density estimation to obtain cluster centers and a similarity matrix, identity clustering results are generated based on the cluster centers, and a similarity index is generated using the similarity matrix.

4. The AI toy voiceprint recognition interaction method according to claim 3, characterized in that: The step of inputting the identity clustering result and the similarity index into a preset credibility determination model to obtain user identity information and a credibility level, and verifying the user identity information based on the credibility level to obtain a verification result includes: Extracting the feature center and distribution radius of the category according to the identity clustering result, and generating an identity confidence envelope using the feature center and the distribution radius; Correlating and matching the similarity index with historical user interaction records to obtain a matching result, and calculating a behavior consistency score and a temporal stability score based on the matching result; wherein the historical user interaction record refers to a historical set of user interaction sessions recorded by the AI toy for the same spatial voiceprint feature vector over multiple time periods, including behavior trajectory data, semantic response labels, tone and emotion patterns, command execution feedback, and voice environment information; Inputting the behavior consistency score, the time stability score, and the identity confidence envelope into a preset credibility determination model, outputting user identity information and a credibility level, and performing hierarchical screening on the user identity information according to the credibility level to construct an identity credibility zone boundary; The identity trusted area boundary is matched and verified with the preset permission configuration in the AI toy, and the verification result is output.

5. The AI toy voiceprint recognition interaction method according to claim 1, characterized in that: The steps of performing identity-weighted screening on the target speech signal using the verification result to obtain semantic intent information and contextual state information, constructing a behavior label based on the semantic intent information, and constructing an emotion label based on the contextual state information include: Extracting semantic segment boundaries and temporal structures based on the target speech signal, constructing a preliminary semantic candidate graph using the semantic segment boundaries, and constructing a context path structure based on the temporal structure; Performing identity-weighted screening on the preliminary semantic candidate graph using the verification result to obtain a semantic intent candidate vector, and generating context state information associated with a context path structure based on the semantic intent candidate vector; The semantic intent candidate vector and the user identity information are jointly embedded in a preset dual-channel attention gating network to generate semantic intent information, a context consistency map is constructed based on the context state information, and the context consistency map is combined with the semantic segment position and turn order to perform contextual reasoning to obtain a contextual reasoning result; wherein, the semantic segment position refers to the start and end time index of the semantic unit in the overall speech in the target speech signal and the position of its syntactic structure; the turn order refers to the turn number information constructed according to the relationship between the speech timing and the response during the alternating dialogue between the user and the AI toy; Analyze the intonation emotion distribution and emotional tendency factors through the contextual reasoning results, the semantic segment boundaries and the temporal structure; The semantic intention information is jointly matched with the intonation emotion distribution to obtain a behavior label, and the contextual reasoning result is fused with the emotion tendency factor to generate an emotion label.

6. The AI toy voiceprint recognition interaction method according to claim 5, characterized in that: The step of analyzing the intonation emotion distribution and the emotional tendency factor through the contextual reasoning result, the semantic segment boundary and the temporal structure includes: constructing semantic temporal alignment data according to the semantic segment boundaries and the temporal structure, and extracting the intonation variation envelope and silence interval distribution of the speech waveform based on the semantic temporal alignment data; Extracting tone change features using the intonation change envelope, and extracting pause features based on the silent gap distribution; The tone change feature and the pause feature are fused with the contextual reasoning result to extract the tone emotion distribution and the emotional tendency factor.

7. The AI toy voiceprint recognition interaction method according to claim 1, characterized in that: The step of generating corresponding multimodal response instructions based on the user identity information, behavior tags, and emotion tags, and inputting the multimodal response instructions into the control module of the AI toy includes: Input the behavior label into a preset intention classification library for matching to obtain corresponding candidate semantic units; input the emotion label into a preset expression and action library for matching to obtain corresponding visual feedback resources; and extract candidate image elements and action trajectories based on the visual feedback resources; Consistently fusing the candidate semantic units, the candidate image elements, and the action trajectory to generate a multimodal candidate response instruction set; The current interaction permission level is obtained according to the user identity information, the multimodal candidate response instruction set is screened and prioritized based on the interaction permission level to obtain a multimodal response instruction, and the multimodal response instruction is input into a control module of the AI toy.

8. An AI toy voiceprint recognition interactive device, characterized in that: include: An acquisition module is used to acquire a target speech signal to be recognized, extract an original audio data set corresponding to the target speech signal, and generate sound field estimation parameters and background noise characteristics based on the original audio data set; an extraction module, configured to extract voiceprint features from the target speech signal using the sound field estimation parameters and the background noise characteristics to obtain a voiceprint feature vector, perform feature clustering on the voiceprint feature vector and a preset children's voiceprint vector set to obtain an identity clustering result and a similarity index; a verification module, configured to input the identity clustering result and the similarity index into a preset credibility determination model to obtain user identity information and a credibility level, and verify the user identity information based on the credibility level to obtain a verification result; A decoding module, configured to perform identity-weighted screening on the target speech signal using the verification result to obtain semantic intent information and contextual state information, construct a behavior label based on the semantic intent information, and construct an emotion label based on the contextual state information; A generation module is used to generate corresponding multimodal response instructions based on the user identity information, behavior tags and emotion tags, and input the multimodal response instructions into the control module of the AI toy.

9. An electronic device, characterized in that: The device comprises a memory and a processor, wherein the memory stores a computer program that can be run on the processor, and when the processor executes the computer program, the AI toy voiceprint recognition interaction method according to any one of claims 1 to 7 is implemented.

Citation Information

Cited By

  • Learning method and system based on developed child toy

    CN120656492A

  • A learning method and system based on developmental toys for children

    CN120656492B

  • Auxiliary turning-over equipment for nursing old people and voice control method thereof

    CN120808768A

  • User behavior data processing method and device, electronic equipment and storage medium

    CN120853548A

  • Personnel calling behavior identification and positioning system based on audio analysis

    CN121171211A