Multi-modal Adaptive Child Voice Interaction Method for Educational Service Robots

Through multimodal adaptive technology, the accuracy of children's speech recognition in noisy and multi-speaking environments is solved, and the efficient and stable voice interaction of educational robots in complex environments is achieved, adapting to the pronunciation characteristics of different children and optimizing in real time.

CN120164478BActive Publication Date: 2025-07-22北京爱宾果科技有限公司

Patent Information

Application Number
CN202510642427.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-19
Publication Date
2025-07-22
Estimated Expiration
2045-05-19

AI Technical Summary

Technical Problem

The existing voice recognition system is difficult to accurately recognize children's voice in noisy and multi-speaking environments, especially in educational and family scenarios, with misunderstandings and missed recognition problems.

Method used

Using technical means of dynamic acquisition and beamforming of multi-array microphones, adaptive noise suppression, multi-speaker separation and speaker identification, children's exclusive model training and online fine-tuning, and context semantic correction and multi-modal fusion, we use end-to-end separation of networks and deep speaker embedding models, and combine cameras and environmental sensors for multi-modal alignment and correction.

Benefits of technology

It significantly improves the focus ability and recognition accuracy of children's voice in noisy environments, reduces misjudgment and confusion, achieves a stable and natural voice interaction effect, and adapts to the pronunciation characteristics of different children through self-learning optimization models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120164478B_ABST
    Figure CN120164478B_ABST
Patent Text Reader

Abstract

The present invention discloses a multi-modal adaptive child voice interaction method for educational service robots, which relates to the technical field of speech signal processing. Through technical means such as dynamic acquisition and beamforming by multi-array microphones, adaptive noise suppression, separation and identification of multiple speakers, training and online fine-tuning of child-exclusive models, as well as context semantic correction and multi-modal fusion, the target speech focusing and recognition accuracy are significantly improved; through the end-to-end network to distinguish multiple-person speech tracks and the model optimization combining the child's vocal range characteristics, the rapid resolution of multi-source speech scenarios is realized; at the same time, by using real-time semantic correction and interaction feedback, the correct vocabulary and pronunciation information are traced back to the speaker database and the child voice model, and supplemented by the multi-modal collaboration of cameras and positioning sensors, the beam is quickly adjusted and interference is suppressed when the child moves randomly and the environmental noise suddenly increases, providing a natural and smooth voice interaction experience in classroom and home scenarios.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of voice signal processing, and particularly to a multi-modal adaptive child voice interaction method for an educational service robot. Background Technique

[0002] In today's fields of education and home companionship, more and more scenarios require voice interactive assisted teaching devices or robots to provide real-time answers to questions and interactive guidance.

[0003] Taking classroom teaching as an example, teachers often hope to use robots to assist in answering students' questions or organizing group discussions. In the home scenario, parents can also let the robot tutor their children to complete homework and language training. However, such scenarios often have environmental interference and situations where multiple people speak at the same time, such as random student speeches in the classroom, parallel communication between children and parents at home, etc. Coupled with the characteristics of children's pronunciation itself, such as high-pitched intonation, fast speech rate, and unclear articulation, traditional single-channel or adult model-dominated voice processing systems frequently misinterpret and miss-recognize children's instructions or answers, thus seriously affecting the fluency of human-computer interaction and teaching efficiency.

[0004] After retrieval, an electronic device for realizing voice signal recognition is disclosed in the application publication number CN110364166A, including: a microphone array for collecting audio signals; a plurality of processors connected to the microphone array; each processor is paired with a beamformer and a voice recognition module. Among them, each beamformer is used to perform beamforming processing on the audio signal in a set of multiple different target directions respectively to obtain corresponding multiple beam signals; each voice recognition module is used to perform voice recognition on the beam signals output by the paired beamformers respectively to obtain the voice recognition results of each beam signal; one of the processors is configured with a processing module for determining the voice recognition result of the audio signal according to the voice recognition results of each beam signal. By performing beamforming processing in different target directions, at least one target direction is close to the voice signal generation direction, which can improve the accuracy of intelligent voice recognition.

[0005] Combined with the above application scenarios and the content in the prior art:

[0006] Most of the mainstream speech recognition solutions on the market today use adult speech as the core training data and are usually only suitable for quiet environments or relatively single-speaker conditions. Once faced with a real educational scenario with complex noise and multiple speakers speaking in parallel, the system will encounter obvious bottlenecks in such aspects as target speech focusing, sound source separation, and adaptation to children's voice characteristics. On the one hand, the noisy background and multiple people speaking simultaneously result in a low signal-to-noise ratio, making it impossible for the microphone to stably capture the speech of the target child. On the other hand, the special vocal range and oral characteristics of children are significantly different from those of adult models, making it difficult for existing technologies to obtain high-accuracy text decoding results after separating the correct sound track. In addition, the lack of an automatic correction mechanism for teaching content or dialogue context also prevents the system from quickly correcting recognition deviations and keeping up with the spontaneous dialogue needs that children may have at any time.

[0007] Generally speaking, how to achieve accurate, stable, and efficient speech recognition and interaction in an environment with noise, multiple sources, and dominant children's speech becomes the technical problem that this invention focuses on solving. Summary of the Invention

[0008] (I) Technical Problems to be Solved

[0009] In view of the deficiencies of the prior art, this invention provides a multi-modal adaptive children's speech interaction method for educational service robots. Through technical means such as dynamic acquisition and beamforming of multi-array microphones, adaptive noise suppression, separation of multiple speakers and speaker identification, training and online fine-tuning of children's exclusive models, and context semantic correction and multi-modal fusion, the focusing of target speech and the recognition accuracy are greatly improved, thus solving the technical problems recorded in the background art.

[0010] (II) Technical Solutions

[0011] To achieve the above objectives, this invention is realized through the following technical solutions:

[0012] A multi-modal adaptive children's speech interaction method for educational service robots includes, when it is detected that the environmental noise exceeds the threshold value, calling the beamforming and adaptive algorithms of multi-array microphones to apply directional gain and filtering to the multi-channel audio matrix, suppressing the main interference sources and then outputting a preprocessed audio signal;

[0013] After receiving the preprocessed audio signal, using an end-to-end separation network and a speaker embedding model to disassemble the overlapping speech, extracting feature vectors for each separated sound track, and finally outputting multiple independent sound tracks and individual identity markers;

[0014] When the individual identity marker indicates that a certain sound track belongs to a child individual, calling the children's speech recognition model and language model constructed by transfer learning to perform recognition and decoding on the corresponding track in the multiple independent sound tracks, and real-time fine-tuning the parameters based on the cumulative samples to generate a preliminary recognition text;

[0015] If it is detected that the initially recognized text conflicts with the teaching theme or common spoken language, call the context rearrangement function and the difference metric to perform semantic error correction, and inject the user's correction information into the speaker database and the children's speech recognition model to form a closed-loop self-learning;

[0016] When the data of the camera and the environmental sensor can assist in the attribution of the audio track, perform multi-modal alignment mapping and call the cross-modal verification function to check the pose and position differences, and dynamically correct the beam weights and the assignment of individual identity markers.

[0017] Furthermore, in the microphone array channels arranged on the fuselage, synchronously collect audio signals to form a multi-channel audio matrix:

[0018] Introduce an energy detection function based on wavelet transform to perform segmented energy estimation on the time-domain signals of the microphone channels at different time scales, capture the relative distribution characteristics of the environmental noise and the speech mixed signal in different frequency bands; and calculate the time difference between different microphone channels based on the cross-correlation method to obtain an approximate estimate of the spatial orientation of the target sound.

[0019] Furthermore, based on the multi-channel audio matrix and the approximate estimate of the spatial orientation, dynamically adjust the channel weight coefficients through adaptive beamforming to enhance the speech in the target direction, and attenuate the time-frequency coefficients with amplitudes lower than the threshold in the wavelet domain to suppress background noise, and obtain the preprocessed audio signal.

[0020] Furthermore, use an end-to-end speech separation model based on a time-domain convolutional network to perform short-time segmentation on the preprocessed audio signal, and complete multi-speaker speech separation through a learnable encoder and decoder;

[0021] And use an objective function with PIT loss to automatically match the optimal correspondence between the network separation output and the true target, and obtain multiple independent audio tracks.

[0022] Furthermore, adopt a deep speaker embedding model to convert the time-domain or frequency-domain segment signals into fixed-dimensional vectors to represent the speaker characteristics of the current audio track;

[0023] By making a similarity match with the registered embedding vectors in the pre-collected and constructed children's voice feature library, identify the children's identity match, and associate each audio track with the corresponding individual identity marker; if the similarity metric does not exceed the preset threshold, it can be temporarily marked as a new speaker.

[0024] Furthermore, when the multiple independent audio tracks and the individual identity markers output show that a certain audio track comes from a child individual, call the children's speech recognition model obtained by transfer learning from a pre-trained large-scale adult speech model;

[0025] Perform speech decoding on each separated track marked as a children's track in combination with the corresponding language model to obtain preliminary recognized text.

[0026] Furthermore, when the tracks of a certain child individual accumulate to a certain number, perform real-time or offline update on the children's speech recognition model based on the adaptive optimization objective function; perform local or global iterative update on the model parameters, store the updated new parameter model in the model configuration corresponding to the child, and output a temporary recognized text set.

[0027] Furthermore, in combination with the current conversation context, course theme, and common children's spoken language expression library, adopt a context-based re-ranking function to perform semantic re-ranking and correction on the recognized text; through comprehensive evaluation of all tokens, perform re-ranking, replacement, or insertion error correction operations to generate a corrected text result.

[0028] Furthermore, add the corresponding audio and correct text marks in the user correction feedback dataset to the pre-constructed speaker database, and perform offline or real-time fine-tuning on the children's speech recognition model after updating the children's exclusive training set;

[0029] Cumulatively manage the collected user correction data, trigger batch update after the long-term memory coefficient reaches the preset threshold, and output the final recognized text after correction and feedback update.

[0030] Furthermore, establish a cross-modal spatio-temporal alignment mechanism at the acquisition level to synchronize visual information, environmental sensor readings, and audio signals, and use a deviation amplification judgment function to evaluate the consistency of vision and positioning coordinates, and screen out or correct unreliable readings to obtain multi-modal alignment data.

[0031] Furthermore, dynamically adjust the beamforming weight vector of the robot microphone array according to the multi-modal alignment data, and convert it into the time-domain weight of each channel through inverse STFT; compare the noise source direction, and additionally introduce an attenuation factor in the weight calculation to suppress the energy in the interference direction; use the updated weight to weight and synthesize the multi-channel original signal to obtain the updated preprocessed audio.

[0032] Furthermore, perform cross-modal verification on the recognized text through visual mouth closure detection, head orientation, and noise emergencies, where: measure the overall consistency of the speaker activities corresponding to the track with the visual and positioning data in terms of posture and position through the cross-modal verification score. If the cross-modal verification score is lower than expected, assign a higher correction priority to the suspicious sentence segment or reassign it to the correct speaker track, and synchronously update the speaker mark.

[0033] (III) Beneficial effects

[0034] The present invention provides a multi-modal adaptive child voice interaction method for educational service robots, which has the following beneficial effects:

[0035] In a noisy and multi-speaker environment, a complete child voice interaction system is constructed through five steps, significantly improving the robot's ability to focus on the voice of the target speaker and the recognition accuracy;

[0036] Output a preprocessed audio signal through a multi-array microphone and adaptive beamforming , reducing environmental interference for subsequent recognition; Subsequently, use an end-to-end separation network and a speaker embedding algorithm to separate multi-person voices in the preprocessed audio signal and generate multiple independent audio tracks , and combine individual identity markers to achieve precise management of who is speaking; On this basis, use a child voice recognition model to recognize the separated child voice track, improve the fitness through transfer learning and online fine-tuning, and output text results and continuously accumulate recognition experience;

[0037] Use means such as a context rearrangement function and a difference metric function to perform semantic correction and error backfilling on the text results , increasing the self-learning and self-correction capabilities: At the same time, trace the user correction information back to the speaker database in the second step and the child recognition model in the third step, continuously strengthening the robustness of multi-speaker separation and child voice recognition;

[0038] By integrating multi-modal information such as cameras and environmental sensors, realize dynamic perception of the child's face orientation, bone position, and noise source, and use a cross-modal verification scoring function to check the matching degree between the speaker's posture and the voice track, further reducing the cases of mismatching and misrecognition;

[0039] Not only enhances the capture of the target child's voice at the microphone beam level, but also implements personalized adaptation in the multi-person separation and exclusive recognition links; At the same time, supplemented by links such as context correction, interaction correction, and multi-modal fusion, a scalable high-precision child voice recognition system is formed; It can significantly reduce voice misjudgment and multi-person track confusion in a noisy environment, with both real-time performance and accuracy, and can be continuously optimized through continuously accumulated feedback data, enabling the educational robot to obtain a stable, natural, and targeted voice interaction effect in complex classrooms or home interactions. Brief Description of the Drawings

[0040] Figure 1 is a schematic flow diagram of the multi-modal adaptive child voice interaction method for the educational service robot of the present invention. Detailed Embodiments

[0041] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0042] Please refer to Figure 1 , the present invention provides a multi-modal adaptive child voice interaction method for an educational service robot, including,

[0043] Step 1: When it is determined that the local signal-to-noise ratio of the multi-channel audio matrix is significantly low based on the environmental noise threshold, based on microphone array beamforming and adaptive noise suppression, perform filtering gain adjustment on the multi-channel audio matrix in the time-frequency domain, and dynamically weaken the energy in non-target directions in combination with the estimated target azimuth and noise characteristic parameters in real time to generate a preprocessed audio signal focused on the effective sound source ;

[0044] The first step includes the following contents:

[0045] Step 101: Multi-channel synchronous acquisition and environmental feature extraction

[0046] Among the microphone array channels arranged on the fuselage, synchronously acquire audio signals to form a multi-channel audio matrix:

[0047] , where represents the time-domain audio signal received by the th microphone at time ; here, it is necessary to extract the main environmental noise distribution characteristics and preliminary localization information of each direction sound source from the multi-channel audio matrix to lay a foundation for subsequent beamforming and noise suppression; among them,

[0048] To realize the recognition and quantification of the noise distribution, an energy detection function based on wavelet transform is introduced:

[0049]

[0050] where represents the wavelet basis function with a scale of , and is used to control the resolution of the wavelet function in the time domain and frequency domain;

[0051] Thus, the time-domain signals of the microphone channels are segmented and energy-estimated at different time scales, and the relative distribution characteristics of the environmental noise and the speech mixture signals in different frequency bands are captured; Subsequently, for the preliminary localization of the sound sources in all directions, the time difference between different microphone channels is calculated based on the cross-correlation method or the Generalized Cross-Correlation (GCC) method to obtain an approximate estimation of the spatial orientation of the target sound.

[0052] Subsequently, for the preliminary localization of the sound sources in all directions, the time difference between different microphone channels is calculated based on the cross-correlation method or the Generalized Cross-Correlation (GCC) method to obtain an approximate estimation of the spatial orientation of the target sound. which is used to represent the between channel and channel with respect to the angle information of the target sound source at time

[0053] When in use, by introducing wavelet transform energy detection, the distribution of environmental noise in different time and frequency ranges can be captured more precisely, with higher recognition; by adopting multi-scale wavelet energy detection to replace the traditional single-point estimation in the time domain or frequency domain, the relative changes of noise distribution and target speech energy can be monitored separately in different frequency band ranges, so as to achieve a more refined recognition of complex environments.

[0054] Step 102, Adaptive Beamforming and Noise Suppression

[0055] Based on the multi-channel audio matrix and the approximate estimation of the spatial orientation , here, the speech signal in the target direction is gain-adjusted through adaptive beamforming, where:

[0056] In specific implementation, the beam synthesis operation is performed on each frame of audio signal in segments to obtain a single-channel audio stream synthesized by all channels according to weights :

[0057]

[0058] where represents the beam weight coefficient corresponding to channel , which is used to perform speech signal gain at the target azimuth angle . The beam weight coefficient will be dynamically adjusted during the execution to suppress the signal contribution in the non-target azimuth (determined by the target azimuth angle ) and enhance the audio signal from the target azimuth angle;

[0059] To further reduce background noise, the following noise suppression method based on wavelet domain threshold can also be combined to attenuate the time-frequency coefficients with amplitudes lower than the threshold and retain the main speech features:

[0060] ,

[0061] In the formula, represents the adaptive noise suppression threshold;

[0062] represents the suppression coefficient , which can attenuate the low-energy noise frequency band;

[0063] The above formula can be understood as setting a noise threshold in the wavelet domain or a similar time-frequency decomposition domain. If the amplitude of the beam output signal is lower than this threshold, it is regarded as noise component and suppressed. If it is higher than the threshold, it is retained to ensure speech clarity.

[0064] After obtaining the final output audio after beamforming and noise suppression , it can be packaged as a preprocessed audio signal ;

[0065] When in use, through adaptive beamforming, the speech signal in a specific direction can be highlighted and the interference in other directions can be attenuated in a noisy environment and a multi-person interaction scenario; introducing threshold-based noise suppression can adaptively attenuate the low-energy noise components in the wavelet or similar time-frequency domain and improve the purity of the target speech; combining beamforming with wavelet domain threshold suppression, rather than simply based on time-domain or frequency-domain simple gain control, can more accurately distinguish the target speech from random noise. Through the effective processing of this step, the interference of environmental noise and speech in other directions can be greatly reduced.

[0066] Step 2: Receive the preprocessed audio signal After that, use the end-to-end separation network and the speaker embedding model to perform multi-speaker disassembly on it, and generate speaker features and individual identity markers for each sound track, and output relatively pure multi-channel independent sound tracks ;

[0067] The said Step 2 includes the following contents:

[0068] Step 201: Multi-speaker end-to-end separation

[0069] After receiving the preprocessed audio signal , perform multi-speaker speech separation on the mixed audio through the deep end-to-end separation network, and output a set of relatively independent speaker signals:

[0070]

[0071] where represents the th separated speaker signal, is the number of speakers detected in the current scenario;

[0072] To achieve stable and high-precision end-to-end speech separation, a structure based on a time-domain convolutional network (such as Conv-TasNet) is adopted. After short-time block processing of the preprocessed audio signal it separates the time-domain waveform through a learnable encoder and decoder;

[0073] If an objective function with PIT loss can be used during the training phase to mitigate the impact of speaker order uncertainty on the separation effect. Specifically, let the output of the separation network be the estimated sound wave and the reference true target signal (if available during training) be Let:

[0074]

[0075] where represents the set of all permutations of speaker sequences, represents one possible permutation mapping, represents the loss metric function for measuring separation accuracy (such as measured based on speech intelligibility or scale-invariant signal distortion ratio), so as to automatically match the optimal correspondence between the network separation output and the true target, and obtain higher separation accuracy;

[0076] After completing this step, the obtained multi-channel independent audio tracks will be used as the input object for subsequent speaker feature recognition;

[0077] The separation network refers to a deep learning model used to split a mixed audio signal into several independent speaker audio tracks, such as end-to-end speech separation networks like Conv-TasNet, DPRNN, Dual-PathRNN, etc.

[0078] When in use, by introducing an end-to-end network and a PIT loss function, it can directly separate high-fidelity single-speaker audio tracks from the mixed signal, reduce separation errors, and no longer rely on fixed spectral masks or manually preset feature extraction processes. The end-to-end mode effectively adapts to the complex characteristics of multi-source aliasing in a noisy environment: enabling subsequent speaker identity extraction to be carried out on a cleaner track, significantly improving the overall recognition accuracy; integrating adaptive beamforming and an end-to-end separation network, and still maintaining stable separation ability in a high-noise, multi-speaker environment.

[0079] Step 202, Individual Feature Extraction and Speaker Identity Determination

[0080] After obtaining the multi-channel independent audio tracks After that, speaker feature extraction and identity determination need to be performed on each audio track separately. Using a deep speaker embedding (such as x-Vector, ECAPA-TDNN) model, the time-domain or frequency-domain segment signals are converted into fixed-dimensional vectors to represent the speaker features of the current audio track:

[0081]

[0082] Among them, represents the feature vector corresponding to the th audio track;

[0083] represents the deep speaker embedding network or the corresponding vector extraction function, which can be specifically implemented as a deep speaker embedding network in the form of ECAPA-TDNN;

[0084] In this extraction process, an adaptive design is made focusing on the differences in formant distribution, pronunciation habits, etc. of children's voices (for example, using the prior distribution of children's speech for network pre-training or transfer learning) to improve the recognition rate of children's voices;

[0085] After obtaining the speaker feature by making a similarity match with the registered embedding vectors in the pre-collected and constructed children's voice feature library (which will be further expanded and optimized in the third step later). If the similarity metric (such as based on cosine distance or multi-channel attention mechanism) exceeds the set threshold, it can be determined that the audio track matches the identity of a certain known child;

[0086] If the similarity metric does not exceed the preset threshold, it can be temporarily marked as a new speaker and wait for confirmation after collecting more voice data in the third step. Finally, each audio track is associated with the corresponding individual identity label :

[0087]

[0088] Among them, represents the speaker identity information of the separated audio track (such as the temporary label of a specific child, teacher, or new user); the output result will be packaged and passed to the third step (children's exclusive speech recognition and adaptive model optimization) so that the subsequent modules can call the customized model for the corresponding audio track and identity;

[0089] During use, through the deep speaker embedding model, discriminative individual feature vectors can be robustly extracted from the separated audio tracks, significantly improving the accuracy of identity determination in the case of multiple people; by comparing with the children's voice feature library and supplemented by adaptive threshold judgment, it can ensure that the audio tracks of certain familiar children or teachers can still be successfully recognized and labeled in complex environments; introducing a training mechanism for the embedding model specifically for children's voice features, rather than directly following the adult voice standard, is more adaptable to children. Dynamically combining with the recognition effect, the known user library and real-time collected data are continuously enriched and updated to meet the needs of multiple children joining the interaction at any time in the classroom or home environment.

[0090] Step 3: If there are multiple independent audio tracks and the individual identity marker jointly indicate that a certain audio track is an individual child, call the children's speech recognition model constructed based on transfer learning or integrated training and the corresponding language model perform recognition and decoding, and perform online fine-tuning on the children's speech recognition model to generate a preliminary recognition text ;

[0091] The content of the above Step 3 includes the following:

[0092] Step 301: Call of the children's exclusive model and transfer adaptation

[0093] When the output multiple independent audio tracks and the individual identity marker indicate that a certain audio track comes from an individual child, at this time, call the children's speech recognition model , the model can be obtained by transfer learning from a pre-trained large-scale adult speech model and perform deep acoustic training on the initially collected children's speech data to form a parameter set adapted to children's voices:

[0094]

[0095] Among them, represents the parameters of the adult model , represents the existing children's speech training corpus;

[0096] is a transfer adaptation function used to perform the transfer and adaptation process (for example, by freezing some network layers and only fine-tuning the high-level parameters, or enhancing the recognition effect of children's voices in the way of multi-model integration).

[0097] The initial children's acoustic model parameters can be obtained through the transfer adaptation function , combined with the corresponding language model Composition of a complete identification system:

[0098]

[0099] Then, for each separated track marked as a children's track Perform speech decoding:

[0100]

[0101] in, Indicates the preliminary recognition text of the corresponding audio track; all preliminary recognition texts It will be temporarily saved as the interim result of this step and wait for adaptive optimization and comprehensive integration in the next step.

[0102] When in use, through transfer learning or multi-model integration, the difficulty of collecting large-scale children's voice data from scratch is reduced, and a recognition model that can adapt to the characteristics of children's voices is quickly obtained; calling a special model for the individual audio track of a child in the recognition process helps to overcome the peculiarities of children's pronunciation (such as higher vocal range, unstable speaking speed, etc.), thereby significantly improving the recognition accuracy; when the adult and child models are merged or switched, the mature general voice features of the adult model are retained, while taking into account the detailed characterization of children's unique features, thereby enhancing adaptability to diverse scenarios.

[0103] Step 302: Real-time / offline fine-tuning and personalized adaptation

[0104] After obtaining the initial recognition text Finally, it also provides a way to continuously collect and fine-tune children's voices to further enhance the recognition effect for children of different ages or with different pronunciation habits. Specifically:

[0105] When a child When a certain number of audio tracks are accumulated, these new voice data can be recorded as children's voice training corpus , based on the following adaptive optimization objective function, for children's speech recognition model To perform real-time or offline updates:

[0106]

[0107] in, Indicates the current track reference text (which can be provided by the teacher or the system's automatic correction mechanism in a subsequent step), is a semantic difference metric (e.g., a comprehensive evaluation based on language modeling and sound sequence alignment);

[0108] During training, the model parameters are Perform local or global iterative updates to adapt to the child's speech characteristics. At this time, an adaptive regularization term can also be introduced according to the diversity of individual pronunciation styles to prevent loss of general performance due to overfitting to a certain child;

[0109] Finally, the updated new parameter model is stored in the model configuration corresponding to the child for preferential use in the next recognition. The result of this step will still output a temporary recognition text set , and can be iterated again when there are new training samples or teacher feedback to ensure that the recognition accuracy continuously improves as the real interaction continues.

[0110] When in use, it can accumulate data and continuously perform personalized fine-tuning for children with different ages, accents, speech rates, or language habits, solve the pain point of a thousand people with a thousand voices, adopt a training strategy of real-time or offline dual modes, improve the compatibility with the actual usage scenarios of users. When teaching time is tight, offline updates can be combined. When high timeliness is required, small-batch real-time updates can be performed. Combined with the identity information determined in the second step, it can smoothly switch between the individual speech exclusive models of each child to avoid recognition confusion caused by mixing parameters in a multi-speaker scenario.

[0111] Furthermore, the definition of the semantic difference metric function can refer to the following content:

[0112] For a given recognition output and the ground truth annotation , denote their acoustic frame sequences as and respectively, and the text (or sub-word / phoneme) sequences as and respectively. Define:

[0113]

[0114] In the formula: is the acoustic feature (such as a certain deep encoding or acoustic frame) of the recognition audio by the model at time ;

[0115] is the ground truth acoustic feature at this time (usually the extraction of correct audio features with manual or semi-automatic annotation); represents a certain distance metric (such as Euclidean norm) between two frame features;

[0116] is the first tokens (which can be sub-words, phonemes, or word-level units) in the text sequence recognized by the model;

[0117] is the first Tokens; Generate the correct token given by the language model (LM) in the context of the existing recognition token The lower the probability, the greater the deviation between the semantics or context prediction and the true value;

[0118] is the weight coefficient of the acoustic alignment term in the overall loss, which is usually related to the acoustic quality requirements; is the weight coefficient of the language model term in the overall loss, which is usually related to the emphasis on text semantic coherence and language correctness;

[0119] It is a multiplication factor used to amplify or reduce acoustic differences. The larger it is, the more sensitive it is to frame feature differences. is the language model difference amplification coefficient, which gives a greater penalty when the language model LM probability is far below 1;

[0120] The above coefficient range generally satisfies positive values: , , , In actual use, the appropriate value can be selected through experiments or hyperparameter search.

[0121] The audio duration or the total time span of the acoustic frame alignment; is the number of tokens in the true sequence (or the recognized sequence after alignment);

[0122] When targeting an individual child Collect enough ground truth comparison data After that, the recognition output With annotation Put the above The function is calculated and the obtained difference metrics are accumulated as the adaptive optimization target , and thus, the model parameter update will minimize The value of can be used to achieve personalized and high-precision child voice recognition; the correct text generated by the teacher or the system automatically corrects can also be regarded as a new Corresponding items, continue through The function strengthens the fine-tuning of the language model to provide more reliable context and acoustic priors for the next round of recognition.

[0123] Step 4: If the text is initially recognized When it is detected that it does not match the course topic or common spoken expressions, the context re-arrangement function is called And the semantic difference metric function , perform probabilistic error correction rearrangement in the dialogue history and keyword environment, and feed user correction information back to the speaker database and children's speech recognition model , forming a closed-loop self-learning mechanism;

[0124] The step 4 includes the following contents:

[0125] Step 401: Rearrangement and correction of context-related semantics

[0126] Get the children's soundtrack The recognized text results Finally, the recognized text is semantically rearranged and corrected based on the current conversation context, course topic, and common children's spoken expression library. Specifically:

[0127] Introducing context-based reordering functions , weighted evaluation is performed on the rationality of the occurrence of each word (or subword, phrase) in the text results and the output order is adjusted or error correction is performed. In form, it can be written as:

[0128]

[0129] Where: , represents the sequence of recognized tokens (e.g., words, subwords, or phonemes);

[0130] Represents context information (which may include course topics, conversation history, and keyword lists), provided by the current teaching content or conversation status recorded in step 3; Modeling language or knowledge base in context Conditional on candidate tokens If the probability is low, it means that the token may not match the expected semantics or teaching scenario;

[0131] Represents the amplification factor, which can be between 0 and 10, and is used to highlight the importance of unreasonable tokens to the result reordering: Reference resources such as possible knowledge bases, question banks or syllabi to intelligently correct or replace suspected erroneous segments;

[0132] Through comprehensive evaluation of all tokens, it can automatically perform rearrangement, replacement or insertion correction operations to generate corrected text results ; If you identify obvious conflicts between children’s spoken language and teaching keywords, you can also refer to reference resources The most similar terms are retrieved from the search results for correction. For example, when children's ambiguous pronunciation is identified as a word that has nothing to do with the course topic, the re-ranking function That is, it will trigger a high penalty value and tend to replace it with a word that better matches the current teaching content;

[0133] When used, by combining the course context with known keywords in the probability assessment, it is possible to more accurately correct ambiguities or mismatches in recognition, reduce irrelevant answers or inadequate expressions; it can quickly identify and correct rare but seriously illogical words, and maintain a high degree of dialogue coherence in complex contexts; the corrected text results produced after completing the context association Can be used for in-depth comparative analysis of teaching interaction or next step (feedback optimization).

[0134] Introducing teaching context knowledge and outline keywords into the semantic rearrangement algorithm, rather than relying solely on the probability of a general language model, greatly improves the adaptability to real classroom / home scenarios; it can effectively filter out major deviations caused by unclear or casual expressions, and ensure the availability and accuracy of text output.

[0135] Step 402: Continuous iteration and data reflow driven by interactive feedback

[0136] After completing the context semantic correction in step 401, a relatively stable corrected text result is obtained. However, in real teaching interactions, teachers, parents or students may find that there are still errors and correct them through voice or text, forming a user correction feedback data set , using this feedback information to perform the following two key operations:

[0137] If an individual identification mark is found If there are errors or correction information for a child's new pronunciation features, the user correction feedback dataset can be The corresponding audio and correct text tags are added to the pre-built speaker database ;

[0138] If the user finds a temporary recognition text set There are systematic deviations in some proprietary words or children's unique accents, and the user correction feedback dataset can be Upload to the adaptive model optimization phase in the third step (such as the adaptive optimization target defined previously) ), added to the children's dedicated training set Perform offline or real-time fine-tuning to make subsequent recognition more accurate;

[0139] Set the long-term memory factor , cumulatively manage the collected user correction data, and trigger batch updates after reaching the preset threshold to ensure that the system resource occupancy and model training overhead are balanced with the actual teaching rhythm, and finally output the corrected and feedback updated final recognized text will become the final result delivered to the learning activity or interaction scenario;

[0140] When in use, immediately backfill the correction opinions into the speaker database and the children's speech recognition model to deeply adapt to the long-term accents, pronunciation habits, and common vocabulary of specific children; the flexible batch update strategy can arrange offline training according to the classroom progress or home usage frequency, which not only ensures teaching continuity but also gradually eliminates recognition weaknesses; when encountering a similar scenario or the same child again in subsequent new interaction rounds, the recognition effect will be significantly improved, realizing a true self-learning closed loop; when integrating user feedback, update the speaker recognition and children's voice recognition simultaneously, not only fixing errors at the specific word or phrase level, but also gradually improving multiple links of the overall speech and language model. Through multiple iterations, it can not only correct existing errors but also actively learn new curriculum concepts or children's oral expressions, laying a foundation for more in-depth language interaction.

[0141] Step Five: When the robot detects the camera video and the output of the environmental sensor helps to improve recognition, perform multi-modal alignment mapping, beam control, and cross-modal verification scoring process, and cooperate with the acoustic information to dynamically update the preprocessed audio signal and the individual identity marker assignment, so that the system maintains higher recognition accuracy and pose matching in scenarios of personnel movement or sudden noise;

[0142] The said Step Five includes the following contents:

[0143] Step 501: Multi-modal data synchronization and spatio-temporal mapping

[0144] To synchronize the processing of visual information, environmental sensor readings, and audio signals, establish a cross-modal spatio-temporal alignment mechanism at the acquisition level. Assume that the robot obtains at the same time frame :

[0145] Camera video stream : including possible skeleton detection, face orientation, and face recognition input;

[0146] Environmental sensor output : including noise intensity detection, indoor positioning system output, auxiliary means such as temperature or light;

[0147] Audio signal or (from the preprocessing or separation phase of the previous step for further processing later).

[0148] Assign accurate timestamps to all acquired modality data and establish a spatial mapping function ,match video frames,audio frames and sensor readings in the same coordinate system to obtain multi-modal alignment data .

[0149] For example,when the camera detects that the head position of a certain child is ,the indoor positioning system can verify its position in the global coordinates ,and the deviation between the two can be used to correct the visual tracking error; combined with multiple independent audio tracks for the estimation of the speaker's orientation,acoustic-visual-environment tripartite fusion can be completed. At this time,a deviation amplification decision function can be defined:

[0150]

[0151] where represents the spatial distance metric; is the amplification factor,

[0152] if the value of approaches 1,indicating that the visual and positioning information is highly consistent,providing a reliable data alignment basis for subsequent multi-modal collaboration;

[0153] When in use,by finely synchronizing and spatio-temporally mapping the camera,environmental sensors,and audio data,the target child can be quickly located and locked in a noisy environment or during multi-person mobile interaction; not only ordinary spatio-temporal synchronization is achieved,but also a spatial mapping function and a deviation amplification decision function are introduced,making full use of the spatial consistency of multi-sensors to screen out or correct unreliable readings; providing a higher-level multi-modal prior for subsequent speaker recognition,noise localization,and speech processing, which can improve the accuracy and robustness of the system in a real teaching scene.

[0154] Step 502,Adaptive microphone array control and noise source suppression

[0155] The multi-modal alignment data can be used to dynamically adjust the beam direction and gain allocation of the robot microphone array and better isolate the noise source. Specifically:

[0156] When the camera detects that a certain child has an obvious movement in the mouth or head orientation,combined with the indoor positioning data When it changes accordingly, update the microphone channel The corresponding beamforming weight coefficient , so that the microphone applies higher sensitivity in the new direction; if the environmental sensor set shows that the noise source at the other end (such as the sound of opening the door, moving of tables and chairs) suddenly increases, automatically trigger the adaptive filter to more strictly suppress the noise in that direction:

[0157]

[0158] Among them, represents the angular distance from the direction of the noise source or the angular difference in the array coordinate system; is the gain coefficient to control the suppression effect, and the exponential form can quickly reduce the weight in the direction of the noise source;

[0159] Specifically: Use the camera face orientation, face position or positioning sensor output in the multimodal alignment data to determine the latest target direction , and at the same time extract the set of main noise source directions from the noise intensity distribution detected by the sensor ;

[0160] In the noise interval without the target speaker being active, sample and calculate the noise covariance matrix from the multi-channel audio matrix :

[0161]

[0162] According to the target direction and the microphone array geometry, generate the sound source steering vector , which is used to describe the time delay / phase relationship of the target sound source in each channel;

[0163] Use the MVDR beam weight to solve the beamforming weight vector assigned to each microphone channel in the frequency domain at time , and convert it to the time domain weight of each channel through the inverse STFT ; Compare the noise source direction , and additionally introduce an attenuation factor in the weight calculation to further suppress the energy in the interference direction;

[0164] Use the updated weight to weighted synthesize the multi-channel original signal to obtain the updated preprocessed audio :

[0165]

[0166] This signal takes into account both vision / positioning-guided target focusing and multi-directional noise suppression, and is used by the subsequent multi-speaker separation module. Based on more visual and sensor information, the adaptability to dynamic scenes can be greatly improved, ensuring that the target voice capture effect is still good when the speaker moves or the noise increases or decreases.

[0167] When in use, the combination of real-time beam control and adaptive filters can accurately track the child who needs to speak when multiple children are interacting, and maintain optimal sound pickup even if he or she moves around or turns his or her head at will; the rapid suppression mechanism for the direction of sudden noise can effectively improve the signal-to-noise ratio, no longer relying solely on the audio itself to estimate the noise, but using the sensor to accurately locate the position of the noise source; the pre-processed audio will be updated The feedback to steps one and two can also further improve the subsequent speech separation and recognition accuracy; the visual orientation, indoor positioning data and noise source information are synchronized for beamforming weight adjustment to avoid the misjudgment of sound sources by pure acoustic algorithms in complex and noisy environments.

[0168] Step 503: Cross-modal verification and recognition result correction

[0169] After achieving multimodal alignment and real-time microphone array control, visual clues or environmental data still need to be used as verification basis in the speech decoding stage, for example:

[0170] When the third step gives a preliminary recognition text of a children's audio track If the camera detects that the speaker is obviously in a closed-mouth state (e.g., not speaking) or the speaker's head is facing others, the text can be considered suspicious, and the corresponding sentence will be given a higher correction priority when semantic correction is performed in the fourth step;

[0171] If a text is found to be suspected to come from another student during correction by a parent or teacher (interactive feedback in step 4), the multimodal spatiotemporal recording can be compared to confirm whether the sound track is more consistent with the skeletal track of another student and to identify the student in the speaker database. Update track markers in ;

[0172] When an abnormal increase in indoor noise intensity is detected, the recognition system can review the audio decoding results of that time period and confirm whether key speech is lost in combination with the visual frame, so as to give priority to checking the accuracy of the corresponding sentence segment in the fourth step of semantic rearrangement;

[0173] A cross-modal validation score can be defined , measures the overall consistency of the speaker's activity corresponding to the audio track over a period of time with the visual,positioning data in terms of posture and position, as follows:

[0174]

[0175] Wherein: is the th child individual identifier; is the moment under visual detection (camera or skeleton tracking) of the vector difference of the child's identity pose information; is the moment under the spatial coordinates given by the positioning or environmental sensor, which is different from the corresponding coordinates of vision;

[0176] is the length of the retrospective time window; , used for weight allocation between the pose difference and the position difference; , is the exponential power, used to amplify the large difference area;

[0177] If the cross-modal verification score is lower than expected, it indicates that the multi-modal detected audio track is inconsistent with the child's activity pose or position, and further correction should be made in the fourth step or a possible recognition chain error should be prompted;

[0178] When used, cross-modal verification can apply information such as vision and positioning to the speech decoding and recognition result determination link, avoiding the scenario of misidentifying the audio track of a quiet person as him speaking; it helps to automatically compensate for the blind spots of the pure audio strategy in a multi-person teaching or accompanying environment: for example, when the voices of multiple children are similar, it is easy to be confused only relying on acoustic features, while combining vision and position can greatly reduce misjudgment; the obtained verification score can also be fed back to the speaker features and children's voice models in steps two and three, continuously correcting who is speaking and what is said, realizing a complete multi-modal closed loop.

[0179] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of this application.

[0180] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0181] In several embodiments provided in the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only for some logical function divisions. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of the devices or units can be in an electrical, mechanical, or other form.

[0182] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0183] As mentioned above, the above is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed in the present application can easily think of changes or substitutions, which should all be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. Multimodal adaptive child voice interaction method for educational service robots, characterized in that: including, when the detected environmental noise exceeds the threshold value, call the multi-array microphone beamforming and adaptive algorithm to apply directional gain and filtering to the multi-channel audio matrix, and output a preprocessed audio signal after suppressing the main interference source; after receiving the preprocessed audio signal, use an end-to-end separation network and a speaker embedding model to disassemble the overlapping speech, and extract feature vectors for each separated sound track, outputting multiple independent sound tracks and individual identity markers; when the individual identity marker indicates that a certain sound track belongs to a child individual, call the child speech recognition model and language model constructed by transfer learning to perform recognition and decoding on the corresponding track in the multiple independent sound tracks, and fine-tune the parameters in real time based on the cumulative samples to generate a preliminary recognition text; if it is detected that the preliminary recognition text conflicts with the teaching theme or common spoken language, call the context rearrangement function and the difference metric to perform semantic error correction, and inject the user's correction information into the speaker database and the child speech recognition model to form a closed-loop self-learning; when the camera and environmental sensor data can assist in the attribution of the sound track, perform multi-modal alignment mapping and call the cross-modal verification function to check the pose and position differences, and dynamically correct the beam weights and individual identity marker assignments; in the microphone array channels arranged on the fuselage, synchronously collect audio signals to form a multi-channel audio matrix: introduce an energy detection function based on wavelet transform to perform segmented energy estimation on the time-domain signals of the microphone channels at different time scales, and capture the relative distribution characteristics of the environmental noise and speech mixed signals in different frequency bands; and calculate the time difference between different microphone channels based on the cross-correlation method to obtain an approximate estimate of the spatial azimuth of the target sound.

2. The multi-modal adaptive child speech interaction method according to claim 1, wherein: Based on the multi-channel audio matrix and the approximate spatial azimuth estimate, adaptively adjust the channel weight coefficients through adaptive beamforming to enhance the speech in the target direction, and attenuate the time-frequency coefficients with amplitudes lower than the threshold in the wavelet domain to suppress background noise, obtaining a preprocessed audio signal.

3. The multi-modal adaptive child speech interaction method according to claim 2, wherein: Use an end-to-end speech separation model based on a time-domain convolutional network to perform short-time segmentation on the preprocessed audio signal, and complete multi-speaker speech separation through a learnable encoder and decoder; and use an objective function with PIT loss to automatically match the optimal correspondence between the network separation output and the true target to obtain multiple independent sound tracks.

4. The multi-modal adaptive child speech interaction method according to claim 3, wherein: Adopt a deep speaker embedding model to convert the time-domain or frequency-domain segment signals into fixed-dimensional vectors to represent the speaker characteristics of the current sound track; Identify the child identity match by making a similarity match with the registered embedding vectors in the pre-collected and constructed child sound feature library, and associate each sound track with the corresponding individual identity marker; if the similarity metric does not exceed the preset threshold, it can be temporarily marked as a new speaker.

5. The multi-modal adaptive child speech interaction method according to claim 4, wherein: When the output multi-channel independent audio tracks and the individual identity markers indicate that a certain audio track comes from a child individual, a child speech recognition model obtained by transfer learning from a pre-trained large-scale adult speech model is called; Combined with the corresponding language model, speech decoding is performed on each separated audio track marked as a child audio track to obtain a preliminary recognition text.

6. The multi-modal adaptive child speech interaction method according to claim 5, wherein: When the audio tracks of a certain child individual accumulate to a certain number, the child speech recognition model is updated in real time or offline based on an adaptive optimization objective function; Perform local or global iterative updates on the model parameters, store the updated new parameter model in the model configuration corresponding to the child, and output a set of temporary recognition texts.

7. The multi-modal adaptive child speech interaction method according to claim 6, wherein: Combined with the current conversation context, course theme and a common child spoken language expression library, a context-based re-ranking function is used to semantically re-rank and correct the recognition text; through comprehensive evaluation of all tokens, perform re-ranking, replacement or insertion error correction operations to generate a corrected text result.

8. The multi-modal adaptive child speech interaction method according to claim 7, wherein: Add the corresponding audio and correct text markers in the user correction feedback dataset to the pre-constructed speaker database, and perform offline or real-time fine-tuning on the child speech recognition model after updating the child-exclusive training set; Cumulatively manage the collected user correction data, trigger batch updates after the long-term memory coefficient reaches a preset threshold, and output the corrected and feedback updated final recognized text.

9. The multi-modal adaptive child speech interaction method according to claim 8, wherein: Establish a cross-modal spatio-temporal alignment mechanism at the acquisition level to synchronize visual information, environmental sensor readings and audio signals, and use a deviation amplification decision function to evaluate the consistency of visual and positioning coordinates, and screen out or correct unreliable readings to obtain multi-modal alignment data.

10. The multi-modal adaptive child speech interaction method according to claim 9, wherein: Dynamically adjust the beamforming weight vector of the robot microphone array according to the multi-modal alignment data, and convert it into the time-domain weights of each channel through inverse STFT; compare the noise source directions, and additionally introduce an attenuation factor in the weight calculation to suppress the energy in the interference direction; Use the updated weights to weighted synthesize the multi-channel original signals to obtain the updated preprocessed audio.

11. The multi-modal adaptive child speech interaction method according to claim 10, wherein: Perform cross-modal verification on the recognition text through visual mouth closure detection, head orientation and noise emergencies, wherein: Measure the overall consistency of the speaker activities corresponding to the audio tracks with the visual and positioning data in terms of posture and position through cross-modal verification scores; If the cross-modal verification score is lower than expected, assign a higher correction priority to the suspicious sentence segments or reassign them to the correct speaker's audio track, and synchronously update the speaker markers.

Citation Information

Patent Citations

  • Electronic device realizing voice signal recognition

    CN110364166A

  • Multi-modal conference data structuring method and device and computer equipment

    CN114298170A

Cited By

  • Children voice expression error recognition and correction method based on comparative learning

    CN121011207A