Multi-mode self-adaptive child voice interaction method for education service robot

Through multi-array microphone and adaptive technology, the method of realizing children's speech recognition in noisy and multi-speaking environments has been solved, and the recognition accuracy and system self-correction capabilities are significantly improved.

CN120164478AActive Publication Date: 2025-06-17北京爱宾果科技有限公司

Patent Information

Application Number
CN202510642427.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-19
Publication Date
2025-06-17
Estimated Expiration
2045-05-19

AI Technical Summary

Technical Problem

The prior art is difficult to achieve accurate, stable and efficient child speech recognition and interaction in noisy and multi-speaker environments, especially when background noise and multiple people speak at the same time.

Method used

The technical means of dynamic acquisition and beamforming of multi-array microphones, adaptive noise suppression, multi-speaker separation and speaker identification, children's exclusive model training and online fine-tuning, and context semantic correction and multimodal fusion are adopted.

Benefits of technology

It significantly improves the target speech focus and recognition accuracy, enhances the system's self-learning and self-correction capabilities, reduces mismatch and misunderstanding situations, and achieves a stable, natural and targeted voice interaction effect in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120164478A_ABST
    Figure CN120164478A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-modal self-adaptive child voice interaction method for an educational service robot, relates to the technical field of voice signal processing, and provides a multi-modal self-adaptive child voice interaction method for an educational service robot through multi-array microphone dynamic acquisition and beam forming, self-adaptive noise suppression, multi-speaker separation and speaker identification, and child exclusive model training and online fine tuning. Technical means such as context semantic correction and multi-modal fusion are adopted, so that the target voice focusing and recognition accuracy is greatly improved; multi-person voice tracks are distinguished through an end-to-end network, and rapid distinguishing of a multi-source voice scene is realized in combination with model optimization of child range features; meanwhile, correct vocabularies and pronunciation information are traced back to a speaker library and a child sound model by using real-time semantic correction and interactive feedback, and multi-mode cooperation of a camera and a positioning sensor is used for assisting, so that beams are quickly adjusted and interference is suppressed when a child moves at will and environmental noise suddenly increases, and the sound quality of the child is improved. And natural and smooth voice interaction experience is provided in class and family scenes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of voice signal processing, and particularly to a multi-modal adaptive child voice interaction method for an educational service robot. Background Art

[0002] In today's education and home companionship fields, more and more scenarios require voice interactive auxiliary teaching devices or robots to provide real-time question answering and interactive guidance.

[0003] Taking classroom teaching as an example, teachers often hope to use robots to assist in answering students' questions or organizing group discussions. In the home scenario, parents can also let the robot tutor their children to complete homework and language training. However, such scenarios often have environmental interference and multiple people speaking at the same time, such as random student speeches in the classroom, parallel communication between children and parents at home, etc. Coupled with the characteristics of children's pronunciation itself, such as high-pitched intonation, fast speech rate and unclear articulation, the traditional single-channel or adult model-dominated voice processing system frequently misinterprets and fails to recognize children's instructions or answers, thus seriously affecting the fluency of human-computer interaction and teaching efficiency.

[0004] After retrieval, an electronic device for realizing voice signal recognition is disclosed in the application publication number CN110364166A, including: a microphone array for collecting audio signals; a plurality of processors connected to the microphone array; each processor is paired with a beamformer and a voice recognition module, wherein each beamformer is used to perform beamforming processing on the audio signal in a set of multiple different target directions respectively to obtain corresponding multiple beam signals; each voice recognition module is used to perform voice recognition on the beam signals output by the paired beamformers respectively to obtain the voice recognition results of each beam signal; one of the processors is configured with a processing module for determining the voice recognition result of the audio signal according to the voice recognition results of each beam signal. This method performs beamforming processing in different target directions, so that at least one target direction is close to the voice signal generation direction, which can improve the accuracy of intelligent voice recognition.

[0005] Combined with the above application scenarios and the content in the prior art: Most of the mainstream speech recognition solutions on the market today use adult speech as the core training data and are usually only suitable for quiet environments or relatively single speaker conditions. Once faced with a real educational scenario with complex noise and multiple speakers speaking in parallel, the system will encounter obvious bottlenecks in such aspects as target speech focusing, sound source separation, and adaptation to children's voice characteristics. On the one hand, the noisy background and multiple people speaking simultaneously result in a low signal-to-noise ratio, making it impossible for the microphone to stably capture the speech of the target child. On the other hand, the special vocal range and oral characteristics of children are significantly different from those of adult models, making it difficult for existing technologies to obtain a high-accuracy text decoding result after separating the correct sound track. In addition, the lack of an automatic correction mechanism for teaching content or dialogue context also prevents the system from quickly correcting recognition deviations and keeping up with the spontaneous dialogue needs that children may have at any time.

[0006] Generally speaking, how to achieve accurate, stable, and efficient speech recognition and interaction in an environment with noise, multiple sources, and dominant children's speech becomes the key technical problem to be solved by the present invention. Summary of the Invention

[0007] (I) Technical problems to be solved In view of the deficiencies of the prior art, the present invention provides a multi-modal adaptive children's speech interaction method for educational service robots. Through technical means such as dynamic acquisition and beamforming of multi-array microphones, adaptive noise suppression, separation of multiple speakers and speaker identification, training and online fine-tuning of children's exclusive models, and context semantic correction and multi-modal fusion, the focusing of the target speech and the recognition accuracy rate are greatly improved, thereby solving the technical problems recorded in the background art.

[0008] (II) Technical solutions To achieve the above objectives, the present invention is realized through the following technical solutions: A multi-modal adaptive children's speech interaction method for educational service robots, including, when it is detected that the environmental noise exceeds the threshold value, calling the beamforming and adaptive algorithms of multi-array microphones to apply directional gain and filtering to the multi-channel audio matrix, suppressing the main interference sources and then outputting a preprocessed audio signal; After receiving the preprocessed audio signal, using an end-to-end separation network and a speaker embedding model to disassemble the overlapping speech, extracting feature vectors for each separated sound track, and finally outputting multiple independent sound tracks and individual identity markers; When the individual identity marker indicates that a certain sound track belongs to a child individual, calling the children's speech recognition model and language model constructed by transfer learning to perform recognition and decoding on the corresponding track in the multiple independent sound tracks, and real-time fine-tuning the parameters on the basis of cumulative samples to generate a preliminary recognition text; If it is detected that the initially recognized text conflicts with the teaching theme or common spoken language, call the context rearrangement function and the difference metric to perform semantic error correction, and inject the user's correction information into the speaker database and the children's speech recognition model to form a closed-loop self-learning; When the camera and environmental sensor data can assist in audio track attribution, perform multi-modal alignment mapping and call the cross-modal verification function to check the pose and position differences, and dynamically correct the beam weights and individual identity marker assignments.

[0009] Furthermore, in the microphone array channels arranged on the fuselage, synchronously collect audio signals to form a multi-channel audio matrix: Introduce an energy detection function based on wavelet transform to perform segmented energy estimation on the time-domain signals of the microphone channels at different time scales, capture the relative distribution characteristics of the environmental noise and speech mixed signals in different frequency bands; and calculate the time difference between different microphone channels based on the cross-correlation method to obtain an approximate estimate of the spatial orientation of the target sound.

[0010] Furthermore, based on the multi-channel audio matrix and the approximate estimate of the spatial orientation, dynamically adjust the channel weight coefficients through adaptive beamforming to enhance the speech in the target direction, and attenuate the time-frequency coefficients with amplitudes lower than the threshold in the wavelet domain to suppress background noise, and obtain the preprocessed audio signal.

[0011] Furthermore, use an end-to-end speech separation model based on a time-domain convolutional network to perform short-time segmentation on the preprocessed audio signal, and complete multi-speaker speech separation through a learnable encoder and decoder; And use an objective function with PIT loss to automatically match the optimal correspondence between the network separation output and the true target, and obtain multiple independent audio tracks.

[0012] Furthermore, adopt a deep speaker embedding model to convert the time-domain or frequency-domain segment signals into fixed-dimensional vectors to represent the speaker characteristics of the current audio track; Identify the child identity match by performing similarity matching with the registered embedding vectors in the pre-collected and constructed children's voice feature library, and associate each audio track with the corresponding individual identity marker; if the similarity metric does not exceed the preset threshold, it can be temporarily marked as a new speaker.

[0013] Furthermore, when the multiple independent audio tracks and individual identity markers output show that a certain audio track comes from a child individual, call the children's speech recognition model obtained by transfer learning from a pre-trained large-scale adult speech model; Combine the corresponding language model to perform speech decoding on each separated audio track marked as a child audio track to obtain the initially recognized text.

[0014] Further, when the audio tracks of a certain child individual accumulate to a certain number, the child speech recognition model is updated in real-time or offline based on the adaptive optimization objective function; the model parameters are iteratively updated locally or globally, and the updated new parameter model is stored in the model configuration corresponding to the child, and a temporary recognition text set is output.

[0015] Further, in combination with the current conversation context, course theme, and common children's spoken language expression library, a context-based re-ranking function is adopted to semantically re-rank and correct the recognized text; through the comprehensive evaluation of all tokens, re-ranking, replacement, or insertion error correction operations are performed to generate a corrected text result.

[0016] Further, the corresponding audio and correct text markers in the user correction feedback dataset are added to the pre-constructed speaker database, and the child speech recognition model is fine-tuned offline or in real-time after updating the child-exclusive training set; The collected user correction data is cumulatively managed, and batch updates are triggered after the long-term memory coefficient reaches the preset threshold, and the final recognized text after correction and feedback update is output.

[0017] Further, a cross-modal spatio-temporal alignment mechanism is established at the acquisition level to synchronize visual information, environmental sensor readings, and audio signals, and a deviation amplification judgment function is used to evaluate the consistency of vision and positioning coordinates, and unreliable readings are screened or corrected to obtain multi-modal alignment data.

[0018] Further, the beamforming weight vector of the robot microphone array is dynamically adjusted according to the multi-modal alignment data, and it is converted into the time-domain weight of each channel through inverse STFT; the noise source directions are compared, and an attenuation factor is additionally introduced in the weight calculation to suppress the energy in the interference direction; the updated weight is used to weighted synthesize the multi-channel original signals to obtain the updated preprocessed audio.

[0019] Further, cross-modal verification is performed on the recognized text through visual mouth closure detection, head orientation, and noise emergencies, where: the cross-modal verification score measures the overall consistency of the speaker activities corresponding to the audio tracks with the visual and positioning data in terms of posture and position. If the cross-modal verification score is lower than expected, higher correction priorities are assigned to the suspicious sentence segments or they are re-allocated to the correct speaker's audio track, and the speaker markers are updated synchronously.

[0020] (III) Beneficial effects The present invention provides a multi-modal adaptive child speech interaction method for an educational service robot, which has the following beneficial effects: In a noisy and multi-speaker environment, a complete child speech interaction system is constructed through five major steps, significantly improving the focusing ability and recognition accuracy of the robot for the speech of the target speaker; Output a preprocessed audio signal through a multi-array microphone and adaptive beamforming to reduce environmental interference for subsequent recognition. , and then, use an end-to-end separation network and a speaker embedding algorithm to separate the voices of multiple people from the preprocessed audio signal and generate multiple independent audio tracks , and combine individual identity markers to achieve precise management of who is speaking; on this basis, use a children's speech recognition model to recognize the separated children's voice tracks, improve the fitness through transfer learning and online fine-tuning, and output text results and continuously accumulate recognition experience; Use context rearrangement functions and difference metric functions and other means to perform semantic correction and error backfilling on the text results, increasing the self-learning and self-correction capabilities: at the same time, backtrack the user correction information to the speaker database in the second step and the children's recognition model in the third step to continuously strengthen the robustness of multi-speaker separation and children's voice recognition; By integrating multi-modal information such as cameras and environmental sensors, achieve dynamic perception of children's facial orientation, bone position, and noise sources, and use a cross-modal verification scoring function to check the matching degree between the speaker's posture and the voice track, further reducing the cases of mismatch and misrecognition; Not only enhance the capture of the target child's voice at the microphone beam level, but also implement personalized adaptation in the multi-person separation and exclusive recognition links; at the same time, supplemented by links such as context correction, interactive correction, and multi-modal fusion, a scalable high-precision children's voice recognition system is formed; it can significantly reduce speech misjudgment and multi-person track confusion in a noisy environment, with both real-time performance and accuracy, and can be continuously optimized through the continuously accumulated feedback data, enabling the educational robot to obtain a stable, natural, and targeted voice interaction effect in a complex classroom or home interaction. Description of the Drawings

[0021] Figure 1 It is a schematic flow diagram of the multi-modal adaptive children's voice interaction method for the educational service robot of the present invention. Detailed Embodiments

[0022] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0023] Please refer toFigure 1 , the present invention provides a multi-modal adaptive child voice interaction method for educational service robots, including Step 1: When it is determined that the local signal-to-noise ratio of the multi-channel audio matrix is significantly low according to the environmental noise threshold, based on microphone array beamforming and adaptive noise suppression, perform filtering gain adjustment on the multi-channel audio matrix in the time-frequency domain, and dynamically weaken the energy in non-target directions by combining the estimated target azimuth and noise characteristic parameters in real time to generate a preprocessed audio signal focused on the effective sound source ; The content of Step 1 is as follows: Step 101: Multi-channel synchronous acquisition and environmental feature extraction Among the microphone array channels arranged on the fuselage, synchronously acquire audio signals to form a multi-channel audio matrix: , where represents the time-domain audio signal received by the th microphone at time ; here, it is necessary to extract the main environmental noise distribution characteristics and preliminary positioning information of each direction sound source from the multi-channel audio matrix to lay a foundation for subsequent beamforming and noise suppression; among them, To realize the recognition and quantification of the noise distribution, an energy detection function based on wavelet transform is introduced:

[0024] where represents the wavelet basis function with scale , is used to control the resolution of the wavelet function in the time domain and frequency domain; Thus, perform segmented energy estimation on the time-domain signals of the microphone channels at different time scales, and capture the relative distribution characteristics of the environmental noise and speech mixed signals in different frequency bands; Subsequently, for the preliminary positioning of each direction sound source, calculate the time difference between different microphone channels based on the cross-correlation method or the generalized cross-power spectrum (GCC) method to obtain an approximate estimate of the spatial azimuth of the target sound , which is used to represent the included angle information about the target sound source between channel and channel at time ; During use, by introducing wavelet transform energy detection, the distribution of environmental noise in different time and frequency ranges can be captured more precisely, with higher recognition accuracy; by using multi-scale wavelet energy detection to replace traditional single-point estimation in the time domain or frequency domain, the relative changes in noise distribution and target speech energy can be monitored separately in different frequency band ranges, thereby achieving more refined recognition of complex environments.

[0025] Step 102, Adaptive beamforming and noise suppression Based on the multi-channel audio matrix and the approximate estimation of spatial orientation , here, the speech signal in the target direction is amplified through adaptive beamforming, where: In specific implementation, the beam synthesis operation is performed on each frame of the audio signal in segments to obtain a single-channel audio stream synthesized by all channels according to weights :

[0026] Among them, represents the beam weight coefficient corresponding to channel , which is used to amplify the speech signal at the target azimuth angle . The beam weight coefficient will be dynamically adjusted during the execution process to suppress the signal contribution from non-target azimuths (determined by the target azimuth angle ) and enhance the audio signal from the target azimuth angle; To further reduce background noise, the following noise suppression method based on wavelet domain threshold can also be combined to attenuate the time-frequency coefficients with amplitudes lower than the threshold and retain the main speech features: ,

[0027] In the formula, represents the adaptive noise suppression threshold; represents the suppression coefficient , which can attenuate the low-energy noise frequency bands; The above formula can be understood as setting a noise threshold in the wavelet domain or a similar time-frequency decomposition domain. If the amplitude of the beam output signal is lower than the threshold, it is regarded as noise component and suppressed. If it is higher than the threshold, it is retained to ensure speech clarity.

[0028] After obtaining the final output audio after beamforming and noise suppression , it can be packaged as a preprocessed audio signal ; In use, through adaptive beamforming, in a noisy environment and a multi-person interaction scenario, it is possible to highlight the speech signals in a specific direction and attenuate the interference from other directions; introducing threshold-based noise suppression can adaptively attenuate low-energy noise components in the wavelet or similar time-frequency domain, improving the purity of the target speech; combining beamforming with wavelet-domain threshold suppression, rather than simply based on time-domain or frequency-domain gain control, can more accurately distinguish the target speech from random noise. Through the effective processing of this step, the interference of environmental noise and speech from other directions can be greatly reduced.

[0029] Step 2: Receive the preprocessed audio signal After that, use an end-to-end separation network and a speaker embedding model to perform multi-speaker decomposition on it, and generate speaker features and individual identity markers for each sound track, and output relatively pure multi-channel independent sound tracks ; The above Step 2 includes the following contents: Step 201: Multi-speaker end-to-end separation After receiving the preprocessed audio signal , perform multi-speaker speech separation on the mixed audio through a deep end-to-end separation network, and output a set of relatively independent speaker signals:

[0030] Among them, represents the th separated speaker signal, is the number of speakers detected in the current scenario; To achieve stable and high-precision end-to-end speech separation, a structure based on a time-domain convolutional network (such as Conv-TasNet) is used to perform short-time block division on the preprocessed audio signal , and then separate the time-domain waveform through a learnable encoder and decoder; If an objective function with PIT loss can be used during the training phase to mitigate the impact of speaker order uncertainty on the separation effect. Specifically, the output of the separation network can be set as the estimated sound wave , and the reference true target signal (if available during training) is . Let:

[0031] Among them, represents the set of all speaker sequence permutations, represents one possible permutation mapping, A loss metric function for measuring separation accuracy (e.g., measured based on speech intelligibility or scale-invariant signal distortion ratio) is used to automatically match the optimal correspondence between the network separation output and the true target, obtaining higher separation accuracy; After this step, the obtained multiple independent audio tracks will be used as the input objects for subsequent speaker feature recognition; The separation network refers to a deep learning model used to split a mixed audio signal into several independent speaker audio tracks, such as end-to-end speech separation networks like Conv-TasNet, DPRNN, Dual-PathRNN, etc.

[0032] When in use, by introducing an end-to-end network and the PIT loss function, high-fidelity individual speaker audio tracks can be directly separated from the mixed signal, reducing separation errors and no longer relying on fixed spectral masks or manually preset feature extraction processes. The end-to-end mode effectively adapts to the complex characteristics of multi-source aliasing in a noisy environment: enabling subsequent speaker identity extraction to be performed on cleaner tracks, significantly improving the overall recognition accuracy; integrating adaptive beamforming and an end-to-end separation network, and still maintaining stable separation ability in high-noise, multi-speaker environments.

[0033] Step 202: Individual feature extraction and speaker identity determination After obtaining multiple independent audio tracks it is necessary to perform speaker feature extraction and identity determination on each audio track respectively. Using a deep speaker embedding (such as x-Vector, ECAPA-TDNN) model, the time-domain or frequency-domain segment signals are converted into fixed-dimensional vectors to represent the speaker features of the current audio track:

[0034] where represents the feature vector corresponding to the th audio track; represents a deep speaker embedding network or the corresponding vector extraction function, which can be specifically implemented as a deep speaker embedding network in the form of ECAPA-TDNN; During this extraction process, an adaptive design is made for the differences in formant distribution, pronunciation habits, etc. of children's voices (for example, using the prior distribution of children's speech for network pre-training or transfer learning) to improve the recognition of children's voices; Obtain the speaker features After that, by making a similarity match with the registered embedded vectors in the pre-collected and constructed children's voice feature library (which will be further expanded and optimized in the third step later), if the similarity metric (such as based on cosine distance or multi-channel attention mechanism) exceeds the set threshold, it can be determined that the audio track matches a certain known child identity; If the similarity metric does not exceed the preset threshold, it can be temporarily marked as a new speaker and wait for confirmation after collecting more voice data in the third step. Finally, each audio track is associated with the corresponding individual identity label :

[0035] Among them, represents the speaker identity information of the separated audio track (such as the temporary label of a specific child, teacher, or new user); the output result will be packaged and passed to the third step (children's exclusive speech recognition and adaptive model optimization) so that subsequent modules can call the customized model for the corresponding audio track and identity; When in use, through the deep speaker embedding model, discriminative individual feature vectors can be robustly extracted from the separated audio tracks, greatly improving the identity determination accuracy in the case of multiple people; by comparing with the children's voice feature library and supplemented by adaptive threshold judgment, it can ensure that the audio tracks of some familiar children or teachers can still be successfully recognized and labeled in a complex environment; introducing a training mechanism for the embedding model specifically for children's voice features, rather than directly following the adult voice standard, is more adaptable to children. Dynamically combining with the recognition effect, the known user library and real-time collected data are continuously enriched and updated to meet the needs of multiple children joining the interaction at any time in the classroom or home environment.

[0036] Step Three. If multiple independent audio tracks and the individual identity label jointly indicate that a certain audio track is a child individual, call the children's speech recognition model constructed based on transfer learning or integrated training and the corresponding language model to perform recognition and decoding, and make online fine-tuning on the children's speech recognition model to generate the preliminary recognition text ; The said Step Three includes the following contents: Step 301. Children's exclusive model call and transfer adaptation When the output multiple independent audio tracks and the individual identity label show that a certain audio track comes from a child individual, at this time, call the children's speech recognition model , the model can be based on a large-scale pre-trained adult speech model The model is obtained through transfer learning, and deep acoustic training is performed on the initially collected children's voice data to form a parameter set adapted to children's voices:

[0037] in, Indicates adult model Parameters, Represents the existing children's speech training corpus; It is a migration adaptation function used to perform the migration and adaptation process (for example, by freezing some network layers and only fine-tuning high-level parameters, or enhancing the recognition effect of children's voices by integrating multiple models).

[0038] The initial child acoustic model parameters can be obtained by migrating the adaptation function , combined with the corresponding language model Composition of a complete identification system:

[0039] Then, for each separated track marked as a children's track Perform speech decoding:

[0040] in, Indicates the preliminary recognition text of the corresponding audio track; all preliminary recognition texts It will be temporarily saved as the interim result of this step and wait for adaptive optimization and comprehensive integration in the next step.

[0041] When in use, through transfer learning or multi-model integration, the difficulty of collecting large-scale children's voice data from scratch is reduced, and a recognition model that can adapt to the characteristics of children's voices is quickly obtained; calling a special model for the individual audio track of a child in the recognition process helps to overcome the peculiarities of children's pronunciation (such as higher vocal range, unstable speaking speed, etc.), thereby significantly improving the recognition accuracy; when the adult and child models are merged or switched, the mature general voice features of the adult model are retained, while taking into account the detailed characterization of children's unique features, thereby enhancing adaptability to diverse scenarios.

[0042] Step 302: Real-time / offline fine-tuning and personalized adaptation After obtaining the initial recognition text Finally, it also provides a way to continuously collect and fine-tune children's voices to further enhance the recognition effect for children of different ages or with different pronunciation habits. Specifically: When a child When the number of audio tracks accumulates to a certain amount, these new voice data can be recorded as children's voice training corpus , for the children's speech recognition model based on the following adaptive optimization objective function to perform real-time or offline updates:

[0043] wherein, represents the reference text for the current audio track (which can be provided by the teacher or the system's automatic correction mechanism in subsequent steps), is a semantic difference metric function (such as a comprehensive evaluation based on language modeling and sound pause sequence alignment); During training, the model parameters will be iteratively updated locally or globally to adapt to the speech characteristics of the child. At this time, an adaptive regularization term can also be introduced according to the diversity of individual pronunciation styles to prevent the loss of general performance due to overfitting to a certain child; Finally, the updated new parameter model will be stored in the model configuration corresponding to the child for preferential use in the next recognition. The result of this step will still output a temporary recognition text set , which can be iterated again when there are new training samples or teacher feedback to ensure that the recognition accuracy continues to improve as the real interaction continues.

[0044] When in use, it can accumulate data and continuously perform personalized fine-tuning for children with different ages, accents, speech rates or language habits, solve the pain point of thousands of voices, adopt a real-time or offline dual-mode training strategy, improve the compatibility with the actual usage scenarios of users. When teaching time is tight, offline updates can be combined. When high timeliness is required, small-batch real-time updates can be performed. Combined with the identity information determined in the second step, it can smoothly switch between the individual speech exclusive models of each child to avoid recognition confusion caused by mixing parameters in multi-speaker scenarios.

[0045] Furthermore, the definition of the semantic difference metric function can refer to the following content: For the given recognition output and the ground truth annotation , denote their acoustic frame sequences as and respectively, and the text (or sub-word / phoneme) sequences as and respectively. Define:

[0046] In the formula: is at time In the above, the model recognizes the acoustic features of the audio (such as some kind of deep encoding or acoustic frame); For this time Ground truth acoustic features on (usually the extraction of correct audio features with manual or semi-automatic annotations); Some distance measure representing the features of two frames (e.g., Euclidean norm); In the text sequence recognized by the model, tokens (which can be subwords, phonemes, or word-level units); is the first Tokens; Generate the correct token given by the language model (LM) in the context of the existing recognition token The lower the probability, the greater the deviation between the semantics or context prediction and the true value; is the weight coefficient of the acoustic alignment term in the overall loss, which is usually related to the acoustic quality requirements; is the weight coefficient of the language model term in the overall loss, which is usually related to the emphasis on text semantic coherence and language correctness; It is a multiplication factor used to amplify or reduce acoustic differences. The larger it is, the more sensitive it is to frame feature differences. is the language model difference amplification coefficient, which gives a greater penalty when the language model LM probability is far below 1; The above coefficient range generally satisfies positive values: , , , In actual use, the appropriate value can be selected through experiments or hyperparameter search.

[0047] The audio duration or the total time span of the acoustic frame alignment; is the number of tokens in the true sequence (or the recognition sequence after alignment); When targeting an individual child Collect enough ground truth comparison data After that, the recognition output With annotation Put the above The function is calculated and the obtained difference metrics are accumulated as the adaptive optimization target , and thus, the model parameter update will minimize The value of can be used to achieve personalized and high-precision child voice recognition; the correct text generated by the teacher or the system automatically corrects can also be regarded as a new The corresponding item continues through fine-tuning the function-enhanced language model part to provide more reliable context and acoustic priors for the next round of recognition.

[0048] Step 4: If the initially recognized text is detected to be inconsistent with the course theme or common spoken expressions, call the context rearrangement function and the semantic difference metric function , perform probabilistic error correction and rearrangement in the dialogue history and keyword environment, and backfill the user correction information into the speaker database and the children's speech recognition model , forming a closed-loop self-learning mechanism; The content of the above Step 4 includes the following: Step 401: Situation-related semantic rearrangement and correction After obtaining the text result recognized from the children's audio track , combine the current dialogue context, course theme, and common children's spoken expression library to perform semantic rearrangement and correction on the recognized text. Specifically: Introduce a context-based reordering function , perform weighted evaluation on the reasonableness of the occurrence of each word (or sub-word, phrase) in the text result and adjust the output order or perform error correction. Formally, it can be written as:

[0049] where: represents the recognized token sequence (such as words, sub-words, or phonemes); represents the context information (which can include the course theme, dialogue history, keyword list), provided by the current teaching content or dialogue state recorded in Step 3; is the reasonableness score of the candidate token under the context by the language or knowledge base model. If this probability is low, it means that the token may not match the expected semantics or teaching scenario; represents the amplification factor, and its value can be between 0 and 10, used to highlight the importance of unreasonable tokens for result rearrangement: is the possible knowledge base, exercise bank, teaching syllabus, or other reference resources for intelligent correction or replacement of suspected error fragments; Through the comprehensive evaluation of all tokens, rearrangement, replacement, or insertion error correction operations can be automatically performed to generate the corrected text result ; when it is recognized that there is an obvious conflict between children's spoken language and teaching keywords, it can also be retrieved from the reference resources​ Retrieve the most similar entries in the dictionary for correction. For example, when a child's slurred pronunciation is recognized as a word that has nothing to do with the course theme, the re-ranking function will trigger a high penalty value and tend to replace it with a word that better matches the current teaching content; When in use, by combining the course context with known keywords in the probability assessment, it is possible to more accurately correct the ambiguity or mismatch in recognition, reducing the situation of answering off-topic or failing to convey the intended meaning; it can quickly identify and correct rare but severely illogical words, maintaining a high degree of dialogue coherence in complex contexts; the corrected text result generated after completing the context association can be used for teaching interaction or in-depth comparative analysis in the next step (feedback optimization).

[0050] Introducing teaching context knowledge and syllabus keywords into the semantic re-ranking algorithm, rather than relying solely on the probabilities of general language models, greatly improves the adaptability to real classroom / home scenarios; it can effectively filter out major deviations caused by unclear speech or casual expressions, ensuring the usability and accuracy of the text output.

[0051] Step 402: Continuous iteration and data reflux driven by interactive feedback After completing the context semantic correction in step 401, a relatively stable corrected text result is obtained , but in real teaching interactions, teachers, parents, or the students themselves may still find errors and correct them via voice or text, forming a user correction feedback dataset , and use this feedback information to perform the following two key operations: If it is found that the individual identity marker is incorrect or there is correction information for a certain child's new pronunciation feature, the corresponding audio and correct text markers in the user correction feedback dataset can be added to the pre-constructed speaker database ; If the user discovers systematic deviations in the temporary recognition text set in some proper nouns or the unique accents of children, the user correction feedback dataset can be uploaded to the adaptive model optimization link in the third step (as defined by the adaptive optimization goal ), and supplemented to the child-specific training set for offline or real-time fine-tuning to make subsequent recognition more accurate; Set the long-term memory coefficient , cumulatively manage the collected user correction data, and trigger batch updates after reaching the preset threshold to ensure that the system resource occupancy and model training overhead are balanced with the actual teaching rhythm, and finally output the corrected and feedback updated final recognized text will become the final result delivered to the learning activity or interaction scenario; When in use, immediately backfill the correction opinions into the speaker database and the children's speech recognition model in it, it is possible to deeply adapt to the long-term accent, pronunciation habits, and common vocabulary of specific children; the flexible batch update strategy can arrange offline training according to the classroom progress or home usage frequency, which not only ensures teaching continuity but also gradually eliminates recognition weaknesses; when encountering similar scenarios or the same child again in subsequent new interaction rounds, the recognition effect will be significantly improved, realizing a true self-learning closed loop; when integrating user feedback, update the speaker recognition and children's voice recognition at the same time, not only repair the errors at the specific word or phrase level, but also gradually improve multiple links of the overall speech and language model. Through multiple iterations, it can not only correct existing errors but also actively learn new curriculum concepts or children's oral expressions, laying a foundation for deeper language interaction.

[0052] Step Five. When the robot detects the camera video and the output of the environmental sensor which helps to improve recognition, perform multi-modal alignment mapping, beam control, and cross-modal verification scoring process, and cooperate with the acoustic information to dynamically update the preprocessed audio signal and individual identity marking assignment, so that the system can maintain higher recognition accuracy and pose matching degree in scenarios of personnel movement or sudden noise; The said Step Five includes the following content: Step 501. Multi-modal data synchronization and spatio-temporal mapping To synchronize the processing of visual information, environmental sensor readings, and audio signals, establish a cross-modal spatio-temporal alignment mechanism at the acquisition level. Assume that the robot obtains at the same time frame the following: Camera video stream : including possible skeleton detection, face orientation, and face recognition input; Output of the environmental sensor : including noise intensity detection, indoor positioning system output, auxiliary means such as temperature or light; Audio signal or (from the preprocessing or separation stage of the previous step, for subsequent further processing).

[0053] All the acquired modal data are respectively assigned accurate timestamps, and a spatial mapping function is established , so that video frames, audio frames and sensor readings are matched in the same coordinate system, and multi-modal alignment data is obtained .

[0054] For example, when the camera detects that the head position of a certain child is , the indoor positioning system can verify its position in the global coordinates , and the deviation between the two can be used to correct the visual tracking error; combined with multiple independent audio tracks for the estimation of the speaker's orientation, the acoustic-visual-environment tripartite fusion can be completed. At this time, a deviation amplification decision function can be defined :

[0055] Among them, represents the spatial distance metric; is the amplification factor, If the value of approaches 1, it indicates that the visual and positioning information is highly consistent, which can provide a reliable data alignment basis for subsequent multi-modal collaboration; When in use, by finely synchronizing and spatio-temporally mapping the camera, environmental sensors, and audio data, the target child can be quickly located and locked in a noisy environment or during multi-person mobile interaction; not only ordinary spatio-temporal synchronization is achieved, but also a spatial mapping function and a deviation amplification decision function are introduced to fully utilize the spatial consistency of multi-sensors to screen out or correct unreliable readings; provide a higher-level multi-modal prior for subsequent speaker recognition, noise localization, and speech processing, which can improve the accuracy and robustness of the system in a real teaching scene.

[0056] Step 502, Adaptive microphone array control and noise source suppression The multi-modal alignment data can be used to dynamically adjust the beam direction and gain allocation of the robot microphone array, and better isolate the noise source. Specifically: When the camera detects that a certain child has an obvious movement in the mouth or head orientation, combined with the indoor positioning data also changes accordingly, update the beamforming weight coefficient corresponding to the microphone channel , so that the microphone has higher sensitivity in the new direction; if the environmental sensor set shows that the noise source at the other end (such as the sound of opening the door, moving of desks and chairs) suddenly increases, automatically trigger the adaptive filter to more strictly suppress the noise in that direction:

[0057] Among them, represents the angular distance from the direction of the noise source or the angular difference in the array coordinate system; is the gain coefficient for controlling the suppression effect, and the exponential form can rapidly reduce the weight in the direction of the noise source; Specifically: Utilize the camera face orientation, face position, or the output of the positioning sensor in the multi-modal alignment data to determine the latest target direction , and simultaneously extract the set of main noise source directions from the noise intensity distribution detected by the sensor ; In the noise interval without the target speaker being active, sample and calculate the noise covariance matrix from the multi-channel audio matrix :

[0058] According to the target direction and the microphone array geometry, generate the sound source steering vector , which is used to describe the time delay / phase relationship of the target sound source in each channel; Use the MVDR beam weight to solve for the beamforming weight vector assigned to each microphone channel in the frequency domain at time , and convert it to the time domain weight of each channel through the inverse STFT ; Compare with the noise source direction , and additionally introduce the attenuation factor in the weight calculation to further suppress the energy in the interference direction; Utilize the updated weight to perform weighted synthesis on the multi-channel original signal to obtain the updated preprocessed audio :

[0059] This signal takes into account both the target focusing guided by vision / localization and the multi-directional noise suppression, and is provided for the subsequent multi-speaker separation module. Based on more visual and sensor information, the adaptability to dynamic scenes can be significantly improved, ensuring that good target speech capture effects can still be maintained when the speaker moves or the noise increases or decreases.

[0060] When in use, the combination of real-time beam control and adaptive filters can accurately track the child who needs to speak when multiple children are interacting, and maintain optimal sound pickup even if he or she moves around or turns his or her head at will; the rapid suppression mechanism for the direction of sudden noise can effectively improve the signal-to-noise ratio, no longer relying solely on the audio itself to estimate the noise, but using the sensor to accurately locate the position of the noise source; the pre-processed audio will be updated The feedback to steps one and two can also further improve the subsequent speech separation and recognition accuracy; the visual orientation, indoor positioning data and noise source information are synchronized for beamforming weight adjustment to avoid the misjudgment of sound sources by pure acoustic algorithms in complex and noisy environments.

[0061] Step 503: Cross-modal verification and recognition result correction After achieving multimodal alignment and real-time microphone array control, visual clues or environmental data still need to be used as verification basis in the speech decoding stage, for example: When the third step gives a preliminary recognition text of a children's audio track If the camera detects that the speaker is obviously in a closed-mouth state (e.g., not speaking) or the speaker's head is facing others, the text can be considered suspicious, and the corresponding sentence will be given a higher correction priority when semantic correction is performed in the fourth step; If a text is found to be suspected to come from another student during correction by a parent or teacher (interactive feedback in step 4), the multimodal spatiotemporal recording can be compared to confirm whether the sound track is more consistent with the skeletal track of another student and to identify the student in the speaker database. Update track markers in ; When an abnormal increase in indoor noise intensity is detected, the recognition system can review the audio decoding results of that time period and confirm whether key speech is lost in combination with the visual frame, so as to give priority to checking the accuracy of the corresponding sentence segment in the fourth step of semantic rearrangement; A cross-modal validation score can be defined , measures the overall consistency of the speaker's activity corresponding to the audio track over a period of time with the visual,positioning data in terms of posture and position, as follows:

[0062] Where: For the Individual identification of children; For the moment Visual detection (camera or skeleton tracking) for child identification Vector difference of posture information; For the moment The spatial coordinates given by the positioning or environmental sensor differ from the visual coordinates; is the length of the lookback time window; , which is used to perform weight allocation between the attitude difference and the position difference; , is an exponential power, which is used to amplify the large difference area; If the cross-modal verification score is lower than expected, it indicates that the multi-modal detection detects that the audio track is inconsistent with the activity posture or position of the child, and further correction should be made in the fourth step or a possible recognition chain error should be prompted; When in use, cross-modal verification can apply information such as vision and positioning to the speech decoding and recognition result determination links, avoiding the scenario of misidentifying the audio track of a quiet person as him speaking; it helps to automatically compensate for the blind spots of the pure audio strategy in a multi-person teaching or companion environment: for example, when the voices of multiple children are similar, it is easy to be confused only relying on acoustic features, while combining vision and position can greatly reduce misjudgment; the obtained verification score can also be fed back to the speaker characteristics and child voice models in step two and step three, continuously correcting who is speaking and what is said, and realizing a complete multi-modal closed loop.

[0063] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but this implementation should not be considered to exceed the scope of this application.

[0064] Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0065] In several embodiments provided in this application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only for some logical function divisions. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point, the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces, and the indirect couplings or communication connections of the devices or units can be in electrical, mechanical, or other forms.

[0066] The unit described as a separation component may or may not be physically separated. The component displayed as a unit may or may not be a physical unit, that is, it may be located in one place or may be distributed across multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0067] As described above, the above are only specific embodiments of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed in the present application can easily think of changes or substitutions, which should all be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A multimodal adaptive child voice interaction method for an educational service robot, characterized in that: include, When the ambient noise exceeds the limit, the multi-array microphone beamforming and adaptive algorithm is called to apply directional gain and filtering to the multi-channel audio matrix, suppress the main interference source, and then output the pre-processed audio signal; After receiving the pre-processed audio signal, it uses an end-to-end separation network and speaker embedding model to disassemble the aliased speech, extract feature vectors for each separated audio track, and output multiple independent audio tracks and individual identity tags; When the individual identity tag indicates that a certain audio track belongs to a child, the child speech recognition model and language model constructed by transfer learning are called to perform recognition and decoding on the corresponding track in multiple independent audio tracks, and the parameters are fine-tuned in real time based on the accumulated samples to generate preliminary recognition text; If the initial recognition text is detected to be inconsistent with the teaching topic or common spoken language, the context rearrangement function and difference measurement are called to perform semantic error correction, and the user correction information is injected into the speaker database and the children's speech recognition model to form a closed-loop self-learning; When camera and environmental sensor data can assist in track attribution, multimodal alignment mapping is performed and cross-modal verification functions are called to check pose and position differences, dynamically correcting beam weights and individual identity tag assignments.

2. The multimodal adaptive children's voice interaction method according to claim 1, characterized in that: In the microphone array channels arranged on the fuselage, audio signals are collected synchronously to form a multi-channel audio matrix: An energy detection function based on wavelet transform is introduced to perform segmented energy estimation on the time domain signal of the microphone channel at different time scales, capturing the relative distribution characteristics of the mixed signal of ambient noise and speech in different frequency bands. The time difference between different microphone channels is calculated based on the cross-correlation method to obtain an approximate estimate of the spatial orientation of the target sound.

3. The multimodal adaptive children's voice interaction method according to claim 2, characterized in that: Based on the multi-channel audio matrix and the approximate estimation of spatial orientation, the channel weight coefficients are dynamically adjusted through adaptive beamforming to enhance the speech in the target direction, and the time-frequency coefficients with amplitudes below the threshold are attenuated in the wavelet domain to suppress background noise and obtain the preprocessed audio signal.

4. The multimodal adaptive children's voice interaction method according to claim 3, characterized in that: An end-to-end speech separation model based on a time-domain convolutional network is used to divide the preprocessed audio signal into short-term blocks, and multi-speaker speech separation is completed through a learnable encoder and decoder; And use the objective function with PIT loss to automatically match the optimal correspondence between the network separation output and the true target to obtain multiple independent audio tracks.

5. The multimodal adaptive children's voice interaction method according to claim 4, characterized in that: Based on the deep speaker embedding model, the time domain or frequency domain segment signal is converted into a fixed-dimensional vector to represent the speaker characteristics of the current audio track; By matching similarity with the registered embedding vectors in the pre-collected and constructed children's sound feature library, the child's identity match is identified and each audio track is associated with the corresponding individual identity marker; If the similarity measure does not exceed a preset threshold, it can be temporarily marked as a new speaker.

6. The multimodal adaptive children's voice interaction method according to claim 5, characterized in that: When the output multiple independent audio tracks and individual identity tags indicate that a certain audio track comes from a child individual, a child speech recognition model based on transfer learning from a pre-trained large-scale adult speech model is called; Perform speech decoding on each separated audio track marked as a children's track in combination with the corresponding language model to obtain preliminary recognition text.

7. The multimodal adaptive children's voice interaction method according to claim 6, characterized in that: When a certain number of audio tracks of an individual child are accumulated, the child speech recognition model is updated in real time or offline based on the adaptive optimization objective function; The model parameters are updated locally or globally iteratively, the updated new parameter model is stored in the model configuration of the corresponding child, and a temporary recognition text set is output.

8. The multimodal adaptive children's voice interaction method according to claim 7, characterized in that: Combining the current conversation context, course topics and a library of common children's spoken expressions, a context-based reordering function is used to semantically rearrange and correct the recognized text; through a comprehensive evaluation of all tokens, rearrangement, replacement or insertion correction operations are performed to generate a corrected text result.

9. The multimodal adaptive children's voice interaction method according to claim 8, characterized in that: Add the corresponding audio and correct text tags in the user correction feedback dataset to the pre-built speaker database, and perform offline or real-time fine-tuning on the child speech recognition model after updating the child-specific training set; The collected user correction data is accumulated and managed, and batch updates are triggered after the long-term memory coefficient reaches the preset threshold. The corrections are output and the updated final certification text is fed back.

10. The multimodal adaptive children's voice interaction method according to claim 9, characterized in that: A cross-modal spatiotemporal alignment mechanism is established at the acquisition level to synchronize visual information, environmental sensor readings, and audio signals. The deviation amplification judgment function is used to evaluate the consistency of visual and positioning coordinates, and unreliable readings are screened out or corrected to obtain multimodal alignment data.

11. The multimodal adaptive children's voice interaction method according to claim 10, characterized in that: The beamforming weight vector of the robot microphone array is dynamically adjusted based on the multimodal alignment data, and converted into the time domain weight of each channel through the inverse STFT. The direction of the noise source is compared, and an additional attenuation factor is introduced in the weight calculation to suppress the energy in the interference direction. The updated weights are used to perform weighted synthesis on the multi-channel original signals to obtain updated pre-processed audio.

12. The multimodal adaptive children's voice interaction method according to claim 9, characterized in that: Cross-modal verification of recognized text using visual mouth closure detection, head orientation, and noise bursts, including: The cross-modal verification score measures the overall consistency of the speaker activity corresponding to the audio track with the visual and positioning data in terms of posture and position; If the cross-modal verification score is lower than expected, the suspicious segment is given a higher correction priority or reassigned to the correct speaker track, and the speaker tag is updated synchronously.

Citation Information

Patent Citations

  • Electronic device realizing voice signal recognition

    CN110364166A

  • Child accent recognition equipment control method, equipment, storage medium and device

    CN110767240A

  • Speech function automatic evaluation system and method based on speech recognition

    CN113496696A

  • Multi-modal conference data structuring method and device and computer equipment

    CN114298170A

  • Voice recognition method and device for audio and video progressive fusion training in noise environment

    CN119107945A

Cited By

  • Conference summary automatic generation method and system based on OCR technology

    CN120337862A

  • Semantic fingerprint adaptive training method for teaching service robot

    CN120653994A

  • Semantic fingerprint adaptive training method for teaching service robot

    CN120653994B

  • Intelligent terminal speaker accurate recognition method based on voice recognition and intelligent terminal

    CN120853580A

  • Human-computer interaction method of intelligent simulation baby robot based on user behavior pattern recognition

    CN120872153A