Family education interaction device and interaction method for infant education

By using intelligent interactive devices and deep learning algorithms, parents can engage in voice modeling and chat with infants and toddlers, solving the problem of a lack of standardized interaction in traditional family education and improving the language learning outcomes and quality of family education for infants and toddlers.

CN121528068APending Publication Date: 2026-02-13熊凤英
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511538492.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-27
Publication Date
2026-02-13

AI Technical Summary

Technical Problem

Traditional family education methods of voice interaction lack standardization and demonstration. Existing infant and toddler education devices have failed to effectively improve language learning outcomes and cannot provide natural and fluent chat interactions when parents are not present, thus affecting the development of infants and toddlers' language abilities.

Method used

It employs an intelligent interactive host, AR projection equipment, intelligent voice interaction module, wearable sensing device, and parent voice model module, combined with deep learning algorithms and speech recognition technology, to enable parent voice modeling and voice imitation chat interaction, providing a personalized and continuous language learning environment.

Benefits of technology

It improves the language learning outcomes of infants and toddlers, creates a familiar language environment, enhances interactive appeal and learning efficiency, and improves the family education interaction system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121528068A_ABST
    Figure CN121528068A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of infant education equipment, and particularly discloses a family education interaction device and method for infant education, and the device comprises an intelligent interaction host, AR projection equipment, an intelligent voice interaction module, wearable induction equipment, a parent control terminal, a parent voice model module, and a voice imitation chat interaction module. The method has the advantages that the interactivity is high, and the language learning ability is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention application relates to the field of infant and toddler education equipment technology, and specifically discloses an interactive family education device and interactive method for infant and toddler education. Background Technology

[0002] In early childhood education at home, voice interaction is a crucial element. Traditional methods of voice interaction in family education have several shortcomings. Parents often lack standardized and demonstrative guidance when communicating and imparting knowledge to infants and toddlers, resulting in inconsistent quality of the voice information received. Furthermore, most existing interactive educational devices for infants and toddlers are not optimized for parental voice interaction, failing to meet the needs of infants and toddlers for language learning and imitation, thus hindering the healthy development of their language abilities. In addition, existing devices cannot facilitate natural and fluent conversational interaction with infants and toddlers in the absence of parents, making it difficult to consistently create a familiar and comfortable language environment for them.

[0003] Therefore, the inventors have provided an interactive family education device and method for infant and toddler education in order to solve the above problems. Summary of the Invention

[0004] This invention aims to provide an interactive family education device and method for infants and toddlers, which includes a parent voice model module and the function of imitating parent voice to chat and interact with infants and toddlers. By standardizing parent voice demonstration through the parent voice model module and combining it with the parent voice chat interaction module, a more efficient and scientific voice interaction for infants and toddlers in family education can be achieved, improving the language learning effect and the quality of family education for infants and toddlers. Even when parents are not present, it can provide infants and toddlers with a familiar voice interaction experience.

[0005] To achieve the above objectives, the basic solution of the present invention provides an interactive family education device and method for infant and toddler education, comprising: an intelligent interactive host, which serves as a core control unit and has a built-in central processing unit, a storage module, and a communication module; AR projection devices, connected to the intelligent interactive host, are used to project virtual educational content into the real environment; The intelligent voice interaction module includes a microphone and a speaker for voice information acquisition and playback, and works in conjunction with the intelligent interaction host through voice recognition and synthesis technology. Wearable sensing devices designed for infants and young children are equipped with multiple built-in sensors to monitor the physiological and motor data of infants and young children and transmit it to the intelligent interactive host. The parent control terminal connects to the smart interactive host via a network, allowing parents to set information, select educational content, and view data reports. The parent voice model module is connected to the intelligent interactive host and has a built-in voice standard database, voice acquisition unit, voice analysis unit, voice guidance unit, voice imitation guidance unit, pronunciation correction unit, and voice standardization practice unit. The voice-simulation chat interaction module, connected to the intelligent interactive host, includes a voice feature learning unit, a voice synthesis simulation unit, an intelligent dialogue logic unit, and an interaction effect optimization unit.

[0006] Furthermore, the voice standard database stores standard voice demonstration content, the voice acquisition unit collects parents' voice information, the voice analysis unit compares and analyzes parents' voice with standard voice to generate a voice analysis report and sends it to the intelligent interactive host and the parent control terminal, the voice guidance unit provides parents with voice improvement suggestions and demonstration guidance voice based on the voice analysis report, the voice imitation guidance unit retrieves standard voice demonstration segments based on the voice analysis report to guide parents in imitation practice, the pronunciation correction unit collects parents' imitation practice voice in real time and compares it with standard voice for error correction feedback, and the voice standardization practice unit generates personalized voice standardization practice content and provides practice guidance based on parents' voice problems and learning progress.

[0007] Furthermore, the speech feature learning unit is used to extract parent speech feature parameters through deep learning algorithms; the speech synthesis simulation unit is used to simulate and generate parent speech based on the extracted speech feature parameters combined with text-to-speech technology; the intelligent dialogue logic unit is used to generate dialogue content and control speech synthesis based on infant speech information, interactive scenarios, and infant status; and the interaction effect optimization unit is used to collect infant feedback data to optimize the operation of the module.

[0008] Furthermore, the wearable sensing device incorporates sensors including an accelerometer, a heart rate sensor, and a body temperature sensor.

[0009] Furthermore, the communication module supports communication methods such as Bluetooth and Wi-Fi.

[0010] Furthermore, the operations performed by the voice analysis unit include: Parents' voice messages are segmented by phonemes and compared with a standard database using DTW (Digital Transmission Method). Generate a speech analysis report that includes a phoneme accuracy matrix and intonation deviation curve; The voice guidance unit reports: Dynamically demonstrate the articulation posture of the target phoneme using a 3D oral cavity model; Generate targeted improvement training plans, including daily practice duration and key phoneme sequences.

[0011] Furthermore, the speech feature learning unit extracts the following features using a convolutional neural network: Statistical characteristics (mean / variance / extremes) of the fundamental frequency profile of a speech signal; Histogram of the distribution of the first three formants of the speech signal; The speech synthesis simulation unit: By fusing extracted features with text input, synthetic speech with parental voice timbre is generated; It supports emotional transfer, converting neutral tones into target emotions such as excitement / gentleness.

[0012] Furthermore, the parent voice model module also includes: Real-time pronunciation correction unit: During parent imitation practice, the speech stream is segmented using endpoint detection; Error phonemes are marked in real time and a correction prompt tone is played through a directional speaker; Personalized practice units: Generate a training set of confusing phoneme pairs based on historical error patterns (e.g., / l / vs / n / ). Dynamically adjust the difficulty level of the practice (D): D = k (1 - accuracy rate of the last 5 times) + m Number of consecutive practice days (k, m are adaptive weights).

[0013] Furthermore, the following steps are included: S001: Parental Calibration Phase: Collect basic voice samples from parents to build personalized voiceprint models; Identify sets of easily mispronounced phonemes and generate customized courses for parents to learn, correct, and communicate with infants and toddlers using the correct pronunciation; S002: Parent-child interaction stage: AR projection devices display virtual characters to guide dialogue; The voice imitation chat module synthesizes conversation content in real time using parents' voice timbres. When an infant or toddler mispronounces a word, a gentle vibration is triggered on the wearable device as a notification. S003: Optimization Phase Analyzing infants' and toddlers' physiological data to determine their attention levels; Dynamically adjust the dialogue complexity C: C =α Age + β Historical accuracy - γ Real-time heart rate variability (α, β, γ are weighting coefficients) The principle and effect of this basic scheme are as follows: 1. Create a familiar language environment: The module that imitates parents' voices to chat and interact with infants and toddlers can chat and interact with them in the absence of parents, continuously creating a familiar and warm language environment for infants and toddlers, reducing the anxiety caused by parents' absence, and helping infants and toddlers' language abilities to continue to develop in a stable environment. 2. Enhance the personalization of interaction: By extracting the voice features of parents through deep learning algorithms, the imitated voice closely matches the characteristics of parents' voices. At the same time, it combines the infant's state and the interaction scenario to generate dialogue content, making the interaction more personalized, meeting the communication needs of infants and toddlers in different situations, and enhancing the attractiveness and effectiveness of the interaction. 3. Optimize interactive learning effects: The interactive effect optimization unit collects feedback in real time and optimizes module operation, continuously improving the naturalness and effectiveness of the interaction, so that infants and toddlers can better learn language expression and communication skills in the process of interacting with and imitating their parents' voices, thereby improving the efficiency and quality of language learning.

[0014] 4. Improve the family education interaction system: The addition of this module, together with other functional modules such as the parent voice model module, further improves the function of the family education interaction device, forming a complete system from parents learning standardized voice, to parents interacting with infants and toddlers, and then to continuous interaction by imitating parents' voice, comprehensively improving the quality and effectiveness of family education. Attached Figure Description

[0015] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0016] Figure 1 This illustration shows an overall schematic diagram of a family education interactive device and interactive method for infant and toddler education proposed in an embodiment of this application. Detailed Implementation

[0017] To further illustrate the technical means and effects of the present invention in achieving its intended purpose, the following detailed description of the specific implementation methods, structures, features, and effects of the present invention, in conjunction with the accompanying drawings and preferred embodiments, is provided below.

[0018] An interactive device and method for family education in infant and toddler education, implementing, for example... Figure 1As shown, the device includes a smart interactive host, which serves as the core control unit. Its internal central processing unit uses a high-performance ARM architecture chip, enabling it to quickly handle complex computational tasks and ensure the coordinated operation of all modules. The storage module uses a large-capacity solid-state drive, pre-stored with a wealth of early childhood education content, covering nursery rhymes, stories, and cognitive courses of different themes and difficulty levels. The communication module integrates Bluetooth 5.0 and Wi-Fi 6 chips. In a home environment, parents can connect the smart interactive host to the home network via Wi-Fi to achieve remote data transmission with the parent control terminal; simultaneously, the Bluetooth function can be used for quick pairing and data interaction with wearable sensing devices and other short-range devices.

[0019] In actual use, parents plug in the smart interactive host, perform initial network settings through the host's graphical user interface, and connect to the home Wi-Fi network. Afterward, the smart interactive host enters standby mode, waiting to receive data and instructions from other modules and processing them accordingly based on preset programs or user actions.

[0020] The AR projector uses DLP (Digital Light Processing) projection technology, featuring high brightness and high contrast, enabling clear projection of virtual educational content in indoor environments. It connects to the smart interactive host via an HDMI high-definition cable, ensuring stable video signal transmission. The device has a built-in environmental sensor that detects ambient light intensity, space dimensions, and other information in real time, automatically adjusting the brightness, contrast, and projection angle of the projected image.

[0021] For example, when the device is placed in the living room for interactive use, the AR projection equipment detects the distance to the wall and the lighting conditions, and automatically adjusts the projected image to a suitable size and resolution, projecting virtual educational scenes, such as forests and oceans, onto the wall in a three-dimensional and vivid form. Infants and toddlers can interact and learn in a scene that blends reality and virtuality, enhancing the fun and immersion of learning.

[0022] The intelligent voice interaction module features a dual-microphone array design, effectively suppressing ambient noise and accurately capturing the voice information of infants and parents. High-fidelity audio speakers ensure clear and natural sound playback. The module integrates an advanced speech recognition chip and natural language processing algorithms, supporting the recognition of multiple languages ​​and dialects.

[0023] During the interaction, the microphone captures voice signals in real time and converts them into digital signals, which are then transmitted to the intelligent interactive host. The voice recognition chip uses deep learning algorithms, such as a voice recognition model based on the Transformer architecture, to analyze and process the voice signals, converting them into text information. After performing semantic understanding and logical judgment based on the text information, the intelligent interactive host uses speech synthesis technology to convert the response content into a voice signal, which is then played back by the speaker, enabling voice interaction with the user.

[0024] The wearable sensing device is designed as a wristband specifically for infants and toddlers, made of soft, skin-friendly silicone to ensure comfort. It incorporates an accelerometer, heart rate sensor, and body temperature sensor. The accelerometer monitors the infant's movement, such as walking, jumping, and waving; the heart rate sensor detects real-time changes in the infant's heart rate; and the body temperature sensor continuously monitors the infant's body temperature.

[0025] Each sensor transmits data to the smart interactive host via Bluetooth Low Energy technology, sending the collected data to the host at regular intervals (e.g., every 10 seconds). In actual use, once the parent puts the wristband on the infant, the device automatically powers on and pairs with the smart interactive host, beginning to collect and transmit data in real time, providing the smart interactive host with a basis for judging the infant's status.

[0026] The parental control terminal is an application (APP) developed for smartphones or tablets. After installation, parents open the APP, register, and log in to enter the main interface. On the main interface, parents can perform various operations: in the settings module, they can enter the infant's basic information, including name, age, gender, and date of birth; in the educational content selection module, they can choose appropriate educational resources based on the infant's age and developmental stage, such as animal recognition or simple nursery rhymes for a 2-year-old; and in the data viewing module, they can view real-time learning data, physiological data, and reports on the parent's own voice training during the interaction process.

[0027] In addition, parents can remotely control the interactive device to start, pause, and stop via the APP, and adjust interactive parameters such as volume and projection brightness, thus achieving convenient management of the entire interactive process.

[0028] It also includes a voice model module, which includes a built-in voice standard database, voice acquisition unit, voice analysis unit, voice guidance unit, voice imitation guidance unit, pronunciation correction unit, and voice standardization practice unit.

[0029] The speech standard database pre-stores a large amount of standard speech examples, including standard audio recordings of various pronunciations, vocabulary, and sentences needed for language learning by infants and toddlers of different ages, as well as corresponding 3D pronunciation model data. This data was recorded and produced by professional linguists and child education experts to ensure the accuracy and standardization of the speech.

[0030] The voice acquisition unit uses a high-sensitivity MEMS microphone and is integrated into the parent control terminal or intelligent voice interaction module. When parents conduct voice training, they can activate the voice acquisition function through the APP or intelligent voice interaction module. Parents read the specified text content into the microphone, and the voice acquisition unit acquires the voice signal with a high sampling rate (e.g., 44.1kHz) and high precision (16-bit), and transmits the acquired voice data to the intelligent interaction host.

[0031] After receiving the speech data, the speech analysis unit uses the Dynamic Time Warping (DTW) algorithm to segment the parent's speech by phoneme. Then, each phoneme is compared frame-by-frame with the standard pronunciation in the speech standard database, calculating the phoneme accuracy matrix and intonation deviation curve. For example, when analyzing the parent's pronunciation of the word "cat," the unit precisely analyzes the differences between the pronunciation duration, intonation variation, and other parameters of each phoneme and the standard pronunciation, generating a detailed speech analysis report, which is then sent to the smart interactive host and the parent control terminal.

[0032] Based on the speech analysis report, the speech guidance unit retrieves relevant 3D oral cavity models and pronunciation instruction videos from the speech standard database. On the parent control terminal, it uses animation to demonstrate details such as oral cavity posture and tongue position when the target phoneme is pronounced correctly. At the same time, it generates a personalized improvement training plan, specifying the daily practice time, practice content, and key phoneme sequences to help parents improve their children's speech in a targeted manner.

[0033] Based on the speech analysis report, the speech imitation guidance unit selects standard speech demonstration segments related to the speech that parents need to improve from the speech standard database. Using AR projection equipment, the demonstration segments are projected as videos, overlaid with dynamic lip-sync demonstrations, guiding parents to practice imitation. During the imitation process, real-time voice prompts and progress feedback are provided to help parents better grasp correct pronunciation, thus facilitating the teaching of infants' and toddlers' pronunciation.

[0034] When parents practice imitating the pronunciation, the pronunciation correction unit uses endpoint detection technology to segment the speech stream in real time, accurately identifying the start and end points of each pronunciation. Then, it monitors the phonemes in the speech stream in real time. When an incorrect phoneme is detected, it immediately plays a correction prompt tone through the speaker of the intelligent voice interaction module, and at the same time, it marks the inaccurate pronunciation part with a conspicuous color on the parent's control terminal, helping parents to discover and correct pronunciation errors in a timely manner.

[0035] The pronunciation standardization practice unit generates a personalized training set of confused phoneme pairs based on parents' historical pronunciation error patterns, such as "z" and "zh", "l" and "n", etc. Simultaneously, an adaptive difficulty adjustment algorithm (D=k) is employed. (1 - accuracy rate of the last 5 times) + m The number of consecutive practice days (where k and m are adaptive weights) dynamically adjusts the practice difficulty based on the parents' practice progress. During the practice, real-time voice encouragement and guidance are provided to motivate parents to continuously improve their pronunciation.

[0036] The voice-mimicking chat interaction module includes a voice feature learning unit, a voice synthesis simulation unit, an intelligent dialogue logic unit, and an interaction effect optimization unit.

[0037] The speech feature learning unit utilizes a convolutional neural network (CNN) to perform in-depth analysis of the large amount of speech data collected by the parent speech acquisition unit. First, the speech signal is preprocessed, including noise reduction and normalization. Then, through multiple convolutional and pooling layers, core acoustic features such as the statistical features (mean, variance, and extreme values) of the fundamental frequency profile and the histogram of the distribution of the first three formants are extracted. After encoding these feature parameters, a personalized speech feature model for parents is constructed, and the model is continuously optimized and updated using newly acquired speech data to improve its accuracy.

[0038] The speech synthesis simulation unit deeply integrates the extracted parent's voice feature parameters with the input text content. Employing deep learning-based text-to-speech (TTS) technologies, such as WaveNet and Tacotron, it adjusts the timbre, intonation, and speech rate of the synthesized speech based on the voice feature parameters, generating highly realistic synthesized speech with a parent's voice. Simultaneously, it supports emotion transfer, capable of converting neutral intonation into different emotional states such as excitement, gentleness, and encouragement, based on the dialogue scenario and emotional needs, making the synthesized speech more natural and vivid.

[0039] The intelligent dialogue logic unit incorporates a knowledge graph-based dialogue strategy engine. Upon receiving voice input from an infant or toddler through the intelligent voice interaction module, it first performs speech recognition and semantic understanding to extract key information. Then, combining the current AR interaction scenario, the infant's / toddler's state data (provided by wearable sensing devices), and relevant knowledge from the knowledge graph, it uses reasoning algorithms to generate appropriate dialogue content. For example, when an infant or toddler asks "Why is the sun round?", the intelligent dialogue logic unit retrieves relevant scientific knowledge from the knowledge graph and generates a simple and easy-to-understand answer based on the infant's / toddler's cognitive level, while simultaneously controlling the speech synthesis simulation unit to output in a manner that mimics the parent's voice.

[0040] The interaction effect optimization unit evaluates the interaction effect through multimodal data collection and analysis. On the one hand, it collects the infant's and toddler's voice responses and analyzes the relevance, completeness, and language expression ability of the response content; on the other hand, it uses a camera in the living room (if present) to recognize the infant's and toddler's facial expressions and movements to determine their emotional state and level of attention; simultaneously, it combines physiological data collected by wearable sensing devices to comprehensively evaluate the interaction effect. Based on the evaluation results, reinforcement learning algorithms are used to optimize and adjust the learning model of the voice feature learning unit and the dialogue strategy of the intelligent dialogue logic unit, continuously improving the quality and effect of imitating parental voice chat interaction.

[0041] The process of using this invention is as follows: (a) Parental Calibration Phase Parents open the parent control terminal APP and select the "Parent Calibration" function module on the main interface. After entering this module, following the system prompts, parents put on headphones and read aloud the standard voice text provided in the APP into their phone's microphone. The reading process lasts approximately 5-10 minutes. The voice acquisition unit collects the parent's voice data in real time and transmits it to the smart interactive host.

[0042] The intelligent interactive host sends voice data to the voice analysis unit of the parent voice model module. The voice analysis unit uses the aforementioned analysis algorithm to perform phoneme-level analysis of the parent's voice and identify the set of phonemes that the parent is prone to making mistakes. At the same time, it uses voiceprint recognition technology to construct a personalized voiceprint model for the parent.

[0043] Based on the analysis results, the voice guidance unit generates customized voice training courses, including detailed practice plans, instructional videos, and demonstration audio, which are then pushed to the parent control terminal. Parents learn and practice according to the course schedule, gradually mastering correct pronunciation through continuous imitation and correction, and using standardized pronunciation to communicate with infants and toddlers in daily life.

[0044] (II) Parent-child interaction stage Parents select the "Start Interaction" button on the parent control terminal to activate the entire interactive device. The AR projection device then begins to work, projecting pre-selected virtual educational scenes into the real environment, such as projecting a virtual fairytale town onto the living room floor. The intelligent voice interaction module plays a welcome message and interactive guidance to attract the infant's and toddler's attention.

[0045] During the interactive process, parents and infants participate together in the virtual scene. Parents communicate with their infants through the intelligent voice interaction module, introducing objects in the scene, telling stories, and so on. At the same time, the parent voice model module collects the parents' voice information in real time, analyzes it, and provides guidance to help parents continuously improve their voice expression.

[0046] When parents need to leave, the voice-simulation chat interaction module automatically activates. The infant continues to play in the virtual environment and interact with the device via voice, such as asking "What's in the castle?" After receiving the infant's voice information, the intelligent dialogue logic unit combines the current scene and the infant's state to generate appropriate dialogue content. The voice synthesis simulation unit then synthesizes this dialogue content into audio that mimics the parent's voice and plays it back, enabling continuous chat interaction with the infant.

[0047] During this process, if the infant makes a pronunciation error, the wearable sensing device detects specific changes in the infant's voice signal (such as abnormal frequency fluctuations during pronunciation), triggering the device's gentle vibration prompt function. At the same time, the virtual character in the AR scene will provide pronunciation guidance in the form of animation to help the infant correct their pronunciation.

[0048] (III) Effect Optimization Phase Throughout the interaction, wearable sensors continuously collect the infant's physiological and motor data and transmit it to the intelligent interactive host in real time. The intelligent interactive host analyzes and processes the data, using a preset algorithm to determine the infant's level of concentration. For example, if the infant's heart rate is stable, their movements are small, and they remain in the same position for a long time, their concentration is judged to be low; if their heart rate is fast, their movements are frequent, and they interact more with the virtual scene, their concentration is judged to be high.

[0049] At the same time, the intelligent interactive host combines the infant's historical learning data (such as previous pronunciation accuracy, knowledge mastery, etc.) and uses a dynamic adjustment algorithm (C=α age+β historical accuracy-γ). Real-time heart rate variability (with α, β, and γ as weighting coefficients) automatically adjusts the complexity of the dialogue content. When infants and toddlers have high concentration and good historical learning performance, the vocabulary and sentence complexity in the dialogue are appropriately increased; when concentration is low or learning is difficult, the difficulty of the dialogue is reduced to adapt to the infants' and toddlers' learning pace and improve educational effectiveness.

[0050] After the interaction concludes, the intelligent interactive host comprehensively organizes and analyzes all data from the interaction process, including parent voice training data, infant learning data, and interaction effect evaluation data. A detailed interaction report is generated and sent to the parent control terminal. Parents can view the report to understand their infant's learning progress and the results of their own voice training, providing important reference for subsequent family education. It also provides data support for algorithm optimization and functional improvement of the device. This invention has the advantages of improving the effectiveness of infant language learning and the quality of family education.

[0051] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention in any way. Although the present invention has been disclosed above with reference to preferred embodiments, it is not intended to limit the present invention. Any person skilled in the art can make some modifications or alterations to the above-disclosed technical content to create equivalent embodiments without departing from the scope of the present invention. Any simple modifications, equivalent changes and alterations made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention shall still fall within the scope of the present invention.

Claims

1. A family education interactive device for infant and toddler education, characterized in that, include: The intelligent interactive host, as the core control unit, has a built-in central processing unit, storage module, and communication module; AR projection devices, connected to the intelligent interactive host, are used to project virtual educational content into the real environment; The intelligent voice interaction module includes a microphone and a speaker for voice information acquisition and playback, and works in conjunction with the intelligent interaction host through voice recognition and synthesis technology. Wearable sensing devices designed for infants and young children are equipped with multiple built-in sensors to monitor the physiological and motor data of infants and young children and transmit it to the intelligent interactive host. The parent control terminal connects to the smart interactive host via a network, allowing parents to set information, select educational content, and view data reports. The parent voice model module is connected to the intelligent interactive host and has a built-in voice standard database, voice acquisition unit, voice analysis unit, voice guidance unit, voice imitation guidance unit, pronunciation correction unit, and voice standardization practice unit. The voice-simulation chat interaction module, connected to the intelligent interactive host, includes a voice feature learning unit, a voice synthesis simulation unit, an intelligent dialogue logic unit, and an interaction effect optimization unit.

2. The family education interactive device for infant and toddler education according to claim 1, characterized in that, The speech standard database stores standard speech demonstration content. The speech acquisition unit collects parents' speech information. The speech analysis unit compares and analyzes parents' speech with standard speech to generate a speech analysis report and sends it to the intelligent interactive host and the parent control terminal. The speech guidance unit provides parents with speech improvement suggestions and demonstration speech based on the speech analysis report. The speech imitation guidance unit retrieves standard speech demonstration segments based on the speech analysis report to guide parents in imitation practice. The pronunciation correction unit collects parents' imitation practice speech in real time and compares it with standard speech for error correction feedback. The speech standardization practice unit generates personalized speech standardization practice content and provides practice guidance based on parents' speech problems and learning progress.

3. The family education interactive device for infant and toddler education according to claim 1, characterized in that, The speech feature learning unit is used to extract parent speech feature parameters through deep learning algorithms; the speech synthesis simulation unit is used to simulate and generate parent speech based on the extracted speech feature parameters and text-to-speech technology; the intelligent dialogue logic unit is used to generate dialogue content and control speech synthesis based on infant speech information, interactive scenarios and infant status; the interaction effect optimization unit is used to collect infant feedback data to optimize the operation of the module.

4. The family education interactive device for infant and toddler education according to claim 1, characterized in that, The wearable sensing device has built-in sensors including an accelerometer, a heart rate sensor, and a body temperature sensor.

5. The family education interactive device for infant and toddler education according to claim 1, characterized in that, The communication module supports communication methods such as Bluetooth and Wi-Fi.

6. The family education interactive device for infant and toddler education according to claim 1, characterized in that, The operations performed by the speech analysis unit include: Parents' voice messages are segmented by phonemes and compared with a standard database using DTW (Digital Transmission Method). Generate a speech analysis report that includes a phoneme accuracy matrix and intonation deviation curve; The voice guidance unit reports: Dynamically demonstrate the articulation posture of the target phoneme using a 3D oral cavity model; Generate targeted improvement training plans, including daily practice duration and key phoneme sequences.

7. The family education interactive device for infant and toddler education according to claim 1, characterized in that, The speech feature learning unit extracts the following features through a convolutional neural network: Statistical characteristics (mean / variance / extremes) of the fundamental frequency profile of a speech signal; Histogram of the distribution of the first three formants of the speech signal; The speech synthesis simulation unit: By fusing extracted features with text input, synthetic speech with parental voice timbre is generated; It supports emotional transfer, converting neutral tones into target emotions such as excitement / gentleness.

8. The family education interactive device for infant and toddler education according to claim 1, characterized in that, The parent voice model module also includes: Real-time pronunciation correction unit: During parent imitation practice, the speech stream is segmented using endpoint detection; Error phonemes are marked in real time and a correction prompt tone is played through a directional speaker; Personalized practice units: Generate a training set of confusing phoneme pairs based on historical error patterns (e.g., / l / vs / n / ). Dynamically adjust the difficulty level of the practice (D): D = k (1 - accuracy rate of the last 5 times) + m Number of consecutive practice days (k, m are adaptive weights). An interactive method based on the device of any one of claims 1-8, characterized in that, Includes the following steps: S001: Parental Calibration Phase: Collect basic voice samples from parents to build personalized voiceprint models; Identify sets of easily mispronounced phonemes and generate customized courses for parents to learn, correct, and communicate with infants and toddlers using the correct pronunciation; S002: Parent-child interaction stage: AR projection devices display virtual characters to guide dialogue; The voice imitation chat module synthesizes conversation content in real time using parents' voice timbres. When an infant or toddler mispronounces a word, a gentle vibration is triggered on the wearable device as a notification. S003: Optimization Phase Analyzing infants' and toddlers' physiological data to determine their attention levels; Dynamically adjust the dialogue complexity C: C =α Age + β Historical accuracy - γ Real-time heart rate variability (α, β, γ are weighting coefficients).