Image big data and voice interaction fused preschool education intelligent interaction system

The intelligent preschool education system, which integrates image big data with voice interaction, solves the problems of data fragmentation and poor adaptability, realizes personalized and fun learning experiences and security control, and improves the interactive effect of preschool education and parental participation.

CN121838577APending Publication Date: 2026-04-10HARBIN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
HARBIN UNIV
Filing Date
2026-01-13
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

Existing intelligent systems for preschool education suffer from problems such as data fragmentation, poor adaptability, unnatural interaction, and insufficient security control, making it difficult to meet the educational needs of preschool children.

Method used

By deeply integrating image big data with voice interaction, we collect children's image, voice, and behavioral data, and use deep learning algorithms and children's voice recognition models to build personalized learning profiles, achieve dynamic content adaptation and security control, and provide multi-dimensional feedback and parent interaction.

Benefits of technology

It enables comprehensive perception and accurate judgment of children's learning status, provides personalized and engaging interactive content, improves learning efficiency and safety, promotes parental involvement, and forms a closed-loop education model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121838577A_ABST
    Figure CN121838577A_ABST
Patent Text Reader

Abstract

The invention discloses an image big data and voice interaction fused preschool education intelligent interaction system, and relates to the technical field of preschool education intelligent interaction. The data acquisition module acquires image, voice and behavior data of children and integrates static education resources; the image processing module is used for extracting emotion and action characteristics of children and optimizing images; the voice interaction module optimizes voice recognition and semantic understanding of children and supports multiple rounds of dialogues; the intelligent adaptation module constructs a personalized learning portrait and dynamically adjusts the content; the content generation module fuses the data and generates multi-type interaction content; the interactive feedback module feeds back the learning effect in real time and provides suggestions; and the security management and control module filters bad information to guarantee use security and data privacy. According to the method, the pain point of traditional preschool education is broken through, accurate personalized interaction is realized through deep fusion of images and voice, the voice interaction fluency and the learning interest of children are improved, safety management and control and parent collaboration are enhanced, and preschool education is more targeted, interesting and safe.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of preschool education intelligent interaction, and particularly relates to a preschool education intelligent interaction system integrating image big data and voice interaction. BACKGROUND

[0002] Preschool education is a key stage for children's cognitive development, habit formation and ability cultivation. The core demand is to stimulate children's interest in learning through interesting and personalized interactive methods, and to balance education and experience. With the penetration of digital technology in the field of education, preschool education products have gradually transformed from traditional picture books and toys to electronic and intelligent ones. Learning devices equipped with image recognition and voice interaction functions have emerged, trying to break the limitations of traditional one-way indoctrination. Early products mainly focus on single functions, such as image recognition for painting comments or voice interaction for content on-demand. The data collection dimension is single, and it is impossible to form a comprehensive perception of children's learning state. In recent years, integrated products have begun to emerge, but most of them still remain at the level of functional superposition and fail to achieve deep integration of image big data and voice interaction, making it difficult to accurately capture children's behavior characteristics and real needs.

[0003] The existing preschool education intelligent system has many outstanding shortcomings. The data layer is obviously fragmented, with image collection focusing on single action recognition and voice interaction only meeting basic instruction response. The data of the two are not effectively linked, resulting in one-sided judgment of children's cognitive level and interest preferences, and inability to build a complete learning profile. In terms of adaptability, most products use a unified content push mode, ignoring individual differences in age, cognitive ability and interest points of preschool children. The difficulty of the content either exceeds the children's acceptance range causing frustration or is too simple leading to attention loss. The voice interaction technology lacks targeting, and existing models are mostly based on adult voice training, with low adaptation to issues such as unclear pronunciation, inaccurate tone and dialect accent caused by the immature development of children's vocal organs, low recognition accuracy and semantic understanding deviation, making it difficult to achieve natural and smooth multi-round dialogue and affecting the interactive experience.

[0004] There are obvious defects in the safety control and feedback mechanism. Some products lack strict content review mechanism, may contain inappropriate information for preschool children, and lack effective control over children's use time and eye health. In addition, the existing system lacks closed-loop feedback and parent linkage mechanism, the learning effect evaluation is superficial, and it cannot accurately locate the weak links of children's knowledge and cannot synchronize the learning situation to parents in time, resulting in low participation of parents in the education process, and it is difficult to form a collaborative mode of "system guidance-children learning-parent supervision". These problems together make it difficult for existing products to balance education, interest and individualization, and cannot fully meet the development needs of preschool children, restricting the deep application of intelligent technology in the field of preschool education, and an intelligent interactive system that can realize deep fusion of image and voice data, accurate adaptation, safe control and closed-loop linkage is urgently needed. SUMMARY

[0005] The image big data and voice interaction fusion preschool education intelligent interactive system provided by the present application solves the problems mentioned in the prior art.

[0006] In order to achieve the above purpose, the present application adopts the following technical scheme: an image big data and voice interaction fusion preschool education intelligent interactive system, comprising: A data acquisition module acquires children's image, voice and behavior data, and uploads them in real time through wireless transmission, and integrates static content data of a preschool education resource library; An image processing module uses a deep learning algorithm to extract children's facial key points, emotional features and action posture features, separates the subject and background through image segmentation technology, identifies concentration, emotional state and action standardization, processes fuzzy images and extracts effective features; A voice interaction module uses a children's voice recognition special model to process voice data, optimizes recognition bias, analyzes children's demand expression and learning questions through semantic understanding algorithm, and generates voice response using child-friendly voice synthesis technology; An intelligent adaptation module analyzes children's age, cognitive level, learning progress and interest preference based on output data, constructs a personalized learning portrait of children in combination with the requirements of the preschool education outline, and dynamically adjusts the difficulty, type and presentation form of interactive content; A content generation module fuses image features and voice demand data to generate interactive preschool education content; An interactive feedback module analyzes children's learning process data in real time, gives immediate feedback through various forms, generates a learning effect visualization report, and provides targeted learning suggestions; A safety control module establishes a multi-level content review mechanism to filter bad information, sets a learning time threshold to send eye protection reminders at regular intervals, limits connection with unfamiliar devices, records children's use trajectory and stores it in encrypted form.

[0007] Furthermore, it also includes a module for precise assessment of children's cognitive levels, which calculates a quantitative value of cognitive level through weighted fusion of multi-dimensional data. The calculation expression is as follows: ,in Quantifying children's cognitive level Age-weighted coefficient This represents the basic cognitive score corresponding to a child's physiological age. For learning performance weighting coefficients, The score is a combination of the degree of knowledge mastery and learning efficiency. Here, F represents the weighting coefficient for interactive feedback, and F is the score for interaction participation and feedback quality. For interest preference weighting coefficients, A score is assigned based on the match between click frequency and dwell time.

[0008] Furthermore, it also includes a data purification and optimization module, which uses a multi-level filtering algorithm to process image, voice, and behavioral data. The first level filters invalid data, the second level removes abnormal data, and the third level corrects interfering data. Missing data is supplemented using interpolation methods, and data format and magnitude are unified through data standardization processing.

[0009] Furthermore, it also includes a dynamic calculation module for interactive content suitability, which adjusts the content's suitability for children in real time using a quantitative model. The calculation expression is as follows: ,in To adapt metrics to content, To adjust the weighting coefficient according to the content difficulty, The difficulty level of the interactive content is scored. This is the correction factor for cognitive bias. Quantifying children's cognitive level Score the cognitive level of the content objectives. Match weight coefficients to interactive needs. A score is given for matching content type with children's interests and preferences.

[0010] Furthermore, it also includes a personalized learning path planning module. Based on children's cognitive level assessment results, learning performance data, and interest preference profiles, combined with the preschool education stage goals and knowledge point association map, it uses a path optimization algorithm to generate customized learning paths, track learning progress and knowledge point mastery in real time, and dynamically adjust path nodes and content arrangements.

[0011] Furthermore, it also includes a voice interaction optimization module, which constructs a children's pronunciation deviation correction model, identifies pronunciation error types through voice feature comparison and analysis, generates targeted pronunciation training content, and optimizes the acoustic feature extraction algorithm of the voice recognition model.

[0012] Furthermore, it also includes an image behavior deep analysis module, which tracks the trajectory of children's movement changes by comparing consecutive frame images, uses attention detection algorithms to analyze children's gaze focus and attention duration, identifies poor learning states, and combines voice interaction data to determine the reasons for children's emotional fluctuations in learning and generate state adjustment suggestions.

[0013] Furthermore, it also includes a parent-side linkage module, which synchronizes system data with parents in real time, pushes children's learning reports, usage time statistics, ability development assessment results, sets learning plans, filters interactive content, enables parent-child interaction mode, and provides family education guidance resources and communication skills suggestions.

[0014] Furthermore, it also includes a security control upgrade module, establishing a database of harmful information features, using a combination of image recognition and voice semantic analysis to review content; setting up a hierarchical permission management mechanism, regularly auditing data security, scanning for system vulnerabilities, and encrypting transmission and storage.

[0015] Furthermore, it also includes a system iteration and optimization module, which collects children's interactive feedback, parents' usage evaluations, and operational data of each module. It uses machine learning algorithms to analyze system performance shortcomings and insufficient content adaptation, optimizes image recognition models, voice interaction algorithms, and cognitive assessment formula parameters, updates the content of the preschool education resource library, and upgrades intelligent adaptation strategies and interactive feedback mechanisms.

[0016] Compared with existing technologies, the beneficial effects of this invention are: The intelligent interactive system for preschool education that integrates image big data and voice interaction in this invention fundamentally solves the core pain points of existing products, such as data fragmentation, poor adaptability, unnatural interaction, and insufficient security control. It provides preschool children with an intelligent learning experience that is educational, fun, and personalized, and promotes the upgrading of intelligent interactive technology in preschool education.

[0017] The system achieves comprehensive collection and deep integration of image, voice, and behavioral data through its data acquisition module, breaking through the limitations of traditional products' single and fragmented data. The image processing module accurately extracts features such as children's facial expressions and postures, while the voice interaction module optimizes children's speech recognition and semantic understanding abilities. The collaborative data from both modules supports intelligent adaptation and content generation, making the analysis of children's learning status more comprehensive and the judgment more accurate, avoiding decision-making biases caused by incomplete data. This multi-dimensional data fusion model provides a solid data foundation for personalized learning, making educational interactions more targeted.

[0018] The collaborative operation of the intelligent adaptation and content generation modules completely changes the traditional "one-size-fits-all" content push model. The system constructs personalized learning profiles based on children's age, cognitive level, and interests, and adjusts the difficulty, type, and presentation of content based on dynamically calculated content suitability. This ensures the content is both appropriate for children's current developmental level and moderately challenging, effectively stimulating their learning interest and initiative. The generated interactive content covers various types, including audiovisual, hands-on, and thinking training. Through storytelling and gamification, knowledge delivery is integrated into fun interactions, allowing children to accumulate knowledge and improve their abilities in a relaxed and enjoyable atmosphere.

[0019] Optimization of voice interaction and image behavior analysis significantly improves the naturalness and effectiveness of interaction. The voice recognition and correction model, optimized for the pronunciation characteristics of preschool children, reduces recognition errors caused by unclear pronunciation and accents, supporting smooth multi-turn dialogues and allowing children to freely express their needs and ask questions through voice. The image behavior deep analysis module tracks children's movements and concentration in real time, promptly identifying negative learning emotions and adjusting them through content switching and engaging guidance to help children maintain a good learning state and improve learning efficiency and experience.

[0020] A comprehensive security control mechanism safeguards children's use. Multi-level content review and a dynamically updated database of harmful information characteristics accurately filter out harmful content; learning time thresholds and eye protection reminders guide children to use the device appropriately; encrypted transmission and storage technology and tiered access control comprehensively protect children's personal information and learning data security, eliminating parents' concerns about usage safety.

[0021] The parent-facing interface and iterative system optimization form a continuously improving educational ecosystem. The parent-facing interface synchronizes children's learning reports and ability assessment results in real time, supporting parental participation in learning plan development and parent-child interaction. This constructs a collaborative education model of "system guidance + parental participation," promoting children's all-round development. The system iteration and optimization module continuously optimizes algorithm models, updates educational resources, and upgrades interactive strategies by collecting multi-dimensional feedback data, ensuring the system's functionality and performance are constantly improved and adaptable to the dynamic changes in children's growth and preschool education needs. Overall, the system achieves a comprehensive upgrade of intelligent interaction in preschool education through deep integration of images and voice, personalized adaptation, security control, closed-loop feedback, and collaborative linkage. This provides children with a higher-quality, safer, and more targeted learning experience, providing strong support for the digital transformation of preschool education. Attached Figure Description

[0022] Figure 1 This is a schematic block diagram of a preschool education intelligent interactive system that integrates image big data and voice interaction, as proposed in this invention. Figure 2 A diagram showing the comparison of the accuracy of voice interaction recognition for children; Figure 3 A diagram illustrating the trend of children's cognitive development; Figure 4 A schematic diagram illustrating the comprehensive performance evaluation of a preschool education intelligent system; Figure 5 A diagram showing the comparison of children's attention span in different interactive scenarios; Figure 6 This diagram illustrates the correlation between content suitability and children's learning interests. Detailed Implementation

[0023] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0024] In the description of this invention, it should be understood that the terms "center," "longitudinal," "lateral," "length," "width," "thickness," "upper," "lower," "front," "rear," "left," "right," "vertical," "horizontal," "top," "bottom," "inner," "outer," "clockwise," and "counterclockwise," etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are only for the convenience of describing this invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on this invention.

[0025] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include one or more of the stated features. In the description of this invention, "a plurality of" means two or more, unless otherwise explicitly specified. Furthermore, the terms "installed," "connected," and "linked" should be interpreted broadly; for example, they may refer to a fixed connection, a detachable connection, or an integral connection; they may refer to a mechanical connection or an electrical connection; they may refer to a direct connection or an indirect connection through an intermediate medium; and they may refer to the internal connection of two components. Those skilled in the art can understand the specific meaning of the above terms in this invention based on the specific circumstances. The invention will now be described in further detail with reference to the accompanying drawings.

[0026] Reference Figures 1 to 6 A preschool education intelligent interactive system integrating image big data and voice interaction, comprising: The data acquisition module uses a high-definition camera to capture children's facial expressions, movement trajectories, and interactive scene images; a high-fidelity microphone to capture children's pronunciation, expression of needs, and interactive voice; and touch device sensors to collect behavioral data such as learning time, operation frequency, and click preferences. It uses wireless transmission technology to achieve real-time data upload and integrates static content data such as picture books, nursery rhymes, animations, and games from the preschool education resource library, covering data acquisition needs for multiple scenarios such as cognitive learning, language expression, artistic creation, and parent-child interaction. The image processing module uses deep learning algorithms to extract key facial features, emotional characteristics, and action posture features of children in images. It separates the child subject from the background in the interactive scene through image segmentation technology, identifies the child's focus, emotional state, and the standardization of actions, performs noise reduction and enhancement processing on blurry images, and accurately extracts effective image features related to learning interaction, providing image data support for subsequent intelligent adaptation. The voice interaction module uses a special children's voice recognition model to process the collected voice data, optimizes the recognition deviation caused by problems such as inaccurate pronunciation of dialects, analyzes children's needs and learning questions through semantic understanding algorithms, and generates friendly and natural voice responses using child-friendly voice synthesis technology. It supports multi-turn dialogue interaction and realizes functions such as voice command control, content on demand, and learning Q&A. The intelligent adaptation module analyzes children's age, cognitive level, learning progress, and interest preferences based on the output data of the data acquisition module, image processing module, and voice interaction module. Combined with the requirements of the preschool education syllabus, it constructs a personalized learning profile for children and dynamically adjusts the difficulty, type, and presentation format of interactive content to make the content fit the developmental needs of children. The content generation module integrates image features and voice demand data to generate interactive preschool education content, including interactive stories that adjust the plot based on children's action feedback, pinyin nursery rhymes that correct pronunciation accuracy in real time, art creation tools that provide intelligent comments based on drawing operations, and fun question-and-answer games generated based on voice commands. The content formats cover a variety of types, including audiovisual interaction, hands-on operation, and thinking training. The interactive feedback module analyzes children's image, voice, and behavioral data in real time during the learning process, and provides immediate feedback through voice encouragement, animated rewards, and points accumulation. It generates a visual report on learning effectiveness, clarifies the mastery of knowledge points and weak areas, provides targeted learning suggestions for children, and enhances their learning enthusiasm and sense of participation. The safety management module establishes a multi-level content review mechanism to filter harmful information such as violence, vulgarity, and superstition; sets learning time thresholds and sends timed reminders for eye protection and rest; restricts connections from unfamiliar devices to ensure data transmission security; records children's usage patterns and stores them in encrypted form to prevent personal information leakage and ensure the safety of preschool children.

[0027] This invention also includes a module for accurately assessing children's cognitive levels, which calculates a quantitative value of cognitive level through weighted fusion of multi-dimensional data. The calculation expression is as follows: ,in The quantitative value for children's cognitive level ranges from 0 to 100. The age weighting coefficient ranges from 0.2 to 0.3. The basic cognitive score for children corresponding to their physiological age ranges from 0 to 100. The weighting coefficient for learning performance is set between 0.3 and 0.4. The combined score for knowledge mastery and learning efficiency ranges from 0 to 100. The weighting coefficient for interactive feedback ranges from 0.2 to 0.25. The scores for interaction engagement and feedback quality range from 0 to 100. The interest preference weighting coefficient ranges from 0.1 to 0.2. The score range for matching click frequency and dwell time is 0 to 100. The weighted fusion of multi-dimensional data enables accurate quantification of cognitive level, providing a scientific basis for decision-making by the intelligent adaptation module.

[0028] This invention also includes a data purification and optimization module, which uses a multi-level filtering algorithm to process the collected image, voice, and behavioral data. The first level filters invalid data caused by equipment failure, the second level removes abnormal data caused by environmental noise interference, and the third level corrects interference data caused by children's misoperation and random pronunciation. Missing data is supplemented using an interpolation method based on the behavioral patterns of preschool children. Data standardization processing unifies the data format and magnitude, improves data quality and the accuracy of subsequent analysis, and provides reliable data support for the efficient operation of each module.

[0029] This invention also includes a dynamic calculation module for interactive content suitability, which adjusts the suitability of content for children through a quantitative model. The calculation expression is as follows: ,in The metric for content adaptation ranges from 0 to 100. The weighting coefficient for content difficulty is set between 0.3 and 0.4. The difficulty level score for interactive content ranges from 0 to 100. The correction factor for cognitive bias ranges from 0.25 to 0.35. The quantitative value for children's cognitive level ranges from 0 to 100. The score for the cognitive level of the content objective ranges from 0 to 100. The weighting coefficient for matching interactive needs ranges from 0.25 to 0.35. The matching score between content type and children's interests ranges from 0 to 100. The dynamic calculation results guide the content generation module to adjust the content parameters so that the interactive content is both suitable for cognitive level and has a moderate challenge.

[0030] This invention also includes a personalized learning path planning module. Based on children's cognitive level assessment results, learning performance data, and interest preference profiles, combined with the goals of preschool education and the knowledge point association map, a path optimization algorithm is used to generate a customized learning path. The path includes the learning order of core knowledge points, the recommendation of supporting interactive content, and the ability improvement training program. The learning progress and knowledge point mastery are tracked in real time, and the path nodes and content arrangements are dynamically adjusted to make the learning process gradual and highlight the key points, helping children systematically build a knowledge system.

[0031] This invention also includes a voice interaction optimization module, which addresses the problems of unclear pronunciation and inaccurate tones caused by the immature development of preschool children's speech organs. It constructs a children's pronunciation deviation correction model, identifies the types of pronunciation errors through voice feature comparison and analysis, generates targeted pronunciation training content, and helps children standardize their pronunciation through methods such as follow-up reading comparison and fun pronunciation correction. At the same time, it optimizes the acoustic feature extraction algorithm of the voice recognition model, expands the recognition coverage of children's pronunciation variations, and improves the success rate of voice interaction in complex environments.

[0032] This invention also includes an image behavior depth analysis module, which tracks the trajectory of children's movement changes by comparing consecutive frame images, uses an attention detection algorithm to analyze children's gaze focus and attention duration, identifies poor learning states such as inattentiveness, irritability, and fatigue, combines voice interaction data to determine the reasons for children's emotional fluctuations in learning, generates state adjustment suggestions, and helps children maintain a good learning state and improve learning efficiency and experience through interactive content switching, rest reminders, and fun guidance.

[0033] This invention also includes a parent-side linkage module, which uses a dedicated application to achieve real-time data synchronization between the system and the parent's end. It pushes children's learning reports, usage time statistics, and ability development assessment results to parents, supports parents in setting learning plans, filtering interactive content, and enabling parent-child interaction modes, provides family education guidance resources and communication skills suggestions, and builds a collaborative education model that combines system guidance with parental participation to promote the all-round development of preschool children.

[0034] This invention also includes a security control upgrade module, which establishes a dynamically updated database of harmful information features, strengthens content review by combining image recognition and voice semantic analysis, accurately filters harmful content such as sensitive words, inappropriate images, and dangerous behavior guidance, sets up a hierarchical permission management mechanism to restrict children's access to system settings, data storage and other functions, conducts regular data security audits and system vulnerability scans, and uses encrypted transmission and storage technologies to ensure the security and privacy of children's personal information and learning data.

[0035] This invention also includes a system iteration and optimization module, which collects operational data from each module, children's interactive feedback data, and parents' usage evaluation data. It uses machine learning algorithms to analyze system performance shortcomings and insufficient content adaptation, optimizes image recognition models, voice interaction algorithms, and cognitive assessment formula parameters, updates the content of the preschool education resource library, and upgrades intelligent adaptation strategies and interactive feedback mechanisms to achieve continuous iteration of system functions and performance, and continuously improves the interactive experience and teaching effectiveness of preschool education.

[0036] The following two examples further illustrate specific embodiments of the present invention: Example 1: Application of interactive scenarios in kindergarten senior class language cognition classroom This embodiment is applied to a language cognition course for senior kindergarten students, targeting 30 children aged 5-6. The core objective of the course is to improve children's pinyin recognition, vocabulary accumulation, and language expression abilities through intelligent interaction. At the same time, it helps teachers to monitor each child's learning status in real time, enabling personalized teaching guidance and solving problems such as uneven interaction, untimely feedback, and insufficient content suitability in traditional classrooms.

[0037] The data acquisition module comprehensively covers all dimensions of classroom interaction data. High-definition cameras deployed at the front and sides of the classroom capture real-time image data of children's facial expressions, hand gestures, posture, and gestures for recognizing pinyin cards, with a sampling frequency of 15 frames per second to ensure continuous motion tracking. High-fidelity microphones are evenly distributed throughout the classroom to collect children's pinyin pronunciation, answers to questions, and group discussions, with a sampling frequency of 16kHz to effectively filter out ambient noise. Touchscreen sensors record children's frequency of clicking pinyin cards, dragging word combinations, completing interactive quizzes, learning duration, and incorrect click locations. Wireless transmission technology synchronously uploads real-time data to the system platform, while also integrating static content data from the preschool education resource library related to language cognition in older children, such as pinyin picture books, fun nursery rhymes, and vocabulary games, covering multiple learning scenarios including pinyin recognition, word matching, sentence creation, and storytelling.

[0038] The data purification and optimization module performs precise processing on the collected data. The first level filters invalid data caused by camera obstruction or microphone malfunction; the second level removes abnormal data caused by external classroom noise and interference from other children's speech; and the third level corrects interference data caused by children accidentally touching the tablet or making random sounds. For missing individual action frame data, interpolation methods based on the behavioral patterns of preschool children in the classroom are used to supplement them. Data standardization processing unifies image pixel format, audio format, and behavioral data volume, improving data quality and the accuracy of subsequent analysis, providing reliable support for the efficient operation of each module.

[0039] The image processing module performs in-depth analysis of the acquired image data. Deep learning algorithms are used to extract key facial points from children, including eye opening and closing, mouth shape, and facial muscle state, identifying emotional characteristics such as happiness, focus, irritability, and fatigue. By comparing consecutive frames, the module tracks the trajectories of actions such as raising hands, sitting posture, and card display, analyzing the standardization of these actions and the level of interactive participation, for example, judging whether the hand-raising action is standard and the sitting posture is correct. Image segmentation technology is used to separate the child from the classroom background, eliminating environmental interference such as desks, chairs, and blackboards. Noise reduction and enhancement are applied to blurry images caused by insufficient lighting, accurately extracting effective image features related to language learning interaction, providing data support for intelligent adaptation.

[0040] The voice interaction module is optimized for classroom language learning needs. It employs a dedicated children's speech recognition model to process speech data such as pinyin repetition and vocabulary reading, expanding the recognition coverage of children's pronunciation variations and optimizing recognition errors caused by issues such as difficulty distinguishing between retroflex and alveolar consonants and confusion between front and back nasal sounds in children from southern China. Semantic understanding algorithms analyze the speech content of children's answers and learning questions, such as recognizing children asking about the pronunciation of a pinyin or the meaning of a word. Child-friendly speech synthesis technology generates friendly and natural voice responses with lively intonation and moderate speed, supporting multi-turn dialogue interaction. This enables voice command control of pinyin card display, content playback, and learning Q&A functions, improving the smoothness of classroom interaction.

[0041] The precise assessment module for children's cognitive levels quantifies children's language cognitive abilities. It uses a formula... ,in We set the age weighting coefficient to 0.25. The basic cognitive score corresponds to the child's physiological age; for a 5.5-year-old child, the A score is set to 85. We set 0.35 as the weighting coefficient for learning performance. A child's pinyin recognition accuracy, vocabulary mastery, and learning efficiency are combined to form a score. The child's pinyin recognition accuracy is 88%, vocabulary mastery is 82%, and learning efficiency is 85%. The overall score is calculated as follows: Take 85; We set 0.22 as the weighting coefficient for interactive feedback. Based on scores for classroom participation and feedback quality, including the number of times a child raised their hand and the quality of their answers, this child actively raised their hand 6 times and answered questions accurately 4 times. Take 90; We set 0.18 as the weighting coefficient for interest preferences. To score the child based on the frequency of clicks and the duration of time spent on content such as pinyin nursery rhymes and vocabulary games, the child showed a preference for pinyin nursery rhymes. Take 88. Calculate... =0.25×85+0.35×85+0.22×90+0.18×88=86.64, which corresponds to a relatively high cognitive level and is suitable for advanced language cognition content.

[0042] The intelligent adaptation module and content generation module work together. Based on cognitive assessment results, image processing, and voice interaction data, a personalized learning profile is constructed for the child: 5.5 years old, solid foundation in pinyin cognition, moderate vocabulary accumulation, preference for audio content, and active classroom participation. In accordance with the preschool education syllabus's requirements for language cognition in senior kindergarten classes, the difficulty of interactive content is dynamically adjusted, upgrading from basic pinyin recognition to pinyin word formation exercises. The content generation module integrates image and voice data to generate interactive pinyin nursery rhymes. When children repeat after the rhymes, the system identifies pronunciation errors through voice feature comparison and analysis, corrects errors in real time, and demonstrates standard pronunciation. It also generates word-matching interactive games, adjusting the game difficulty based on the child's dragging and dropping actions; new words are unlocked after three correct matches. Finally, it generates sentence creation interactive tasks, guiding children to create simple sentences through voice or text input, with the system providing intelligent feedback.

[0043] The interactive feedback module provides real-time output of learning outcomes. It offers immediate feedback through voice encouragement (e.g., accurate pinyin pronunciation) and animated rewards (e.g., accumulating stars and redeeming class badges with points). After the lesson, a visual report on learning outcomes is generated, clearly showing the child's mastery of knowledge points such as pinyin recognition accuracy, vocabulary collocation correctness, and sentence creation completeness. It also identifies weaknesses such as confusion between nasal and non-nasal sounds and provides targeted learning suggestions, such as listening to and practicing rhymes comparing nasal and non-nasal sounds. The report is also synchronized to the teacher's end, helping them accurately understand each child's learning status and adjust subsequent teaching plans accordingly.

[0044] The security management module ensures safe classroom use. A multi-level content review mechanism filters out inappropriate vocabulary and images for children, sets a 40-minute threshold for classroom learning time, and sends eye-protection rest reminders every 20 minutes, guiding children to look into the distance and relax. It restricts unfamiliar devices from connecting to the classroom system, encrypts and stores children's learning data to prevent personal information leaks, and ensures safe classroom use. The parent-side integration module pushes children's classroom learning reports, participation statistics, and ability development assessment results to parents through a dedicated application. It also provides family-based pinyin practice resources and parent-child interaction suggestions, invites parents to participate in children's after-school language reinforcement, and builds a collaborative education model between the classroom and home.

[0045] Table 1: Comparison of Interactive Effects in Language Cognition Classrooms of Kindergarten Senior Classes Evaluation index Traditional classroom teaching performance Invention system performance Classroom interaction participation Unbalanced Balanced Pinyin recognition accuracy General Excellent Voice command recognition success rate Lower Higher Personalized guidance coverage Limited Comprehensive Parental participation synergy Lower Higher Children's interest in learning General Strong Table 1 clearly demonstrates the significant advantages of this invention's system in kindergarten classroom settings. Traditional classroom teaching is limited by teacher energy, resulting in uneven participation, some children lacking opportunities to express themselves, limited personalized guidance coverage, and difficulties for parents in monitoring classroom learning. This invention's system, through comprehensive data collection and deep integration, achieves accurate cognitive assessment and personalized content adaptation for each child, ensuring classroom interaction covers all children. The optimized voice interaction module improves recognition success rate, while engaging content generation stimulates learning interest, significantly improving pinyin recognition accuracy. The parent-side linkage module bridges the gap between classroom and home education, enhancing parental involvement and collaboration, forming a comprehensive language cognition development system that effectively compensates for the shortcomings of traditional classrooms.

[0046] Example 2: Application of parent-child interaction scenarios in families with children aged 3-4 This embodiment is applied to parent-child interaction in families with children aged 3-4. The core requirement is to assist parents in carrying out scientific preschool education through intelligent interaction, covering areas such as language expression, artistic perception, and cognitive exploration. It solves problems such as parents' lack of professional educational resources, monotonous interaction methods, and inability to accurately grasp the child's developmental level, and achieves a unity of educational and fun in parent-child interaction.

[0047] The data acquisition module is adapted to the characteristics of family interaction scenarios. A high-definition camera deployed in the living room captures image data such as facial expressions during parent-child reading, gestures during craft activities, and body language during interactive games, with a sampling frequency of 10 frames per second, balancing data accuracy and device power consumption. A high-fidelity microphone captures children's voices while reading picture books, asking questions, giving game commands, and parents' guiding voices, with a sampling frequency of 16kHz, optimizing for interference filtering from TV sound and outside noise in the home environment. Touchscreen sensors record children's learning time, frequency of operation, and preferred content types when clicking on picture book animations, playing cognitive games, and coloring activities. Wireless transmission technology enables real-time data upload, integrating static content data from the preschool education resource library for 3-4 year olds, including parent-child picture books, nursery rhymes, craft tutorials, and cognitive games, covering multiple scenarios such as language development, color recognition, hands-on skills, and parent-child collaboration.

[0048] The data purification and optimization module specifically processes data from home scenarios. The first level filters out invalid data caused by improper camera angles or microphones being too far away. The second level removes abnormal data caused by sudden noises or pet interference in the home environment. The third level corrects interference data caused by children accidentally touching devices or making unintentional sounds. For missing brief action or voice data, an interpolation method based on the behavioral patterns of preschool children in their homes is used to supplement it. Data standardization ensures a unified data format and volume, guaranteeing the accuracy and reliability of subsequent analysis.

[0049] The image processing module accurately analyzes family interaction image data. It uses deep learning algorithms to extract emotional features from children's faces, such as smiles and frowns, to determine their level of liking for the interactive content. By comparing consecutive frames, it tracks children's movements such as origami, coloring, and building blocks, analyzing their motor coordination and fine motor skills. Image segmentation technology separates children, parents, and the family background, focusing on core parent-child interactions and extracting relevant features for educational interaction, providing image support for intelligent adaptation.

[0050] The voice interaction module optimizes the family-child dialogue experience. Addressing the characteristics of 3-4 year old children's immature pronunciation and unclear articulation, a pronunciation deviation correction model is constructed, focusing on optimizing the recognition of easily confused initials such as b / p and m / n. Through voice feature comparison and analysis, pronunciation error types are identified, generating targeted pronunciation training content. Methods such as follow-along reading comparison and fun pronunciation correction help children standardize their pronunciation. Semantic understanding algorithms accurately interpret children's simple questions, such as "What color is this?", and expressions of need, such as "I want to hear a nursery rhyme." Using child-friendly voice synthesis technology, friendly responses are generated, supporting multi-turn parent-child interactive dialogues and enabling voice command control of picture book playback, content switching, game launch, and other functions.

[0051] The interactive content adaptation dynamic calculation module adjusts parent-child interactive content in real time. It uses a formula... ,in We set 0.35 as the content difficulty adaptation weight coefficient. Based on the difficulty level score of the interactive content, the parent-child picture book "The Kingdom of Colors" is suitable for children aged 3-4. Take 80; We set 0.3 as the cognitive bias correction factor. To quantify children's cognitive level, a multi-dimensional assessment is conducted. =75; Based on the cognitive level of the content objectives, the goal of this picture book is to improve color recognition. Take 78; We set 0.35 as the weighting coefficient for matching interactive needs. A score was assigned to match content type with children's interests and preferences; children preferred color-related content. Take 90. Calculate... =0.35×80-0.3×|75-78|+0.35×90=58.6, the adaptation is good, the system maintains the interactive form of the picture book, while fine-tuning the content details and adding more color Q&A interaction.

[0052] The intelligent adaptation module constructs a personalized learning profile for each child. Based on collected image, voice, and behavioral data, it analyzes children aged 3.5 years, with strong color recognition abilities, early language development, moderate fine motor skills, and a preference for interactive content combining visual and auditory elements. In line with the requirements of the early childhood education syllabus, the presentation format of interactive content is dynamically adjusted. For example, parent-child picture books are designed as animated videos with voice prompts, craft tutorials include step-by-step voice guidance, and cognitive games feature simple and easy-to-use touch interactions.

[0053] The content generation module integrates multi-dimensional data to generate interactive parent-child content. An interactive coloring tool adjusts color matching suggestions based on children's coloring actions; when a child chooses red to color an apple, the system provides a voice prompt suggesting that red and green look better together. A parent-child Q&A game is generated based on voice commands; for example, if a parent asks for round objects, the system guides the child to click on round objects in the picture and provides voice encouragement. A story continuation task is generated based on parent-child reading interactions; after a picture book story ends, the system invites children to continue the ending using voice or simple actions, providing intelligent feedback and adding interesting details.

[0054] The interactive feedback module enhances parent-child interaction in real time. It provides immediate feedback through voice praise (e.g., "Baby's coloring is so even!"), animated rewards (e.g., "Flowers bloom!"), and points accumulation for exclusive parent-child interaction privileges. After each interaction, a visual report on learning outcomes is generated, clearly showing the child's development in color recognition, language expression, and fine motor skills. It also identifies weaknesses in language expression and provides targeted suggestions for parent-child interaction, such as more story retelling practice.

[0055] The enhanced security management module ensures safe use in the home. It establishes a dynamically updated database of harmful information characteristics and employs a combination of image recognition and voice semantic analysis to strengthen content review, accurately filtering sensitive words and inappropriate images. A daily study time threshold of 30 minutes is set, and an eye-protection rest reminder is sent every 10 minutes to guide children to look into the distance or perform eye exercises. Children's access to system settings and data storage functions is restricted, and encrypted transmission and storage technologies are used to protect children's personal information and learning data. Regular data security audits are conducted. The parent-child interaction module allows parents to set learning plans, filter interactive content, and enable parent-child interaction modes. It provides family education guidance resources and communication skills suggestions to help parents improve the professionalism of parent-child interactions.

[0056] The system iteration and optimization module collects family interaction data, children's feedback, and parents' evaluations. It uses machine learning algorithms to analyze system performance shortcomings and insufficient content adaptation, optimizes the image recognition model's recognition accuracy for parent-child interaction actions, the voice interaction algorithm's adaptation to young children's pronunciation, and the formula parameters for cognitive assessment and content adaptation. It also updates the content of the preschool education resource library, upgrades the intelligent adaptation strategy and interactive feedback mechanism, and continuously improves the family parent-child interaction experience and educational effectiveness.

[0057] Table 2: Comparison of Parent-Child Interaction Effects in Families with Children Aged 3-4 Evaluation index Traditional parent-child interaction performance Invention system performance Interactive content education Insufficient Sufficient Children's participation concentration Lower Higher Parental guidance professionalism General Higher Children's ability to improve Slow Significant Interactive form richness Single Rich Data security level General Higher Table 2 highlights the core value of this invention's system in family-based parent-child interaction scenarios. Traditional parent-child interactions often rely on simple toys or picture books, lacking educational content, having limited interactive formats, and requiring parents to provide professional guidance, making it difficult to accurately promote children's ability development. This invention's system, through comprehensive data collection and deep integration, generates interactive content that is both educational and engaging, enriching parent-child interaction formats. Intelligent adaptation and dynamic calculation of content suitability ensure that the content aligns with children's developmental levels, improving engagement and focus. Professional guidance resources and interaction suggestions provided to parents enhance the professionalism of parental guidance, making parent-child interactions more targeted and resulting in significant improvement in children's abilities. A comprehensive security control mechanism safeguards data security and children's healthy use, eliminating parental concerns and providing scientific, efficient, and safe intelligent support for family-based preschool education.

[0058] Reference Figure 2This diagram visually demonstrates the core advantages of the system in children's speech recognition, addressing the pain point of traditional systems' poor adaptability to the speech of young children. Traditional preschool education systems' speech recognition models are mostly based on adult speech training, which is insufficiently adapted to the characteristics of 3-6 year old children's immature pronunciation, unclear articulation, and inaccurate tones. The recognition accuracy rate for 3-year-old children is only 62%, making natural interaction difficult. The system of this invention optimizes the recognition model for children's pronunciation characteristics, focusing on solving the problem of easily confused initials and finals, while expanding the coverage of children's pronunciation variations. The recognition accuracy rate exceeds 90% for all age groups, reaching 94% in mixed-age scenarios, significantly improving the fluency and accuracy of speech interaction, allowing children to freely express their needs through speech and truly achieve natural multi-turn dialogue interaction.

[0059] Reference Figure 3 The diagram clearly demonstrates the continuous improvement of children's cognitive abilities by the system of this invention, distinguishing it from traditional preschool education products that lack closed-loop optimization. In the initial stage, the quantified value of a child's cognitive level is only 65. Traditional products often use a fixed content delivery model, making it difficult to dynamically adjust according to the child's learning progress, resulting in slow improvement. The system of this invention uses a precise cognitive level assessment module to periodically quantify the child's cognitive status and dynamically calculates personalized learning content based on content suitability. After one month, the quantified value rises to 72, and subsequently, with continuous system adaptation and optimization, it reaches 91 after four months. This data-driven closed-loop learning model allows children's cognitive abilities to improve gradually, while precisely addressing knowledge gaps, achieving personalization and systematic learning in preschool education.

[0060] Reference Figure 4 This diagram comprehensively reflects the overall performance advantages of the system of this invention throughout the entire preschool education process, breaking through the predicament of traditional systems' weak single function and unbalanced overall performance. Traditional preschool education intelligent systems, due to technological limitations, generally score low in core dimensions, especially content adaptation capability (only 3 points) and parent interaction capability (only 2 points), failing to meet the personalized learning and home-school collaboration needs of preschool children. The system of this invention, through the deep integration of image big data and voice interaction, combined with dynamic adaptation, parent-side interaction, and multi-level security control mechanisms, achieves scores exceeding 9 points in all dimensions, with content adaptation capability reaching 9.6 points and voice recognition capability reaching 9.5 points. It realizes a leap in full-process capability from "perception-analysis-adaptation-interaction-control," fully adapting to the diversified needs of preschool education.

[0061] Reference Figure 5This diagram visually highlights the significant effect of the system of this invention in improving children's concentration during learning, addressing the pain points of traditional preschool education's monotonous interactive formats and lack of engagement. Traditional interactive modes often employ one-way instruction or simple operations, easily leading to boredom in children. For example, children's focus time for picture book reading is only 8 minutes, and for pinyin learning, it's a mere 5 minutes. The system of this invention integrates image and voice data to generate interactive content, such as picture books with storylines adjusted based on children's actions, pinyin nursery rhymes with real-time pronunciation correction, and intelligent feedback art creation tools. This integrates learning into fun and interactive activities, increasing focus time in each scenario to approximately 20 minutes, and reaching 28 minutes in parent-child game scenarios. This sustained state of focus allows children to absorb knowledge more deeply, significantly improving the effectiveness of preschool education.

[0062] Reference Figure 6 The figure clearly reveals the positive correlation between content suitability and children's learning interest, demonstrating the core value of the dynamic adaptation mechanism of this invention. When the content suitability is only 45, the child's learning interest score is only 4 points. At this point, the content difficulty deviates significantly from the child's cognitive level, easily leading to frustration. As the suitability increases to 55, the interest score rises to 6 points. With continued improvement in suitability, the interest score reaches the maximum of 10 points at 85. This invention's system, through an interactive content suitability dynamic calculation module, adjusts the content difficulty and type in real time, keeping the suitability consistently high and fundamentally stimulating children's learning initiative. This data-driven precise adaptation avoids the loss of interest caused by the "one-size-fits-all" content push of traditional systems, allowing children to maintain their desire to explore and their enthusiasm within appropriately matched learning content.

[0063] The above are merely preferred embodiments of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.

Claims

1. A preschool education intelligent interactive system integrating image big data and voice interaction, characterized in that, Includes the following modules: The data acquisition module collects children's image, voice, and behavioral data, uploads them in real time via wireless transmission, and integrates static content data from the preschool education resource database. The image processing module uses deep learning algorithms to extract key points, emotional features, and action posture features of children's faces. It separates the subject from the background through image segmentation technology, identifies focus, emotional state, and action standardization, reduces noise in blurred images, and extracts effective features. The voice interaction module uses a special model for children's voice recognition to process voice data, optimize recognition errors, analyze children's needs and learning questions through semantic understanding algorithms, and generate voice responses using child-friendly voice synthesis technology. The intelligent adaptation module analyzes children's age, cognitive level, learning progress, and interest preferences based on output data, and constructs a personalized learning profile for children in accordance with the requirements of the preschool education syllabus, dynamically adjusting the difficulty, type, and presentation format of interactive content; The content generation module integrates image features and voice demand data to generate interactive preschool education content; The interactive feedback module analyzes children's learning process data in real time, provides immediate feedback in various forms, generates visual reports on learning outcomes, and offers targeted learning suggestions. The security management module establishes a multi-level content review mechanism to filter inappropriate information, sets learning time thresholds to send timed reminders for eye protection and rest, restricts connections from unfamiliar devices, and records and encrypts children's usage patterns.

2. The preschool education intelligent interactive system integrating image big data and voice interaction according to claim 1, characterized in that, It also includes a module for precise assessment of children's cognitive levels, which calculates a quantitative value of cognitive level through weighted fusion of multi-dimensional data. The calculation expression is as follows: ,in Quantifying children's cognitive level Age-weighted coefficient This represents the basic cognitive score corresponding to a child's physiological age. For learning performance weighting coefficients, The score is a combination of the degree of knowledge mastery and learning efficiency. Here, F represents the weighting coefficient for interactive feedback, and F is the score for interaction participation and feedback quality. For interest preference weighting coefficients, A score is assigned based on the match between click frequency and dwell time.

3. The preschool education intelligent interactive system integrating image big data and voice interaction according to claim 1, characterized in that, It also includes a data purification and optimization module, which uses a multi-level filtering algorithm to process image, voice and behavioral data. The first level filters invalid data, the second level removes abnormal data, and the third level corrects interfering data. Missing data is supplemented by interpolation methods, and data format and magnitude are unified through data standardization processing.

4. The preschool education intelligent interactive system integrating image big data and voice interaction according to claim 1, characterized in that, It also includes a dynamic calculation module for interactive content suitability, which adjusts the content's suitability for children in real time using a quantitative model. The calculation expression is: ,in To adapt metrics to content, To adjust the weighting coefficient according to the content difficulty, The difficulty level of the interactive content is scored. This is the correction factor for cognitive bias. Quantifying children's cognitive level Score the cognitive level of the content objectives. Match weight coefficients to interactive needs. A score is given for matching content type with children's interests and preferences.

5. The preschool education intelligent interactive system integrating image big data and voice interaction according to claim 1, characterized in that, It also includes a personalized learning path planning module, which uses path optimization algorithms to generate customized learning paths based on children's cognitive level assessment results, learning performance data, interest and preference profiles, combined with the preschool education stage goals and knowledge point association map, and tracks learning progress and knowledge point mastery in real time, dynamically adjusting path nodes and content arrangements.

6. The preschool education intelligent interactive system integrating image big data and voice interaction according to claim 1, characterized in that, It also includes a voice interaction optimization module, which builds a children's pronunciation deviation correction model, identifies pronunciation error types through voice feature comparison and analysis, generates targeted pronunciation training content, and optimizes the acoustic feature extraction algorithm of the voice recognition model.

7. The preschool education intelligent interactive system integrating image big data and voice interaction according to claim 1, characterized in that, It also includes an image behavior deep analysis module, which tracks the trajectory of children's movement changes by comparing consecutive frame images, uses attention detection algorithms to analyze children's gaze focus and attention duration, identifies poor learning states, and combines voice interaction data to determine the reasons for children's emotional fluctuations in learning and generate state adjustment suggestions.

8. The preschool education intelligent interactive system integrating image big data and voice interaction according to claim 1, characterized in that, It also includes a parent-side interaction module that synchronizes system data with parents in real time, pushes children's learning reports, usage time statistics, ability development assessment results, sets learning plans, filters interactive content, enables parent-child interaction mode, and provides family education guidance resources and communication skills suggestions.

9. The intelligent interactive system for preschool education integrating image big data and voice interaction according to claim 1, characterized in that, It also includes a security control upgrade module, which establishes a database of harmful information features, uses a combination of image recognition and voice semantic analysis to review content, sets up a hierarchical permission management mechanism, regularly audits data security, scans for system vulnerabilities, and encrypts transmission and storage.

10. The intelligent interactive system for preschool education integrating image big data and voice interaction according to claim 1, characterized in that, It also includes a system iteration and optimization module, which collects children's interactive feedback, parents' evaluations, and operational data of each module. It uses machine learning algorithms to analyze system performance shortcomings and content adaptation deficiencies, optimizes image recognition models, voice interaction algorithms, and cognitive assessment formula parameters, updates the content of the preschool education resource library, and upgrades intelligent adaptation strategies and interactive feedback mechanisms.