Child learning adaptive interaction method combining software and hardware, electronic device and medium
By using multimodal data collection and analysis through the sensor panel and cloud service platform, the problem of disconnect between physical operation and digital content in children's learning products has been solved. This enables real-time assessment and personalized learning paths, improving learning interest and effectiveness, and supporting parent-child collaboration.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- ENCHENG CULTURE & EDUCATION TECHNOLOGY (DALIAN) CO LTD
- Filing Date
- 2026-03-17
- Publication Date
- 2026-06-09
AI Technical Summary
Existing smart learning products for children suffer from several problems in their hardware and software integration: a disconnect between physical operation and digital content, difficulties in synchronizing multimodal data, an inability to assess children's cognitive status in real time, and a lack of parent-child collaboration. These issues lead to decreased learning interest and inaccurate assessments.
By operating the sensor panel to identify physical markers in real time, combined with children's terminal devices and cloud service platforms, multimodal data can be collected and time-series aligned in real time. This allows for the fusion analysis of children's emotional state, learning intentions, and knowledge mastery, dynamically adjusting learning content and feedback to form an adaptive interactive closed loop.
It enables children to have an immersive and interactive learning experience, accurately assess their cognitive status, personalize their learning paths, improve their learning participation and effectiveness, and support parent-child collaborative learning.
Smart Images

Figure CN122175746A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the fields of educational artificial intelligence and multimodal human-computer interaction technology, specifically to a software and hardware integrated adaptive interaction method, electronic device, and medium for children's learning. Background Technology
[0002] With the development of artificial intelligence and the Internet of Things, early childhood education products are gradually evolving towards intelligence and interactivity. Children aged 3-8 are in a critical period of transition from concrete to abstract thinking, and their learning highly relies on hands-on activities, multi-sensory collaboration, and immediate feedback. However, existing smart learning products for children mainly suffer from the following two limitations in terms of technological implementation: One type is purely software-based educational applications, which primarily interact through touchscreens. While these products offer rich digital content, they lack hands-on experience and fail to meet children's physiological and psychological needs to build cognition through touching and manipulating physical objects. This can easily lead to distraction and decreased interest in learning.
[0003] Another category is physical hardware devices, such as reading pens, smart building blocks, and early education machines. While these products offer a tangible, hands-on experience, they mostly employ a one-way triggering mechanism, resulting in the following technical bottlenecks: physical operations and digital content are largely triggered unidirectionally, failing to perceive children's continuous operational behaviors in real time, leading to a disconnect between physical interaction and digital feedback; the hardware devices lack a collaborative mechanism with visual and voice acquisition units, resulting in clock deviations and inconsistent sampling frequencies, hindering accurate alignment of multimodal data over time; existing products mostly use general intelligent models, which are poor at recognizing children's unique emotional expressions and leaps in thought; assessment of knowledge mastery relies heavily on after-class tests, failing to achieve seamless, real-time evaluation during natural operation; learning content libraries are mostly statically preset, unable to dynamically adjust based on children's real-time status; and parents cannot obtain real-time information about their children's learning status, lacking an effective parent-child collaboration mechanism.
[0004] In summary, existing technologies have not yet formed a complete technical solution that can deeply integrate software and hardware to achieve multimodal perception and understanding, seamless assessment, dynamic content expansion, and parent-child collaboration. There is an urgent need for a software and hardware integrated adaptive interactive method and system for children's learning that can systematically solve the above problems. Summary of the Invention
[0005] The purpose of this disclosure is to provide a software and hardware integrated adaptive interactive method, electronic device and medium for children's learning, in order to solve the technical problems of deep coupling and real-time mapping of software and hardware, synchronization and accurate alignment of heterogeneous multimodal data, deep understanding of children's emotions and cognition, non-intrusive learning assessment and accurate diagnosis, dynamic expansion of personalized learning paths, and support for parent-child collaboration and home-school co-education.
[0006] To achieve the above objectives, in a first aspect, this disclosure provides a hardware-software integrated adaptive interactive method for children's learning. The adaptive interactive system for children's learning includes an operation sensing panel, multiple physical identifiers, a children's terminal device, and a cloud service platform. The operation sensing panel is communicatively connected to both the children's terminal device and the cloud service platform. The cloud service platform stores a learning content library, and each physical identifier corresponds to a learning content stored in the learning content library. The method is applied to the operation sensing panel, and the method includes the following steps: S1 initial learning content loading: Identify the target entity identifier placed on the operation sensing panel, determine the target learning content corresponding to the target entity identifier from the learning content library of the cloud service platform, and display the target learning content on the child terminal device; S2 Multimodal Data Acquisition and Timing Alignment: The system collects physical operation data of children on multiple entity identifiers based on the target learning content in real time, and obtains children's visual and speech data from the children's terminal device; using the system time of the children's terminal device as a global time reference, the physical operation data, the visual data, and the speech data are time-aligned to generate a time-aligned multimodal dataset, which includes physical operation data, visual data, and speech data. S3 Multimodal Fusion Analysis: Based on the multimodal dataset, the child's current emotional state, learning intention, mastery level of each knowledge point in the target learning content, and weak knowledge points in the target learning content are determined, and a child's cognitive state result including the emotional state, learning intention, mastery level, and weak knowledge points is generated. S4 Adaptive Interactive Feedback: Based on the child's cognitive state results, a target task difficulty suitable for the child's current state is determined, and personalized learning content that matches the target task difficulty and the learning intention is determined, so that when the physical identifier placed on the operation sensing panel is recognized again, the personalized learning content is displayed on the child's terminal device.
[0007] In a second aspect, this disclosure provides an electronic device including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the computer program, implements the method described in the first aspect.
[0008] Thirdly, this disclosure provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method described in the first aspect.
[0009] Compared with the prior art, the beneficial effects of this disclosure are: By operating the sensor panel to perceive the placement and operation of physical markers in real time, and dynamically loading corresponding learning content based on the cloud mapping table, the limitations of traditional devices' one-way triggering are broken. Every operation of the child on the physical object can drive changes in digital content in real time, forming a closed-loop embodied interactive experience, which significantly improves the immersion and participation in learning.
[0010] By using the system time of children's terminal devices as a global benchmark, the physical operation, visual, and voice data are time-series aligned to ensure that multimodal data at the same moment can be accurately correlated, providing a high-quality data foundation for subsequent fusion analysis and avoiding cognitive misjudgments caused by data misalignment.
[0011] By comprehensively analyzing children's emotional state, learning intentions, knowledge mastery, and weaknesses through aligned multimodal data, this approach overcomes the adaptability bottleneck of general intelligent models in children's scenarios, achieving a comprehensive and in-depth understanding of children's cognitive state.
[0012] By identifying the child's cognitive state, the difficulty of the task is dynamically adjusted, and personalized learning content is accurately selected from the cloud-based learning content library. When the child interacts with the physical marker again, adaptive feedback is output in real time, forming an adaptive closed loop of "perception-understanding-decision-feedback", which significantly improves the effect of personalized learning. Attached Figure Description
[0013] Figure 1 A flowchart illustrating the hardware and software integrated adaptive interactive learning method for children provided in this disclosure; Figure 2 This is a flowchart illustrating the multimodal data acquisition and time-series alignment steps provided in this disclosure; Figure 3 This is a flowchart illustrating the multimodal fusion analysis steps provided in this disclosure; Figure 4 A flowchart illustrating the adaptive interactive feedback steps provided in this disclosure; Detailed Implementation
[0014] To further illustrate the technical means and effects adopted by this application to achieve the intended purpose of the invention, the following, in conjunction with the accompanying drawings and preferred embodiments, details the specific implementation, structure, features, and effects of the hardware-software integrated adaptive interactive method, electronic device, and medium for children's learning proposed in this application. In the following description, different "one embodiment" or "another embodiment" do not necessarily refer to the same embodiment. Furthermore, specific features, structures, or characteristics in one or more embodiments can be combined in any suitable form.
[0015] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains.
[0016] The following description, in conjunction with the accompanying drawings, details the specific scheme of the hardware-software integrated adaptive interactive method and system for children's learning provided in this application.
[0017] This disclosure provides an embodiment of a hardware and software integrated adaptive interactive method, electronic device, and medium for children's learning. The adaptive interactive system for children's learning includes: an operation sensing panel, multiple physical markers, a children's terminal device, a guardian's terminal device, and a cloud service platform. The operation sensing panel is communicatively connected to the children's terminal device and the cloud service platform, respectively.
[0018] The operation sensing panel is made of 300×200mm ABS+PC material and is divided into grid-shaped sensing areas. In this embodiment, there are 24 grids of 50×50mm each, with a ring antenna arranged below each sensing area and an electromagnetic shielding cover to reduce crosstalk between adjacent antennas. It is used to sense the position and identity information of physical markers placed on it in real time.
[0019] The physical identifier is a carrier with a built-in sensing chip, such as a passive RFID chip. Each physical identifier stores a unique identifier, which is associated with learning content stored in the learning content library of the cloud service platform through a cloud mapping table. Specifically, the cloud service platform maintains a physical identifier-learning content mapping table to record the binding relationship between each physical identifier's identifier and the learning content. When a child places a physical identifier on the operating sensing panel, the system reads its identifier and queries the mapping table to dynamically load the corresponding learning content from the learning content library. This association can be dynamically adjusted by the cloud based on the child's learning progress, cognitive state, or parental settings, thereby enabling personalized expansion of learning content without replacing the hardware.
[0020] Children's terminal devices can be, for example, tablet computers, which are equipped with an image acquisition unit, an audio acquisition unit, and a display and interaction unit, used to collect visual and audio data and output learning content and interactive feedback, respectively.
[0021] The cloud service platform stores a learning content library and a learning intent library, and communicates with the operation sensor panel, children's terminal device, and guardian's terminal device. The connection with the operation sensor panel and children's terminal device enables real-time loading of learning content, collection and synchronization of multimodal data, and distribution of personalized feedback. The communication connection with the guardian's terminal device is used to push children's learning reports, cognitive development status, learning behavior summaries, and parent-child interaction recommendations to the guardian, allowing parents to understand their children's learning progress in real time and receive data-supported co-education suggestions, strengthening effective family participation and guidance in children's learning.
[0022] Please see Figure 1 Based on the above system, the hardware-software integrated adaptive interactive learning method for children described in this embodiment includes the following steps in sequence: S1 initial learning content loading, S2 multimodal data acquisition and temporal alignment, S3 multimodal fusion analysis, and S4 adaptive interactive feedback. The following provides a detailed description of each step: S1 Initial Learning Content Loading Steps When a child places a physical identifier on the operating sensor panel, the initialization process is initiated. The operating sensor panel identifies the target physical identifier placed on it using its built-in sensor array, such as an RFID sensor array, and reads the unique identifier of the target physical identifier. Based on this identifier, the operating sensor panel sends a content request to the cloud service platform. The cloud service platform, according to a pre-stored physical identifier-learning content mapping table, retrieves and determines the target learning content corresponding to the target physical identifier from the learning content library. Subsequently, the cloud service platform sends the matched target learning content to the child's terminal device, and the child's terminal device displays the target learning content through its display interaction unit, completing the initial loading of learning content. To achieve efficient identification, according to another embodiment of this disclosure, identifying a target entity identifier placed on an operation sensing panel includes: The frequency of children's operation in each sensing area of the operation sensing panel within a preset period is statistically analyzed. For each sensing area, the percentage of operation frequency in the sensing area is determined. If the percentage of operation frequency is greater than or equal to a frequency threshold, the sensing area is determined to be a hotspot sensing area; if the percentage of operation frequency is less than the frequency threshold, the sensing area is determined to be a non-hotspot sensing area. The scanning period of the hotspot sensing area is determined to be a first period and the scanning period of the non-hotspot sensing area is determined to be a second period, wherein the first period is shorter than the second period.
[0023] This dynamic adaptive scanning strategy divides regions by heat, using a shorter scanning cycle for high-frequency operating areas to ensure real-time interaction, and a longer scanning cycle for low-frequency areas to reduce invalid scans. This improves scanning efficiency and reduces total system power consumption while ensuring real-time recognition.
[0024] S2 Multimodal Data Acquisition and Timing Alignment Steps Please see Figure 2 After loading the initial learning content, children begin interacting with physical markers based on the target learning content. To comprehensively perceive the children's learning status, this step collects their physical operation data, visual data, and speech data in real time. Addressing the data timing misalignment issue caused by clock deviations and inconsistent sampling frequencies in multi-source acquisition devices, this step uses the system time of the child's terminal device as the global time reference to perform time-series alignment processing on the collected multimodal data. This eliminates time axis deviations between different modalities, generating a time-aligned multimodal dataset, providing a precise data foundation for subsequent multimodal fusion analysis. Specifically, this includes the following two core execution steps: S21 Multimodal Data Acquisition and S22 Time-Series Alignment Processing.
[0025] S21 Multimodal Data Acquisition Real-time collection of multimodal data related to children's learning status, including: Physical operation data: Collected in real time by the operation sensor panel, recording children's operation events such as placing, moving, and removing physical markers.
[0026] Visual data: Collected by the camera of the child's terminal device, recording visual information such as the child's facial expressions, gaze direction, and body movements.
[0027] Voice data: Collected by the microphone of the child's terminal device, recording the voice content, tone, and speed of the child's speech.
[0028] S22 Timing Alignment Processing Physical operation data, visual data, and voice data suffer from clock skew and inconsistent acquisition frequencies due to different acquisition devices, resulting in inaccurate synchronization of multimodal data along the timeline. To address this issue, this step employs a weighted multi-stage time-series alignment method, sequentially performing timestamp unification calibration, filling in missing values for data from different frequencies, and weighted synchronization error compensation, ultimately generating a strictly time-aligned multimodal dataset. Specifically, it includes the following sub-steps: S221 Timestamp Unified Calibration Based on the reference time of the child terminal device system To establish a global time reference, the timestamps of the three types of data acquisition devices—operation sensor panels, cameras, and microphones—are uniformly calibrated. Finally, the local timestamps of all data are uniformly calibrated to the global time reference. The calibration formula is as follows: Fixed time difference The calculation formula is: in, The value is used to collect the device type variable. (Operation sensor panel), vis (camera), audio (microphone), For the original local timestamp of the i-th frame of data from the corresponding operation sensor panel, camera, and microphone device; This refers to the fixed time difference between the corresponding operation sensor panel, camera, microphone device, and child terminal device. The local time for the first data collection by the corresponding operation sensor panel, camera, and microphone device. The globally unified timestamp after calibration of the i-th frame of data from the corresponding operation sensor panel, camera, and microphone device, i.e. .
[0029] S222 Different Frequency Data Missing Value Completion Due to differences in acquisition frequency, the calibrated physical operation data, visual data, and voice data are prone to data gaps at the preset target fusion time point. This step uses a linear interpolation method to fill in the missing values in the calibrated data to ensure data continuity. The interpolation formula is as follows: in, For the target time The corresponding complete feature values; The feature values of adjacent calibrated frames; The calibration timestamps corresponding to adjacent frames, and satisfying the following conditions: .
[0030] It should be understood that the physical operation data is a globally unified timestamp calibrated using the k-th frame data from the operation sensor panel. The globally unified timestamp after calibration with the (k+1)th frame data This is used to fill in missing values using the interpolation formula mentioned above. For visual data, it utilizes a globally unified timestamp calibrated from the k-th frame of the camera data. The globally unified timestamp after calibration with the (k+1)th frame data The missing values are filled in using the interpolation formula mentioned above. For speech data, a globally unified timestamp calibrated from the k-th frame of data from the microphone device is used. The globally unified timestamp after calibration with the (k+1)th frame data It is used to fill in missing values using the interpolation formula mentioned above.
[0031] This step ensures that all modalities have continuously available data at any given time, avoiding the impact of missing data on the accuracy of subsequent multimodal fusion analysis.
[0032] S223 Weighted Synchronization Error Compensation After unified timestamp calibration and missing value completion, although the physical operation data, visual data, and speech data already possess continuous time series, their original sampling times differ. Directly using interpolated estimates for fusion would introduce temporal misalignment. Therefore, this step introduces a weighted synchronization error compensation mechanism, resampling the three types of data based on a unified reference timestamp to generate time-aligned multimodal sample points.
[0033] First, based on the importance of each modality to children's cognitive assessment, differentiated importance weights were assigned to physical manipulation data, visual data, and phonological data, denoted as [weights to be filled in]. The weights are non-negative. The weight values can be preset according to the actual application scenario. For example, operational data directly reflects the mastery of knowledge points and has the highest weight; visual data reflects emotional state and has the second highest weight; and voice data reflects learning intention and has the lowest weight.
[0034] For any given moment to be analyzed, find the data from the physical operation data, visual data, and voice data respectively. nearest calibrated timestamp Then, a unified reference timestamp is calculated using a weighted average, as shown in the formula: The denominator of this formula automatically normalizes the weights, resulting in an "optimal common time reference." Data with larger weights has a more accurate original sampling time reference. The greater the impact, the closer the reference time will be to the actual sampling time of the important data, thus prioritizing its accuracy in subsequent compensation.
[0035] Obtain a unified reference timestamp Then, through linear interpolation, physical operation data, visual data, and voice data are respectively generated in [the following contexts / systems]. The feature values at each time step are used to form a truly synchronized multimodal sample point. This process is repeated for all time steps that need to be analyzed to obtain a temporally aligned multimodal dataset for the entire learning process.
[0036] By employing a multi-stage time alignment mechanism based on the terminal system time, including unified timestamp calibration, missing value completion for inter-frequency data, assignment of importance weights, unified reference timestamp for weighted average calculation, and weighted synchronization error compensation, the problem of clock deviation and sampling frequency inconsistency between multi-source acquisition devices is effectively solved. This ensures accurate alignment of physical operation, visual, and voice data in the time dimension, providing a high-quality data foundation for subsequent multimodal fusion analysis and fundamentally avoiding cognitive misjudgments caused by data misalignment.
[0037] S3 Multimodal Fusion Analysis Steps Please see Figure 3 This step aims to gain a deep understanding of children's cognitive states from the aligned multimodal dataset, providing a basis for decision-making in subsequent adaptive interactions. Addressing technical challenges such as the variability of children's emotional expressions, the jump in learning intentions, and the difficulty in quantifying knowledge point mastery in real time, this step employs multi-algorithm collaborative analysis to identify children's current emotional states and learning intentions, quantify their mastery of each knowledge point, and pinpoint weak knowledge points, ultimately generating a structured result of children's cognitive state. Specifically, it includes the following four sub-steps: S31 Children's Emotional State Identification, S32 Children's Learning Intention Identification, S33 Determining Knowledge Point Mastery and Locating Weak Knowledge Points, and S34 Generating Children's Cognitive State Results.
[0038] S31 Children's Emotional State Recognition To accurately identify children's emotional states during the learning process, this step uses a weighted fusion strategy based on visual and speech data from a multimodal dataset.
[0039] First, based on aligned visual and speech data, unimodal emotion confidence scores are calculated using pre-trained models. Specifically, the visual unimodal confidence score... Facial features were extracted using a MobileNetV2 model finely tuned based on a dataset of 150,000 facial expressions from children aged 3-8 years. These features were then processed through a fully connected layer and a Softmax function to calculate the speech single-modal confidence score. The speech features are extracted using the MFCC algorithm and then calculated using a fully connected layer and the Softmax function.
[0040] Subsequently, modality validity assessment factors and emotion fluctuation correction factors were introduced, and the child's current overall emotional confidence level was calculated using a weighted fusion formula. The formula is as follows: in, These are the preset weighting coefficients for the visual and speech modalities, respectively. Typically, the visual modal is given a higher weight because facial expressions are the core carrier of emotion in children. The magnitude of the weight directly affects the contribution of each modality to the final emotion recognition result: the larger the weight, the more significant the impact of the modality's confidence level on the fusion result.
[0041] This is the modal validity criterion factor, used to determine whether modal data at a certain moment is valid, and is defined as follows: This factor uses a threshold to eliminate low-confidence single-modal data, avoiding interference from invalid data caused by device obstruction, environmental noise, etc., and improving the robustness of recognition.
[0042] The emotion fluctuation correction factor, used to calibrate the bias between visual and speech modal confidence scores, is defined as follows: in, This is the fluctuation correction coefficient, used to control the degree of penalty imposed on the fusion result by modal inconsistency. When the confidence levels of the visual and speech modalities differ significantly, A decrease in the overall confidence level after fusion corresponds to a decrease in the overall confidence level, reflecting the inconsistency between modalities and preventing errors in overall emotion recognition due to misjudgment of a single modality. Conversely, when the confidence levels of the two modalities are equal, A value close to 1 indicates a more reliable fusion result.
[0043] This weighted fusion strategy, by introducing validity assessment and fluctuation correction factors, effectively avoids interference from invalid data caused by device occlusion, environmental noise, etc., on the fusion results. It also calibrates the confidence bias between the visual and speech modalities, preventing overall emotion recognition errors due to misjudgment of a single modality. This ensures that emotion recognition remains stable and accurate across children's varied facial expressions and speech. The final output is a comprehensive confidence score for emotion. It can be used for subsequent cognitive state analysis: if the emotional confidence level is high and points to positive emotions such as focus and excitement, the system can appropriately increase the difficulty of the task to stimulate learning potential; if negative emotions such as frustration and confusion are identified, the system will automatically reduce the difficulty of the task or output encouraging feedback to avoid children's frustration and improve their learning experience and participation.
[0044] S32 Children's Learning Intention Recognition To accurately understand children's learning intentions during the operation process, this step is based on an aligned multimodal dataset, which integrates physical operation data and voice data for collaborative analysis.
[0045] First, physical operation feature vectors and speech feature vectors are extracted from the multimodal dataset. Specifically, the physical operation feature vectors... The operation sequence encoding of entity identifiers is used to represent the degree of matching between physical operation behavior and candidate intentions in the learning intention database; the speech feature vector is obtained from the operation sequence encoding of entity identifiers and includes information such as knowledge point labels, operation location, and operation timing. Extracted from speech data using a semantic understanding model, it is used to characterize the degree of matching between speech content and each candidate intent.
[0046] For each candidate learning intent stored in the learning intent library The cosine similarity between the corresponding physical operation feature vector and the speech feature vector is calculated as a measure of the consistency between the intention and the child's current behavior. The cosine similarity formula is as follows: in, For physical operation feature vectors, For speech feature vectors, To unify the dimension of feature vectors; These are the nth eigenvalues of the two eigenvectors, and their similarity is... The value ranges from [0,1]. The closer the value is to 1, the higher the probability that the operation and the speech expression point to the same intention, and vice versa.
[0047] In the calculation of cosine similarity, the dimension of the feature vector... The dimensionality determines the representational power of the features. Higher dimensionality allows for richer information carrying, helping to capture children's leaps and symbolic learning intentions, but also increases computational complexity. Optimizing the feature extraction network can improve... It fully reflects the deep semantic connection between children's physical manipulation and speech, thereby improving the accuracy of intent recognition.
[0048] After obtaining the similarity of each candidate intent, a preset similarity threshold is introduced. θ The determination is made based on the cosine similarity of a candidate intent. If the intention is greater than or equal to the threshold θ, then the intention is determined as the child's current learning intention. Threshold θ The threshold setting directly affects the performance of intent recognition: a higher threshold results in higher accuracy but may miss some genuine intents; a lower threshold increases recall but may introduce false positives. In practical applications, the threshold can be dynamically adjusted according to product requirements to balance precision and recall.
[0049] The identified learning intentions will serve as an important component of the child's cognitive state, used for content selection and difficulty adjustment in subsequent adaptive interactive feedback. For example, if the child's current intention is identified as "exploring new knowledge," the system can prioritize recommending extension content; if the intention is "repeated practice," it will focus on pushing reinforcement tasks.
[0050] By integrating physical manipulation features with speech semantic features and performing cross-modal cosine similarity matching, this system effectively captures the implicit, symbolic learning intentions of children during physical manipulation, overcoming the limitations of traditional single-modal intention understanding. Cosine similarity calculation quantifies the consistency between manipulative behavior and speech expression, enabling the system to accurately distinguish between different intentions such as "exploring new knowledge" and "repeated practice." The preset similarity threshold can be dynamically adjusted according to product strategies, balancing accuracy and recall. The identified learning intentions, as an important component of children's cognitive state, provide reliable intention dimension input for content selection and difficulty adjustment in subsequent adaptive interactive feedback, making personalized recommendations more aligned with children's current learning needs.
[0051] S33 Determining the Mastery of Knowledge Points and Identifying Weak Knowledge Points To quantify children's mastery of each knowledge point in a seamless and real-time manner during the learning process, and to accurately identify weak areas, this step is based on an aligned multimodal dataset, which integrates physical operation data and difficulty coefficients for modeling and analysis.
[0052] First, for each knowledge point in the target learning content... The number of times children correctly manipulated the physical markers corresponding to the knowledge point during the activity was counted. and number of erroneous operations The total number of operations is At the same time, a preset difficulty level is associated with each operation. ,in Indicates the first This operation. This indicates the knowledge point corresponding to the operation. The difficulty level is pre-defined by the learning content library. The larger the value, the higher the difficulty of the task represented by the operation. The value range is usually [0.1, 1.0].
[0053] Based on the above statistics, the knowledge points are obtained through weighted calculation. initial level of mastery The formula is as follows: in, This represents the sum of the difficulty coefficients of all correct operations; This represents the sum of the difficulty coefficients of all erroneous operations; The denominator is the average difficulty coefficient of all operations for the knowledge point. This is the sum of the difficulty levels of all operations. This is the error penalty coefficient, a constant greater than 1, used to impose additional penalties on erroneous operations to reflect the negative impact of errors on mastery. The larger the value, the more significant the reduction in mastery caused by the same erroneous operation. The function ensures that the initial mastery level does not exceed 100%.
[0054] This formula is based on the difficulty-weighted sum of correct operations, subtracts the difficulty-weighted sum of incorrect operations after punishment, and then divides by the total difficulty-weighted sum to obtain a normalized percentage of mastery. The underlying logic is: the more correct operations performed and the higher the difficulty, the higher the mastery; the more incorrect operations performed and the higher the difficulty, the lower the mastery. By introducing a difficulty coefficient weighting, the contribution of correct operations on high-difficulty tasks to mastery is made greater, more accurately reflecting children's actual abilities.
[0055] Based on the initial mastery level, a continuous error penalty factor is further introduced to correct the initial mastery level, resulting in knowledge points. The final level of mastery is determined by the formula: in, This is a consecutive error penalty factor, the value of which is related to the number of consecutive errors. Related. This embodiment uses a piecewise function definition: The purpose of the consecutive error penalty factor is to identify potential weaknesses in children when they make multiple consecutive mistakes on the same knowledge point. The system then considers this a possible comprehension obstacle and significantly lowers their mastery assessment score, for example, by multiplying it by 0.5. This helps identify potential weaknesses earlier and avoids misjudgments based on accidental errors. If the number of consecutive errors is small, the initial mastery score remains unchanged. This factor effectively enhances sensitivity to learning difficulties.
[0056] Finally, for each knowledge point ,judge Is it less than the preset mastery threshold? In this embodiment It can be adjusted according to the characteristics of different knowledge points. If If so, mark that knowledge point as a weak knowledge point.
[0057] The mastery level and weak areas of each knowledge point obtained will serve as an important component of the child's cognitive status results, directly used for adjusting task difficulty and selecting personalized learning content in subsequent adaptive interactive feedback. For example, when a knowledge point is identified as a weak point, the system will prioritize recommending reinforcement content for that knowledge point and appropriately reduce the task difficulty; conversely, for knowledge points with a high level of mastery, the system can provide more challenging extension content.
[0058] By statistically analyzing correct and incorrect operation data in real time during natural operation and combining it with a difficulty coefficient weighted calculation, a seamless learning assessment is achieved. This approach does not interfere with children's normal operation while accurately quantifying their mastery of each knowledge point. A continuous error penalty factor is introduced to correct the initial mastery level. When children make multiple consecutive errors on the same knowledge point, the system can identify potential learning difficulties earlier, improving diagnostic sensitivity. Finally, knowledge points below a preset threshold are marked as weak areas, providing a precise quantitative basis for adjusting task difficulty and selecting personalized content in subsequent adaptive interactive feedback. This allows the system to deliver targeted reinforcement content and enhance learning outcomes.
[0059] S34 Children's Cognitive Status Results After identifying children's emotional state, learning intentions, knowledge mastery, and weaknesses, this step systematically integrates the results of the above multi-dimensional analysis to generate structured results of children's cognitive state, providing a unified and interpretable basis for subsequent adaptive interactive feedback.
[0060] The generated results of children's cognitive states will be passed to the S4 adaptive interactive feedback step, serving as the core input for determining task difficulty and selecting personalized content.
[0061] Through the aforementioned structured integration, the system achieves the abstraction and transformation from raw multimodal data to high-dimensional cognitive representations, enabling the subsequent adaptive engine to make refined and dynamic interactive decisions based on the child's current comprehensive state. This result can also be synchronized to the cloud service platform to update children's learning characteristic data, supporting the generation of learning reports and the optimization of parent-child interaction recommendations, thereby realizing a data-driven, closed-loop process and the continuous evolution of personalized learning paths.
[0062] S4 Adaptive Interactive Feedback Steps Please see Figure 4After obtaining structured results of the child's cognitive state, this step enters the adaptive interactive feedback stage. The core objective of this stage is to dynamically adjust the difficulty of subsequent learning tasks based on the child's current cognitive level, emotional state, and learning intentions, and to select personalized learning content that matches these levels. When the child interacts with the physical marker again, the system not only selects personalized learning content based on their cognitive state but also combines this with the current emotion recognition results, outputting the content in real time through multimodal methods such as screen animation and voice prompts, thereby achieving a deep integration of physical operation and digital feedback. Specifically, this includes the following three sub-steps: determining the difficulty of the target task, selecting personalized learning content, and outputting multimodal interactive feedback. Each sub-step is explained in detail below.
[0063] Determining the difficulty of the S41 objective task Dynamic adjustment of task difficulty is a core element of adaptive learning. Its design must adhere to the principles of "smooth transition" and "positive reinforcement," meaning that while avoiding sudden increases in difficulty that could lead to frustration in children, it's also crucial to prevent excessively low difficulty that could cause boredom, thus guiding children to remain within their "zone of proximal development." To this end, this method proposes an emotion-mastery nonlinear coupling algorithm. By integrating a child's cognitive level (i.e., their mastery of knowledge points) with their emotional state (i.e., positive emotions such as focus and excitement), it generates a target task difficulty that aligns with their current developmental stage.
[0064] First, based on the level of mastery of each knowledge point in the children's cognitive state results, the average mastery level of the target learning content is calculated. The scores were then normalized to the [0,1] interval to reflect the child's overall proficiency with the current learning content. Simultaneously, the confidence score for focused emotion was extracted from the emotion state recognition results. confidence level of excitement Both values range from [0,1], with higher values indicating stronger positive emotions.
[0065] Based on this, the difficulty of the fundamental target is calculated through nonlinear coupling. : in, This is the base difficulty coefficient for the current knowledge point, preset by the learning content library, representing the standard difficulty of the knowledge point without any adjustment factors. This is a difficulty adjustment coefficient, which controls the overall impact of emotion and mastery on difficulty. The higher the value, the more sensitive the difficulty becomes to changes in mastery and positive emotions. In practical applications, this can be adjusted according to product strategy. Parameter To determine the mastery offset, 0.6 was chosen as the baseline threshold based on the cognitive principle in child development psychology that "a mastery level of 60% is generally considered a preliminary sign of understanding." When M > 0.6, the offset is positive, indicating that the child has the foundation to accept higher challenges; when M < 0.6, the offset is negative, meaning that the difficulty should be appropriately reduced to consolidate the foundation. The reason for using multiplication instead of addition for the positive emotion coupling term is that the simultaneous presence of two positive emotions, focus and excitement, has a synergistic reinforcing effect on learning: only when children are both focused and excited is it appropriate to moderately increase the difficulty to stimulate their potential; if only one emotion is high, such as blind excitement without focus, it is not advisable to blindly increase the difficulty. Multiplicative coupling can effectively reflect this synergistic effect; only when both are high will the product increase significantly, thereby driving the increase in difficulty.
[0066] Basic target difficulty The calculations demonstrate that the positive driving force behind the increase in difficulty comes from the combination of "achieving mastery" and "high levels of positive emotions." Otherwise, the difficulty remains stable or decreases, ensuring that children are always in a positive learning zone.
[0067] To avoid excessive fluctuations in difficulty that could interfere with children, it is necessary to... The target task difficulty is obtained by smoothing the process. : in, This represents the actual difficulty coefficient of the previous interactive task. The smoothing mechanism stipulates that the difficulty adjustment range for a single instance cannot exceed ±20%, and the difficulty coefficient is limited to the range [0.3, 1.5]. The lower limit of 0.3 ensures that the task has at least basic challenge, avoiding boredom due to overly simple tasks; the upper limit of 1.5 prevents the task from being too difficult and exceeding the child's ability. This range was determined through extensive prior data analysis of children's behavior, balancing safety and motivation.
[0068] The final target task difficulty This will serve as one of the key inputs for selecting personalized learning content, and will be used together with the child's learning intentions and weak knowledge points to match the most suitable learning content from the learning content library.
[0069] S42 Personalized Learning Content Selection Determine the difficulty of the target task Next, this step aims to select personalized learning content from the cloud-based learning content library that best matches the child's learning needs and emotional characteristics, based on the child's current cognitive state. The selection process comprehensively considers the following input factors: the difficulty of the target task. Children's learning intentions List of weak knowledge points The system also incorporates real-time state features of children extracted from multimodal datasets. The selection process comprises two core steps: multimodal feature attention fusion and content matching calculation. Through these two steps, the system can accurately match children's multimodal behavioral features with candidate content in the learning content library, achieving adaptive content recommendation.
[0070] S421 Multimodal Feature Attention Fusion To comprehensively characterize a child's current learning state, features extracted from different modalities need to be effectively fused. However, the contribution of different modalities to the current learning state is not constant. For example, when a child is expressing an intention using language, the phonological modality should have a higher weight; when a child is focused on observing a screen, the visual modality is more crucial. Therefore, this step introduces an attention mechanism to dynamically calculate the weights of each modal feature, generating a comprehensive feature vector that adaptively highlights key modal information.
[0071] Specifically, from the time-aligned multimodal dataset generated in step S2, three modal feature vectors are extracted for the target time, i.e., the current analysis time: Visual feature vectors : Derived from the intermediate layer output of the facial expression recognition model, with dimensions of It encodes visual information such as children's facial expressions and gaze direction.
[0072] Speech feature vector : Derived from the intermediate layer output of the speech intent recognition model, dimension It encodes semantic information such as children's speech content and intonation.
[0073] Physical operation feature vector : Obtained by encoding the operation sequence of entity identifiers using a Long Short-Term Memory (LSTM) network, with a dimension of It encodes the knowledge point tags, operation order, operation frequency and other behavioral patterns of the operation.
[0074] To fuse the aforementioned feature vectors into a unified representation, an attention score is first calculated for each modality using a multilayer perceptron (MLP), and then the attention weights are obtained by normalization using a softmax function. : in, It is a trainable multilayer perceptron that maps input features to a scalar score, reflecting the importance of that modality feature to the current learning state. Attention weights. satisfy The larger the value, the more critical the modality is to the representation of the child's state at the current moment.
[0075] After obtaining the attention weights, the feature vectors of each modality are weighted and concatenated to form a comprehensive fusion feature vector. : in, This represents a vector concatenation operation, which joins the three weighted feature vectors along their feature dimensions to form a vector with dimension 1. The comprehensive feature vector is generated. This vector retains the original information of each modality and strengthens the modal features most relevant to the current state through an attention mechanism, enabling subsequent content matching to focus on the child's core behavioral performance.
[0076] The key to the attention mechanism lies in its dynamic adaptability. As a child's learning state changes, such as shifting from focused operation to seeking help through language, the attention weights automatically adjust to ensure that the fused features always highlight the most informative modality. This mechanism significantly improves the relevance and robustness of feature representation, laying a solid foundation for the accurate calculation of subsequent content matching.
[0077] The generated integrated feature vector The digital representation of the child's current state is used to calculate the matching degree with the feature vectors of each candidate content in the learning content library, thereby selecting the learning content that best meets the child's personalized needs.
[0078] S422 Content Matching Calculation and Filtering In obtaining the comprehensive fusion feature vector of the child's current state Next, this step aims to select personalized learning content from the cloud-based learning content library that best matches the child's cognitive state, learning intentions, and the difficulty of the target task. The selection process consists of two stages: first, a preliminary screening is conducted based on learning intentions and weak knowledge points to narrow down the candidate pool; then, a fine screening is performed using a matching degree calculation formula to determine the final recommended content.
[0079] The preliminary screening of candidate content specifically involves: each candidate content in the cloud-based learning content library. All have pre-defined feature vectors This vector encodes multiple attributes of the content, including: knowledge point tags and preset difficulty level. Information such as interactive formats, games, nursery rhymes, age-appropriate content, and corresponding entity identifiers (IDs).
[0080] The system first identifies the child's current learning intention based on step S32. I And the list of weak knowledge points identified in step S33 W extracts all content related to I from the learning content library. or W The relevant candidate content forms a candidate content set C. candidateThis initial screening process ensures that subsequent refined screening focuses on knowledge areas that the child is currently interested in or needs to reinforce, thus improving recommendation efficiency.
[0081] The matching degree calculation is specifically as follows: for the candidate content set C candidate For each candidate content C, calculate its integrated feature vector with the child's current state. Matching degree The calculation formula is as follows: in, The comprehensive fusion feature vector of the child's current state is obtained by step S421 through an attention mechanism, which integrates the visual feature vector. Speech feature vectors Physical operation feature vector It is formed by weighted fusion, with the following dimensions. This vector comprehensively reflects a child's behavioral patterns, emotional state, and learning intentions at the current moment, and is a digital representation of the child's state.
[0082] It is the feature vector of candidate content C The transpose of . and They share the same dimensions, encoding the semantic information and pedagogical attributes of the content. The similarity in the feature space is calculated by taking the dot product of the two values. Theoretically, the value can be any real number, but after feature normalization, a larger value indicates that the child's state is closer to the semantics of the content. The value range is usually between [−1,1] or [0,1]. This part reflects the degree of semantic matching between the content and the child's current interests and needs.
[0083] It is the target task difficulty coefficient determined in step S41, with a value range of [0.3, 1.5], representing the difficulty target after the system adaptively adjusts it according to the child's cognitive state and emotional level. This is the preset difficulty level of candidate content C, pre-defined by the content library, with a value range of [missing information]. Consistency reflects the inherent challenge of the content itself. It is the absolute difference between the difficulty of the target task and the difficulty of the candidate content, with a value range of [0, 1.2]. The smaller the difference, the more the difficulty of the content matches the child's current ability level.
[0084] It's the difficulty matching, when and When the two are completely equal, this item is 1, indicating a perfect match in difficulty; when the difference is at its maximum, this item is -0.2, indicating a severe mismatch in difficulty. The larger this item is, the better the difficulty fit.
[0085] This is the difficulty matching weight coefficient, which is set to 0.3 in this embodiment to balance the contributions of semantic matching and difficulty matching to the final matching degree. The higher the coefficient, the stronger the influence of difficulty factors on content selection; conversely, the lower the coefficient, the more emphasis is placed on semantic similarity. This coefficient can be dynamically adjusted according to product strategy; for example, it can be appropriately increased during the consolidation phase. To enhance difficulty adaptation, the difficulty level can be reduced during the exploration phase. To encourage content diversity.
[0086] Match It combines semantic similarity and difficulty suitability; the higher the value, the better the candidate content matches the child's current overall state. The value range is influenced by both the dot product term and the difficulty term.
[0087] System traverses candidate content set C candidate The matching degree of each piece of content is calculated, and the content with the highest matching degree is selected as the final personalized learning content. If multiple pieces of content have similar matching degrees, the content that can specifically make up for weak knowledge points and is suitable for children's age characteristics is selected first to enhance learning effectiveness.
[0088] S43 Personalized Learning Content Output After selecting personalized learning content, the goal of this step is to deeply integrate this content with children's physical activities and output it in a multimodal format at appropriate times, thereby achieving closed-loop adaptive interactive feedback. The output stage must ensure a clear binding relationship between the content and physical markers, a reliable triggering mechanism, and the ability to provide empathetic interaction based on the child's real-time emotional state to enhance the learning experience.
[0089] The selected personalized learning content is not isolated but associated with specific physical identifiers. Specifically, the system maintains a dynamic mapping table in the cloud, recording the unique identifier (ID) of each physical identifier and its corresponding learning content. When step S42 determines that the personalized learning content is the optimal recommended content, if the physical identifier corresponding to the personalized learning content already exists in the current operating environment (i.e., the identifier is already in the child's hands), the system directly establishes a temporary or permanent binding between the identifier's ID and the personalized learning content in the cloud mapping table. If the personalized learning content requires a new physical identifier, the system can prompt the parent or child to pick up the corresponding identifier through the child's terminal device, for example, by displaying "Please place the bear building block on the panel" on the screen. The binding is completed after the identifier is placed.
[0090] This binding mechanism ensures that learning content is no longer fixed to a specific piece of hardware, but rather dynamically associated with the identifier ID through cloud mapping. When a child subsequently places the same identifier on the operating sensor panel again, the system can load the corresponding personalized content from the cloud in real time based on the ID, achieving adaptive expansion of "same hardware, different content".
[0091] When a child places a bound physical identifier on the operation sensor panel, the panel quickly reads the identifier's ID using a dynamic adaptive scanning strategy and uploads it to the child's terminal device. The child's terminal device then requests corresponding personalized content from the cloud based on the ID. The cloud service platform then sends the personalized content to the child's terminal device, which displays the personalized content through its display interaction unit.
[0092] During the personalized learning content output process, the system continuously monitors the child's emotional changes through the emotion recognition results in step S31, and dynamically adjusts the feedback strategy to provide empathetic interaction: If the system detects that a child is experiencing negative emotions such as frustration or confusion, it will automatically perform one or more of the following actions: temporarily reduce the difficulty of the current task, play encouraging voice messages, provide visual cues or demonstration animations to guide the child to operate correctly.
[0093] If the system recognizes positive emotions such as high concentration and excitement in children, it can appropriately expand the advanced content: automatically unlock hidden challenges or Easter eggs after the task is completed, add reward animations or praise voices, and recommend more exploratory extension content.
[0094] Regarding the determination of target task difficulty, an emotion-mastery nonlinear coupling algorithm is used to co-model children's emotional state and their level of knowledge mastery, resulting in a positive correlation between target task difficulty, mastery level, and positive emotions. When children have a high level of mastery and are in a focused, excited, or other positive emotional state, the system appropriately increases the task difficulty to stimulate their learning potential; conversely, it maintains a stable level or reduces the difficulty to avoid frustration and ensure that the learning task is always within the child's "zone of proximal development."
[0095] In terms of personalized learning content selection, the system first extracts candidate content from the learning content library based on the difficulty of the target task, learning intention, and weak knowledge points for initial screening. Then, it extracts visual, speech, and physical operation features from a multimodal dataset and dynamically weights and fuses them using an attention mechanism to generate a comprehensive fusion feature vector that adaptively highlights key modal information. Finally, it calculates the matching degree between each candidate content and this feature vector, selecting the content with the highest matching degree as personalized learning content. This mechanism ensures that content selection simultaneously considers the suitability of the target task difficulty, the matching degree of learning intention, and the targeted reinforcement of weak knowledge points, guaranteeing a precise fit between the recommended content and the child's current overall state and significantly improving the effectiveness of personalized learning.
[0096] S5 Data Closed Loop and Cloud Collaboration Steps This disclosed hardware-software integrated adaptive interactive learning method for children, based on the four core steps mentioned above, adds cloud data synchronization, dynamic expansion of learning content, and learning progress reports and parent-child interactive recommendation push, realizing a closed-loop data process and home-school collaboration. The specific implementation is as follows: S51 Children's Learning Data Desensitization and Synchronization The system anonymizes locally generated multi-source data, retaining only structured data that cannot identify individuals, including: operation event records (timestamps and locations of marker placement, movement, and removal), knowledge point mastery statistics (final mastery level of each knowledge point), and aggregated indicators of ability dimensions (scores in dimensions such as language, logic, spatial reasoning, and emotional perception). This protects children's privacy.
[0097] The aggregated metrics for the capability dimension are calculated through the following steps: First, the system pre-defines a core competency framework, including language skills, logical thinking, spatial awareness, and emotional cognition. Each knowledge point in the learning content library is pre-labeled with its corresponding competency dimension.
[0098] Within each evaluation period, the cloud service platform performs weighted aggregation of the final mastery level of each knowledge point calculated in step S33, categorized by dimension. Each knowledge point is assigned a different weight coefficient based on its importance within its respective dimension.
[0099] The scores for each dimension are ultimately presented in the learning report in the form of radar charts, trend curves, etc., intuitively showing the child's development level in various abilities. This mechanism transforms fragmented knowledge mastery into interpretable and comparable ability dimension indicators, making it easier for parents and educators to quickly grasp the child's overall development status.
[0100] The anonymized data is transmitted to the cloud service platform via an encryption protocol. The cloud platform updates its stored data on children's learning characteristics, i.e., children's cognitive profiles, based on the received data. All cloud data is statically stored using encryption algorithms, strictly adhering to the principles of data minimization and access control policies to ensure data security throughout its entire lifecycle.
[0101] S52 Physical Identifiers - Dynamically Linked Learning Content The cloud service platform maintains a mapping table between physical identifiers and learning content, recording the unique identifier of each physical identifier and its binding relationship with the learning content in the learning content library. To enable personalized expansion of learning content, the platform employs a hierarchical dynamic association mechanism, dynamically adjusting the mapping relationship based on updated children's learning characteristic data. Specifically: The content in the learning content library is divided into multiple levels such as basic level, intermediate level and challenge level according to the theme, difficulty and ability dimension. Each level contains several content units. Based on the child’s current cognitive profile, such as the level of mastery of each knowledge point, the score of the ability dimension, and the learning preferences, the most suitable content unit for the child’s current development level is dynamically matched for each entity identifier. Each time new learning data for children is received, the platform reassesses the child's status and triggers an update to the mapping table.
[0102] S53 Learning Report Generation and Parent-Child Interaction Recommendation Push Based on updated children's learning characteristics data, the cloud service platform generates visual content using the following algorithms and pushes it to the guardian's terminal: Learning Report Generation: This process aggregates children's mastery of various knowledge points according to preset ability dimensions and calculates a comprehensive score for each dimension. In this embodiment, four ability dimensions are set: language, logic, spatial reasoning, and emotional perception. The score for each dimension is obtained by weighted averaging of the mastery of all knowledge points within that dimension. The final learning report generates visualizations such as radar charts and trend curves, intuitively displaying the child's ability development and trends.
[0103] Parent-child interaction recommendation generation: Combining a list of children's weak knowledge points, interest tags, and historical recommendation records, personalized parent-child interaction recommendations are generated through a rule engine or collaborative filtering algorithm. The recommendation results include the interaction theme, goal, step suggestions, and necessary entity markers. For example, if a child's mastery of the "spatial geometry" knowledge point is below the threshold, but they have recently shown interest in building block activities, then the interaction program "Parent-child building blocks - recognizing shapes" is recommended.
[0104] The generated learning reports and parent-child interaction recommendations are pushed to the guardian's terminal device connected to the cloud via an encrypted protocol. Parents can understand their child's learning progress and ability profile in real time, and receive actionable suggestions for home-school collaboration, enabling them to participate more effectively in their child's learning process with the assistance of artificial intelligence.
[0105] Data anonymization and encrypted transmission ensure children's privacy and security while synchronizing local learning data to the cloud, building a continuously updated cognitive profile of children and providing a data foundation for personalized services. Based on this profile, the cloud mapping relationship between physical identifiers and learning content is dynamically adjusted to achieve personalized expansion of learning content, with 100% hardware reuse rate, significantly reducing the cost of content expansion. The cognitive profile is transformed into a visual learning report and an actionable parent-child interaction recommendation, which is pushed to the guardian's terminal, enabling parents to understand their children's learning progress in real time and receive precise co-education guidance.
[0106] This embodiment provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the hardware-software combined adaptive interactive learning method for children as described above.
[0107] This embodiment provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the hardware-software combined adaptive interactive learning method for children as described above.
[0108] The above description is merely a preferred embodiment of this disclosure and is not intended to limit the scope of protection of this disclosure. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.
Claims
1. A hardware-software integrated adaptive interactive learning method for children, characterized in that, The children's adaptive interactive learning system includes an operation sensing panel, multiple physical identifiers, a children's terminal device, and a cloud service platform; the operation sensing panel is communicatively connected to both the children's terminal device and the cloud service platform; the cloud service platform stores a learning content library, and each physical identifier corresponds to a learning content stored in the learning content library; the method is applied to the operation sensing panel, and the method includes the following steps: S1 initial learning content loading: Identify the target entity identifier placed on the operation sensing panel, determine the target learning content corresponding to the target entity identifier from the learning content library of the cloud service platform, and display the target learning content on the child terminal device; S2 Multimodal Data Acquisition and Timing Alignment: The system collects physical operation data of children on multiple entity identifiers based on the target learning content in real time, and obtains children's visual and speech data from the children's terminal device; using the system time of the children's terminal device as a global time reference, the physical operation data, the visual data, and the speech data are time-aligned to generate a time-aligned multimodal dataset, which includes physical operation data, visual data, and speech data. S3 Multimodal Fusion Analysis: Based on the multimodal dataset, the child's current emotional state, learning intention, mastery level of each knowledge point in the target learning content, and weak knowledge points in the target learning content are determined, and a child's cognitive state result including the emotional state, learning intention, mastery level, and weak knowledge points is generated. S4 Adaptive Interactive Feedback: Based on the child's cognitive state results, a target task difficulty suitable for the child's current state is determined, and personalized learning content that matches the target task difficulty and the learning intention is determined, so that when the physical identifier placed on the operation sensing panel is recognized again, the personalized learning content is displayed on the child's terminal device.
2. The method according to claim 1, characterized in that, The adaptive interactive learning system for children also includes a guardian terminal device, which is communicatively connected to the cloud service platform. The method further includes: synchronizing at least one of the multimodal dataset, the child's cognitive state results, and the personalized learning content as child learning data to the cloud service platform to update the child learning feature data stored on the cloud service platform; the cloud service platform adjusts the association between the entity identifier and the learning content in the learning content library based on the updated child learning feature data, and generates a learning progress report and parent-child interaction recommendations based on the updated child learning feature data, and pushes the learning progress report and interaction recommendations to the guardian terminal.
3. The method according to claim 1, characterized in that, The time alignment process includes: using the system time of the child terminal device as a global time reference, performing unified timestamp calibration on the physical operation data, the visual data, and the voice data respectively; The calibrated physical operation data, visual data, and voice data are filled with missing values for different frequencies. Based on the importance of the completed physical operation data, visual data, and speech data to children's cognitive assessment, differentiated importance weights are assigned to the physical operation data, visual data, and speech data, respectively. Based on the calibrated timestamps of the physical operation data, the visual data, and the voice data, and their respective importance weights, a unified reference timestamp is calculated by weighted average. Based on the unified reference timestamp, weighted synchronization error compensation is performed on the physical operation data, the visual data, and the voice data to generate a time-aligned multimodal dataset.
4. The method according to claim 1, characterized in that, Determining the child's current emotional state includes: Single-modal emotion confidence and validity determination are performed on the visual data and speech data in the multimodal dataset to identify valid visual data and valid speech data. The effective visual data and the effective speech data are weighted and fused according to preset weights to obtain a fusion result. The fusion result is then calibrated using an emotion fluctuation correction factor to obtain the child's current emotional state.
5. The method according to claim 1, characterized in that, The cloud service platform stores a learning content library and a learning intention library. The learning intention library stores multiple candidate learning intentions. Determining a child's current learning intention includes: Based on the multimodal dataset, physical operation sequence feature vectors and speech semantic feature vectors are extracted respectively. The physical operation sequence feature vectors are used to characterize the degree of matching between the physical operation data and multiple candidate learning intentions in the learning intention library, and the speech semantic feature vectors are used to characterize the degree of matching between the speech data and the multiple candidate learning intentions. For each candidate learning intention, the cosine similarity between the matching degree of the physical operation data and the candidate learning intention and the matching degree of the voice data and the candidate learning intention is determined. If the cosine similarity is greater than or equal to the similarity threshold, the candidate learning intention is determined as the child's current learning intention.
6. The method according to claim 1, characterized in that, Determining the level of mastery of each knowledge point among the multiple knowledge points included in the target learning content, and identifying the weak knowledge points in the target learning content, including: Based on the multimodal dataset, the correct and incorrect operation data of children on the entity icons corresponding to the target learning content are statistically analyzed; The initial level of mastery is obtained by weighting the correct and incorrect operation data and their corresponding difficulty coefficients. The initial mastery level is corrected by a continuous error penalty factor to determine the child's mastery level of each knowledge point among the multiple knowledge points included in the target learning content. For each of the aforementioned knowledge points, if the level of mastery of the knowledge point is less than a preset threshold, then the knowledge point is identified as a weak knowledge point in the target learning content.
7. The method according to claim 1, characterized in that, Determine the difficulty level of the target task to be appropriate for the child's current state, including: Based on the emotional state and mastery level in the children's cognitive state results, the difficulty of the target task is determined through nonlinear coupling, wherein the difficulty of the target task is positively correlated with the mastery level and emotional state; The determination of personalized learning content that matches the difficulty of the target task and the learning intention includes: Based on the target task difficulty, learning intention, and weak knowledge points, candidate learning content is extracted from the learning content library; Visual features, speech features, and physical operation features are extracted from the multimodal dataset; The visual features, speech features, and physical operation features are fused using an attention mechanism to obtain a comprehensive fusion feature vector of the child's current state; Calculate the matching degree between each candidate learning content and the integrated feature; determine the candidate learning content with the highest matching degree as the personalized learning content.
8. The method according to claim 1, characterized in that, The identification of the target entity marker placed on the operation sensing panel includes: The frequency of children's operation in each sensing area of the operation sensing panel within a preset period is statistically analyzed. For each sensing area, the percentage of operation frequency in the sensing area is determined. If the percentage of operation frequency is greater than or equal to a frequency threshold, the sensing area is determined to be a hotspot sensing area; if the percentage of operation frequency is less than the frequency threshold, the sensing area is determined to be a non-hotspot sensing area. The scanning period of the hotspot sensing area is determined to be a first period and the scanning period of the non-hotspot sensing area is determined to be a second period, wherein the first period is shorter than the second period.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the software and hardware combined adaptive interactive method for children's learning as described in any one of claims 1 to 8.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the software and hardware combined adaptive interactive method for children's learning as described in any one of claims 1 to 8.