An educational system for multi-dimensional perception and intention recognition

Through the multi-dimensional perception and intention recognition education system, the high-fidelity space-time interaction engine and video sequence generation module are used, combined with three-dimensional Gaussian sputtering and neural radiation field rendering technology, the shortcomings of existing equipment in terms of interaction, cognition, security and content ecology are solved, and an immersive, personalized and secure educational experience is achieved.

CN120164149BActive Publication Date: 2025-08-29SHENZHEN EFERCRO ELECTRONIC TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510640195.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-19
Publication Date
2025-08-29
Estimated Expiration
2045-05-19

AI Technical Summary

Technical Problem

The existing children's education equipment has significant shortcomings in terms of interaction dimension, cognitive depth, security protection and content ecology. The interaction is single, cognitive depth, weak security protection and rigid content ecology, making it difficult to achieve immersive, personalized and dynamic interactions.

Method used

The education system of multi-dimensional perception and intention recognition is adopted, and the high-fidelity space-time interaction engine, video sequence generation module and audio-visual generation module are used, combined with three-dimensional Gaussian sputtering and neural radiation field rendering technology, to achieve efficient mapping and interaction from physical input to dynamic digital twins, and safe management and adaptive learning are carried out through the intelligent control center.

Benefits of technology

It realizes high-fidelity and consistent timing immersive interaction, refined content recommendation and security management, supports personalized content creation and editing, and improves the interactive capabilities and security of the equipment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120164149B_ABST
    Figure CN120164149B_ABST
Patent Text Reader

Abstract

The present application discloses an educational system for multi-dimensional perception and intention recognition, which includes: a high-fidelity spatiotemporal interaction engine for performing posture estimation and picture card recognition on image streams using a dual-attention ray scoring network, obtaining the physical interaction mode between children and equipment and picture cards and the picture card recognition results, and generating a high-fidelity three-dimensional model based on the picture card recognition results in combination with three-dimensional Gaussian sputtering / neural radiation field rendering technology; a video sequence generation module for generating a video sequence based on a high-fidelity three-dimensional model using a spatiotemporal collaborative network when a video sequence needs to be generated; an audio-visual generation module for inputting text descriptions and video sequences into an audio generation model to obtain target audio and video; a compensation module for correspondingly editing the target audio and video during dynamic teaching or children's creation. Through the above method, efficient mapping and interaction from physical input to dynamic, time-consistent digital twins can be achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of educational systems, and in particular to an educational system for multi-dimensional perception and intention recognition. Background Art

[0002] Existing children's educational devices have significant deficiencies in terms of interaction dimensions, cognitive depth, safety protection, and content ecology:

[0003] Single interaction dimension: Traditional devices mostly remain at the level of static graphics or preset animations, lacking real-time perception and dynamic response to children's physical posture, environmental context, and intrinsic motivations, making it difficult to achieve truly immersive and personalized interaction.

[0004] Insufficient cognitive depth: Learning content is mostly one-way indoctrination, which cannot effectively distinguish the intrinsic interests and external incentives behind children's behavior, resulting in superficial personalized recommendations and guidance and inefficient knowledge internalization.

[0005] Passive security protection: Most devices only focus on content filtering and usage duration. They have weak defenses against spoofing attacks on access authentication (such as facial recognition) and black box detection and adversarial attacks against the AI ​​model itself, posing potential security risks.

[0006] Rigid content ecosystem: Lacking the ability to dynamically generate, edit, and integrate multiple modalities (especially time-series audio and video), the content format is single, making it difficult to meet the growing demand for personalized and contextualized learning, and also limiting the secondary creation of parents / educators. Summary of the Invention

[0007] Based on the above background, the core technical issues that need to be solved are upgraded to:

[0008] How can we build a high-fidelity, time-consistent, and multimodally collaborative immersive cognitive environment? This requires moving beyond static graphics and simple animations to achieve efficient mapping and interaction from physical input (image cards, gestures) to dynamic, time-consistent digital twins (3D models, video, and audio).

[0009] How can we achieve proactive, adaptive, intelligent control and guidance based on a deep understanding of motivation? This requires going beyond simple rules and regulations to gain a deeper understanding of the internal and external drivers of children's behavior. This, combined with digital twin simulation, allows for refined, forward-looking content recommendations and safety management.

[0010] How can we build an end-to-end proactive security defense system that covers the entire process from access to interaction? We need to defend against multi-dimensional threats, from identity forgery to model detection / attacks, especially black-box query attacks during the interaction process.

[0011] How can we empower devices with the capabilities of dynamic content generation, editing, and multimodal integration while ensuring consistency of content (especially quantity)? This requires integrating generative AI with editing technologies to support on-demand creation of personalized content.

[0012] The multi-dimensional perception and intention recognition educational system provided in this application can achieve efficient mapping and interaction from physical input (picture cards, posture) to dynamic, time-consistent digital twins (3D models, video, audio).

[0013] In the first aspect, the present application provides an educational system for multi-dimensional perception and intention recognition, which includes: a high-fidelity spatiotemporal interaction engine, which is used to use a dual-attention ray scoring network to perform posture estimation and picture card recognition on the collected image stream, obtain the physical interaction mode between the child and the device and the picture card, and the picture card recognition results, and generate a high-fidelity three-dimensional model based on the picture card recognition results in combination with three-dimensional Gaussian sputtering / neural radiation field rendering technology; a video sequence generation module, which is used to use a spatiotemporal collaborative network to generate a video sequence based on the high-fidelity three-dimensional model when a video sequence needs to be generated; an audio-visual generation module, which is used to input text descriptions and video sequences into an audio generation model, generate audio that is synchronized in time and matches the meaning of the video sequence, and obtain the target audio and video; a compensation module, which is used to edit the target audio and video accordingly during dynamic teaching or children's creation.

[0014] Among them, the high-fidelity spatiotemporal interaction engine is also used to combine device performance level, user preferences and content metadata to select three-dimensional Gaussian sputtering or neural radiation field rendering technology to generate high-fidelity three-dimensional models.

[0015] Among them, in the process of displaying high-fidelity three-dimensional models, the high-fidelity spatiotemporal interaction engine is also used to switch between three-dimensional Gaussian sputtering and neural radiation field rendering technologies based on the monitored performance indicators.

[0016] Among them, the video sequence generation module is also used to obtain the first feature map of the current frame, the second feature map of the previous frame, and the third feature map of the next frame; obtain the forward attention map based on the first feature map and the second feature map, and obtain the backward attention map based on the first feature map and the third feature map; perform collaborative attention calculation on the forward attention map and the backward attention map to obtain the collaborative attention map; obtain the previous context information by weighted summing the collaborative attention map with the second feature map, and obtain the subsequent context information by weighted summing the collaborative attention map with the third feature map; fuse the previous context information, the subsequent context information and the first feature map to obtain the current frame features; generate a video sequence based on the current frame features.

[0017] Among them, the video sequence generation module includes: a motion prediction unit, which is used to predict the motion information between frames, and the motion information is input into the spatiotemporal collaborative network as a condition; a spatiotemporal feature fusion and propagation unit, which is used to splice the feature maps of the previous and next frames in the channel dimension and send them to the subsequent processing layer, as well as to add the features of the previous and next frames at the element level or perform weighted summation according to the attention weight, and use a gating mechanism to control the propagation and update of historical information, and use a convolution kernel that can process both spatial and temporal dimensions at the same time.

[0018] Among them, the audio-visual generation module is also used to extract and align features of text descriptions and video sequences to obtain aligned visual features and text features; and input the visual features and text features into the audio generation model to generate audio that is synchronized in time with the video sequence and matches the meaning, thereby obtaining the target audio and video.

[0019] Among them, the compensation module includes: a feature compensation unit, which is used to extract features of the original image during dynamic teaching or children's creation to obtain initial features, and interact the quantity and object information in the text prompt with the initial features to obtain a compensated feature vector, and fuse the compensated feature vector and the initial features to obtain enhanced image features; a quantity perception attention unit, which is used to extract information about each object from the enhanced image features, and interact the attention of each object information with the current feature map to obtain quantity perception features, and inject the quantity perception features into the corresponding feature map.

[0020] The education system also includes: an intelligent management and control center, which uses an internal / external factor decoupling framework to analyze children's interaction logs, learning choices, and usage time; as well as monitoring interaction patterns with cloud-based artificial intelligence services to detect potential black box adversarial attack attempts in real time.

[0021] Among them, the intelligent management and control center is also used to use digital twins for simulation prediction and adjust strategies in advance.

[0022] Among them, the education system also includes: adaptive learning and full-process security protection modules, which are used for personalized adaptive learning, multi-level active safety and continuous health monitoring.

[0023] The beneficial effects of the present application are as follows: Different from the existing technology, the multi-dimensional perception and intention recognition education system provided by the present application includes: a high-fidelity spatiotemporal interaction engine, which is used to use a dual-attention ray scoring network to perform posture estimation and picture card recognition on the collected image stream, obtain the physical interaction mode between the child and the device and the picture card, and the picture card recognition result, and generate a high-fidelity three-dimensional model based on the picture card recognition result in combination with three-dimensional Gaussian sputtering / neural radiation field rendering technology; a video sequence generation module, which is used to use a spatiotemporal collaborative network to generate a video sequence based on the high-fidelity three-dimensional model when a video sequence needs to be generated; an audio-visual generation module, which is used to input text descriptions and video sequences into the audio generation model to generate audio that is synchronized in time and matches the video sequence in meaning, to obtain the target audio and video; a compensation module, which is used to edit the target audio and video during dynamic teaching or children's creation, and can achieve efficient mapping and interaction from physical input (picture card, posture) to dynamic, time-consistent digital twins (3D model, video, audio), as well as realize refined and forward-looking content recommendation and security management, and support on-demand creation of personalized content. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for describing the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For those skilled in the art, other drawings can be obtained based on these drawings without inventive efforts. Among them:

[0025] Figure 1 This is a schematic diagram of the structure of an embodiment of the multi-dimensional perception and intention recognition education system provided by the present application;

[0026] Figure 2 This is a structural diagram of an embodiment of a video sequence generation module provided by the present application;

[0027] Figure 3 It is a structural diagram of an embodiment of a compensation module provided in this application. DETAILED DESCRIPTION

[0028] The technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. It will be understood that the specific embodiments described herein are only used to explain the present application, rather than to limit the present application. It should also be noted that, for ease of description, only some, rather than all, structures related to the present application are shown in the drawings. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application.

[0029] References herein to "embodiments" mean that a particular feature, structure, or characteristic described in connection with the embodiments may be included in at least one embodiment of the present application. The appearance of this phrase in various places in the specification does not necessarily refer to the same embodiment, nor does it constitute an independent or alternative embodiment that is mutually exclusive of other embodiments. It is understood, both explicitly and implicitly, by those skilled in the art that the embodiments described herein may be combined with other embodiments.

[0030] See Figure 1 , Figure 1 1 is a schematic diagram of an embodiment of an educational system for multi-dimensional perception and intention recognition provided by the present application. The educational system 100 includes: a high-fidelity spatiotemporal interaction engine 10, a video sequence generation module 20, an audio-visual generation module 30, and a compensation module 40.

[0031] The high-fidelity spatiotemporal interaction engine 10 is used to use a dual-attention ray scoring network to perform posture estimation and picture card recognition on the collected image stream, obtain the physical interaction mode between the child and the device, and the picture card, and the picture card recognition results, and generate a high-fidelity three-dimensional model based on the picture card recognition results in combination with three-dimensional Gaussian sputtering / neural radiation field rendering technology.

[0032] In some embodiments, the high-fidelity spatiotemporal interaction engine 10 is further used to select three-dimensional Gaussian sputtering or neural radiation field rendering technology to generate a high-fidelity three-dimensional model in combination with device performance level, user preference settings and content metadata.

[0033] In some embodiments, during the display of a high-fidelity three-dimensional model, the high-fidelity spatiotemporal interaction engine 10 is further configured to switch between three-dimensional Gaussian sputtering and neural radiation field rendering techniques based on monitored performance indicators.

[0034] In some embodiments, the high-fidelity spatiotemporal interaction engine 10 primarily involves physical world perception and gesture understanding, as well as dynamic content activation and timing generation.

[0035] The perception of the physical world and posture understanding are mainly reflected in: using a dual-attention ray scoring network, a monocular visible light camera can be used to estimate the six-degree-of-freedom posture of children / devices in real time with high precision, and accurately capture the physical interaction methods (distance, angle, and viewing method) between children and devices and picture cards.

[0036] Dynamic content activation and time-series generation are mainly reflected in: after image card recognition, high-fidelity three-dimensional models are generated by combining three-dimensional Gaussian sputtering / neural radiation field rendering technology.

[0037] Among them, 3D Gaussian sputtering does not rely on complex neural networks for volume rendering, but uses a large number (tens of thousands or even millions) of 3D Gaussian functions (primitives) with specific properties (3D position, covariance matrix i.e. shape / rotation, color, opacity) to accurately fit and represent 3D objects or scenes.

[0038] When a children's educational device successfully identifies a specific physical card (e.g., a card with an apple on it) within its field of view using an image recognition algorithm (e.g., a convolutional neural network (CNN)-based classifier or object detection model), the system triggers a critical real-time 3D content generation and rendering process. The goal of this process is to instantly render a highly relevant, visually realistic, and interactive 3D digital model on the device screen.

[0039] The rendering process mainly includes content association and model loading, real-time rendering, and view synthesis and coloring.

[0040] Content association and model loading are primarily achieved by using the image card recognition result ("apple") as a query index. Based on this index, the system quickly loads the corresponding 3D Gaussian parameter set for "apple" (including the position, shape, color, opacity, and other data of all Gaussian primitives) from a pre-built and optimized model library. This model library is typically generated offline through multi-view image training or conversion from 3D scan data.

[0041] Real-time rendering is mainly reflected in: utilizing the high parallel computing capabilities of modern graphics processing units (GPUs) and a specially designed differentiable Gaussian rasterization rendering pipeline to efficiently "sputter" or project millions of loaded three-dimensional Gaussian primitives into the two-dimensional screen space.

[0042] View synthesis and shading are primarily manifested in the following: the rendering pipeline performs depth sorting and alpha blending on the projected Gaussian primitives based on the current viewing angle (possibly dynamically adjusted based on real-time child / device pose data), calculating the final color and opacity of each pixel to produce a high-quality, hole-free 2D image. The rendering process also incorporates learned color information (typically represented by spherical harmonics to support view-dependent effects) and opacity to achieve realistic lighting and texture effects.

[0043] The main advantages of 3D Gaussian sputtering technology are: fast rendering speed, often achieving real-time or even faster-than-real-time frame rates, making it ideal for applications requiring immediate interactivity. Model representation is relatively intuitive, allowing for modifications to specific areas (such as color changes and local deformations).

[0044] The neural radiance field is primarily implemented by training one or more multi-layer perceptron (MLP) neural networks to learn a continuous 3D scene function. This function takes as input the coordinates of a 3D point (x, y, z) and a viewing direction (θ, φ), and outputs the volume density (σ) and view-dependent color radiance (c) at that point.

[0045] The rendering process mainly includes content association and model loading, view-based volume rendering, sampling and query, and color integration.

[0046] The main features of content association and model loading are: the image card recognition result ("apple") is also used as an index, and the system loads the weights of the pre-trained neural radiance field network model associated with "apple". This model has been trained offline using a large number of rendering losses (comparing the rendered image with the real multi-view image).

[0047] The main feature of view-based volume rendering is that when an image with a specific viewpoint needs to be rendered, the system will send a large number of rays from the virtual camera through the scene.

[0048] Sampling and querying are primarily performed as follows: along each ray, the system samples multiple points in 3D space. The coordinates and ray direction of each sampled point are input into the loaded neural radiance field network, and the volume density and color of that point are queried.

[0049] Color integration involves numerically integrating the color and density of all sample points along the ray path using classic volume rendering equations (alpha synthesis) to calculate the color of the ray as it ultimately reaches the camera. The calculated results of all the rays together form the final rendered 2D image.

[0050] The advantages of neural radiation fields are mainly reflected in: they can usually generate extremely high quality, visual Figure 1 The model representation (network weights) is relatively compact.

[0051] In some embodiments, a neural radiation field and a three-dimensional Gaussian sputtering technique are adaptively selected.

[0052] The adaptive selection strategy is as follows:

[0053] a. Offline content preparation and metadata tagging.

[0054] Dual / multiple representation preparation: For each image card (or its corresponding 3D object) in the content library, during the content production phase, not only should an optimized 3D Gaussian sputtering representation (baseline standard) be generated, but one or more highly optimized neural radiation field variant representations should also be generated for core content with high visual complexity, special optical effects, or particular importance.

[0055] Metadata tagging: Associate detailed metadata with each card content, including content complexity, visual importance, interactivity requirements, available formats, and performance benchmarks.

[0056] Content Complexity: Indicates whether the object contains complex transparency / reflection / refraction effects (high complexity) or is primarily a diffuse surface (low complexity).

[0057] Visual Importance: Marks whether the content is a core teaching point and requires extremely high visual details.

[0058] Interaction requirements: Mark the main interaction modes expected (e.g., static observation, rapid perspective changes, possible future editing needs).

[0059] Available Formats: Lists the formats for which this content is available (e.g., 3D Gaussian Sputtering Optimization, Neural Radiance Field Optimization, 3D Gaussian Sputtering Basics).

[0060] Performance Benchmark: You can benchmark the rendering performance of different formats on reference hardware in an offline phase as reference data.

[0061] b. Device capability assessment and user preference settings. This includes device analysis during initial launch / installation and decision logic during image card loading.

[0062] Device analysis at first boot / installation:

[0063] When the app is first launched, the device's hardware capabilities are evaluated (GPU model and performance, memory size, CPU performance, etc.), and the device can be divided into different performance levels (for example: high, medium, low).

[0064] User / Parent Preferences: Provides simple settings options to allow users (or parents) to select preferred modes, such as:

[0065] Performance priority: Force the use of 3D Gaussian sputtering.

[0066] Balanced Mode: Uses 3D Gaussian sputtering by default, but try higher quality options when conditions allow.

[0067] Quality First: On high-end devices, neural radiation field variants are tried first for supported content.

[0068] The decision logic when loading a chart is as follows:

[0069] When an image card is identified, the device's performance level, user preferences, and content metadata are combined to make the following decisions:

[0070] For low-end devices: Force loading and using 3D Gaussian Sputtering Basic or 3D Gaussian Sputtering Optimized.

[0071] For mid-range / high-end devices:

[0072] Performance priority mode: loads and uses 3D Gaussian sputtering optimization.

[0073] Balanced / Quality Priority Mode:

[0074] Check the metadata for available formats.

[0075] If an optimized neural radiation field variant exists:

[0076] Checks if Content Complexity is High or Visual Importance is High.

[0077] Check whether the interaction requirements are primarily "static observation".

[0078] Checks whether the user preference is "Quality First" or "Balanced".

[0079] If most of the above conditions are met and the device is high-end: initially choose to load neural radiation field optimization.

[0080] Otherwise: Select Load 3D Gaussian Sputtering Optimization (default option).

[0081] If only 3D Gaussian sputtering format exists: load and use 3D Gaussian sputtering optimization.

[0082] c. Runtime status monitoring and dynamic switching (runtime adaptation).

[0083] Performance monitoring: Real-time monitoring of key performance indicators: rendering frame rate (FPS), GPU / CPU load, memory usage, and device temperature.

[0084] Dynamically switch triggers:

[0085] Performance degradation: If the current rendering frame rate is continuously lower than the preset threshold (for example, <25 FPS), or the device temperature is too high, a frequency reduction risk is triggered.

[0086] For interaction mode changes: if the user changes from static viewing to fast mobile device (requiring high FPS).

[0087] Regarding system resource pressure: If other high-priority tasks are started in the background, resources need to be released.

[0088] For battery level / power mode: Enter low power mode or energy saving mode.

[0089] Switching logic:

[0090] Downgrade Switching: If Neural Radiance Field Optimization is currently in use and any of the above conditions are triggered, the system should automatically and seamlessly switch to pre-loaded or on-demand fast-loading 3D Gaussian Sputtering Optimization. This is the most common switching scenario to ensure smooth operation.

[0091] The video sequence generation module 20 is used to generate a video sequence based on a high-fidelity three-dimensional model using a spatiotemporal collaborative network when a video sequence needs to be generated.

[0092] In some embodiments, the video sequence generation module 20 is also used to obtain the first feature map of the current frame, the second feature map of the previous frame, and the third feature map of the next frame; obtain a forward attention map based on the first feature map and the second feature map, and obtain a backward attention map based on the first feature map and the third feature map; perform collaborative attention calculation on the forward attention map and the backward attention map to obtain a collaborative attention map; obtain previous context information by weighted summing the collaborative attention map with the second feature map, and obtain subsequent context information by weighted summing the collaborative attention map with the third feature map; fuse the previous context information, the subsequent context information and the first feature map to obtain the current frame features; generate a video sequence based on the current frame features.

[0093] In some embodiments, see Figure 2 The video sequence generation module 20 includes: a motion prediction unit 21 and a spatiotemporal feature fusion and propagation unit 22.

[0094] The motion prediction unit 21 is used to predict motion information between frames, and the motion information is input into the spatiotemporal collaborative network as a condition.

[0095] The spatiotemporal feature fusion and propagation unit 22 is used to splice the feature maps of the previous and next frames in the channel dimension and send them to the subsequent processing layer, and to add the features of the previous and next frames at the element level or perform weighted summation according to the attention weight, and to use the gating mechanism to control the propagation and update of historical information, and to use a convolution kernel that can process both spatial and temporal dimensions at the same time.

[0096] In this embodiment, a spatiotemporal collaborative network is introduced to generate video sequences (such as story animations) to ensure temporal continuity and content consistency across frames and suppress flickering and mutations.

[0097] After children's educational devices recognize image cards and generate high-fidelity 3D models, a key application scenario is generating dynamic video sequences based on these models, such as telling a story, demonstrating a process, or creating an animation. Directly rendering the 3D model frame by frame (even if the model itself is fixed) or using simple interpolation methods often fails to meet the requirements for high-quality video and is prone to the following problems:

[0098] Timing incoherence: The movement of objects is abrupt and jumpy, lacking natural transitions.

[0099] Inconsistent content: The appearance of objects (such as lighting and texture details) changes slightly between adjacent frames without any reason, the background is unstable, or the identity characteristics of objects drift.

[0100] Visual artifacts: Flicker (rapid fluctuations in brightness or color) or sudden changes (sudden changes in an object's position, posture, or appearance).

[0101] To overcome these challenges, a video sequence generation module20 was constructed using a spatiotemporal collaborative network that can perceive and exploit inter-frame temporal dependencies. Its goal is to fully consider the information of the previous and next frames when generating each frame, thereby ensuring the smoothness, stability, and consistency of the entire video sequence.

[0102] The key technical mechanisms include cross-frame information dependency modeling, bidirectional temporal attention / collaboration mechanism, spatiotemporal feature fusion and propagation, motion / optical flow guidance, and consistency loss function design.

[0103] The cross-frame information dependency modeling is introduced as follows:

[0104] Core Principle: Generating the video content at frame t (F_t) depends not only on the state at the current time t (such as the 3D model's pose, camera position, and lighting settings), but also explicitly on information from past (F_{t-1}, F_{t-2}, ...) and future (F_{t+1}, F_{t+2}, ...) frames. This contextual awareness is fundamental to ensuring coherence.

[0105] Implementation: Usually embedded in the architecture design of video generation networks (such as U-Net variants based on Diffusion Model, generators of video GANs, or Transformer-based video models), cross-frame information interaction is achieved through specific layers or modules.

[0106] The bidirectional temporal attention / cooperation mechanism is introduced as follows:

[0107] Mechanism Details: Similar to the Transformer in natural language processing, this approach introduces an attention mechanism in the temporal dimension. When generating the tth frame, the mechanism calculates the correlation weights between the features of the current frame and those of the preceding, following, and adjacent frames (or even further away).

[0108] The importance of "bidirectionality": not only should we pay attention to past frames to inherit motion trends and content status, but we should also pay attention to future frames to foresee upcoming actions or changes, thereby making smoother transitions.

[0109] "Synergy" effect: The attention calculations of previous and next frames can be designed to influence each other, so that past and future information can more effectively cooperate in the generation of the current frame and jointly determine which temporal information is most important.

[0110] Application: The attention mechanism can act on pixel-level features, high-level semantic features, or latent space representations, dynamically aggregating the most relevant temporal contextual information to guide the generation of details in the current frame.

[0111] For example, a bidirectional temporal collaborative attention module can be set up. The details are as follows:

[0112] 1. Module goals and motivations:

[0113] The core goal of the bidirectional temporal co-attention module is to address temporal consistency issues in video sequence processing, particularly in generation or fusion tasks. As previously mentioned, processing video frames independently can easily lead to issues such as flickering and content jumps. The bidirectional temporal co-attention module aims to enable the feature representation of the current frame (t) to simultaneously perceive and utilize information from both past (t-1) and future (t+1) frames. It then intelligently aggregates this temporal context through a co-attention mechanism, resulting in smoother and more coherent video results.

[0114] 2. Core idea: bidirectionality and collaboration.

[0115] The bidirectional nature of this module lies in the fact that, unlike traditional RNNs that only consider past information or perform only local inter-frame processing (such as simple 3D convolution), the bidirectional temporal co-attention module explicitly looks both forward (the relationship between t and t-1) and backward (the relationship between t and t+1). This is crucial for understanding motion continuity, smooth lighting transitions, and avoiding sudden appearances or disappearances of content.

[0116] The key aspect of synergy lies in the fact that this is more than simply combining forward and backward attention. The key innovation of the bidirectional temporal collaborative attention module lies in its assumption of a mutually dependent and reinforcing relationship between forward and backward attention. In other words, the degree to which a region in the current frame pays attention to past frames should influence its attention to future frames, and vice versa. This "synergistic" mechanism enables better focus on truly important and continuous regions of change in the temporal sequence.

[0117] 3. The implementation process is as follows:

[0118] Get the feature map of the current frame t, the feature map of the previous frame t-1, and the feature map of the next frame t+1:

[0119] Step 1: Generate query, key, and value.

[0120] The current frame feature map is queried through linear transformation (or other mapping).

[0121] The key and value are obtained by linearly transforming the feature map of the previous frame.

[0122] The feature map of the next frame is linearly transformed to obtain the key K and value V.

[0123] Here, Q, K, and V are concepts in the standard self-attention mechanism, representing different functional roles of features.

[0124] Step 2: Calculate the forward and backward attention maps.

[0125] The forward attention map is constructed by calculating the similarity (usually the dot product) between the query Q of the current frame and the key K of the previous frame, and normalizing it through the Softmax function to obtain a weight map indicating "how much attention each position in the current frame should pay to the corresponding position information in the previous frame". For example, the following formula can be used:

[0126] A^{t-1} = softmax( (Q^t)(K^{t-1})^T / sqrt(d_k) ).

[0127] The backward attention map is constructed as follows: Similarly, the similarity between the current frame query Q and the next frame key K is calculated and normalized by Softmax to obtain a weight map of "how much the current frame should pay attention to the next frame information". For example, the following formula can be used:

[0128] A^{t+1} = softmax( (Q^t)(K^{t+1})^T / sqrt(d_k) ).

[0129] sqrt(d_k) is a scaling factor used to stabilize training.

[0130] Step 3: Co-attention computation.

[0131] Multiply the forward and backward attention maps element-wise. A region will only retain a high value in the multiplied map if it receives high attention in both the forward and backward attention maps. This highlights regions that are consistently important or undergo smooth transitions over time.

[0132] The multiplication result is normalized again through Softmax to obtain the final collaborative attention map. For example, the following formula can be used:

[0133] A_{co} = softmax(A^{t-1} ⊙ A^{t+1}).

[0134] The collaborative attention map embodies a more refined attention distribution that incorporates forward and backward bidirectional dependencies.

[0135] Step 4: Weighted aggregation and feature enhancement.

[0136] Use the collaborative attention map to perform a weighted sum of the value V of the previous frame and the value V of the next frame. This is equivalent to extracting the most important context information based on collaborative attention. For example, the following formula can be used:

[0137] Context_{t-1} =A_{co} * V^{t-1} (matrix multiplication).

[0138] Context_{t+1} =A_{co} * V^{t+1} (matrix multiplication).

[0139] The extracted contextual information Context_{t-1} and Context_{t+1} (usually processed by a feed-forward neural network FFN) is fused with the original query Q (or some transformation thereof) of the current frame. This fusion can be done by addition, concatenation, or other methods. For example, the following formula can be used:

[0140] F̂^t_f =Q^t ⊕ FFN(Context_{t-1}) ⊕ FFN(Context_{t+1})

[0141] F̂^t_f is the current frame feature enhanced by the bidirectional temporal collaborative attention module, which contains the bidirectional temporal context information after intelligent aggregation.

[0142] The advantages of the bidirectional temporal collaborative attention module are as follows:

[0143] Improve temporal coherence: By explicitly utilizing information from previous and next frames and using collaborative attention to emphasize persistently important regions, flickering and content mutations are effectively suppressed.

[0144] Enhanced dynamic perception: For moving objects or changing scene elements, it can better capture their continuous trajectories and appearance changes.

[0145] Intelligent information filtering: The collaborative attention mechanism can help the model ignore interference information that only appears briefly in one direction (forward or backward) or is irrelevant, and focus on the temporal context that truly helps maintain consistency.

[0146] Architectural flexibility: The bidirectional temporal co-attention module can be embedded as a module into the encoder or decoder stages of various video processing networks, especially for enhancing feature representations that require temporal awareness.

[0147] Summary: The Bidirectional Temporal Co-Attention Module utilizes innovative bidirectional temporal information processing and a co-attention mechanism, enabling the model to intelligently and interdependently reference the content of preceding and following frames when processing each frame in a video sequence. It computes forward and backward attention maps, collaboratively interacts through element-wise multiplication and softmax, and finally uses the resulting co-attention map to weightedly aggregate the value information of the preceding and following frames and incorporate it into the feature representation of the current frame. This design significantly improves the temporal coherence and content consistency of the generated videos.

[0148] The fusion and propagation of spatiotemporal features are introduced as follows:

[0149] Mechanism details: In the middle layers of the video generation network, feature representations from adjacent frames are explicitly fused. This can be achieved through various methods, such as feature concatenation, feature addition / weighting, recurrent structures, and 3D convolution / spatiotemporal convolution.

[0150] Feature splicing is mainly reflected in: splicing the feature maps of the previous and next frames in the channel dimension and sending them to the subsequent processing layer.

[0151] Feature addition / weighting is mainly reflected in: adding the features of the previous and next frames at the element level or performing weighted summation based on the attention weight.

[0152] The loop structure is mainly reflected in: introducing a gating mechanism similar to RNN or LSTM to control the propagation and update of historical information.

[0153] 3D convolution / space-time convolution is mainly reflected in the use of convolution kernels that can process both spatial and temporal dimensions at the same time.

[0154] Its role is to ensure that when generating the current frame, the network "sees" and utilizes the spatial and appearance information of the previous and next frames, thereby enforcing the continuity of content (such as texture and lighting).

[0155] Motion / optical flow guidance is introduced as follows:

[0156] Mechanism: An additional module can be introduced to predict inter-frame motion information (such as optical flow). Inputting the predicted motion information into the generative network as a condition can more accurately guide the movement of pixels or features, generating motion trajectories that are more consistent with physical laws.

[0157] The design of the consistency loss function is introduced as follows:

[0158] Objective: When training a video generation model, in addition to ensuring the quality of single-frame generation, a loss term that enforces temporal consistency must be added.

[0159] Common forms include inter-frame difference loss, optical flow consistency loss, and variational consistency loss.

[0160] The inter-frame difference loss is mainly reflected in: penalizing excessive pixel differences, feature differences or perceptual differences between adjacent frames in the generated video sequence (such as using LPIPS loss).

[0161] The optical flow consistency loss is mainly reflected in the following: if optical flow is used, it can be required that the image / feature warped from the previous frame to the current frame according to the optical flow should be as similar as possible to the current frame image / feature directly generated.

[0162] The variational consistency loss is mainly reflected in: it is a more refined loss that can distinguish dynamic foreground and static background, enforces higher temporal smoothness on background areas, and allows foreground objects to change according to their motion trajectory, but requires their changes to be consistent with the source video (if applicable) or the expected animation path.

[0163] The purpose is to enable the model to learn to generate videos that are inherently temporally coherent by explicitly optimizing these consistency goals during training.

[0164] Summary: By deeply integrating the concepts of spatiotemporal collaborative networks, specifically introducing cross-frame information dependency modeling, bidirectional temporal attention / cooperation mechanisms, spatiotemporal feature fusion and propagation, and employing a consistency loss function as a constraint during training, we can effectively address the problems of temporal incoherence, content inconsistency, flickering, and sudden changes associated with independent frame generation. This will enable children's educational devices to generate smooth, stable, and highly consistent story animations or teaching demonstration videos, significantly enhancing the sense of immersion and the quality of the learning experience.

[0165] The audio-visual generation module 30 is used to input the text description and video sequence into the audio generation model, generate audio that is synchronized with the video sequence in time and matches the meaning, and obtain the target audio and video.

[0166] Among them, the audio-visual generation module 30 is also used to extract and align features of the text description and the video sequence to obtain aligned visual features and text features; and input the visual features and text features into the audio generation model to generate audio that is synchronized in time with the video sequence and matches the meaning, thereby obtaining the target audio and video.

[0167] In some embodiments, text-assisted audio-visual generation technology is used to automatically generate high-quality audio (sound effects, narration) that is precisely synchronized and semantically matched with the video content (and an optional large language model to generate text descriptions), thereby achieving audio-visual collaboration.

[0168] In children's educational devices, simply generating visual content (such as animations and 3D model demonstrations) is often insufficient to provide a complete immersive experience. Sound (including ambient sound effects, object movement sounds, character voices, background music, and narration) is a crucial component in conveying information and conveying emotion. Traditional video dubbing or sound effects production processes typically require manual effort, which is costly, time-consuming, and difficult to align with dynamically generated or changing visual content in real time.

[0169] Video-to-audio generation technology aims to automatically generate corresponding audio from video images. However, simple V2A models may face the following challenges:

[0170] Semantic ambiguity: It is sometimes difficult to accurately determine what sound should be generated based on visual information alone. For example, when you see a person opening their mouth, it could mean speaking, singing, yawning, or even a silent expression.

[0171] Lack of high-level abstract understanding: It is difficult to generate audio that requires understanding of storylines or abstract concepts, such as narration.

[0172] Difficulty in fine-tuning the generated audio based on specific needs (e.g., emphasizing a sound effect or changing the tone of narration).

[0173] To solve these problems, this application introduces text assistance, using text description as a powerful auxiliary information source to guide audio generation together with video content.

[0174] The core approach lies in text-enhanced audio-visual co-generation. For example, by fusing visual information and text descriptions, the audio generation model is provided with richer and more explicit semantic and temporal constraints, thereby generating high-quality audio that is precisely synchronized with the video content in time and highly matched in meaning.

[0175] The technical implementation process is as follows:

[0176] a. Multimodal input preparation:

[0177] Video input: video clips to be dubbed or dynamically generated video sequences.

[0178] Text input: This is crucial auxiliary information. Text can come from a variety of sources and forms: human annotations / scripts, automatic video descriptions, Large Language Model (LLM) generated descriptions / narrations, and simple keywords / tags.

[0179] Manual annotation / scripting is the most direct method, providing detailed scene descriptions, action instructions, dialogue lines, narration text, etc.

[0180] Automatic video description primarily involves leveraging advanced video understanding models (possibly based on Transformers or multimodal large language models) to automatically generate natural language descriptions of video content. These descriptions can encompass scenes, objects, actions, and even emotions or atmosphere.

[0181] Large language models (LLMs) generate descriptions and narrations by combining video content with pre-defined educational objectives or story outlines, leveraging large language models (such as GPT-4 and LLaMA) to generate more creative and educationally relevant text descriptions or complete narration transcripts. This offers enormous potential for personalized and dynamic content generation.

[0182] Simple keywords / tags are mainly reflected in: It can also be more concise text information, such as "bird singing", "car engine sound", "cheerful background music", etc.

[0183] b. Multimodal feature extraction and alignment. Specifically includes feature extraction and cross-modal alignment.

[0184] Feature extraction mainly involves using pre-trained models to extract feature representations of video, text, and audio (if necessary, as training targets or references).

[0185] Video encoder: such as SlowFast and ViT variants, extracts a visual feature sequence Ev containing spatiotemporal information.

[0186] Text encoder: Such as BERT, T5, and CLIP's text tower, which extracts the semantic features of the text to represent El.

[0187] Audio encoder: such as PANNs and CLAP, extracts the feature sequence Ea of the audio Mel-spectrogram (used as the target during training).

[0188] Cross-modal alignment is a key step in achieving audiovisual-text synergy. This involves mapping the extracted Ev and El (and Ea during training) into a shared latent space using methods such as contrastive learning. By maximizing the similarity between matching visual-text, visual-audio, and text-audio sample pairs, and minimizing the similarity between mismatched pairs, features describing the same semantic concepts across modalities are placed close together in the latent space. This ensures that textual information can effectively "understand" and "associate" with the corresponding visual content.

[0189] c. Text-assisted audio generation model:

[0190] Model architecture: Usually a conditional generation model is used, such as a conditional diffusion model (latent diffusion model LDM is a possible architecture used in TA-V2A), a conditional GAN, or a sequence-to-sequence model based on Transformer.

[0191] Conditional Injection: The aligned visual features Ev and textual features El are input into the generative model as conditions. Injection can be done in a variety of ways, such as cross-attention, conditional embedding concatenation / addition, and multiple guidance.

[0192] Cross attention is mainly reflected in: in the decoder of the U-Net or Transformer of the Diffusion model, the intermediate features in the audio generation process are focused on Ev and El through the cross attention mechanism, thereby generating corresponding audio details based on the visual content and text description.

[0193] Conditional embedding concatenation / addition is mainly reflected in: concatenating or adding the global representation or time-step-related representation of Ev and El with the noise or intermediate state in the generation process to guide the generation direction.

[0194] The main feature of multiple guidance is that during the inference (sampling) phase of the Diffusion model, guidance signals based on visual conditions, text conditions, and even negative text conditions (unwanted sounds) can be independently calculated and weighted together to achieve more refined control over the generated results.

[0195] Generation objective: The goal of the model is to generate an audio representation (such as a mel-spectrogram ẑ0) that is highly consistent with the input Ev and El in both temporal and semantic terms.

[0196] d. Audio Synthesis: This involves converting the audio representation (such as a Mel-spectrogram) output by the generative model into an audible waveform file through a vocoder (such as HiFi-GAN or DiffWave).

[0197] The audio-visual synergy effects and advantages are:

[0198] Precise Timing: A powerful video encoder and cross-modal alignment mechanism ensure that generated audio events (such as footsteps and collisions) are precisely timed to the corresponding actions in the video.

[0199] Enriched semantic matching: The introduction of text descriptions greatly enhances semantic understanding capabilities:

[0200] For example, it can generate specific sound effects: If the text clearly states "cat meow", the model can generate the cat meow sound instead of just guessing based on vague visuals.

[0201] For example, it can generate narration / dialogue: If the text input is a narration or dialogue script, the model can generate the corresponding speech (which may require the combination of TTS technology).

[0202] For example, it can generate ambient sound / background music: the text description "cheerful scene" can guide the model to generate background music or ambient sound that matches the mood.

[0203] For example, disambiguation: for visually ambiguous actions, text can provide clear explanations and guide the generation of correct audio.

[0204] High-quality audio output: Advanced generative models (such as the Diffusion Model) and vocoders can produce high-fidelity, low-noise audio waveforms.

[0205] Flexible Control and Personalization: By modifying the input text description, users can easily adjust the generated audio content, style, or emphasis, achieving a highly personalized audio-visual experience. For example, they can change the tone of the narration or add specific ambient sound effects.

[0206] Summary: Text-assisted audio-visual generation technology achieves precise temporal synchronization and close semantic alignment of audio generation by deeply integrating video content and text descriptions, leveraging cross-modal alignment and conditional generative models. It automatically generates high-quality sound effects, narration, or background music that seamlessly collaborate with (dynamically generated) video content, significantly enhancing the immersiveness, information delivery efficiency, and personalized creative capabilities of children's educational devices. It is a key technical support for building a complete audio-visual experience. The introduction of large language models further enhances the generation capabilities and flexibility of text descriptions, making content creation more intelligent and automated.

[0207] The compensation module 40 is used to edit the target audio and video during dynamic teaching or children's creation.

[0208] In some embodiments, see Figure 3The compensation module 40 includes: a feature compensation unit 41 and a quantity-aware attention unit 42.

[0209] The feature compensation unit 41 is used to extract features from the original image during dynamic teaching or children's creation to obtain initial features, and to interact the quantity and object information in the text prompt with the initial features to obtain a compensated feature vector, and to fuse the compensated feature vector and the initial features to obtain enhanced image features.

[0210] The quantity-aware attention unit 42 is used to extract information about each object from the enhanced image features, and to perform attention interaction between each object information and the current feature map to obtain quantity-aware features, and to inject the quantity-aware features into the corresponding feature map.

[0211] In some embodiments, the fusion of the quantity-aware attention mechanism (the quantity-aware attention unit 42) and the feature compensation module (the feature compensation unit 41) supports quantity-preserving multi-object editing (such as adding or removing objects, changing attributes) of generated content (images / scenes) for dynamic teaching or creation.

[0212] In children's educational devices, simply displaying content statically is not enough. Dynamic teaching (for example, adding and removing objects when demonstrating addition and subtraction) or allowing children to perform simple creative tasks (for example, adding or removing toys from a scene) can greatly enhance the interactivity and fun of learning. However, existing image / scene editing technologies, especially those based on large-scale pre-trained models (such as Stable Diffusion), often encounter the following challenges when dealing with multi-object scenes:

[0213] Quantity Distortion: When editing commands require increasing or decreasing the quantity of a specific object, the model may not be able to accurately control the final generated quantity, resulting in incorrect quantity (too much or too little).

[0214] Attribute Bleeding / Loss: When editing one object in the scene, its attributes (color, shape, pose, etc.) may "bleed" into other objects, causing the attributes of other objects to be lost or changed in unexpected ways.

[0215] Identity Inconsistency: During the editing process, the identity characteristics of an object may drift and no longer look like the original object.

[0216] Dependence on auxiliary tools: Some methods rely on auxiliary tools such as masks to specify the editing area, which increases the complexity of the operation and is not suitable for children or ordinary users.

[0217] To address these issues, we introduce a feature compensation module (feature compensation unit 41) and a quantity-aware attention mechanism (quantity-aware attention unit 42). The core idea is to improve the model's feature representation and attention mechanism, enabling it to clearly perceive and maintain the number of objects in the scene without relying on additional auxiliary tools (such as masks), and to ensure that the attributes of each object are independent and clearly separable during the editing process.

[0218] The technical process is as follows:

[0219] a. Feature Compensation Module - Improves the separation of object attributes:

[0220] Objective: To address the attribute confusion problem caused by pre-trained models (such as CLIP's image encoder) when extracting multi-object image features, ensuring that the attribute features of each object are clear and independent, laying the foundation for subsequent quantitative perception and precise editing.

[0221] The mechanism is as follows:

[0222] Input: The original image I and a text prompt cq containing the object name and quantity information (e.g., “three red apples and two green pears”). This prompt is the key input of FeCom.

[0223] Feature extraction: The image encoder of the CLIP model is used to extract the initial features of the image CLIP(I).

[0224] Feature Interaction and Compensation: A feature attention mechanism is designed. This mechanism uses the quantity and object information in the text prompt cq as a query and interacts with the image features CLIP(I). This interaction aims to identify object attributes that are not fully expressed or distinguished in the image features and generate a compensating feature vector Ic. This compensation vector contains the object-specific attribute information that needs to be "enhanced" or "clarified."

[0225] Feature enhancement: The compensated feature Ic is fused with the original image feature CLIP(I) (e.g., weighted addition) to obtain the enhanced image feature Ig (Ig = CLIP(I) + λIc). The attributes of each object in Ig are more clearly represented and separated, reducing the possibility of attribute confusion in subsequent processing.

[0226] Function: The feature compensation module actively "compensates" for the image encoder's shortcomings in distinguishing multi-object attributes by leveraging prior knowledge (object names and quantities) in textual prompts, generating high-quality feature representations that are more suitable for multi-object editing.

[0227] b. Quantity-aware attention mechanism - achieving quantity preservation and precise editing:

[0228] Goal: Enable the subsequent generation / editing network (usually a U-Net-based Diffusion model) to explicitly "perceive" the existence and quantity of each independent object in the scene, maintain this quantitative relationship during the generation process, and accurately apply the editing instructions to the target objects.

[0229] The mechanism is as follows:

[0230] Input: Enhanced image features Ig (from the feature compensation module) and the noise / intermediate feature map z^t during the diffusion process (at a specific layer, such as the 4th U-Net block B4 selected by the feature compensation module).

[0231] Object Information Extraction: QTTN first includes an extraction module (e.g., a fully connected layer FC or a small convolutional network) to further extract information about each individual object from the enhanced features Ig. This is not just spatial location information, but more importantly, instance-level information that can distinguish "this is an apple" from "this is another apple."

[0232] Attention interaction: Perform attention interaction (e.g., cross-attention) on the extracted object instance information and the current noise / feature map z^t. z^t provides the query (Query, Qz). The extracted object instance information Et(Ig) provides the key (Key, Kg) and value (Value, Vg). By calculating the softmax (Qz * Kg^T / sqrt(dk)) * Vg), each spatial position of z^t dynamically aggregates information from different object instances based on its relevance to those instances.

[0233] Quantity-aware information injection: The quantity-aware feature Vnew after attention-weighted aggregation is injected into the feature map of this layer of U-Net (for example, through addition or channel concatenation post-processing: z^{t+1}_layer =Attention(Qz,Ki,Vi) + βVnew).

[0234] Purpose: The quantity-aware attention module enables the U-Net to "know" how many objects should be in the scene, their approximate location, and their attributes at each step of the denoising (i.e., generation / editing) process. This explicit quantity and instance-awareness enables the model to:

[0235] Keep Quantity: If the editing command does not involve quantity changes, the model tends to maintain the original number of objects.

[0236] Precise control of quantities: If the instructions involve increases or decreases, the model can execute more accurately based on the injected quantity information.

[0237] Apply properties precisely: Apply property changes (such as color) precisely to one or more specified object instances without affecting other objects.

[0238] c. Application in dynamic teaching or creation:

[0239] Dynamic teaching: For example, in mathematics enlightenment, the teacher or system can use text instructions (such as "There are 3 apples on the table now, put 2 more apples") to drive the change of the number of apples in the scene, and use narration (from text-assisted audio and video generation) to explain the addition process.

[0240] Children's Creation: Children can use simple instructions (possibly converted into text through voice recognition) to add, delete or modify objects in virtual scenes (such as "change the red building block to a yellow one" or "add another car"), engaging in exploratory learning and creative expression.

[0241] The technical effects are as follows:

[0242] Number-preserving editing: The core advantage is the ability to precisely maintain or control the number of objects during the editing process.

[0243] Attribute fidelity and separation: Effectively prevents interference with other objects when editing one object, maintaining the independence and consistency of attributes.

[0244] No auxiliary tools required: The editing process does not rely on the user to provide additional information such as masks, which simplifies the operation process.

[0245] Improve interactivity and creativity: Provides powerful technical support for dynamic teaching demonstrations and children's independent creation, enriching the possibilities of educational applications.

[0246] Summary: By integrating a feature compensation module and a quantification-aware attention mechanism, children's educational devices can achieve quantity-preserving multi-object editing when processing image or scene content. The feature compensation module utilizes textual cues to enhance the separation of object attributes, while the quantification-aware attention mechanism provides the generative network with explicit quantity and instance awareness. This combination enables the device to respond to editing commands such as adding or removing objects or changing specific object attributes while maintaining the stability of the attributes of other objects in the scene and the accuracy of the total number of objects (unless the command requires a quantity change), greatly enhancing dynamic teaching demonstrations and children's creative expression.

[0247] Multimodal information fusion: Drawing on multimodal fusion strategies, it integrates multi-source information such as vision (camera, screen), hearing (microphone, speaker), posture (posture estimation algorithm output), and interaction (touch screen / buttons) to build a more comprehensive scene understanding.

[0248] In some embodiments, the education system 100 also includes: an intelligent management and control center, which adopts an internal / external factor decoupling framework to analyze children's interaction logs, learning choices, and usage time; and monitors the interaction patterns with cloud-based artificial intelligence services to detect potential black box adversarial attack attempts in real time.

[0249] Among them, the intelligent management and control center is also used to use digital twins for simulation prediction and adjust strategies in advance.

[0250] The motivation-driven intelligent management and control center is mainly reflected in: deep user profiling and intention understanding, intelligent decision-making and flexible control, and safety monitoring of the interaction process.

[0251] Deep user profiling and intent understanding are mainly reflected in: using an internal / external factor decoupling framework to analyze children's interaction logs, learning choices, usage time and other data to distinguish whether their behavior is based on stable intrinsic interests or situation-driven external factors (such as specific time, parental requirements).

[0252] Intelligent decision-making and flexible control (software-defined networking / digital twin concept application) are mainly reflected in:

[0253] Drawing on the principles of software-defined networking and digital twin integrated detection, a software-defined control plane and digital twins of children's learning status and preferences are constructed. The control logic (parent app) leverages motivation analysis derived from factor decoupling and digital twin simulation results to implement differentiated and dynamic permission management and content recommendations. For example, learning behaviors driven by intrinsic interest can be given more time, while behaviors driven by extrinsic factors (such as solely for task completion) can be appropriately guided or restricted. Simulation predictions (such as predicting learning fatigue and interest shifts) can be made using the digital twin to proactively adjust strategies.

[0254] Security monitoring of the interaction process is mainly reflected in: deploying query update analysis mechanisms, monitoring interaction patterns with cloud-based artificial intelligence services (such as recommendation systems and content generation) (analyzing the incremental similarity of query sequences), and detecting potential black-box adversarial attack attempts in real time, rather than relying solely on input filtering.

[0255] In some embodiments, the education system 100 also includes: an adaptive learning and full-process security protection module for personalized adaptive learning, multi-level active security, and continuous health monitoring.

[0256] Personalized adaptive learning is mainly reflected in: combining federated learning with motivation decoupling capabilities to build a more accurate and explanatory personalized recommendation engine, and dynamically adjust the learning path and content difficulty.

[0257] Multi-level active security is mainly reflected in: access security and interaction security.

[0258] Access security: If facial recognition is used for user authentication or personalized switching, the integrated dual alignment (domain and modality) facial anti-counterfeiting technology effectively resists spoofing attacks such as photos, videos, and 3D masks, especially ensuring robustness in multi-domain (different lighting, devices) and multi-modal (visible light + depth / infrared sensor if available) scenarios.

[0259] Interaction security: Protect core AI models (recommendations, content generation, motivation analysis, etc.) from malicious detection or manipulation by detecting black-box query attacks during interactions.

[0260] Continuous health monitoring is mainly reflected in: retaining and possibly strengthening the original plan's eye tracking and distance monitoring, combined with precise posture data to achieve more reliable eye health management and intervention.

[0261] The technical effect lies in the core creativity of the feedback loop formed by the deep integration of perception, dynamic content generation, user understanding and security protection. This enables the system to simultaneously achieve: extreme realism, deep personalization, proactive security protection and powerful creative empowerment.

[0262] Here are some key synergy examples:

[0263] Hyper-realistic interactive scenarios → driving deeper cognitive understanding and personalization:

[0264] Technologies involved: high-fidelity pose estimation + adaptive 3D rendering (3D Gaussian sputtering / neural radiation field) + decoupling of internal and external motivations + child digital twins.

[0265] Collaborative Approach: The system not only "sees" the child (pose estimation), but also accurately understands the child's six-degree-of-freedom (6DoF) interactions with the physical card / device in real time. This interaction triggers the generation of highly realistic digital twin content (3D Gaussian scattering / neural radiation field), intelligently selecting the optimal rendering technology based on device capabilities and interaction context. This rich interaction data (pose, distance, viewing angle, duration, etc.) provides the motivation decoupling framework with a far stronger signal than simple clicks or usage time, effectively distinguishing whether a child's behavior is driven by intrinsic interest or extrinsic task motivation. This detailed understanding updates the child's digital twin model in real time.

[0266] Unobvious Effects: This synergy elevates the dimension of user understanding from traditional 2D screen interaction to refined three-dimensional interactive understanding based on the physical world and in a high-fidelity, dynamically generated environment. The "realism" here isn't just eye candy; it provides higher-quality, more physically meaningful data for motivation analysis. This enables the digital twin to become a more precise mapping of the child's state, enabling personalization that extends beyond content recommendations to interactive style, learning cadence, and even the fidelity of visual presentation (for example, when the system determines that a child has a strong intrinsic desire to explore, it may switch to the more computationally intensive but more effective neural radiation field rendering). This enables a leap from "adaptive learning" to "context-aware, motivation-driven adaptive reality."

[0267] Dynamic content generation and editing → Enable interactive, quantifiable, embodied experiential learning:

[0268] Technologies involved: pose estimation + 3D Gaussian sputtering / neural radiation field generation + temporal consistency video generation + text-assisted audio and video generation + quantity-aware content editing.

[0269] Collaborative approach: The child shows an "apple" card (pose estimation triggers recognition and loading of the 3D Gaussian sputtering / neural radiation field model). The system can receive a command (which may come from speech recognition or be generated by a large model based on learning objectives), such as "Show me what happens if I add 2 more apples." The quantity-aware editing module accurately adds the specified number of objects to the 3D scene while maintaining the identity characteristics of the original objects. The timing consistency module ensures that the animation of the new apples appears smoothly and naturally without flickering. Text-assisted audio generation synchronously outputs the commentary narration ("Look, 3 apples plus 2 apples, now there are 5 apples!").

[0270] Unobvious Effects: This combination enables instant, physically interactive, and quantifiable concept demonstrations. Rather than pre-made animations, these interactive simulations are dynamically generated in direct response to the child's physical movements and current learning needs. Integrating quantitative editing capabilities directly into the high-fidelity rendering and animation pipeline allows abstract concepts (such as mathematical operations) to be visualized concretely, accurately, and coherently in an immersive environment, freshly "triggered" by the child's own hands. This is something that cannot be dynamically achieved with static images or simple preset animations. The accompanying audio generation further enhances this multi-sensory, interpretive learning loop.

[0271] Motivation-aware digital twins → Proactively shaping interactive experiences and security strategies:

[0272] Technologies involved: motivation decoupling + digital twin + adaptive rendering / content selection + dynamic content editing + interactive query mode monitoring.

[0273] Collaborative Approach: The digital twin constructed based on motivation decoupling not only passively reflects the child's status but also enables prediction and simulation. An intelligent control platform, similar to a software-defined network, can "query" the digital twin: "If this slightly difficult concept is introduced now, based on the child's current intrinsic motivation level, how likely is it that he or she will lose interest?" Based on the simulation results, the system may decide to postpone or simplify the concept or, if the child's intrinsic interest is high, proactively enhance the visuals (such as switching to higher-quality rendering) or increase interactivity to further stimulate it. Interaction query monitoring also leverages the contextual information of the digital twin. For example, if the digital twin shows that a child is deeply engaged in exploring a model, a series of repetitive, patterned exploratory interaction queries will make the system more likely to identify this as a potential black-box attack attempt (against the AI ​​model) than if the child were exploring randomly.

[0274] Less obvious effects: The digital twin has evolved into a proactive simulation engine for personalized and safety decisions. The system's control decisions (content recommendations, difficulty adjustments, interaction modes, etc.) are proactive and context-aware, based on simulations based on a deep understanding of the child. Security protection is no longer simply passive filtering; it has become context-sensitive, leveraging the same user model to assess the rationality of interaction patterns. This creates an intelligent system that proactively manages learning engagement and potential safety risks based on a unified, dynamic user model.

[0275] End-to-end, context-aware security system → Ensures the trustworthiness of immersive generative experiences:

[0276] Technologies involved: dual alignment face anti-counterfeiting + interactive query mode monitoring + posture estimation + full-link dynamic content generation (3D Gaussian sputtering / neural radiation field / video / audio / editing).

[0277] Collaboration: Robust identity authentication (face recognition) ensures system access security and incorporates multimodal and multi-environmental factors. More importantly, interactive query monitoring protects the generative AI model at the core of the system.

[0278] The AI ​​models themselves—including 3D renderers, video / audio synthesizers, and content editors—are protected from malicious detection, manipulation, or reverse engineering during user interaction. Physical perception data, such as pose estimation, provides crucial contextual evidence: Does the user's query align with their current physical pose, attention focus, and interaction object?

[0279] Creative / Unobvious Effects: Security is seamlessly woven into the entire closed loop, from access to interaction to content generation. It specifically addresses and addresses the new attack surfaces introduced by the introduction of powerful generative AI. While traditional security may focus on input filtering or output auditing, this solution proactively monitors the behavioral patterns of interactions with AI models and leverages physical context (posture) and user state (derived from digital twins) to help distinguish legitimate exploration from malicious probing. This builds the necessary trust foundation for the deployment and application of cutting-edge generative AI technologies in sensitive, highly interactive scenarios like children's education.

[0280] In the several embodiments provided in this application, it should be understood that the disclosed methods and devices can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the modules or units is merely a logical functional division. In actual implementation, other division methods may be used, such as combining or integrating multiple units or components into another system, or ignoring or not implementing certain features.

[0281] If the integrated units in the other embodiments described above are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) or a processor to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes various media that can store program code, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0282] The above description is only an implementation method of the present application and does not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made using the contents of the description and drawings of this application, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present application.

Claims

1. A multi-dimensional perception and intention recognition education system, characterized by: The education system includes: A high-fidelity spatiotemporal interaction engine, which uses a dual-attention ray scoring network to perform pose estimation and card recognition on the captured image stream. This engine then derives the physical interaction patterns between the child, the device, and the card, as well as the card recognition results. Based on these card recognition results, a high-fidelity 3D model is generated using 3D Gaussian sputtering / neural radiation field rendering technology. A video sequence generation module, configured to generate a video sequence based on the high-fidelity three-dimensional model using a spatiotemporal collaborative network when a video sequence needs to be generated; An audio-visual generation module, configured to input the text description and the video sequence into an audio generation model, generate audio that is synchronized with the video sequence in time and matches the meaning, and obtain the target audio and video; A compensation module, used for editing the target audio and video accordingly during dynamic teaching or children's creation; The video sequence generation module is further configured to obtain a first feature map of a current frame, a second feature map of a previous frame, and a third feature map of a subsequent frame; obtain a forward attention map based on the first feature map and the second feature map, and obtain a backward attention map based on the first feature map and the third feature map; perform collaborative attention calculation on the forward attention map and the backward attention map to obtain a collaborative attention map; perform weighted summation of the collaborative attention map with the second feature map to obtain previous context information, and perform weighted summation of the collaborative attention map with the third feature map to obtain subsequent context information; fuse the previous context information, the subsequent context information, and the first feature map to obtain current frame features; and generate the video sequence based on the current frame features; The video sequence generation module includes: A motion prediction unit, configured to predict motion information between frames, wherein the motion information is input as a condition into the spatiotemporal collaborative network; The spatiotemporal feature fusion and propagation unit is used to concatenate the feature maps of the previous and next frames in the channel dimension and feed them into the subsequent processing layer. It also performs element-wise addition of the features of the previous and next frames or performs weighted summation based on the attention weights. It also uses a gating mechanism to control the propagation and update of historical information and uses convolution kernels that can process both spatial and temporal dimensions. The video sequence generation module obtains the first feature map of the current frame t, the second feature map of the previous frame t-1, and the third feature map of the next frame t+1: Obtain a query by linearly transforming or otherwise mapping the first feature map of the current frame; The second feature map of the previous frame is linearly transformed to obtain a key and a value; The third feature map of the next frame is linearly transformed to obtain the key K and value V.

2. The educational system according to claim 1, wherein: The high-fidelity spatiotemporal interaction engine is also used to select three-dimensional Gaussian sputtering or neural radiation field rendering technology to generate the high-fidelity three-dimensional model in combination with device performance level, user preference settings and content metadata.

3. The educational system according to claim 1, wherein: During the display of the high-fidelity three-dimensional model, the high-fidelity spatiotemporal interaction engine is further configured to switch between the three-dimensional Gaussian sputtering and neural radiation field rendering techniques based on monitored performance indicators.

4. The educational system according to claim 1, wherein: The audio-visual generation module is also used to extract and align features of the text description and the video sequence to obtain aligned visual features and text features; and input the visual features and the text features into the audio generation model to generate audio that is synchronized in time with the video sequence and matches the meaning, thereby obtaining the target audio and video.

5. The educational system according to claim 1, wherein: The compensation module includes: A feature compensation unit is used to extract features from the original image during dynamic teaching or children's creation to obtain initial features, interact the quantity and object information in the text prompt with the initial features to obtain a compensated feature vector, and fuse the compensated feature vector with the initial features to obtain enhanced image features; The quantity-aware attention unit is used to extract information about each object from the enhanced image features, and to perform attention interaction between each object information and the current feature map to obtain quantity-aware features, and inject the quantity-aware features into the corresponding feature map.

6. The educational system according to claim 1, wherein: The educational system also includes: The intelligent management and control center uses a framework for decoupling internal and external factors to analyze children's interaction logs, learning choices, and usage time; as well as monitor interaction patterns with cloud-based artificial intelligence services to detect potential black box adversarial attack attempts in real time.

7. The educational system according to claim 6, characterized in that The intelligent management and control center is also used to use digital twins for simulation prediction and adjust strategies in advance.

8. The educational system according to claim 1, wherein: The educational system also includes: Adaptive learning and full-process security protection module for personalized adaptive learning, multi-level active security, and continuous health monitoring.

Citation Information

Patent Citations

  • Controllable video generation method and system based on multi-modal fusion

    CN119091362A

  • Audio-driven three-dimensional digital human generation method and system based on neural radiation field

    CN119888023A