Education system for multi-dimensional perception and intention recognition
By building an immersive cognitive environment with high fidelity, consistent timing, and multi-modal collaboration, and using multi-dimensional perception and intention recognition technology, the shortcomings of existing educational equipment in terms of interaction dimensions, cognitive depth, security protection and content ecology are solved, efficient mapping and interaction of physical input to digital twins are achieved, and refined content recommendation and security management are provided.
Patent Information
- Application Number
- CN202510640195.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-19
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2045-05-19
AI Technical Summary
The existing children's education equipment has significant shortcomings in terms of interaction dimension, cognitive depth, security protection and content ecology, making it difficult to achieve truly immersive and personalized interaction.
By building an immersive cognitive environment with high fidelity, consistent timing, and multi-modal collaboration, using multi-dimensional perception and intention recognition technology, it realizes efficient mapping and interaction from physical input to dynamic digital twins, and combines deep motivation understanding for intelligent management and security management.
It realizes efficient mapping and interaction from physical input to dynamic digital twins, provides refined and forward-looking content recommendation and security management, and supports on-demand creation of personalized content, improving the interactivity and security of the education system.
Smart Images

Figure CN120164149A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of education systems, and particularly to an education system for multi-dimensional perception and intention recognition. Background Art
[0002] Existing children's educational devices have significant deficiencies in terms of interaction dimension, cognitive depth, safety protection, and content ecosystem: Single interaction dimension: Traditional devices mostly stay at the level of static graphics and texts or preset animations, lacking real-time perception and dynamic response to children's physical postures, environmental situations, and internal motivations, and it is difficult to achieve true immersive and personalized interactions.
[0003] Insufficient cognitive depth: Learning content is mostly one-way indoctrination, unable to effectively distinguish the internal interests and external incentives behind children's behaviors, resulting in personalized recommendations and guidance being superficial and the efficiency of knowledge internalization being low.
[0004] Passive safety protection: Most devices only focus on content filtering and usage duration, and are weak in defending against spoofing attacks on access authentication (such as face recognition) and black-box probing and adversarial attacks against the artificial intelligence model itself, presenting potential safety risks.
[0005] Rigid content ecosystem: Lack the capabilities of dynamic generation, editing, and multi-modal fusion (especially sequential audio and video), with a single content form, difficult to meet the growing personalized and contextual learning needs, and also restricting the secondary creation of parents / educators. Summary of the Invention
[0006] Based on the above background, the core technical problems to be solved are upgraded to: How to construct an immersive cognitive environment with high fidelity, consistent timing, and multi-modal collaboration? It is necessary to break through static graphics and texts and simple animations, and achieve efficient mapping and interaction from physical inputs (picture cards, postures) to dynamic and timing-consistent digital twins (3D models, videos, audios).
[0007] How to achieve active and adaptive intelligent control and guidance based on in-depth motivation understanding? It is necessary to go beyond simple rule restrictions, deeply understand the internal and external driving factors of children's behaviors, and combine digital twin simulation to achieve refined and forward-looking content recommendation and safety management.
[0008] How to construct an end-to-end active safety defense system covering the entire process from access to interaction? It is necessary to defend against multi-dimensional threats from identity forgery to model probing / attack, especially against black-box query attacks during the interaction process.
[0009] How to endow the device with the capabilities of dynamic content generation, editing, and multimodal fusion while ensuring the consistency of content (especially the quantity)? It is necessary to integrate generative artificial intelligence and editing technologies to support the on-demand creation of personalized content.
[0010] The education system for multi-dimensional perception and intention recognition provided by this application can achieve efficient mapping and interaction from physical inputs (picture cards, gestures) to dynamic and temporally consistent digital twins (3D models, videos, audio).
[0011] In a first aspect, this application provides an education system for multi-dimensional perception and intention recognition. The education system includes: a high-fidelity spatio-temporal interaction engine for performing pose estimation and picture card recognition on the collected image stream using a dual attention ray scoring network to obtain the physical interaction methods between children and the device and picture cards, and the picture card recognition results, and generating a high-fidelity 3D model according to the picture card recognition results in combination with the 3D Gaussian sputtering / neural radiance field rendering technology; a video sequence generation module for generating a video sequence according to the high-fidelity 3D model using a spatio-temporal collaboration network when a video sequence needs to be generated; an audio-visual generation module for inputting the text description and the video sequence into an audio generation model to generate audio that is temporally synchronized and meaning-matched with the video sequence to obtain the target audio-visual; a compensation module for editing the target audio-visual during dynamic teaching or children's creation.
[0012] Among them, the high-fidelity spatio-temporal interaction engine is also used to select the 3D Gaussian sputtering or neural radiance field rendering technology to generate a high-fidelity 3D model in combination with the device performance level, user preference settings, and metadata of the content.
[0013] Among them, during the process of displaying the high-fidelity 3D model, the high-fidelity spatio-temporal interaction engine is also used to switch between the 3D Gaussian sputtering and neural radiance field rendering technologies according to the monitored performance metrics.
[0014] Among them, the video sequence generation module is also used to obtain the first feature map of the current frame, the second feature map of the previous frame, and the third feature map of the next frame; obtain the forward attention map according to the first feature map and the second feature map, and obtain the backward attention map according to the first feature map and the third feature map; perform collaborative attention calculation on the forward attention map and the backward attention map to obtain the collaborative attention map; respectively weighted sum the collaborative attention map with the second feature map to obtain the previous context information, and respectively weighted sum the collaborative attention map with the third feature map to obtain the subsequent context information; fuse the previous context information, the subsequent context information, and the first feature map to obtain the current frame feature; generate a video sequence according to the current frame feature.
[0015] Among them, the video sequence generation module includes: a motion prediction unit for predicting the motion information between frames, and the motion information is input as a condition into the spatio-temporal cooperation network; a spatio-temporal feature fusion and propagation unit for concatenating the feature maps of the front and back frames in the channel dimension, sending them to the subsequent processing layer, and performing element-wise addition of the features of the front and back frames or weighted summation according to the attention weights, and using a gating mechanism to control the propagation and update of historical information, and using a convolution kernel that can process both spatial and temporal dimensions.
[0016] Among them, the video-audio generation module is also used for feature extraction and alignment of the text description and the video sequence to obtain the aligned visual features and text features; and inputting the visual features and text features into the audio generation model to generate audio that is synchronized with the video sequence in time and meaning-matched, so as to obtain the target video-audio.
[0017] Among them, the compensation module includes: a feature compensation unit for extracting features from the original image during dynamic teaching or children's creation to obtain initial features, and interacting the quantity and object information in the text prompt with the initial features to obtain a compensated feature vector, and fusing the compensated feature vector and the initial features to obtain enhanced image features; a quantity-aware attention unit for extracting information about each object from the enhanced image features, and performing attention interaction between each object information and the current feature map to obtain quantity-aware features, and injecting the quantity-aware features into the corresponding feature maps.
[0018] Among them, the education system further includes: an intelligent management and control center for analyzing children's interaction logs, learning selections, and usage durations by adopting an internal / external factor decoupling framework; and monitoring the interaction mode with the cloud artificial intelligence service to detect potential black-box adversarial attack attempts in real time.
[0019] Among them, the intelligent management and control center is also used for simulation prediction using digital twins to adjust the strategy in advance.
[0020] Among them, the education system further includes: an adaptive learning and full-process security protection module for personalized adaptive learning, multi-level proactive security, and continuous health monitoring.
[0021] The beneficial effects of this application are as follows: Different from the prior art, the education system for multi-dimensional perception and intention recognition provided by this application includes: a high-fidelity spatio-temporal interaction engine, which is used to perform pose estimation and card recognition on the collected image stream by using a dual attention ray scoring network, obtain the physical interaction methods between children, devices, and cards and the card recognition results, and according to the card recognition results, generate a high-fidelity 3D model by combining the three-dimensional Gaussian sputtering / neural radiance field rendering technology; a video sequence generation module, which is used to generate a video sequence by using a spatio-temporal cooperation network according to the high-fidelity 3D model when a video sequence needs to be generated; an audiovisual generation module, which is used to input the text description and the video sequence into an audio generation model to generate audio that is temporally synchronized and meaning-matched with the video sequence, obtaining the target audiovisual; a compensation module, which is used to edit the target audiovisual during dynamic teaching or children's creation, and can realize efficient mapping and interaction from physical input (cards, poses) to dynamic and temporally consistent digital twins (3D models, videos, audio), as well as realize refined and forward-looking content recommendation and security management, and support the on-demand creation of personalized content. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] In order to more clearly illustrate the technical solutions in the embodiments of this application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of this application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings. Among them: Figure 1 It is a schematic structural diagram of an embodiment of the education system for multi-dimensional perception and intention recognition provided by this application; Figure 2 It is a schematic structural diagram of an embodiment of the video sequence generation module provided by this application; Figure 3 It is a schematic structural diagram of an embodiment of the compensation module provided by this application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0023] The following will clearly and completely describe the technical solutions in the embodiments of this application with reference to the drawings in the embodiments of this application. It can be understood that the specific embodiments described herein are only used to explain this application, rather than limiting this application. Additionally, it should be noted that for the sake of description, only the parts related to this application rather than all the structures are shown in the drawings. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of this application.
[0024] References herein to "embodiments" mean that the particular features, structures, or characteristics described in connection with the embodiments can be included in at least one embodiment of the present application. The phrase appears in various places in the specification and does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art will explicitly and implicitly understand that the embodiments described herein can be combined with other embodiments.
[0025] Refer to Figure 1 , Figure 1 FIG. is a schematic structural diagram of an embodiment of a multi-dimensional perception and intention recognition education system provided by the present application. The education system 100 includes: a high-fidelity spatio-temporal interaction engine 10, a video sequence generation module 20, an audio-visual generation module 30, and a compensation module 40.
[0026] The high-fidelity spatio-temporal interaction engine 10 is used to perform pose estimation and card recognition on the collected image stream by using a dual attention ray scoring network, obtain the physical interaction modes between the child and the device and the card, and the card recognition result, and generate a high-fidelity three-dimensional model according to the card recognition result in combination with the three-dimensional Gaussian sputtering / neural radiance field rendering technology.
[0027] In some embodiments, the high-fidelity spatio-temporal interaction engine 10 is further used to select the three-dimensional Gaussian sputtering or neural radiance field rendering technology to generate a high-fidelity three-dimensional model in combination with the device performance level, user preference settings, and metadata of the content.
[0028] In some embodiments, during the process of displaying the high-fidelity three-dimensional model, the high-fidelity spatio-temporal interaction engine 10 is further used to switch between the three-dimensional Gaussian sputtering and neural radiance field rendering technologies according to the monitored performance metrics.
[0029] In some embodiments, the high-fidelity spatio-temporal interaction engine 10 is mainly involved in physical world perception and pose understanding, as well as dynamic content activation and timing generation.
[0030] Physical world perception and pose understanding are mainly reflected in: using a dual attention ray scoring network, the six-degree-of-freedom pose of the child / device can be accurately estimated in real time through a monocular visible light camera, and the physical interaction modes (distance, angle, viewing method) between the child and the device and the card can be accurately captured.
[0031] Dynamic content activation and timing generation are mainly reflected in: after card recognition, a high-fidelity three-dimensional model is generated in combination with the three-dimensional Gaussian sputtering / neural radiance field rendering technology.
[0032] Among them, 3D Gaussian sputtering does not rely on complex neural networks for volume rendering. Instead, it uses a large number (thousands, even millions) of 3D Gaussian functions (primitives) with specific attributes (3D position, covariance matrix i.e., shape / rotation, color, opacity) to accurately fit and represent 3D objects or scenes.
[0033] When the children's educational device successfully identifies a specific physical card (e.g., a card with an apple pattern) within its field of view through its integrated high-resolution camera using image recognition algorithms (e.g., a classifier or object detection model based on the convolutional neural network CNN), the system will trigger a crucial real-time 3D content generation and rendering process. The goal of this process is to instantaneously present a highly relevant, visually realistic, and interactive 3D digital model of the card content (apple) on the device screen.
[0034] The rendering process mainly includes content association and model loading, real-time rendering, and view synthesis and shading.
[0035] Content association and model loading are mainly reflected in: the card recognition result ("apple") serves as a query index. Based on this index, the system quickly loads the 3D Gaussian parameter set corresponding to "apple" (including data such as the positions, shapes, colors, and opacities of all Gaussian primitives) from a pre-built and optimized model library. This model library is usually obtained offline through multi-view image training or 3D scan data conversion.
[0036] Real-time rendering is mainly reflected in: leveraging the high parallel computing power of modern graphics processing units (GPUs), through a specially designed differentiable Gaussian rasterization rendering pipeline, millions of loaded 3D Gaussian primitives are efficiently "sputtered" or projected onto the 2D screen space.
[0037] View synthesis and shading are mainly reflected in: the rendering pipeline will perform depth sorting and alpha blending on the projected Gaussian primitives according to the current viewing perspective (which may be dynamically adjusted in combination with the real-time pose data of the child / device obtained), calculate the final color and opacity of each pixel, and form a high-quality, hole-free 2D image. The rendering process will consider the learned color information (usually represented by spherical harmonics to support view-dependent effects) and opacity to achieve realistic lighting and texture effects.
[0038] The advantages of 3D Gaussian sputtering technology are mainly reflected in: fast rendering speed, usually able to achieve real-time or even super-real-time frame rates, which is very suitable for applications that require instant interaction. The model representation is relatively intuitive, and specific areas can be modified (such as changing colors, local deformations).
[0039] Among them, the neural radiance field is mainly reflected in: training one or more multi-layer perceptron (MLP) neural networks to learn a continuous three-dimensional scene function. This function takes the coordinates of a three-dimensional space point (x, y, z) and an observation direction (θ, φ) as inputs, and outputs the volume density (σ) of the point and the view-dependent color radiance (c).
[0040] The rendering process mainly includes content association and model loading, view-based volume rendering, sampling and querying, and color integration.
[0041] Content association and model loading are mainly reflected in: the graphics card recognition result ("apple") is also used as an index, and the system loads the pre-trained neural radiance field network model weights associated with "apple". This model has been trained offline through a large number of rendering losses (comparing the rendered image with the real multi-view image).
[0042] View-based volume rendering is mainly reflected in: when rendering an image of a specific view, the system emits a large number of rays from the virtual camera through the scene.
[0043] Sampling and querying are mainly reflected in: along each ray, the system samples multiple three-dimensional space points. The coordinates of each sampled point and the ray direction are input into the loaded neural radiance field network to query the volume density and color of the point.
[0044] Color integration is mainly reflected in: using the classical volume rendering equation, numerically integrating (alpha compositing) the colors and densities of all sampled points along the ray path to calculate the color of the ray finally reaching the camera. The calculation results of all rays together constitute the final rendered two-dimensional image.
[0045] The advantages of the neural radiance field are mainly reflected in: usually being able to generate extremely high-quality and highly consistent rendering results, especially being good at handling complex lighting, reflection, and translucency effects. The model representation (network weights) is relatively compact. Figure 1 In some embodiments, an adaptive selection is made between the neural radiance field and the three-dimensional Gaussian splatting technology.
[0046] The adaptive selection strategy is as follows:
[0047] a. Offline content preparation and metadata tagging. a. Offline content preparation and metadata tagging.
[0048] Dual / multiple representation preparation: For each graphic card (or its corresponding three-dimensional object) in the content library, during the content production stage, not only an optimized three-dimensional Gaussian splatting representation (baseline standard) should be generated, but also one or more highly optimized neural radiance field variant representations should be considered for those core contents with high visual complexity, special optical effects, or particular importance.
[0049] Metadata tagging: Associate detailed metadata with each card content, including: content complexity, visual importance, interaction requirements, available formats, and performance benchmarks.
[0050] Content complexity: Tag whether the object contains complex transparency / reflection / refraction effects (high complexity) or is mainly a diffuse surface (low complexity).
[0051] Visual importance: Tag whether the content is a core teaching point with extremely high requirements for visual details.
[0052] Interaction requirements: Tag the expected main interaction modes (e.g., static observation, rapid perspective transformation, possible future editing requirements).
[0053] Available formats: List the representation formats in which this content is ready (e.g., 3D Gaussian sputtering optimization, neural radiance field optimization, 3D Gaussian sputtering base).
[0054] Performance benchmarks: The rendering performance of different formats on reference hardware can be benchmarked during the offline phase as reference data.
[0055] b. Device capacity assessment and user preference settings. Include device analysis at the first startup / installation and decision-making logic when loading the card.
[0056] Device analysis at the first startup / installation: When the application is first started, evaluate the hardware capabilities of the device (GPU model and performance, memory size, CPU performance, etc.). And the device can be divided into different performance levels (e.g., high, medium, low).
[0057] User / parent preference settings: Provide simple setting options to allow the user (or parent) to select a preference mode, e.g.: Performance priority: Force the use of 3D Gaussian sputtering.
[0058] Balanced mode: By default, use 3D Gaussian sputtering, but try higher-quality options when conditions permit.
[0059] Quality priority: On high-end devices, preferentially try neural radiance field variants for supported content.
[0060] The decision-making logic when loading the card is as follows: When the card is recognized, make a decision by combining the device performance level, user preference settings, and the metadata of the content: For low-end devices: Force load and use 3D Gaussian sputtering base or 3D Gaussian sputtering optimization.
[0061] For mid / high-end devices: Performance - Priority Mode: Load and use 3D Gaussian sputtering optimization.
[0062] Balanced / Quality - Priority Mode: Check available formats in the metadata.
[0063] If there is an optimized neural radiance field variant: Check whether the content complexity is "high" or the visual importance is "high".
[0064] Check whether the interaction requirement is mainly "static observation".
[0065] Check whether the user preference is "quality - priority" or "balanced".
[0066] If most of the above conditions are met and the device is high - end: Initially select to load neural radiance field optimization.
[0067] Otherwise: Select to load 3D Gaussian sputtering optimization (default option).
[0068] If only the 3D Gaussian sputtering format exists: Load and use 3D Gaussian sputtering optimization.
[0069] c. Runtime State Monitoring and Dynamic Switching (Runtime Adaptation).
[0070] Performance Monitoring: Real - time monitor key performance indicators: rendering frame rate (FPS), GPU / CPU load, memory occupancy, device temperature.
[0071] Dynamic Switching Triggers: For performance degradation: If the current rendering frame rate continuously drops below a preset threshold (e.g., <25 FPS), or the device temperature is too high, triggering the risk of frequency reduction.
[0072] For interaction mode change: If the user changes from static observation to quickly moving the device (requiring high FPS).
[0073] For system resource pressure: If other high - priority tasks are started in the background and resources need to be released.
[0074] For battery power / supply mode: Enter the low - battery mode or energy - saving mode.
[0075] Switching Logic: Downgrade Switching: If the current neural radiance field optimization is in use and any of the above conditions is triggered, the system should automatically and seamlessly switch to the pre - loaded or on - demand quickly loaded 3D Gaussian sputtering optimization. This is the most common switching scenario to ensure smoothness.
[0076] The video sequence generation module 20 is used to generate a video sequence according to the high - fidelity 3D model by using the spatio - temporal cooperation network when a video sequence needs to be generated.
[0077] In some embodiments, the video sequence generation module 20 is also used to obtain a first feature map of the current frame, a second feature map of the previous frame, and a third feature map of the next frame; obtain a forward attention map based on the first feature map and the second feature map, and obtain a backward attention map based on the first feature map and the third feature map; perform collaborative attention calculation on the forward attention map and the backward attention map to obtain a collaborative attention map; obtain previous context information by weighted summing the collaborative attention map with the second feature map, and obtain subsequent context information by weighted summing the collaborative attention map with the third feature map; fuse the previous context information, the subsequent context information, and the first feature map to obtain current frame features; generate a video sequence based on the current frame features.
[0078] In some embodiments, see Figure 2 The video sequence generation module 20 includes: a motion prediction unit 21 and a spatiotemporal feature fusion and propagation unit 22.
[0079] The motion prediction unit 21 is used to predict the motion information between frames, and the motion information is input into the spatiotemporal collaborative network as a condition.
[0080] The spatiotemporal feature fusion and propagation unit 22 is used to splice the feature maps of the previous and next frames in the channel dimension and send them to the subsequent processing layer, and to add the features of the previous and next frames at the element level or to perform weighted summation according to the attention weight, and to use the gating mechanism to control the propagation and update of historical information, and to use a convolution kernel that can process both spatial and temporal dimensions at the same time.
[0081] In this embodiment, a spatiotemporal collaborative network is introduced to ensure temporal continuity and content consistency across frames and suppress flickering and mutations when generating video sequences (such as story animations).
[0082] After children's educational equipment recognizes picture cards and generates high-fidelity 3D models, an important application scenario is to generate dynamic video sequences based on this model, such as telling a story, demonstrating a process, or making an animation. Directly rendering 3D models frame by frame (even if the model itself is fixed) or using simple interpolation methods often cannot meet the requirements of high-quality video and are prone to the following problems: Timing incoherence: objects move in a jerky and jumpy manner, lacking natural transitions.
[0083] Inconsistent content: The appearance of objects (such as lighting, texture details) changes slightly between adjacent frames without any reason, the background is unstable, or the identity characteristics of objects drift.
[0084] Visual artifacts: Flicker (rapid fluctuations in brightness or color) or sudden changes (sudden changes in the position, posture, or appearance of objects).
[0085] To overcome these challenges, a video sequence generation module 20 that can sense and utilize the temporal dependencies between frames is constructed through a spatio-temporal collaboration network. Its goal is to fully consider the information of the frames before and after each frame when generating it, so as to ensure the smoothness, stability, and consistency of the entire video sequence.
[0086] The key technical mechanisms include cross-frame information dependence modeling, bidirectional temporal attention / collaboration mechanism, spatio-temporal feature fusion and propagation, motion / optical flow guidance, and consistency loss function design.
[0087] The cross-frame information dependence modeling is introduced as follows: Core principle: Generating the video content (F_t) of the t-th frame not only depends on the state at the current moment t (such as the pose of the 3D model, the camera position, and the lighting settings), but also must explicitly depend on the information of its past (F_{t - 1}, F_{t - 2},...) and future (F_{t + 1}, F_{t + 2},...) frames. This context awareness is the basis for ensuring coherence.
[0088] Implementation method: Usually embedded in the architecture design of a video generation network (such as a U-Net variant based on the Diffusion Model, the generator of a video GAN, or a Transformer-based video model), cross-frame information interaction is achieved through specific layers or modules.
[0089] The bidirectional temporal attention / collaboration mechanism is introduced as follows: Mechanism details: Similar to the Transformer in natural language processing, an attention mechanism is introduced in the time dimension. When generating the t-th frame, this mechanism calculates the correlation weights between the features of the current frame and the features of adjacent frames (or more distant frames) before and after.
[0090] The importance of "bidirectional": Not only should we pay attention to past frames to inherit the motion trend and content state; we should pay more attention to future frames to anticipate upcoming actions or changes, so as to make a smoother transition.
[0091] The "collaboration" effect: It can be designed so that the attention calculations of the frames before and after affect each other, enabling the information of the past and future to more effectively collaborate in the generation of the current frame and jointly determine which temporal information is the most important.
[0092] Application: The attention mechanism can act on pixel-level features, high-level semantic features, or latent space representations, dynamically aggregating the most relevant temporal context information to guide the generation of the details of the current frame.
[0093] For example, a bidirectional temporal collaboration attention module can be set. The specific introduction is as follows: 1. Module objectives and motivations: The core objective of the bidirectional temporal collaborative attention module is to address the temporal consistency issue in video sequence processing (especially in generation or fusion tasks). As mentioned before, processing video frames independently can easily lead to problems such as flickering and content jumps. The bidirectional temporal collaborative attention module aims to enable the feature representation of the current frame (t) to simultaneously perceive and utilize the information of its past (t - 1) and future (t + 1) frames, and intelligently aggregate this temporal context through a collaborative attention mechanism to generate smoother and more coherent video results.
[0094] 2. Core idea: Bidirectional and collaborative.
[0095] Among them, the bidirectional aspect is mainly reflected in: Different from only considering past information (such as traditional RNN) or only performing local inter-frame processing (such as simple 3D convolution), the bidirectional temporal collaborative attention module explicitly looks forward (the relationship between t and t - 1) and backward (the relationship between t and t + 1) simultaneously. This is crucial for understanding the continuity of motion, the smooth transition of lighting, and avoiding the sudden appearance or disappearance of content.
[0096] The collaborative aspect is mainly reflected in: This is not just simply combining forward attention and backward attention. The key innovation of the bidirectional temporal collaborative attention module is that it believes there is an interdependent and enhancing relationship between forward attention and backward attention. That is to say, the degree of attention of a certain region in the current frame to the past frame should affect its attention to the future frame, and vice versa. This "collaborative" mechanism can better focus on the truly important and continuous changing regions in time series.
[0097] 3. The implementation process is as follows: Obtain the feature map of the current frame t, as well as the feature maps of the previous frame t - 1 and the next frame t + 1: Step 1: Generate Query, Key, and Value.
[0098] Obtain the query by linearly transforming (or other mappings) the current frame feature map.
[0099] Obtain the key and value by linearly transforming the previous frame feature map.
[0100] Obtain the key K and value V by linearly transforming the next frame feature map.
[0101] Here, Q, K, and V are concepts in the standard self-attention mechanism, representing different functional roles of features.
[0102] Step 2: Calculate the forward and backward attention maps.
[0103] The forward attention map is obtained as follows: Calculate the similarity (usually dot product) between the query Q of the current frame and the key K of the previous frame, and normalize it through the Softmax function to obtain a weight map representing "how much each position in the current frame should focus on the corresponding position information in the previous frame". For example, the following formula can be used: A^{t - 1}=softmax((Q^t)(K^{t - 1})^T / sqrt(d_k)).
[0104] The backward attention map is obtained as follows: Similarly, calculate the similarity between the query Q of the current frame and the key K of the next frame, and normalize it through Softmax to obtain a weight map representing "how much the current frame should focus on the information of the next frame". For example, the following formula can be used: A^{t + 1}=softmax((Q^t)(K^{t + 1})^T / sqrt(d_k)).
[0105] sqrt(d_k) is a scaling factor used to stabilize training.
[0106] Step 3: Co - attention calculation.
[0107] Element - wise multiply the forward attention map and the backward attention map. Only when a region receives high attention in both the forward and backward attentions will it maintain a high value in the multiplied map. This highlights those regions that are continuously important or undergo smooth transitions over time.
[0108] Normalize the multiplication result through Softmax again to obtain the final co - attention map. For example, the following formula can be used: A_{co}=softmax(A^{t - 1}⊙A^{t + 1}).
[0109] The co - attention map reflects a more refined attention distribution that incorporates bidirectional dependencies between the front and back.
[0110] Step 4: Weighted aggregation and feature enhancement.
[0111] Use the co - attention map to perform weighted summation on the values V of the previous frame and the values V of the next frame respectively. This is equivalent to extracting the most important context information according to the co - attention. For example, the following formula can be used: Context_{t - 1}=A_{co}*V^{t - 1} (matrix multiplication).
[0112] Context_{t + 1}=A_{co}*V^{t + 1} (matrix multiplication).
[0113] The extracted context information Context_{t-1} and Context_{t+1} (usually processed by a feed-forward neural network FFN first) are fused with the original query Q (or some transformation thereof) of the current frame. The fusion methods can be addition, concatenation and post-processing, etc. For example, the following formula can be used: F̂^t_f =Q^t ⊕ FFN(Context_{t-1}) ⊕ FFN(Context_{t+1}) F̂^t_f is the feature of the current frame enhanced by the bidirectional temporal collaborative attention module, which contains the bidirectional temporal context information after intelligent aggregation.
[0114] The advantages of the bidirectional temporal collaborative attention module are as follows: Improve temporal coherence: By explicitly using the information of the previous and next frames and using collaborative attention to emphasize continuously important regions, it effectively suppresses flickering and content mutations.
[0115] Enhance dynamic perception: For moving objects or changing scene elements, it can better capture their continuous trajectories and appearance changes.
[0116] Intelligent information screening: The collaborative attention mechanism can help the model ignore the interference information that only appears briefly or is irrelevant in one direction (forward or backward), and focus on the temporal context that is truly helpful for maintaining consistency.
[0117] Architectural flexibility: The bidirectional temporal collaborative attention module can be embedded as a module into the encoder or decoder stage of various video processing networks, especially for enhancing feature representations that require temporal perception.
[0118] Summary: Through innovative bidirectional temporal information processing and collaborative attention mechanism, the bidirectional temporal collaborative attention module enables the model to intelligently and interdependently refer to the content of its previous and next frames when processing each frame of the video sequence. It calculates the forward and backward attention maps, conducts collaborative interaction through element-wise product and Softmax, and finally uses the obtained collaborative attention map to weighted aggregate the value information of the previous and next frames and integrate it into the feature representation of the current frame. This design significantly improves the temporal coherence and content consistency of the generated video.
[0119] The introduction of spatio-temporal feature fusion and propagation is as follows: Detailed mechanism: At the middle layer of the video generation network, the feature representations from adjacent frames are explicitly fused. This can be achieved in various ways, such as feature concatenation, feature addition / weighting, recurrent structures, 3D convolution / spatio-temporal convolution.
[0120] Feature concatenation is mainly reflected in: concatenating the feature maps of the previous and next frames in the channel dimension and feeding them into the subsequent processing layer.
[0121] Feature addition / weighting is mainly reflected in: performing element-wise addition of the features of the previous and next frames or weighted summation according to attention weights.
[0122] The loop structure is mainly reflected in: introducing a gating mechanism similar to RNN or LSTM to control the propagation and update of historical information.
[0123] 3D convolution / spatiotemporal convolution is mainly reflected in: using a convolution kernel that can process both spatial and temporal dimensions simultaneously.
[0124] The role is to ensure that when generating the current frame, the network "sees" and utilizes the spatial and appearance information of the previous and next frames, thereby enforcing the continuity of content (such as texture, lighting).
[0125] The introduction of motion / optical flow guidance is as follows: Mechanism: An additional module can be introduced to predict the motion information between frames (such as the optical flow field). Taking the predicted motion information as a conditional input into the generation network can more precisely guide the movement of pixels or features and generate a more physically consistent motion trajectory.
[0126] The design of the consistency loss function is introduced as follows: Objective: When training a video generation model, in addition to ensuring the generation quality of a single frame, a loss term that enforces temporal consistency must be added.
[0127] Common forms include inter-frame difference loss, optical flow consistency loss, and variational consistency loss.
[0128] The inter-frame difference loss is mainly reflected in: penalizing excessive pixel differences, feature differences, or perceptual differences (such as using the LPIPS loss) between adjacent frames in the generated video sequence.
[0129] The optical flow consistency loss is mainly reflected in: if optical flow is used, it can be required that the image / features warped from the previous frame to the current frame according to the optical flow be as similar as possible to the directly generated current frame image / features.
[0130] The variational consistency loss is mainly reflected in: This is a more refined loss that can distinguish between dynamic foregrounds and static backgrounds, enforce higher temporal smoothness in the background region, while allowing foreground objects to change according to their motion trajectories, but requires their changes to be consistent with the source video (if applicable) or the expected animation path.
[0131] The role is to make the model learn to generate videos that inherently have temporal coherence by explicitly optimizing these consistency objectives during training.
[0132] Summary: By deeply integrating the idea of spatio-temporal collaboration network, especially introducing cross-frame information dependence modeling, bidirectional temporal attention / collaboration mechanism, spatio-temporal feature fusion and propagation, and using a consistency loss function for constraint in training, the problems of temporal incoherence, content inconsistency, flickering and mutation caused by independent frame generation can be effectively solved. This will enable children's educational devices to generate smooth, stable, and highly consistent story animations or teaching demonstration videos, significantly enhancing the immersion and the quality of the learning experience.
[0133] The audio-visual generation module 30 is used to input the text description and the video sequence into the audio generation model to generate audio that is temporally synchronized and meaning-matched with the video sequence, obtaining the target audio-visual content.
[0134] Among them, the audio-visual generation module 30 is also used to extract and align the features of the text description and the video sequence to obtain the aligned visual features and text features; and input the visual features and text features into the audio generation model to generate audio that is temporally synchronized and meaning-matched with the video sequence, obtaining the target audio-visual content.
[0135] In some embodiments, using text-assisted audio-visual generation technology, high-quality audio (sound effects, narration) that is precisely synchronized and semantically matched with the video content (and optionally text descriptions generated by large language models) is automatically generated based on the video content to achieve audio-visual collaboration.
[0136] In children's educational devices, simply generating visual content (such as animations, 3D model demonstrations) is often not enough to provide a complete immersive experience. Sound (including environmental sound effects, object movement sounds, character voices, background music, narration, etc.) is an important part of information transmission and emotional rendering. Traditional video dubbing or sound effect production processes usually require manual completion, which is costly, time-consuming, and difficult to match with dynamic or changing visual content in real time.
[0137] Video-to-audio generation technology aims to automatically generate corresponding audio according to video frames. However, a pure V2A model may face the following challenges: Semantic ambiguity: It is sometimes difficult to accurately determine what kind of sound should be generated based solely on visual information. For example, seeing a person opening their mouth may be speaking, singing, yawning, or even making a silent expression.
[0138] Lack of high-level abstract understanding: It is difficult to generate audio that requires understanding the story plot or abstract concepts, such as narration.
[0139] Difficulty in fine control: It is difficult to fine-tune the generated audio according to specific requirements (such as emphasizing a certain sound effect or changing the tone of the narration).
[0140] To address these issues, this application introduces text assistance, using text descriptions as a powerful auxiliary information source to jointly guide the generation of audio along with video content.
[0141] The core method lies in text-enhanced video-audio collaborative generation. By fusing visual information and text descriptions, more abundant and explicit semantic and temporal constraints are provided for the audio generation model, thereby generating high-quality audio that is precisely synchronized with the video content in time and highly matched in meaning.
[0142] The technical implementation process is as follows: a. Multi-modal input preparation: Video input: The video clip to be dubbed or the dynamically generated video sequence.
[0143] Text input: This is the key auxiliary information. The text can have various sources and forms: such as manual annotation / script, automatic video description, descriptions / narratives generated by large language models (LLMs), simple keywords / tags.
[0144] Manual annotation / script is the most direct way, providing detailed scene descriptions, action explanations, dialogue lines, narrative texts, etc.
[0145] Automatic video description is mainly reflected in: using advanced video understanding models (possibly based on Transformer or multi-modal large language models) to automatically generate natural language descriptions of video content. The descriptions can cover scenes, objects, actions, and even possible emotions or atmospheres.
[0146] Descriptions / narratives generated by large language models (LLMs) are mainly reflected in: combining video content and preset educational goals or story outlines, and using large language models (such as GPT-4, LLaMA, etc.) to generate more creative and teaching-demand-compliant text descriptions or complete narrative manuscripts. This offers great potential for personalized and dynamic content generation.
[0147] Simple keywords / tags are mainly reflected in: It can also be more concise text information, such as "bird calls", "car engine sounds", "lively background music", etc.
[0148] b. Multi-modal feature extraction and alignment. Specifically, it includes feature extraction and cross-modal alignment.
[0149] Feature extraction is mainly reflected in: respectively using pre-trained models to extract the feature representations of video, text, and (if necessary, as the training target or reference) audio.
[0150] Video encoder: Such as SlowFast, ViT variants, to extract the visual feature sequence Ev containing spatio-temporal information.
[0151] Text Encoder: A text tower such as BERT, T5, or CLIP extracts the semantic feature representation El of the text.
[0152] Audio Encoder: Such as PANNs or CLAP, it extracts the feature sequence Ea of the audio mel spectrogram (used as the target during training).
[0153] Cross-modal alignment is a crucial step in achieving audiovisual-text collaboration. It is mainly reflected in: Using methods such as contrastive learning to map the extracted Ev, El (and Ea during training) into a shared latent space. By maximizing the similarity of matching video-text, video-audio, and text-audio sample pairs and minimizing the similarity of mismatched sample pairs, the features describing the same semantic concept between different modalities are made to approach each other in the latent space. This ensures that text information can effectively "understand" and "relate" to the corresponding visual content.
[0154] c. Text-Assisted Audio Generation Model: Model Architecture: Usually, a conditional generation model is adopted, such as a Conditional Diffusion Model (the Latent Diffusion Model LDM is an architecture that may be used in TA-V2A), a Conditional GAN, or a Transformer-based sequence-to-sequence model.
[0155] Condition Injection: The aligned visual feature Ev and text feature El are used as conditions and input into the generation model. The injection methods can be diverse: For example, cross-attention, conditional embedding concatenation / addition, and multi-guided.
[0156] Cross-attention is mainly reflected in: In the U-Net of the Diffusion model or the decoder of the Transformer, let the intermediate features during audio generation pay attention to Ev and El through the cross-attention mechanism, so as to generate corresponding audio details according to the visual content and text description.
[0157] Conditional embedding concatenation / addition is mainly reflected in: Concatenating or adding the global representation or time-step related representation of Ev and El with the noise or intermediate state during the generation process to guide the generation direction.
[0158] Multi-guided is mainly reflected in: During the inference (sampling) stage of the Diffusion model, the guidance signals based on visual conditions, text conditions, and even negative text conditions (undesired sounds) can be calculated independently and weighted and combined to achieve finer control of the generation results.
[0159] Generation Objective: The goal of the model is to generate an audio representation (e.g., a mel-spectrogram ẑ0) that is highly consistent with the input Ev and El both temporally and semantically.
[0160] d. Audio Synthesis: This is mainly reflected in: converting the audio representation (such as Mel-spectrogram) output by the generative model into a final audible waveform file through a vocoder (such as HiFi-GAN, DiffWave, etc.).
[0161] The audio-visual synergy effect and advantages are: Precise timing synchronization: Powerful video encoders and cross-modal alignment mechanisms ensure that generated audio events (such as footsteps and collisions) are precisely timed to the corresponding actions in the video.
[0162] Enriched semantic matching: The introduction of text description greatly enhances the semantic understanding ability: For example, specific sound effects can be generated: if the text explicitly states "cat's meow", the model can generate the sound of a cat's meow instead of just guessing based on vague visuals.
[0163] For example, it can generate narration / dialogue: If the text input is a narration or dialogue script, the model can generate the corresponding speech (which may require the combination of TTS technology).
[0164] For example, it can generate ambient sound / background music: the text description "cheerful scene" can guide the model to generate background music or ambient sound that matches the mood.
[0165] For example, disambiguation: for visually ambiguous actions, text can provide clear explanations, guiding the generation of correct audio.
[0166] High-quality audio output: Advanced generative models (such as the Diffusion Model) and vocoders can produce high-fidelity, low-noise audio waveforms.
[0167] Flexible control and personalization: By modifying the input text description, users can easily adjust the generated audio content, style or emphasis to achieve a highly personalized audio-visual experience. For example, they can change the tone of the narration or add specific environmental sound effects.
[0168] Summary: Text-assisted audio-visual generation technology achieves precise timing synchronization and high semantic matching of audio generation by deeply integrating video content and text description, using cross-modal alignment and conditional generation models. It can automatically generate high-quality sound effects, narration or background music, and seamlessly collaborate with (dynamically generated) video content, greatly improving the immersion, information transmission efficiency and personalized creation capabilities of children's educational equipment, and is a key technical support for building a complete audio-visual experience. The introduction of large language models further enhances the generation capability and flexibility of text descriptions, making content creation more intelligent and automated.
[0169] The compensation module 40 is used to edit the target audio and video accordingly during dynamic teaching or children's creation.
[0170] In some embodiments, see Figure 3 The compensation module 40 includes: a feature compensation unit 41 and a quantity-aware attention unit 42.
[0171] The feature compensation unit 41 is used to extract features from the original image during dynamic teaching or children's creation to obtain initial features, and to interact the quantity and object information in the text prompt with the initial features to obtain a compensated feature vector, and to fuse the compensated feature vector and the initial features to obtain enhanced image features.
[0172] The quantity-aware attention unit 42 is used to extract information about each object from the enhanced image features, and to perform attention interaction between each object information and the current feature map to obtain quantity-aware features, and to inject the quantity-aware features into the corresponding feature map.
[0173] In some embodiments, a quantity-aware attention mechanism (quantity-aware attention unit 42) and a feature compensation module (feature compensation unit 41) are integrated to support quantity-preserving multi-object editing (such as adding or removing objects, changing attributes) of generated content (images / scenes) for dynamic teaching or creation.
[0174] In children's educational devices, it is not enough to simply display content statically. Dynamic teaching (e.g., adding and subtracting objects when demonstrating addition and subtraction) or allowing children to make simple creations (e.g., adding or removing toys from the scene) can greatly enhance the interactivity and fun of learning. However, existing image / scene editing techniques, especially those based on large-scale pre-trained models (such as Stable Diffusion), often encounter the following challenges when dealing with multi-object scenes: Quantity Distortion: When the editing command requires increasing or decreasing the quantity of a specific object, the model may not be able to accurately control the final generated quantity, resulting in incorrect quantity (more or less).
[0175] Attribute Bleeding / Loss: When editing an object in a scene, its attributes (color, shape, pose, etc.) may "bleed" onto other objects, or cause the attributes of other objects to be lost or change unexpectedly.
[0176] Identity Inconsistency: During the editing process, the identity characteristics of an object may drift, and it no longer appears to be the original object.
[0177] Dependence on auxiliary tools: Some methods rely on auxiliary tools such as masks to specify the editing area, which increases the complexity of the operation and is not suitable for children or ordinary users.
[0178] To solve the above problems, a Feature Compensation Module (Feature Compensation Unit 41) and a Quantity-Aware Attention Mechanism (Quantity-Aware Attention Unit 42) are introduced. The core idea is to improve the feature representation and attention mechanism of the model without relying on additional auxiliary tools (such as masks), enabling it to clearly perceive and maintain the quantity of objects in the scene and ensuring that the attributes of each object are independent and clearly distinguishable during the editing process.
[0179] The technical process is as follows: a. Feature Compensation Module - Enhancing the separation of object attributes: Objective: To solve the problem of attribute confusion caused by pre-trained models (such as the image encoder of CLIP) when extracting image features of multiple objects, ensuring that the attribute features of each object are clear and independent, and laying a foundation for subsequent quantity perception and precise editing.
[0180] The mechanism is as follows: Input: The original image I and a text prompt cq containing object names and quantity information (e.g., "three red apples and two green pears"). This prompt is the key input to FeCom.
[0181] Feature extraction: Use the image encoder of models such as CLIP to extract the initial features of the image CLIP(I).
[0182] Feature interaction and compensation: Design a feature attention mechanism. This mechanism uses the quantity and object information in the text prompt cq as queries (Query) to interact with the image features CLIP(I). The interaction process aims to identify the object attribute information that is not fully expressed or distinguishable in the image features and generate a compensation feature vector Ic. This compensation vector contains the object-specific attribute information that needs to be "strengthened" or "clarified".
[0183] Feature Enhancement: The compensated feature Ic is fused with the original image feature CLIP(I) (such as weighted summation) to obtain the enhanced image feature Ig (Ig = CLIP(I) + λIc). The attribute representation of each object in Ig is clearer and more separated, reducing the possibility of attribute confusion in subsequent processing.
[0184] Function: The feature compensation module actively "compensates" for the deficiency of the image encoder in distinguishing multi-object attributes by leveraging the prior knowledge (object names and quantities) in the text prompt, generating a high-quality feature representation that is more suitable for multi-object editing.
[0185] b. Quantity-Aware Attention Mechanism - Achieving Quantity Preservation and Precise Editing: Objective: Enable the subsequent generation / editing network (usually a U-Net-based Diffusion model) to explicitly "perceive" the existence and quantity of each independent object in the scene and maintain this quantity relationship during the generation process, while precisely applying the editing instructions to the target object.
[0186] The mechanism is as follows: Input: The enhanced image feature Ig (from the feature compensation module) and the noise / intermediate feature map z^t during the Diffusion process (at a specific layer, such as the 4th U-Net block B4 selected by the feature compensation module).
[0187] Object Information Extraction: QTTN first includes an extraction module (such as a fully connected layer FC or a small convolutional network) to further extract information about each independent object from the enhanced feature Ig. This is not only spatial location information, but more importantly, instance-level information that can distinguish "this is an apple" and "this is another apple".
[0188] Attention Interaction: Perform attention interaction (such as cross-attention) between the extracted object instance information and the current noise / feature map z^t. z^t provides the query (Query, Qz). The extracted object instance information Et(Ig) provides the key (Key, Kg) and value (Value, Vg). By calculating softmax(Qz * Kg^T / sqrt(dk)) * Vg, each spatial position of z^t will dynamically aggregate information from these object instances according to its association degree with different object instances.
[0189] Quantity perception information injection: Inject the quantity perception feature Vnew after attention-weighted aggregation into the feature map of this layer of the U-Net (for example, through post-processing such as addition or channel concatenation: z^{t + 1}_layer = Attention(Qz, Ki, Vi) + βVnew).
[0190] Function: The quantity perception attention module enables the U-Net to "know" at each step of denoising (i.e., generating / editing) how many objects should be in the scene, approximately where each object is located, and what attributes it should have. This explicit quantity and instance perception ability enables the model to, when executing editing instructions (such as "turn one of the apples blue" or "add an apple"): Maintain quantity: If the editing instruction does not involve a change in quantity, the model tends to maintain the original number of objects.
[0191] Precisely control quantity: If the instruction involves an increase or decrease, the model can execute more accurately based on the injected quantity information.
[0192] Precisely apply attributes: Precisely apply attribute changes (such as color) to one or more specified object instances without affecting other objects.
[0193] c. Application to dynamic teaching or creation: Dynamic teaching: For example, in early math education, a teacher or system can drive the change in the number of apples in a scene through text instructions (such as "there are 3 apples on the table now, put 2 more apples"), and explain the addition process with a voiceover (generated from text-assisted video and audio).
[0194] Children's creation: Children can add, delete, or modify objects in a virtual scene through simple instructions (possibly converted from speech recognition into text) (such as "change the red building block to yellow", "add another car"), for exploratory learning and creative expression.
[0195] The technical effects are as follows: Quantity-preserving editing: The core advantage is the ability to precisely maintain or control the number of objects during the editing process.
[0196] Attribute fidelity and separation: Effectively prevent interference with other objects when editing one object, and maintain the independence and consistency of attributes.
[0197] No need for auxiliary tools: The editing process does not rely on additional information provided by the user such as masks, simplifying the operation process.
[0198] Enhance interactivity and creativity: Provide strong technical support for dynamic teaching demonstrations and children's independent creation, enriching the possibilities of educational applications.
[0199] Summary: The children's educational device integrating the feature compensation module and the quantity-aware attention mechanism can achieve quantity-preserving multi-object editing when processing image or scene content. The feature compensation module enhances the separation of object attributes using text prompts, while the quantity-aware attention mechanism enables the generation network to have explicit quantity and instance awareness capabilities. The combination of these two enables the device to respond to editing instructions such as adding or removing objects and changing specific object attributes, while maintaining the stability of other object attributes in the scene and the accuracy of the total number of objects (unless the instruction requires a change in quantity), greatly enhancing the ability of dynamic teaching demonstrations and children's creative expression.
[0200] Multi-modal information fusion: Drawing on multi-modal fusion strategies, integrate multi-source information such as vision (camera, screen), audition (microphone, speaker), gesture (output of gesture estimation algorithm), and interaction (touch screen / button) to build a more comprehensive scene understanding.
[0201] In some embodiments, the educational system 100 further includes: an intelligent control center for analyzing children's interaction logs, learning choices, and usage duration using an internal / external factor decoupling framework; and monitoring the interaction mode with cloud artificial intelligence services to detect potential black-box adversarial attack attempts in real time.
[0202] Among them, the intelligent control center is also used to perform simulation prediction using digital twins and adjust strategies in advance.
[0203] The motivation-driven intelligent control center is mainly reflected in: deep user profiling and intention understanding, intelligent decision-making and flexible control, and security monitoring of the interaction process.
[0204] Deep user profiling and intention understanding are mainly reflected in: using an internal / external factor decoupling framework to analyze data such as children's interaction logs, learning choices, and usage duration, and distinguishing whether their behavior is due to stable internal interests or situation-driven external factors (such as specific time, parental requirements).
[0205] Intelligent decision-making and flexible control (application of software-defined network / digital twin concepts) are mainly reflected in: Drawing on the idea of software-defined network / digital twin integrated detection, construct a software-defined control plane and a digital twin of children's learning status / preferences. The control logic (parent-side APP) realizes differentiated and dynamic permission management and content recommendation based on the motivation analysis output by factor decoupling and the digital twin simulation results. For example, give more lenient time to learning behaviors driven by internal interests and appropriately guide or restrict behaviors driven by external factors (such as just to complete tasks). Use digital twins for simulation prediction (such as predicting learning fatigue, interest transfer) and adjust strategies in advance.
[0206] The security monitoring of the interaction process is mainly reflected in: deploying a query update analysis mechanism to monitor the interaction patterns with cloud artificial intelligence services (such as recommendation systems, content generation) (analyzing the incremental similarity of query sequences), and detecting potential black-box adversarial attack attempts in real time, rather than relying solely on input filtering.
[0207] In some embodiments, the education system 100 further includes: an adaptive learning and full-process security protection module for personalized adaptive learning, multi-level proactive security, and continuous health monitoring.
[0208] Personalized adaptive learning is mainly reflected in: combining federated learning with the ability to decouple motivation to build a more accurate and interpretable personalized recommendation engine, and dynamically adjusting the learning path and content difficulty.
[0209] Multi-level proactive security is mainly reflected in: access security and interaction security.
[0210] Access security: If face recognition is used for user authentication or personalized switching, integrate dual alignment (domain and modality) face anti-counterfeiting technology to effectively resist spoofing attacks such as photos, videos, and 3D masks, and ensure robustness especially in multi-domain (different lighting, devices) and multi-modal (visible light + depth / infrared sensors if any) scenarios.
[0211] Interaction security: Detect black-box query attacks during the interaction process to protect the core artificial intelligence models (recommendation, content generation, motivation analysis, etc.) from being maliciously detected or manipulated.
[0212] Continuous health monitoring is mainly reflected in: retaining and possibly strengthening the original eye movement tracking and distance monitoring, and combining precise pose data to achieve more reliable eye health management and intervention.
[0213] The technical effect lies in that the core creativity lies in the feedback loop formed by the deep integration among perception, dynamic content generation, user understanding, and security protection. This enables the system to simultaneously achieve: extreme realism, deep personalization, proactive security protection, and powerful creative empowerment.
[0214] The following are some examples of key synergy effects: Surreal interaction scenarios → driving deeper cognitive understanding and personalization: Involved technologies: high-fidelity pose estimation + adaptive 3D rendering (3D Gaussian splatter / neural radiance field) + decoupling of internal and external motivations + digital twin of children.
[0215] Collaborative approach: The system doesn't just "see" the child (pose estimation), but precisely understands in real-time the six-degree-of-freedom (6DoF) interaction between the child and the physical card / device. This interaction triggers the generation of extremely realistic digital twin content (3D Gaussian splatting / Neural Radiance Field), and intelligently selects the optimal rendering technology based on device performance and interaction context. These rich interaction data (pose, distance, viewing angle, duration, etc.) provide much stronger signals for the motivation decoupling framework than simple clicks or usage duration, enabling more effective discrimination between whether the child's behavior stems from internal interest or external task-driven. This refined understanding updates the child's digital twin model in real-time.
[0216] Non-obvious effect: This collaboration elevates the dimension of user understanding from traditional 2D screen interaction to refined 3D interaction understanding in a high-fidelity dynamically generated environment based on the physical world. The "realism" here is not just visual candy; it provides higher-quality and more physically meaningful data for motivation analysis. This makes the digital twin a more accurate mapping of the child's state, enabling personalization that extends beyond content recommendation to interaction style, learning rhythm, and even the fidelity of visual presentation (e.g., when the system determines that the child has a strong internal exploration desire, it may switch to the more computationally intensive but better-performing Neural Radiance Field rendering), achieving a leap from "adaptive learning" to "context-aware, motivation-driven adaptive reality".
[0217] Dynamic content generation and editing → Achieving interactive and quantifiable embodied experiential learning: Technologies involved: Pose estimation + 3D Gaussian splatting / Neural Radiance Field generation + Temporally consistent video generation + Text-assisted audio-visual generation + Quantity-aware content editing.
[0218] Collaborative approach: The child shows an "apple" card (pose estimation triggers recognition and loading of the 3D Gaussian splatting / Neural Radiance Field model). The system can receive an instruction (possibly from speech recognition or generated by a large model based on learning goals), such as "Show what happens when adding 2 more apples". The quantity-aware editing module precisely adds the specified number of objects in the 3D scene while maintaining the identity characteristics of the original objects. The temporally consistent module ensures that the animation of the new apples appearing is smooth and natural without flickering. Text-assisted audio generation synchronously outputs the explanatory narration ("Look, 3 apples plus 2 apples, now there are 5 apples!").
[0219] Non - obvious effect: This combination enables instant generation, physics - based interaction, and quantifiable concept demonstrations. It is not a pre - made animation but an interactive simulation that is dynamically generated in direct response to the child's physical actions and current learning needs. Integrating quantity - aware editing capabilities directly into the high - fidelity rendering and animation process allows abstract concepts (such as mathematical operations) to be concretely, accurately, and coherently visualized in an immersive environment that the child has just "triggered" physically, which cannot be dynamically achieved by static pictures or simple pre - set animations. The accompanying audio generation further makes it a multi - sensory, explanatory learning loop.
[0220] Motivation - aware digital twin → Actively shaping the interaction experience and security strategy: Involved technologies: Motivation decoupling + Digital twin + Adaptive rendering / content selection + Dynamic content editing + Interactive query pattern monitoring.
[0221] Collaboration method: The digital twin constructed based on the motivation decoupling result not only passively reflects the child's state but can also make predictions and simulations. An intelligent control platform similar to software - defined networks can "query" the digital twin: "If this slightly more difficult concept is introduced now, based on the child's current intrinsic motivation level, what is the likelihood that he / she will lose interest?" According to the simulation results, the system may decide to postpone, simplify the concept, or, when the child's intrinsic interest is high, instead actively enhance the visual effect (such as switching to higher - quality rendering) or increase interactivity to further stimulate. At the same time, interactive query monitoring also utilizes the context information of the digital twin. For example, if the digital twin shows that the child is immersed in a deep exploration of a certain model, and a series of repetitive, patterned exploratory interactive queries occur at this time, the system is more likely to judge it as a potential (against the AI model) black - box attack attempt than in a random exploration state.
[0222] Non - obvious effect: The digital twin evolves into an active simulation engine for personalized and security decision - making. The control decisions of the system (content recommendation, difficulty adjustment, interaction mode, etc.) are forward - looking and context - aware, based on simulation and deduction on the basis of a deep understanding of the child. Security protection is no longer just passive filtering but becomes context - sensitive, using the same user model to evaluate the rationality of interaction behavior patterns. This constructs an intelligent system that can actively manage learning engagement and potential security risks based on a unified, dynamic user model.
[0223] End - to - end, context - aware security system → Ensuring the trustworthiness of the immersive generation experience: Involved technologies: Dual - aligned face anti - forgery + Interactive query pattern monitoring + Pose estimation + Full - link dynamic content generation (3D Gaussian sputtering / Neural Radiance Fields / video / audio / editing).
[0224] Collaborative approach: Robust identity authentication (face anti-counterfeiting) ensures the security of the system's entry point and incorporates multi-modal and multi-environment factors. More crucially, the interactive query monitoring protects the system's core generative AI model itself.
[0225] The AI model itself - including 3D renderers, video / audio synthesizers, content editors, etc. - is protected from being maliciously probed, manipulated, or reverse-engineered during its interaction with users. Physical perception data such as pose estimation provides important contextual evidence: Does the user's query instruction match their current physical pose, attention focus, and interaction object? Creative / non-obvious effects: Security is seamlessly woven into the entire closed-loop from access to interaction to content generation. It specifically focuses on and addresses the new attack surfaces brought about by the introduction of powerful generative AI. Traditional security may focus more on input content filtering or output result auditing, but this solution actively monitors the behavior patterns of interactions with the AI model itself and uses physical context (pose) and user status (from digital twins) to assist in differentiating between legitimate exploration and malicious probing. This builds the necessary trust foundation for the deployment and application of cutting-edge generative AI technologies in sensitive and highly interactive scenarios such as children's education.
[0226] In several implementation manners provided by this application, it should be understood that the disclosed methods and devices can be implemented in other ways. For example, the device implementation manners described above are merely illustrative. For example, the division of the modules or units is only a logical function division, and there can be other division manners in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed.
[0227] If the integrated units in the above-mentioned other implementation manners are implemented in the form of software function units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor to execute all or part of the steps of the methods described in various implementation manners of this application. The aforementioned storage medium includes: USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs, etc., which can store program codes.
[0228] The above are only the embodiments of the present application, and do not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present application, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present application.
Claims
1. A multi-dimensional perception and intention recognition education system, characterized in that: The educational system includes: A high-fidelity spatiotemporal interaction engine, which uses a dual-attention ray scoring network to perform posture estimation and card recognition on the collected image stream, obtains the physical interaction mode between the child and the device and the card, and the card recognition result, and generates a high-fidelity 3D model based on the card recognition result in combination with 3D Gaussian sputtering / neural radiation field rendering technology; A video sequence generation module, used to generate a video sequence according to the high-fidelity three-dimensional model using a spatiotemporal collaborative network when a video sequence needs to be generated; An audio-visual generation module, used to input the text description and the video sequence into an audio generation model, generate audio that is synchronized with the video sequence in time and matches the meaning, and obtain the target audio and video; The compensation module is used to edit the target audio and video accordingly during dynamic teaching or children's creation.
2. The educational system according to claim 1, characterized in that The high-fidelity spatiotemporal interaction engine is also used to select three-dimensional Gaussian sputtering or neural radiation field rendering technology to generate the high-fidelity three-dimensional model in combination with device performance level, user preference settings and content metadata.
3. The educational system according to claim 1, characterized in that During display of the high-fidelity three-dimensional model, the high-fidelity spatiotemporal interaction engine is also used to switch between the three-dimensional Gaussian sputtering and neural radiation field rendering techniques according to monitored performance indicators.
4. The educational system according to claim 1, characterized in that The video sequence generation module is also used to obtain a first feature map of the current frame, a second feature map of the previous frame, and a third feature map of the next frame; obtain a forward attention map according to the first feature map and the second feature map, and obtain a backward attention map according to the first feature map and the third feature map; perform collaborative attention calculation on the forward attention map and the backward attention map to obtain a collaborative attention map; obtain previous context information by weighted summing the collaborative attention map with the second feature map, and obtain subsequent context information by weighted summing the collaborative attention map with the third feature map; fuse the previous context information, the subsequent context information, and the first feature map to obtain the current frame feature; The video sequence is generated according to the current frame features.
5. The educational system according to claim 4, characterized in that The video sequence generation module comprises: A motion prediction unit, used for predicting motion information between frames, wherein the motion information is input into the spatiotemporal collaborative network as a condition; The spatiotemporal feature fusion and propagation unit is used to concatenate the feature maps of the previous and next frames in the channel dimension and send them to the subsequent processing layer, as well as to add the features of the previous and next frames at the element level or perform weighted summation according to the attention weight, and to use the gating mechanism to control the propagation and update of historical information, and to use convolution kernels that can process both spatial and temporal dimensions at the same time.
6. The educational system according to claim 1, characterized in that The audio-visual generation module is also used to extract and align features of the text description and the video sequence to obtain aligned visual features and text features; and input the visual features and the text features into the audio generation model to generate audio that is synchronized in time with the video sequence and matches the meaning, thereby obtaining the target audio and video.
7. The educational system according to claim 1, characterized in that The compensation module comprises: A feature compensation unit is used to extract features from the original image during dynamic teaching or children's creation to obtain initial features, and to interact the quantity and object information in the text prompt with the initial features to obtain a compensated feature vector, and to fuse the compensated feature vector with the initial features to obtain enhanced image features; The quantity-aware attention unit is used to extract information about each object from the enhanced image features, and to perform attention interaction between each object information and the current feature map to obtain quantity-aware features, and to inject the quantity-aware features into the corresponding feature map.
8. The educational system according to claim 1, characterized in that The educational system also includes: The intelligent management and control center adopts an internal / external factor decoupling framework to analyze children's interaction logs, learning choices, and usage time; as well as monitor the interaction patterns with cloud-based artificial intelligence services to detect potential black-box adversarial attack attempts in real time.
9. The educational system according to claim 8, characterized in that The intelligent management and control center is also used to use digital twins for simulation prediction and adjust strategies in advance.
10. The educational system according to claim 1, characterized in that The educational system also includes: Adaptive learning and full-process safety protection module for personalized adaptive learning, multi-level active safety and continuous health monitoring.
Citation Information
Patent Citations
Controllable video generation method and system based on multi-modal fusion
CN119091362A
Education experience system based on virtual reality technology
CN119694171A
Audio-driven three-dimensional digital human generation method and system based on neural radiation field
CN119888023A
Automated video generation
US12176007B1
Cited By
Digital exhibition hall system and visitor guiding method thereof
CN120950165A