Multimodal ai-based interactive animation generation method and system for children's drawings, electronic device and medium
By preprocessing and semantically understanding children's paintings using multimodal AI, a consistent sequence of animation frames is generated and interactive functions are configured. This solves the problem that existing technologies cannot analyze and transform children's abstract paintings, enabling the generation of interactive animations and enhancing the fun and engagement of educational scenarios.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- GUANGZHOU61LEARN INFORMATION TECH CO LTD
- Filing Date
- 2026-04-10
- Publication Date
- 2026-06-09
AI Technical Summary
Existing technologies cannot deeply analyze the inner meaning of children's abstract paintings, making it difficult to transform them into interactive animations. They lack the fun and interactivity required for educational scenarios, hindering the formation of a creative closed loop.
By preprocessing and semantically understanding children's drawings using multimodal AI, a consistent sequence of animation frames is generated, and interactive functions, including touch, voice, and gesture control, are configured to enable real-time interaction between children and the animation.
In-depth analysis of children's abstract paintings, efficiently transforming them into interactive animations, supporting children's natural interaction, and enhancing educational value and creative enthusiasm.
Smart Images

Figure CN122176126A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the fields of artificial intelligence and education technology, specifically to a method, system, electronic device, and medium for generating interactive animations of children's paintings based on multimodal AI. Background Technology
[0002] In early childhood education, painting, as a core activity for children to express their inner world, develop creativity, and cultivate concrete thinking, has been widely recognized for its value. Children convey emotions and imagination through their artwork, and creative extension activities based on these drawings are crucial for building a complete learning experience. However, current digital processing tools for children's artwork are extremely limited, mainly confined to image scanning, basic storage, and simple editing operations, such as photographing and archiving paper artwork, performing basic color adjustments, or outlining lines. While these tools can digitally preserve artwork, they cannot deeply analyze its inherent meaning, let alone transform static images into dynamic content. Especially for young children, whose artwork often exhibits highly abstract characteristics, such as simplified lines, unbalanced proportions, and free use of color, existing tools encounter bottlenecks at the semantic understanding level. They cannot accurately identify the core image in the artwork, nor can they extract key visual features, making it difficult for children to extend their creative intentions into an interactive, dynamic experience, thus weakening the sustained motivation and educational value of their creations.
[0003] In recent years, generative artificial intelligence and multimodal AI technologies have made significant progress in image recognition and content generation, theoretically supporting semantic parsing and dynamic content creation. However, the design logic of mainstream AI platforms on the market focuses on adult standardized drawings or professional design scenarios, and their training data and recognition mechanisms are difficult to adapt to the non-standardized expressions of children's drawings. For example, these platforms often misclassify children's simple line drawings as irrelevant objects, or ignore children's unique creative logic during feature extraction, resulting in semantic comprehension bias. At the same time, their output often lacks the fun and interactivity required for educational scenarios, failing to meet children's expectations for dynamic feedback. On the other hand, while traditional children's animation tools are simple to operate and have user-friendly interfaces, such as template-based collage software, they heavily rely on manual editing of animation frames. This is not only inefficient, leading to inconsistencies in the images, but also lacks intelligent semantic understanding capabilities, preventing children from achieving autonomous transformation from drawings to animations. This disconnect hinders the formation of a creative closed loop; children cannot directly witness the dynamic evolution of their drawings, nor can they deepen their learning experience through real-time interaction. Summary of the Invention
[0004] The purpose of this application is to provide a method, system, electronic device and medium for generating interactive animations of children's paintings based on multimodal AI, which can deeply analyze children's abstract paintings, efficiently transform them into interactive animations, support children's natural interaction and enhance educational value.
[0005] This application provides a method for generating interactive animations of children's drawings based on multimodal AI, including: The system collects images of children's paintings, preprocesses them, and outputs standardized image data that is compatible with the input of a third-party multimodal AI model. The preprocessing includes image normalization, size adjustment, noise removal, and data augmentation, while preserving the original line and color features of the children's paintings. Standardized image data is input into a third-party multimodal AI model to achieve semantic understanding of children's abstract paintings, extract the core images and key features in children's abstract paintings, and output image feature description text and image outline data. The image feature description text and image outline data are input into the image model, and combined with the types of movements that children like to generate continuous image frames. The image model uses fixed seeds, shared latent variables and Controlnet pose anchoring technology to ensure the consistency of images between frames. The system sorts and processes the transitions between frames of continuously moving images, synthesizes a basic animation, and then combines the basic animation with the background template library to generate a scene-based animation. Configure interactive functions that adapt to children's operating habits for scene-based animations, and output the configured interactive animations to the display terminal to realize real-time interaction between children and interactive animations.
[0006] Furthermore, the images of children's drawings are preprocessed, including: Brightness calibration and shadow removal were performed on the collected images of children's paintings to eliminate light interference caused by the collection environment; The pixel values of the children's paintings are scaled to a preset range to complete the normalization process, and the size of the children's paintings is adjusted to fit the fixed input size of the third-party multimodal AI model. Gaussian filtering was used to remove noise from children's artwork images, and random slight rotation, cropping, and color enhancement were performed on the images for data augmentation.
[0007] Furthermore, standardized image data is input into a third-party multimodal AI model to achieve semantic understanding of children's abstract paintings, extract core images and key features from the paintings, and output image feature description text and image outline data, including: The third-party multimodal AI model is a general multimodal large model that uses standardized image data to extract visual features and understand semantics from children's abstract paintings. Image features of standardized image data are extracted by the visual encoder of a third-party multimodal AI model, and the semantic understanding ability of the third-party multimodal AI model is combined to identify the core images and key attributes in children's abstract paintings. Based on image features and semantic understanding results, generate image outline mask data or edge coordinates corresponding to children's abstract paintings, and organize and output structured image feature description text and image outline data.
[0008] Furthermore, the text describing the image features and the image outline data are input into the image model, and combined with the types of movements that children like to generate continuous image frames of movements, including: The ComfyUI node-based workflow based on Z-Image-Turbo is used as the inference architecture for the raw image model. The API is called in batches through Python scripts, and the action description prompts are adjusted frame by frame according to the types of actions that children like to perform, so as to achieve gradual changes in actions. A fixed starting seed is set for the raw image model and incremented frame by frame. The determinism of random noise is controlled by sharing latent variables, and the action pose is anchored by Controlnet's OpenPose or Depth module. Control the sampling steps of the raw image model, keep the frame rate of the image frames at 24-30 frames per second, and set the number of image frames with continuous action according to the animation duration requirements to ensure that the outline and color of the image between frames are highly consistent with the original artwork of the children.
[0009] Furthermore, the image frames with continuous motion are sorted into a frame sequence and inter-frame transitions are processed to synthesize a basic animation. This basic animation is then combined with a background template library to blend with the background, generating a scene-based animation, including: The frame sequence of continuously moving image frames is sorted, and inter-frame transition effects such as fade-in / fade-out and dissolve are added to synthesize the basic animation. The background template library includes background templates such as grassland, forest, and playground that are adapted to children's aesthetics. It supports manual selection or automatic AI matching of backgrounds based on the characteristics of the core image and the colors of children's paintings. Embed the basic animation into the selected background, adjust the size and position of the basic animation to achieve a natural blend, and add scene dynamic effects such as fluttering leaves and shimmering light spots to the background to generate a scene-based animation.
[0010] Furthermore, interactive functions adapted to young children's operating habits are configured for the scene-based animations, including: The configured interactive functions include one or more of touch control, voice command control, and gesture control; Touch control triggers preset action switching for animated characters on the touch display terminal, with a response time of ≤0.5 seconds; Voice command control involves collecting children's voice commands and converting them into control signals using a recognition API that adapts to the characteristics of children's voices, thereby driving the animated characters to perform corresponding actions. Gesture control uses a camera or depth sensor to recognize a child's waving hand commands, converts these commands into control signals, and controls the animated character to complete movement actions.
[0011] This application also provides a multimodal AI-based interactive animation generation system for children's artwork, including: The artwork acquisition and preprocessing module is used to acquire images of children's artwork, normalize, resize, remove noise, and enhance the data of the images, output standardized image data, and preserve the original creative characteristics of the children's artwork. The multimodal AI understanding module is used to input standardized image data into a third-party multimodal AI model, receive the visual feature extraction results and semantic understanding results returned by the third-party multimodal AI model, and organize and output image feature description text and image outline data. The motion frame generation module has a built-in raw image model and ComfyUI node-based workflow. It is used to receive image feature description text and image outline data, and generate a sequence of image frames with continuous motion and consistent image. The animation compositing module is used to process image frame sequences to synthesize basic animations. It combines the background template library to achieve the fusion of basic animations and backgrounds, and outputs scene-based animations. The interactive control module is used to configure one or more interactive functions among touch, voice, and gesture for scene-based animations, converting children's interactive instructions into control signals to drive the animated characters to make corresponding responses. The display terminal module is used to display the configured interactive animation, receive children's interactive operations and provide feedback on the animation response results, and is compatible with kindergarten all-in-one machines and tablet terminal devices.
[0012] Furthermore, the artwork acquisition and preprocessing module includes an image acquisition unit and a preprocessing unit: The image acquisition unit supports both paper artwork scanning and electronic artwork import. It uses a high-definition scanner and a high-definition camera as image acquisition devices to ensure that the acquired images of children's artwork are clear and without distortion. The preprocessing unit integrates an image processing algorithm library to achieve integrated processing of brightness calibration, shadow removal, normalization, size adjustment, noise removal, and data augmentation for children's paintings.
[0013] Furthermore, the motion frame generation module is also used for: The image frame sequence is kept consistent across frames by using a fixed seed, shared latent variables, and ControlNet pose anchoring techniques. Control the sampling steps of the raw image model to keep the frame rate of the image frame sequence between 24-30 frames / second, and the frame generation efficiency is not less than 1.5 seconds / frame; The generated image frame sequence has a resolution adapted to kindergarten terminal devices, providing high-quality materials for the animation synthesis module.
[0014] Furthermore, the animation compositing module is also used for: The background template library contains at least 20 background templates with themes such as grassland, forest, playground, and ocean, which are suitable for children's aesthetics. It supports manual selection or AI automatic background matching. The image frame sequence is sorted, inter-frame transition effects are added, the basic animation is blended with the selected background, scene dynamic effects are added to the blended animation, and scene-based animation is output.
[0015] Furthermore, the interactive control module includes a touch recognition submodule, a voice recognition submodule, and a gesture recognition submodule: The touch recognition submodule recognizes touch control signals based on touch events from the display terminal; The speech recognition submodule uses cloud APIs or local lightweight models adapted to the characteristics of young children's speech, and supports keyword wake-up and voice command parsing; The gesture recognition submodule captures images of children's gestures through a camera and calls the MediaPipe lightweight gesture recognition algorithm to recognize gesture control signals; The interactive control module responds to children's interactive commands in ≤0.5 seconds to ensure real-time interaction.
[0016] This application also provides an electronic device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to run the steps of the above-described method for generating interactive animations of children's paintings based on multimodal AI. The electronic device also includes a communication interface and an object storage client. The communication interface supports RESTful API data interaction and is used to establish connections with third-party multimodal AI model services and speech recognition services. The object storage client is used to connect to image and audio file storage services to realize the storage and access of animation frames, background materials and final animation files.
[0017] This application also provides a medium, which is a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it performs the steps of the above-described method for generating interactive animations of children's paintings based on multimodal AI. The computer-readable storage medium is a non-transitory storage medium that supports encrypted data storage and access control to ensure the security of children's artwork data and generated animation content.
[0018] Compared with the prior art, this application has the following beneficial effects: By preprocessing to preserve the original features of children's paintings, using multimodal AI to understand the semantics of abstract paintings, generating a consistent sequence of animation frames, compositing animations, and adding interactive functions, this technology solves the problem that existing technologies cannot analyze children's abstract paintings and transform them into interactive animations. It can deeply analyze children's abstract paintings, efficiently transform them into interactive animations, support children's natural interaction, and enhance educational value. Attached Figure Description
[0019] Figure 1 This is a schematic diagram illustrating the method for generating interactive animations of children's drawings, using the vidu2 image as an example, as an embodiment of this application.
[0020] Figure 2 This application provides a schematic diagram illustrating a method for generating interactive animations of children's artwork, using Vidu2 (vidu2-image) as an example. Figure 3 This application provides a schematic diagram illustrating a method for generating interactive animations of children's artwork, using Vidu2 (vidu2-image) as an example. Figure 4 This is a schematic diagram of a method for generating interactive animations of children's paintings based on multimodal AI, provided in an embodiment of this application.
[0021] Figure 5 This is a schematic diagram of the structure of a multimodal AI-based interactive animation generation system for children's paintings, provided in an embodiment of this application. Detailed Implementation
[0022] The technical solutions of this application will now be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are merely some embodiments of this application, and not all embodiments. The components of this application described and shown in the accompanying drawings can generally be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of this application provided in the accompanying drawings is not intended to limit the scope of the claimed application, but merely to illustrate selected embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without inventive effort are within the scope of protection of this application.
[0023] It should be noted that similar reference numerals and letters in the following figures indicate similar items; therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures. Furthermore, in the description of this application, terms such as "first," "second," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.
[0024] Existing tools for digitizing children's artwork primarily focus on image scanning, storage, and basic editing, lacking capabilities in semantic understanding, creative animation transformation, and extended interactive experiences. For abstract drawings by young children, existing tools struggle to achieve accurate interpretation and dynamic transformation. Furthermore, current AI image and animation generation platforms are typically designed for adult-oriented artwork, exhibiting poor accuracy in semantic understanding and feature extraction of children's drawings, and failing to meet the needs of early childhood education scenarios. Traditional animation production tools suffer from low creation efficiency, poor image consistency, and the inability to achieve real-time interaction.
[0025] It should be noted that the third-party multimodal AI models mentioned in the context include, but are not limited to, OpenAI GPT-4o / GPT-4V, Google Gemini 3.1 / 2.5Pro, Anthropic Claude 3 Opus / Sonnet, Meta LLaMA4 / Llama 3.2 Vision (open source), Alibaba's Tongyi Qianwen Qwen 3.5-Omni / Qwen-VL, Zhipu GLM-4V / GLM-4, Baidu Wenxin Yiyan 6.0 (ERNIE 4.0), ByteDance Doubao / BAGEL (open source), Tencent Hunyuan, Shengshu Vidu2 (vidu2-image / vidu2-video), Midjourney / Stable Diffusion 3 / FLUX.1 / 2, and Pika Labs / Runway Gen-3.
[0026] Taking Vidu 2 (vidu2-image) as an example, the following method for generating interactive animations of children's drawings is implemented, such as... Figures 1-3 As shown.
[0027] like Figure 4 As shown, this application proposes a method for generating interactive animations of children's drawings based on multimodal AI, the method comprising: Collect images of children's paintings, preprocess the images, and output standardized image data that is compatible with the input of a third-party multimodal AI model. The preprocessing includes image normalization, size adjustment, noise removal, and data augmentation, while preserving the original line and color features of the children's paintings. The standardized image data is input into a third-party multimodal AI model to achieve semantic understanding of children's abstract paintings, extract the core images and key features in the children's abstract paintings, and output image feature description text and image outline data. The image feature description text and the image outline data are input into the image model, and combined with the types of movements that children like to generate continuous image frames. The image model uses fixed seeds, shared latent variables and Controlnet pose anchoring technology to ensure the consistency of images between frames. The image frames of the continuous action are sorted into a frame sequence and inter-frame transitions are processed to synthesize a basic animation. The basic animation is then combined with a background template library to achieve the fusion of the basic animation with the background, generating a scene-based animation. Configure interactive functions adapted to children's operating habits for this scenario-based animation, and output the configured interactive animation to the display terminal to realize real-time interaction between children and the interactive animation.
[0028] First, the images of children's artwork are acquired and preprocessed. Image acquisition can be achieved in various ways; for example, paper artwork can be photographed with a regular digital camera, or digital artwork files can be directly imported. The acquired raw images may have uneven lighting, inconsistent sizes, noise, etc. Therefore, these images need to be preprocessed. Preprocessing may include simple scaling to adjust the size of the image, or manually adjusting brightness and contrast to improve image quality. In addition, basic image filtering algorithms can be used to remove some noise, and simple image flipping or cropping can be performed to increase the data volume. After the above processing, the image data is converted into a standardized format to meet the input requirements of subsequent multimodal AI models, while preserving the original line and color characteristics of the children's artwork.
[0029] In some implementations, preprocessed, standardized image data is input into a multimodal AI model. This model is responsible for semantic understanding of the children's abstract drawings. For example, a general image recognition model can be used, capable of identifying common objects or shapes in the image and outputting corresponding labels. By manually filtering and organizing these labels, the core image in the drawing can be initially identified. Simultaneously, image segmentation tools can be used to manually outline the contour of the core image, generating simple contour data. Finally, the identified image and contour information are organized into text descriptions and contour data.
[0030] Based on this, the obtained image feature description text and image outline data are input into the raw image model to generate image frames with continuous motion. The raw image model can be a basic diffusion model that generates images based on text prompts. To achieve motion, the motion description prompts for each frame can be manually written, and images are generated frame by frame. This method may result in differences in appearance, color, or style between images generated in different frames. To generate continuous animation, these independent image frames need to be arranged in a preset order.
[0031] The generated, sequentially moving image frames are then processed to synthesize a base animation. This can include simply stitching all the image frames together to form a video sequence. To enhance the animation's visual appeal, simple transition effects can be manually added between frames, such as direct cuts or rapid fade-ins and fade-outs. Simultaneously, to enrich the animation scene, a background library containing a small number of generic backgrounds (such as solid color backgrounds or simple landscape images) can be used to select a background. The synthesized base animation is then overlaid onto the selected background, and the size and position of the animated figures are manually adjusted to blend them with the background, thus generating a scene-based animation.
[0032] Finally, the generated scene-based animation is configured with interactive features and output to a display terminal. The interactive features can be configured by setting click events for specific areas of the animation; when a user clicks on that area, the animation plays the next pre-set segment. The display terminal can be any device with screen output capabilities, such as a regular computer monitor or television. In this way, young children can engage in simple interactions with the animation.
[0033] By introducing multimodal AI to semantically understand children's abstract paintings, the limitations of traditional tools in interpreting artwork are overcome. Combining raw image models to generate continuous and consistent animation frames solves the problems of poor generation quality from existing AI platforms and low efficiency from traditional tools. Through animation synthesis and interactive configuration, children's paintings are transformed into scene-based interactive animations, enabling real-time interaction between children and self-made animations, thereby stimulating children's creative enthusiasm and enhancing the fun and engagement of educational scenarios.
[0034] If the complexity of the image acquisition environment and the stringent requirements of subsequent multimodal AI models on the quality of input data are not fully considered during the preprocessing of collected children's artwork images, problems such as uneven lighting, inconsistent sizes, and noise interference may occur. These problems will affect the AI model's accurate understanding of the children's artwork and may even alter the original artistic features of the artwork, thereby reducing the effectiveness and consistency of subsequent animation generation.
[0035] In some implementations, specific steps are proposed for preprocessing the children's artwork images to ensure the quality and consistency of the input data. Specifically, firstly, brightness calibration and shadow removal are performed on the acquired children's artwork images. When acquiring children's artwork images, limitations such as shooting equipment and lighting conditions may lead to uneven brightness and local shadows. Brightness calibration aims to adjust the overall brightness distribution of the image to achieve a uniform and moderate level. Shadow removal identifies and eliminates shadow areas in the image through image processing algorithms, such as methods based on local contrast enhancement, multi-scale Retinex algorithms, or deep learning shadow removal models. These processes effectively eliminate light interference from the acquisition environment, ensuring the visual quality and color reproduction of the artwork images, providing clear and unbiased input for subsequent AI understanding.
[0036] Secondly, the pixel values of the children's artwork images are scaled to a preset range for normalization, and the size of the images is adjusted to fit the fixed input size of the third-party multimodal AI model. Normalization scales the pixel values to a preset range, such as [0,1] or [-1,1]. This helps stabilize the neural network training process and prevents gradient explosion or vanishing caused by excessively large or small pixel values. Size adjustment involves uniformly scaling or cropping children's artwork images of different sizes to the fixed input size required by the third-party multimodal AI model, such as 224x224 or 512x512 pixels. This ensures that all input data has consistent dimensions, meets the input specifications of the AI model, and avoids model errors or performance degradation due to size mismatch.
[0037] Furthermore, Gaussian filtering is used to remove noise from the children's artwork images, followed by random slight rotation, cropping, and color enhancement data augmentation operations. Gaussian filtering is a commonly used linear smoothing filter that blurs the image by weighted averaging of image pixels, effectively removing high-frequency noise such as speckle noise. Its principle is to use a Gaussian function as weights, with pixels closer to the center pixel having a higher weight. Data augmentation operations expand the training dataset by introducing random variations, improving the model's generalization ability. Random slight rotation simulates the artwork's appearance from different angles, cropping simulates attention to local details, and color enhancement (such as adjusting brightness, contrast, and saturation) simulates different lighting or color conditions. These operations increase data diversity without altering the core content of the artwork, making the model more robust to various input variations while avoiding over-enhancement that could distort the original line and color features.
[0038] In some possible embodiments, it is proposed to input the standardized image data into a third-party multimodal AI model to achieve semantic understanding of children's abstract paintings, extract the core images and key features in the children's abstract paintings, and output image feature description text and image outline data. Specifically, the third-party multimodal AI model is a general multimodal large model that only uses the standardized image data to achieve visual feature extraction and semantic understanding of children's abstract paintings without additional training or modification; the visual encoder of the third-party multimodal AI model extracts image features from the standardized image data, and combines the semantic understanding capability of the third-party multimodal AI model to identify the core images and key attributes in the children's abstract paintings; based on the image features and semantic understanding results, image outline mask data or edge coordinates corresponding to the children's abstract paintings are generated, and the structured image feature description text and image outline data are compiled and output.
[0039] Specifically, the third-party multimodal AI model is configured as a general-purpose multimodal large model, characterized by the fact that it does not require additional model training or fine-tuning for the specific style of children's paintings. These models are typically pre-trained on large-scale, multi-domain data, possessing strong cross-modal understanding and generalization capabilities, and can directly process different types of image and text inputs. By leveraging its inherent visual and language understanding capabilities, the model can directly analyze children's paintings, significantly reducing the complexity and cost of development and deployment, and quickly adapting to different styles of children's paintings, avoiding the tedious process of customized training for abstract paintings.
[0040] In some implementations, the standardized image data is converted into high-dimensional image feature vectors using the visual encoder of the third-party multimodal AI model. The visual encoder is the part of the multimodal AI model specifically responsible for processing image input; it can extract rich visual information from the original image, including texture, shape, color, and spatial layout. Subsequently, combined with the semantic understanding capabilities of the third-party multimodal AI model, these image features are analyzed in depth to identify the core images (e.g., small animals, people, objects, etc.) and their key attributes (e.g., color, posture, emotion, size, etc.) in the children's abstract paintings. This process represents the model's transformation from pixel-level visual information to conceptual-level semantic understanding, a crucial step in accurately capturing the deeper meaning of children's paintings.
[0041] Based on the aforementioned image features and semantic understanding results, the system can generate image outline mask data or edge coordinates corresponding to the children's abstract paintings. Image outline mask data is a pixel-level binary image used to accurately mark the area of the core image in the painting, achieving effective separation of foreground and background. Edge coordinates are a set of object boundary points extracted through image processing algorithms, accurately describing the geometric shape of the image. Simultaneously, the system organizes and outputs the semantic understanding results as structured image feature description text. This description text typically uses a standardized format (such as JSON or XML) and includes detailed attribute information such as the core image's name, color, action, and emotion. The image outline data refers to the generated mask data or edge coordinates. This structured and precise data output provides high-quality, controllable input for subsequent image generation models, ensuring the accuracy and consistency of the animation generation process.
[0042] In some possible embodiments, a specific implementation method is proposed, which involves inputting image feature description text and image outline data into the raw image model, and combining this with the types of actions that children prefer to generate continuous image frames of actions. Specifically, this application adopts the ComfyUI node-based workflow based on Z-Image-Turbo as the inference architecture of the raw image model. Z-Image-Turbo is a high-performance image generation model that can provide fast and high-quality image generation capabilities. The ComfyUI node-based workflow provides a visual and modular model inference and customization environment, allowing the complex generation process to be decomposed into a series of configurable nodes, facilitating the flexible combination of different generation steps and control parameters. By calling the API in batches through Python scripts, the animation generation process can be automated and programmatically controlled, thereby finely adjusting the action description prompts frame by frame according to the types of actions that children prefer to take, in order to achieve gradual changes in the animated actions. For example, by gradually modifying the action description in the prompts of consecutive frames, such as from "standing" to "slightly raising the hand" and then to "fully waving the hand," the model can be guided to generate a smoothly transitioning sequence of actions.
[0043] To further ensure consistency in image quality across frames, this application sets a fixed starting seed for the raw image model and increments it frame by frame. In the diffusion model, the seed determines the initial random noise distribution. A fixed starting seed ensures controllable randomness at the beginning of the generated sequence, while frame-by-frame incrementing introduces subtle changes that contribute to the continuity of movement while maintaining overall consistency. Simultaneously, by sharing latent variables to control the determinism of random noise, it means that most of the random noise in the latent space is shared when generating a series of image frames. This greatly stabilizes the outline and color of the main image, effectively avoiding image jumps or flickering caused by each frame starting with completely independent random noise. Furthermore, by combining Controlnet's OpenPose or Depth modules, the animation character's pose can be precisely anchored. The OpenPose module extracts poses from skeletal keypoint information, while the Depth module uses depth information to control the 3D structure. By using preset poses (such as skeletal keypoint sequences or depth map sequences) as conditional inputs to Controlnet, the raw image model is forced to generate images according to these poses, thus ensuring the accuracy, coherence, and natural smoothness of the animation.
[0044] To optimize generation efficiency and animation quality, this application also controls the number of sampling steps for the raw image model. An appropriate number of sampling steps can improve generation speed while ensuring image generation quality. Simultaneously, controlling the frame rate of the image frames to 24-30 frames per second is a commonly used frame rate range in film and animation production, ensuring the visual smoothness and continuity of the final animation and conforming to human visual perception. Depending on the animation duration requirements, the number of consecutive image frames can be set to ensure the integrity of the animation content. Through the synergistic effect of these technical measures, it is ultimately possible to ensure that the images in the generated animation frames maintain a high degree of consistency with the original children's artwork in terms of outline and color, making the animated characters faithful to the original creative intent.
[0045] In some possible embodiments, a method is proposed to perform frame sequence sorting and inter-frame transition processing on image frames with continuous motion, synthesize a basic animation, and combine it with a background template library to achieve the fusion of the basic animation and the background to generate a scene-based animation. The steps include: sorting the image frames with continuous motion and adding fade-in / fade-out and dissolve inter-frame transition effects to synthesize the basic animation; the background template library contains background templates such as grassland, forest, and playground that are adapted to children's aesthetics, supporting manual selection or automatic background matching by AI based on the characteristics of the core image and the colors of the children's paintings; embedding the basic animation into the selected background, adjusting the size and position of the basic animation to achieve natural fusion, and adding scene dynamic effects such as fluttering leaves and flickering light spots to the background to generate the scene-based animation.
[0046] Specifically, sequentially ordering the frames of continuously moving images ensures the logical order and temporal continuity of the animation playback. This is typically managed using timestamps or frame indexes to ensure the images are presented in the correct order. Adding fade-in / fade-out and dissolve transitions between frames smooths the transitions between different images, avoiding abrupt visuals and improving the overall smoothness and visual comfort of the animation. Fade-in / fade-out effects can be achieved by gradually adjusting the image's transparency or brightness, while dissolve effects usually involve pixel-level blending of two frames, causing one frame to gradually disappear while another gradually appears. The introduction of these transition effects makes the visual experience of the animation more natural and coherent. Through the above processing, a series of continuously moving and smoothly transitioned image frames are combined to form a complete basic animation.
[0047] The background template library is a pre-set collection of digital images containing various background templates tailored to young children's aesthetic preferences, such as grasslands, forests, and playgrounds. These templates are carefully designed in terms of color, composition, and elements to align with children's cognitive characteristics and visual preferences. Manual selection is supported, meaning users can choose a suitable background from the library based on their own preferences or specific needs. AI-powered automatic background matching based on the characteristics of the core image and the colors of the child's artwork means the system can intelligently analyze the core image extracted from the child's drawing (e.g., if the core image is a fish, an ocean background might be matched) and its overall color style (e.g., if the drawing is predominantly green, a forest or grassland background might be matched), and recommend or automatically select the most suitable background template from the library accordingly. This is typically achieved through image feature extraction, semantic analysis, and pre-set matching rules or machine learning models.
[0048] Embedding the base animation into a selected background involves overlaying the foreground (base animation) onto the background image. Adjusting the size and position of the base animation ensures visual harmony and proportion between the foreground and background, making it appear naturally within the scene and avoiding unnatural situations such as the foreground being too large or too small, or improperly positioned. This typically involves scaling and panning the animation. Adding dynamic scene effects such as fluttering leaves and shimmering light spots to the background further enhances the scene's liveliness and immersion. Fluttering leaves can be simulated using particle systems or achieved through pre-rendered animation overlays; shimmering light spots can be achieved by adjusting lighting effects or overlaying dynamic textures. These dynamic effects enliven the static background, providing richer visual stimulation for young children. Through this series of processes, the base animation and dynamic background are organically combined to ultimately generate a scene-based animation with a complete scene and rich dynamic details.
[0049] By meticulously sequencing the continuous motion frames and introducing various inter-frame transition effects such as fade-in / fade-out and dissolve, the system effectively solves the problems of stuttering or abrupt transitions that may occur during animation playback, significantly improving the smoothness and visual coherence of the basic animation. Simultaneously, a background template library adapted to children's aesthetics is constructed, and an AI-powered automatic background matching function is provided. This transforms background selection from a purely manual operation into an intelligent recommendation of the most suitable scene based on the core image characteristics and color style of the child's artwork, greatly enriching the animation's visual expressiveness and reducing the user's selection burden. Furthermore, the precise adjustment of the size and position of the basic animation and the selected background ensures a natural blend between the foreground and background, avoiding visual disharmony. Adding dynamic scene effects such as fluttering leaves and flickering light spots to the background further enhances the animation's vividness and immersion, making the generated scene-based animation more vibrant and better able to attract children's attention, stimulating their interest in exploration and interaction.
[0050] Some implementations propose a method for generating interactive animations from children's artwork based on multimodal AI, capable of transforming children's drawings into vivid, scene-based animations. However, if the generated scene-based animations are presented only in a passive viewing format, it may be difficult to fully stimulate children's interest in participation and their desire for active exploration, thus limiting the interactivity and educational value of the animations. To enable children to interact with the animation content more naturally and intuitively, and to enhance their immersion and learning experience, it is necessary to configure the animations with interactive methods that are more in line with children's cognitive and operational habits.
[0051] In some implementations, interactive functions adapted to young children's operating habits are proposed for configuring the scene-based animation. Specifically, the configured interactive functions include one or more of touch control, voice command control, and gesture control. The touch control involves touching the animated character on the display terminal to trigger a preset action switch, with a response time of ≤0.5 seconds. The voice command control involves collecting the child's voice commands and converting the voice commands into control signals through a recognition API adapted to the characteristics of the child's voice, driving the animated character to perform corresponding actions. The gesture control involves recognizing the child's waving commands through a camera or depth sensor, converting the waving commands into control signals, and controlling the animated character to complete movement actions.
[0052] Interactive features refer to the ability to allow users to interact with digital content through specific actions and receive feedback. For young children, these features need to be designed to be intuitive, easy to understand, and easy to operate. One or more interactive methods can be selectively integrated depending on the actual application scenario and the hardware configuration of the display terminal. For example, touch control is preferred on tablets or touchscreen all-in-ones; voice command control is more convenient on smart speakers or devices with microphones; and gesture control provides a more natural contactless interaction for devices with cameras or depth sensors.
[0053] Touch control provides a direct and intuitive way to interact, allowing young children to change the state of animated characters or trigger specific behaviors by touching them on the display screen. When a user touches the screen, the system captures the coordinates of the touch event. By determining whether these coordinates fall within the rendering area of the animated character, the system can identify the touched character. Each animated character can have one or more preset action sequences, such as jumping, rotating, or making a sound. When a touch event is detected, the system switches the animated character's current action according to preset logic. A response time of ≤0.5 seconds is required to ensure the immediacy of the interaction and prevent children from losing interest due to delays. This is typically achieved by optimizing event handling mechanisms, reducing animation loading time, and using an efficient rendering pipeline.
[0054] Voice command control allows young children to interact with animations through verbal commands, a natural and convenient method for those unfamiliar with written language or complex operations. The system collects children's speech via the built-in microphone of the display terminal or an external microphone. The collected speech data is sent to a speech recognition API for processing. Considering the characteristics of young children's pronunciation (such as limited vocabulary, non-standard pronunciation, and slow speech speed), this API requires specialized training and optimization to improve the accuracy of speech recognition. The recognition API converts speech into text, and the system then parses the text content, mapping it to preset control signals, such as "jump," "run," and "laugh." These control signals are then used to drive the animated characters to perform corresponding preset actions.
[0055] Gesture control provides a non-contact interaction method, allowing young children to control animations through body movements (such as waving), increasing the fun and physical engagement of the interaction. The system uses the display terminal's built-in camera or an external depth sensor to capture images or depth information of the child's hands. This image or depth data is input into a gesture recognition algorithm. This algorithm is trained to recognize specific gestures, such as "waving." When a wave command is recognized, the system converts it into preset control signals, such as "move left" or "move right." These control signals are then used to adjust the position of the animated character in the scene, realizing the movement action. To improve the accuracy and robustness of recognition, gesture recognition algorithms are usually combined with machine learning models and trained under different lighting conditions and backgrounds.
[0056] Existing tools for digitizing children's artwork primarily focus on image scanning, storage, and basic editing, lacking capabilities in semantic understanding, creative animation transformation, and extended interactive experiences. For abstract drawings by young children, existing tools struggle to achieve accurate interpretation and dynamic transformation. Furthermore, current AI image and animation generation platforms are typically designed for adult-oriented artwork, exhibiting poor accuracy in semantic understanding and feature extraction of children's drawings, and failing to meet the needs of early childhood education scenarios. Traditional animation production tools suffer from low creation efficiency, poor image consistency, and the inability to achieve real-time interaction.
[0057] The innovation of this embodiment lies in combining multimodal AI models with motion frame generation in a standardized data flow manner, while introducing fixed seeds, shared latent variables, and Controlnet pose anchoring technology. This achieves high-precision semantic understanding and image feature extraction of children's abstract paintings, and ensures high consistency of images between animation frames, thus achieving the effect of efficiently transforming children's abstract paintings into interactive animations.
[0058] like Figure 5 As shown, this embodiment also discloses a multimodal AI-based interactive animation generation system for children's artwork, including: The artwork acquisition and preprocessing module is used to acquire images of children's artwork, normalize, resize, remove noise, and enhance the data of the images, output standardized image data, and preserve the original creative features of the children's artwork. The multimodal AI understanding module is used to input the standardized image data into a third-party multimodal AI model, receive the visual feature extraction results and semantic understanding results returned by the third-party multimodal AI model, and organize and output image feature description text and image outline data. The action frame generation module has a built-in raw image model and ComfyUI node-based workflow. It is used to receive the image feature description text and the image outline data, and generate a sequence of image frames with continuous action and consistent image. The animation compositing module is used to process the image frame sequence to synthesize a basic animation, and combines the background template library to achieve the fusion of the basic animation with the background, outputting a scene-based animation; The interactive control module is used to configure one or more interactive functions among touch, voice, and gesture for the scene-based animation, converting the child's interactive instructions into control signals to drive the animated character to make corresponding responses. The display terminal module is used to display the configured interactive animation, receive children's interactive operations and provide feedback on the animation response results, and is compatible with kindergarten all-in-one machines and tablet terminal devices.
[0059] In some possible embodiments, a specific implementation of the artwork acquisition and preprocessing module is proposed. However, during its implementation, it is necessary to ensure that the image preprocessing both eliminates interference from the acquisition environment and fully preserves the original creative characteristics of the children. Specifically, the artwork acquisition and preprocessing module includes an image acquisition unit and a preprocessing unit. The image acquisition unit supports both scanning of paper artworks and importing electronic artworks, and uses a high-definition scanner and a high-definition camera as image acquisition devices to ensure that the acquired images of children's artworks are clear and distortion-free. The preprocessing unit integrates an image processing algorithm library to achieve integrated processing of brightness calibration, shadow removal, normalization, size adjustment, noise removal, and data enhancement for children's artwork images.
[0060] In some implementations, the motion frame generation module ensures the consistency of image frame sequences across frames by using fixed seeds, shared latent variables, and ControlNet pose anchoring technology; it controls the sampling steps of the raw image model to maintain the frame rate of the image frame sequence at 24-30 frames per second, with a frame generation efficiency of no less than 1.5 seconds per frame; the resolution of the generated image frame sequence is adapted to kindergarten terminal devices, providing high-quality materials for the animation compositing module. The animation compositing module's background template library contains at least 20 background templates suitable for children's aesthetics, including grasslands, forests, playgrounds, and oceans, supporting manual selection or AI-automatic background matching; it sorts the image frame sequences, adds inter-frame transition effects, merges the basic animation with the selected background, adds scene dynamic effects to the merged animation, and outputs scene-based animation.
[0061] The interactive control module includes a touch recognition submodule, a voice recognition submodule, and a gesture recognition submodule: the touch recognition submodule recognizes touch control signals based on touch events from the display terminal; the voice recognition submodule uses a cloud API or a local lightweight model adapted to the characteristics of young children's voices, supporting keyword wake-up and voice command parsing; the gesture recognition submodule captures images of children's gestures through a camera and calls the MediaPipe lightweight gesture recognition algorithm to recognize gesture control signals; the interactive control module's response time to children's interactive commands is ≤0.5 seconds, ensuring real-time interaction.
[0062] In some possible embodiments, the above-mentioned painting acquisition and preprocessing module is proposed to include an image acquisition unit and a preprocessing unit.
[0063] The image acquisition unit is responsible for acquiring the original image data of children's artwork. This unit supports both scanning of paper artwork and importing digital artwork, designed to accommodate artwork from different sources. For paper artwork, a high-definition scanner can be used for digitization, ensuring that lines, colors, and details are accurately captured, avoiding distortion caused by low resolution or color deviation. For digital artwork, direct import of digital files is supported, such as artwork created using a tablet or drawing software. In this case, a high-definition camera can be used as an acquisition device to photograph or capture the digital artwork on the screen, similarly ensuring image clarity and integrity. By using a high-definition scanner or high-definition camera as the image acquisition device, it is ensured that the acquired images of children's artwork are clear and distortion-free, providing high-quality raw data for subsequent processing.
[0064] The preprocessing unit integrates an image processing algorithm library to perform a series of optimization processes on the acquired images of children's artwork. This library includes various image processing functions, enabling integrated processing of brightness calibration, shadow removal, normalization, resizing, noise removal, and data augmentation. Brightness calibration and shadow removal aim to eliminate uneven ambient lighting or shadow interference that may occur during acquisition, restoring the true colors and details of the artwork. Normalization scales the pixel values of the image to a preset range to eliminate brightness differences between different images, making them suitable for model input requirements. Resizing unifies the image to a fixed input size suitable for multimodal AI models, ensuring stable model processing. Noise removal effectively removes noise from the image using techniques such as Gaussian filtering, improving image clarity. Data augmentation operations, such as random slight rotation, cropping, and color enhancement, increase data diversity without altering the core features of the artwork, improving the generalization ability and robustness of subsequent AI models. These processing steps are integrated into an algorithm library, achieving an automated and integrated processing flow, ensuring the efficiency and standardization of preprocessing.
[0065] In some possible embodiments, the motion frame generation module is further configured to: ensure the inter-frame image consistency of the image frame sequence by using a fixed seed, shared latent variables, and Controlnet pose anchoring technology; control the sampling steps of the raw image model to keep the frame rate of the image frame sequence between 24-30 frames / second, with a frame generation efficiency of not less than 1.5 seconds / frame; and ensure that the resolution of the generated image frame sequence is adapted to the kindergarten terminal device, providing high-quality materials for the animation synthesis module.
[0066] Specifically, to ensure the consistency of image appearance across frame sequences, this application employs fixed seeds, shared latent variables, and ControlNet pose anchoring techniques. Fixed seeds, in generative models (such as diffusion models), ensure a consistent starting point for each generation process by setting an unchanging initial random value, which helps maintain the core visual features of the generated images. Shared latent variables, within the model's latent space, allow different frames to share the same set of abstract, compressed feature representations, thus maintaining a high degree of consistency in the image's inherent visual attributes (such as color, texture, and shape) even as the image's pose changes. ControlNet pose anchoring techniques, for example, utilize OpenPose or Depth modules, providing skeletal pose or depth information as additional conditions to precisely guide the image generation model to produce images conforming to specific poses, while ensuring visual consistency between the generated and original images. The combined use of these techniques effectively avoids unnatural jumps or deformations in animated characters between consecutive frames, ensuring the visual continuity of the animated characters.
[0067] To optimize animation smoothness and generation efficiency, this application controls the sampling steps of the raw image model and maintains the frame rate of the image frame sequence at 24-30 frames per second, while ensuring a frame generation efficiency of no less than 1.5 seconds per frame. The sampling steps of the raw image model refer to the number of iterations in the diffusion model to gradually denoise and generate a clear image from a noisy image. By finely adjusting the sampling steps, generation speed can be effectively balanced while ensuring image quality. Maintaining a frame rate of 24-30 frames per second is a recognized standard range in the animation field for providing a smooth visual experience, making the animation appear natural and fluid. A frame generation efficiency of no less than 1.5 seconds per frame ensures that the system maintains a high response speed when processing a large number of frames, meeting the real-time requirements of interactive animation.
[0068] In some implementations, the resolution of the generated image frame sequence is also ensured to be compatible with the kindergarten terminal equipment, thereby providing high-quality materials for the animation compositing module. Resolution compatibility means that the pixel size of the generated image frames matches the display capabilities of commonly used display terminal equipment in kindergartens (such as all-in-one machines and tablet computers). This avoids problems such as image blurring, pixelation, or stretching distortion caused by resolution mismatch, ensuring that the animation presents a clear and delicate visual effect on the target device, providing a high-quality visual foundation for subsequent animation compositing.
[0069] In some possible embodiments, the animation synthesis module is proposed to also be used for: the background template library containing at least 20 background templates of grassland, forest, amusement park, and ocean that are suitable for children's aesthetics, supporting manual selection or AI automatic background matching; sorting the image frame sequence, adding inter-frame transition effects, merging the synthesized basic animation with the selected background, adding scene dynamic effects to the merged animation, and outputting the scene-based animation.
[0070] Specifically, the background template library aims to provide a rich variety of scene choices to meet children's imaginative needs for different environments. "Adapting to children's aesthetics" means that these background templates are carefully designed in terms of color, composition, and elements, typically using bright, saturated colors, simple and cute graphic elements, and scenes common in children's daily lives, such as meadows, forests, playgrounds, and oceans, to ensure visual appeal and cognitive comprehensibility. At least 20 templates ensure the richness of the animation scenes, avoiding repetition and monotony. The system supports manual selection or AI-automated background matching. Manual selection allows users (such as parents or teachers) to choose a suitable background template based on the theme of the child's artwork or personal preferences. AI-automated background matching analyzes the image features of the child's artwork (such as main color tone and line style) and semantic understanding results (such as the category of the core image and emotional tendency) to intelligently recommend or directly select the background template that best matches the artwork's mood, thereby improving matching accuracy and efficiency and reducing manual intervention. During animation synthesis, the image frame sequence is sorted to ensure the logical coherence of the animation actions, allowing the image frames to play in the correct sequence. Simultaneously, adding inter-frame transition effects, such as fade-in / fade-out, dissolve, and erase, smoothly connects adjacent image frames, eliminating the abruptness between frames and making the animation playback smoother and more natural. Blending the synthesized basic animation with the selected background involves overlaying the foreground animation onto the background template and adjusting its size, position, lighting, etc., to achieve harmony between the foreground and background, making the animation appear as if it were situated within a real scene. To further enhance the animation's vividness, dynamic scene effects are added to the blended animation. This involves adding small, looping animated elements to the static background to enhance the scene's liveliness and immersion. For example, animations of falling leaves can be added to a forest background, butterflies fluttering or flowers swaying to a grassy background, ripples on water or fish swimming to an ocean background, or shimmering light spots to a playground background. These dynamic effects make the entire scene more vibrant and attract children's attention. Finally, the system outputs a complete scene-based animation file with rich backgrounds and dynamic effects, which can be played directly on the display terminal module for children to watch and interact with.
[0071] This embodiment also presents implementation details of the interactive control module. In some possible embodiments, an interactive control module is proposed to configure interactive functions adapted to young children's operating habits for scene-based animations, and to convert children's interactive commands into control signals to drive the animated characters to respond accordingly. However, in practical applications, ensuring that the interactive control module can efficiently and accurately recognize various interactive commands from young children, such as touch, voice, and gestures, and optimizing it for children's unique operating habits and cognitive characteristics to provide a smooth, real-time interactive experience, is a technical problem that needs further resolution.
[0072] To address the aforementioned issues, the interactive control module of this application includes a touch recognition submodule, a voice recognition submodule, and a gesture recognition submodule. The touch recognition submodule recognizes touch control signals based on touch events from the display terminal; the voice recognition submodule employs a cloud-based API or a local lightweight model adapted to the characteristics of young children's speech, supporting keyword wake-up and voice command parsing; the gesture recognition submodule captures images of children's gestures through a camera and uses the MediaPipe lightweight gesture recognition algorithm to recognize gesture control signals; furthermore, the interactive control module's response time to children's interactive commands is ≤0.5 seconds to ensure real-time interaction.
[0073] Specifically, the interactive control module, as the core component for system-user interaction, integrates multiple recognition sub-modules to receive and parse different forms of interactive commands from the user. This modular design helps improve the system's flexibility and scalability, enabling it to support multiple interaction modes simultaneously and meet the needs of different scenarios and user preferences. The interactive control module can be a software layer responsible for coordinating and managing the data flow and control logic of each recognition sub-module. It receives the recognition results from the sub-modules, converts them into unified control signals, and sends them to the animation generation or display module to drive the animated character to perform corresponding actions.
[0074] The touch recognition submodule is specifically responsible for handling touch operations performed by users through the display terminal. It converts physical touch events into control signals that the system can understand, which is key to achieving intuitive and direct interaction. The touch recognition submodule listens to the touch event API (Application Programming Interface) provided by the operating system of the display terminal (such as a capacitive or resistive screen). When a user touches the screen, the operating system generates a touch event, containing information such as the coordinates of the touch point, touch pressure, and touch time. The touch recognition submodule captures these events and parses them according to preset interaction logic (e.g., clicking a specific area of an animated character to trigger an action, or swiping the screen to switch scenes), generating corresponding control signals. For example, it can be set to trigger a "jump" action control signal when the touch point falls within the bounding box of an animated character.
[0075] The speech recognition submodule enables the system to understand children's verbal commands, providing a natural and convenient interaction method, especially suitable for children who are not yet proficient in fine motor skills. The speech recognition submodule can utilize cloud APIs adapted to the characteristics of children's speech or local lightweight models. For cloud APIs, it can integrate APIs provided by mainstream speech recognition service providers (such as Baidu AI Open Platform, iFlytek Open Platform, Google Cloud Speech-to-Text, etc.). These APIs typically possess powerful speech recognition capabilities and can be optimized for specific groups (such as children) through customized training. The system uploads the collected children's speech data to the cloud API for recognition, receives the returned text results, and then parses the commands using Natural Language Processing (NLP) technology. For local lightweight models, localized lightweight speech recognition models can be deployed for scenarios with high real-time requirements or unstable network environments. These models are typically optimized, consume fewer resources, and have a fast recognition speed. For example, end-to-end deep learning-based models, such as Conformer and RNN-T, can be used and trained using a corpus containing a large amount of children's speech data to improve adaptability to characteristics such as inaccurate pronunciation and uneven speech rate in young children. Keyword wake-up functionality can be implemented through a preset wake-up word detection model, initiating the complete voice command parsing process only when a wake-up word is detected, thus saving resources and improving recognition accuracy.
[0076] The gesture recognition submodule allows young children to interact with animations through body movements, providing an intuitive and expressive non-contact interaction method that helps stimulate their participation. The gesture recognition submodule continuously captures real-time video streams of the child using a built-in or external camera. Then, it calls the MediaPipe lightweight gesture recognition algorithm. MediaPipe is an open-source machine learning solution framework developed by Google, providing various pre-trained models, including a gesture recognition model. This model can detect key hand points in images in real time and recognize predefined gestures (such as waving, clenching a fist, pointing, etc.) based on the position and posture of these key points. Its "lightweight" nature means it can run efficiently on resource-constrained devices. The gesture recognition submodule calls MediaPipe's API, inputting image frames captured by the camera into the model, which outputs the recognized gesture category and confidence level. Based on the recognition results, the system maps specific gestures to corresponding control signals; for example, when a "waving" gesture is recognized, the animated character performs a "goodbye" action.
[0077] To ensure real-time interaction, the response time of the interactive control module to children's interactive commands is controlled to within 0.5 seconds. Achieving low-latency response requires system optimization at both the hardware and software levels. At the hardware level, high-performance processors, fast storage, and low-latency display devices can be used. At the software level, the computational efficiency of the recognition algorithm is optimized, and the overhead of data transmission and processing is reduced. For example, for speech recognition, local lightweight models are prioritized or cloud API calling strategies are optimized; for gesture recognition, MediaPipe itself is lightweight and efficient. In addition, the system can use asynchronous processing and multi-threading technology to process different tasks in parallel, avoiding blocking the main thread. Through comprehensive optimization and testing of the recognition speed, data transmission speed, and animation rendering speed of each submodule, the latency of the entire link from user input to animation response is ensured to be controlled within 0.5 seconds.
[0078] Existing tools for digitizing children's artwork primarily focus on image scanning, storage, and basic editing, lacking capabilities in semantic understanding, creative animation transformation, and extended interactive experiences. For abstract drawings by young children, existing tools struggle to achieve accurate interpretation and dynamic transformation. Furthermore, current AI image and animation generation platforms are typically designed for adult-oriented artwork, exhibiting poor accuracy in semantic understanding and feature extraction of children's drawings, and failing to meet the needs of early childhood education scenarios. Traditional animation production tools suffer from low creation efficiency, poor image consistency, and the inability to achieve real-time interaction.
[0079] This embodiment also discloses an electronic device, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements the steps of the above-described method for generating interactive animations of children's paintings based on multimodal AI. The electronic device also includes a communication interface and an object storage client. The communication interface supports RESTful API data interaction and is used to establish connections with third-party multimodal AI model services and speech recognition services. The object storage client is used to connect to image and audio file storage services to realize the storage and access of animation frames, background materials and final animation files.
[0080] This embodiment also discloses a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the steps of the above-described method for generating interactive animations of children's paintings based on multimodal AI. This computer-readable storage medium is a non-transitory storage medium that supports encrypted data storage and access control to ensure the security of children's artwork data and generated animation content.
[0081] The above description is merely an embodiment of this application and is not intended to limit the scope of protection of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of protection of this application.
Claims
1. A multi-modal AI-based infant drawing interactive animation generation method, characterized in that, include: Collect images of children's paintings, preprocess the images, and output standardized image data that is adapted to the input of a third-party multimodal AI model; The preprocessing includes image normalization, resizing, noise removal, and data augmentation, while preserving the original line and color characteristics of the children's drawings; The standardized image data is input into a third-party multimodal AI model to achieve semantic understanding of children's abstract paintings, extract the core images and key features in the children's abstract paintings, and output image feature description text and image outline data. The image feature description text and the image outline data are input into the image model, and combined with the types of movements that children like to generate image frames with continuous movements; the image model adopts fixed seed, shared latent variables and Controlnet pose anchoring technology to ensure the consistency of images between frames; The image frames with continuous action are sorted into a frame sequence and processed between frames to synthesize a basic animation. Then, the basic animation is combined with the background template library to achieve the fusion of the basic animation and the background, generating a scene-based animation. Configure interactive functions adapted to children's operating habits for the scene-based animation, and output the configured interactive animation to the display terminal to realize real-time interaction between children and the interactive animation.
2. The multi-modal AI-based infant drawing interactive animation generation method according to claim 1, characterized in that, The preprocessing of the children's artwork images includes: Brightness calibration and shadow removal were performed on the collected images of children's paintings to eliminate light interference caused by the collection environment; The pixel values of the children's artwork images are scaled to a preset range to complete the normalization process, and the size of the children's artwork images is adjusted to fit the fixed input size of the third-party multimodal AI model. The noise in the children's paintings is removed by Gaussian filtering, and the paintings are then subjected to random slight rotation, cropping, and color enhancement data augmentation operations.
3. The method for generating interactive animations of children's paintings based on multimodal AI according to claim 1, characterized in that, The process involves inputting the standardized image data into a third-party multimodal AI model to achieve semantic understanding of children's abstract paintings, extracting the core images and key features from the paintings, and outputting image feature description text and image outline data, including: The third-party multimodal AI model is a general multimodal large model that uses the standardized image data to extract visual features and understand semantics from children's abstract paintings. The visual encoder of the third-party multimodal AI model extracts the image features of the standardized image data, and the semantic understanding ability of the third-party multimodal AI model is combined to identify the core image and key attributes in the children's abstract paintings. Based on the image features and semantic understanding results, generate the image outline mask data or edge coordinates corresponding to the children's abstract paintings, and organize and output the structured image feature description text and the image outline data.
4. The method for generating interactive animations of children's paintings based on multimodal AI according to claim 1, characterized in that, The step of inputting the image feature description text and the image outline data into the image model, and generating continuous image frames of actions based on the types of actions that children like, includes: The ComfyUI node-based workflow based on Z-Image-Turbo is used as the inference architecture for the raw image model. The API is called in batches through Python scripts, and the action description prompts are adjusted frame by frame according to the types of actions that children like to perform, so as to achieve gradual changes in the actions. A fixed starting seed is set for the raw image model and incremented frame by frame. The determinism of random noise is controlled by sharing latent variables, and the action pose is anchored by the Controlnet's OpenPose or Depth module. The sampling steps of the raw image model are controlled to keep the frame rate of the image frames between 24-30 frames per second. The number of consecutive image frames for the action is set according to the animation duration requirements to ensure that the outline and color of the image between frames are highly consistent with the original artwork of the children.
5. The method for generating interactive animations of children's paintings based on multimodal AI according to claim 1, characterized in that, The process of sorting and processing the frame sequence of consecutive image frames to synthesize a basic animation, and then combining the basic animation with the background template library to generate a scene-based animation, includes: The image frames with continuous action are sorted into a frame sequence, and inter-frame transition effects such as fade-in / fade-out and dissolve are added to synthesize the basic animation. The background template library contains background templates such as grassland, forest, and playground that are adapted to children's aesthetics. It supports manual selection or automatic background matching by AI based on the characteristics of the core image and the colors of the children's paintings. The basic animation is embedded into the selected background, and the size and position of the basic animation are adjusted to achieve a natural blend. Scene dynamic effects such as fluttering leaves and flickering light spots are added to the background to generate the scene animation.
6. The method for generating interactive animations of children's paintings based on multimodal AI according to claim 1, characterized in that, The provision of interactive functions adapted to children's operating habits for the scene-based animation includes: The configured interactive functions include one or more of touch control, voice command control, and gesture control; The touch control involves touching the animated image on the display terminal to trigger a preset action switch, with a response time of ≤0.5 seconds; The voice command control involves collecting children's voice commands and converting them into control signals using a recognition API that adapts to the characteristics of children's voices, thereby driving the animated character to perform corresponding actions. The gesture control involves recognizing a child's waving hand commands through a camera or depth sensor, converting the waving hand commands into control signals, and controlling the animated character to complete movement actions.
7. A multimodal AI-based interactive animation generation system for children's artwork, characterized in that, include: The artwork acquisition and preprocessing module is used to acquire images of children's artworks, normalize, resize, remove noise, and enhance the data of the images, output standardized image data, and retain the original creative features of the children's artworks. The multimodal AI understanding module is used to input the standardized image data into a third-party multimodal AI model, receive the visual feature extraction results and semantic understanding results returned by the third-party multimodal AI model, and organize and output image feature description text and image outline data. The action frame generation module has a built-in raw image model and ComfyUI node-based workflow. It is used to receive the image feature description text and the image outline data, and generate an image frame sequence with continuous action and consistent image. The animation compositing module is used to process the image frame sequence to synthesize a basic animation, and combine the basic animation with the background template library to achieve the fusion of the basic animation with the background, and output a scene-based animation. The interactive control module is used to configure one or more interactive functions among touch, voice, and gesture for the scene animation, and to convert the child's interactive instructions into control signals to drive the animated character to make corresponding responses. The display terminal module is used to display the configured interactive animation, receive children's interactive operations and provide feedback on the animation response results, and is compatible with kindergarten all-in-one machines and tablet terminals.
8. The interactive animation generation system for children's paintings based on multimodal AI according to claim 7, characterized in that, The interactive control module includes a touch recognition submodule, a voice recognition submodule, and a gesture recognition submodule: The touch recognition submodule recognizes touch control signals based on the touch events of the display terminal; The speech recognition submodule uses a cloud API or a local lightweight model adapted to the characteristics of children's speech, and supports keyword wake-up and voice command parsing; The gesture recognition submodule captures images of children's gestures through a camera and calls the MediaPipe lightweight gesture recognition algorithm to recognize gesture control signals; The interactive control module responds to children's interactive commands in ≤0.5 seconds, ensuring real-time interaction.
9. An electronic device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method for generating interactive animations of children's paintings based on multimodal AI as described in any one of claims 1 to 6; The electronic device also includes a communication interface and an object storage client. The communication interface supports RESTful API data interaction and is used to establish connections with third-party multimodal AI model services and speech recognition services. The object storage client is used to connect to image and audio file storage services to realize the storage and access of animation frames, background materials and final animation files.
10. A medium, a computer-readable storage medium, having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the method for generating interactive animations of children's paintings based on multimodal AI as described in any one of claims 1 to 6; The computer-readable storage medium is a non-transitory storage medium that supports encrypted data storage and access control, ensuring the security of children's artwork data and generated animation content.