AI-based child animation background automatic generation and point position labeling method and system
By constructing a background material template library and a large language model to generate background prompt words, combining image generation and semantic segmentation models, the automatic generation and point annotation of children's animation backgrounds are realized, solving the problem of time-consuming, labor-intensive and poor interactive animation background production, and improving production efficiency and flexibility.
Patent Information
- Application Number
- CN202510398943.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-01
- Publication Date
- 2025-07-04
AI Technical Summary
In the prior art, animation background production is time-consuming and labor-intensive, difficult to quickly respond to changes in project requirements, and inaccurate generation of prompt words for non-professional users, poor interaction between characters and backgrounds, and lack flexibility and scalability.
Build a template library for children's animation background material, extract natural language descriptions through large language models, generate background prompt words, and use image generation models and semantic segmentation models to automatically generate background pictures and mark points to realize automatic background generation and point annotation.
It improves the efficiency of animation background production, reduces the professional and technical threshold, enhances the universality and ease of use of the system, ensures coordinated interaction between characters and background, supports flexible adjustment of scene elements, and meets diverse creative needs.
Smart Images

Figure CN120259493A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of computer-aided automatic generation of animations, and particularly relates to a method and system for automatically generating and point-position marking of children's animation backgrounds in natural language understanding, image generation, computer vision, and animation material libraries. Background Art
[0002] The production of animated background pictures is a key link in the creation of children's animations and game content. The background not only shows the location and atmosphere where the story takes place, but also stipulates the spatial range of character activities, directly affecting the narrative effect and overall visual experience of the animation. Currently, the creation of backgrounds for two-dimensional animations or games mainly relies on manual drawing and presetting by professional art designers and scene designers. For example, Adobe Photoshop or Clip Studio Paint is used for background design. Although this traditional production method can ensure the beauty and detail performance of background pictures, it also has obvious limitations.
[0003] First of all, the traditional background design process is time-consuming and laborious, and it is difficult to quickly respond to changes in project requirements. Especially in the production of children's animations, it is often necessary to make immediate adjustments according to different scenes and plots, and the existing manual drawing methods are difficult to provide sufficient flexibility. For example, when using Autodesk Maya or Blender for 3D modeling and animation production, changes in background design may require remaking the entire scene. In addition, the scalability and modifiability of background pictures are also limited. The backgrounds drawn by professional designers and the character plots developed under the corresponding backgrounds are usually fixed, and users are greatly restricted in customizing creative content.
[0004] Secondly, with the development of artificial intelligence technology and the progress in the fields of natural language understanding, image generation, and computer vision, new possibilities have been brought to the creation of animated backgrounds. Through AI technology, animated background pictures corresponding to the input natural language description can be automatically generated, thus significantly reducing the technical threshold for background creation. At the same time, currently, simply using AI technology to generate background pictures cannot be directly applied to animation generation, and it is also necessary to deeply understand the background pictures for use in animation generation. The specific challenges are as follows:
[0005] 1) There are challenges in the use of prompts for non-professional users in specific scenarios: When non-professional users use AI to generate backgrounds, they usually describe the scenarios they want by inputting prompts. However, due to the non-professional selection and description methods of prompts, the generated images sometimes do not match the expectations. Especially in complex scenarios, such as backgrounds with high requirements for light and shadow, perspective, hierarchy, etc., users need to have certain visual expression abilities and animation production knowledge, which is a challenge for ordinary users. Therefore, how to guide non-professional users to accurately describe their needs or further optimize prompt processing through AI has become one of the difficult problems in the field of animation background generation.
[0006] 2) There are challenges in the interactivity between characters and backgrounds, and in-depth understanding of background image information is required: In animations, characters not only coexist with the background but also need to interact. For example, the position and movement trajectory of characters in the background need to be coordinated with the objects in the background. For example, a character may pass through a forest or have physical interactions with objects in a room. This requires AI to not only generate static background images but also analyze and understand the deep information in the background, such as the spatial layout, sense of distance, and front-back relationship of objects, to ensure the natural and reasonable dynamic interaction between characters and the background. Summary of the Invention
[0007] Aiming at the challenges of difficult generation of scene prompts in AI painting content creation and the difficulty in understanding the backgrounds generated by AI, the purpose of the present invention is to propose a method and system for automatically generating and point-position marking of children's animation backgrounds based on AI, laying a foundation for the automatic generation of children's animations.
[0008] The specific technical solution to achieve the purpose of the present invention is as follows:
[0009] A method for automatically generating and point-position marking of children's animation backgrounds based on AI includes the following steps:
[0010] Step 1: Construct a template library of children's animation background materials;
[0011] Step 2: Generate background prompts; Extract background information from the natural language text input by children, and perform understanding and relevance expansion on the extracted background descriptions to obtain background prompts;
[0012] Step 3: Generate background images; Match the background prompts with the background material template library to generate background images;
[0013] Step 4: Mark the point positions of the background image; Mark the point positions of the background areas on the generated background image to generate a point-position configuration file of the background image.
[0014] Furthermore, the background material template library described in Step 1 includes a material information table, a material template mapping table, and a point layout table. The material information table records the name, category, and size of the background material. The material template mapping table contains the mapping between the background material name and the templatized and structured background description. The point layout table records the mapping relationship between the background material and the background area points.
[0015] Furthermore, the specific construction process of Step 1 is as follows: First, for various materials with a children's style, they are classified according to the main scenery in the materials into geographical categories, water area categories, and artificial building categories. Among the background materials in the children's animation style, the geographical category has natural land scenery as the main scenery, including hillsides, ground, and grasslands; the water area category has water scenery as the main scenery, including riversides, streamsides, and bridge sides; the artificial building category includes indoor and outdoor scenes of houses, such as artificial facilities, including amusement parks and ski resorts. Secondly, appropriate background pictures are selected according to the categories respectively. The common principle for selecting the three categories of materials is to ensure that there is a background area available for the character to move horizontally. Finally, a material information table, a material template mapping table, and a point layout table are constructed. The material information table records the name, category, and size of the material. The material template mapping table records the mapping between the material and the templatized description of the material, where the templatized description of the material includes identified scenery, scenery priority, attributes, position, relationship, dynamics, lighting conditions, as well as composition rules, emotional color, theme, style, and description consistency. The point layout table includes the point information of the material and the scenery within the material.
[0016] Furthermore, Step 2 is specifically as follows: Build a meta-model of the children's animation story based on the scene requirements of the children's animation; design corresponding information extraction prompt words to drive the large language model to extract the background information of the story from the natural language described by the children's input story; combine the domain knowledge in the meta-model and embed it into the background information expansion prompt words to perform scene-related expansion on the extracted background information to obtain the final background prompt words as the input for the next step.
[0017] Further, the background area points described in step four include: key points for character activities and points for the default arrangement of characters; the key points for character activities are the core positions where the character can interact in the animation, including ground points, water area points, sky points, and object points; among them, the ground points refer to the points on the area where the character can walk, run, or jump on the ground, and these points are usually the basic points in the scene; the water area points are the points where the character can move in the water, involving swimming, floating, or water animation effects; the sky points are the points where the character can fly or jump in the air, applicable to characters with flight ability; the object points are the points where the character can interact with objects in the scene, applicable to climbing, grabbing, and pushing actions; the points for the default arrangement of characters are the initial or default positions of the character in the scene, including the ground entry point, sky entry point, and screen exit point; among them, the ground entry point is the position where the character enters the screen from the ground, usually serving as the starting point of the animation; the sky entry point is the point where the character enters the scene from the sky, applicable to scenes flying in from the air; the screen exit point is the point where the character leaves the screen, used to end the animation or the character's exit action.
[0018] Further, step four is specifically as follows: for the generated background image, design segmentation prompt words, use a semantic segmentation model to obtain the bounding box information of various scenery in the background, design a point-taking strategy, take points for the character to move in each scenery area, and comprehensively obtain the scenery point information to generate a point configuration file and persistently store it.
[0019] A system for implementing the above method, the system includes: an interaction layer, a function layer, a knowledge layer, and a data persistence layer, where:
[0020] The interaction layer is responsible for handling the interaction operations between children and the system, including a touch interaction module and a material display module; the touch interaction module is used to obtain the click and text input interaction information of children; the material display module displays the template materials in the background material library to children in the form of a thumbnail list, supporting children to select and browse.
[0021] The function layer realizes the core functions of automatic generation of animation backgrounds and point marking, specifically including a prompt word expansion module, an image generation module, a point marking module, and a point configuration file generation module; the prompt word expansion module extracts and expands background information based on the natural language text provided by child users to generate background prompt words for background generation; the image generation module combines the specific materials in the background material template library to perform image-to-image operations, or uses text-to-image to generate background images when the materials cannot match the background prompt words; the point marking module is responsible for scene recognition and area marking of the generated background image, generating scene labels and marking the activity points in the scene area; the point configuration file generation module comprehensively processes the point information and performs persistent storage.
[0022] The support layer includes machine learning applications that support the functions of the upper-layer modules, including image generation models, large language models, and semantic segmentation models. Among them, the image generation models include Stable Diffusion based on the latent diffusion model, the LoRA model fine-tuned using children's animation-style background materials, and the ControlNet plugin with image-guided input. The large language model refers to the Spark Cognitive Model called through the API interface, and the semantic segmentation model adopts the Segment Anything Model (SAM).
[0023] The data persistence layer includes a background material template library and various intermediate results and configuration data generated during the system operation.
[0024] The background material template library contains a material information table, a material template mapping table, and a point layout table. The material information table records the name of the background material, the category of the material, and the size. The material template mapping table contains the mapping between the name of the background material and the templated and structured background description. The point layout table records the mapping relationship between the background material and the background area points.
[0025] The various intermediate results and configuration data generated during the system operation include user interaction data files, background generation prompt data files, generated background image files, and scene recognition and area annotation data files. Among them, the user interaction data files record the interaction operations of child users, including clicks and text inputs, ensuring that the system can track the user's interaction behavior and provide a reference for personalized animation generation. The background generation prompt data files save the background prompts generated by the prompt expansion module, which are used to guide the image generation module to select or generate background images that meet the user's intentions, ensuring the accuracy and consistency of background generation. The generated background image files store the background images generated by the image generation module, which meet the user's intentions and can be stored for a long time for subsequent use or user access. The scene recognition and area annotation data files include the mapping relationship between the scene labels generated by the point annotation module and the scene point information.
[0026] The present invention provides an artificial intelligence-based method and system for automatically generating animation backgrounds, aiming to solve the problems of low efficiency, dependence on manual work, lack of interactivity and scalability in the existing background production methods. Compared with the prior art, the present invention has the following beneficial effects:
[0027] 1. The present invention uses artificial intelligence technology to realize the automatic generation of background images, avoiding the cumbersome process of traditional manual drawing by professional artists and improving the efficiency of animation background production.
[0028] 2. The present invention realizes the understanding and expression of the input intention of non-professional users by guiding the user to input a natural language description and combining with a prompt word optimization module, reduces the professional technical threshold of animation background production, enables ordinary users to efficiently complete the construction of animation scenes, and improves the versatility and usability of the system.
[0029] 3. The present invention not only automatically generates static background images, but also analyzes the spatial structure, object layout and hierarchical relationship in the background, so as to realize the coordinated interaction between the character and the background in terms of position, path, action, etc., and improves the coherence and immersion of animation generation.
[0030] 4. The background images generated by the present invention have the support of structured semantic information, can flexibly adjust scene elements according to the plot development or user needs, improve the reuse efficiency of background resources, and meet the diverse and personalized animation creation needs. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] Figure 1 is the flowchart of the method of the present invention;
[0032] Figure 2 is an example diagram of the template description structure of the background materials of the present invention;
[0033] Figure 3 is a schematic diagram of the meta-model in the field of children's animation background of the present invention;
[0034] Figure 4 is the system framework diagram of the present invention;
[0035] Figure 5 is an example diagram of the generated result of the background image of the present invention;
[0036] Figure 6 is an example diagram of the marked result of the background area points of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0037] The present invention will be described in detail below with reference to the drawings and embodiments.
[0038] A method for automatically generating and point marking of children's animation backgrounds based on AI according to the present invention specifically realizes the steps as follows:
[0039] The first step is to construct a background material template library
[0040] The background material template library includes the knowledge necessary for background generation and scene point detection. Its structure is shown in Tables 3, 4, and 5, and it includes three main parts: the material information table, the material template mapping table, and the point arrangement table. Among them, the material information table contains information related to background materials. The material template mapping table is a mapping between the material name and the templated and structured expression of the background material. The point arrangement table includes the mapping relationship between the point configuration information of the scene area and the background material.
[0041] 1) As shown in Table 3, the material information table is used to establish the mapping relationship between the natural language name and the background material. The fields included in the table structure are material name, material path, material size, and material category. The background materials are classified and the actual storage path of the background materials is saved. The primary key is a specific name visible to the user and is used to directly retrieve a specific material in the corresponding background material library. For example, the key is ["Forest Cottage"], and the material path value is " / assets / forestHouse.png".
[0042] Table 3 Background Information Table of the Background Material Template Library in the Present Invention
[0043]
[0044] 2) As shown in Table 4, the material template mapping table contains the templated and structured expression for each specific material. By designing a key-value structure that conforms to the JSON Schema specification, various attributes of the material are standardized, and the order of the attributes extracted from the material and the nested relationship between the attributes are clarified. As Figure 2 shown, this JSON Schema defines a complex structure for describing the content of an image. Among them, scene recognition includes main scenes and secondary scenes. The main scene determines the theme of a background image. As long as the main scene can be matched during subsequent background matching, the matching can be completed. Generally speaking, the fields include scene recognition, priority, attributes, position, relationship, dynamics, lighting conditions, as well as composition rules, emotional color, theme, style, and description consistency, etc. Through this structure, the content and characteristics of an image can be described in detail, providing a standardized data format for image analysis and processing.
[0045] Table 4 Material Template Mapping Table of the Background Material Template Library in the Present Invention
[0046]
[0047] 3) As shown in Table 5, the point layout table stores the default information required for displaying materials, that is, the point configuration data for the background area. The point data for the background area includes the detailed information of different areas in each background image and the data of special points. The area is usually approximately represented by a rectangle, recording the central position, length, and width of the area; while the special point data indicates some specific coordinates, such as the entrance and exit positions of the area. The material template library uses a non-relational database structure to store the specific background material files. In the divided scenery frame area, the number of points to be taken is determined according to the size of the area, and a detailed point mapping table within the scenery area is established, generating a series of candidate points within each scenery frame area. For this purpose, a minimum grid spacing needs to be set for this area first, and this spacing determines the fineness of the grid division, that is, the minimum distance between adjacent points. According to this minimum grid spacing, the scenery frame area is divided into uniform grids, and each grid intersection point is a potential point. According to the specific size, shape of the area and the number of points to be arranged within this area, these grid points are screened. For overly dense or unreasonable points, they can be adjusted or removed according to actual needs. Finally, a point list containing reasonable points is generated, and these points are recorded in the point mapping table of the scenery area in the form of detailed coordinates. After the point list is generated, the next step is to select the specific points actually applied to the animation scene from it.
[0048] Table 5 Point Layout Table of the Background Material Template Library in the Present Invention
[0049]
[0050] Step 2: Generate background prompts
[0051] The meta-model in the field of children's animation backgrounds is a conceptual class diagram structure designed to support the generation and optimization of children's animation backgrounds, used to qualitatively describe concepts such as scenes, background features, etc., and the association relationships between concepts. The meta-model is as Figure 3 shown, and the specific concept definitions are as follows:
[0052] 1. Scene class: Represents a specific background scene in the animation. Each scene contains multiple sub-elements, such as background features, areas, and scenery.
[0053] 2. Background feature class: Used to describe the characteristics of the scene, including time, weather, emotional tone, etc.
[0054] 3. Scenery class: Describes the main elements in the scene, which are the inseparable and elements that need to be presented in the animation background.
[0055] 4. Scenery description class: Describes the landscape feature information within the area, such as water scenery, regional scenery, or man-made scenery, etc. It is a further refinement and explanation of the scenery class.
[0056] 5. Region class: Used to divide and describe the spatial structure within a scene. Each scene is divided into several regions.
[0057] 6. Position class: The relative positions of the scenery within a region, such as the edge of a forest, beside a river, under a willow tree, etc.
[0058] 7. Region feature class: Describes the features of a region in detail, such as terrain, climate, vegetation, etc.
[0059] 8. Scenery association class: Describes the association relationships between sceneries, such as semantic adjacency, complementarity, opposition, etc. in terms of regional position or scenery description.
[0060] 9. Association rule class: Defines the association rules based on region features and scenery associations.
[0061] Extracting background information from the natural language input by children includes dividing story sentences, background information extraction, and associative expansion. The natural language input by children is a complete story text. The story text is sliced into story sentences according to different backgrounds as it unfolds in chronological order according to the plot. Story sentences are the basic elements that make up a story. They can be a single sentence or a group of sentences, including their descriptions of characters, actions, scenes, and interactions, as long as these sentences share the same background information. Understand the background information for each story sentence. Based on the background generation requirements of children's animations, design corresponding information extraction prompt words to drive the large language model to extract background-related information from the story sentences. This step is constructed through the prompt word framework of instruction + text, and techniques such as few-shot prompting and chain of thought are used. The complete prompt words are shown in Table 1.
[0062] Table 1 Prompt word content for generating background prompt words of the present invention
[0063]
[0064] The specific framework of the prompt word is, Prompt = <concept> + <instruction> + <context>+ <input> . The Concept section elaborates on the basic concepts such as story sentences and background prompts in the prompt words. Instruction represents the specific domain tasks expected to be executed by the model. Context represents the external information or additional context that can be used to add few-shot prompts and the intermediate reasoning process of the chain of thought to guide the language model to respond better. Input represents the original prompt words from the user.
[0065] The overall design idea of the prompt words is as follows: First, extract all possible scenic elements from the story sentences and eliminate irrelevant content such as characters and actions. Combine the meta-model to understand the overall background information and use the scene class to define the overall framework of the background. Based on the regional feature class, extract the content related to spatial division, such as "by the river" and "in the forest". The background feature class captures background dimension information such as time, weather, and emotional tone, such as "the morning mist".
[0066] Next, expand and enrich the extracted background information. One is to describe the scenery in detail. The scenery class and the scenery description class define the basic elements and refined features of the scenery, which helps to standardize the detail expansion. The regional feature class clarifies the characteristics of natural and artificial landscapes within the region, providing a direction for supplementing background information. Based on the structure of the scenery description class, add details such as the color, material, and state of the scenery. For example, "the grassland" is expanded to "the emerald green grassland under the sun". Use the regional features and combine with the regional feature class to add natural attributes to the background area, such as "dense shrubs are planted on both sides of the riverbank". Inject background features and adjust the description style through the background feature class. For example, add adjectives such as "quiet" and "lively" according to the emotional tone.
[0067] The related expansion of scenery information is based on the scenery association class: Through the semantic or spatial association rules defined by the scenery association class, add adjacent scenery to the background. The scenery association class provides the logical association relationship between sceneries, such as complementary and adjacent. The association rule class helps to constrain the rationality of the expansion at the rule level. For example, "by the river" can be associated with "willows on the bank" or "bridge". Application of the association rule class: Combine the association rule class to generate reasonable association items through regional features and scenery associations. Coordination of background features: Ensure that the newly added scenery is consistent with the overall background features. For example, do not add a "snowing" scene in summer.
[0068] Text tokenization and correction: Perform tokenization on the background description, correct grammar errors, and ensure language fluency. Check whether the generated prompt words contain characters or animals. If so, remove them again. According to the requirements of the prompt words, convert the background information into the input format suitable for the generation model.
[0069] The third step is to generate the background image
[0070] There are two specific ways to generate the background, namely text-to-image and image-to-image. When the template matching degree is high, the image-to-image method is adopted. At this time, specific materials in the new background material library need to be used as the input source to generate the background. When any material in the existing material library cannot be hit by the background prompt word, or the matching degree is low, the LoRA model trained by the materials in the existing material library is combined, and the text-to-image method is used to generate pictures similar to the style of the existing children's animation background materials. No matter which specific background generation method is used, template matching needs to be carried out first, and the decision is made according to the matching result. Therefore, the detailed process of template matching will be described first.
[0071] First, maintain a template list during runtime, extract the fields of the main scenery and secondary scenery from the structured JSON fields of the material template library, and load them into the template list. Generally speaking, the instruction text for background matching, the complete template list, and the background prompt word are used as the prompt word input for the large model. As shown in Table 2, the prompt word framework here is, Prompt = <concept> + <instruction>+<Input>. The Concept is used to illustrate key concepts such as <background prompt words>, <template list>, <background description>, etc.; Instruction represents a series of rules that the expected model is to execute, and Input represents the background prompt words of the previous step.
[0072] Table 2 Prompt word content for background template matching in the present invention
[0073]
[0074] The background matching process is carried out hierarchically to improve the matching efficiency and accuracy. First, theme matching is performed. The themes included in the template are "nature", "sky", "amusement", etc. Ensure that the theme described in natural language is consistent with the background theme. Materials with a high background theme matching degree will be preferentially selected. After ensuring the theme consistency, other theme branches are excluded, and the matching is further refined to specific scenery. Specifically, a group of background pictures matching the theme are screened out from the material library, and then the most suitable background is selected from them to ensure that the background contains the core scenery related to the main keywords in the natural language description. For common scenery features such as the sky and the sun as secondary scenery, they can be used as secondary candidates for matching. In addition, attention should also be paid to the consistency of the attributes of the main and secondary scenery, including color, size, shape, etc., to ensure the harmony and unity of the overall visual effect. During the scenery matching process, not only physical attributes should be considered, but also the consistency of emotional color and style should be taken into account to ensure that the selected background meets the requirements of the plot in terms of emotional expression and artistic style. To further improve the matching degree, it is also necessary to consider whether the position and composition rules of the background are consistent with the position and composition requirements in the natural language description. This consideration will help ensure that the overall layout of the background meets the needs of the story narrative. Season matching is another important factor. When matching the background, the season to which the background belongs is inferred based on the characteristics of the scenery, and it is ensured that this season is consistent with the season in the description to improve the matching accuracy. Comprehensively considering the matching degrees of background theme, scenery, emotion, style, position composition, etc., the most suitable background material for the natural language description is selected. According to the result of template matching, it can be known the matching degree between the background prompt words and the template description in the material library. If the matching degree is relatively high, the image-to-image method is used for the next step. If the current material library is lacking, the text-to-image method is used to complete the background generation.
[0075] The idea of generating backgrounds through image-to-image is as follows: In the field of image generation, image generation models such as Stable Diffusion usually have a certain degree of uncontrollability when generating images, which means it is difficult for users to precisely guide the model to generate images with specific layouts or structures. To address this issue, conditional control is exerted on the details of the picture layout through ControlNet. The background template image after background matching is used as the input of ControlNet, combined with background prompts. On this basis, processing is carried out according to the above methods and steps.
[0076] The idea of generating backgrounds through text-to-image is as follows: In the case where the background material library cannot cover all children's story scenarios, there may be a problem of mismatch between the material resources and the natural language intentions of child users. Therefore, the LoRA (Low-Rank Adaptation) technology is used to fine-tune an image generation model (such as the Stable Diffusion model) with a specific weight version to train and generate backgrounds with specific styles, especially those that match the children's style of the material library. The stable diffusion weight model fine-tuned in this process is stable_diffusion_sdxl_base_1.0. The entire process includes preprocessing of the dataset, training, and model inference during text-to-image.
[0077] The dataset uses the materials in the background material library as the basis to train a LoRA model with a natural children's style. First, using the picture preprocessing function in the image generation model, all pictures are cropped to the specification of 768 x 448. Then, a large language model is used for picture annotation. During the fine-tuning process, the annotation text of the pictures can be regarded as a trigger word for a certain background style. The features corresponding to the trigger word can be migrated after training, while the features not recorded by the trigger word will be retained. To ensure that the trained model has strong transfer and generalization capabilities, important features such as the main scenery, lighting, and color are annotated in detail. Through this method, it is expected that the trained model can generate images that conform to the children's style and have strong generalization capabilities. During the fine-tuning process, the maximum number of training epochs is set to 10, the model is automatically saved every 2 epochs, the batch size is 1, the learning rate of the U-Net is 0.0001, the learning rate of the text encoder is 0.00001, the total learning rate is 0.0001, and the video memory is RTX 4090 (24GB).
[0078] During the inference process of background generation, the fine-tuned LoRA model is combined with the generated background prompts as the input to perform text-to-image operations, and new information is added to the background material library for the newly generated background material pictures at this time.
[0079] Step 4: Mark the points in the background area
[0080] This step includes constructing segmentation prompt words, performing semantic segmentation using a semantic segmentation model such as the SAM model, taking points in the area by combining the segmentation results with the point-taking strategy, and generating a point configuration file. First, construct the prompt words for semantic segmentation to guide the model to segment specific scenic objects, such as "sky, ground, water". Secondly, after receiving the prompt word instructions, the semantic segmentation model identifies the positions of the corresponding scenic objects and returns the corresponding masks and two-dimensional bounding box coordinates. Specifically, the result of semantic segmentation is a list containing multiple scenic object labels and their bounding box information. Then, when traversing the segmentation results, the system generates a point configuration file based on the bounding box information and object labels of the scenic objects. The divided scenic areas are gridified and divided into equal grids according to the unit length W0. The center point of each grid is used as the position of the grid. Determine the number of points according to the specific size and shape of the bounding box area, and screen these grid points. For overly dense or unreasonable points, they can be adjusted or removed according to actual needs to generate a point list containing a reasonable number of points. Finally, determine the type of points according to the object labels, and record the point list with point types in the point mapping table of the scenic area in the form of detailed coordinates to form a point configuration file.
[0081] Refer to Figure 4 , the AI-based automatic background generation and point annotation system for children's animations of the present invention includes: an interaction layer, a function layer, a knowledge layer, and a data persistence layer, where:
[0082] The interaction layer is responsible for handling the interaction operations between children and the system, including a touch interaction module and a material display module; the touch interaction module is used to obtain the click and text input interaction information of children; the material display module displays the template materials in the background material library to children in the form of a thumbnail list, supporting children to select and browse;
[0083] The function layer realizes the core functions of automatic animation background generation and point annotation, specifically including a prompt word expansion module, an image generation module, a point annotation module, and a point configuration file generation module; the prompt word expansion module extracts and expands background information based on the natural language text provided by child users to generate background prompt words for background generation; the image generation module combines the specific materials in the background material template library to perform image-to-image operations, or uses text-to-image to generate background images when the materials cannot match the background prompt words; the point annotation module is responsible for identifying and regionally annotating the generated background images, generating scenic object labels and annotating the active points in the scenic areas; the point configuration file generation module synthesizes the point information and performs persistent storage;
[0084] The support layer includes machine learning applications that support the functions of the upper-layer modules, including image generation models, large language models, and semantic segmentation models. Among them, the image generation models include Stable Diffusion based on latent diffusion models, the LoRA model fine-tuned using children's animation-style background materials, and the ControlNet plugin with image-guided input. The large language model refers to the Spark Cognitive Model called through the API interface, while the semantic segmentation model uses the Segment Anything Model (SAM).
[0085] The data persistence layer includes a background material template library and various intermediate results and configuration data generated during system operation.
[0086] The background material template library contains a material information table, a material template mapping table, and a point layout table. Among them, the material information table records the name of the background material, the category of the material, and the size. The material template mapping table contains the mapping between the background material name and the templated and structured background description. The point layout table records the mapping relationship between the background material and the background area points.
[0087] The various intermediate results and configuration data generated during system operation include user interaction data files, background generation prompt data files, generated background image files, and scene recognition and area annotation data files. Among them, the user interaction data file records the interaction operations of child users, including clicks and text inputs, to ensure that the system can track the user's interaction behavior and provide a reference for personalized animation generation. The background generation prompt data file stores the background prompts generated by the prompt expansion module, which are used to guide the image generation module to select or generate background images that meet the user's intentions, ensuring the accuracy and consistency of background generation. The generated background image file stores the background images generated by the image generation module, which meet the user's intentions and can be stored for long-term use or user access. The scene recognition and area annotation data file includes the mapping relationship between the scene labels generated by the point annotation module and the scene point information.
[0088] Embodiment
[0089] Next, taking a story sentence as an example, the background automatic generation and background area point detection of the present invention will be introduced in detail. The natural language description from children is: "There is green grass around the small house in the forest, and a hot air balloon is hanging in the sky". Refer to Figure 1 , the automatic generation and point detection of the children's animation background in the embodiment of the present invention mainly include the following steps:
[0090] The first step is to build a background material template library
[0091] The background material library contains knowledge and information about background materials. It includes three parts: the material information table, the material template mapping table, and the point layout table.
[0092] 1) Establish a background material template library. The background material template library includes the material information table, the material template mapping table, and the point layout table. First, for various material templates with a children's style, they are divided into geographical categories, water area categories, and artificial building categories. Geographical categories have the main scenery features of natural land scenery and include horizontal roads suitable for characters to move on. Common backgrounds include "honeycomb", "forest", "garden", "grassland", "cave", "country road", "park", "hillside", "desert", "rocky plateau", "paddy field", "bamboo forest", and "crater", etc. These scenes usually show diverse environments in nature and provide rich exploration and interaction spaces. Water area categories mainly revolve around the water environment and include background materials that can create a waterside atmosphere or are related to water. Common scenes include "beside the bridge", "lake", "pond", "stream side", "waterfall", "beach", "brook", "river bend", "swamp", etc. Artificial building categories focus on the environmental background built by humans and are suitable for depicting daily life, urban, or indoor scenes. Common materials include "amusement park", "birthday party", "hospital", "room", "library", "police station", "classroom", "mall", "railway station", "swimming pool", "art museum", "café", etc. These scenes provide a social background for the character's story line and can set off different emotional and behavioral scenes, such as parties, learning, entertainment, and work, etc.; Second, each material template has a corresponding template description, including the identified scenery, scenery priority, attributes, location, relationship, dynamics, lighting conditions, as well as composition rules, emotional color, theme, style, and description consistency; The point layout table includes information related to the points of materials and the scenery within the materials. The following provides a more detailed description of the construction of each table.
[0093] 2) Establish a material information table. Establish a mapping of background material names. For example: [{key: [["forest cottage"]], value: "forestHouse"}, {key: [["riverside"]], value: "riverBank"}, {key: [["garden"]], value: "garden"}, {key: [["bathroom"]], value: "bathroom"}] and so on. After establishing the name mapping, the name mapping table of the materials binds these names to the material paths in the specific database. The material information table stores the background area point information of the background materials. The categories of the points include the ground, the sky, the water area, and can also include general items. For example, for the "pond" material, three areas of "ground", "river water", and "sky" are defined, and two types of marked points, "entrance and exit points, center points", are marked for each area.
[0094] 3) Establish a mapping table for background material templates. A template is a more detailed structured description of the material. The specific format is a structure defined by a JSON Schema, which contains multiple fields to describe various aspects of the background material. The main fields include basic information such as the title of the picture (title), text description (text description), theme (theme), etc. For the above "pond" material, its template is: {"title":"lake.png","text description":"This is a cartoon landscape painting with a pond and several totem poles. The pond is blue, surrounded by green grass and trees. The totem poles are red with various patterns.","theme":"Nature","scene recognition":{"main scenes":["pond","totem pole"],"secondary scenes":["grass","trees"]},"scene description":[{"scene name":"pond","priority":1,"shape":"rectangle","color":"blue","texture":"smooth","position":"in the middle of the picture","scale comparison":"larger"},{"scene name":"totem pole","priority":2,"shape":"columnar","color":"red","texture":"rough","position":"on both sides of the pond","scale comparison":"smaller"}],"size relationship between scenes":"The pond is larger and the totem pole is smaller","position relationship between scenes":"The pond is in the middle of the picture and the totem poles are on both sides of the pond","lighting conditions":"Sunny","composition rule":"Central composition","emotional color":"cheerful","style":"cartoon"}. Each background material has a corresponding template expression.
[0095] 4) Establish a point layout table. For each material template, the point information that can be included is: ground point, water area point, sky point, object point, ground entry point, sky entry point, picture exit point.
[0096] The second step is to generate background prompts
[0097] Extract background information from the natural language input by children, and combine it with the domain meta-model of the children's animation background to understand and expand the relevance of the extracted background description, so as to obtain a complete background prompt for background generation.
[0098] 1) Taking the natural language description mentioned above as an example, the content is "There is a small house in the forest surrounded by green grass, and there is a hot air balloon hanging in the sky". Use this natural language description as the <input> part in the prompt word framework, and call the API interface of the Spark Cognitive Model. You can ask the model to print the thinking details during debugging, as well as whether the application of the rules meets the intention.
[0099] 2) Scenery recognition:
[0100] Excluding irrelevant information, all the scenes in the input are recognized as: forest, hut, grass, sky, hot air balloon;
[0101] 3) Describe each scene in detail: Forest: Dense forest, tall trees, green leaves, sunlight shining through the gaps between the trees; Hut: A wooden hut with a brown tile roof, a small step in front of the door, and white curtains on the windows; Grass: Green grass, dotted with colorful flowers, and the grass leaves swaying slightly in the wind; Sky: The blue sky, a few white clouds floating, the sky is vast and boundless
[0102] Hot Air Balloon: A colorful hot air balloon, floating in the air, drifting gently with the breeze;
[0103] 4) Think of other possible scenes and expand them: Forest: There is a stream at the edge of the forest. The water is so clear that you can see the bottom and occasionally you can see pebbles in the water; Cottage: There is a garden next to the cottage, which is full of colorful flowers and a path leading to the forest; Grassland: There is a lonely tree in the corner of the grassland, and there is a stone bench under the tree for people to rest; Sky: There are a few birds flying in the sky, and occasionally a few fallen leaves fall;
[0104] 5) Add relative relationships in the scenery: the cabin is located at the edge of the forest, close to the stream; the stream flows by the cabin and eventually flows into the depths of the forest; the garden is next to the cabin, and you can walk into the forest along the path; the lonely tree is located in a corner of the lawn, and the stone bench is under the tree; the hot air balloon floats above the cabin in the sky, and birds fly from a distance, crossing the shadow of the hot air balloon;
[0105] 6) Construct a reasonable visual scene and form a correct grammatical structure. There is a wooden hut on the edge of a dense forest. The green grass around the hut is dotted with colorful flowers. The garden next to the hut is full of bright flowers, and a path winds its way into the deep forest. In a corner of the grass, there is a lonely tree, and under the tree there is a stone bench for people to rest. A clear stream flows at the edge of the forest and eventually disappears into the depths of the forest. There are a few white clouds floating in the sky, and colorful hot air balloons slowly rise into the sky. A few birds fly over the shadow of the hot air balloon and fly away;
[0106] 7) Re-identify and remove people and animals in the background prompt words, identification and removal: Identify a new animal "bird". Start to regenerate from rule 1;
[0107] 8) Regeneration: Identify and remove the people and animals in the prompt. After removal, the prompt is: At the edge of a dense forest, there is a wooden cottage. The green grass around the cottage is dotted with colorful small flowers. The garden beside the cottage is full of bright flowers, and a path winds towards the deep forest. At one corner of the grass, there is a lonely tree, and a stone bench is placed under the tree for people to rest. A clear stream flows at the edge of the forest and finally disappears into the depths of the forest. There are a few white clouds floating in the sky, and a colorful hot air balloon slowly ascends into the sky.
[0108] 1) Natural language description as input
[0109] Take the example text "The grass around the cottage in the forest is green, and there is a hot air balloon hanging in the sky" as part of the prompt. When calling the cognitive large model interface during debugging, require the large model to print its application of the rules and the details of the processing process to check whether the logic conforms to the user's intention.
[0110] 2) Scenery recognition
[0111] Exclude information irrelevant to the background from the input, and only retain the scenery related to the background. The recognition results are as follows:
[0112] Scenery set: forest, cottage, grassland, sky, hot air balloon.
[0113] 3) Refined description of scenery
[0114] Combined with the scenery description class, make a more detailed description of each scenery to ensure visual realizability and emotional saturation:
[0115] Forest: A dense forest with tall and straight trees, the leaves are emerald green and dripping, and the sunlight shines through the gaps between the trees, forming mottled light and shadow.
[0116] Cottage: A wooden cottage with a dark brown tiled roof, the wall shows the texture of the log, and a potted blooming flower is placed on the windowsill.
[0117] Grassland: Green grassland, with little wildflowers dotted among the grass blades. When the gentle breeze blows, the grass sway gently with the wind.
[0118] Sky: A blue sky with a few cotton-like white clouds floating leisurely, and the sky is vast and boundless.
[0119] Hot air balloon: A colorful hot air balloon with patterns mainly in bright red, yellow and blue, hanging high in the sky, slowly floating, and the ropes hanging under the basket sway with the wind.
[0120] 4) Expansion of scenery association
[0121] Based on the association rules of existing scenery, deduce the possible relevant scenery and expand the scene hierarchy:
[0122] Forest: There is a gurgling brook at the edge of the forest. The water in the brook is crystal clear, and there are low shrubs and mosses growing beside it.
[0123] Cottage: There is a simple garden beside the cottage. Colorful flowers are planted in the garden, and a winding path extends from it, leading deep into the forest.
[0124] Meadow: An ancient big tree stands at one corner of the meadow, and a wooden bench is placed under the tree for people to rest.
[0125] Sky: Colorful flags can be faintly seen fluttering under the hot air balloon, adding dynamic elements to the sky.
[0126] 5) Relative relationships of the scenery
[0127] Combining the regional structure and the association of the scenery, clarify the relative positions of each element:
[0128] The cottage is located at the edge of the forest, just a few steps away from the brook.
[0129] The brook winds around the cottage, flows deep into the forest, and converges with the bushes.
[0130] The garden is adjacent to the cottage, and the path leads to the forest along it.
[0131] The meadow surrounds the cottage, and the big tree at one corner of the meadow forms a natural shelter.
[0132] The hot air balloon is suspended in the sky, and its shadow is projected between the cottage and the meadow.
[0133] 6) Scene construction and grammar adjustment
[0134] Combining the above refined content and relative relationships, reconstruct the descriptive text:
[0135] At the edge of the dense forest stands a simple and ancient cottage. Around the cottage is a green meadow dotted with wildflowers, and the fragrance of the flowers wafts in the wind. The brook winds through the forest edge, bypasses the cottage, making a crisp sound of flowing water. The garden beside the cottage is full of colorful flowers, and a winding path starts from the garden and leads into the deep forest. At one corner of the meadow stands an ancient big tree, and the wooden bench under the tree stands quietly, as if waiting for the return of someone. The sky is blue and vast, with a few white clouds floating leisurely, and a colorful hot air balloon slowly rises into the sky, casting its shadow on the meadow, painting a peaceful and harmonious landscape painting.
[0136] 7) Eliminate irrelevant information and regenerate the prompt
[0137] Through rule verification, it is found that no new characters or animals are added, and the background description meets the requirements, so there is no need to regenerate the background prompt.
[0138] Step 3: Background image generation
[0139] The key point of this step is to use the background prompt to match the materials in the local background template material library, and use the pre-set prompt to make the large language model perform template matching. <input> Represents the background prompt of the previous step. Generally speaking, the instruction text for background matching, the complete template list, and the background prompt are used as the prompt input for the large model. Taking the background prompt obtained in the previous step as an example, matching the background material library, the final result obtained is: {key: [[\"Forest cottage\"]], value: \"forestHouse\"}. Therefore, this material is used as the input basis for the image generation model, combined with the background prompt, and the finally generated background is as Figure 5 shown.
[0140] Step 4: Point position marking for the background image area
[0141] Use the background image generated in the previous step as the input. First, check the width and height (w and h) of the image. After the image is uploaded, use a semantic segmentation model such as the SAM (Segment Anything Model) model to perform semantic segmentation on the image. In the way of prompt segmentation, the input text_prompt variable indicates the object to be segmented, such as \"sky, ground, water\". The returned segmentation result is a list containing multiple object labels and their bounding boxes. When traversing the segmentation result, the label of each object (item['label']) determines how the system will generate and process the point positions in this area.
[0142] Generate corresponding points based on the size of the bounding box of the scene (item['box']). The key variables include the starting and ending width and height of the area (height1, height2, width1, width2), and the number of points (width_num). The generated point coordinates are first represented in a normalized form (screen ratio) and then converted to the actual screen coordinates through a function. Each generated point is marked according to its type (multiple types are defined in the PointType enumeration class, such as DEFAULT_POINT, SKY_POINT, etc.). The system packs the generated points and related information into a JSON object, and the key fields include backgroundName (background image name), initPositionList (point list), and pointMarks (point mark list). These fields ensure the integrity and reusability of the background data. Finally, persistently store the generated point configuration file. And the point information for a specific background image can be displayed according to the point configuration file.< / instruction> < / concept> < / context> < / instruction> < / concept>
Claims
1. An AI-based automatic generation and point marking method for children's animation backgrounds, characterized in that, It includes the following steps: Step 1: Construct a template library for background materials of children's animations; Step 2: Generate background prompts; Extract background information from the natural language text input by children, understand and expand the relevant background descriptions for the extracted background descriptions to obtain background prompts; Step 3: Generate background images; Match the background prompts with the background material template library to generate background images; Step 4: Mark the positions of the background image points; Mark the position points of the background area for the generated background image to generate a position configuration file for the background image.
2. The method for automatically generating a children's animation background and marking the point positions in the background area according to claim 1, characterized in that, The background material template library described in Step 1 includes a material information table, a material template mapping table, and a point position arrangement table; The material information table records the name, category, and size of the background material; The material template mapping table contains the mapping between the background material name and the templated and structured background description; The point position arrangement table records the mapping relationship between the background material and the background area point positions.
3. The method for automatically generating a children's animation background and marking points as claimed in claim 1, wherein, For Step 1, the specific construction process includes: First, classify various materials with a children's style according to the scenery in the materials, divided into geographical categories, water area categories, and artificial building categories; The geographical category takes the natural scenery of the land as the scenery, including hillsides, ground, and grasslands; The water area category takes the water scenery as the scenery, including riversides, streamsides, and bridge sides; The artificial building category includes the indoor and outdoor scenes of houses, including amusement parks and ski resorts; Second, select background images according to the categories respectively. The common principle for selecting the three types of materials is to ensure that there is a background area for the character to move horizontally; Finally, construct a material information table, a material template mapping table, and a point position arrangement table; The material information table records the name, category, and size of the material; The material template mapping table records the mapping between the material and the templated description of the material, where the templated description of the material includes the recognized scenery, scenery priority, attributes, position, relationship, dynamics, lighting conditions, as well as composition rules, emotional color, theme, style, and description consistency; The point position arrangement table includes the point position information of the material and the scenery within the material.
4. The method for automatically generating a children's animation background and point position marking according to claim 1, characterized in that, For Step 2, specifically: Construct a meta-model of a children's animation story based on the scene requirements of the children's animation; Design corresponding information extraction prompt words to drive the large language model to extract the background information of the story from the natural language described in the story input by children; Combine the domain knowledge in the meta-model and embed it into the background information expansion prompt words to perform scene-related expansion on the extracted background information to obtain the final background prompt words as the input for the next step.
5. The method for automatically generating a children's animation background and marking points as claimed in claim 1, wherein The background area points described in Step 4 include: key points for character activities and points for the default arrangement of characters; the key points for character activities are the core positions where the character can interact in the animation, including ground points, water area points, sky points, and object points; among them, the ground points refer to the points on the area where the character can walk, run, or jump on the ground, which are the basic points in the scene; the water area points are the points where the character can move in the water, involving swimming, floating, or water animation effects; the sky points are the points where the character can fly or jump in the air, applicable to characters with flight abilities; the object points are the points where the character can interact with objects in the scene, applicable to climbing, grabbing, and pushing actions; the points for the default arrangement of characters are the initial or default positions of the character in the scene, including the ground entry point, sky entry point, and screen exit point; among them, the ground entry point is the position where the character enters the screen from the ground, serving as the starting point of the animation; the sky entry point is the point where the character enters the scene from the sky, applicable to scenes where the character flies in from the air; the screen exit point is the point where the character leaves the screen, used to end the animation or the character's exit action.
6. The automatic generation and point marking method for the children's animation background according to claim 1, characterized in that, Step 4 is specifically as follows: For the generated background image, design segmentation prompt words, use a semantic segmentation model to obtain the bounding box information of various types of scenery in the background, design a point-taking strategy, take points where the character can move within each scenery area, and comprehensively obtain the scenery point information to generate a point configuration file and persistently store it.
7. A system for implementing the method according to claim 1, characterized in that, The system includes: Interaction layer, function layer, knowledge layer, and data persistence layer, where: The interaction layer is responsible for handling the interaction operations between children and the system, including a touch interaction module and a material display module; the touch interaction module is used to obtain the click and text input interaction information of children; the material display module displays the template materials in the background material library to children in the form of a thumbnail list, supporting children to select and browse. The function layer realizes the automatic generation of the animation background and point annotation, specifically including a prompt word expansion module, an image generation module, a point annotation module, and a point configuration file generation module; the prompt word expansion module extracts and expands background information based on the natural language text provided by the child user to generate background prompt words for background generation; the image generation module combines the specific materials in the background material template library to perform image-to-image operations, or uses text-to-image to generate background images when the materials cannot match the background prompt words; the point annotation module is responsible for identifying and annotating the scenery in the generated background image, generating scenery labels and annotating the activity points in the scenery area; the point configuration file generation module comprehensively obtains the point information and performs persistent storage. The support layer includes machine learning applications that support the functions of the upper-layer modules, including image generation models, large language models, and semantic segmentation models. Among them, the image generation models include Stable Diffusion based on the latent diffusion model, the LoRA model fine-tuned using children's animation style background materials, and the ControlNet plugin with image-guided input. The large language model refers to the Spark Cognition Model called through the API interface, and the semantic segmentation model uses the Segment Anything Model. The data persistence layer includes a background material template library and various intermediate results and configuration data generated during system operation. The background material template library contains a material information table, a material template mapping table, and a point layout table. The material information table records the name of the background material, the category of the material, and the size. The material template mapping table contains the mapping between the background material name and the templatized and structured background description. The point layout table records the mapping relationship between the background material and the background area points. The various intermediate results and configuration data generated during system operation include user interaction data files, background generation prompt data files, generated background image files, and scene recognition and area annotation data files. Among them, the user interaction data file records the interaction operations of child users, including clicks and text inputs, to ensure that the system can track the user's interaction behavior and provide reference for personalized animation generation. The background generation prompt data file stores the background prompts generated by the prompt expansion module, which are used to guide the image generation module to select or generate background images that meet the user's intentions, ensuring the accuracy and consistency of background generation. The generated background image files store the background images generated by the image generation module, which meet the user's intentions and can be stored for long-term use or user access. The scene recognition and area annotation data file includes the mapping relationship between the scene labels generated by the point annotation module and the scene point information.