An AI agent-based intangible cultural heritage animation IP intelligent generation method
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SICHUAN TOURISM UNIV
- Filing Date
- 2026-06-22
- Publication Date
- 2026-08-07
AI Technical Summary
然而,非遗文化具有深厚的内涵与严格的符号规范,现有的通用内容生成技术往往缺乏对非遗文化语义的深度理解与合规性校验,导致生成的动漫IP内容容易偏离非遗文化的核心内涵,甚至出现违反内容安全规范的表述
[0041]通过结构化解析提取非遗文化特征关键词,并对其进行文化语义与内容安全双重合规校验,原理上利用语言模型与非遗知识库的匹配度计算识别偏差与违规,从而在生成源头过滤了偏离非遗内涵和不安全的内容,实现了生成内容在文化语义上的准确性与内容安全上的合规性。通过分层约束注入方式构建包含文化约束层、任务约束层和输出约束层的生成提示词,原理上将文化内涵、生成任务与风格规格解耦并分层融合,使得多模态生成模型在调度时能够受到多层条件的精准引导,实现了生成内容在多模态转换中文化特征与风格的一致性。
Smart Images

Figure FT_1 
Figure FT_2
Abstract
Description
Technical Field
[0001] This invention relates to the technical field of artificial intelligence and digital content, and in particular to a method for intelligent generation of intangible cultural heritage animation IPs based on AI Agent. Background Technology
[0002] Currently, in the field of digital protection and dissemination of intangible cultural heritage, animation IP content is typically generated directly using manual creation or simple generative models. However, intangible cultural heritage possesses profound connotations and strict symbolic norms. Existing general content generation technologies often lack a deep understanding of the semantics of intangible cultural heritage and compliance verification, leading to generated animation IP content that easily deviates from the core connotations of intangible cultural heritage and even contains expressions that violate content security regulations. Furthermore, most existing generation methods rely on single models for direct generation, lacking multimodal collaborative scheduling and multi-layered constraint control mechanisms. This makes it difficult to guarantee the cultural accuracy of intangible cultural heritage patterns and forms during multimodal transformation and to perform closed-loop iterative optimization based on user feedback. Consequently, the final generated intangible cultural heritage animation IP content suffers from serious deficiencies in cultural accuracy, stylistic consistency, and content compliance. Therefore, there is an urgent need for an intelligent generation solution for intangible cultural heritage animation IP that can ensure accurate cultural connotations, content security and compliance, and support multimodal collaboration and iterative optimization. Summary of the Invention
[0003] The present invention aims to at least partially solve one of the technical problems in the related art.
[0004] Therefore, the purpose of this invention is to propose an intelligent generation method for intangible cultural heritage animation IP based on AI Agent. This method performs structured analysis and compliance verification of intangible cultural heritage information, generates multimodal content based on hierarchical constraint injection and AI Agent collaborative scheduling, and iteratively optimizes the method by combining user feedback. This achieves intelligent generation and closed-loop improvement of intangible cultural heritage animation IP in terms of cultural accuracy, content security and compliance, and multimodal consistency.
[0005] To achieve the above objectives, this invention proposes an intelligent generation method for intangible cultural heritage animation IPs based on AI Agents, comprising the following steps:
[0006] S1: Receive intangible cultural heritage information input and generation target, perform structured parsing on the intangible cultural heritage information input, and extract keywords of intangible cultural heritage characteristics;
[0007] S2: Perform compliance verification on the intangible cultural heritage characteristic keywords, filter out non-compliant keywords and output a set of compliant keywords;
[0008] S3: Based on the set of compliant keywords and the generation target, construct and generate prompt words through a hierarchical constraint injection method;
[0009] S4: Based on the generated prompts, the multimodal generation model is coordinated and scheduled by the AI Agent to generate multimodal content of intangible cultural heritage animation IP;
[0010] S5: Receive user feedback on the multimodal content, adjust the prompt word constraint parameters based on the feedback and re-execute multimodal generation until the preset output conditions are met.
[0011] In addition, the AI Agent-based intelligent generation method for intangible cultural heritage animation IP proposed according to the present invention may also have the following additional technical features:
[0012] Specifically, the compliance verification of keywords representing intangible cultural heritage characteristics includes:
[0013] The intangible cultural heritage characteristic keywords were subjected to cultural semantic compliance verification and content security compliance verification, respectively.
[0014] Among them, cultural semantic compliance verification is used to identify deviation keywords that deviate from the connotation of intangible cultural heritage, and content security compliance verification is used to identify illegal keywords that violate content security regulations.
[0015] Specifically, the verification of cultural semantic compliance and content security compliance of intangible cultural heritage characteristic keywords includes:
[0016] Semantic analysis of the intangible cultural heritage characteristic keywords is performed using a language model, and the matching degree between the intangible cultural heritage characteristic keywords and the standard cultural features in the intangible cultural heritage knowledge base is calculated.
[0017] When the matching degree is lower than the preset compliance threshold, the corresponding keyword is determined to be a non-compliant keyword;
[0018] For keywords that do not comply with cultural semantics, generate alternative keywords that retain cultural connotations; for keywords that do not comply with content security, generate alternative expressions that comply with security standards.
[0019] Specifically, the method of generating prompt words through hierarchical constraint injection includes:
[0020] The set of compliance keywords is used as the cultural constraint layer, the generation target is used as the task constraint layer, and the style specification parameters are used as the output constraint layer.
[0021] By integrating the cultural constraint layer, the task constraint layer, and the output constraint layer, a generation prompt word containing multiple layers of constraints is constructed.
[0022] Specifically, the generation of multimodal content for intangible cultural heritage animation IPs through the collaborative scheduling of a multimodal generation model using an AI Agent includes:
[0023] Based on the generated prompt words, the AI Agent sequentially schedules the text generation model, image generation model, and video generation model;
[0024] The text generation model generates text content for the intangible cultural heritage animation IP, the image generation model generates character images and scene images based on the text content, and the video generation model generates animation videos based on the character images and scene images.
[0025] Specifically, in the process of generating character images and scene images based on the text content, the image generation model integrates a structural control network to apply structural constraints to the generated images through at least one control method among edge detection, depth estimation, or pose estimation, so as to maintain the cultural accuracy of intangible cultural heritage patterns and forms.
[0026] Specifically, receiving user feedback on the multimodal content, adjusting the prompt word constraint parameters based on the feedback, and re-performing multimodal generation includes:
[0027] Receive multi-dimensional feedback from users on the multimodal content, including cultural accuracy evaluation, style preference adjustment, and content modification suggestions;
[0028] Based on the multi-dimensional feedback, the prompt word constraint parameters are adjusted, and the multimodal generation model is re-scheduled and generated through the AI Agent.
[0029] The structured parsing of intangible cultural heritage information input, and the extraction of intangible cultural heritage characteristic keywords, include:
[0030] The intangible cultural heritage information input is decomposed into intangible cultural heritage item metadata, cultural symbol set and skill element description;
[0031] Extract corresponding intangible cultural heritage characteristic keywords from the metadata of the intangible cultural heritage projects, the set of cultural symbols, and the description of the skill elements.
[0032] A smart generation system for intangible cultural heritage animation IP based on AI Agent, comprising:
[0033] The information input and parsing module is used to receive intangible cultural heritage information input and generation target, perform structured parsing on the intangible cultural heritage information input, and extract keywords of intangible cultural heritage characteristics;
[0034] The keyword compliance verification module is used to verify the compliance of the intangible cultural heritage characteristic keywords, filter out non-compliant keywords, and output a set of compliant keywords.
[0035] The prompt word construction module is used to construct and generate prompt words based on the set of compliant keywords and the generation target through a hierarchical constraint injection method;
[0036] The multimodal collaborative generation module is used to generate multimodal content of intangible cultural heritage animation IP by collaboratively scheduling the multimodal generation model through the generated prompt words and AI Agent.
[0037] The user feedback iteration module is used to receive user feedback information on the multimodal content, adjust the prompt word constraint parameters based on the feedback information, and re-execute multimodal generation until the preset output conditions are met.
[0038] The keyword compliance verification module outputs the verified keywords to the prompt word construction module as input to the cultural constraint layer, and the keywords that fail the verification are filtered out and compliant alternative suggestions are generated and sent back to the prompt word construction module.
[0039] The prompt word construction module integrates the cultural constraint layer, the task constraint layer corresponding to the generation target, and the output constraint layer corresponding to the style specification parameters to construct a generation prompt word containing multiple layers of constraints and output it to the multimodal collaborative generation module.
[0040] Compared with the prior art, the technical solution provided by the present invention has the following beneficial effects:
[0041] By extracting key words representing intangible cultural heritage features through structured parsing and performing dual compliance checks on cultural semantics and content security, the system utilizes the matching degree between a language model and an intangible cultural heritage knowledge base to calculate and identify deviations and violations. This filters out content that deviates from the essence of intangible cultural heritage and is unsafe at the generation source, achieving accuracy in cultural semantics and compliance in content security. A layered constraint injection approach is used to construct generation prompts containing cultural constraint layers, task constraint layers, and output constraint layers. This decouples and integrates cultural connotations, generation tasks, and style specifications in a layered manner, allowing the multimodal generation model to be precisely guided by multiple conditions during scheduling. This ensures consistency of cultural features and style in the generated content during multimodal transformation.
[0042] By sequentially scheduling text, image, and video generation models through an AI Agent, and integrating a structural control network to apply structural constraints during image generation, the system utilizes edge detection, depth estimation, or pose estimation to lock the physical structure of intangible cultural heritage patterns and forms. This avoids distortion or loss of intangible cultural heritage symbols during multimodal generation, achieving high fidelity in the visual presentation of these patterns and forms. Through iterative generation by receiving multi-dimensional user feedback and adjusting prompt constraints, a feedback-based closed-loop optimization mechanism is constructed. This allows the generated results to gradually approach the user's cultural accuracy assessment and style preferences, achieving a high degree of matching between the final output content and user expectations.
[0043] Additional aspects and advantages of the invention will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of the invention. Attached Figure Description
[0044] The above and / or additional aspects and advantages of the present invention will become apparent and readily understood from the following description of the embodiments taken in conjunction with the accompanying drawings, wherein:
[0045] Figure 1 This is a flowchart of the intelligent generation method for intangible cultural heritage animation IP based on AI Agent according to the present invention;
[0046] Figure 2 This is a flowchart of the AI Agent-based intelligent generation system for intangible cultural heritage animation IPs according to the present invention.
[0047] As shown in the figure: 1. Information input and parsing module; 2. Keyword compliance verification module; 3. Prompt word construction module; 4. Multimodal collaborative generation module; 5. User feedback iteration module. Detailed Implementation
[0048] Embodiments of the present invention are described in detail below, examples of which are illustrated in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain the invention, and should not be construed as limiting the invention. Rather, embodiments of the invention include all variations, modifications, and equivalents falling within the spirit and scope of the appended claims.
[0049] The following describes, with reference to the accompanying drawings, the method for intelligent generation of intangible cultural heritage animation IPs based on AI Agent according to an embodiment of the present invention.
[0050] Example 1
[0051] like Figure 1 As shown, this embodiment provides a method for intelligent generation of intangible cultural heritage animation IPs based on AI Agents. This method establishes a closed-loop logic from input parsing to feedback iteration, achieving intelligent generation and closed-loop improvement of intangible cultural heritage animation IPs in terms of cultural accuracy, content security and compliance, and multimodal consistency. It should be understood that this embodiment is merely illustrative and not restrictive. Those skilled in the art can make various changes or substitutions to the execution order and specific implementation methods of the steps without departing from the spirit of the invention.
[0052] Step S1: Receive intangible cultural heritage information input and generation target, perform structured parsing on the intangible cultural heritage information input, and extract keywords of intangible cultural heritage characteristics.
[0053] Specifically, intangible cultural heritage (ICH) information input refers to original materials related to ICH provided by users, which can take the form of at least one of multiple modalities, such as text descriptions, audio explanations, reference images, or video clips. The generation target refers to the specific form of the animation IP that the user expects to output, such as a short video, a set of character design illustrations, or a script outline. The core of structured parsing of ICH information input lies in deconstructing the chaotic original multimodal data into machine-readable structured tags, thereby extracting cultural characteristic keywords that represent the core connotations of ICH. The "structured parsing" referred to in this application refers to the algorithmic process of deconstructing and classifying unstructured data according to preset semantic dimensions, rather than simple full-text keyword extraction. Through this step, the system can understand the user's intent and the connotations of ICH from the source, laying a data foundation for subsequent compliance verification and prompt word construction.
[0054] Step S2: Perform compliance verification on the intangible cultural heritage characteristic keywords, filter out non-compliant keywords and output a set of compliant keywords.
[0055] Specifically, compliance verification is a defensive screening mechanism that ensures generated content does not deviate from the essence of intangible cultural heritage and complies with content security regulations. Since the original information input by users may contain cognitive biases or inappropriate expressions, directly using it to generate content can easily lead to distortions in cultural semantics or the emergence of illegal content. Therefore, the system needs to perform dual verification on the extracted keywords, identifying and removing those biased keywords that deviate from the connotations of intangible cultural heritage and those that violate content security regulations, while outputting a cleaned and corrected set of compliant keywords. The "compliance verification" referred to in this application not only covers conventional sensitive word blacklist matching but also emphasizes the identification of cultural connotation biases based on semantic understanding, representing a deep semantic compliance screening. This step acts as a "firewall" in the generation pipeline, effectively blocking the spread of harmful semantics to downstream generation models.
[0056] Step S3: Based on the set of compliant keywords and the generation target, generate prompt words through a hierarchical constraint injection method.
[0057] Specifically, the hierarchical constraint injection approach is a prompt word engineering strategy that decouples and hierarchically integrates cultural connotations, generation tasks, and style specifications. Traditional prompt word construction often uses single long text concatenation, making it difficult for the model to distinguish the priority of different constraints during generation, easily leading to conflicts where cultural features are overridden by style requirements. Hierarchical constraint injection, on the other hand, maps the compliant keyword set, generation target, and style specification parameters to different constraint levels and integrates them according to a pre-defined logical relationship, constructing generation prompt words containing multiple layers of constraints. The "hierarchical constraint injection" referred to in this application refers to the technical means of hierarchically encoding generation requirements of different dimensions according to their impact on cultural accuracy and injecting them into the model's guidance parameters. In this way, the system provides precise and hierarchically distinct guidance signals for the multimodal generation model, ensuring a high degree of consistency of the generated content under multi-dimensional constraints.
[0058] Step S4: Based on the generated prompt words, the multimodal generation model is coordinated and scheduled by the AI Agent to generate multimodal content of intangible cultural heritage animation IP.
[0059] Specifically, in this embodiment, the AI Agent acts as a multimodal scheduling hub. Multimodal generation models typically include independently running engines such as text generation, image generation, and video generation models. Without a unified scheduling logic, these models can easily operate independently, leading to a disconnect between textual narrative and visual presentation. The AI Agent, based on the task constraints in the generated prompts, schedules each model sequentially according to a specific temporal logic, ensuring that the output of the upstream model serves as the input constraint for the downstream model. This achieves coherent generation from text script to character design, and finally to animated videos. The "AI Agent" referred to in this application is an intelligent agent scheduling program with capabilities for environmental perception, task decomposition, tool invocation, and execution feedback. Its core value lies in breaking the limitations of direct output from a single model and constructing a multimodal collaborative pipeline. Through the central scheduling of the AI Agent, the system achieves morphological coherence and semantic consistency of intangible cultural heritage animation IPs during the multimodal transformation process.
[0060] Step S5: Receive user feedback on the multimodal content, adjust the prompt word constraint parameters based on the feedback, and re-execute multimodal generation until the preset output conditions are met. (Right 1)
[0061] Example 2
[0062] This embodiment builds upon the basic method and flow provided in Embodiment 1, further developing the structured parsing and dual compliance verification steps for intangible cultural heritage information input. It should be understood that this embodiment is merely illustrative and not restrictive; those skilled in the art can make various changes or substitutions to the specific implementation of the parsing dimensions and verification algorithms without departing from the spirit of the invention.
[0063] Step S1: Receive intangible cultural heritage information input and generation target, perform structured parsing on the intangible cultural heritage information input, and extract keywords of intangible cultural heritage characteristics.
[0064] Specifically, the structured parsing of intangible cultural heritage (ICH) information input and the extraction of ICH cultural feature keywords include: decomposing the ICH information input into ICH project metadata, cultural symbol set, and skill element description; and extracting corresponding ICH cultural feature keywords from the ICH project metadata, the cultural symbol set, and the skill element description.
[0065] Step S2: Perform compliance verification on the intangible cultural heritage characteristic keywords, filter out non-compliant keywords, and output a set of compliant keywords. (Right 1)
[0066] Specifically, the compliance verification of intangible cultural heritage characteristic keywords includes: performing cultural semantic compliance verification and content security compliance verification on the intangible cultural heritage characteristic keywords respectively; wherein, cultural semantic compliance verification is used to identify deviation keywords that deviate from the connotation of intangible cultural heritage, and content security compliance verification is used to identify illegal keywords that violate content security regulations; keywords that fail the verification are filtered out, and compliant alternative suggestions are generated.
[0067] This application employs dual verification rather than simple blacklist keyword filtering because intangible cultural heritage is highly context-dependent and semantically ambiguous. Simple keyword filtering can only block literal violations but cannot identify "hidden deviations" where the literal content is compliant but the semantics are off. For example, a user might mistakenly enter "ghostly shadow" as a symbolic description of shadow puppetry. This term might not trigger a conventional security blacklist on a literal level, but it seriously deviates from the core essence of shadow puppetry—the art of "light and shadow"—and carries superstitious connotations. Therefore, the system must identify such deviations through cultural semantic compliance verification, while simultaneously using content security compliance verification to block explicit violations such as violence and vulgarity, thus establishing a layered defense from cultural connotation to security regulations.
[0068] Furthermore, the step of performing cultural semantic compliance verification and content security compliance verification on intangible cultural heritage characteristic keywords includes: performing semantic analysis on the intangible cultural heritage characteristic keywords using a language model, calculating the matching degree between the intangible cultural heritage characteristic keywords and standard cultural features in the intangible cultural heritage knowledge base; when the matching degree is lower than a preset compliance threshold, the corresponding keyword is determined to be a non-compliant keyword; for keywords with non-compliant cultural semantics, generating alternative keywords that retain cultural connotations; for keywords with non-compliant content security, generating alternative expressions that comply with security standards.
[0069] Example 3
[0070] This embodiment builds upon the basic method flow provided in Embodiment 1, further expanding on the prompt word construction step. It should be understood that this embodiment is merely illustrative and not restrictive; those skilled in the art can make various changes or substitutions to the specific implementation of the constraint hierarchy and fusion logic without departing from the spirit of the invention.
[0071] Step S3: Based on the set of compliant keywords and the generation target, generate prompt words through a hierarchical constraint injection method.
[0072] Specifically, the method of constructing generation prompt words through hierarchical constraint injection includes: using the set of compliant keywords as a cultural constraint layer, the generation target as a task constraint layer, and style specification parameters as an output constraint layer; and integrating the cultural constraint layer, the task constraint layer, and the output constraint layer to construct generation prompt words containing multiple layers of constraints.
[0073] Example 4
[0074] This embodiment builds upon the basic method flow provided in Embodiment 1, and further elaborates on the multimodal generation steps. It should be understood that this embodiment is merely illustrative and not restrictive; those skilled in the art can make various changes or substitutions to the specific implementation of scheduling timing and structural control without departing from the spirit of the invention.
[0075] Step S4: Based on the generated prompt words, the multimodal generation model is coordinated and scheduled by the AI Agent to generate multimodal content of intangible cultural heritage animation IP.
[0076] Specifically, the process of generating multimodal content for intangible cultural heritage animation IPs by collaboratively scheduling multimodal generation models through an AI Agent includes: based on the generated prompts, sequentially scheduling a text generation model, an image generation model, and a video generation model through an AI Agent; wherein, the text generation model generates text content for the intangible cultural heritage animation IPs, the image generation model generates character images and scene images based on the text content, and the video generation model generates animation videos based on the character images and scene images.
[0077] Furthermore, in the process of generating character images and scene images based on the text content, the image generation model integrates a structural control network to apply structural constraints to the generated images through at least one control method among edge detection, depth estimation, or pose estimation, so as to maintain the cultural accuracy of intangible cultural heritage patterns and forms.
[0078] The "structural control network" referred to in this application is a neural network module that imposes hard constraints on the physical spatial structure of the generated image by introducing additional conditional input branches during the inference process of an image generation model. Its core mechanism lies in the fact that, in each step of the model's denoising computation, the structural control network fuses the structural constraint signal with the text prompt word signal, forcing the model to adhere to a pre-defined physical structural blueprint while fitting the text semantics, thereby preventing the model from deviating from the target shape during free denoising.
[0079] Example 5
[0080] This embodiment builds upon the basic method flow provided in Embodiment 1, further elaborating on the feedback iteration steps. It should be understood that this embodiment is merely illustrative and not restrictive; those skilled in the art can make various changes or substitutions to the specific implementation of the feedback dimension and parameter mapping logic without departing from the spirit of the invention.
[0081] Step S5: Receive user feedback on the multimodal content, adjust the prompt word constraint parameters based on the feedback, and re-execute multimodal generation until the preset output conditions are met. (Right 1)
[0082] Specifically, receiving user feedback on the multimodal content, adjusting prompt word constraint parameters based on the feedback, and re-executing multimodal generation includes: receiving multi-dimensional feedback from the user on the multimodal content, including cultural accuracy evaluation, style preference adjustment, and content modification suggestions; adjusting the prompt word constraint parameters based on the multi-dimensional feedback, and re-executing the multimodal generation model through AI Agent collaborative scheduling.
[0083] Example 6
[0084] This embodiment applies the AI Agent-based intelligent generation method for intangible cultural heritage animation IPs provided in embodiments 1 to 5 to a real intangible cultural heritage project application scenario, fully demonstrating the closed-loop operation from input parsing to feedback iteration.
[0085] It should be understood that this embodiment is illustrative only and not restrictive. Those skilled in the art can make various changes or substitutions to the type and generation goal of specific intangible cultural heritage projects without departing from the spirit of the invention.
[0086] In this application scenario, the user's intangible cultural heritage information input is set to relevant materials and descriptions of "Shaanxi shadow puppetry," and the generation goal is "to create a short shadow puppet martial arts film." After receiving the intangible cultural heritage information input and the generation goal, the system performs the following processing steps:
[0087] Step S1: Receive intangible cultural heritage information input and generation target, perform structured parsing on the intangible cultural heritage information input, and extract keywords of intangible cultural heritage characteristics.
[0088] Step S2: Perform compliance verification on the intangible cultural heritage characteristic keywords, filter out non-compliant keywords and output a set of compliant keywords.
[0089] Specifically, the system performs cultural semantic compliance verification and content security compliance verification on the extracted intangible cultural heritage characteristic keywords.
[0090] In this input, due to cognitive bias, the user mistakenly described the manipulation technique of shadow puppetry as "puppet string manipulation" in the original material. The system uses a language model to perform semantic analysis on the keyword "puppet string manipulation" and calculates its matching degree with the standard cultural characteristics of "Shaanxi shadow puppetry" in the intangible cultural heritage knowledge base.
[0091] Step S3: Based on the set of compliant keywords and the generation target, generate prompt words through a hierarchical constraint injection method.
[0092] Specifically, the system uses the verified and corrected set of compliant keywords as the cultural constraint layer, the generated goal "to produce a short shadow puppet martial arts film" as the task constraint layer, and the style specification parameters as the output constraint layer.
[0093] The system integrates the cultural constraint layer, the task constraint layer, and the output constraint layer to construct a generation prompt word containing multiple layers of constraints. During the integration process, the system detected that the phrase "martial arts short film" in the task constraint layer might trigger the model's free association with "exaggerated fighting and 3D stereoscopic effect," which potentially conflicts with the "2D flatness and side profile" in the cultural constraint layer. Based on the principle of prioritizing cultural accuracy, the system automatically enhances the weight coefficients of "side profile and light and shadow projection" in the cultural constraint layer during the parameterized integration of prompt words. At the same time, it downgrades the visual expectation of "martial arts" in the task constraint layer to "traditional martial arts" that conforms to the physical laws of shadow puppetry, ensuring that the generated martial arts short film does not deviate from the form norms of shadow puppetry in its action design.
[0094] Step S4: Based on the generated prompt words, the multimodal generation model is coordinated and scheduled by the AI Agent to generate multimodal content of intangible cultural heritage animation IP.
[0095] Specifically, based on the generated prompt words, the system sequentially schedules the text generation model, image generation model, and video generation model through the AI Agent.
[0096] First, the text generation model generates a script outline that conforms to the narrative norms of shadow puppetry and the plot of martial arts, based on multi-layered constraint prompts. Then, the image generation model generates character images and scene images based on the text content. During this process, the system integrates a structural control network, applying structural constraints to the generated images through edge detection control to maintain the cultural accuracy of intangible cultural heritage patterns and forms.
[0097] Step S5: Receive user feedback on the multimodal content, adjust the prompt word constraint parameters based on the feedback, and re-execute multimodal generation until the preset output conditions are met.
[0098] Specifically, the system receives multi-dimensional feedback from users on the initially generated multimodal content, including cultural accuracy evaluation, style preference adjustment, and content modification suggestions.
[0099] After watching the initial video, users submitted the following feedback: In terms of cultural accuracy, users felt that "the side profile features and light and shadow projection effects of the shadow puppets are very accurate"; in terms of style preference adjustment, users suggested that "the colors of the picture should be more traditional warm yellow tones"; and in terms of content modification suggestions, users pointed out that "the characters' martial arts movements are not agile enough and appear stiff".
[0100] Based on the multi-dimensional feedback, the system adjusts the prompt word constraint parameters and re-schedules the multimodal generation model to generate the product through the AI Agent.
[0101] Example 7
[0102] To more clearly demonstrate the irreplaceability of the key features of this invention, two sets of comparative examples are provided below. It should be understood that these embodiments are merely illustrative and not restrictive; the analysis of counterexamples aims to establish a creative defense and refute the tendency to directly apply conventional generative models to the field of intangible cultural heritage.
[0103] Example 8
[0104] This embodiment provides an intelligent generation system for intangible cultural heritage animation IPs based on AI Agents. This system complements the method embodiment in terms of dimensions, focusing on the structural connections between modules and the binding of data flow. It should be understood that this embodiment is merely illustrative and not restrictive. Those skilled in the art can make various changes or substitutions to the module division and hardware implementation without departing from the spirit of the invention.
[0105] The system includes: an information input and parsing module, used to receive intangible cultural heritage (ICH) information input and generation target, perform structured parsing of the ICH information input, and extract ICH cultural characteristic keywords; a keyword compliance verification module, used to perform compliance verification on the ICH cultural characteristic keywords, filter out non-compliant keywords, and output a set of compliant keywords; a prompt word construction module, used to construct generation prompt words based on the set of compliant keywords and the generation target through a hierarchical constraint injection method; a multimodal collaborative generation module, used to generate multimodal content of ICH animation IP based on the generation prompt words and through AI Agent collaborative scheduling of multimodal generation models; and a user feedback iteration module, used to receive user feedback information on the multimodal content, adjust the prompt word constraint parameters based on the feedback information, and re-execute multimodal generation until the preset output conditions are met.
[0106] Regarding the temporal logic and interaction relationships between modules, the core of this system lies in the data flow interaction mechanism between the keyword compliance verification module and the prompt word construction module. The keyword compliance verification module outputs verified keywords to the prompt word construction module as input to the cultural constraint layer, while keywords that fail verification are filtered and compliant alternative suggestions are generated and sent back to the prompt word construction module. The prompt word construction module integrates the cultural constraint layer, the task constraint layer corresponding to the generation target, and the output constraint layer corresponding to the style specification parameters to construct generated prompt words containing multiple layers of constraints and output them to the multimodal collaborative generation module.
[0107] In the description of this specification, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this invention, "a plurality of" means at least two, such as two, three, etc., unless otherwise explicitly specified.
[0108] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the present invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.
[0109] Although embodiments of the present invention have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those skilled in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of the present invention.
Claims
1. A method for intelligently generating intangible cultural heritage animation IPs based on AI Agents, characterized in that, Includes the following steps: S1: Receive intangible cultural heritage information input and generation target, perform structured parsing on the intangible cultural heritage information input, and extract keywords of intangible cultural heritage characteristics; S2: Perform compliance verification on the intangible cultural heritage characteristic keywords, filter out non-compliant keywords and output a set of compliant keywords; S3: Based on the set of compliant keywords and the generation target, construct and generate prompt words through a hierarchical constraint injection method; S4: Based on the generated prompts, the multimodal generation model is coordinated and scheduled by the AI Agent to generate multimodal content of intangible cultural heritage animation IP; S5: Receive user feedback on the multimodal content, adjust the prompt word constraint parameters based on the feedback and re-execute multimodal generation until the preset output conditions are met.
2. The method for intelligent generation of intangible cultural heritage animation IP based on AI Agent according to claim 1, characterized in that, The compliance verification of keywords related to intangible cultural heritage features includes: The intangible cultural heritage characteristic keywords were subjected to cultural semantic compliance verification and content security compliance verification, respectively. Among them, cultural semantic compliance verification is used to identify deviation keywords that deviate from the connotation of intangible cultural heritage, and content security compliance verification is used to identify illegal keywords that violate content security regulations.
3. The method for intelligent generation of intangible cultural heritage animation IP based on AI Agent according to claim 1, characterized in that, The verification of cultural semantic compliance and content security compliance of keywords related to intangible cultural heritage features includes: Semantic analysis of the intangible cultural heritage characteristic keywords is performed using a language model, and the matching degree between the intangible cultural heritage characteristic keywords and the standard cultural features in the intangible cultural heritage knowledge base is calculated. When the matching degree is lower than the preset compliance threshold, the corresponding keyword is determined to be a non-compliant keyword; For keywords that do not comply with cultural semantics, generate alternative keywords that retain cultural connotations; for keywords that do not comply with content security, generate alternative expressions that comply with security standards.
4. The method for intelligent generation of intangible cultural heritage animation IP based on AI Agent according to claim 1, characterized in that, The method of generating prompt words through hierarchical constraint injection includes: The set of compliance keywords is used as the cultural constraint layer, the generation target is used as the task constraint layer, and the style specification parameters are used as the output constraint layer. By integrating the cultural constraint layer, the task constraint layer, and the output constraint layer, a generation prompt word containing multiple layers of constraints is constructed.
5. The method for intelligent generation of intangible cultural heritage animation IP based on AI Agent according to claim 1, characterized in that, The generation of multimodal content for intangible cultural heritage animation IPs through the collaborative scheduling of a multimodal generation model using an AI Agent includes: Based on the generated prompt words, the AI Agent sequentially schedules the text generation model, image generation model, and video generation model; The text generation model generates text content for the intangible cultural heritage animation IP, the image generation model generates character images and scene images based on the text content, and the video generation model generates animation videos based on the character images and scene images.
6. The method for intelligent generation of intangible cultural heritage animation IP based on AI Agent according to claim 5, characterized in that, In the process of generating character images and scene images based on the text content, the image generation model integrates a structural control network to apply structural constraints to the generated images through at least one control method among edge detection, depth estimation, or pose estimation, so as to maintain the cultural accuracy of intangible cultural heritage patterns and forms.
7. The method for intelligent generation of intangible cultural heritage animation IP based on AI Agent according to claim 1, characterized in that, The step of receiving user feedback on the multimodal content, adjusting the cue word constraint parameters based on the feedback, and re-performing multimodal generation includes: Receive multi-dimensional feedback from users on the multimodal content, including cultural accuracy evaluation, style preference adjustment, and content modification suggestions; Based on the multi-dimensional feedback, the prompt word constraint parameters are adjusted, and the multimodal generation model is re-scheduled and generated through the AI Agent.
8. The method for intelligent generation of intangible cultural heritage animation IP based on AI Agent according to claim 1, characterized in that, The structured parsing of intangible cultural heritage information input, and the extraction of intangible cultural heritage characteristic keywords, include: The intangible cultural heritage information input is decomposed into intangible cultural heritage item metadata, cultural symbol set and skill element description; Extract corresponding intangible cultural heritage characteristic keywords from the metadata of the intangible cultural heritage projects, the set of cultural symbols, and the description of the skill elements.
9. A smart generation system for intangible cultural heritage animation IP based on AI Agent, characterized in that, include: The information input and parsing module is used to receive intangible cultural heritage information input and generation target, perform structured parsing on the intangible cultural heritage information input, and extract keywords of intangible cultural heritage characteristics; The keyword compliance verification module is used to verify the compliance of the intangible cultural heritage characteristic keywords, filter out non-compliant keywords, and output a set of compliant keywords. The prompt word construction module is used to construct and generate prompt words based on the set of compliant keywords and the generation target through a hierarchical constraint injection method; The multimodal collaborative generation module is used to generate multimodal content of intangible cultural heritage animation IP by collaboratively scheduling the multimodal generation model through the generated prompt words and AI Agent. The user feedback iteration module is used to receive user feedback information on the multimodal content, adjust the prompt word constraint parameters based on the feedback information, and re-execute multimodal generation until the preset output conditions are met.
10. The method for intelligent generation of intangible cultural heritage animation IP based on AI Agent according to claim 9, characterized in that, The keyword compliance verification module outputs the verified keywords to the prompt word construction module as input to the cultural constraint layer, and the keywords that fail the verification are filtered out and compliant alternative suggestions are generated and sent back to the prompt word construction module. The prompt word construction module integrates the cultural constraint layer, the task constraint layer corresponding to the generation target, and the output constraint layer corresponding to the style specification parameters to construct a generation prompt word containing multiple layers of constraints and output it to the multimodal collaborative generation module.