Methods, related devices and computer programs for generating animations

By determining the layout and content of animation areas based on demand description information, and combining the rendering engine and large language model to generate animations, the problem of low efficiency and poor quality in existing animation generation technologies is solved, achieving efficient and low-cost animation generation, and improving the automation and visual effects of animations.

CN122336082APending Publication Date: 2026-07-03SHANGHAI HODE INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHANGHAI HODE INFORMATION TECH CO LTD
Filing Date
2026-03-10
Publication Date
2026-07-03

AI Technical Summary

Technical Problem

Existing technologies struggle to generate dynamic graphic animations efficiently, cost-effectively, and with high quality, resulting in limited range of animation content variation, severe homogenization of visual effects, and a tendency for layout errors and content truncation.

Method used

Based on the requirements description information of the target animation, the area layout, content and display actions are determined, rendering code is generated, and the rendering engine is used to process the code to bypass pixel prediction. The Flexbox and Grid algorithms are used to determine the position of visual elements, and the animation is generated by combining a large language model.

Benefits of technology

It improves the automation level and animation quality of animation generation, adapts to content layout, avoids visual element overlap and layout disorder, and enhances the diversity and expressiveness of animation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122336082A_ABST
    Figure CN122336082A_ABST
Patent Text Reader

Abstract

This application provides a method, related apparatus, and computer program product for generating animation. Based on a requirement description of the target animation, the application determines the area layout, content, and display actions for each area within the target animation; based on the area layout, content, and display actions, it generates rendering code; and a rendering engine processes the rendering code to render the target animation. This enables the animation generation process and the generated animation to better adapt to the content layout, providing content and actions, thereby improving the automation level of animation generation while enhancing the animation quality.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular to a method, apparatus, electronic device, computer-readable medium, and computer program product for generating animation. Background Technology

[0002] With the rapid development of video platforms and self-media industries, in order to provide viewers with higher quality content and help them better understand data and its changing trends, content providers and uploaders often choose to use dynamic graphics as an animation format to display data and other content, in order to verify the authenticity and accuracy of the content and help viewers better understand the content they are expected to express.

[0003] Motion graphics (MG), also known as dynamic graphic animation, is a form of expression that uses dynamic presentation of visual elements such as graphics, text, and colors to convey design concepts. It can transform data into intuitive graphics, helping users to more clearly understand data changes and trends.

[0004] Therefore, how to generate dynamic graphics animations more efficiently, at a lower cost, and with higher quality is a matter of concern and urgent need. Summary of the Invention

[0005] This application provides a method, apparatus, electronic device, computer-readable storage medium, and computer program product for generating animations. Based on the requirements for the animation, content, layout, and actions are planned in advance, and corresponding code is generated. Then, animation rendering is performed based on this code. This approach bypasses "pixel prediction" code and directly generates and renders the animation, enabling the animation generation process and the generated animation to better adapt to the content layout and provide content and actions. This not only improves the automation level of animation generation but also enhances the animation quality of the generated animation.

[0006] One aspect of this application provides a method for generating animation, comprising: determining the area layout, area content, and display actions for each area content of the target animation based on the requirement description information for the target animation; generating rendering code based on the area layout, area content, and display actions; and processing the rendering code using a rendering engine to render the target animation.

[0007] Another aspect of this application provides an apparatus for generating animation, comprising: a demand processing module configured to determine, based on demand description information for a target animation, the region layout of each region included in the target animation, the region content of each region, and the display actions for the content of each region; a code generation module configured to generate rendering code based on the region layout, region content, and display actions; and an animation rendering module configured to process the rendering code using a rendering engine to render the target animation.

[0008] In another aspect of this application, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the method for generating animation as provided above.

[0009] Another aspect of this application provides a computer-readable storage medium having computer program instructions stored thereon, which can be executed by a processor to implement the method for generating animations as provided above.

[0010] In another aspect of this application, there is a computer program product including a computer program having stored computer program instructions thereon, which, when executed by a processor, can implement the method for generating animation as provided above.

[0011] The solution provided in this application, based on the requirement description information for the target animation, determines the area layout of each region included in the target animation, the area content of each region, and the display actions for the content of each region; based on the area layout, area content, and display actions, rendering code is generated; the rendering engine processes the rendering code to render the target animation. Therefore, the animation generation process and the generated animation can better adapt to the content layout, providing content and actions, thereby improving the automation level of animation generation while enhancing the animation quality of the generated animation. Attached Figure Description

[0012] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0013] Other features, objects, and advantages of this application will become more apparent from the following detailed description of non-limiting embodiments with reference to the accompanying drawings:

[0014] Figure 1A flowchart illustrating a process for generating an animation, as provided in one embodiment of this application;

[0015] Figure 2 A flowchart illustrating another process for generating animation provided in one embodiment of this application;

[0016] Figure 3 A flowchart illustrating the process of generating an animation in a specific application scenario, provided as another embodiment of this application;

[0017] Figure 4 A schematic diagram of the structure of an apparatus for generating animation provided in an embodiment of this application;

[0018] Figure 5 This is a schematic diagram of the structure of an electronic device suitable for implementing the solutions in the embodiments of this application.

[0019] The same or similar reference numerals in the accompanying drawings represent the same or similar parts. Detailed Implementation

[0020] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0021] In a typical configuration of this application, the terminal and the service network devices each include one or more processors (CPUs), input / output interfaces, network interfaces, and memory.

[0022] Memory may include non-persistent storage in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.

[0023] Computer-readable media include permanent and non-permanent, removable and non-removable media, which can store information by any method or technology. Information can be computer program instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, read-only optical disc (CD-ROM), digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transfer medium that can be used to store information accessible by a computing device.

[0024] As discussed above, how to generate motion graphics animations more efficiently, at a lower cost, and with higher quality is a matter of concern and urgent need.

[0025] In some solutions, to automate the generation of animations (e.g., motion graphics) based on user needs, pre-set animation project files (e.g., Lottie files) can be used as templates. Then, after the user provides their specific requirements, relevant text and images can be determined based on those requirements, and the animation can be rendered and generated by programmatically replacing the text, images, or simple color attributes in the animation project files.

[0026] However, this method severely limits the range of variation in animation content because most of the animation content comes from pre-defined templates. This makes it difficult to meet diverse user needs and results in highly similar and homogenized animations, as all generated animations utilize templates. Furthermore, since the animation and pixel content corresponding to the template are fixed and cannot be changed, animations can only be generated through pixel prediction. This leads to a high probability of layout errors and content truncation in the generated animations, significantly impacting animation quality.

[0027] To address this, this application provides a method for generating animations. This method, based on a requirement description of the target animation, determines the area layout, content, and display actions for each area's content within the target animation. Based on the area layout, content, and actions, rendering code is generated. A rendering engine is then used to process the rendering code to render the target animation. This allows the animation generation process and the generated animation to better adapt to the content layout, providing both content and actions, thus improving both the automation level and the animation quality.

[0028] In practical scenarios, the execution entity of this method can be a user device, a device composed of a user device and a network device integrated through a network, or an application running on the aforementioned devices. User devices include, but are not limited to, various terminal devices such as computers, mobile phones, tablets, smartwatches, and wristbands. Network devices include, but are not limited to, network hosts, single network servers, multiple network server sets, or cloud computing-based computer sets. Here, the cloud consists of a large number of hosts or network servers based on cloud computing. Cloud computing is a type of distributed computing, consisting of a virtual computer composed of a group of loosely coupled computer sets.

[0029] When the executing entity is software, it can be installed in the electronic devices listed above. It can be implemented as multiple software programs or software modules, or as a single software program or software module, without specific limitations.

[0030] Furthermore, the acquisition, storage, use, processing, transportation, provision, and disclosure of any type of information involved in the technical solutions disclosed herein, such as user personal information (e.g., user-provided demand description information as discussed later in this disclosure), comply with relevant laws and regulations and do not violate public order and good morals.

[0031] Figure 1 The present application illustrates a process 100 for generating animation, which includes at least the following processing steps:

[0032] (Step) S101, Based on the requirement description information for the target animation, determine the area layout of each area included in the target animation, the area content of each area, and the display actions for the content of each area.

[0033] In embodiments of this application, a user can send a requirement description to the executing entity, describing their expectations for the animation. The requirement description typically describes the content the user needs to display in the animation, such as the metrics and data types to be displayed, their specific values, and the user's interpretation and understanding of these metrics and data.

[0034] In practice, the implementing entity can provide a user interface for obtaining requirement description information, so that users can interact with the implementing entity by filling in information on the interface to complete the process and purpose of providing requirement description information to the implementing entity.

[0035] Accordingly, after receiving the requirement description information, the executing entity can parse it to determine the areas included in the target animation, the layout of each area, the content of each area, and the display actions for the content of each area.

[0036] Regions are typically associated with the structure of the content to be displayed. In other words, each region can correspond to a specific type of information or content, allowing different regions to be used to present various types of information and content within the animation, based on their unique display functions. For example, regions may include areas for numerical descriptions, numerical displays, explanatory information presentation, and so on.

[0037] Accordingly, the content displayed in the areas (i.e., the area content) can be the specific content that needs to be displayed, corresponding to the display functions of these areas (e.g., explanatory information, numerical values, animation themes, etc.). For example, the direct subject can directly use the user-provided text content included in the requirement description information as the area content of the corresponding area.

[0038] Correspondingly, display actions can be understood as those presentation and display actions associated with a region and its content. Examples include the display of content within a region, actions taken during the display process to achieve dynamic visual effects, and actions to exit or change content within a region. For instance, display actions may include "presenting content within a region," dynamically adjusting at least one visual element within the displayed content (e.g., shaking, flashing, etc.) to achieve dynamic visual effects, and switching actions to switch between two or more groups of content belonging to the same region (e.g., using one region to progressively present complete, two-level, or multi-level content).

[0039] Typically, in this step, the implementing entity can determine the location of each area based on the requirements description information, determine the content of each area based on the function of each area, and adaptively select display actions according to the content of the area.

[0040] In some embodiments, for the layout between regions, the arrangement order between the regions of each function and the size of each region can be pre-selected, so that after the execution subject actually determines the regions, it can determine the position and layout of each region involved based on such arrangement order.

[0041] In some embodiments, the executing entity can also determine the size of each region based on the amount of content in each region after dividing the regions, and then dynamically determine the location occupied by each region based on the size of each region. Accordingly, in such cases, in order to improve the quality of region division, an optional location range, size constraint, etc., can be pre-defined for each type of region so that the regions can be laid out more reasonably.

[0042] In some embodiments, display actions can be pre-set and configured based on the specific circumstances of the area content, so that the executing entity can determine the required display actions after determining the area content. For example, for area content that displays data trends in the form of a "curve," the corresponding display action could be shaking, vibrating, or "flowing upwards" to allow users to more intuitively understand the area content and the "trend of change" it potentially represents. Typically, display actions can include actions in the preparatory stage (e.g., initial presentation or entry of visual elements), actions in the change stage (e.g., continuous change stage of visual elements, such as continuous growth or change of the curve), and actions in the result stage (e.g., the final fixed presentation stage of visual elements) to differentiate the various presentation and display states of the area content.

[0043] In some embodiments, in order to improve the execution subject's ability to understand the requirement description information, and the ability to process, determine the region, region layout, region content, and display actions based on the requirement description information, a parsing agent can be pre-trained so that the execution subject can call the trained parsing agent to quickly and accurately complete the understanding of the requirement description information.

[0044] An intelligent agent is a system or entity capable of perceiving its environment and autonomously making decisions and executing tasks. It can judge and select based on its own goals and changes in the external environment, and take actions to achieve a specific purpose. For example, an intelligent agent can be an aggregation of one or more different models, trained using training samples corresponding to the intended use. For instance, in the scenario described above, where regions, their layouts, content, and display actions are determined based on requirement description information, the intelligent agent can be trained using sample requirement description information and corresponding sample regions, layouts, content, and display actions to enable it to process requirement description information. For example, it can break down the required regions based on the specific content of the sample requirement description information, adaptively lay them out, and determine the required content and display actions. This allows subsequent execution entities to efficiently, accurately, and as expected complete the process of determining the layout, content, and display actions of each region in the target animation based on requirement description information by invoking the intelligent agent.

[0045] In some embodiments, considering that the purpose of motion graphics animation may often be to use data and presented content to express certain trends or potential situations, that is, the broadcaster may often want to use animation to "metaphorically" express certain trends or patterns (e.g., those patterns and trends expressed or explained using explanatory information), in such cases, in order to more intuitively display the area content and to display the area content in conjunction with the user's explanation or description, the executing entity may choose to first determine the area layout of each area included in the target animation based on the requirement description information for the target animation.

[0046] Then, the implementing entity can determine visual metaphor information based on the semantic information of the requirement description information (for example, based on the semantic information of the situation caused by data changes, determine the "metaphorical information" that may be more desired to highlight the "changed data" and the actual changes that the "changed data" has).

[0047] Then, the implementing entity, as an alternative, selects to determine the content of each area and the display actions for each area's content based on the requirement description information and visual metaphor information, rather than simply using the formal content directly read from the requirement description information to determine the display content and actions (e.g., the data directly recorded in the requirement description information and the desired way to display the data). For example, after determining the area content based on the requirement description information, the implementing entity can highlight those area contents that are associated with the metaphorical content indicated by the visual metaphor information (e.g., data that can metaphorically represent trends) by highlighting, enlarging, or otherwise demonstrating content that is not directly described by the user.

[0048] For example, the implementing entity can also use specific display actions to showcase areas of content associated with the metaphorical content indicated by visual metaphorical information (e.g., data that metaphorically represent trends) to highlight and emphasize the changes and impacts they bring. For instance, this could involve using unique entry actions, unique dynamic visual elements, and so on. For example, for data changes that disrupt or negatively impact existing structures or states, dynamic visual elements such as "explosion" or "damage" can be added to metaphorically represent that such data changes have caused a negative impact.

[0049] In some embodiments, "visual metaphorical information" can also be associated with the presentation purpose of the animation. For example, visual metaphorical information can be associated with two presentation purposes: emphasizing changes in data under a specific data type or emphasizing comparisons between data types. This allows the implementing entity to differentiate between these two presentation purposes by adjusting the presentation actions to represent the "metaphor" based on changes in data within the content area, or by adjusting the presentation actions to represent the "metaphor" based on data types within the content area (e.g., the data type that is more advantageous in the comparison dimension).

[0050] In some embodiments, in addition to utilizing the aforementioned agents capable of jointly determining regions, region layouts, region content, and display actions, it is also possible to train a separate agent for each of these functions or parts, and achieve a similar purpose by jointly invoking these agents. This ensures that each agent is trained and maintained independently, reducing the difficulty and cost of agent deployment.

[0051] S102 generates rendering code based on area layout, area content, and display actions;

[0052] In the embodiments of this application, after determining the area layout, area content and display action based on the above-described step S101, the execution entity can generate rendering code based on them in this step.

[0053] In some embodiments, the rendering code can be in the format of HyperText Markup Language (HTML), Cascading Style Sheets (CSS), JavaScript, or other standards that can be edited into a "web page" format using build tools. This allows the animation to be rendered as a "web page" rather than "predicting pixels." This avoids pixel overlap due to a lack of awareness of physical layout, thus improving the quality of the generated animation.

[0054] In some embodiments, the executing entity can generate rendering code by invoking a large model, such as a Large Language Model (LLM). For example, the executing entity can encapsulate the region layout, region content, and corresponding display actions into parameters and "description information" that can be understood by the large model, instructing the large model to generate rendering code that satisfies and implements the region layout, region content, and display actions to render a "web page." An LLM is an artificial intelligence model designed to understand and generate human language, and based on its understanding, the LLM can perform corresponding processing operations to obtain the corresponding processing results. For example, upon obtaining the region layout, region content, and display actions, the LLM can generate such rendering code after understanding the instruction (e.g., generating rendering code that satisfies the region layout, region content, and display actions in the rendered "web page" or "animation").

[0055] LLMs can be trained on large amounts of text data and can perform a wide range of tasks, including text summarization, translation, sentiment analysis, and more. LLMs are characterized by their large scale, typically including a large number of parameters to help them learn complex patterns in language data. These models are often based on deep learning architectures, such as transformers, which helps them provide better processing performance on various natural language tasks. In embodiments of this disclosure, the implementing agent can leverage a generative large language model (e.g., LLM) as a policy model to process descriptive information, generating processing policies for handling risk information to improve the speed and quality of generated processing policies, thereby providing more effective processing strategies faster.

[0056] In some embodiments, during the process of generating rendering code (e.g., calling LLM to generate rendering code), the executing entity may also be constrained in the way the display position of each visual element displayed in the final target animation is determined, so as to avoid the use of predicted pixel position in the rendering code to determine the display position, which would lead to overlapping, overflow, etc. of visual elements.

[0057] In some embodiments, the execution entity may be constrained to use a flexible box algorithm or a mesh algorithm to determine the display position of visual elements in the target animation, or in other words, constrained to use a flexible box algorithm or a mesh algorithm to generate the part of the rendering code corresponding to the display position of visual elements in the target animation.

[0058] The Flexbox algorithm aims to make the arrangement, alignment, and allocation of space for web page elements more flexible and efficient through a simple yet powerful set of properties. Flexbox simplifies the implementation of complex layouts, better adapting to different screen sizes and container dimensions. By setting a container to `display: flex`, Flexbox transforms its child elements into flex items, allowing items to be arranged in different directions and ways without using floats or complex positioning rules. Compared to grid algorithms, Flexbox is more advantageous in one-dimensional layouts.

[0059] Unlike Flexbox, which primarily focuses on one-dimensional layout, the Grid algorithm can handle horizontal and vertical arrangements simultaneously and perform layout in both dimensions, providing more powerful and intuitive layout capabilities in two-dimensional layout scenarios.

[0060] Therefore, by combining the specific dimensions of the animation layout, Flexbox or Grid can be selected accordingly to predict and determine the display position of visual elements in the target animation, thereby improving the quality of the generated target animation.

[0061] In some embodiments, to make the target animation more tailored to the user's needs and expression habits, multiple different styles can be pre-selected and configured. Then, after maintaining and configuring the corresponding graphics code for each style, a target rendering library is formed.

[0062] Accordingly, the implementing entity can provide these styles to the (broadcaster) user, allowing them to further select the desired animation style when generating the animation. For example, the aforementioned acquisition interface can provide controls associated with style selection, enabling the user to select and indicate the target style to be used.

[0063] Subsequently, if the executing entity receives a selection instruction for the target style, it can respond by calling the graphics code corresponding to the target style from the target rendering library (e.g., the shape, size, color, etc. that the visual elements should have, as described in code).

[0064] Then, during the process of generating rendering code based on the area layout, area content, and display actions, the executing entity may, as an alternative, choose to use graphic code as source material to generate rendering code based on the area layout, area content, and display actions. That is, for those parts of the rendering code involving visual elements, if graphic code can be used, the executing entity can use that graphic code as a constraint and reference to generate that part, so that the corresponding parts of the generated rendering code are associated with that graphic code and its image. This ensures that after the animation is rendered using the rendering code, the visual elements in the animation can be presented in this "target style."

[0065] In some embodiments, when generating the portion of rendering code corresponding to the region content, the executing entity may similarly select and constrain pre-configured algorithms and code in a code library to instruct how to draw vector paths, so that the rendering code can instruct these algorithms and code to complete the calculation of data and the drawing of visual elements (e.g., curves, polylines), thereby improving the accuracy of the content in the region content.

[0066] S103 uses the rendering engine to process the rendering code and render the target animation.

[0067] In the embodiments of this application, after the executing entity obtains the rendering code based on the above S102, it can choose to use a rendering engine (e.g., a "browser") to process the rendering code and render the target animation.

[0068] Subsequently, the animation generation method provided in this application, based on the requirement description information for the target animation, determines the area layout of each region included in the target animation, the area content of each region, and the display actions for the content of each region; based on the area layout, area content, and display actions, it generates rendering code; and uses a rendering engine to process the rendering code to render the target animation. Therefore, the animation generation process and the generated animation can better adapt to the content layout, providing content and actions, thereby improving the automation level of animation generation while enhancing the animation quality of the generated animation.

[0069] In some embodiments, in order to enhance the content depth and enrich the content of the target animation, the executing entity may also choose to use at least two "narrative units" to compose the target animation, so that multiple contents from different directions and angles can be combined into the target animation through the narrative units.

[0070] To better understand this situation, please refer to the following: Figure 2 Let's have a discussion. Figure 2 This application illustrates another animation generation process 200 provided in an embodiment of the present application, which includes at least the following processing steps:

[0071] S201, Based on the requirement description information for the target animation, at least two narrative storyboard units are identified;

[0072] Specifically, in this step, after reading the requirement description information, the executing entity can segment the information based on its semantic information and then divide the "requirement" corresponding to the requirement description information into at least two independent narrative storyboard units according to the segmentation results. That is, each segmentation result is assigned to a separate narrative storyboard unit.

[0073] S202, determine the layout of each area corresponding to each narrative storyboard unit, the content of each area corresponding to each area, and the corresponding display actions for the content of each area.

[0074] Specifically, similar to what was discussed in S101 above, in this step, the executing entity can determine the regional layout of each area in each narrative storyboard unit, the regional content of each area, and the corresponding display actions for the content of each area.

[0075] That is, each "narrative storyboard unit" can be treated as an independent whole to complete the process discussed in S101 above, so as to determine the regional layout of each area corresponding to the narrative storyboard unit, the regional content of each area, and the corresponding display actions for the content of each area.

[0076] S203 generates rendering code based on area layout, area content, and display actions;

[0077] S204 uses the rendering engine to process the rendering code and render the target animation.

[0078] The above S203-S204 and such Figure 1 The S102-S103 shown are the same. For the same parts, please refer to the corresponding parts of the previous embodiment. They will not be repeated here.

[0079] In some embodiments, during the execution of S203 described above, the executing entity may further select to determine the unit connection actions between each narrative storyboard unit (e.g., overlaying the area content of a new narrative storyboard unit onto the previous, historical area content, etc.), and then further select, based on the unit connection actions between narrative storyboard units, and the corresponding area layout, area content, and display actions for each narrative storyboard unit, to generate rendering code. This allows the rendering code to also record the switching and connection actions between narrative storyboard units, enabling the rendered target animation to clearly and smoothly switch between narrative storyboard units using these unit connection actions, thereby improving the quality of the target animation.

[0080] In some embodiments, in order to meet different user needs in different scenarios, improve the quality of the target animation, and enhance the automation of the target animation generation, the executing entity may also generate an audio stream associated with the content of the target animation.

[0081] Since embodiments that do not involve narrative storyboard units can be approximated as cases where multiple narrative storyboard units are specifically divided and determined, but there is only one "narrative storyboard unit," to reduce repetitive explanations, this discussion and explanation will only focus on embodiments that include a "narrative storyboard unit." For example, in process 200 above, after the executing entity divides the narrative storyboard units, it can generate corresponding audio streams for each narrative storyboard unit based on the content of the respective regions it includes. For example, this audio stream can be generated based on a pre-determined, user-selected timbre, used for reading aloud or introducing the content of the region. For example, the executing entity can call a Text-to-Speech (TTS) module to generate this audio stream based on the content of the region.

[0082] Then, after generating the audio stream, the execution entity can add or combine it to the previously generated target animation to form the final target animation with audio.

[0083] Similarly, in some embodiments, if it is necessary to present subtitles or other content in the target animation, the executing entity can also generate subtitle information corresponding to the regional content and audio stream through a text recognition module, and add it to the target video accordingly to form a target animation with more information and richer content.

[0084] In some embodiments, if an audio stream is added to the target animation, the executing entity can refer to the audio stream when generating and determining the display action to determine the specific action parameters of the display action, including duration parameters and action frequency parameters. This ensures that when displaying, for example, a narrative storyboard unit, the duration and speed of the audio stream can match the duration and frequency of the display action, improving the display effect of the target animation.

[0085] The discussion will also take the "narrative storyboard unit" example. In this case, the executing entity can first generate the corresponding audio stream for each narrative storyboard unit.

[0086] Then, based on the audio length of the audio stream, the action parameters corresponding to each display action are determined, namely the duration parameter and the action frequency parameter mentioned above. For example, the executing entity can use an animation library, such as GSAP (GreenSock Animation Platform), to write timeline code corresponding to the length of the audio stream. Then, based on the length, the importance of the area content, the richness of the content, etc., the animation speed and delay of the display action are determined accordingly, so that the animation and audio are matched in terms of length and speed (for example, the action frequency of visual elements is matched with the speech rate of the audio stream).

[0087] Then, the actions with such action parameters are used to complete the actions mentioned above, such as generating rendering code in S203.

[0088] Subsequently, after the target animation is generated and rendered, the executing entity can respond by combining the audio streams corresponding to each narrative storyboard unit into the target animation, so that when the target animation has "audio content", it can connect and match the display effect of "audio content" and image content (e.g., displaying actions).

[0089] Based on any of the above embodiments, after generating the rendering code, the executing entity can also choose to "inspect" the rendering code to ensure its quality, thereby improving the quality of the target animation rendered based on the rendering code.

[0090] Accordingly, in some embodiments, the executing entity may choose to check and detect whether the rendering code includes rendering sub-code that calls (or calls to) deprecated library functions. For example, it may statically analyze the generated abstract syntax tree to determine whether it contains rendering sub-code that calls obsolete or incompatible deprecated library functions (Deprecated APIs). If such sub-code exists, the executing entity may respond by determining the updated library function based on the mapping relationship of the deprecated library functions.

[0091] Then, the rendering sub-code is updated using update library functions. This ensures the quality of the rendering code.

[0092] Similarly, in some embodiments, the executing entity can also detect various times in the rendering sub-code, and when, for example, the audio duration does not match the animation duration (e.g., the total duration corresponding to the display action), it can accelerate the animation portion by adjusting the animation duration, for example by injecting a time compression algorithm (Time Scale) into the rendering code, so that the audio and animation durations match.

[0093] In some embodiments, if at least two narrative scene units are involved, the executing entity can also independently encapsulate the rendering (sub)code corresponding to each narrative scene unit in the rendering code to form different domains, thereby avoiding mutual interference between them. For example, the rendering subcode corresponding to each narrative scene unit can be wrapped in an immediately invoked function expression and injected with semicolons and strict mode declarations to prevent variable pollution between different narrative scene units and ensure code safety when splicing multiple segments.

[0094] To enhance understanding, this application also provides a specific implementation scheme based on a particular application scenario. Please refer to it. Figure 3 , Figure 3 This is a flowchart of an animation generation process 300 implemented in a specific application scenario, as provided in an embodiment of this application.

[0095] In process 300, for example, the entire process of generating animation can be performed by a combination of user equipment (e.g., terminal equipment) and network equipment (server). For example, the user equipment can provide an interface 310 to collect the user's requirement description information, and the network equipment can complete the subsequent steps and actions after obtaining the requirement description information.

[0096] For example, in process 300, the user can also select a "style". Therefore, in the acquisition interface 310, in addition to the fill box 313 for collecting and allowing the user to fill in the "requirement description information", a style selection control 312 can also be included so that the user can provide and indicate the style they want to use by triggering the style selection control 312 (for example, the user can obtain the indicator control corresponding to the candidate style by triggering the style selection control 312, and select and indicate the "target style" to be used by clicking the indicator control of the candidate style).

[0097] In such cases, it is also possible to select the presentation style description box 311 associated with the style selection control 312 in the acquisition interface 310, so as to use the style description box 311 to indicate the currently selected "(target) style".

[0098] In addition, the interface 310 can also provide instruction information 314 associated with the fill box 313 to indicate the function of the fill box 313, so as to help users understand the function of the fill box 313.

[0099] For example, a user can provide a requirement description information 315 through the fill box 313. Accordingly, after the user provides and submits the requirement description information 315, the executing entity (e.g., a network device) can respond to this by executing S301 to determine regions 321, 322...32N (where N is a positive integer) based on the requirement description information 315, as well as the layout 331 of the regions, the region content 332 of the regions, and the display action 333 of the region content.

[0100] Next, since process 300 involves "style", the execution subject (e.g., network device) can first execute S302 to call the graphics code 335 corresponding to the style (e.g., "style A" shown in the figure) from the rendering library 330.

[0101] Then, the executing entity (e.g., a network device) can continue to execute S303, using the graphics code 335 as material, and based on the region layout 331, the region content 332, and the display action 333 of the region content, to generate rendering code 340.

[0102] Next, the execution entity (e.g., a network device) can continue to execute S304 to utilize the rendering engine 350 to process the rendering code 340 and render the animation 360.

[0103] This application also provides an apparatus for generating animation, the structure of which is as follows: Figure 4 The apparatus 400 shown includes: a demand processing module 410, configured to determine the area layout, area content, and display actions for each area content of the target animation based on demand description information for the target animation; a code generation module 420, configured to generate rendering code based on the area layout, area content, and display actions; and an animation rendering module 430, configured to process the rendering code using a rendering engine and render the target animation.

[0104] This embodiment exists as a device embodiment corresponding to the above method embodiment. The device for generating animation provided in this embodiment enables the animation generation process and the generated animation to better adapt to the content layout, provide content and actions, and improve the animation quality of the generated animation while improving the automation level of animation generation.

[0105] In some embodiments, the demand processing module 410 includes: a unit splitting submodule, configured to determine at least two narrative storyboard units based on the demand description information for the target animation; and a demand processing submodule, configured to determine the area layout of each area corresponding to each narrative storyboard unit, the area content of each area corresponding to each area, and the display action for the content of each area corresponding to each area.

[0106] In some embodiments, the code generation module 420 is further configured to generate rendering code based on the unit connection actions between narrative storyboard units, and the corresponding area layout, area content and display actions of each narrative storyboard unit.

[0107] In some embodiments, the apparatus 400 further includes: an audio stream generation module configured to generate corresponding audio streams for each narrative storyboard unit; an action parameter determination module configured to determine action parameters corresponding to each display action based on the audio length of the audio stream, wherein the action parameters include duration parameters and action frequency parameters; and an audio stream combination module configured to combine the audio streams corresponding to each narrative storyboard unit into the target animation in response to rendering the target animation.

[0108] In some embodiments, the demand processing module 410 includes: a region layout determination submodule, configured to determine the region layout of each region included in the target animation based on the demand description information for the target animation; a metaphor information determination submodule, configured to determine visual metaphor information based on the semantic information of the demand description information; and a content and action determination submodule, configured to determine the region content of each region and the display action for the content of each region based on the demand description information and the visual metaphor information.

[0109] In some embodiments, the portion of the rendering code corresponding to the display position of a visual element in the target animation is determined based on an elastic box algorithm or a mesh algorithm.

[0110] In some embodiments, the apparatus 400 further includes: a graphics code invocation module configured to invoke graphics code corresponding to the target style from a target rendering library in response to receiving a selection instruction for a target style; and a code generation module 420 further configured to generate rendering code based on the graphics code as material, and on the area layout, area content, and display action.

[0111] In some embodiments, the apparatus 400 further includes: an update library function determination module, configured to determine an update library function based on a mapping relationship of expired library functions in response to rendering sub-code in the rendering code including a call to an expired library function; and a library function update module, configured to update the rendering sub-code using the update library function.

[0112] Based on the same concept, this application also provides an electronic device, a readable storage medium, and a computer program product. The method corresponding to the electronic device can be the animation generation method in the foregoing embodiments, and its problem-solving principle is similar to that method. The electronic device provided in this application includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the methods and / or technical solutions of the foregoing embodiments of this application.

[0113] Electronic devices can be user devices, or devices composed of user devices and network devices integrated through a network, or applications running on the aforementioned devices. User devices include, but are not limited to, various terminal devices such as computers, mobile phones, tablets, smartwatches, and wristbands. Network devices include, but are not limited to, network hosts, single network servers, multiple network server sets, or cloud computing-based computer sets, and can be used to implement some processing functions when setting an alarm clock. Here, the cloud consists of a large number of hosts or network servers based on cloud computing. Cloud computing is a type of distributed computing, consisting of a virtual computer composed of a group of loosely coupled computer sets.

[0114] Figure 5 The diagram illustrates the structure of an electronic device suitable for implementing the methods and / or technical solutions in the embodiments of this application. The electronic device 500 includes a Central Processing Unit (CPU) 501, which can perform various appropriate actions and processes based on a program stored in a Read Only Memory (ROM) 502 or a program loaded from a storage portion 508 into a Random Access Memory (RAM) 503. The RAM 503 also stores various programs and data required for system operation. The CPU 501, ROM 502, and RAM 503 are interconnected via a bus 504. An Input / Output (I / O) interface 505 is also connected to the bus 504.

[0115] The following components are connected to I / O interface 505: an input section 506 including a keyboard, mouse, touchscreen, microphone, infrared sensor, etc.; an output section 507 including a cathode ray tube (CRT), liquid crystal display (LCD), LED display, OLED display, etc., and speakers, etc.; a storage section 508 including one or more computer-readable media such as hard disk, optical disk, magnetic disk, semiconductor memory, etc.; and a communication section 509 including a network interface card such as a LAN (Local Area Network) card, modem, etc. The communication section 509 performs communication processing via a network such as the Internet.

[0116] In particular, the methods and / or embodiments in this application can be implemented as computer software programs. For example, the embodiments disclosed in this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowchart. When the computer program is executed by the central processing unit (CPU) 501, it performs the functions defined in the methods of this application.

[0117] Another embodiment of this application provides a computer-readable storage medium and a computer program product having computer program instructions stored thereon, which can be executed by a processor to implement the methods and / or technical solutions of any one or more embodiments of this application described above.

[0118] Specifically, this embodiment may employ any combination of one or more computer-readable media. A computer-readable medium may be a computer-readable signal medium or a computer-readable storage medium. A computer-readable storage medium may be, for example, a system, apparatus, or device that is, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor, or any combination thereof. More specific examples of computer-readable storage media (a non-exhaustive list) include: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this document, a computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.

[0119] Computer-readable signal media may include data signals propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. Computer-readable signal media may also be any computer-readable medium other than computer-readable storage media, capable of sending, propagating, or transmitting programs for use by or in connection with an instruction execution system, apparatus, or device.

[0120] Program code contained on a computer-readable medium may be transmitted using any suitable medium, including but not limited to wireless, wire, optical fiber, RF, etc., or any suitable combination thereof.

[0121] Computer program code for performing the operations of this application can be written in one or more programming languages ​​or a combination thereof. Programming languages ​​include object-oriented programming languages—such as Java, Smalltalk, and C++—as well as conventional procedural programming languages—such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0122] The flowcharts or block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of devices, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-specific system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0123] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0124] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of modules and units is only a logical functional division, and in actual implementation, there may be other division methods. Taking units as examples, multiple units or page components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be indirect coupling or communication connection between apparatuses or units through some interfaces, and may be electrical, mechanical, or other forms.

[0125] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0126] Furthermore, the functional modules and units in the various embodiments of this application can be integrated into one processing module or unit, or each module or unit can exist physically separately, or two or more units can be integrated into one module or unit. The integrated unit can be implemented in hardware or in the form of hardware plus software functional modules and units.

[0127] The integrated modules and units implemented as software functional modules and units described above can be stored in a computer-readable storage medium. These software functional modules and units, stored in a storage medium, include several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) or processor to execute some steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0128] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.

[0129] Furthermore, it is clear that the word "comprising" does not exclude other units or steps, and the singular does not exclude the plural. Multiple units or devices recited in a device claim may also be implemented by a single unit or device through software or hardware. The terms "first," "second," etc., are used to indicate names and do not indicate any specific order.

Claims

1. A method for generating animation, characterized in that, include: Based on the requirement description information for the target animation, the regional layout of each area included in the target animation, the regional content of each area, and the display actions for the content of each area are determined. Based on the area layout, the area content, and the display action, generate rendering code; The rendering engine processes the rendering code to render the target animation.

2. The method according to claim 1, characterized in that, The process of determining the regional layout of each area included in the target animation, the regional content of each area, and the display actions for the content of each area, based on the requirement description information for the target animation, includes: Based on the requirements description information for the target animation, at least two narrative storyboard units are identified; The layout of each area corresponding to each narrative storyboard unit, the content of each area, and the display actions for the content of each area are determined respectively.

3. The method according to claim 2, characterized in that, The step of generating rendering code based on the region layout, the region content, and the display action includes: Based on the unit connection actions between the narrative storyboard units, and the respective regional layout, regional content, and display actions of each narrative storyboard unit, rendering code is generated.

4. The method according to claim 2, characterized in that, The method further includes: Generate corresponding audio streams for each of the aforementioned narrative storyboard units; Based on the audio length of the audio stream, determine the action parameters corresponding to each of the display actions, wherein the action parameters include duration parameters and action frequency parameters; In response to rendering the target animation, the audio streams corresponding to each of the narrative storyboard units are correspondingly combined into the target animation.

5. The method according to claim 1, characterized in that, The process of determining the regional layout of each area included in the target animation, the regional content of each area, and the display actions for the content of each area, based on the requirement description information for the target animation, includes: Based on the requirement description information for the target animation, the regional layout of each area included in the target animation is determined; Based on the semantic information of the aforementioned requirement description information, visual metaphor information is determined; Based on the demand description information and the visual metaphor information, the content of each area and the display actions for the content of each area are determined.

6. The method according to claim 1, characterized in that, The portion of the rendering code corresponding to the display position of the visual element in the target animation is determined based on the elastic box algorithm or the mesh algorithm.

7. The method according to claim 1, characterized in that, The method further includes: In response to receiving a selection instruction for a target style, the system retrieves the graphics code corresponding to the target style from the target rendering library; and The step of generating rendering code based on the region layout, the region content, and the display action includes: Using the graphic code as material, rendering code is generated based on the area layout, the area content, and the display action.

8. The method according to any one of claims 1-7, characterized in that, The method further includes: In response to the rendering code including rendering sub-code that calls expired library functions, the update library function is determined based on the mapping relationship of the expired library functions; The rendering sub-code is updated using the update library function.

9. An apparatus for generating animation, characterized in that, include: The requirement processing module is configured to determine the regional layout of each region included in the target animation, the regional content of each region, and the display actions for the content of each region based on the requirement description information for the target animation. The code generation module is configured to generate rendering code based on the area layout, the area content, and the display action. The animation rendering module is configured to use the rendering engine to process the rendering code and render the target animation.

10. An electronic device, the electronic device comprising: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1 to 8.

11. A computer-readable medium having stored thereon computer program instructions that can be executed by a processor to implement the method as described in any one of claims 1 to 8.

12. A computer program product comprising a computer program that, when executed by a processor, implements the method as described in any one of claims 1 to 8.