A generative animation shot building method, system, device and storage medium
Patent Information
- Application Number
- CN202610971632.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-01
- Publication Date
- 2026-09-22
AI Technical Summary
[0006]本申请提供一种生成式动画分镜构建方法、系统、设备及存储介质,用以解决现有技术中多角色动画分镜生成过程中存在的站位重叠、透视比例失真、动作拖影以及镜头边界错分的问题
本申请首先通过获取包含对话文本和动作指令的剧本数据以及镜头拆分规则数据和参数确定规则数据,为后续处理提供基础输入;然后基于导演视角知识数据库对剧本数据进行语义解析,生成包含环境信息和光影信息的视觉数据,从而将抽象的文学描述转化为具体的视觉要素;
Smart Images

Figure CN122798960A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of animation production technology, and in particular to a generative animation storyboard construction method, system, device and storage medium. Background Technology
[0002] With the rapid development of the digital content industry, the animation production field has placed higher demands on the efficiency of storyboard generation. Generative models, due to their powerful semantic understanding and image generation capabilities, are gradually being applied to the automated construction process from text scripts to visual storyboards. Such methods are expected to significantly shorten the pre-production cycle of animation and reduce labor costs.
[0003] In existing animation storyboard generation solutions, large language models are typically used to directly parse the script text. The model outputs a textual storyboard table containing shot numbers, shot descriptions, and scene descriptions. Some solutions further input the generated text descriptions into an image generation model to obtain the corresponding storyboard images. These methods have achieved the conversion from text to visual to a certain extent.
[0004] Currently, the existing solutions closest to this application mainly fall into three categories: The first category directly inputs the script text into a large language model, which generates a storyboard table with zero or few samples. This type of solution does not include constraints on character spatial coordinates and cannot handle the problem of overlapping positions of multiple characters. The second category uses a text-image multimodal model to parse the script and generate storyboard data containing visual descriptions. Although this type of solution introduces visual features, it lacks a depth-of-field layer segmentation mechanism, and the perspective relationship between characters remains unclear. The third category is based on a rule engine to divide the script into shots according to dialogue turns. The shot segmentation of this type of solution relies entirely on predefined rules and lacks the ability to analyze peak motion amplitude and segment shot attribution probability, resulting in inaccurate shot boundary segmentation in continuous action scenes. None of the above three types of solutions jointly constrain character spatial coordinates, depth-of-field layers, peak motion amplitude, and shot attribution probability. Therefore, in multi-character interaction scenes, there are problems such as overlapping positions, perspective distortion, motion blur, and missegmentation of shot boundaries.
[0005] However, when dealing with multi-character interaction scenes, such solutions often struggle to accurately depict the spatial relationships between characters, leading to overlapping positions or distorted perspective. When describing continuous actions, the storyboards generated by the model often have visual flaws such as motion blur or limb distortion. Furthermore, the consistency of the art style between different shots in the same scene is difficult to guarantee, making it difficult for the output storyboard data to directly meet the standardized requirements of subsequent animation production. Therefore, existing technologies suffer from insufficient compatibility between the quality of generated storyboard data and production standards. Summary of the Invention
[0006] This application provides a generative animation storyboard construction method, system, device, and storage medium to solve the problems of overlapping positions, perspective distortion, motion blur, and misaligned shot boundaries in the generation of multi-character animation storyboards in the prior art.
[0007] To address the aforementioned technical problems, in a first aspect, this application provides a generative animation storyboard construction method, comprising: Acquire script data, shot splitting rule data, and parameter determination rule data for the target animation. The script data includes dialogue text for multiple characters and action instructions for each character. Based on a pre-built director's perspective knowledge database, the script data is semantically parsed to generate the visual data of the target animation, which includes environmental information and lighting information. Based on the environmental information in the visual data and the directional description information in the action instructions, a virtual coordinate space is constructed to determine the spatial coordinate information of each character. The depth information is determined based on the correspondence between the depth layer number and the depth blur parameter, and combined with the character height information and the character's line of sight direction information to form scene scheduling data. Based on the scene scheduling data, continuous action information is extracted from the action instructions of each character, the action amplitude value is calculated, and the instantaneous posture at which the action amplitude value reaches its peak is selected to generate key action freeze data. The visual data, scene scheduling data, and key motion freeze data are input into a pre-built large language model. The large language model determines the shot attribution probability distribution of the input segments based on the shot splitting rule data, and sets the splitting boundary when the shot attribution label of adjacent input segments changes. Based on the parameter determination rule data, the model generates shot sequences, shot parameters, angle parameters, and scene description text, and outputs the shot breakdown data.
[0008] Secondly, this application provides a generative animation storyboard construction system, comprising: The acquisition module is used to acquire script data, shot splitting rule data and parameter determination rule data of the target animation. The script data includes dialogue text of multiple characters and action instructions for each character. The parsing module is used to perform semantic parsing on the script data based on a pre-set director's perspective knowledge database, and generate visual data for the target animation, including environmental information and lighting information. The construction module is used to construct a virtual coordinate space based on the environmental information in the visual data and the orientation description information in the action instructions, determine the spatial coordinate information of each character, determine the depth information based on the correspondence between the depth layer number and the depth blur parameter, and combine the character height information and the character's line of sight information to form scene scheduling data. The extraction module is used to extract continuous action information from the action instructions of each character based on the scene scheduling data, calculate the action amplitude value, and select the instantaneous posture when the action amplitude value reaches its peak to generate key action freeze data. The segmentation module is used to input the visual data, the scene scheduling data, and the key action freeze data into a pre-built large language model, so that the large language model determines the shot attribution probability distribution of the input segment according to the shot segmentation rule data, sets the segmentation boundary when the shot attribution label of adjacent input segments changes, and generates shot sequence, shot type parameters, angle parameters, and screen description text according to the parameter determination rule data, and outputs the shot segmentation data.
[0009] Thirdly, this application provides an electronic device, comprising: Memory, used to store computer programs; A processor for executing the computer program to implement the steps of the generative animation storyboard construction method as described in the first aspect above.
[0010] Fourthly, this application provides a computer-readable storage medium storing a computer program that, when executed by a processor, can implement the steps of the generative animation storyboard construction method described in the first aspect above.
[0011] The technical solution provided in this application has the following beneficial effects: This application first obtains script data containing dialogue text and action instructions, as well as shot splitting rule data and parameter determination rule data, to provide basic input for subsequent processing; then, based on the director's perspective knowledge database, it performs semantic analysis on the script data to generate visual data containing environmental information and lighting information, thereby transforming abstract literary descriptions into concrete visual elements. Next, scene scheduling data containing spatial coordinate information, depth information, character height information, and character gaze direction information is constructed based on visual data and orientation description information to clarify the spatial hierarchy of the character in the picture. Then, continuous motion information is extracted from motion instructions based on scene scheduling data, motion amplitude values are calculated, and peak instantaneous posture is selected to generate key motion freeze data, converting the dynamic process into a static freeze description. Finally, visual data, scene scheduling data, and key motion freeze data are input into the big language model. The big language model divides the shot sequence according to the shot splitting rules and determines the shot parameters, angle parameters, and image description text for each shot sequence, outputting structured shot data. This results in shot outputs that have clear spatial hierarchy, clear action expression, and reasonable shot scheduling.
[0012] Furthermore, this application converts visual data, scene scheduling data, and key action freeze-frame data into embedding vectors corresponding to input segments through the input layer of a large language model. A self-attention layer assigns weights to the embedding vectors corresponding to the input segments based on shot segmentation rules to obtain the probability distribution of each input segment belonging to different shot units. A sequence segmentation layer divides the input segment sequence into multiple shot units based on this probability distribution and combines them into a shot sequence. Then, based on parameter determination rules, it determines the shot size and angle parameters from the number of characters, character spatial coordinates, character gaze direction, and character height information corresponding to each shot unit. Finally, the output layer merges the environmental information, lighting information, character posture information, and character spatial coordinates information corresponding to each shot unit into a scene description text. The shot size parameters, angle parameters, and scene description text of each shot unit are then combined according to the segmentation order to form shot data.
[0013] Furthermore, this process enables end-to-end conversion from unstructured input to structured storyboard data, making the boundaries of shot sequences more precise. The determination of shot size and angle parameters is closely related to the content of the scene, and the final output storyboard data is improved in terms of shot logic coherence and parameter standardization.
[0014] These or other aspects of this application will become more apparent in the following description of the embodiments. Attached Figure Description
[0015] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0016] Figure 1 A flowchart illustrating a generative animation storyboard construction method provided in this application embodiment; Figure 2 This is a schematic diagram illustrating a specific implementation of a generative animation storyboard construction method provided in this application. Figure 3 This is a schematic diagram of a generative animation storyboard construction system provided in an embodiment of this application. Detailed Implementation
[0017] In the field of automated generation of animation storyboards, existing technical solutions typically rely on large language models to directly parse the script text, outputting textualized storyboard tables. Some solutions further input textual descriptions into image generation models to obtain storyboard images. However, when dealing with scenes involving multiple character interactions, the models struggle to accurately depict the spatial relationships between characters, easily resulting in overlapping positions or perspective distortion. When describing continuous actions, the generated storyboard images often have visual flaws such as motion blur or limb distortion. Furthermore, the consistency of style between different shots in the same scene is difficult to guarantee, making it difficult for the output storyboard data to directly meet the standardized requirements of subsequent animation production.
[0018] To address the aforementioned issues, this application proposes a generative animation storyboard construction method. This method first performs semantic parsing of script data using a director's perspective knowledge database to generate visual data containing environmental and lighting information. Then, based on the visual data and orientation descriptions, it constructs scene scheduling data containing spatial coordinates, depth of field, character height, and character gaze direction, clarifying the spatial hierarchy of characters within the frame. Next, it extracts continuous motion information from action instructions, calculates motion amplitude values, and selects peak moments to generate key motion freeze-frame data. Finally, the visual data, scene scheduling data, and key motion freeze-frame data are input into a large language model. The large language model then divides the shot sequence according to shot splitting rules and determines shot size parameters, angle parameters, and scene description text, outputting structured storyboard data. This effectively solves the problems of ambiguous spatial hierarchy in storyboard data, distorted motion expression, and incoherent shot logic in existing technologies.
[0019] To enable those skilled in the art to better understand the present application, the present application will be further described in detail below with reference to the accompanying drawings and specific embodiments. Obviously, the described embodiments are merely some embodiments of the present application, and not all embodiments. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0020] A flowchart illustrating a generative animation storyboard construction method provided in this application is shown below. Figure 1 As shown, the method includes: Step 101: Obtain the script data, shot splitting rule data, and parameter determination rule data of the target animation. The script data includes dialogue text of multiple characters and action instructions for each character.
[0021] In step 101, the target animation refers to the animated work for which storyboard construction is to be performed. Script data, shot splitting rule data, and parameter determination rule data constitute the input objects of the method of this application; script data refers to the original text material used to describe the animated storyline, scene location, and character performance, wherein dialogue text records the language interaction content between multiple characters, action instructions describe the physical actions, orientation, position, or emotional state set for each character, and scene description information records the location, landmark objects, or spatial atmosphere where the plot takes place; shot splitting rule data is used to determine shot boundaries, and parameter determination rule data is used to determine shot size parameters and angle parameters; in optional embodiments, reference sample data can be used to train or calibrate a large language model.
[0022] In this embodiment, script data, shot splitting rule data, and parameter determination rule data of the target animation are first obtained from the user terminal or storage device. The script data contains dialogue text between multiple characters and action instructions set for each character. The shot splitting rule data and parameter determination rule data serve as the basis for subsequent shot splitting and parameter determination by the large language model. In an optional embodiment, reference sample data can also be obtained to train or calibrate the large language model.
[0023] Step 102: Based on a pre-set director's perspective knowledge database, perform semantic parsing on the script data to generate visual data for the target animation. The visual data includes environmental information and lighting information.
[0024] The director's perspective knowledge database is constructed by collecting film and television professional literature, animation director's creative notes, and classic animation storyboard examples. This database includes fields for scene description keywords, scene environment descriptions, emotion identifiers, action type identifiers, lighting direction, and lighting intensity. The specific data table structure is as follows: the primary key field is the record number (record_id); the scene description keyword field (scene_keyword) includes scene words such as forest, path, indoor, street, courtyard, and riverbank; and the scene environment description field (scene_description) is in natural language. The text includes eight categories for the emotion tag field (emotion_tag): joy, sadness, anger, fear, sarcasm, surprise, calm, and tension. The action type tag field (action_type_tag) includes eight categories for the action type: running, jumping, sitting, standing, waving, falling, turning, and climbing. The light direction field (light_direction) includes six categories for the light type: front lighting, side lighting, backlighting, side-backlighting, top lighting, and eye-level lighting. The light intensity field (light_intensity) includes four categories for the light intensity: strong light, medium light, natural soft light, and weak light.
[0025] The matching process employs the following priority rules: First, scene keywords are extracted from scene description information or directional carrier words in action instructions within the script data, and then matched against the scene environment description field according to the scene description keyword field. When there is no exact match for the scene keywords, semantic similarity calculation is used for approximate matching, with a semantic similarity threshold set at 0.75. If the similarity is below this threshold, a general scene environment description is returned. The emotion identifier field and action type identifier field are used to match or correct lighting direction information, lighting intensity information, and scene atmosphere parameters, but are not used alone to determine the scene location. When there are conflicts in the matching results, the matching result of the scene description information is used as the basis for determining the scene environment description information, and the matching result of the emotion information and action type information is used as the basis for determining the lighting parameters.
[0026] A complete data record example is as follows: record_id is DR001, scene_keyword is mountain path, scene_description is a description of the mountain path scene, including the winding dirt road, roadside weeds and wildflowers, dappled tree shadows and the faint outline of distant mountains, emotion_tag is mockery, action_type_tag is sit down, light_direction is eye-level light, and light_intensity is natural soft light.
[0027] Typical data examples include: when the scene description keyword is "street", the corresponding scene environment description field stores a description of a rainy day street scene, including a gloomy sky, wet cobblestone road, and closed wooden doors and windows; when the emotion identifier is "sad", it can be used to correct the atmosphere of the picture and the light intensity; when the action type identifier is "running", the corresponding light direction field can store side backlighting and the light intensity field can store strong light.
[0028] Visual data refers to a structured set of information used to describe the visual elements of a scene. This visual data includes environmental information and lighting information. Environmental information refers to the description of background environmental elements that match the scene description information. This environmental information includes scene location, atmosphere characteristics, vegetation type, architectural style, etc. Lighting information refers to a set of parameters used to describe the angle and brightness of light. This lighting information includes lighting direction information and lighting intensity information.
[0029] In this embodiment, step 102 includes the following process: Step 1021: Parse the script data to obtain the scene description information carried by the script data, the emotional information carried by the dialogue text, and the action type information carried by the action instructions.
[0030] In step 1021, scene description information refers to the description of location, landmarks, and spatial atmosphere parsed from the location carrier words in the script scene annotations or action instructions; emotion information refers to the description of emotional tendencies parsed from the dialogue text; and action type information refers to the category of physical behavior parsed from the action instructions.
[0031] In this embodiment, the scene annotations and action instructions in the script data are first semantically parsed to extract scene description information by identifying location nouns, landmark nouns, and directional carrier words. Then, the dialogue text is semantically parsed to extract emotional information by identifying adjectives or interjections expressing emotional tendencies. Next, the action instructions are semantically parsed to extract action type information by identifying verb phrases in the action instructions. For example, the information of a mountain path and a roadside landmark is extracted from "mountain path" or "a large rock by the roadside," the emotion of appreciation is extracted from "it's so beautiful," and the action type of sitting down is extracted from "slumping down breathlessly." The parsed scene description information, emotional information, and action type information are used as input data for subsequent matching steps.
[0032] In practical applications, assuming the scene description information in the script data is "a mountain path with a large rock by the roadside", the scene keywords extracted after semantic parsing of this scene description information are "mountain path" and "large rock by the roadside"; assuming a dialogue text is "Little Hammer Star God: I'm tired already", the emotional information extracted after semantic parsing of this dialogue text is sarcasm; assuming an action instruction in the script data is "slumped on a large rock by the roadside, covered in sweat", the action type information extracted after semantic parsing of this action instruction is "sit down".
[0033] Step 1022: Match the scene environment description information corresponding to the scene description information from the director's perspective knowledge database, and use the matched scene environment description information as environment information.
[0034] In step 1022, the scene environment description information refers to the complete description text of the background environment elements that match the scene description information.
[0035] In this embodiment of the application, the scene description information parsed in step 1021 is used as the matching basis to query a preset director's perspective knowledge database. The director's perspective knowledge database stores the correspondence between scene description keywords and scene environment description information in advance. For example, "street" corresponds to street scene description, and "garden" corresponds to garden scene description. The corresponding scene environment description information is matched from the director's perspective knowledge database according to the parsed scene description information, and the scene environment description information is used as environment information.
[0036] In practical applications, assuming that the scene description information obtained from step 1021 is a mountain path and a large rock by the roadside, the scene environment description information corresponding to this scene description information is searched in the director's perspective knowledge database. The matched scene environment description information is the scene description of the mountain path, including details such as the winding dirt road, roadside weeds and wildflowers, and dappled tree shadows.
[0037] Step 1023: Based on the emotion information and the action type information, match the lighting direction information and lighting intensity information from the director's perspective knowledge database, and use the matched lighting direction information and lighting intensity information as the lighting and shadow information.
[0038] In step 1023, the illumination direction information refers to the description of the angle of light illumination, and the illumination intensity information refers to the quantitative description of the brightness of the light.
[0039] In this embodiment, the emotional information and action type information parsed in step 1021 are used as the matching basis to query a preset director's perspective knowledge database. This director's perspective knowledge database stores the correspondence between emotional information, action type information, lighting direction information, and lighting intensity information. For example, tension can correspond to side backlighting, and sitting down can correspond to soft light and medium lighting intensity. Based on the parsed emotional information and action type information, the corresponding lighting direction information and lighting intensity information are matched from the director's perspective knowledge database, and the lighting direction information and lighting intensity information are used as light and shadow information.
[0040] In practical applications, assuming that the emotional information obtained from step 1021 is sarcasm and the action type information is sitting down, the corresponding lighting direction information and lighting intensity information are searched in the director's perspective knowledge database. The matching lighting direction information is eye level light and the matching lighting intensity information is natural soft light.
[0041] Step 1024: Combine the environmental information and the light and shadow information into visual data.
[0042] In this embodiment of the application, the scene environment description information obtained by matching in step 1022 is used as environment information, and the lighting direction information and lighting intensity information obtained by matching in step 1023 are used as light and shadow information. The environment information and the light and shadow information are combined to form complete visual data. The visual data is used to describe the visual elements of the target animation in a specific scene.
[0043] In practical applications, the environmental information of the mountain path matched in step 1022 is combined with the light and shadow information of eye level light and natural soft light matched in step 1023 to form visual data containing environmental details and lighting parameters.
[0044] Through the above steps, this application transforms the scene descriptions in the script data into specific environmental information, and combines them with emotion and action descriptions to determine lighting information, providing a structured visual element foundation for subsequent scene scheduling and shot generation.
[0045] Step 103: Based on the environmental information in the visual data and the orientation description information extracted from the action instructions, construct scene scheduling data. The scene scheduling data includes the spatial coordinate information, depth information, character height information, and character gaze direction information of multiple characters in the scene.
[0046] Among them, the preset logical constraint rules refer to the set of rules set in advance to guide the spatial layout of the screen. The preset logical constraint rules include the correspondence between different depth layer numbers and depth blur parameters; the scene scheduling data refers to the set of structured information used to describe the spatial layout of the characters in the screen. The scene scheduling data includes the spatial coordinate information, depth information, character height information and character gaze direction information of multiple characters in the screen; the spatial coordinate information refers to the spatial coordinate point of the character in the screen; the depth information refers to the sharpness range parameter of the character in the depth direction of the screen.
[0047] Examples of the preset logical constraint rules are shown in Table 1: Table 1. Examples of Preset Logical Constraint Rules
[0048]
[0049] It should be noted that the "unit" mentioned in Table 1 above and in subsequent embodiments of this application refers to normalized coordinate units, that is, logical coordinate values after normalization based on the screen width and height. The lower left corner of the screen corresponds to the origin (0, 0), and the upper right corner of the screen corresponds to the coordinate point (maximum screen width, maximum screen height). When outputting scene scheduling data to the subsequent rendering engine or image generation model, it is necessary to convert the normalized coordinate units to pixel coordinates according to the target output resolution. The conversion process is as follows: pixel coordinate X equals the number of pixels of the output image width divided by the maximum screen width and then multiplied by the normalized coordinate X; pixel coordinate Y equals the number of pixels of the output image height divided by the maximum screen height and then multiplied by the normalized coordinate Y. The sharp radius value in the depth-of-field blurring parameters also uses normalized coordinate units, and the blurring intensity value is a dimensionless ratio with a value range of 0 to 1.0, where 0 represents complete sharpness and 1.0 represents the maximum blurring degree.
[0050] The above example is only one example of this application. In practical applications, the depth-of-field blurring parameters corresponding to different depth-of-field layer numbers can also be set according to requirements. This application does not limit this.
[0051] In this embodiment, step 103 includes the following process: Step 1031: Based on the environmental information in the visual data, extract the location description information corresponding to each character from the action instructions of each character in the script data.
[0052] In step 1031, the orientation description information refers to the set of words parsed from the action instructions that represent the relative positional relationship of the characters.
[0053] In this embodiment of the application, environmental information is first extracted from the visual data generated in step 102. This environmental information includes the spatial structural features of the scene. Then, based on this environmental information, the directional description information corresponding to each character is extracted from the action instructions of each character in the script data. This directional description information includes words such as front, back, left, right, above, and below that indicate relative spatial position. The extracted directional description information is used as the basis for determining the spatial coordinates of the character in subsequent steps.
[0054] In practical applications, suppose the action instructions for the character Gou Dan in the script data include "slump down on a big rock by the roadside" and the action instructions for the character Xiao Chui Xing Shen include "stand next to". The location description information extracted from the action instructions of these two characters are "on a big rock" and "next to".
[0055] Step 1032: Based on the orientation description information and combined with the scene spatial structure in the environmental information, determine the spatial coordinate information of each character in the screen, and use the spatial coordinate information as the position information.
[0056] The scene spatial structure refers to the set of elements extracted from environmental information to describe the geometric layout of the scene. This scene spatial structure includes the spatial positional relationships of fixed objects and landmark elements in the scene, such as the direction of road extension, the placement of stones, the distribution range of trees, and the outline of mountains. The scene spatial structure comes from further analysis of environmental information in visual data. By identifying the scene location type, landmark object name, and relative positional relationship between objects described in the environmental information, this information is integrated into structured spatial layout data that can be used for spatial coordinate calculation.
[0057] Spatial coordinate information refers to the two-dimensional coordinates of the character in the virtual coordinate space, which is a Cartesian coordinate system with the lower left corner of the screen as the origin.
[0058] Step 1032 may specifically include the following steps: A1: Extract the boundary range parameters and reference point coordinate parameters of the scene spatial structure from the environmental information.
[0059] In step A1, the boundary range parameter refers to the set of values used to limit the spatial range that the image can accommodate. The boundary range parameter includes the maximum image width and the maximum image height. The reference point coordinate parameter refers to the coordinate values of the reference point used as a position reference.
[0060] In this embodiment of the application, the boundary range parameters of the scene spatial structure and the reference point coordinate parameters are extracted from the environmental information contained in the visual data generated in step 102. The boundary range parameters are used to limit the maximum range in which a character can be placed in the picture, and the reference point coordinate parameters are used as a reference point to determine the relative position of the character.
[0061] In practical applications, assuming the scene described by the environmental information is a mountain path, the boundary range parameters extracted from this environmental information are: a maximum screen width of 100 units and a maximum screen height of 80 units. The reference point coordinate parameters are: the center point coordinates of the screen are (50, 40).
[0062] A2: Based on the boundary range parameters and the reference point coordinate parameters, a virtual coordinate space is constructed, which defines the spatial range in the screen that can accommodate the character.
[0063] In step A2, the virtual coordinate space refers to an abstract coordinate system used to represent the spatial range that can accommodate characters in the screen at the logical level. The virtual coordinate space establishes a two-dimensional Cartesian coordinate system with the lower left corner of the screen as the origin. The horizontal axis is the X-axis, which represents the horizontal extension range of the screen from left to right, and the vertical axis is the Y-axis, which represents the vertical extension range of the screen from bottom to top.
[0064] In this embodiment of the application, a virtual coordinate space is constructed based on the boundary range parameters and reference point coordinate parameters extracted in step A1 to characterize the space range that can accommodate the character in the screen. The virtual coordinate space establishes a Cartesian coordinate system with the lower left corner of the screen as the origin. The range of values for the horizontal coordinate is determined by the maximum value of the screen width, and the range of values for the vertical coordinate is determined by the maximum value of the screen height.
[0065] In practical applications, based on the maximum screen width of 100 units, the maximum screen height of 80 units, and the reference point coordinate parameters (50, 40), a virtual coordinate space is constructed with an abscissa range of 0 to 100 and a ordinate range of 0 to 80. The reference point is located at the coordinate point (50, 40) in this virtual coordinate space.
[0066] A3: Extract the relative positional relationship of the character with respect to the coordinate parameters of the reference point from the orientation description information of each character.
[0067] In step A3, the relative positional relationship refers to the spatial orientation relationship between the character and the reference point, which includes the offset direction and offset distance.
[0068] In this embodiment of the application, for each character, the relative positional relationship of the character with respect to the coordinate parameters of the reference point is analyzed from the orientation description information extracted in step 1031. Specifically, the offset direction is determined by analyzing the direction words in the orientation description information, and the offset distance is determined by analyzing the distance words in the orientation description information.
[0069] The mapping of directional descriptions to offset direction and distance follows these rules: Directional words "above," "above," and "top" are mapped to a positive Y-axis offset; directional words "below," "below," and "bottom" are mapped to a negative Y-axis offset; directional words "left," "left side," and "left" are mapped to a negative X-axis offset; directional words "right," "right side," "right," and "beside" are mapped to a positive X-axis offset; directional words "in front" and "ahead" are mapped to a negative Y-axis offset; and directional words "behind" and "behind" are mapped to a positive Y-axis offset. When the directional description does not contain explicit distance quantifiers, the default offset distance is calculated as 20% of the maximum screen width, i.e., a screen width of 100 units corresponds to a default offset distance of 20 units. When the directional description contains close-range words such as "adjacent" or "close to," the offset distance is calculated as 10% of the maximum screen width. When the directional description contains distant-range words such as "far away" or "distant," the offset distance is calculated as 40% of the maximum screen width.
[0070] In practical applications, assuming the location description of the character Gou Dan is on the big rock, and the reference point is the center point of the screen (50, 40), the relative position of the character with respect to the reference point is obtained by parsing that the offset direction is downward and the offset distance is 15 units; assuming the location description of the character Xiao Chui Xing Shen is next to, the relative position of the character with respect to the reference point is obtained by parsing that the offset direction is to the right and the offset distance is 20 units.
[0071] A4: Based on the relative positional relationship, determine the spatial coordinate point corresponding to each character in the virtual coordinate space, and use the spatial coordinate point of each character as the spatial coordinate information corresponding to the character.
[0072] In this embodiment of the application, based on the relative positional relationship of each character with respect to the reference point coordinate parameters obtained in step A3, the spatial coordinate point corresponding to each character is calculated in the virtual coordinate space constructed in step A2. Specifically, the reference point coordinate parameters are superimposed with the offset direction and offset distance in the relative positional relationship to obtain the spatial coordinate point of each character, and the spatial coordinate point is used as the spatial coordinate information corresponding to the character. The spatial coordinate point is represented in the form of coordinate pairs.
[0073] In practical applications, the reference point coordinates are (50, 40). The relative position of the character Gou Dan is that the offset direction is downward and the offset distance is 15 units. The calculated spatial coordinates of the character Gou Dan are (50, 25). The relative position of the character Xiao Chui Xing Shen is that the offset direction is to the right and the offset distance is 20 units. The calculated spatial coordinates of the character Xiao Chui Xing Shen are (70, 40).
[0074] It should be noted that the Y-coordinate value used in this embodiment to determine the order of characters in the depth direction is applicable to standard scene composition scenarios with a level view. That is, characters with smaller Y-coordinate values are closer to the observer, and characters with larger Y-coordinate values are farther away from the observer. For non-standard view scenarios such as overhead or under-the-head views, the Z-axis coordinate can be extended in the virtual coordinate space. By increasing the depth dimension value, the order of characters in the depth direction is determined by the Z-axis coordinate value. The smaller the Z-axis coordinate value, the closer the characters are to the observer, and the larger the Z-axis coordinate value, the farther away the characters are from the observer.
[0075] Step 1033: Determine the front-to-back order of each character in the depth direction of the screen based on the spatial coordinate information of each character.
[0076] In step 1033, the front-to-back order information refers to a value used to represent the order in which the characters are arranged in the depth direction of the screen. The smaller the value in the front-to-back order information, the closer the character is to the foreground of the screen, and the larger the value, the farther the character is from the foreground of the screen.
[0077] In this embodiment of the application, based on the ordinate value in the spatial coordinate information of each character determined in step A4, the front-back order information of each character in the depth direction of the screen is determined, wherein the character with the smaller ordinate value is closer to the foreground in the screen, and the character with the larger ordinate value is farther away from the foreground in the screen; the front-back order information is assigned to each character in ascending order of ordinate value.
[0078] In practical applications, the spatial coordinates of the character Gou Dan are (50, 25), and the spatial coordinates of the character Xiao Chui Xing Shen are (70, 40). Since the vertical coordinate value of Gou Dan (25) is less than the vertical coordinate value of Xiao Chui Xing Shen (40), the front-back order information of Gou Dan is determined to be 1, indicating that he is closer to the foreground, and the front-back order information of Xiao Chui Xing Shen is determined to be 2, indicating that he is farther away from the foreground.
[0079] Step 1034: Assign a corresponding depth layer number to each character based on the aforementioned sequence information.
[0080] In step 1034, the depth layer number refers to an identifier used to identify the layer in which the character is located in the depth direction of the image.
[0081] In this embodiment of the application, based on the order information of each character determined in step 1033, a corresponding depth layer number is assigned to each character in ascending order of the order information. The character with the smallest order information value is assigned the smallest depth layer number, and the character with the largest order information value is assigned the largest depth layer number.
[0082] In practical applications, based on the order information of character Gou Dan (1) and character Xiao Chui Xing Shen (2), the depth layer number is assigned to character Gou Dan as 1 and character Xiao Chui Xing Shen as 2.
[0083] Step 1035: Based on the depth blurring parameters defined for different depth layer numbers in the preset logical constraint rules, determine the depth range information corresponding to each character, and use the depth range information as the depth information.
[0084] In step 1035, the depth-of-field blur parameter refers to a numerical parameter used to describe the range of sharpness of the character. The depth-of-field blur parameter includes a sharpness radius value and a blur intensity value; the depth-of-field range information refers to the distance range parameter for the character to remain in sharp image in the picture.
[0085] In this embodiment of the application, the depth-of-field blurring parameters defined for different depth-of-field layer numbers in the preset logical constraint rules are obtained. The preset logical constraint rules pre-store the correspondence between depth-of-field layer numbers and depth-of-field blurring parameters. According to the depth-of-field layer number assigned to each character in step 1034, the depth-of-field blurring parameter corresponding to the depth-of-field layer number is found from the preset logical constraint rules. The depth-of-field blurring parameter is used as the depth-of-field range information corresponding to the character, and the depth-of-field range information is used as the depth information.
[0086] Among them, the sharpness radius value in the depth of field blurring parameters is used to control the distance range at which the character remains completely sharp in the image. This sharpness radius value is directly mapped to the bokeh radius parameter in the subsequent image generation model or rendering engine. The blur intensity value is used to control the degree of blur beyond the sharpness radius range. This blur intensity value is directly mapped to the Gaussian blur coefficient (sigma) in the rendering engine. The specific conversion process is as follows: sigma is obtained by multiplying the blur intensity value by the maximum blur radius in pixels, where the maximum blur radius in pixels is a preset rendering parameter with a default value of 10 pixels.
[0087] In practical applications, the preset logical constraint rules specify that the depth-of-field blurring parameters corresponding to depth-of-field layer number 1 are a sharp radius of 5 units and a blur intensity of 0, while the depth-of-field blurring parameters corresponding to depth-of-field layer number 2 are a sharp radius of 4 units and a blur intensity of 0.3. Based on the depth-of-field layer number 1 of the character Gou Dan, the depth-of-field range information for this character is determined to be a sharp radius of 5 units and a blur intensity of 0. Based on the depth-of-field layer number 2 of the character Xiao Chui Xing Shen, the depth-of-field range information for this character is determined to be a sharp radius of 4 units and a blur intensity of 0.3.
[0088] Step 1036: Determine the character height and line of sight for each character.
[0089] In this embodiment, the character height information is used to describe the vertical position and height ratio of the character in the screen, and can be determined based on the longitudinal coordinate value in the character spatial coordinate information and the ratio of the character's height to the screen height; the character gaze direction information is used to describe the direction in which the character's eyes or face are facing, and can be determined based on the speaking character in the dialogue text, the character being pointed to, and the orientation description information in the action command.
[0090] Step 1037: Combine the spatial coordinates, depth of field, height, and line of sight of each character into scene scheduling data.
[0091] In this embodiment of the application, the spatial coordinate information of each character determined in step A4 is used as position information, and the depth range information of each character determined in step 1035 is used as depth information. The position information, depth information, character height information and character line of sight information are combined according to the character to form complete scene scheduling data. The scene scheduling data is used to describe the spatial layout relationship of multiple characters in the screen.
[0092] In practical applications, the spatial coordinates (50, 25) of the character Gou Dan, the depth information "clarity radius of 5 units and blur intensity of 0", the character height information "the height of the main body of the character accounts for 45% of the height of the screen", and the character's gaze direction information "looking at Xiao Chui Xing Shen" are combined; the spatial coordinates (70, 40) of the character Xiao Chui Xing Shen, the depth information "clarity radius of 4 units and blur intensity of 0.3", the character height information, and the character's gaze direction information are combined to form scene scheduling data containing the spatial layout information of the two characters.
[0093] This application combines the environmental information in the visual data with the orientation description in the script data through the above steps to construct scene scheduling data that includes character spatial coordinates, depth of field, character height, and character gaze direction, providing a precise spatial layout basis for subsequent key action freezes and shot generation.
[0094] Step 104: Based on the scene scheduling data, extract continuous action information from the action instructions of each character to generate key action freeze data corresponding to each character. The key action freeze data records the posture description information of the character at a single instant.
[0095] Among them, continuous motion information refers to the complete process data used to describe the actions performed by a character over a period of time. This continuous motion information includes continuous action verbs and the corresponding action trajectories of the continuous action verbs. Key motion freeze data refers to the posture description information used to characterize the character at a specific moment during the action. This key motion freeze data includes the limb position information and facial expression information corresponding to the freeze posture.
[0096] The action template library refers to a pre-built database used to map natural language action descriptions into structured posture parameters. Each record in the action template library contains the following fields: action type identifier (action_type, such as running, sitting, jumping, waving, etc.), start pose parameters (start_pose, including the angle values of each major joint and the spatial displacement values of the limb ends), mid pose parameters (mid_pose, same as above), and end pose parameters (end_pose, same as above). The joint angle values cover 10 degrees of freedom, including the shoulder, elbow, hip, knee, and ankle joints. The limb displacement values include 6 dimensions of displacement, including the displacement of the end of the hands and the displacement of the trunk center of gravity. For example, in the template record with the action type identifier "sitting", the start pose parameters are a knee joint angle of 0 degrees and a trunk center of gravity displacement of 0 cm, the mid pose parameters are a knee joint angle of 30 degrees and a trunk center of gravity drop of 15 cm, and the end pose parameters are a knee joint angle of 90 degrees and a trunk center of gravity drop of 40 cm. Natural language action commands are first mapped to the corresponding action type identifier in the action template library through verb phrase recognition. Then, the starting posture parameters, intermediate posture parameters, and ending posture parameters are read from the template record corresponding to the action type identifier as joint angle and limb displacement data of each instant in the action trajectory.
[0097] In this embodiment, step 104 includes the following process:
[0098] Step 1041: Based on the position information of each character in the scene scheduling data, parse the action instructions to identify the continuous action verbs in the action instructions and the action trajectories corresponding to the continuous action verbs.
[0099] In step 1041, continuous action verbs refer to verbs used to describe a character performing continuous actions. These continuous action verbs include words that indicate the process of an action, such as running, jumping, waving, and sitting down. The action trajectory refers to the complete change process of the action corresponding to the continuous action verb on the time axis. The action trajectory includes the starting posture, the intermediate posture, and the ending posture of the action.
[0100] In this embodiment, the position information of each character in the scene scheduling data generated in step 103 is first obtained. This position information is used to determine the spatial position of the character in the screen. Then, for each character, the action instructions corresponding to the character are extracted from the script data. The action instructions are semantically parsed to identify the continuous action verbs in the action instructions and the action trajectories corresponding to the continuous action verbs. Specifically, the continuous action verbs are determined by identifying the verb phrases in the action instructions. The identified continuous action verbs are matched in a preset action template library. The starting posture parameters, intermediate posture parameters, and ending posture parameters are read from the matched template records as the action trajectory corresponding to the action. The action trajectory contains the complete process posture from the start to the end of the action.
[0101] Step 1042: Combine the continuous action verbs and the action trajectories into continuous action information.
[0102] In this embodiment of the application, the continuous action verbs identified in step 1041 and the corresponding action trajectories are combined to form the continuous action information corresponding to the character. The continuous action information fully records the verb type of the character's action and the entire process of the action's posture changes on the timeline.
[0103] Step 1043: Select the instantaneous posture corresponding to the maximum amplitude of the action from the action trajectory in the continuous action information as the freeze posture.
[0104] In step 1043, the frozen posture refers to the single instantaneous posture selected from the motion trajectory that is representative or has the largest range of motion.
[0105] The maximum amplitude of movement refers to the quantitative indicator of the peak value of limb extension or movement tension selected from the movement trajectory. This amplitude of movement is calculated by analyzing the change in joint angle and limb displacement distance at each instant in the movement trajectory.
[0106] In the specific calculation, the standardized values of joint angle change and limb displacement distance are weighted and summed to obtain the motion amplitude value of each instantaneous posture. The general formula for motion amplitude is: ,in Indicates the amplitude of the movement. This represents the standardized value of the change in joint angle. This represents the standardized value of limb displacement distance, with a joint angle weighting coefficient of 0.6 and a limb displacement weighting coefficient of 0.4.
[0107] In the numerical example, assuming that the change in joint angle of the initial posture of bending over and kneeling is 0 degrees and the limb displacement distance is 0 cm, the value of the movement amplitude is 0.6×0+0.4×0=0; in the intermediate posture of the body sinking, the knee joint bending angle increases by 30 degrees. The formula for calculating the change in joint angle after standardization is shown in Equation (1).
[0108] .
[0109] In formula (1): This represents the standardized value of the change in joint angle. This represents the maximum possible change in joint angle. The default value is 90 degrees. This default value is an empirically calibrated value based on statistics of human joint range of motion and is applicable to common limb movements such as sitting, bending over, and waving. Indicates the change in joint angle;
[0110] Substituting into equation (1) yields The formula for calculating the limb displacement distance after standardization is given as follows: the trunk descent distance increases by 15 cm. See formula (2).
[0111] .
[0112] In formula (2): This represents the standardized value of limb displacement distance. Indicates the maximum possible displacement distance of a limb. This indicates the distance of limb displacement. The default value is 30 centimeters. This default value is an empirically calibrated value based on statistics of the range of motion of the human torso and is applicable to action types involving torso displacement such as sitting, squatting, and bending over.
[0113] Substitution (2) The amplitude of the movement is 0.6×0.33+0.4×0.5=0.198+0.2=0.398. In the final posture of the buttocks touching the stone, the knee joint bending angle remains unchanged at 30 degrees. After standardization, the change in joint angle is 0, the torso descent distance is 0 cm, and the limb displacement distance is 0. Therefore, the amplitude of the movement is 0.6×0+0.4×0=0. The intermediate posture of the body sinking corresponding to the maximum amplitude of the movement of 0.398 is selected as the freeze posture.
[0114] In this embodiment of the application, the motion amplitude value of each instantaneous posture in the motion trajectory generated in step 1042 is analyzed from the motion trajectory of the continuous motion information. The motion amplitude value is used as a quantitative indicator to characterize the degree of extension of the character's limbs or the tension of the motion. The instantaneous posture corresponding to the maximum motion amplitude value is selected from the motion trajectory as the freeze posture. The freeze posture is the instantaneous state that is most visually impactful or most representative of the characteristics of the action during the action.
[0115] Step 1044: Combine the limb position information and facial expression information corresponding to the frozen posture into key action frozen data corresponding to the character.
[0116] In step 1044, limb position information refers to a set of parameters used to describe the posture of the character's limbs and torso in space, and facial expression information refers to a set of parameters used to describe the shape of the character's facial features.
[0117] In this embodiment of the application, the limb position information and facial expression information corresponding to the freeze-frame posture selected in step 1043 are combined to form the key motion freeze-frame data corresponding to the character. The key motion freeze-frame data is used to characterize the complete posture description of the character at a specific moment during the action.
[0118] Through the above steps, this application transforms the continuous action process in the character's action instructions into a representative single-moment freeze-frame posture, providing a clear and explicit basis for action expression in subsequent shot generation.
[0119] Step 105: Input the visual data, the scene scheduling data, and the key action freeze data into the pre-constructed large language model. The large language model divides the input data into shot sequences according to the shot splitting rule data, and determines the shot size parameters, angle parameters, and scene description text in each shot sequence to output the shot data.
[0120] The large language model adopts a generative pre-trained model based on the Transformer architecture. This model consists of an input layer, a multi-head self-attention layer, a feedforward neural network layer, a layer normalization module, a sequence partitioning layer, and an output layer stacked in sequence. The multi-head self-attention layer contains 12 attention heads for parallel calculation of the association weights between elements in the input data. The feedforward neural network layer consists of two fully connected sub-layers and uses Gaussian error linear units as activation functions. The sequence partitioning layer is a sub-structure added to the standard Transformer architecture specifically for segmenting shot boundaries based on the attribution probability distribution.
[0121] The specific structure of the sequence partitioning layer is as follows: The sequence partitioning layer takes the hidden vector corresponding to each token output by the multi-head self-attention layer as input, with an input dimension of 768. After passing through a fully connected layer, the 768-dimensional hidden vector is mapped to an output vector with the number of lens units. Then, after passing through the Softmax activation function, the probability distribution of the lens units corresponding to each token is obtained. The output dimension is equal to the preset maximum number of lens units, with a default value of 20. During the inference phase, the lens unit number with the highest probability value in the probability distribution of each token is taken as the lens assignment label of that token. The lens assignment labels of all tokens are traversed, and the splitting boundary is set at the position where the lens assignment labels of adjacent tokens change.
[0122] The training process of the model is as follows: First, paired sample data containing script input and standard storyboard output are collected. The training dataset contains a total of 5000 pairs of samples, of which the training set accounts for 80% (4000 pairs), the validation set accounts for 10% (500 pairs), and the test set accounts for 10% (500 pairs). The input fields of each pair of samples include script_text, action_instruction, environment_tag, role_position, and key_pose. The output fields include shot_id, shot_boundary, shot_size_label (shot size label, values are long shot, full shot, medium shot, close-up, extreme close-up), camera_angle_label (angle label, values are eye-level, top-down, low-angle, side-angle), and storyboard_text. The annotation specification is as follows: each sample is independently annotated by 3 annotators with animation director experience, and the final label is determined by majority vote.
[0123] After preprocessing the sample data, it is input into the model, and a multi-task loss function can be used to simultaneously constrain shot boundaries, shot classification, angle classification, and image description text generation. The above training process is a conventional technique for model implementation, and this application does not depend on a specific training set size, a specific optimizer, a specific hidden layer dimension, or a specific maximum number of shot units; the relevant values are only examples and can be adjusted according to model size, image ratio, or business needs.
[0124] It should be noted that the number of attention heads, the activation function type of the feedforward neural network layer, and the dimension of the hidden layer in the Transformer architecture described above can be adjusted according to the actual needs of the scenario. However, the model must include a sequence partitioning unit for outputting the probability distribution of shot attribution, and the shot must be partitioned according to the mechanism of setting the splitting boundary according to the change of adjacent attribution labels. Other equivalent neural network structures that meet the above core mechanism requirements can replace the Transformer architecture described above.
[0125] Shot splitting rules refer to the logical criteria defined from the corresponding rule data for determining how to divide a continuous storyline into multiple shots. These shot splitting rules include the following five types of triggering conditions and corresponding shot boundary rules: The first type is the dialogue character switching triggering condition; when the speaking character in the dialogue text changes, the shot boundary is set at the character switching position. The second type is the action peak triggering condition; when the action amplitude value in continuous action information reaches a peak and then drops below 50% of the peak, the shot boundary is set at the peak position. The third type is the spatial position change triggering condition; when the character's spatial coordinate information undergoes a displacement exceeding 30% of the maximum screen width, the shot boundary is set at the displacement position. The fourth type is the emotional transition triggering condition; when adjacent... When the emotional information category of a dialogue text segment changes, a shot boundary is set at the emotional transition point; the fifth category is time jump triggering conditions, when descriptive words indicating time span appear in the script data, a shot boundary is set at the time jump point; the priority of the above five triggering conditions from high to low is time jump, emotional transition, dialogue character switching, action peak, and spatial position change; when multiple triggering conditions are met simultaneously at the same position, only one segmentation boundary is set; the above explicit rules serve as the basis for generating shot boundary labels for training samples during the model training phase, and as post-processing constraints for the output results of the sequence segmentation layer during the model inference phase; when the segmentation boundary output by the model is inconsistent with the segmentation boundary of the explicit rules, the result of the explicit rules is used for correction;
[0126] A shot sequence refers to a structured sequence of multiple shot units arranged in the order of plot development; a shot size parameter refers to a parameter used to describe the field of view of a shot, including types such as long shot, full shot, medium shot, close-up, and extreme close-up; an angle parameter refers to a parameter used to describe the shooting angle of a shot, including types such as eye-level, overhead, low-angle, and side view; and a picture description text refers to natural language text used to describe the content of the shot.
[0127] The parameter determination rule refers to the logical criteria defined from the corresponding rule data for determining the shot size and angle parameters based on the image content. Specifically, the parameter determination rule includes the following mapping relationships: Shot size parameter determination rule: When the number of characters in a shot unit is equal to 1 and the height of that character occupies 30% to 60% of the image height, the shot size parameter is determined to be medium shot; when the number of characters is equal to 1 and the height of that character occupies more than 60% of the image height, the shot size parameter is determined to be close-up; when the number of characters is equal to 1 and the height of that character occupies more than 80% of the image height, the shot size parameter is determined to be extreme close-up; when the number of characters is greater than or equal to 2 and the lateral distance between each character exceeds 40% of the maximum image width, the shot size parameter is determined to be wide shot; when the number of characters is greater than or equal to 2 and the lateral distance between each character does not exceed 40% of the maximum image width, the shot size parameter is determined to be medium shot; when the shot unit does not contain any characters and only contains environmental information, the shot size parameter is determined to be long shot. The rules for determining the angle parameters are as follows: when the character's gaze direction information in the camera unit is looking straight ahead, the angle parameter is determined to be level; when the character's gaze direction information is two characters looking at each other, the angle parameter is determined to be side view; when the character's height information is located in the upper third of the screen, the angle parameter is determined to be upward view; when the character is located in the lower third of the screen, the angle parameter is determined to be downward view.
[0128] Storyboard data refers to a structured set of information used to describe the storyboard script of an animation. This storyboard data includes shot parameters, angle parameters, and scene description text corresponding to each of the multiple shot units. The shot units are arranged and combined according to the plot development order to form a standardized storyboard table that can be directly used in the subsequent animation production process.
[0129] In this embodiment, step 105 includes the following process, such as... Figure 2 As shown:
[0130] Step 1051: Input the visual data, the scene scheduling data, and the key motion freeze data into the large language model, and convert the visual data, the scene scheduling data, and the key motion freeze data into embedding vectors through the input layer of the large language model.
[0131] In step 1051, the input fragment refers to script statements, action instruction fragments, character scheduling fields, key pose fields, or combinations thereof; the embedding vector refers to the result of converting the input fragment into a numerical vector representation, which is used for subsequent computational processing by the large language model.
[0132] In this embodiment of the application, the visual data generated in step 102, the scene scheduling data generated in step 103, and the key action freeze data generated in step 104 are first input into the pre-constructed large language model according to the input segment sequence; through the input layer of the large language model, each input segment is converted into a corresponding embedding vector, so that the originally unstructured text data is converted into a structured numerical vector that can be calculated by the model.
[0133] In practical applications, assuming that the input fragments include environmental information such as a mountain path, lighting information such as eye level light and natural soft light, spatial coordinates of the character Gou Dan in the scene scheduling data such as (50, 25) and depth information such as a sharp radius of 5 units and a blur intensity of 0, and the frozen posture of the character Gou Dan in the key action freeze data such as a body sinking and arms hanging down, the input layer converts these input fragments into corresponding embedding vectors respectively.
[0134] Step 1052: Using the self-attention layer of the large language model, weights are assigned to the embedding vectors corresponding to each input segment based on the shot splitting rule data, to obtain the probability distribution of the attribution of each input segment to different shot units.
[0135] In step 1052, the attribution probability distribution refers to the set of probability values for each input segment belonging to each shot unit, and the sum of the probability values corresponding to the same input segment in the attribution probability distribution is 1.
[0136] In this embodiment of the application, the embedding vectors corresponding to each input segment generated in step 1051 are input to the self-attention layer of the large language model; the self-attention layer assigns weights to each input segment according to the shot splitting rule data. Specifically, the self-attention layer calculates the degree of correlation between each input segment and combines the logic of shot boundary division in the shot splitting rule data to generate a probability value for each input segment to belong to different shot units, and finally obtains the probability distribution of each input segment to belong to different shot units.
[0137] In practical applications, assuming the input segment sequence contains 5 input segments corresponding to the character Gou Dan's action description, the self-attention layer calculates based on the shot splitting rule data that the probability of the first 2 input segments belonging to shot unit 1 is 0.9 and the probability of belonging to shot unit 2 is 0.1, and the probability of the last 3 input segments belonging to shot unit 1 is 0.2 and the probability of belonging to shot unit 2 is 0.8.
[0138] Step 1053: Through the sequence segmentation layer of the large language model, determine the shot attribution label of each input segment according to the attribution probability distribution, and set the segmentation boundary when the shot attribution label of adjacent input segments changes.
[0139] Step 1053 may specifically include the following steps: B1: Obtain the shot attribution label corresponding to each input segment in the attribution probability distribution, wherein the shot attribution label indicates the shot unit number to which the input segment belongs.
[0140] In step B1, the shot attribution label is an identifier used to identify which shot unit the input segment belongs to, and the value of the shot attribution label is the shot unit number.
[0141] In this embodiment of the application, from the attribution probability distribution generated in step 1052, the lens unit number corresponding to the maximum attribution probability of each input segment is determined, and the lens unit number is used as the lens attribution label of the input segment; specifically, for each input segment, the probability values of its belonging to each lens unit are compared, and the lens unit number with the largest probability value is selected as the lens attribution label of the input segment.
[0142] In practical applications, assuming the probability of the first two input segments belonging to shot unit 1 is 0.9 and the probability of belonging to shot unit 2 is 0.1, then the shot unit 1 is the label for the first two input segments; the probability of the last three input segments belonging to shot unit 1 is 0.2 and the probability of belonging to shot unit 2 is 0.8, then the shot unit 2 is the label for the last three input segments.
[0143] B2: Traverse each input segment in the input segment sequence, and when the shot attribution label of two adjacent input segments changes, set a split boundary between the two adjacent input segments.
[0144] In step B2, the dividing boundary refers to the marker point used to identify the division position of the lens unit.
[0145] In this embodiment of the application, each input segment is traversed sequentially according to the order of the input segment sequence. For each pair of adjacent input segments, the lens attribution labels of the two input segments are compared. If the lens attribution labels of two adjacent input segments are different, a split boundary is set between the two input segments. The split boundary is used to identify the division position of the lens unit.
[0146] In practical applications, the five input segments in the input segment sequence are traversed. The first two input segments are both labeled with shot unit 1. The second and third input segments are adjacent input segments. The second input segment is labeled with shot unit 1, while the third input segment is labeled with shot unit 2. Since they are different, a split boundary is set between the second and third input segments.
[0147] B3: Divide the input segment sequence into multiple segments according to the segmentation boundary, and treat each segment as a shot unit.
[0148] In step B3, a segment refers to a subset of data consisting of consecutive input segments between adjacent split boundaries.
[0149] In this embodiment of the application, the input segment sequence is divided into multiple consecutive segments according to the segmentation boundary set in step B2. Each segment consists of input segments located between two adjacent segmentation boundaries. Each segment is treated as a shot unit, and each shot unit corresponds to a storyboard frame.
[0150] In practical applications, the input segment sequence is divided into two segments based on the splitting boundary set between the second and third input segments. The first segment contains the first two input segments, and the second segment contains the last three input segments. The first segment is designated as shot unit 1, and the second segment is designated as shot unit 2.
[0151] B4: Obtain the arrangement order of each of the lens units in the input segment sequence, and arrange the multiple lens units sequentially according to the arrangement order to obtain the lens sequence.
[0152] In this embodiment, the order of each shot unit in the input segment sequence is obtained. This order is determined by the order in which the shot units appear in the input segment sequence. Multiple shot units are arranged sequentially according to this order to form a complete shot sequence. The order of the shot units in this shot sequence is consistent with the time sequence of the plot development.
[0153] In practical applications, lens unit 1 is located at the beginning of the input segment sequence, and lens unit 2 is located at the end. Lens unit 1 and lens unit 2 are arranged in order to obtain a shot sequence containing two lens units.
[0154] Step 1054: For each shot unit in the shot sequence, determine the rule data according to the parameters, and determine the shot unit's framing parameters and angle parameters from the number of characters, character spatial coordinate information, character gaze direction information, and character height information corresponding to the shot unit.
[0155] In step 1054, the parameter determination rule refers to the logical criteria defined from the corresponding rule data for determining the shot parameters and angle parameters based on the content of the scene. The character's gaze direction refers to the direction information of the character's eyes looking at, and the character's height information refers to the vertical position information of the character in the scene.
[0156] In this embodiment of the application, for each shot unit in the shot sequence obtained in step 1053, features are extracted from the number of characters, spatial coordinate information of characters, direction of gaze of characters and height information of characters corresponding to the shot unit according to the parameter determination rule data, and the shot parameters and angle parameters corresponding to the shot unit are determined.
[0157] In practical applications, for shot unit 1, there is one character, whose spatial coordinates are (50, 25), whose gaze direction is looking straight ahead, and whose height is 25 units. Based on the parameter determination rule data, the shot unit's framing parameter is determined to be medium shot and its angle parameter is eye level. For shot unit 2, there are two characters, whose spatial coordinates are (50, 25) and (70, 40) respectively, whose gaze direction is looking at each other, and whose heights are 25 units and 40 units respectively. Based on the parameter determination rule data, the shot unit's framing parameter is determined to be full shot and its angle parameter is side view.
[0158] Step 1055: Through the output layer of the large language model, the environmental information, lighting information, character posture information and character spatial coordinate information corresponding to each shot unit are merged into the screen description text of the shot unit, and the shot parameters, angle parameters and screen description text of each shot unit are combined into storyboard data according to the division order of the shot units.
[0159] In this embodiment of the application, the result after processing in step 1054 is input to the output layer of the large language model; for each shot unit, the output layer merges the environmental information, lighting information, character posture information and character spatial coordinate information corresponding to the shot unit to form the screen description text of the shot unit; then the shot parameters, angle parameters and screen description text of each shot unit are combined according to the division order of the shot units to finally output complete structured storyboard data.
[0160] In practical applications, the output layer combines the environmental information corresponding to lens unit 1 (mountain path), the lighting information (eye-level light and natural soft light), the character posture information (body sinking and arms hanging down), and the character spatial coordinate information (50, 25) into a screen description text. This screen description text is then combined with the shot type parameter (medium shot) and the angle parameter (eye-level) to form the storyboard data for the first lens unit. Similarly, the storyboard data for the second lens unit is generated, and the data from the two lens units are combined in sequence to form the complete storyboard data.
[0161] Through the above steps, this application collaboratively inputs visual data, scene scheduling data, and key action freeze-frame data into a large language model. By utilizing the attention mechanism, sequence segmentation capability, and parameter determination capability of the large language model, it automatically completes the segmentation of shot sequences, the determination of shot size parameters and angle parameters, and the generation of scene description text, ultimately outputting structured storyboard data.
[0162] Figure 3 This is a schematic diagram of the structure of a generative animation storyboard construction system provided in an embodiment of this application, as shown below. Figure 3 As shown, the system includes: The acquisition module 31 is used to acquire script data, shot splitting rule data and parameter determination rule data of the target animation. The script data includes dialogue text of multiple characters and action instructions for each character.
[0163] The parsing module 32 is used to perform semantic parsing on the script data based on a preset director's perspective knowledge database, and generate visual data for the target animation, the visual data including environmental information and lighting information.
[0164] Module 33 is used to determine the character's spatial coordinates, depth of field, height, and line of sight based on the environmental information in the visual data and the orientation description information in the action instructions, and combine them into scene scheduling data.
[0165] The extraction module 34 is used to extract continuous motion information from motion instructions based on scene scheduling data, calculate motion amplitude values, and select the instantaneous posture at which the motion amplitude value reaches its peak to generate key motion freeze data.
[0166] The segmentation module 35 is used to input visual data, scene scheduling data, and key action freeze data into the pre-built large language model, so that the large language model determines the shot ownership probability distribution of the input segment based on the shot segmentation rule data, sets the segmentation boundary when the shot ownership label of adjacent input segments changes, and generates shot sequence, shot type parameters, angle parameters, and screen description text based on the parameter determination rule data, and outputs the shot segmentation data.
[0167] The generative animation storyboard construction system of this application embodiment is used to implement the aforementioned generative animation storyboard construction method. Therefore, the specific implementation of the generative animation storyboard construction system can be found in the embodiment section of the generative animation storyboard construction method above. The specific implementation can be referred to the description of the corresponding embodiment, and will not be repeated here.
[0168] This application also provides an electronic device, comprising: a memory for storing a computer program; and a processor for executing the computer program to implement the steps of any of the generative animation storyboard construction methods described above.
[0169] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of any of the generative animation storyboard construction methods described above.
[0170] In one exemplary embodiment, the aforementioned computer-readable storage medium may include, but is not limited to, various media capable of storing computer programs, such as USB flash drives, read-only memory, random access memory, portable hard drives, magnetic disks, or optical disks.
[0171] The embodiments of this application also provide a computer program product, which includes a computer program that, when executed by a processor, implements the steps in any of the above embodiments of the generative animation storyboard construction method.
[0172] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0173] The foregoing has provided a detailed description of a generative animation storyboard construction method, system, device, and storage medium provided in this application. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the embodiments above are merely for the purpose of helping to understand the method and its core ideas. It should be noted that those skilled in the art can make various improvements and modifications to this application without departing from its principles, and these improvements and modifications also fall within the protection scope of this application.
Claims
1. A generative animation storyboard construction method, characterized in that, include: Acquire script data, shot splitting rule data, and parameter determination rule data for the target animation. The script data includes dialogue text for multiple characters and action instructions for each character. Based on a pre-built director's perspective knowledge database, the script data is semantically parsed to generate the visual data of the target animation, which includes environmental information and lighting information. Based on the environmental information in the visual data and the directional description information in the action instructions, a virtual coordinate space is constructed to determine the spatial coordinate information of each character. The depth information is determined based on the correspondence between the depth layer number and the depth blur parameter, and combined with the character height information and the character's line of sight direction information to form scene scheduling data. Based on the scene scheduling data, continuous action information is extracted from the action instructions of each character, the action amplitude value is calculated, and the instantaneous posture at which the action amplitude value reaches its peak is selected to generate key action freeze data. The visual data, scene scheduling data, and key motion freeze data are input into a pre-built large language model. The large language model determines the shot attribution probability distribution of the input segments based on the shot splitting rule data, and sets the splitting boundary when the shot attribution label of adjacent input segments changes. Based on the parameter determination rule data, the model generates shot sequences, shot parameters, angle parameters, and scene description text, and outputs the shot breakdown data.
2. The method according to claim 1, characterized in that, The step of inputting the visual data, the scene scheduling data, and the key motion freeze data into a pre-built large language model includes: The visual data, the scene scheduling data, and the key action freeze-frame data are input into the large language model according to the input segment sequence. The input layer of the large language model converts each input segment into a corresponding embedding vector. By using the self-attention layer of the large language model, and based on the shot splitting rule data, weights are assigned to the embedding vectors corresponding to each input segment to obtain the probability distribution of the belonging to different shot units for each input segment. The sequence segmentation layer of the large language model determines the shot attribution label of each input segment based on the attribution probability distribution, and sets a segmentation boundary when the shot attribution labels of two adjacent input segments change. The input segment sequence is divided into multiple shot units according to the segmentation boundary, and the multiple shot units are combined into a shot sequence according to the original input order. Through the output layer of the large language model, the scene parameters, angle parameters, and image description text corresponding to each shot unit are output according to the parameters to determine the rule data. The shot parameters, angle parameters, and scene description text of each shot unit are combined into storyboard data according to the division order of the shot units.
3. The method according to claim 1, characterized in that, The step involves determining the spatial coordinates, depth of field, height, and gaze direction of each character based on the environmental information in the visual data and the orientation description information in the action commands, and combining these into scene scheduling data, including: Based on the environmental information in the visual data, the location description information corresponding to each character is extracted from the action instructions of each character in the script data; Based on the orientation description information and combined with the scene spatial structure in the environmental information, the spatial coordinate information of each character in the screen is determined, and the spatial coordinate information is used as the position information. Based on the spatial coordinate information of each character, determine the front-to-back order of each character in the depth direction of the screen; Based on the aforementioned sequence information, assign a corresponding depth layer number to each character; Based on the depth-of-field blurring parameters defined for different depth-of-field layer numbers in the preset logical constraint rules, the depth-of-field information corresponding to each character is determined. The preset logical constraint rules include the correspondence between depth-of-field layer numbers and depth-of-field blurring parameters. The character's height information is determined based on the character's spatial coordinates, and the character's line of sight is determined based on the character's interaction relationships and orientation descriptions. The spatial coordinates, depth of field, height, and line of sight of each character are combined into scene scheduling data.
4. The method according to claim 1, characterized in that, The pre-built director's perspective knowledge database performs semantic parsing on the script data to generate visual data for the target animation. This visual data includes environmental information and lighting information, including: The script data is parsed to obtain the scene description information carried by the script data, the emotional information carried by the dialogue text, and the action type information carried by the action instructions; Match the scene environment description information corresponding to the scene description information from the director's perspective knowledge database, and use the matched scene environment description information as the environment information; Based on the emotion information and the action type information, the lighting direction information and lighting intensity information are matched from the director's perspective knowledge database, and the matched lighting direction information and lighting intensity information are used as lighting and shadow information; The environmental information and the light and shadow information are combined into visual data.
5. The method according to claim 3, characterized in that, The determination of the spatial coordinates of each character in the scene based on the orientation description information and the scene spatial structure in the environmental information includes: Extract the boundary range parameters and reference point coordinate parameters of the scene spatial structure from the environmental information; Based on the boundary range parameters and the reference point coordinate parameters, a virtual coordinate space is constructed, which defines the spatial range in the screen that can accommodate the character. The relative positional relationship of the character with respect to the coordinate parameters of the reference point is parsed from the orientation description information of each character. Based on the relative positional relationship, the spatial coordinate point corresponding to each character is determined in the virtual coordinate space, and the spatial coordinate point of each character is used as the spatial coordinate information corresponding to the character.
6. The method according to claim 2, characterized in that, The step of determining the shot attribution label for each input segment based on the attribution probability distribution, and setting a segmentation boundary when the shot attribution labels of two adjacent input segments change, includes: Obtain the shot attribution label corresponding to each input segment in the attribution probability distribution, wherein the shot attribution label indicates the shot unit number to which the input segment belongs; Traverse each input segment in the input segment sequence, and when the shot attribution label of two adjacent input segments changes, set a split boundary between the two adjacent input segments; The input segment sequence is divided into multiple segments according to the segmentation boundary, and each segment is treated as a shot unit. Obtain the arrangement order of each of the lens units in the input segment sequence, and arrange the multiple lens units sequentially according to the arrangement order to obtain the lens sequence.
7. The method according to claim 1, characterized in that, The step of extracting continuous motion information from the action commands of each character based on the scene scheduling data, calculating the motion amplitude value, and selecting the instantaneous posture at which the motion amplitude value reaches its peak to generate key motion freeze data includes: Based on the position information of each character in the scene scheduling data, the action instructions are parsed to identify the continuous action verbs in the action instructions and the action trajectories corresponding to the continuous action verbs; The continuous action verbs and the action trajectories are combined into continuous action information; The amplitude of the movement is calculated based on the changes in joint angles and the distance of limb displacement at each instant in the movement trajectory. Select the instant when the amplitude of the motion reaches its peak from the motion trajectory as the freeze-frame posture; The limb position information and facial expression information corresponding to the frozen posture are combined into key motion frozen data corresponding to the character.
8. A generative animation storyboard construction system, characterized in that, include: The acquisition module is used to acquire script data, shot splitting rule data and parameter determination rule data of the target animation. The script data includes dialogue text of multiple characters and action instructions for each character. The parsing module is used to perform semantic parsing on the script data based on a pre-set director's perspective knowledge database, and generate visual data for the target animation, including environmental information and lighting information. The construction module is used to construct a virtual coordinate space based on the environmental information in the visual data and the orientation description information in the action instructions, determine the spatial coordinate information of each character, determine the depth information based on the correspondence between the depth layer number and the depth blur parameter, and combine the character height information and the character's line of sight information to form scene scheduling data. The extraction module is used to extract continuous action information from the action instructions of each character based on the scene scheduling data, calculate the action amplitude value, and select the instantaneous posture when the action amplitude value reaches its peak to generate key action freeze data. The segmentation module is used to input the visual data, the scene scheduling data, and the key action freeze data into a pre-built large language model, so that the large language model determines the shot attribution probability distribution of the input segment according to the shot segmentation rule data, sets the segmentation boundary when the shot attribution label of adjacent input segments changes, and generates shot sequence, shot type parameters, angle parameters, and screen description text according to the parameter determination rule data, and outputs the shot segmentation data.
9. An electronic device, characterized in that, include: Memory, used to store computer programs; A processor, configured to implement the steps of a generative animation storyboard construction method as described in any one of claims 1 to 7 when executing the computer program.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, enables a generative animation storyboard construction method as described in any one of claims 1 to 7.