Method, device, equipment, storage medium and program product for generating split view video
Patent Information
- Application Number
- CN202611200381.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-07
- Publication Date
- 2026-09-18
AI Technical Summary
[0011] The method, apparatus, electronic device, computer-readable storage medium, and computer program product for generating storyboard videos provided in this disclosure firstly generate object constraints based on the participating objects and the constraint relationships between them recorded in the scene description information; then, using the object constraints, generate a set of placement information for the participating objects in the storyboard description information split from the scene description information; next, update the storyboard description information using the placement information set corresponding to the storyboard description information to generate updated storyboard description information; finally, call a video generation model to process the updated storyboard description information and generate the storyboard video.
Smart Images

Figure CN122783709A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the fields of artificial intelligence technology, such as information processing, content generation, and intelligent directing, and in particular to methods, apparatuses, electronic devices, computer-readable storage media, and computer program products for generating storyboard videos. Background Technology
[0002] With the development of AI-generated content technology, text generation, image generation, and video generation models have been gradually applied to scenarios such as AI short dramas, advertising creativity, virtual character interaction, game previews, and video material production.
[0003] For example, in short drama generation scenarios, upstream large-scale models can typically generate relatively complete scene descriptions and shot descriptions based on user-inputted themes, plot settings, character settings, or creative needs. Then, subsequent video generation models process the shot descriptions to generate the corresponding storyboard videos.
[0004] Therefore, considering that the quality of generated storyboard videos is closely related to the quality of storyboard description information, it is worthwhile and urgent to focus on how to improve the quality of storyboard description information and how to reduce the cost for users to configure storyboard description information so that users can complete the generation of video and other content with lower operating costs. Summary of the Invention
[0005] This disclosure provides a method, apparatus, electronic device, computer-readable storage medium, and computer program product for generating storyboard videos.
[0006] In a first aspect, embodiments of this disclosure propose a method for generating storyboard videos, comprising: generating object constraints based on the participating objects recorded in the scene description information and the constraint relationships between the participating objects; using the object constraints to generate a set of placement information of the participating objects in the storyboard description information split from the scene description information; updating the storyboard description information using the placement information set corresponding to the storyboard description information to generate updated storyboard description information; and calling a video generation model to process the updated storyboard description information to generate a storyboard video.
[0007] Secondly, embodiments of this disclosure propose an apparatus for generating storyboard videos, comprising: an object constraint generation unit configured to generate object constraints based on the participating objects recorded in the scene description information and the constraint relationships between the participating objects; a placement information generation unit configured to use the object constraints to generate a set of placement information for the participating objects in the storyboard description information split from the scene description information; a description information updating unit configured to update the storyboard description information using the placement information set corresponding to the storyboard description information to generate updated storyboard description information; and a storyboard video generation unit configured to call a video generation model to process the updated storyboard description information and generate a storyboard video.
[0008] Thirdly, embodiments of this disclosure provide an electronic device comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to implement the method for generating storyboard video as described in any implementation of the first aspect.
[0009] Fourthly, embodiments of this disclosure provide a non-transitory computer-readable storage medium storing computer instructions that enable a computer, when executed, to implement a method for generating storyboard video as described in any implementation of the first aspect.
[0010] Fifthly, embodiments of this disclosure provide a computer program product including a computer program that, when executed by a processor, can implement a method for generating storyboard videos as described in any implementation of the first aspect.
[0011] The method, apparatus, electronic device, computer-readable storage medium, and computer program product for generating storyboard videos provided in this disclosure firstly generate object constraints based on the participating objects and the constraint relationships between them recorded in the scene description information; then, using the object constraints, generate a set of placement information for the participating objects in the storyboard description information split from the scene description information; next, update the storyboard description information using the placement information set corresponding to the storyboard description information to generate updated storyboard description information; finally, call a video generation model to process the updated storyboard description information and generate the storyboard video.
[0012] This disclosure enables the construction of reusable object constraints across storyboards using scene and scene description information. When subsequently planning the coordinates of participating objects in each storyboard, these object constraints uniformly define the positional relationships, coordinate ranges, and spatial layouts of the objects. This ensures consistency in the coordinate system and positional logic across different storyboard descriptions, preventing spatial relationship shifts or discontinuities caused by independently determining coordinates in individual storyboard descriptions. Consequently, it improves the spatial consistency and overall quality of the generated storyboard videos.
[0013] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description
[0014] Other features, objects, and advantages of this disclosure will become more apparent from the following detailed description of non-limiting embodiments with reference to the accompanying drawings: Figure 1 This is an exemplary system architecture to which this disclosure can be applied; Figure 2 A flowchart illustrating a process for generating storyboard video, provided as an embodiment of this disclosure; Figure 3 A flowchart illustrating a process for generating action description information, provided in an embodiment of this disclosure; Figure 4 A flowchart illustrating another process for generating storyboard video provided in this embodiment of the disclosure; Figure 5 A flowchart illustrating the process of generating storyboard video in a specific application scenario, provided as an embodiment of this disclosure; Figure 6 A structural block diagram of an apparatus for generating storyboard videos provided in an embodiment of this disclosure; Figure 7 This is a schematic diagram of the structure of an electronic device suitable for performing a method for generating storyboard video, as provided in an embodiment of this disclosure. Detailed Implementation
[0015] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding; these should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description. It should be noted that, unless otherwise specified, the embodiments and features described in this disclosure can be combined with each other.
[0016] Furthermore, the acquisition, storage, use, processing, transportation, provision, and disclosure of any type of information involved in the technical solutions disclosed herein, such as user personal information (e.g., images containing facial objects as discussed later in this disclosure), comply with relevant laws and regulations and do not violate public order and good morals.
[0017] Figure 1 An exemplary system architecture 100 is shown, in which embodiments of methods, apparatuses, electronic devices, and computer-readable storage media for generating storyboard videos can be applied.
[0018] like Figure 1 As shown, system architecture 100 may include terminal devices 101, 102, and 103, a network 104, and a server 105. Network 104 serves as the medium for providing communication links between terminal devices 101, 102, and 103 and server 105. Network 104 may include various connection types, such as wired or wireless communication links, or fiber optic cables, etc.
[0019] Users can use terminal devices 101, 102, and 103 to interact with server 105 via network 104 to receive or send messages, etc. Various applications for enabling information communication between the terminal devices 101, 102, and 103 and server 105 can be installed. These applications include video generation applications, storyboard description information generation applications, and instant messaging applications.
[0020] Terminal devices 101, 102, and 103 and server 105 can be either hardware or software. When terminal devices 101, 102, and 103 are hardware, they can be various electronic devices with displays, including but not limited to smartphones, tablets, laptops, and desktop computers. When terminal devices 101, 102, and 103 are software, they can be installed in the aforementioned electronic devices, and can be implemented as multiple software programs or software modules, or as a single software program or software module; no specific limitation is made here. When server 105 is hardware, it can be implemented as a distributed server cluster composed of multiple servers, or as a single server. When server 105 is software, it can be implemented as multiple software programs or software modules, or as a single software program or software module; no specific limitation is made here.
[0021] Server 105 can provide various services through its built-in applications. For example, a video generation application can generate a segmented video consisting of at least one storyboard video based on user-provided segmented description information. When running this application, server 105 can achieve the following: First, it receives segmented description information from terminal devices 101, 102, and 103 via network 104. Then, server 105 generates object constraints based on the participating objects and their constraints recorded in the segmented description information. Next, server 105 uses these object constraints to generate a set of placement information for the participating objects in the storyboard description information. Then, server 105 updates the storyboard description information using the placement information set corresponding to the storyboard description information, generating updated storyboard description information. Finally, server 105 calls a video generation model to process the updated storyboard description information and generate the storyboard video.
[0022] It should be noted that, in addition to being obtained from terminal devices 101, 102, and 103 via network 104, the scene description information can also be pre-stored locally on server 105 through various means. Therefore, when server 105 detects that this data is already stored locally (e.g., when starting to process previously stored video generation tasks), it can choose to directly retrieve this data from locally. In this case, the exemplary system architecture 100 may also exclude terminal devices 101, 102, and 103 and network 104.
[0023] Since parsing scene description information to generate object relationships may require significant computing resources and capabilities, the methods for generating storyboard videos provided in subsequent embodiments of this disclosure are generally executed by a server 105 with strong computing power and abundant computing resources. Correspondingly, the device for generating storyboard videos is also generally located within the server 105. However, it should also be noted that when terminal devices 101, 102, and 103 also possess sufficient computing power and resources, they can also perform the aforementioned calculations performed by the server 105 through video generation applications installed on them, thereby outputting the same results as the server 105. Especially when multiple terminal devices with different computing capabilities exist simultaneously, if the video generation application determines that the terminal device it is using has strong computing power and sufficient remaining computing resources, it can allow the terminal device to perform the aforementioned calculations, thereby appropriately reducing the computing load on the server 105. Accordingly, the device for generating storyboard videos can also be located within terminal devices 101, 102, and 103. In this case, the exemplary system architecture 100 may also exclude the server 105 and the network 104.
[0024] It should be understood that Figure 1The number of terminal devices, networks, and servers shown is merely illustrative. Depending on implementation needs, any number of terminal devices, networks, and servers can be included.
[0025] Please refer to Figure 2 , Figure 2 A flowchart of a process for generating storyboard video provided in an embodiment of this disclosure includes process 200.
[0026] Process 200 specifically includes the following steps: Step 201: Generate object constraints based on the participating objects and the constraint relationships between them recorded in the scene description information; In embodiments of this disclosure, this step is intended to be performed by the entity executing the method for generating storyboard video (e.g., Figure 1 The server 105 shown first obtains the screen description information, for example, the screen description information provided by the terminal devices 101, 102, and 103 mentioned above.
[0027] Scene description information describes the content that should be included in a scene. Typically, a scene can consist of one or more shots within the same scene. In other words, a scene can be understood as a general description of a plot unfolding in a specific scene, while a shot can correspond to a sub-plot segment or shot unit within that scene or scene. In some scenarios, a scene can also be described as a sub-scene, setting, or sequence of scenes, which can be video content units composed of at least one shot.
[0028] In practice, the segment description information can be language or instruction information provided by the user in natural language form, or it can be structured information obtained by the upstream model based on the user's natural language instructions, as discussed above.
[0029] For example, the scene description information can be recorded in units of individual scenes, so that the subsequent execution entity can generate the corresponding scene video by providing these scene description information to the video generation model.
[0030] Correspondingly, storyboard description information can also be such "structured information".
[0031] In some embodiments, the fields used in the structured information of the storyboard description information can correspond to the storyboard identifier, the sequence information of the storyboard in the scene, the type of shot used in the storyboard, the camera movement method of the storyboard, the camera position of the shot in the storyboard, the participating objects (e.g., human objects, building objects, prop objects, etc.), the state of the participating objects (e.g., appearance color, facial expressions, posture), the duration of the storyboard, the scene (content) that the storyboard is expected to describe, etc. This allows the subsequent video generation model to accurately understand the content to be generated and its parameters by extracting the corresponding fields from this structured information.
[0032] In some embodiments, if the scene description information is described in natural language, the executing entity can process the scene description information by calling a text analysis model (e.g., Large Language Mode, or LLM) to divide it into multiple scenes and output the scene description information in the form of formatted information for each scene, so as to form a formatted scene description information composed of multiple scene description information.
[0033] It should be noted that, as discussed above, the scene description information can be obtained directly from the local storage device by the aforementioned execution entity, or it can be obtained from a non-local storage device (e.g., Figure 1 The information can be obtained from the terminal devices 101, 102, and 103 shown. The local storage device can be a data storage module located within the execution entity, such as a server hard drive. In this case, the screen description information can be quickly read locally. The non-local storage device can also be any other electronic device configured to store data, such as user terminals. In this case, the execution entity can obtain the required screen description information by sending an acquisition command to the electronic device.
[0034] After obtaining the scene description information, the executing entity can read the participating objects and the constraint relationships between the participating objects recorded in the scene description information in this step.
[0035] Typically, the participating object can be, for example, a character, an animal, a scene prop, or any other person or object that appears in a scene or shot. Constraints, on the other hand, can be used to characterize the relationships between different participating objects, or between the same participating object in different shots within a scene.
[0036] For example, if the semantics of the scene description information records that character A runs from the stairwell entrance to the left side of the stairwell landing, then in this case, the "constraint relationship" can be that character A should be at the stairwell entrance in the first scene (that is, character A's position in the first scene is constrained by being adjacent to the "stairwell entrance" as a participating object).
[0037] In some embodiments, the executing entity can construct a constraint graph, using participating objects as nodes and the constraints between participating objects as edges connecting the nodes, to record the aforementioned constraints. Accordingly, this constraint graph can be understood as the aforementioned "object constraints." For example, these object constraints can be expressed as object constraints... G To express, and that G Specifically, it can be ,in, V This indicates the participating object as a node, and E It indicates the constraint relationship between participating objects.
[0038] In some embodiments, in addition to the characters and props actually displayed in the scene (e.g., a virtual scene) mentioned above, the participating objects can also be "shooting participants" such as "(virtual) lenses" or "(virtual) cameras" used to capture scene content. This allows the final video presentation effect to be considered simultaneously when determining the placement and related information of the participating objects based on object constraints, thereby avoiding unexpected position or orientation jumps due to unreasonable camera placement, and inter-shot image distortion caused by coordinate system confusion.
[0039] Similarly, "action" can also be included as a participating object, so that characters and roles can be represented in the object constraints by the "edge" with the participating object as "action". This allows the situation of the character performing the action to be taken into account when determining the placement and information later, thus avoiding screen errors caused by the character's movement (e.g., the character moving in the wrong direction, drifting, or conflicting movement logic in different scenes due to unreasonable placement).
[0040] Accordingly, the above V, E It can then be represented as follows: , ,in, C For participants of the character (role) type, R For participants that appear as props or set elements in the scene, K As a participant in the lens type, T for
[0041] The participants in "sports" type activities, and These can respectively represent the constraint relationships between characters, This indicates the constraints between characters and props / scenery in the setting. This indicates the constraint relationship between the characters and the camera. This indicates a continuity constraint relationship between two actions.
[0042] Step 202: Using object constraints, generate a set of placement information for the objects involved in the scene description information extracted from the scene description information; In the embodiments of this disclosure, after generating object constraints based on step 201 above, the executing entity can use these object constraints to determine the placement information of each participating object in the storyboard included in the scene description information, thereby forming a placement information set corresponding to the storyboard and its storyboard description information (i.e., a "set" formed by the placement information of each participating object under the storyboard). That is, the executing entity can determine the corresponding placement information set for each storyboard description information, so that the placement information set can be used subsequently to describe and constrain the placement information such as the position and orientation of the participating objects included in the storyboard.
[0043] Accordingly, because this set of placement information is generated based on object constraints, the placement information in this set will not cause conflicts between storyboards.
[0044] The placement information set can include placement information for each participating object (e.g., its position and orientation within the scene). In practice, this placement information can be understood as the position and orientation that each participating object should have in the storyboard, so that the subsequent video generation model is constrained by this placement information set. In other words, the scene and environment composition described by the placement information set is used to generate the video. For example, the executing entity can first determine the placement information of each participating object in the first storyboard, and then, through semantic understanding of the storyboard description information and the aforementioned constraints (e.g., if character A moves in the first storyboard so that its position changes from the stairwell entrance to the stairwell landing in the second storyboard, then character A in the second storyboard should be on the stairwell landing), understand whether the participating object has changed its position or orientation in the storyboard, and the specific details of the change. If such a change exists, the executing entity can determine the placement information that the participating object should have in the next storyboard based on this change.
[0045] For example, the executing entity can construct an equation that includes the participating objects associated with the above constraints, and then determine and assign "positions" to each participating object by solving the equations to form the placement information of the participating object.
[0046] For example, taking a storyboard as an example, for a specific character, the actor can determine the character's "starting position" and orientation (e.g., towards other characters they are communicating with) within the storyboard based on the character's positional relationship with other characters, the character's subsequent actions, and the character's position relative to other scenery and props in the scene. This starting position and orientation are then used as the character's placement information. Similarly, by constructing a system of equations, the positions and orientations of each participating object can be determined.
[0047] In some embodiments, when generating the corresponding placement information set for the first storyboard, the executing agent can establish a coordinate system to describe the position and orientation. For example, the executing agent can select a reference object in the first storyboard (e.g., a participating object that appears in all storyboards and does not move) to construct the coordinate system. Then, in each storyboard, the executing agent can assign corresponding "coordinates" to the position of each participating object based on the size of the reference object, the size of the scene that the storyboard can present, and the size of the participating object, to characterize, for example, the placement information of the position of each reference object included in the storyboard.
[0048] For example, the position in "Placement Information" can be represented as follows: {"Person A": [3.0, 0.0, 0.0], "Person B": [-1.5, 0.0, 0.0], "Stairwell Entrance": [3.5, 0.0, 0.0], "Staircase Landing": [-1.5, 0.0, 0.0], "Descending Stairs": [-3.0, 0.0, -1.0]}.
[0049] Step 203: Update the storyboard description information using the placement information set corresponding to the storyboard description information to generate updated storyboard description information; In the embodiments of this disclosure, after generating and determining the placement information set based on step 202 above, the executing entity can update the placement information (set) to the corresponding scene description information of each scene. This allows the subsequent video generation model to understand and be constrained by the position and orientation of each participating object in the scene through this placement information, preventing the video generation model from guessing and generating placement information for individual scenes based on the "interpretation" of the text information in the scene description information. That is, the video generation model can use the newly added coordinates, orientations, etc. in the updated scene description information to understand how to place each participating object, instead of guessing the position and orientation of each participating object based on semantic understanding such as "adjacent" or "left side," which would lead to conflicts between scenes due to differences in understanding of coordinate systems, distances, orientations, etc.
[0050] For example, in this step, after the execution entity has broken down the scene description information and determined the scene description information corresponding to each scene in the above steps, it uses the placement information set to write the position, orientation, changes in position, and changes in orientation of the participating objects in that scene into the scene description information in the form of coordinates and direction vectors. This forms the updated scene description information corresponding to that scene, and the scene description information is used to provide the placement information (set) of each participating object in that scene, so that the subsequent video generation model can determine the position and orientation of each participating object according to the constraints of the placement information. For example, the execution entity can directly write this placement information and the corresponding participating object into the scene description information, so that the scene description information can directly and specifically indicate the position and orientation (e.g., the coordinates and orientation) of the participating object in that scene.
[0051] It should be understood that if the above-mentioned storyboard description information already includes information such as coordinates and direction vectors, then in such a case, if it differs from the placement information, or in other words, if it cannot meet the requirements of the object constraints, then the executing entity can choose to update it by replacing these coordinates and direction vectors so that they can meet the global constraints between storyboards that belong to the scene.
[0052] In some embodiments, before such a replacement is expected to be performed, the executing entity may also choose to report the conflict and present it to the user, so that the user can choose whether to actually perform the replacement, or request that the storyboard description information of other parts of the object constraints that are not satisfied be adjusted according to the coordinates, direction vectors, etc. provided by the user, so that the user can dynamically and personally select the update strategy based on the actual situation, and ensure the quality of the update of the storyboard description information.
[0053] In some embodiments, if the participating object undergoes a change in at least one of its position or orientation, the executing entity may choose to add the position and orientation of the participating object in the first video frame of the segment, and its position and orientation in the last video frame of the segment, to the segment description information. This forms position and orientation constraints on a segment-by-segment basis, allowing the participating object to appear more coherently and smoothly across segments without "drifting."
[0054] Correspondingly, this approach also enables the coordinate system modeling of each scene as a whole, thus avoiding the positional confusion that would occur if the subsequent video generation model independently modeled the coordinate system when processing the scene description information of each scene separately.
[0055] Step 204: Call the video generation model to process and update the storyboard description information, and generate the storyboard video.
[0056] In the embodiments of this disclosure, after updating the storyboard description information, the executing entity can call the video generation model to process the updated storyboard description information in order to generate a storyboard video corresponding to the storyboard description information.
[0057] The video generation model can be a large, trained model, an AI model trained on massive amounts of data with a large parameter scale and strong generalization ability, enabling it to possess one or more capabilities such as text understanding, image understanding, video understanding, content generation, cross-modal conversion, and logical reasoning. Thus, the executing entity can invoke this video generation model based on prompts and scene description information, leveraging its understanding of the content and semantics contained in the scene description information to generate a scene video that matches the scene description information in content and semantics.
[0058] Accordingly, if there are multiple storyboards, the executing entity can similarly utilize the video generation model to process and update the description information of each storyboard separately. Then, after generating the required individual storyboard videos, they are combined into the corresponding scene videos.
[0059] The method for generating storyboard videos provided in this disclosure first generates object constraints based on the participating objects and the constraints between them recorded in the scene description information. Then, using these object constraints, a set of placement information for the participating objects in the scene description information is generated. Next, the scene description information is updated using the placement information set corresponding to the scene description information, generating updated scene description information. Finally, a video generation model is called to process the updated scene description information, generating the storyboard video. This approach utilizes scene description information for scenes and scenes to construct reusable object constraints across scenes. When subsequently planning the coordinates of participating objects in each scene, the positional relationships, coordinate ranges, and spatial layouts of the objects are uniformly limited based on these object constraints. This ensures consistency in the coordinate system and positional logic within the scene description information of different scenes, avoiding spatial relationship shifts or discontinuities caused by independently determining coordinates in a single scene description. Therefore, it improves the spatial consistency and overall quality of the generated storyboard videos.
[0060] In some embodiments, for a storyboard or storyboard description, when the placement information set is solved based on object constraints, at least two satisfactory solutions may be obtained. Alternatively, for a storyboard description, at least two placement information sets may be obtained that meet the requirements (e.g., these two placement information sets have different camera positions under the same character, scene prop placement positions and orientations). In such cases (i.e., the storyboard description is associated with at least two placement information sets), the executing entity can respond by generating a placement score for the placement information set relative to the storyboard description information based on a consistency constraint function.
[0061] The consistency constraint function may include one or at least two constraint parameters to evaluate the performance of each set of placement information in the constraint dimensions corresponding to these parameters. For example, the constraint parameters may include at least one of the following: a narrative axis consistency parameter characterizing the degree of conformity with the storyline of the act description information. A shot consistency parameter used to characterize the coherence of shots used in each shot within a scene. Anchor point consistency parameter, used to characterize the degree of consistency in the relative positional relationships between participating objects. Cross-scene coherence consistency parameter used to characterize the coherence of placement information sets with adjacent scenes. And the penalty parameter that causes overlapping participants. .
[0062] In some embodiments, if the consistency constraint function is constructed based on constraint parameters of at least two aspects, then in such cases, the consistency constraint parameters can be actually constructed in a weighted summation manner, so that the weight of attention to a specific constraint dimension in the scoring process can be flexibly adjusted by adjusting the weights in different scenarios.
[0063] For example, for episode description information x Its placement score It can be determined based on the following formula (1): (1) in, , , , , Then they can represent the weighting coefficients respectively.
[0064] Then, the executing entity can choose to use the target placement information set with the highest placement score to update the storyboard description information that is split from the scene description information, and generate updated storyboard description information.
[0065] Therefore, based on placement scores, multiple optional or available placement information sets can be quantitatively evaluated and compared. In the case of multiple candidate placement results, the placement information set with higher scores and better overall effect can be selected as the final placement result. This avoids the problem of unstable placement effect caused by determining the placement result based on a single rule or random method, and improves the quality of the generated placement information and results.
[0066] In some embodiments, when updating the storyboard description information, the executing entity may constrain or guide the video generation process by adding placement information for the "camera" (e.g., the camera's position, orientation, etc.).
[0067] Accordingly, when determining the placement information of a shooting object, such as a camera, the executing entity can determine its placement information based on the shooting object it is shooting, so as to avoid the shooting object being assigned to a position or orientation that is detached from the shooting object, thus affecting the quality of video generation.
[0068] Accordingly, in such a case, the executing entity can first read the reference position and reference orientation of the first participating object (e.g., the "speaking person") that has already been generated in the placement information set when generating the placement information set of the shooting participants.
[0069] Then, the executing entity can determine the placement information of the subjects being photographed by "inverse calculation" based on this reference position and reference vector. For example, taking the reference position of the first subject as... The reference orientation is (For example, it can be a direction vector) In the case of this, the executing entity can determine the distance between the shooting participant and the first participant using the following formulas (2) and (3). and based on this Determine the final location of the subjects in space: (2) (3) in, For rotational compensation, As weight, This is the distance coefficient.
[0070] As for "orientation", the executing entity can determine the orientation of the shooting subject based on the reverse of the reference orientation of the first participating subject (or, the "reverse" adjusted based on the storyboard description information).
[0071] This allows the executing entity to generate and optimize the placement information of the subjects being filmed by considering its actual use in subsequent video generation, shot composition, object presentation, or scene rendering, rather than solely based on static positional relationships or a single placement rule. This ensures that the determined placement information is more aligned with subsequent processing needs and actual application scenarios, improving the adaptability between the placement information of special subjects like those being filmed and the video generation process.
[0072] In some embodiments, as discussed above, in a scene or shot, there may be a motion participant that moves across shots. For example, the participant may be performing a continuous action, such as running, in two consecutive shots. In such cases, the executing entity can also coordinate the motion process, or motion flow, of the motion participant in each shot description based on the scene description information to ensure that the motion participant's actions are coherent between shots.
[0073] For example, in some embodiments, the executing entity may, alternatively or in place of step 204 above, further utilize the placement information set and the action description information used to indicate the actions that should be implemented in the storyboard and can be connected with other storyboards to update the storyboard description information to generate updated storyboard description information.
[0074] To better understand the process of generating action description information, you can also refer to... Figure 3 Let's have a discussion. Figure 3 A flowchart of a process for generating action description information provided in an embodiment of this disclosure includes process 300.
[0075] Process 300 specifically includes the following steps: Step 301: In response to determining the second participating object as a motion participating object based on the scene description information, determine the motion flow of the second participating object based on the scene description information; Specifically, as discussed above, there may be instances of movement of participating objects in scene-by-scene and shot-by-shot videos. Accordingly, such objects exhibiting "movement" can be described as motion participating objects. In such cases, the executing entity can determine whether a second participating object belonging to a motion participating object is included or exists by reading the specific description in the scene description information.
[0076] Next, if such a second participating object exists, the executing entity can respond to it and determine the action flow of the second participating object based on the scene description information.
[0077] In practice, motion flows can be determined based on a pre-defined set of usable actions, such as movement and stopping. In some embodiments, to more accurately describe the motion, these actions can be broken down into finer-grained steps, such as dividing movement into initiation and preparation for movement. This allows subsequent motion description information provided as motion flows and motion dimensions to more accurately describe and constrain the executed actions, enabling the video generation model to better understand the performed actions.
[0078] Step 302: Based on the action flow, determine the action stage of the scene description information extracted from the scene description information; Specifically, the executing entity can determine a series of actions to be executed based on the action flow, and assign each action to a corresponding scene, forming one or more action stages within these scenes. Then, the executing entity can map these stages, based on the "scene," to the scene description information derived from the scene description information. That is, it associates and maps the actions (sets) belonging to a scene to the scene description information of that scene.
[0079] In some embodiments, in this step, the executing entity may first select to extract the actions from the action flow and assign them to corresponding storyboards based on the scene in which they are executed (e.g., based on whether the distance or angle generated by their displacement exceeds the limit of a storyboard). For example, after assigning an action to a storyboard, the executing entity may continuously assign subsequent actions to that storyboard until the storyboard can no longer carry new actions.
[0080] For example, for the action flow of "character A pushes open the door and enters the stairwell", the executing entity can continuously assign the two actions of pushing open the door and moving in the action flow to the first storyboard, and then assign the action of "after entering the door and standing in the stairwell" that the current storyboard cannot continue to carry to the second storyboard.
[0081] After the allocation is completed, the executing entity can determine the actions assigned to each "storyboard" by taking "storyboard" as the unit, and then determine the actions belonging to the same storyboard by constructing a storyboard-action reverse lookup relationship, so as to bind the actions to the "storyboard".
[0082] Then, the executing entity can determine the action stage of the scene description information, which is split from the scene description information, based on actions belonging to the same target scene—that is, actions bound to the same scene. In some scenarios, this action stage can also be understood as action information that describes the complete state of the participating object, based on the combination of action sub-flows used to form the action flow and situations not recorded in the action flow, such as standing or not moving.
[0083] Therefore, by combining the position of the storyboard in the overall action flow and the action content it contains, the action stage of the storyboard description information corresponding to the target storyboard can be determined, avoiding misjudgments caused by judging solely based on the storyboard order or a single action keyword, improving the accuracy and stability of action stage determination, and enhancing the coherence of subsequent storyboard description information updates and action transitions between storyboards.
[0084] Step 303: Based on the action stage in which the storyboard description information is located, generate corresponding action description information for the storyboard description information.
[0085] Specifically, after determining the action phase corresponding to each scene, the executing entity can generate corresponding action description information based on the action phase involved in that scene within the scene description information. For example, the action description information may include the action performed, the result of the action (e.g., the starting point and ending point of movement), etc.
[0086] In some embodiments, the action description information may further describe the action performed by the participating object in more detail. For example, the "action" may be specifically indicated by describing the positional changes of the skeletal points of the character object.
[0087] In some embodiments, if a storyboard involves at least two actions, the executing agent may choose to respond by generating a first action descriptor for each action to describe the action, so as to provide a precise, fine-grained description and constraint for each action.
[0088] Furthermore, based on this, second action descriptor information can be generated to describe the first frame action in the first frame image and the last frame action in the last frame image of the storyboard corresponding to the storyboard description information.
[0089] Then, the executing entity obtains the action description information by combining the first action description sub-information and the second action description sub-information.
[0090] Therefore, while enabling the action description information to accurately and granularly describe the situation of each action, the second action description information can also be used to connect and constrain the action results between each scene, so as to ensure the continuity of action between scenes.
[0091] In some embodiments, for the second action description information, the executing entity can use the last frame action in the last frame image of the previous segment adjacent to the current segment as the first frame action in the first frame image of the current frame, thereby achieving better connection between segment actions, avoiding drifting of segmented video between segments, and improving the overall quality of segmented video.
[0092] In some embodiments, if the storyboard description information does not provide content at the "image frame" granularity—for example, if the storyboard description information provides content in the form of first parts and second parts—then the executing entity may also choose to generate the second action description information based on the sequence number and part number. For example, the action result in the content of the part with the largest sequence number can be equivalently used as the "last frame action" to generate the second action description information.
[0093] Subsequently, the executing entity can update the storyboard description information using the placement information set and action description information as described above (for example, by adding more specific action description information such as the action being performed and the result of the action to the storyboard description information), and generate updated storyboard description information, which will not be repeated here.
[0094] In some embodiments, the storyboard description information update process discussed above mainly involves refining the scene, composition, and other aspects of the storyboard video through textual descriptions. Given that existing video generation models typically also possess image understanding capabilities, the executing entity can further generate a panoramic scene image and use it as visual reference information for the video generation model. That is, the executing entity can generate a corresponding panoramic scene image based on the aforementioned set of placement information, reflecting or realizing the object placement relationships indicated by the set of placement information. This transforms the relatively abstract object positions, orientations, and relative relationships in the placement information set into more intuitive visual reference information, allowing the video generation model to reference the overall layout and object distribution of the target scene when subsequently invoked.
[0095] Furthermore, the executing entity can also use the panoramic view of the scene to input or constrain the video generation model, so as to reduce problems such as object position offset, inconsistent spatial relationships, and scene layout distortion during the video generation process, thereby improving the consistency between the video generation result and the placement information set, and enhancing the image stability and scene performance quality of the generated video.
[0096] Accordingly, during step 204, the executing entity may, as an alternative, first generate a scene panorama based on the updated storyboard description information, using the placement information set corresponding to the updated storyboard description information. For example, the executing entity may generate a scene panorama, such as an equidistant bar chart, based on the coordinates and orientation in the placement information set using 3D modeling.
[0097] For example, the executing entity can project the 3D position of each non-central reference object in the storyboard into a panoramic orientation, and divide it into quadrants such as front, right, back, and left according to azimuth angle. For example, for a participating object, the executing entity can use " yaw "" indicates the horizontal azimuth angle of the reference object relative to the center of the scene, using " pitch "" indicates the vertical pitch angle of the reference object relative to the center of the scene. yaw "and" pitch "It can be determined based on the following formulas (4) and (5) respectively: (4) (5) in, x Indicates the position in the left and right directions. y Indicates the position in the forward and backward directions. z Indicates the position in the vertical direction; atan2 It is an arctangent function with quadrant determination, used to distinguish between front, back, left, and right positions.
[0098] Formula (4) is used to determine the horizontal position of the participating object in the panoramic view of the scene. For example, yaw A value close to 0° indicates the participant is directly in front; close to 90° indicates the participant is on the right; close to 180° or -180° indicates the participant is behind; and close to -90° indicates the participant is on the left.
[0099] Formula (5) is used to determine the vertical position of the participating object in the panoramic view of the scene. sqrt ( x² + y² ) This represents the horizontal distance of the participating object from the center of the scene. If z The larger, pitch The higher up the z-axis, the smaller the z-axis. pitch The lower down.
[0100] Then, during the process of utilizing and calling video generation, the executing entity can further call the video generation model based on the reference of the scene panorama to process and update the storyboard description information and generate storyboard videos.
[0101] In some embodiments, to facilitate the retrieval of the scene panorama and prevent the (updated) storyboard description information from being located in the form of structured information, "text," or "code" due to the addition of such a scene panorama, the executing entity may choose to store the scene panorama at a target location after generating it, and generate an access link for the target location. Then, the executing entity can populate the updated storyboard description information with the access link, so that when the video generation model is subsequently called to process the updated storyboard description information and generate the storyboard video based on the reference of the scene panorama, the video generation model can retrieve the scene panorama through the access link and use it as a reference to process the updated storyboard description information and generate the storyboard video.
[0102] In some embodiments, in addition to providing a panoramic view of the scene, in order to better provide a reference for the video generation model, the executing entity can also segment the panoramic view of the scene according to the target viewpoint to generate a slice reference map. This allows the generated slice reference map to be used to constrain the video generation model more finely from the selected viewpoint.
[0103] Therefore, to provide a clearer discussion of the embodiments that provide these "images" in form, it is also possible to combine Figure 4 Let me explain. Figure 4 This disclosure provides a flowchart of another process for generating storyboard video, including process 400.
[0104] Process 400 specifically includes the following steps: Step 401: Generate object constraints based on the participating objects and the constraint relationships between them recorded in the scene description information; Step 402: Using object constraints, generate a set of placement information for the objects involved in the scene description information extracted from the scene description information; Step 403: Update the storyboard description information using the placement information set corresponding to the storyboard description information to generate updated storyboard description information; The above steps 401-403 and as follows Figure 2 The steps 201-203 shown are the same. For the same parts, please refer to the corresponding parts of the previous embodiment. They will not be repeated here.
[0105] Step 404: For the updated storyboard description information, generate a scene panorama using the placement information set corresponding to the updated storyboard description information; Specifically, as discussed above, for a specific storyboard and updated storyboard description information, the executing entity can use the corresponding placement information set of the updated storyboard description information (i.e. the placement information set used to generate the updated storyboard description information) to generate a scene panorama through, for example, 3D modeling, which will not be repeated here.
[0106] Step 405: Divide the panoramic scene into slices according to the target viewpoint and generate slice reference images; Specifically, as discussed above, the executing entity can segment the panoramic scene image according to a predetermined target perspective to generate a sliced reference image. For example, the executing entity can generate an 8-direction sliced reference image by segmenting it in 8 directions that are evenly divided into 360°. For example, the "8 directions" can be the front corresponding to 0°, the right front corresponding to 45° to the right of the front, the right corresponding to 90° clockwise from the front, and the subsequent 8 directions of right rear, rear, left rear, left, and left front every 45° interval.
[0107] In some embodiments, since the size of the panoramic view may change accordingly with the scene, in order to ensure that each slice reference image can provide sufficient and high-quality reference content, an angle threshold can also be set to constrain the "field of view" before the slice reference image is segmented.
[0108] Accordingly, if the field of view of the current target view is greater than or equal to the angle threshold, the executing entity can respond by updating the target view by compressing the lateral view of the target view to obtain the updated target view.
[0109] For example, the implementing entity can determine the lateral projection scaling factor based on the following formula (6). S and utilize this S Compress the lateral view of the target: (6) in, S This represents the horizontal projection scaling factor. This indicates the horizontal angle between the current sampling point and the direction of the slice center. d The compression parameter representing the Panini projection is usually a positive number. x' This represents the horizontal coordinate component after projection. z' This represents the projected forward depth component. and Used to convert angles into direction vectors.
[0110] Then, the executing entity can segment the panoramic view of the scene according to the updated target perspective to generate sliced reference images. This ensures the quality of the content in the sliced reference images by constraining the field of view of the target perspective.
[0111] Step 406: Call the video generation model to process and update the storyboard description information based on the scene panorama and slice reference images, and generate the storyboard video.
[0112] Specifically, as mentioned above, the executing entity can provide the scene panoramic view and slice reference image to the video generation model. For example, the scene panoramic view and slice reference image can be provided to the video generation model by obtaining a link, so that the video generation model can refer to the scene panoramic view and slice reference image to more clearly understand the placement relationship, spatial layout and screen composition of each participating object in the storyboard, so as to make more accurate and constrained use of the above information when generating storyboard video.
[0113] In this way, the video generation model no longer relies solely on inferring and parsing the natural language semantics of the storyboard descriptions to understand the video content. Instead, it can combine more intuitive visual reference information to generate corresponding storyboard videos. This not only reduces model comprehension bias and minimizes errors such as incorrect object placement, inconsistent layout, or misaligned composition, but also improves the efficiency of the video generation model, enhancing the accuracy, stability, and image quality of the generated storyboard videos.
[0114] Based on any of the above embodiments, the execution subject may, as an alternative or alternative, first determine and detect whether there are description conflicts (e.g., position conflicts and / or orientation conflicts) in different scenes or scene description information during the process of generating object constraints (e.g., executing step 201).
[0115] If there is a description conflict between the first and second scene description information extracted from the scene description information, the executing entity will respond by selecting the participating objects and the constraint relationships between the participating objects recorded in the scene description information to generate object constraints.
[0116] Therefore, the execution entity only generates and determines global constraints when there are conflicts between storyboards in the storyboard. This avoids unnecessary constraint generation and subsequent processing when there are no conflicts between the storyboard description information. This not only reduces computational overhead and processing resource consumption, but also improves the overall processing efficiency of the storyboard video generation process.
[0117] Based on any of the above embodiments, the updated storyboard description information provided by the executing entity, which includes placement information, may include composition information (or composition summary) for indicating the placement information of each participating object in the storyboard, so that the subsequent video generation model can "understand" the position and orientation of each participating object based on the composition information.
[0118] When performing composition analysis and determining composition information, the executing entity can set a reference object (e.g., a person) corresponding to the "composition camera and composition angle" based on the "camera position facing the subject" and use dot product to determine the angle type of the reference object facing the camera lens.
[0119] For example, the direction vector of the camera pointing to the participating object. v, The intermediate quantity used to distinguish the front, side, and back of the participating object. q, And the intermediate amount used to distinguish the left and right orientation of the camera. r It can be determined based on the following formulas (7), (8) and (9) respectively: (7) (8) (9) in, This refers to the location of the participating object in space. The position of the camera in space. Let this be the orientation vector of the participating object in space. Let be the rightward direction vector of the camera.
[0120] To enhance understanding, this disclosure also provides an implementation scheme for a specific application scenario. Please refer to the following for details. Figure 5 . Figure 5 A flowchart illustrating the process of generating storyboard video in a specific application scenario provided by embodiments of this disclosure includes process 500. For example, process 500 may be generated by server 105 (not in the above-described architecture 100) in the architecture 100. Figure 5 (As shown in the image below) it is implemented as the "executing entity".
[0121] In process 500, the “screen” corresponding to the screen description information 510 can be composed of 3 scenes, and the screen description information 510 can include scene description information 511, 512 and 513.
[0122] It should be understood that the number of scene descriptions mentioned above is a choice made for the convenience of the example and is not intended to impose any limit on the number.
[0123] Accordingly, for ease of understanding, the following examples will mainly focus on updating the storyboard description information 511 and generating the storyboard video 570 using the updated storyboard description information 511', and will not repeat the explanation of the process of updating the storyboard description information 512 and 513.
[0124] Next, server 105 can first execute S501 to generate object constraints 520 based on the participating objects and the constraint relationships between the participating objects recorded in the episode description information 510.
[0125] Then, server 105 can execute S502 to generate a set 530 of placement information for the objects participating in the storyboard description information using object constraints 520.
[0126] Furthermore, server 105 can also determine, based on the scene description information 510, whether it includes a second participating object that is a motion participant. If it does, server 105 can respond by executing S504 to determine the action flow 540 of the second participating object based on the scene description information 510.
[0127] For example, the action flow 540 may include a series of actions 1, 2, ..., N (where N is a positive integer).
[0128] It should be understood that the descriptions of S502 and S504 are not intended to restrict the order in which these two steps are executed. In different embodiments, S502 and S504 can be executed in any order, or in parallel or simultaneously.
[0129] Following S504, server 105 can continue executing S505 to determine the action stage of storyboard description information 511 based on action flow 540. For example, action stage 545, consisting only of "Action 1" and "Action 2" in action flow 540, can be attributed to the storyboard corresponding to storyboard description information 511. Accordingly, server 105 can bind the actions involved in action stage 545 to the storyboard corresponding to storyboard description information 511 by constructing a reverse lookup relationship.
[0130] Then, after S505, server 105 can further execute S506 to generate corresponding action description information 545 based on the action stage of the storyboard description information 511, so as to describe the specific action situation (e.g., changes in the body pose of the participants, changes in position, etc.).
[0131] Furthermore, after S502, the server 105 can also generate a scene panorama 550 using the placement information set 530 by executing S507, so that the scene panorama 550 can be used as a reference when calling the video generation model (e.g., the video generation model 560 mentioned below).
[0132] Next, server 105 can continue to execute S508 after S506 to use the placement information set 530 and action description information 545 corresponding to the storyboard description information 511 to update the storyboard description information 511 by writing the placement information set 530 and action description information 545 to the storyboard description information 511, and generate updated storyboard description information 511'.
[0133] Finally, after obtaining the updated storyboard description information 511' and the scene panorama 550, server 105 can choose to execute S509 to call video generation model 560 to process the updated storyboard description information 511' based on the reference of scene panorama 550, and generate the aforementioned storyboard video 570 to complete the generation process of storyboard video 570.
[0134] Further reference Figure 6 As an implementation of the methods shown in the above figures, this disclosure provides an embodiment of an apparatus for generating storyboard videos, which is similar to... Figure 2 Corresponding to the method embodiments shown, this device can be specifically applied to various electronic devices.
[0135] like Figure 6 As shown, the device 600 for generating storyboard videos in this embodiment may include: an object constraint generation unit 601, a placement information generation unit 602, a description information update unit 603, and a storyboard video generation unit 604. The object constraint generation unit 601 is configured to generate object constraints based on the participating objects recorded in the scene description information and the constraint relationships between them; the placement information generation unit 602 is configured to use the object constraints to generate a set of placement information for the participating objects in the storyboard description information derived from the scene description information; the description information update unit 603 is configured to update the storyboard description information using the placement information set corresponding to the storyboard description information, generating updated storyboard description information; and the storyboard video generation unit 604 is configured to call a video generation model to process the updated storyboard description information and generate the storyboard video.
[0136] In this embodiment, the specific processing and technical effects of the object constraint generation unit 601, placement information generation unit 602, description information update unit 603, and storyboard video generation unit 604 in the device 600 for generating storyboard videos can be found in the following references. Figure 2 The relevant descriptions of steps 201-204 in the corresponding embodiments will not be repeated here.
[0137] In some optional implementations of this embodiment, the description information update unit 603 includes: a placement score generation subunit, configured to generate a placement score for the placement information set relative to the scene description information based on a consistency constraint function in response to the association of at least two placement information sets with the scene description information; and a description information update subunit, configured to update the scene description information using the target placement information set with the highest placement score, thereby generating updated scene description information.
[0138] In some optional implementations of this embodiment, the device 600 further includes: a motion flow determination unit, configured to determine the motion flow of the second participating object based on the scene description information in response to determining the second participating object as a motion participating object based on the scene description information; an motion stage determination unit, configured to determine the motion stage of the storyboard description information extracted from the scene description information based on the motion flow; a motion description information generation unit, configured to generate corresponding motion description information for the storyboard description information based on the motion stage of the storyboard description information; and a description information update unit 603, further configured to update the storyboard description information using the placement information set and the motion description information to generate updated storyboard description information.
[0139] In some optional implementations of this embodiment, generating corresponding action description information based on the action stage in which the storyboard description information is located includes: in response to the action stage in which the storyboard description information is located including at least two actions, generating first action descriptor information for describing the action for each action; and generating second action descriptor information for describing the first frame action in the first frame image and the last frame action in the last frame image of the storyboard corresponding to the storyboard description information.
[0140] In some optional implementations of this embodiment, the action stage determination unit includes: a storyboard allocation subunit, configured to determine the storyboard associated with each action in the action flow based on the action flow; an action determination subunit, configured to determine the actions belonging to the same storyboard by constructing a storyboard-action reverse lookup relationship; and an action stage determination subunit, configured to determine the action stage of the storyboard description information corresponding to the target storyboard, which is split from the scene description information, based on the actions belonging to the same target storyboard.
[0141] In some optional implementations of this embodiment, the placement information of the shooting participants in the placement information set is determined based on the following method: reading the reference position and reference orientation of the first participating object that has been generated in the placement information set; and using object constraints, based on the reference position and reference orientation, generating the placement information of the shooting participants in the storyboard description information that is split from the scene description information.
[0142] In some optional implementations of this embodiment, the storyboard video generation unit 604 includes: a panoramic image generation subunit, configured to generate a scene panoramic image using a set of placement information corresponding to the updated storyboard description information for updating storyboard description information; and a storyboard video generation subunit, configured to call a video generation model to process the updated storyboard description information based on a reference to the scene panoramic image, and generate a storyboard video.
[0143] In some optional implementations of this embodiment, the apparatus 600 further includes: a panoramic image maintenance unit configured to store a scene panoramic image to a target location and generate an acquisition link for the target location; a panoramic image backfilling unit configured to backfill the acquisition link to update the storyboard description information; and a storyboard video generation subunit further configured to call a video generation model to acquire the scene panoramic image based on the acquisition link, and based on a reference to the scene panoramic image, process and update the storyboard description information to generate a storyboard video.
[0144] In some optional implementations of this embodiment, the apparatus 600 further includes: a slice reference image generation unit, configured to segment the scene panorama according to the target viewpoint and generate a slice reference image; the storyboard video generation subunit is further configured to call a video generation model to process and update the storyboard description information based on the reference of the scene panorama and the slice reference image, and generate a storyboard video.
[0145] In some optional implementations of this embodiment, the slice reference image generation unit includes: a viewpoint update subunit, configured to update the target viewpoint by compressing the lateral viewpoint of the target viewpoint in response to the field of view of the target viewpoint being greater than or equal to an angle threshold, thereby obtaining an updated target viewpoint; and a slice reference image generation subunit, configured to segment the scene panorama according to the updated target viewpoint to generate a slice reference image.
[0146] In some optional implementations of this embodiment, the object constraint generation unit 601 is further configured to generate object constraints based on the participating objects recorded in the scene description information and the constraint relationships between the participating objects, in response to a description conflict between the first scene description information and the second scene description information split from the scene description information.
[0147] This embodiment is a device embodiment corresponding to the method embodiment described above. The device for generating storyboard videos provided in this embodiment can construct reusable object constraints across storyboards using scene and scene description information. When subsequently planning the coordinates of participating objects in each scene, the device uniformly limits the positional relationships, coordinate ranges, and spatial layout of the objects based on these object constraints. This ensures the consistency of the coordinate system and positional logic in the storyboard description information of different scenes, avoiding spatial relationship shifts or discontinuities caused by independently determining coordinates in a single scene description. Therefore, it can improve the spatial consistency and overall quality of the generated storyboard videos.
[0148] According to embodiments of this disclosure, this disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0149] Figure 7 A schematic block diagram of an example electronic device 700 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0150] like Figure 7 As shown, device 700 includes a computing unit 701, which can perform various appropriate actions and processes based on a computer program stored in read-only memory (ROM) 702 or a computer program loaded into random access memory (RAM) 703 from storage unit 708. The RAM 703 may also store various programs and data required for the operation of device 700. The computing unit 701, ROM 702, and RAM 703 are interconnected via bus 704. Input / output (I / O) interface 705 is also connected to bus 704.
[0151] Multiple components in device 700 are connected to I / O interface 705, including: input unit 706, such as keyboard, mouse, etc.; output unit 707, such as various types of monitors, speakers, etc.; storage unit 708, such as disk, optical disk, etc.; and communication unit 709, such as network card, modem, wireless transceiver, etc. Communication unit 709 allows device 700 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0152] The computing unit 701 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 701 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 701 performs the various methods and processes described above, such as the method of generating storyboard video. For example, in some embodiments, the method of generating storyboard video may be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 708. In some embodiments, part or all of the computer program may be loaded and / or installed on device 700 via ROM 702 and / or communication unit 709. When the computer program is loaded into RAM 703 and executed by computing unit 701, one or more steps of the method of generating storyboard video described above may be performed. Alternatively, in other embodiments, the computing unit 701 may be configured to perform the method of generating storyboard video by any other suitable means (e.g., by means of firmware).
[0153] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0154] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0155] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0156] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0157] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.
[0158] Computer systems can include clients and servers. Clients and servers are generally geographically separated and typically interact via communication networks. The client-server relationship is established by computer programs running on the respective computers and having a client-server relationship with each other. Servers can be cloud servers, also known as cloud computing servers or cloud hosts, which are hosting products within the cloud computing service ecosystem to address the management difficulties and weak business scalability inherent in traditional physical hosts and Virtual Private Servers (VPS) services. Servers can also be categorized as distributed system servers or servers incorporating blockchain technology.
[0159] According to the technical solution of this disclosure, scene description information can be used to construct reusable object constraints across scenes. When subsequently planning the coordinates of participating objects in each scene, these object constraints uniformly define the positional relationships, coordinate ranges, and spatial layouts of the objects. This ensures consistency in the coordinate system and positional logic within the scene description information of different scenes, avoiding spatial relationship shifts or discontinuities caused by independently determining coordinates in a single scene description. Therefore, the spatial consistency and overall quality between generated scene videos can be improved.
[0160] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution provided in this disclosure can be achieved, and this is not limited herein.
[0161] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.
Claims
1. A method for generating storyboard video, comprising: Based on the participating objects and the constraint relationships between them recorded in the scene description information, object constraints are generated; Using the object constraints, a set of placement information for the participating objects in the scene description information extracted from the scene description information is generated; The storyboard description information is updated using the set of placement information corresponding to the storyboard description information, thereby generating updated storyboard description information; The video generation model is invoked to process the updated storyboard description information and generate a storyboard video.
2. The method according to claim 1, wherein, The step of updating the storyboard description information using the placement information set corresponding to the storyboard description information to generate updated storyboard description information includes: In response to the storyboard description information being associated with at least two placement information sets, a placement score is generated for the placement information sets relative to the storyboard description information based on a consistency constraint function; Using the set of target placement information with the highest placement score, the storyboard description information is updated to generate updated storyboard description information.
3. The method according to claim 1, further comprising: In response to determining that the second participating object is a motion participating object based on the scene description information, the motion flow of the second participating object is determined based on the scene description information; Based on the action flow, determine the action stage of the scene description information extracted from the scene description information; Based on the action stage in which the storyboard description information is located, corresponding action description information is generated for the storyboard description information. as well as The step of updating the storyboard description information using the placement information set corresponding to the storyboard description information to generate updated storyboard description information includes: The storyboard description information is updated using the placement information set and the action description information corresponding to the storyboard description information, thereby generating updated storyboard description information.
4. The method according to claim 3, wherein, Based on the action stage in which the storyboard description information is located, corresponding action description information is generated, including: In response to the action phase in which the storyboard description information is located including at least two actions, for each of the actions, a first action descriptor information for describing the action is generated; Generate second action descriptor information to describe the first frame action in the first frame image and the last frame action in the last frame image of the storyboard corresponding to the storyboard description information.
5. The method according to claim 3, wherein, The step of determining the action stage of the scene description information extracted from the scene description information based on the action flow includes: Based on the action flow, determine the storyboard associated with each action in the action flow; By constructing a shot-action reverse lookup relationship, the actions belonging to the same shot are determined; Based on the actions belonging to the same target storyboard, determine the action stage of the storyboard description information that is split from the scene description information corresponding to the target storyboard.
6. The method according to claim 1, wherein, The placement information of the subjects participating in the shooting in the placement information set is determined based on the following method: Read the reference position and reference orientation of the first participating object that has been generated in the placement information set; Using the object constraints, based on the reference position and the reference orientation, the placement information of the shooting participants in the shot description information split from the scene description information is generated.
7. The method according to claim 1, wherein, The process of calling the video generation model to process the updated storyboard description information and generate a storyboard video includes: For the updated storyboard description information, a scene panorama is generated using the placement information set corresponding to the updated storyboard description information; The video generation model is invoked based on the reference of the panoramic image of the scene to process the updated storyboard description information and generate a storyboard video.
8. The method according to claim 7, further comprising: The panoramic image of the scene is stored at the target location, and a link to access the target location is generated. The obtained link is then populated back into the updated storyboard description information; And by calling a video generation model based on a reference to the scene panorama, processing the updated storyboard description information, and generating a storyboard video, including: The video generation model is invoked to obtain the scene panorama based on the obtained link, and the scene panorama is referenced to process the updated storyboard description information to generate a storyboard video.
9. The method according to claim 7, further comprising: The panoramic view of the scene is segmented according to the target perspective to generate a sliced reference image; The video generation model, based on a reference to the scene panorama, processes the updated storyboard description information to generate a storyboard video, including: The video generation model is invoked to process the updated storyboard description information based on the panoramic view of the scene and the slice reference image, and to generate a storyboard video.
10. The method according to claim 9, wherein, The step of segmenting the panoramic scene image according to the target viewpoint to generate a sliced reference image includes: In response to the target view's field of view being greater than or equal to an angle threshold, the target view is updated by compressing the target view's lateral field of view to obtain an updated target view. The scene panorama is segmented according to the updated target perspective to generate a sliced reference image.
11. The method according to any one of claims 1-10, wherein, The generation of object constraints based on the participating objects and the constraint relationships between them recorded in the scene description information includes: In response to a description conflict between the first and second scene description information extracted from the scene description information, object constraints are generated based on the participating objects and the constraint relationships between them recorded in the scene description information.
12. An apparatus for generating storyboard video, comprising: The object constraint generation unit is configured to generate object constraints based on the participating objects and the constraint relationships between them recorded in the episode description information; The placement information generation unit is configured to use the object constraints to generate a set of placement information for the participating objects in the storyboard description information split from the scene description information; The description information update unit is configured to update the storyboard description information using the placement information set corresponding to the storyboard description information, and generate updated storyboard description information; The storyboard video generation unit is configured to call a video generation model to process the updated storyboard description information and generate a storyboard video.
13. An electronic device, comprising: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, which, when executed by the at least one processor, enables the at least one processor to perform the method for generating storyboard video according to any one of claims 1-11.
14. A non-transitory computer-readable storage medium storing computer instructions for causing the computer to perform the method for generating storyboard video according to any one of claims 1-11.
15. A computer program product comprising a computer program that, when executed by a processor, implements the method for generating storyboard video according to any one of claims 1-11.