Method and device for generating a movie from a script

By combining natural language processing and video understanding technologies to optimize the generation process, the problem of automatically generated videos failing to reflect the content of movie scripts was solved, enabling efficient and high-fidelity film production.

CN114697495BActive Publication Date: 2025-09-12TCL TECHNOLOGY GROUP CORPORATION
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202111098830.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-12-29
Filing Date
2021-09-18
Publication Date
2025-09-12
Estimated Expiration
2041-09-18

AI Technical Summary

Technical Problem

In the existing technology, automatically generated videos cannot fully reflect the content of the movie script, resulting in low film production efficiency and insufficient quality.

Method used

Combining natural language processing and video understanding techniques, the problem mapping generation process is optimized, dynamic programming is used to reduce computational complexity, evaluate the fidelity and aesthetic compliance of the video, and iteratively optimize camera settings and character performances to generate high-quality videos.

Benefits of technology

The efficiency of script-to-movie generation is significantly improved, ensuring the fidelity and aesthetic compliance of the generated video with the movie script and reducing computational complexity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114697495B_ABST
    Figure CN114697495B_ABST
Patent Text Reader

Abstract

A method and device for generating a movie from a script, comprising obtaining a movie script, generating a video according to the movie script, optimizing the generated video until conditions are met and outputting the generated video.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer graphics technology, and in particular to a method and device for generating a script into a movie. Background Art

[0002] In the traditional film industry, scriptwriting (i.e., creating a film script) and film production are completely separate processes. New technologies like Write-A-Movie (Write-A-Movie) automatically generate videos based on film scripts, significantly improving film production efficiency. However, these automatically generated videos often fail to fully reflect the content of the film script.

[0003] This paper presents a script-to-movie generation method that incorporates a novel evaluation mechanism that combines the visual comprehensibility of a movie script with its compliance with cinematography guidelines. The script-to-movie generation process is thus mapped into an optimization problem to improve the quality of the automatically generated video. Dynamic programming is also incorporated into the solution of this optimization problem to reduce computational complexity and shorten film production time. Summary of the Invention

[0004] One aspect of the present invention provides a method for generating a movie from a script on a computer device, comprising: obtaining a movie script, generating a video based on the movie script, optimizing the generated video until a passing condition is met; and outputting the optimized video.

[0005] Another aspect of the present invention provides a device for generating a movie from a script. The device includes a memory storing program instructions, and a processor coupled to the memory, the processor configured to execute the program instructions to: obtain a movie script, generate a video based on the movie script, optimize the generated video until a passing condition is met; and output the optimized video.

[0006] The device, wherein the processor is further configured to:

[0007] generating a first action list according to the movie script;

[0008] generating a stage performance according to the actions in the first action list; and

[0009] A video camera is used to capture video of the stage performance.

[0010] The device, wherein the processor is further configured to:

[0011] evaluating a total aesthetic distortion value of the video, the video being captured by a camera from the stage performance;

[0012] generating a second action list based on the video, the video being captured by a camera from the stage performance;

[0013] determining a fidelity error E between the first action list and the second action list; and

[0014] Iteratively optimize the camera settings and character performance to minimize the total aesthetic distortion D, thereby satisfying the passing condition, wherein the passing condition includes satisfying that the fidelity error E is less than or equal to a preset fidelity error threshold Th E Or minimize the number of iterations until the count reaches a preset count threshold.

[0015] The device, wherein:

[0016] Each action in the first action list and the second action list has attributes, and the attributes include subject, action, object, action duration, subject start position, subject end position, subject emotion and action style.

[0017] The device, wherein:

[0018] The first action list is represented by a time-ordered action list {a i |i=1,2,…,N}; and

[0019] The second action list is composed of action lists {a′ i |i=1,2,…,N} represents;

[0020] Among them, a i represents the i-th action object, which includes information about one or more virtual characters in the scene of the stage performance; a′ i is the i-th action object, which includes information about one or more virtual characters in a scene of a stage performance; and N is the total number of action objects performed by multiple characters in multiple scenes of the stage performance.

[0021] The device, wherein:

[0022] The stage performance is t |t=1,2,…,T} represents, where p t is the character's stage performance at time t, where T is the total performance time; and

[0023] Corresponding to a i The stage performance is performed by Indicates that It is action a i duration, and is from the action list {ai |i=1,2,…,N} derived fixed values.

[0024] Other aspects of the present invention also include contents understood by those skilled in the art from the description, claims and drawings of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] The following drawings are based on the embodiments disclosed in the present invention and are examples for illustrative purposes only and are not intended to limit the scope of the present invention.

[0026] Figure 1 A functional schematic diagram of a device for generating a script into a movie according to an embodiment of the present invention is shown.

[0027] Figure 2 A schematic structural diagram of a device for generating a script into a movie according to an embodiment of the present invention is shown.

[0028] Figure 3 A flowchart of a method for generating a movie from a script according to an embodiment of the present invention is shown;

[0029] Figure 4A and Figure 4B A schematic diagram showing camera placement according to an embodiment of the present invention;

[0030] Figure 5 A functional schematic diagram of another device for generating a script into a movie according to an embodiment of the present invention is shown. DETAILED DESCRIPTION

[0031] Reference will now be made in detail to the embodiments of the present invention illustrated in the accompanying drawings. Hereinafter, embodiments of the present invention will be described with reference to the accompanying drawings. Wherever possible, the same reference numerals will be used in the drawings to refer to the same or similar components. It should be understood that the described embodiments are some, but not all, embodiments of the present invention. Based on the disclosed embodiments, a person of ordinary skill in the art may derive other embodiments consistent with the present invention, all of which are within the scope of protection of the present invention.

[0032] The "Write-A-Movie" technology is an adaptive, self-enhancing automatic movie generation framework that automatically generates videos from movie scripts. The present invention provides a script-to-movie generation device that utilizes the latest advances in natural language processing, computer cinematography, and video understanding. The automatic workflow of the script-to-movie generation device greatly reduces the time and knowledge required for the script-to-movie generation process. By combining a novel hybrid objective evaluation mechanism, the video generation process has been mapped to an optimization problem aimed at generating better quality videos, which simultaneously considers the comprehensibility of the visual presentation of the movie script and compliance with cinematography guidelines. Dynamic programming can solve the optimization problem and obtain the optimal solution with the most efficient computational complexity. Therefore, the script-to-movie generation device described in the present invention greatly accelerates the movie production process.

[0033] In the traditional film industry, scriptwriting and filmmaking are completely separate processes. With recent advances in artificial intelligence (AI), a significant portion of the filmmaking process can now be performed by computers. Combining scriptwriting and filmmaking offers direct benefits for all parties involved. Scriptwriters can visualize and edit their work before submitting it. Producers can screen scripts by viewing pre-visualized versions, rather than reading hundreds of pages. The script-to-film generation process must meet two quality requirements. First, the output film must maintain reasonable fidelity to the script. Second, the output film must adhere to film aesthetics and adhere to film codes.

[0034] Therefore, a mechanism is needed to evaluate the fidelity of generated videos to the corresponding movie scripts and, if the assessed fidelity falls below acceptable standards, provide feedback to the animation and cinematography processes for further improvement. Therefore, the computer cinematography process needs to consider not only aesthetics but also the perceptual ability to achieve fidelity to the movie script. While current state-of-the-art video understanding capabilities are not yet sufficient to accurately assess the fidelity of generated videos to movie scripts, it is sufficient for evaluating certain types of movies with less complex scenes and less challenging action recognition.

[0035] In an embodiment of the present invention, the script-to-movie generation device automatically converts the movie script into a movie, such as an animated film. The script-to-movie generation device includes an arbitration mechanism that is supported by video understanding technology and natural language understanding technology. The video understanding technology converts the generated video into a list of executed actions, and the natural language understanding technology converts the movie script into a list of expected actions, so that it can be determined whether the generated video can be understood and the fidelity of the movie script. The evaluation results are then fed back to the stage performance to improve the quality of the generated video. In addition, the aesthetic and fidelity requirements are combined in a unified evaluation framework, and the video quality improvement process is mapped into an optimization problem, which is to achieve the desired video quality by adjusting the camera settings and the character action settings. The optimization problem is designed to be solvable by dynamic programming to reduce computational complexity.

[0036] Figure 1 FIG. 1 is a functional diagram of a device for generating a movie from a script according to an embodiment of the present invention. Figure 1 As shown, the script (ie movie script) is input into the action list generation process to generate an action list arranged in chronological order. The action list includes {a i |i=1,2,…,N} represents a list of expected actions, where a i represents the i-th action object, which includes information about one or more virtual characters in a scene of a stage performance; N is the total number of action objects performed by multiple characters in multiple scenes of the stage performance. i |i=1,2,…,N} is a collection of action objects used to generate character performances during a stage performance. These objects are arranged in chronological order and do not overlap. For example, the characters may be virtual characters in an animated film. In some embodiments, multiple characters execute action objects simultaneously, allowing a single action object to encompass multiple characters in the same scene. For example, two characters may be fighting, or a mother may be hugging her daughter.

[0037] In some embodiments, the action list {a i Each action in |i=1,2,…,N} includes attributes such as subject, action, object, action duration, subject start position, subject end position, subject emotion, and action style. The subject start position is the subject's position when the action starts. The subject end position is the subject's position when the action ends. The default value of the subject emotion is neutral. The default value of the action style is neutral (i.e., no preferred style). The user can select an action style from the following: self-action (when the camera focuses on the subject), multi-action (when the camera focuses on the subject and the object at the same time), and environmental action (when the camera focuses on the environment around the subject, such as the view around the action).

[0038] refer to Figure 1 , the action list {a i |i=1,2,…,N} is input into the stage performance process to generate a video. During the stage performance, the input action list {a i |i=1,2,…,N} is converted into corresponding stage performance data, and the stage performance data is represented by {p t |t=1,2,…,T} represents, where p t is the stage performance of the character at time t, and T is the total performance time. t |t=1,2,…,T} is continuous. However, due to the limitation of computing power, the continuous information is converted into discrete information for camera optimization. The stage performance data is recorded as p for each time unit (e.g. half a second). t In this specification, stage performance, stage performance data and character performance are used interchangeably.

[0039] For the action list {a i For each action in |i=1,2,…,N}, the corresponding performance data is Indicates that It is action a i duration, and is from the action list {a i |i=1,2,…,N} derived fixed values. In some embodiments, different action objects overlap with each other. For example, two events occur at the same time and both need to be shown to the audience. In various scenes, all cameras capture all views of all characters from all angles. The camera optimization process then calculates the best camera path to capture the character performance. The camera optimization process converts the stage performance data {p t |t=1,2,…,T} as input, calculate the optimal camera settings for each time t, the optimal camera settings are given by {c t |t=1,2,…,T} represents. In some embodiments, the camera setting includes at least one of a camera path and a camera parameter. The camera optimization process is performed on the discrete data from time t to T. The camera setting {c t |t=1, 2, ..., T} represents all allowed camera selections for each time slot from time t to T, and for each time slot, only one camera can be selected in the camera optimization process.

[0040] In some embodiments, the camera optimization process identifies a camera path with a minimum distortion D. The distortion D is calculated based on a cost function derived from the cinematographic guidelines. A camera path corresponding to the stage performance data {p is then generated based on the optimized camera settings. t |t=1,2,…,T} corresponding to the video, the video is represented by {f t |t=1,2,…,T} represents.

[0041] Because the camera optimization process only minimizes errors from an aesthetic perspective, the script-to-film generation device of the present invention also considers the fidelity of the generated video to the film script. On the one hand, it is necessary to evaluate the fidelity in an objective measurement. On the other hand, it is necessary to incorporate the measurement of the fidelity into the camera optimization process to minimize aesthetic distortion. Therefore, the generated video is evaluated and output after it meets the passing criteria, and the passing criteria ensure the quality of the output video. If the aesthetics or fidelity of the generated video is determined to be unacceptable, the camera optimization process or the stage performance process is iterated one or more times to generate another video with adjusted camera settings and / or adjusted character performances.

[0042] In some embodiments, when a camera is identified as the cause of the generated video failing to meet a passing condition, the corresponding cost associated with the identified camera is maximized as a subsequent iteration of the camera optimization process or the stage performance process. In other words, the identified camera is removed from the cameras used to capture the stage performance.

[0043] In some embodiments, the video understanding process converts the candidate video {f t |t=1,2,…,T} is used as input to generate another action list, which includes a list of performed actions. The action list identified by the video understanding process is composed of {a′ i |i=1,2,…,N} represents, where a′ i is the i-th action object, which includes information about one or more virtual characters in a scene of a stage performance; N is the total number of action objects performed by multiple characters in multiple scenes of the stage performance. Then, the arbitration process compares the action lists {a i |i=1,2,…,N} and the action list {a′ i |i=1,2,…,N} to obtain the fidelity error E. The fidelity error E is used to quantify the consistency between visual perception and textual meaning, where the visual perception is the visual perception of the generated video and the textual meaning is the textual meaning of the movie script. t|t=1,2,…,T}, the arbitration process also considers the total aesthetic distortion D. t When the total aesthetic distortion D and the fidelity error E given by |t=1,2,…,T} are unqualified, a wider range of acceptable camera settings and acceptable settings of character action performances will be considered to re-optimize the calculation and then re-arbitrate. Repeat the iteration until the candidate video {f t |t=1,2,…,T} is qualified or the iteration count reaches the preset count threshold.

[0044] In some embodiments, after comparing the action list {a i |i=1,2,…,N} and the action list {a′ i |i=1,2,…,N}, the action list {a′ i All actions in |i=1, 2, ..., N} are sorted according to their similarity. If it is necessary to optimize the stage performance, the action with the highest similarity is selected from the sorted list for re-production.

[0045] Figure 5 A functional schematic diagram showing another device for generating a script into a movie according to an embodiment of the present invention is shown. Figure 5 The script-to-movie generation device shown is similar to Figure 1 The script-to-movie generation device shown is similar. The difference lies in whether the video understanding process and the arbitration process are omitted. The specific implementation method can be referred to the previous description and will not be repeated here.

[0046] In an embodiment of the present invention, the script-to-movie generation device leverages recent advances in natural language processing, computational cinematography, and video understanding to significantly reduce the time and knowledge required for the script-to-movie generation process. By incorporating a novel hybrid objective evaluation mechanism that simultaneously considers the visual comprehensibility of the script and compliance with cinematography guidelines, the video generation process is mapped into an optimization problem aimed at generating higher-quality videos. Dynamic programming can be used to solve the optimization problem and obtain the optimal solution with the most efficient computational complexity. Therefore, the script-to-movie generation device of the present invention significantly accelerates the filmmaking process.

[0047] Figure 2 FIG1 shows a schematic diagram of the structure of the device for generating a movie from a script according to some embodiments of the present invention. Figure 2 As shown, computing device 200 includes a processor 202, storage medium 204, display 206, communication module 208, database 210, and peripherals 212, as well as one or more buses 214 coupling the devices together. Certain devices may be omitted and others may be included.

[0048] The processor 202 may be any suitable processor or multiple processors. In addition, the processor 202 may include multiple cores for multi-threading or parallel processing. The processor 202 may execute a sequence of computer program instructions or program modules to perform various processes, such as requesting a user to input director's prompts on a graphical user interface, generating / rendering an animated video, translating director's prompts for editing and optimizing the animated video, etc. The storage medium 204 may include a memory module, such as a read-only memory (ROM), a random access memory (RAM), a flash memory module, an erasable memory, and a large-capacity memory such as a read-only compact disk (CD-ROM), a USB flash drive (U-disk), a hard disk, etc. When the processor 202 runs the storage medium 204, the storage medium 204 may store computer program instructions or program modules for implementing various processes.

[0049] In addition, the communication module 208 may include a network device for establishing a connection through a communication network. The database 210 may include one or more databases for storing certain data (e.g., images, videos, animation materials, etc.) and performing operations on the stored data, such as searching the database and retrieving data.

[0050] The display 206 may include any suitable type of computer display device or electronic device display (e.g., cathode ray tube (CRT) or liquid crystal (LCD) based devices, touch screens, light emitting diode (LED) displays, etc.). The peripherals 212 may include various sensors and other input / output (I / O) devices, such as speakers, cameras, motion sensors, keyboards, mice, etc.

[0051] In operation, the computing device 200 can perform a series of actions to implement the disclosed automatic cinematography method and framework. The computing device 200 can run a terminal or a server, or a combination of the two. The terminal as used herein can refer to any suitable user terminal with a certain computing capability, including collecting director prompts input by the user, displaying a preview video, and editing and optimizing the video. For example, the terminal can be a personal computer (PC), a workstation computer, a server computer, a handheld computing device (tablet), a mobile terminal (mobile phone or smartphone), or any other user-side computing device. The server described herein can refer to one or more server computers configured to provide certain server functions, including determining camera settings for shooting an animated video, generating the animated video based on the camera settings, and editing the animated video by finding the path with the minimum cost function in the image model. The server can also include one or more processors to execute computer programs in parallel. The terminal and / or the server can be configured to provide the structure and functionality for the above-mentioned actions and operations. In some embodiments, some actions can be performed on the server, while other actions can be performed on the terminal.

[0052] The present invention also provides a method for generating a movie from a script. Figure 3 Flowchart showing the method for generating a movie from a script according to some embodiments of the present invention. Figure 2 The computing device 200 shown in FIG. Figure 3 As shown, the method includes the following steps.

[0053] S302, obtaining a movie script.

[0054] Specifically, the movie script is used to generate a video corresponding to the movie script.

[0055] S304: Generate a video according to the movie script.

[0056] Specifically, generating a video according to the movie script includes generating a first action list according to the movie script, generating a stage performance according to each action in the first action list, and shooting the video of the stage performance using one or more cameras. In some embodiments, the first action list is a time-ordered action list, including actions intended to present a visual presentation of the movie script. The first action list is composed of {a i |i=1,2,…,N} means, where a i represents the i-th action object, which includes information about one or more virtual characters in a scene of a stage performance; N is the total number of action objects performed by multiple characters in multiple scenes of the stage performance.

[0057] In some embodiments, the stage performance is t |t=1,2,…,T} indicates that the stage performance is based on the first action list {a i Each action in |i=1,2,…,N} is generated, where p t is the stage performance of the character at time t, and T is the total performance time. In some embodiments, for each action a i for stage performances Indicates that It is action a i duration, and is from the first action list {a i |i=1,2,…,N} derived fixed values.

[0058] In some embodiments, one or more cameras capture images from the stage performance. t |t=1,2,…,T} candidate videos {f t |t=1,2,…,T}. In the stage performance, one or more cameras are deployed as planned and bound to each character.

[0059] S306: Optimize the generated video until it meets the passing condition. The optimization process can be performed based on the aesthetic evaluation and fidelity evaluation of the video.

[0060] Specifically, optimizing the generated video until a pass condition is met includes evaluating a total aesthetic distortion D of the video, the video being shot by one or more of the cameras from the stage performance; generating a second action list based on the video, the video being shot by one or more of the cameras from the stage performance; determining a fidelity error E between the first action list and the second action list, and iteratively optimizing camera settings and character performance to minimize the total aesthetic distortion D, thereby meeting the pass condition. The pass condition includes satisfying that the fidelity error E is less than or equal to a pre-configured fidelity error threshold Th E Or minimize the number of iterations until the count reaches a pre-configured count threshold.

[0061] In some embodiments, the candidate video {f t |t=1,2,…,T}, the total aesthetic distortion D, the candidate video is taken from the stage performance by one or more cameras {p t |t=1,2,…,T} shooting.

[0062] In some embodiments, the total aesthetic distortion D includes the following: tThe role visibility V(c t ). The role visibility V(c t ) by calculating To evaluate, k is the ratio of the size of the character k in the current video frame to the total size of the current video frame. k Indicates how easy it is for the audience to notice the character k in the video frame. t When the camera is in the field of view, the camera c t Treat the role it is bound to as the most important role. I(c t ,k) by the camera c t and the role k, giving different weights to the combination of different roles and different cameras, I(c t ,k) represents the camera c t and the correlation between the role k. I(c t ,k) is a low value, indicating that the character k is the camera c t The more important role is I(c t ,k) the lower the value, the more the character k is to the camera c t is more important.

[0063] In some embodiments, the total aesthetic distortion D also includes the character action A (c t ). The character action A(c t ) describes whether character k is acting at time t. The audience is more likely to notice characters in motion. If character k is acting at time t, the camera c bound to the character k is more likely to be selected. t For example, A(c t ) can be obtained according to the following formula:

[0064]

[0065] In some embodiments, the total aesthetic distortion D also includes the t Camera settings The camera settings By calculation To evaluate, represents the camera position, Represents the shooting direction, represents the action object at time t, and φ C () represents the distortion cost function of the camera setup.

[0066] Different camera setups serve different purposes in filmmaking. For example, when a character is performing a general action, a medium shot is most often used. When a character is performing a specific action, a pan shot, a long shot, a surrounding shot, and a character view shot are usually better choices. On the other hand, different actions may require the camera to be shot from different directions. For example, walking and running actions can be shot from the front and back of the character with minimal distortion. However, speaking actions may be more distorted when shot from the back of the character than from the front and side of the character. Therefore, the distortion of the camera setup depends on the time t from which the action object (i.e., a i ) action type derived from camera c t The derived camera position p and shooting direction d.

[0067] In some embodiments, the total aesthetic distortion D also includes screen continuity S (c t ,c t-1 ), the screen continuity S(c t ,c t-1 ) includes a summary of the position changes of each individual character in the current video frame. The screen continuity S(c t ,c t-1 ) is calculated by To evaluate, where p(k,c t ) indicates that the camera c t The position of character k in the current video frame, p(k,c t-1 ) indicates that the camera c t-1 The position of character k in the current video frame; if character k appears in the camera c t In the view of v(k,c t )=1, otherwise v(k,c t )=0;φ S () is the penalty for the role position change, which is about the role position p(k,c t ) and the character position p(k,c t-1 ) is a nonlinear function of the distance between them.

[0068] Visual-spatial continuity in the video can prevent the video viewer from feeling distortion. For example, cinematography guidelines include the 180-degree rule. The minimum penalty for the character position change is 0, and as the character position p(k, c t ) and the character position p(k,c t-1 ) increases. When character k appears in only one video frame, a maximum penalty of 1 is applied.

[0069] In some embodiments, the total aesthetic distortion D also includes motion continuity M(c t ,c t-1 ), the mobile continuity M(c t ,c t-1 ) includes a change in the direction of the character's movement, the change in the direction of the character's movement being caused by the camera c t The view change is caused by the character action before or after the movement. t ,c t-1 ) is calculated by To evaluate, where m(k,c t ) indicates that the camera c t The character motion direction vector in the current video frame shot, m(k,c t-1 ) indicates that the camera c t-1 The character's motion direction vector in the current video frame captured; φ M () is the penalty for the change of the character's movement direction, which is about the character's moving direction vector m(k,c t ) and the character's moving direction vector m(k,c t-1 ). The penalty increases as the angle between the motion direction vectors increases. When character k appears in only one video frame, a maximum penalty of 1 is applied.

[0070] In some embodiments, the total aesthetic distortion D also includes the shot duration distortion The shot duration distortion By calculation To evaluate, is the average shot duration set for each scene, q is the maximum allowed shot duration, φ U () is the penalty for the shot duration of a video frame that changes the camera in the range [tq,…,t].

[0071] Shot duration is closely related to the degree of audience attention. Generally speaking, the shorter the shot duration, the more intense the content in the video is and the easier it is to attract the audience's attention. In some embodiments, an average shot duration is configured for each scene in the shot duration distribution. In some other embodiments, shot duration configurations learned from existing movies are used for various scenes in the shot duration distribution.

[0072] After summing up various distortions, the total aesthetic distortion ω0, ω1, ω2, ω3, and ω4 are values ​​between 0 and 1, and are weights of each distortion component.

[0073] In some embodiments, the second action list is generated based on the stage performance. Specifically, one or more cameras are taken from the stage performance. t |t=1,2,…,T} shoot candidate videos {f t |t=1,2,…,T}. Then, according to the candidate video {f t |t=1,2,…,T} generates the second action list. The second action list is a list of actions arranged in chronological order, including a list of executed actions. The second action list consists of {a′ i |i=1,2,…,N} represents, where a′ i is the i-th action object, which includes information about one or more virtual characters in a scene of a stage performance; and N is the total number of action objects performed by multiple characters in multiple scenes of the stage performance.

[0074] In some embodiments, a fidelity error E between the first action list and the second action list is determined, and camera settings and character performance are optimized to minimize the total aesthetic distortion D, so that the passing condition (e.g., the fidelity error E is less than or equal to a preset fidelity error threshold Th) is satisfied. E ).

[0075] Specifically, the action similarity between the first action list and the second action list is compared to obtain a fidelity error E. The fidelity error E is used to quantify the consistency between visual perception and textual meaning, wherein the visual perception is the visual perception of the generated video and the textual meaning is the textual meaning of the movie script. t The total aesthetic distortion D is also considered when |t=1, 2, ..., T}. When the fidelity error E is less than or equal to the preset fidelity error threshold Th E When the candidate video {f t |t=1,2,…,T} is qualified. If the candidate video {f t When the total aesthetic distortion D and the fidelity error E given by |t=1,2,…,T} are unqualified, a wider range of acceptable camera settings and acceptable settings for character action performances will be considered to re-optimize the calculation and then recalculate the fidelity error E. Repeat the iteration until the candidate video {f t |t=1,2,…,T} is qualified or the iteration count reaches the preset count threshold.

[0076] In some embodiments, the fidelity error E between the generated video and the movie script can be approximated and evaluated by the difference between a first action list and a second action list, where the second action list is derived from the candidate video via a video understanding process. The video understanding process reads the candidate video and outputs an action list identified from the candidate video. Generally speaking, the video understanding process can be performed as well as a human, and the action list generation process can thoroughly understand the movie script. Therefore, it is feasible to use action list difference comparison to approximate the fidelity error E. The fidelity error is mainly caused by the character performance or the camera shooting process. In the former case, the character performance does not trigger the natural human intuition to reflect the specific actions in the movie script. In the latter case, there is a lack of views that match the specific meaning in the movie script. In practical applications, errors may occur in the video understanding process and the action list generation process. However, in the embodiments of the present invention, errors in the video understanding process and the action list generation process are not considered.

[0077] In some embodiments, the action difference d i Used to represent two related actions i and a′ i The arbitration process uses the GloVe (Global Vectors for Word Representation) word embedding model to generate two description vectors, and then calculates the difference between the two vectors as Where G() is the GloVe word embedding model. Therefore, the fidelity error E can be expressed as Description. By defining the function W() as: when time t is equal to a i The start time, W(t) = d t , otherwise W(t)=0, then the above formula can be transformed into

[0078] In some embodiments, the camera settings are optimized to minimize the overall aesthetic distortion D. Specifically, camera placement is optimized for different lens sizes, different silhouette angles, and different camera heights. Multiple virtual cameras are placed around each rigged character. Each camera maintains its relative position to the rigged character.

[0079] Positioning a camera in three-dimensional space to capture video that satisfies two-dimensional constraints is a 7-degree-of-freedom problem, encompassing the camera's position, orientation, and focal length (i.e., lens dimensions). In practical applications, optimizing in seven dimensions can consume significant computational power. To simplify the problem without loss of generality, we reduce the infinite 7-degree-of-freedom search space to a countable number of discrete camera settings, based on camera positions from classic movies.

[0080] In some embodiments, only cameras with a maximum of two characters are considered, as shots with more characters in view can often be replaced by multiple single-character shots. Consider a shot of two characters with a torus model. Figure 4A and Figure 4B A schematic diagram illustrating camera placement for some embodiments of the present invention. A camera that maintains its relative position to a rigged character during a stage performance is called a Point of View (POV) camera. The POV camera follows the rigged character's head movements.

[0081] In some embodiments, Figure 4A As shown, during the stage performance, each character is bound to 34 cameras. Each camera is marked with an index number. The 34 cameras include 1 POV camera (index number 0), 3 close-up shot (CS) cameras (index numbers 1-3), 20 medium shot (MS) cameras (index numbers 4-23), 2 environment medium shot (MS-S) cameras (index numbers 24-25), 4 full shot (FS) cameras (index numbers 26-29) and 4 long shot (LS) cameras (index numbers 30-33). The profile angle (i.e., shooting direction) of each camera is Figure 4A Indicated by separate dotted arrows. Of the 34 cameras, 8 MS cameras (indexes 4-11) and 2 MS-S cameras are deployed at the character's eye level (e.g. Figure 4B The relative position is as shown in Figure 4A As shown (i.e., the dotted arrows at 0°, 60°, 90°, 120°, 180°, -120°, -90°, -60°); 6 MS cameras (indexes 12-17) are deployed at high angles of the character (e.g. Figure 4B The relative position is as follows Figure 4A As shown (i.e., at 60°, 90°, 120°, -120°, -90°, -60° as indicated by the dotted arrows); the other 6 MS cameras (indexes 18-23) are deployed at low angles of the character (such as Figure 4BThe relative position is as follows Figure 4A As shown (i.e., the dotted arrows at 0°, 60°, 90°, 120°, 180°, -120°, -90°, and -60°). The two MS-S cameras set to observe the environment in front of the character have the following Figure 4A Profile angle shown.

[0082] In some embodiments, the Lagrange multiple method is used to relax the error constraints of the identification so that the shortest path algorithm can be used to solve the problem. The Lagrange cost function is J λ (c t ,a t )=D+λ·E, where λ is the Lagrange multiplier. If there exists λ * Make And E=Th E Then {c * t ,a * t} is the equation The optimal solution of Therefore, the task of solving the above equation is transformed into a simpler task, which is to find the function J that minimizes the Lagrangian cost. λ (c t ,a t ) and choose appropriate Lagrange multipliers to satisfy the constraints.

[0083] In some embodiments, z k =(c k ,a k ) and the cost function G T (z T-q ,…,z T ) is defined as minimizing the fidelity error E and the total aesthetic distortion D up to and including the k-th video frame, where z k-q ,…,z t is the decision vector from the (kq)th video frame to the kth video frame. Therefore, G T (z T-q ,…,z T ) represents the minimum value of the sum of the fidelity error E and the total aesthetic distortion D of all video frames, so

[0084]

[0085] In some embodiments, the key observation for deriving an efficient algorithm is the fact that given q+1 decision vectors z for the (kq-1)th video frame to the (k-1)th video frame k-q-1 ,…,z k-1 , given the cost function G k-1 (z k-q-1 ,…,z k-1 ), for the next decision vector z k The choice of is independent of the previous decision vector z1,z2,…,z k-q-2 This means that the cost function can be recursively expressed as

[0086]

[0087] S308: Output the optimized video.

[0088] Specifically, after the video is optimized according to the present invention, the quality of the optimized video is improved, and the optimized video is output to complete the process from script to movie.

[0089] The recursive expression of the cost function above makes future steps of the optimization process independent of the past steps of the future steps, which is the basis of dynamic programming. This problem can be transformed into a graph theory problem of finding the shortest path in a directed acyclic graph (DAG). The computational complexity of the algorithm is O(T × |Z| q+1 )(where Z is the action list {a i All actions described in |i=1,2,…,N} are performed on the stage {p t The computational complexity depends directly on the value of q. In most cases, q is a small number, so the algorithm is much more efficient than an exhaustive search algorithm with exponential computational complexity.

[0090] In an embodiment of the present invention, the script-to-movie generation method leverages recent advances in natural language processing, computational cinematography, and video understanding to significantly reduce the time and knowledge required to generate a movie from a script. By incorporating a novel hybrid objective evaluation mechanism that simultaneously considers the visual comprehensibility of the movie script and compliance with cinematography guidelines, the video generation process has been mapped into an optimization problem aimed at generating better quality videos. Dynamic programming can solve the optimization problem and obtain the optimal solution with the most efficient computational complexity. Therefore, the script-to-movie generation method of the present invention significantly accelerates the filmmaking process.

[0091] The principles and implementations of the present invention are illustrated in the specification using specific examples. The description of the embodiments is intended to help understand the method and core inventive concept of the present invention. Furthermore, those skilled in the art may change or modify the specific implementation and scope of the present application based on the embodiments of the present invention. Therefore, the contents of the specification should not be construed as limiting the present invention.

Claims

1. A method for generating a movie from a script, characterized in that: include: Get movie scripts; generating a video according to the movie script; Optimizing the generated video until a passing condition is met; and Outputting the generated video; Generating a video according to the movie script includes: generating a first action list according to the movie script; Generate stage performance data according to the elements in the first action list; and extracting a video of the stage performance data using a camera; The video generated by the optimization is optimized until the passing condition is met, including: evaluating a total aesthetic distortion value of the video, the video being extracted from the stage performance data by a camera; generating a second action list according to the video; Determining a fidelity error between the first action list and the second action list E ;and Iteratively optimize camera settings and stage performance data to minimize overall aesthetic distortion D , thereby satisfying the passing condition, wherein the passing condition includes satisfying the fidelity error E Less than or equal to the preset fidelity error threshold Or minimize the number of iterations until the count reaches a preset count threshold.

2. The method according to claim 1, wherein: Each action in the first action list and the second action list has attributes, and the attributes include subject, action, object, action duration, subject start position, subject end position, subject emotion and action style.

3. The method according to claim 1, wherein: The first action list is represented by a time-ordered action list ;and The second action list consists of action lists arranged in chronological order express; in, Indicates the i an action object, wherein the action object includes information of one or more virtual characters in a scene of the stage performance data; It is i an action object, wherein the action object includes information of one or more virtual characters in a scene of the stage performance data; N It is the total number of action objects performed by multiple characters in multiple scenes of the stage performance data.

4. The method according to claim 3, wherein: The stage performance data is used Indicates that For time t The character's stage performance data, T is the total performance time; and correspond The stage performance data is provided by Indicates that It's action duration, and is from the action list Fixed value for export.

5. The method according to claim 4, characterized in that: for , the optimized camera settings are given by indicates; and correspond Video by express.

6. The method according to claim 5, characterized in that The evaluating a total aesthetic distortion value of the video, the video being extracted from the stage performance data by a camera, comprises: Role k In the camera settings Role visibility under By calculation To evaluate, is the character in the current video frame k The ratio of the size of to the total size of the current video frame, Indicates camera and said character k The correlation between The lower the value, the more likely the character is to be k For the camera The importance of is higher; If the camera Bound character k In time t When there is action, move the character Evaluates to 0 otherwise 1; By calculation To evaluate the camera Camera settings ,in represents the camera position, Represents the extraction direction, represents the action object at time t, and a distortion cost function representing the camera setup; By calculation To evaluate screen continuity , the screen continuity Includes a summary of the position changes of each individual character in the current video frame, where Indicated by the camera Extract the character in the current video frame k location, Indicates that the camera Extract the character in the current video frame k position; if the character k Appearing on the camera In the view of ,otherwise ; It is a penalty for the change of the character's position. and character position A nonlinear function of the distance between them; By calculation To assess mobility continuity , the mobile continuity Including the change of the character's moving direction, the change of the character's moving direction is caused by the camera The view change is caused by the character action before or after Indicated by the camera The extracted character motion direction vector in the current video frame, Indicated by the camera The extracted character motion direction vector in the current video frame; It is the penalty for the change in the character's movement direction, which is about the character's moving direction vector The character's moving direction vector a nonlinear function of the difference between ; and By calculation To evaluate the shot duration distortion ,in is the average shot duration set for each scene, q is the maximum allowed shot duration, is the penalty for the shot duration of a video frame that is The camera was changed within the scope; The total aesthetic distortion , 、 、 、 and is a value between 0 and 1, which is the weight of each distortion component.

7. The method according to claim 6, characterized in that Determining a fidelity error between the first action list and the second action list E include: calculate to determine the action differences between the text descriptions of the first action list and the second action list, wherein It is the GloVe word embedding model; Defining a function When the time t equal At the start time of ,otherwise ;and calculate , where T is the total execution time.

8. The method according to claim 7, characterized in that The camera settings are optimized to minimize the overall aesthetic distortion D include: Place multiple cameras around a rigged character, each maintaining its relative position to the rigged character, allowing for optimized camera placement for different shot sizes, different silhouette angles, and different camera heights.

9. The method according to claim 8, characterized in that The iterations optimize the camera settings and stage performance data to minimize the overall aesthetic distortion D , thereby satisfying the passing conditions, including: Make .

10. The method according to claim 9, characterized in that: definition ,in is the Lagrange multiplier; and Will Make Simplified to .

11. The method according to claim 10, characterized in that: definition ; Define the cost function To represent the fidelity error of all video frames E and total aesthetic distortion D The minimum value of the sum; and 。 12. The method according to claim 11, wherein: ;as well as in: Future steps of the optimization process are independent of past steps of said future steps; The optimization process is converted into a graph theory problem of finding the shortest path in a directed acyclic graph; and The computational complexity of the optimization process is , and is more efficient than the exhaustive search algorithm which has exponential computational complexity.

13. A device for generating a movie from a script, characterized in that: include: Memory for storing program instructions; and a processor coupled to the memory, the processor configured to execute program instructions to: Get movie scripts; generating a video according to the movie script; Optimizing the generated video until a passing condition is met; and Outputting the generated video; Generating a video according to the movie script includes: generating a first action list according to the movie script; Generate stage performance data according to the elements in the first action list; and extracting a video of the stage performance data using a camera; The video generated by the optimization is optimized until the passing condition is met, including: evaluating a total aesthetic distortion value of the video, the video being extracted from the stage performance data by a camera; generating a second action list according to the video; Determining a fidelity error between the first action list and the second action list E ;and Iteratively optimize camera settings and stage performance data to minimize overall aesthetic distortion D , thereby satisfying the passing condition, wherein the passing condition includes satisfying the fidelity error E Less than or equal to the preset fidelity error threshold Or minimize the number of iterations until the count reaches a preset count threshold.

Citation Information

Patent Citations

  • Method and device for processing video and electronic equipment

    CN105872857A

  • Dynamic presentation method and device of statement meanings, electronic equipment and storage medium

    CN111246248A

  • Method and system for generating an animated movie

    EP3176787A1

  • Automated cinematographic editing tool

    WO2009055929A1