Methods and devices for generating movie videos using game engines

CN122580686APending Publication Date: 2026-08-14THE HONG KONG POLYTECHNIC UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-07
Publication Date
2026-08-14

AI Technical Summary

Technical Problem

然而,在此类环境中的电影构图设计仍然具有挑战性,因为设计者需要在大量选项中筛选出某时刻帧级别上的理想构图,并确保不同时刻的构图衔接流畅[13]

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122580686A_ABST
    Figure CN122580686A_ABST
Patent Text Reader

Abstract

A computer implementation method for generating movie videos using a game engine and a corresponding computer device are disclosed. A non-transitory computer-readable medium including executable code is also disclosed, which, when executed by the processor of a computing device, causes the device to perform the aforementioned computer implementation method.
Need to check novelty before this filing date? Find Prior Art

Description

Cross-references to related applications

[0001] This application claims priority to U.S. Provisional Application No. 63 / 644109, filed May 8, 2024, with the United States Patent and Trademark Office, the entire contents of which are incorporated herein by reference. Technical Field

[0002] This invention relates generally to film construction, and more specifically to computer-based methods and related equipment for generating movie videos using game engines. Background Technology

[0003] Filmmaking based on game engines Game engine-based filmmaking is a new approach to making film videos that differs significantly from traditional filmmaking methods

[12] . In traditional filmmaking, a key stage in the production process is the shot design stage[5]. In this stage, the composition of each consecutive “beat”

[41] in the story scene is carefully planned, using a “storyboard” as a reference for the final image realization

[23] . Creative shot design transforms a scene from a unique viewing perspective to express the intended emotion (called “focus” in narratology)[1, 10]. To master this skill, novice designers often turn to principles of cinematography that outline basic “shot grammar”

[10] . In contrast, experienced designers tend to expand their design skills through film analysis and drawing inspiration from master examples[6, 40, 49]. This practice aligns with many studies on design creativity that have shown that referencing “design paradigms” from similar or even different fields can foster creativity in designers, even those who have already received creative training[9, 27, 31, 50].

[0004] Game engine-based filmmaking refers to creating film compositions and prototyping them in a virtual 3D environment using real-time game engine-based filmmaking platforms such as UnityTimeline, Unreal Engine Sequencer, and CryENGINE Trackview. Originating in the gaming industry for creating in-game cutscenes, this approach has been increasingly adopted by traditional filmmaking, particularly in pre-production, due to its significant advantages. Compared to traditional filmmaking, which typically relies on fixed storyboards, the "real-time" advantage of game engine-based filmmaking allows for more flexible and cost-effective exploration of creative film compositions.

[0005] Game engine-based filmmakers can easily edit the placement, movement, and animation of virtual characters, enabling them to experiment and iterate on various alternatives from different viewing angles in real time, ultimately determining the most suitable option

[36] . Given this advantage, the traditional filmmaking industry is increasingly using advanced game engines for script rehearsals and exploration of innovative film composition design concepts before putting them into the storyboard

[33] . A prominent example is the production of the animated film EVANGELION: 3.0 + 1.0 THRICE UPON A TIME, which achieved significant commercial success and professional acclaim[3]. However, despite the many advantages that mecha films offer, creating a coherent narrative through compositional sequences in such dynamic environments remains a challenge

[32] , requiring game engine-based filmmakers to explore and evaluate a vast array of design options.

[0006] Creativity support tools The third wave of human-computer interaction (HCI)-oriented creativity research introduced Creativity Support Tools (CSTs), advocating their use to enhance the potential of human creativity [19-21, 48]. Despite extensive research on CSTs, tools for designers and artists engaged in storytelling and content generation remain scarce. Frich et al. highlighted the role of CSTs in enhancing design outcomes, improving design knowledge, and providing new implementation methods [20, 44]. However, in game engine-based filmmaking, the focus has been primarily on the needs of beginners, such as the system proposed by Nicolas et al., which uses a rule-based approach to evaluate the “correctness” of shot design

[14] . However, the need for experienced filmmakers to seek real-time feedback on potential design alternatives has been overlooked. Davis et al.’s “distributed exploratory visualization” model highlights this need, emphasizing that CSTs should facilitate a smooth and low-cost creative exploration and evaluation process

[13] .

[0007] Intelligent cinematography The automation of cinematography tasks in physical and virtual environments has been a focus of numerous autonomous or interactive cinematography tools. For physical scenes in live-action films, a large body of research has proposed methods to suggest coherent sequences of cinematic compositions based on provided video clips. For example, Moorthy et al.

[43] developed a tool that generates a sequence of shots with appropriate timing by utilizing a master shot that covers the entire scene. Arev et al. [4] recommended shot sequences by analyzing and editing multiple sets of video sources to identify narrative focal points in the scene. While many of these methods are applicable to scenarios in animation or digitally constructed films, composing virtual scenes presents unique characteristics and challenges.

[0008] For virtual scenes, a pioneering study

[25] proposed a set of cinematography guidelines based on game engines, in addition to the classic rules [5] summarized by Arijon. Subsequent work [15, 28, 35] has implemented autonomous cinematography on game engine platforms based on these guidelines. However, it can be challenging to formalize all possible rules into computational models, and there is a great deal of variability in how these guidelines are applied in practice. As a result, this rule-based approach is prone to producing a limited variety of results. Recent work has begun to explore data-driven approaches, training machine learning models from data to directly predict cinematic composition. For example, Edirlei et al.

[15] trained an SVM model to select from several pre-defined fixed camera poses in a scene. Jiang et al.

[29] trained a deep learning model to extract camera behavior from reference film clips and reapply it to a given 3D animation. Unfortunately, all of these approaches are fully automated. They do not allow users to specify their preferences and inject their ideas into the results, which is a key requirement identified in the previous section.

[0009] Evin et al.

[17] introduced Cine-AI, an interactive filmmaking tool that uses a game engine to generate film videos in a style similar to those produced by human directors. It allows manual adjustment of the generated results. However, the generated shot sequences are unrelated to the content of the input scene and rely entirely on a predefined set of rules. Therefore, the generated shot sequences are usually of low quality and have problems with poor relevance to the input scene and lack of coherence. Therefore, a lot of manual retouching is still required to produce reasonable results, making it unable to effectively support filmmaking based on game engines. Filmmaking needs to obtain immediate quality feedback at the narrative level

[13] .

[0010] Film composition design Planning the camera's position and movement through scenes of unfolding events to inform the narrative is a fundamental task in filmmaking, often referred to as film composition. Setting the camera's posture—its angle and distance relative to the actors in the scene—determines how the actors and set are framed on screen. Different compositions can convey different meanings and messages to the audience, such as... Figure 1As shown in scenario 100. Designing effective film composition is a challenging task, as it requires extensive knowledge, experience, and creativity in cinematography. As mentioned above, game engine-based platforms are gaining popularity in the filmmaking industry, enabling more effective exploration and evaluation of design alternatives in virtual environments. However, film composition design in such environments remains challenging, as designers need to sift through a large number of options to find the ideal composition at the frame level at a given moment and ensure smooth transitions between compositions at different moments

[13] . Despite the abundance of cinematography rules and reference examples, filmmakers still find it difficult to draw inspiration from them in practice when building on machine-made films. Although there has been some research on automating cinematography in virtual environments, none of the studies have provided an effective solution to support creativity in the process of film composition design.

[0011] Therefore, there is a need for a solution that can address at least one of the problems in the prior art, and / or to provide an option that is useful in the art. Summary of the Invention

[0012] The technology described in this article may involve a method and apparatus for generating movie videos using a game engine.

[0013] According to a first aspect, a computer implementation method for generating movie videos using a game engine is disclosed. The method includes: determining feature vectors for each frame of a three-dimensional (3D) animation video, semantic information of the subject matter of the 3D animation, semantic information of the estimated emotional state conveyed by the 3D animation, and an embedded representation of the camera pose in the 3D animation at a first plurality of time steps, wherein the camera pose is associated with at least two characters in corresponding frames of the video, and each feature vector characterizes the 3D spatial relationship between the two characters, and wherein the embedded representation is a toroidal coordinate representation of the camera pose; based on the feature vectors, performing temporal processing on keyframes of the 3D animation to obtain keyframe feature vectors for each keyframe, wherein each keyframe defines a key moment in the 3D animation and is configured to be located at the center of a plurality of frames within a local window in the 3D animation, and wherein each keyframe feature vector is... The configuration includes contextual information of a local window; processing semantic information of the subject matter of the 3D animation and semantic information of the estimated emotional state through a first feedforward neural network to obtain a first embedding representation set; and processing one-hot vector embedding representations of camera pose and keyframe feature vectors through a second feedforward neural network to obtain a second embedding representation set associated with keyframes at a second plurality of time steps, wherein the one-hot vector embedding representations are based on the embedding representations of camera pose; processing the first and second embedding representation sets through an autoregressive transformer to obtain a third embedding representation set of predicted camera poses associated with two characters at a second plurality of time steps, wherein the first plurality of time steps precede the second plurality of time steps in temporal order; and generating a corresponding conditional probability distribution of the predicted camera pose and multiple probabilities of keyframes marking shot boundaries based on the third embedding representation set through multiple decoders.

[0014] Alternatively or additionally, multiple frames in a local window may include up to 5 frames.

[0015] Alternatively, each keyframe feature vector can be configured as an 1123-dimensional vector.

[0016] Additionally or alternatively, temporal processing of keyframes to obtain keyframe feature vectors may include processing the determined feature vectors using a set of 1D convolutions along the time axis.

[0017] Alternatively, two characters can be selected from multiple characters in a 3D animated video.

[0018] Additionally or alternatively, the two characters may include characters from a 3D animated video, as well as virtual characters introduced as reference characters.

[0019] Additionally or alternatively, the method may also include receiving a 3D animated video as input.

[0020] Alternatively or additionally, multiple first feedforward neural networks and multiple second feedforward neural networks can be implemented using a multilayer perceptron (MLP).

[0021] Additionally or alternatively, multiple decoders may include a multilayer perceptron (MLP) configured with a softmax output layer for class probability prediction, and an MLP configured with a sigmoid output layer.

[0022] Alternatively or additionally, the camera pose may be limited by an embedding representation defined as follows: ,in For camera posture, and For two roles The 2D screen position of the head, and is the angle in the toroidal coordinate system.

[0023] Additionally or alternatively, the method may further include: providing a subset of the predicted camera pose at a given time step, based on the ordinal likelihood value of a subset of the predicted camera pose obtained at that time step, with reference to the conditional probability distribution of the predicted camera pose.

[0024] Additionally or alternatively, the semantic information of the subject matter and the semantic information of the estimated emotional state of the 3D animation can be encoded into their respective one-hot vectors before being processed by the first feedforward neural network.

[0025] Additionally or alternatively, processing the one-hot vector embedding representation of the camera pose and the keyframe feature vector through the second feedforward neural network to obtain the second embedding representation set may include, for each time step: processing the one-hot vector of the corresponding camera pose through the third feedforward neural network; concatenating the processed one-hot vector of the corresponding camera pose with the keyframe feature vector at that time step to obtain a concatenated embedding representation; and providing the concatenated embedding representation to the feedforward neural network in the second feedforward neural network for processing.

[0026] Additionally or alternatively, each feature vector can be defined as ,in, There are two characters 3D distance between heads , It is a connecting role and roles The angle formed by the first line of the head and the second line connecting the shoulders.

[0027] Additionally or alternatively, each conditional probability distribution can be configured to be referenced across multiple quantization categories.

[0028] According to a second aspect, a computing device for generating movie videos using a game engine is disclosed, comprising: one or more memories storing executable code; and one or more processors coupled to the one or more memories and configured to execute the code to enable the device to: The feature vectors of each frame of a 3D animation video, the semantic information of the subject matter of the 3D animation, the semantic information of the estimated emotional state conveyed by the 3D animation, and the embedding representation of the camera pose in the 3D animation at a first plurality of time steps are determined, wherein the camera pose is associated with at least two characters in the corresponding frame of the video, and each feature vector represents the 3D spatial relationship between the two characters, and wherein the embedding representation is a toroidal coordinate representation of the camera pose. Based on feature vectors, time processing is performed on keyframes of 3D animation to obtain keyframe feature vectors for each keyframe. Each keyframe defines a key moment in the 3D animation and is configured to be located at the center of multiple frames within a local window in the 3D animation. Furthermore, each keyframe feature vector is configured to include context information of the local window. The first feedforward neural network processes the semantic information of the subject matter of the 3D animation and the semantic information of the estimated emotional state to obtain a first embedding representation set, and the second feedforward neural network processes the one-hot vector embedding representation of the camera pose and the key frame feature vector to obtain a second embedding representation set associated with the key frame at a second multiple time step, wherein the one-hot vector embedding representation is an embedding representation based on the camera pose. The first and second embedding representation sets are processed by an autoregressive transform to obtain a third embedding representation set of predicted camera poses associated with the two characters at a second plurality of time steps, wherein the first plurality of time steps precede the second plurality of time steps in temporal order; and The corresponding conditional probability distribution for predicting camera pose is generated based on a third embedding representation set by multiple decoders, as well as multiple probabilities for keyframes that mark lens boundaries.

[0029] According to a third aspect, a computing device for generating movie videos using a game engine is disclosed, comprising: A means for determining feature vectors of individual frames of a video of a three-dimensional (3D) animation, semantic information of the subject matter of the 3D animation and semantic information of the estimated emotional state conveyed by the 3D animation, and an embedding representation of camera pose in the 3D animation at a first plurality of time steps, wherein the camera pose is associated with at least two characters in the corresponding frame of the video, and each feature vector characterizes the 3D spatial relationship between the two characters, and wherein the embedding representation is a toroidal coordinate representation of the camera pose. An apparatus for performing time processing on keyframes of a 3D animation based on feature vectors to obtain keyframe feature vectors for each keyframe, wherein each keyframe defines a key moment in the 3D animation and is configured to be located at the center of multiple frames within a local window in the 3D animation, and wherein each keyframe feature vector is configured to include context information of the local window. An apparatus for processing semantic information of the subject matter of a 3D animation and semantic information of estimated emotional state through a first feedforward neural network to obtain a first embedding representation set, and for processing one-hot vector embedding representations of camera pose and keyframe feature vectors through a second feedforward neural network to obtain a second embedding representation set associated with keyframes at a second plurality of time steps, wherein the one-hot vector embedding representation is an embedding representation based on camera pose. A means for processing a first and a second set of embedded representations via an autoregressive transform to obtain a third set of embedded representations of predicted camera poses associated with two characters at a second plurality of time steps, wherein the first plurality of time steps precede the second plurality of time steps in temporal order; and An apparatus for generating, via multiple decoders, a corresponding conditional probability distribution for predicting camera pose based on a third embedded representation set, and multiple probabilities for keyframes marking lens boundaries.

[0030] According to the fourth aspect, a non-transitory computer-readable medium is disclosed, comprising executable code that causes the device to perform the method of the first aspect when the processor of the computing device executes the code.

[0031] Other benefits and advantages of the disclosed aspects will become apparent from the specification and drawings. These benefits and / or advantages are available individually from the various aspects and features of the specification and drawings, and not all of these aspects and features must be provided in order to obtain one or more of these benefits and / or advantages. Attached Figure Description

[0032] In the accompanying drawings, the same reference numerals refer to the same or similarly functional elements in each drawing, which, together with the following detailed description, are incorporated in and form part of this specification, and are used to illustrate various aspects and explain the various principles and advantages as interpreted in this disclosure.

[0033] Figure 1 The illustrations depict example scenarios of how scenes can be constructed from different camera viewpoints to convey different emotions, in some aspects of this disclosure.

[0034] Figure 2 The use of Cinemassist in some aspects of this disclosure is illustrated. This system / software is designed to support and enhance the creative process by generating multiple cinematic composition schemes at the keyframe and scene levels.

[0035] Figure 3 The user interface of Cinemassist is shown in some aspects of this disclosure.

[0036] Figure 4 A camera pose representation based on a toroidal coordinate system is shown in some aspects of this disclosure.

[0037] Figure 5 An exemplary framework for implementing a deep generative model of Cinemassist is shown in some aspects of this disclosure.

[0038] Figure 6 This disclosure illustrates how Cinemassist can be used to generate diverse camera pose sequences to convey the story of an animation.

[0039] Figure 7 The following are examples of training methods shown in some aspects of this disclosure. Figure 5 3D camera pose estimation for datasets of deep generative models.

[0040] Figure 8 The flowchart illustrating a method for generating movie videos on a game engine-based platform is shown in some aspects of this disclosure.

[0041] Figure 9 Some aspects of this disclosure show the confidence ratings reported by participants in user studies after designing film compositions without and with Cinemassist.

[0042] Figure 10 The distribution of participants’ self-satisfaction ratings for each question regarding the usability and ease of use of Cinemassist in user studies is shown in some aspects of this disclosure.

[0043] Figure 11 The results of an expert rating study are shown in some aspects of this disclosure, involving user preference rankings for using standard Unity features (manual), using Cinemassist, and using composition sequences automatically generated by Cinemassist.

[0044] Figure 12 Exemplary film composition sequences created by users with and without Cinemassist, as well as composition sequences automatically generated by Cinemassist, are shown in some aspects of this disclosure.

[0045] Figure 13 This is a block diagram illustrating an exemplary implementation of Cinemassist in some aspects of this disclosure.

[0046] Figure 14 This is a block diagram of a keyframe feature extractor in some aspects of this disclosure.

[0047] Figure 15 This is a block diagram of a camera pose generator in some aspects of this disclosure.

[0048] Figure 16 It is used for implementation in some aspects of this disclosure. Figure 8 A schematic diagram of an exemplary computing device for the method.

[0049] Figure 17 It is used for implementation in some aspects of this disclosure. Figure 8 A schematic diagram of an exemplary computing device. Detailed Implementation

[0050] For the sake of completeness, it is hereby clarified that any reference to the format definition “[X]” in any paragraph of the description of this disclosure shall be interpreted as referring to the corresponding citation “X” in the “References” section of this specification. For example,

[10] refers to citation

[10] listed in the “References” section of this specification, while [21-24] accordingly refers to citations

[21] -

[24] , with necessary modifications.

[0051] The following descriptions are presented, explicitly or implicitly, in the form of algorithms and functions or symbolic representations that operate on data in computer memory. These algorithmic descriptions and function or symbolic representations are means by which those skilled in the art of data processing effectively communicate the essence of their work to other practitioners in the field. An algorithm is generally considered to be a series of self-consistent steps that lead to a desired result. These steps require physical operations on physical quantities, such as electrical, magnetic, or optical signals that can be stored, transmitted, combined, compared, and otherwise manipulated.

[0052] Unless otherwise specified, and as will be apparent from the following text, it should be understood that throughout this disclosure, discussions using terms such as “scan,” “calculate,” “determine,” “replace,” “generate,” “initialize,” “output,” etc., may refer to the actions and processes of a computer system or similar electronic device that manipulate and / or transform data in the computer system, expressed in physical quantities, into other data similarly expressed in physical quantities in the computer system or other information storage, transmission, or display devices.

[0053] This disclosure also discloses apparatus / devices for performing method operations. Such apparatus / devices may be specifically constructed or arranged for a desired purpose, or may include a computer or other device selectively activated or reconfigured by a computer program stored in a computer. The algorithms and displays presented herein are not inherently related to any particular computer or other device. Various machines may be used with computer programs based on the teachings herein. Alternatively, dedicated apparatus for performing method steps may be constructed as appropriate. For completeness, the structure of a conventional computer will also be described below.

[0054] Furthermore, this disclosure implicitly discloses a computer program, as it will be apparent to those skilled in the art that the various steps of the methods described herein can be implemented by computer code / instructions. This computer program is not intended to be limited to any particular programming language or implementation thereof. It should be understood that the teachings contained herein can be implemented using a variety of programming languages ​​and their encoding. Moreover, this computer program is not intended to be limited to any particular control flow. Many other variations of the computer program may exist, using different control flows, without departing from the spirit or scope of this disclosure.

[0055] Furthermore, one or more steps of a computer program may be executed in parallel rather than sequentially. Such a computer program may be stored on any computer-readable medium. Computer-readable media may include storage devices such as magneto / optical disks, memory chips, or other storage devices suitable for interfacing with a general-purpose computer. Computer-readable media may also include wired media (as exemplified by Internet systems) or wireless media (as exemplified by GSM, GPRS, 3G, 4G, 5G, NR mobile communication systems and other wireless communication systems / standards such as Bluetooth, ZigBee, or Wi-Fi). When loaded and executed on such a computer, the computer program effectively forms means for implementing various aspects of this disclosure.

[0056] One or more aspects of this disclosure can also be implemented as associated hardware modules. More specifically, at the hardware level, a module is a functional hardware unit designed to work in conjunction with other components or modules. For example, a module can be implemented using discrete electronic components or can be part of an overall electronic circuit, such as an application-specific integrated circuit (ASIC) or a field-programmable gate array (FPGA). Many other possibilities are known in the art. Those skilled in the art will also understand that the system can be implemented through a combination of hardware and software modules.

[0057] Some aspects of this disclosure provide a (computer-implemented) method and corresponding apparatus for generating movie videos using a game engine. The following discussion sets forth the disclosed subject matter in accordance with various aspects of this disclosure.

[0058] Figure 2Exemplary use of Cinemassist 200 is shown in some aspects of this disclosure. Cinemassist is the proposed system / software for supporting and enhancing the creative process (for film composition, such as in game engine-based filmmaking) by generating various film composition schemes at the keyframe and scene levels. Figure 3 An example of the Cinemassist user interface 300 is depicted.

[0059] Cinemassist is an intelligent creativity support tool configured to assist in the conceptualization and exploration of film composition, targeting not only game engine-based filmmakers but also traditional filmmakers who use game engine-based platforms to develop storyboards during pre-production. To this end, on-site interviews were conducted with three professional game engine-based filmmakers to better understand the game engine-based filmmaking workflow, composition design process, and design challenges. Based on the interview results and a literature review of traditional solutions, Cinemassist was proposed and developed to support the film composition design process.

[0060] Given a 3D animation and additional conditional semantics (such as film theme and expected emotional state), Cinemassist enables users to design 3D camera pose trajectories (i.e., film composing) through multiple iterations. In each iteration, both Cinemassist (in the computer context) and the user (in the human context) can contribute to and influence the design. On the computer side, Cinemassist first suggests a diverse set of composing alternatives that may fit the content of the user-selected keyframes and are coherent with previous composing designs. On the human side, the user turns to the suggested designs for ideas and creates a final composing, which is then provided to Cinemassist to generate suggestions for the next iteration. This forms an interactive workflow that alternates between human decision-making and computer suggestions throughout the design process, allowing users to draw inspiration from a variety of suggested options while enjoying considerable freedom to customize the film composing design.

[0061] This paper presents a user study in which representative users from diverse backgrounds designed film composition sequences with and without Cinemassist. Expert ratings were also used to evaluate user design outcomes. The results demonstrate that Cinemassist can effectively facilitate the user's creative process in various ways, helping to produce better design results and enhancing the performance of users with animation expertise. Furthermore, based on the research findings, several design insights were identified for further iterations of Cinemassist's design and to provide a reference for the design of related creativity support tools for tasks in such real-time 3D environments. According to this disclosure, the contributions made to the development of Cinemassist are described below.

[0062] Contributions: Cinemassist discloses its design principles and system implementation to facilitate the creative process of cinematic composition design in game engine-based filmmaking. Cinemassist's target users include digital game artists / designers creating game cutscenes on game engine-based filmmaking platforms, as well as traditional filmmakers using game engines for storyboarding during pre-production.

[0063] Technical methodological contribution: A deep generative model for Cinemassist is disclosed, trained on existing film data to generate diverse camera pose trajectories to depict the story of a given 3D animation in varied ways. This model is at the core of Cinemassist, enabling unique interactive workflows, the ability to recommend diverse and coherent film composition schemes, and the flexibility to acquire the knowledge needed to guide users without relying on hard-coded rules.

[0064] Design Insights: Based on user research and expert ratings, four design insights were derived that can provide a reference for the design of creativity support tools for design tasks in 3D real-time environments.

[0065] Understanding the Game Engine-Based Film Production Workflow field interviews To better understand the design process and thinking of game engine-based filmmakers, we conducted in-person interviews with three professional game engine-based filmmakers. The three interviewees, P1 (female, junior game designer), P2 (female, junior game designer), and P3 (male, lead game artist), came from a world-leading digital game company. P1 and P2 had approximately 3 to 5 years of professional narrative design experience, using game engine-based filmmaking platforms to prototype cutscenes for MMO games. P3 had over 10 years of professional experience, using game engine-based filmmaking platforms not only to create cutscenes but also to produce 3D digital films. The interviews were face-to-face and lasted approximately 60 minutes. The interview questions were structured around three themes, with no specific restrictions on the answers: the game engine-based filmmaking workflow, film composition design, and the difficulties encountered during the design process. Their feedback is summarized below.

[0066] Game Engine-Based Film Production Workflow: Based on feedback from three respondents, this paper summarizes a common four-stage workflow in digital game cutscenes and 3D digital film production. In the first stage, designers receive text scripts describing the scenes. The second stage involves translating these scripts into concrete scenes, including building the scene environment and planning 3D animation within the scene by specifying character positions, movements, and behaviors using character models and animations. Notably, the planning process is typically based on keyframes representing key moments in the animation. The third stage focuses on designing a series of cinematic compositions to tell the story of the animation, also based on keyframes. It's worth noting that P3 commented, "3D scene and animation production can be separated from cinematic composition design not only in cutscene production but also in traditional film production workflows, as the latter stage can be completed using intermediate works." Finally, in the fourth stage, the scenes are rendered, and the film sequence is merged into video or descriptive script formats.

[0067] Film Composition Design: In the design process of the film composition sequence (the third stage of the workflow mentioned above), all three respondents emphasized their tendency to validate and refine the initial design of each keyframe. Specifically, they controlled the virtual camera to synthesize three main aspects: the position, orientation, and scale of the narrative focus on the screen. Time permitting, respondents also explored alternative designs to pursue more creative results. Evaluation criteria for each design included: fit with the current keyframe scene content, coherence with the previous keyframe, and effective expression of the intended emotion in the animation. Furthermore, all respondents mentioned their tendency to draw inspiration from "masterpieces" that transcend traditional cinematography norms. P1 shared, "I carefully archive film clips that I find interesting for reference in my work." Nevertheless, P1 also acknowledged that understanding the fundamental principles behind these masterworks and applying them to improve their own cinematography is a challenge.

[0068] Challenges: All respondents mentioned the challenges of the film composition stage, especially given the limited production time. P1 and P2 specifically pointed out that adjusting the camera position and orientation to compose each anticipated keyframe was extremely time-consuming and required precise and intensive manual control using a mouse and keyboard. P3 complained, "Trying different alternatives to film composition is crucial, but this exploration is severely constrained by the production schedule."

[0069] Design requirements Based on feedback collected from field interviews, the design requirements for building a tool that intervenes in the film composition design process, namely the third stage of the game engine-based film production workflow, where 3D scenes and animations are ready. First, Cinemassist is configured to significantly enhance the exploration of creative film composition design in a typical game engine-based film production workflow (R1). Second, Cinemassist is configured to provide real-time film composition design suggestions at the keyframe and scene levels

[13] (R2). Third, Cinemassist is configured to recommend a diverse range of feasible design alternatives as inspirational examples to extend traditional design paradigms based on cinematography rules (R3). Fourth, Cinemassist is configured to recommend coherent film compositions that seamlessly integrate with the specific scene context, thereby minimizing the need for manual retouching (R4). Finally, the film composition designs recommended by Cinemassist are displayed in real time for users to quickly evaluate (R5). The design features of Cinemassist will be discussed in the next section, and the corresponding labels R1, R2, R3, R4 and R5 will be used to correspond to the requirements mentioned above.

[0070] The proposed system: CINEMASSIST Based on design requirements extracted from field interviews, Cinemassist, a creativity support software tool designed to enhance creative film composition design exploration, was proposed and implemented. Specifically, Cinemassist implements an interactive paradigm, alternately applying human decision-making and intelligent suggestions throughout the entire film composition design process. Cinemassist's target users include game engine-based filmmakers and traditional filmmakers who conceive and pre-visualize storyboards in game engine environments. This section provides an overview of Cinemassist's functionality and explores its design principles. The Cinemassist user interface is as follows... Figure 3 As shown, it roughly comprises three components: (a) a control panel for configuring input 3D animation and advanced semantics; (b) a design panel for designing cinematic compositions at the frame level; and (c) a storyboard panel for visualizing recommended composition sequences at the scene level. In addition, the user interface provides a scene view for displaying 3D animation in real time and providing suggested camera poses in 3D space, a cinema view for previewing compositions for specific camera poses, and an animation timeline for easy selection of keyframes. Each of these three components will be discussed further below.

[0071] Input configuration: Users can load 3D animated videos into Cinemassist. In the implementation, UnityTimeline is used to prepare the animation for experimental purposes. This tool uses pre-given 3D assets to plan and record the positions, movements and interactions between virtual characters. The animation is displayed in the scene view, and the timeline is displayed below the scene view. Then, in the control panel, users can select two scene objects as characters of interest and select a camera object in the scene as the default control camera (R1). It is worth noting that the reason for focusing on two characters is mainly because the toroidal coordinate system is used to express the camera pose, which assumes that there are two target characters on the screen. In addition, users can select a movie genre category (including "action", "romance" and "thriller") to which the expected animated story belongs, as well as an expected emotional state that the user wants to express to the audience through the movie composition of the animation. According to Plutchik's emotional wheel

[42] , five common expected emotional states are considered, including "happiness", "anger", "surprise", "sadness" and "fear".

[0072] Frame-Level Design Exploration: Before starting to design the cinematic composition for a given animation, users can drag arrows along the timeline to preview the entire sequence. Users can then position the arrows to key moments in the animation, designating them as "keyframes." At this point, users can click the "+Frame" button on the frame-level suggestion panel to capture that keyframe. To design the cinematic composition for this keyframe, in addition to manually finding suitable camera poses using the default control camera, Cinemassist automatically suggests a set of potential camera poses (R2, R3). These suggestions are visually displayed as light bulbs in the scene view, allowing users to explore in 3D space. When a user selects a suggestion, the corresponding 2D composition is immediately displayed in the cinema view for real-time evaluation (R5). Alternatively, users can click the "View" button to expand and display a row of 2D compositions containing all suggested camera poses, sorted by the model's predicted quality score. This feature provides users with a quick visual reference for evaluating the suggestions (R3, R5). Once the desired camera pose is found, users can click the "Record" button to confirm the composition of the current keyframe using that pose.

[0073] The user then proceeds to identify the next keyframe, repeating the system suggestions and user decision-making process described above. Notably, at the end of each iteration, the user's selected options are fed into a deep generative model, enabling it to predict each iteration based on the content of the current keyframe and the previously designed composition, thus generating relevant and coherent composition suggestions. Furthermore, for each keyframe, Cinemassist can indicate whether it represents a shot boundary, i.e., whether a new shot should begin from that boundary, thus helping users better plan their camera in practice. Finally, the user can export the generated composition sequence and quickly assess its overall quality (R5) by clicking the "Play" button to view its rendered animation in cinematic view.

[0074] Scene-Level Design Exploration: In addition to the frame-level design exploration described above, Cinemassist offers a feature that allows users to explore cinematic composition sequences at the scene level. To do this, users first select a range on the timeline in the Storyboard panel by specifying start and end times and setting time intervals. Then, after clicking the "Recommend" button, Cinemassist automatically samples a series of evenly spaced keyframes within the specified time range at the specified time intervals and suggests multiple storyboards (R2, R3, R4). Each storyboard is a cinematic composition sequence, corresponding to one keyframe. Note that the selected range can vary, covering the entire animation or only a small portion. Users can select a satisfactory sequence to export to the Design panel and continue iterating on it through frame-level design exploration. Furthermore, users can export a cinematic composition sequence in progress from the Design panel to the Storyboard panel, where Cinemassist can automatically expand that portion of the sequence into multiple complete sequences (R1) by generating the remaining compositions.

[0075] The proposed model Problem Definition Given a 3D animation and some high-level semantics, including film themes and expected emotional states, the proposed model is configured to generate camera trajectories across the 3D scene, thereby enabling the generated film compositions to well represent animated events and reflect the semantic information of the input. Specifically, the proposed model... A sequence of keyframes sampled from an animation. For the input semantics, where and These represent the subject matter and the expected emotional response, respectively. The proposed model outputs a 3D camera pose sequence. Each keyframe corresponds to one keyframe. To implement Cinemassist, the proposed model primarily considers the following three aspects: The model should be multimodal, capable of generating multiple alternatives from a single input, allowing users to gain inspiration by exploring different options. It should also be highly controllable, meaning it doesn't directly generate the final result but rather provides a mechanism to continuously integrate user decisions into the generation process, thus granting users significant freedom in the design process. Finally, the model should be learnable, enabling it to extract substantial knowledge of film composition directly from the data, rather than relying on fixed rules.

[0076] These considerations are addressed by proposing an autoregressive probabilistic generative model for Cinemassist. Unlike other types of deep generative models, such as Generative Adversarial Networks (GANs) and Variational Autoencoders (VAEs), the autoregressive model generates a sequence of values ​​by relating the prediction of the next value to previous values. This makes the autoregressive model a suitable choice for achieving the desired controlled workflow, in which user decisions (i.e., past values) can be continuously integrated into the generation process, thus influencing future predictions. Formally, the proposed model can be summarized as: aiming to learn a given 3D animation input. and semantics Camera pose sequence Conditional distribution This conditional distribution can be sampled to generate various camera pose sequences (multimodal) consistent with the input, such as... Figure 5 As shown. Furthermore, it can be assumed that this conditional distribution can be decomposed into the following form: .

[0077] In other words, the output camera pose is generated through autoregression. In each iteration... In the middle, the predicted value of the current camera pose. In the manner previously predicted The model uses the proposed frame and camera representations as conditions and as input for the next iteration. This mechanism allows the user to determine the camera pose for the current iteration based on the model's predictions, and the user's decision influences the model's predictions for the next iteration. This gives the user significant control over the generation process (i.e., controllability). Finally, the model is trained using existing movie data without hard-coding any rules (i.e., learnability). The following sections discuss the proposed frame and camera representations, as well as the model's architecture, training details, and generation process.

[0078] Frames and camera representation Each frame in the animation is represented by the 3D pose of one or more characters, each character described by a set of 3D body joint positions. According to

[27] , the camera pose is expressed in a toroidal coordinate system. The toroidal representation is defined in local reference frames of two given targets / selected characters, making it easier to capture the correlation between the 3D character pose and the relative camera pose in the toroidal coordinate system. Given two target characters, the representation of the camera pose is... Defined as: , in , These are the positions of the two characters' heads on the 2D screen. and These are two parameters, angles, in the toroidal representation

[34] , such as Figure 4 The camera pose representation in the diagram is shown schematically. It is worth noting that given the 3D head positions and camera pose representations of two characters, the camera pose in 3D Cartesian space can be reconstructed[8] and used to frame the two characters on the screen. In addition, the toroidal camera representation is defined based on two targets, but the actual 3D animation may contain any number of characters. Therefore, for scenes with more than two characters, the user needs to select two of them as targets, while for single-character scenes, this problem can be circumvented by introducing a virtual “observer” character (as a second character) and treating the head of this character as a camera pointing at the actual character in the scene. This approach is consistent with the “reverse shot” technique

[38] , which is often used in dialogue scenes between two characters, showing each speaker from the other’s perspective, and is considered to work well in practice.

[0079] The continuous camera pose representation is further uniformly quantized as A discrete interval. The camera pose sequence is now represented as a discrete token sequence: ,in It is a special token that indicates the starting position of the sequence. It is the first The model is a one-hot vector representing the pose of each camera. It is trained to predict sequences based on the input in an autoregressive manner for conditional camera pose generation.

[0080] Network Architecture and Training like Figure 5 As shown, the proposed model may include two components: (1) a keyframe feature extractor, configured to learn a feature vector for each keyframe that captures its local contextual information; and (2) a camera pose generator, configured to synthesize a camera pose sequence based on the extracted keyframe features and input semantics.

[0081] Keyframe feature extractor: such as Figure 5 As shown in a, the keyframe feature extractor first uses the character feature extractor. A feature vector is computed for each input animation frame, representing the 3D spatial relationship between two characters in that frame. Similar to the movie character features in

[27] , this vector contains three quantities. ,in It is the 3D distance between the characters' heads. It is the line connecting the character's head and the line connecting the character. (Role The angle between the lines on the shoulders. Then, for each keyframe (in Figure 5 (highlighted in bold in a) is obtained through training the temporal feature extractor. To refine the corresponding feature vectors. The input consists of character feature vectors from five local windows centered on the keyframe. It processes this input using a set of 1D convolutions along the time axis, generating a refined 1123-dimensional feature vector for the keyframe. This refined feature vector enhances the original character features with their local contextual information, based on information identified in the five frames arranged within the associated local windows. The final output of the keyframe feature extractor is a sequence of refined keyframe feature vectors. , used to adjust the camera pose generator.

[0082] Camera pose generator: such as Figure 5 As shown in b, the camera pose generator is implemented using a decoder-only Transformer to generate camera pose sequences in an autoregressive manner. The Transformer architecture was chosen because of its superior capabilities in various sequence modeling tasks, which can be attributed to the powerful performance of its underlying self-attention mechanism in modeling long-term dependencies. According to some aspects of this disclosure, the camera pose generator can learn to capture the temporal dependencies between camera poses in different steps, which ensures the coherence of the generated sequence. For predicting in step... The camera pose generator generates an output embedded representation of the camera pose. The embedding is based on the previous embedding sequence. and the embedding representation of input semantics , As a condition. To obtain In the steps Use encoder The generated camera pose is encoded, and the resulting vector is compared with the keyframe feature vector. The parts are then assembled and provided to the fusion module. .

[0083] and Both can be implemented using a multilayer perceptron (MLP). and By using one-hot encoding to convey the input theme and expected sentiment, two semantic encoders are employed. and Both encoders are implemented using MLP. Then, the output embedding representation... Provided to the camera pose decoder Its generation is in Probability distribution on each quantization category . It is also implemented using an MLP, with a softmax layer configured at the end to predict class probabilities. Furthermore, it uses another decoder. This decoder is also implemented using an MLP and configured with a sigmoid layer at the end, according to... Predict the probability that the current step size is the edge of the lens. .

[0084] During training, the keyframe feature extractor and camera pose generator are on the dataset. The dataset is used for joint training, and this dataset will be detailed in the next section. It includes 3D animation frame by frame, It is the camera pose for each frame. These are the corresponding semantic tags. This is a binary mask vector used to indicate whether a frame is a shot boundary. Both camera pose prediction and shot boundary prediction employ the cross-entropy loss function. The goal is to optimize the combination of these two losses using an Adam optimizer with a learning rate of 0.001. The proposed model was trained for 98 epochs. For each training animation sequence, keyframes were randomly selected at random time points, while ensuring that the interval between two consecutive keyframes was between 1 and 5 seconds. This allows the proposed model to generalize better to handle different keyframe sampling intervals that may be encountered during testing.

[0085] generate After training, the proposed model can be used to generate camera poses in different ways based on given animation and input semantics, thus supporting various features of Cinemassist. First, the model's autoregressive nature allows it to naturally support interactive generation through multiple iterations. Specifically, in each iteration, the model predicts the distribution of all possible camera poses and draws a set of samples from this distribution, sorted by likelihood, as suggested options. The user can then select the desired option, and the model uses the selected option as input to predict the next iteration. This enables frame-level design exploration capabilities on Cinemassist's design dashboard.

[0086] Secondly, high-confidence camera pose sequences are sampled from the model using beam search (e.g., ...). Figure 6 As shown in Example 600 (which annotates different camera trajectories), scene-level design exploration on the storyboard panel is possible without any user intervention. It's worth noting that a beam size of 15 yields the best results compared to other beam size options. Specifically, Figure 6Example 600 shows three sequences of animation generated for two characters, visualized in a 3D scene using three different camera trajectories: blue, cyan, and green. Each trajectory contains three camera poses generated from three keyframes, with both characters belonging to the same keyframe marked with a corresponding timestamp. Solid lines on the trajectory represent camera movement within the same shot, while dashed lines represent camera switching between different shots.

[0087] Third, the proposed model can also accept partial sequences as input and automatically complete the remaining parts. This means that sequences being designed on the design panel can be exported to the storyboard panel, and, as mentioned earlier, suggestions for complete design sequences can also be automatically generated. To enhance the "creativity" of the proposed model, the input to the softmax function of the final layer of the camera pose decoder is divided by a temperature parameter (e.g., 1.0 in the example implementation) to make sampling more random.

[0088] Dataset used to train the model To train the model, a dataset of 18 films was constructed, belonging to three different genre categories (romance, action, and thriller). Films with high IMDb ratings and covering a wide range of cinematography styles were carefully selected for the dataset. Each film was segmented using scene boundary annotations provided by MovieNet

[26] , and then each scene was further segmented into shot sequences using MovieNet's shot detector tool. Figure 7 As shown, in the 3D camera pose estimation of the dataset 700, for each containing For each frame of the scene, the 3D poses of the characters in each frame were extracted using Metrabs

[47] . Notably, even when parts of the characters on the screen are occluded or cropped, Metrabs was able to estimate their full-body 3D poses in the frames. This generated 3D animations of the characters in the scene. .

[0089] Then, the 3D pose of the character in each frame is used to calculate the toroidal camera parameters

[34] , thereby obtaining the camera pose annotation sequence. It is worth noting that the circular camera approach assumes that there are at least two targets in the scene. When a frame contains more than two characters, the two largest characters on the screen are selected as targets according to Hitchcock's cinematography principles

[22] , while for frames with only one character, a virtual observer is introduced as described above. To obtain the subject matter label of the scene, its film subject matter on IMDb is used. The sentiment label of the scene is also estimated by using a text-to-sentiment model [2], which predicts the probability distribution of five basic sentiment categories and then assigns the most likely category to the scene. The semantic label is thus obtained. In addition, each frame is assigned a label indicating whether it is a shot boundary, thereby generating a shot boundary mask. In this way, for each scenario, a value in the form of... The dataset contains 977 samples. By filtering out scenes with no characters or more than 10 characters, these samples were then divided into a 70% and 30% set, with 70% used as the training set and 30% as the test set.

[0090] Based on the above discussion, Figure 8 The flowchart illustrates a method 800 for generating movie videos using a game engine, according to some aspects of this disclosure. Method 800 can be implemented as a computer-executable method. According to some aspects of this disclosure, method 800 can be implemented within and executed by Cinemassist to support all features and functionalities of generating movie videos using a game engine (as referenced above). Figure 5 (as described above). The operation of method 800 can be performed as follows: Figure 16-17 The computer devices 1600, 1700, or components thereof are implemented and executed thereon. In some examples, the computer devices 1600, 1700 may execute a set of instructions to control their functional elements, thereby performing the various functions described below. Additionally or alternatively, the computer devices 1600, 1700 may use dedicated hardware to perform various aspects of the functions described below.

[0091] In step 805, the method may include: determining the feature vectors of each frame of the three-dimensional (3D) animation video, the semantic information of the subject matter of the 3D animation (i.e., Figure 5 The Chinese character is represented as ) and the semantic information conveyed by 3D animation to estimate emotional states (i.e. Figure 5 The Chinese character is represented as The document describes the embedding representation of camera pose in a 3D animation at a first plurality of time steps, where the camera pose is associated with at least two characters in the corresponding frames of the video, and each feature vector characterizes the 3D spatial relationship between the two characters. The 3D animation video can be received as input. It should be understood that the embedding representation of camera pose refers to the toroidal coordinate representation of the camera pose, as described in the "Frames and Camera Representations" section of this disclosure above.

[0092] Step 805 can be performed by a character feature extractor. Execution. Each feature vector can be defined as... ,in, Two characters in a 3D animation 3D distance between heads , It connects to the first selected character. Second selected role The angle formed by the first line of the head and the second line connecting their shoulders.

[0093] The camera pose can be constrained by the following embedded representation: ,in, For camera posture, and For the two selected roles The 2D screen position of the head, and is the angle in the toroidal coordinate system.

[0094] These two characters can be selected from multiple characters (existing in a single frame / scene) in a 3D animated video. In some examples, the two selected characters may include a character from the 3D animated video, as well as a (fictitious) virtual character introduced as a reference "observer" character.

[0095] In step 810, the method may include performing temporal processing on keyframes of the 3D animation based on feature vectors to obtain keyframe feature vectors for each keyframe, wherein each keyframe defines a key moment in the 3D animation and is configured to be located at the center of multiple frames within a local window in the 3D animation, and wherein each keyframe feature vector is configured to include context information of the local window. Step 810 may be performed by a temporal feature extractor. Execution. In 3D animation, multiple frames within a local window can include up to 5 frames, and the feature vector of each keyframe can be configured as an 1123-dimensional vector.

[0096] In this context, temporal processing of multiple keyframes in a 3D animation may include processing a set of 1D convolutions along the time axis to process a defined feature vector.

[0097] It should be understood that, within the context of Cinemassist, steps 805 and 810 can be performed by the keyframe feature extractor. The keyframe feature vector obtained in step 810 corresponds to the refined keyframe feature vector sequence output by the keyframe feature extractor. .

[0098] In step 815, the method may include: using a first feedforward neural network (i.e. Figure 5 As shown and ) processes the semantic information of the subject matter and estimates the semantic information of the emotional state of the 3D animation to obtain a first embedding representation set, and then processes it through a second feedforward neural network (i.e. Figure 5 As shown The camera pose one-hot vector embedding representation and keyframe feature vectors are processed to obtain a second embedding representation set associated with the keyframes at a second set of multiple time steps. The first embedding representation set can be represented as follows: Figure 5 As shown , The second embedding representation set can be represented as Figure 5 As shown It should also be understood that the one-hot vector embedding representation of the camera pose is obtained based on the camera pose toroidal coordinate representation determined in step 805 as the camera pose embedding representation. The “Frames and Camera Representations” section of this disclosure discusses how to obtain the one-hot vector embedding representation of the camera pose.

[0099] Multiple first-feedforward neural networks and multiple second-feedforward neural networks can both be implemented using a multilayer perceptron (MLP). It should be understood that MLP is considered the most basic machine learning model, capable of compressing high-dimensional input features into lower dimensions.

[0100] It should be understood that the semantic information of the subject matter and the semantic information of the estimated emotional state of the 3D animation will be encoded into corresponding one-hot vectors before being processed by the first feedforward neural network.

[0101] Furthermore, it should be understood that the second embedding representation set is obtained by processing the one-hot vector embedding representation of the camera pose and the keyframe feature vectors through a second feedforward neural network. This process may include, for each time step: processing through a third feedforward neural network (i.e., ... Figure 5 As shown The first step involves processing the one-hot vector corresponding to the camera pose; concatenating the processed one-hot vector with the keyframe feature vector at that time step to obtain a concatenated embedding representation; and then providing the concatenated embedding representation to the feedforward neural network in the second feedforward neural network for processing. The third feedforward neural network can also be implemented using an MLP.

[0102] In step 820, the method may include processing a first and a second set of embedded representations via an autoregressive transform to obtain a third set of embedded representations for predicted camera poses associated with the two characters at a second plurality of time steps, wherein the first plurality of time steps precede the second plurality of time steps in temporal order. The third set of embedded representations may be represented as follows: ,like Figure 5 As shown in the image.

[0103] In step 825, the method may include using multiple decoders (i.e. Figure 5 As shown and Based on a third embedding representation set, a corresponding conditional probability distribution for predicting camera pose is generated, along with multiple probabilities for keyframes marking lens boundaries. Each conditional probability distribution is configured to reference multiple quantization categories.

[0104] Multiple decoders include multilayer perceptrons (MLPs) configured with softmax output layers for class probability prediction (i.e., Figure 5 In ), and MLPs with a sigmoid output layer configured at the end (i.e. Figure 5 As shown For clarity, please refer to... Figure 5 Camera pose decoder It is configured to generate a conditional probability distribution for predicting camera pose, while other decoders... Multiple probabilities are arranged to generate keyframes that mark the boundaries of the shot.

[0105] Subsequently, the method may further include: referring to the conditional probability distribution of the predicted camera pose obtained at the time step, and providing a subset of the predicted camera pose at the time step based on the ranking likelihood value of a subset of the predicted camera pose.

[0106] It should be understood that, within the context of Cinemassist, steps 815, 820, and 825 can be performed by the camera pose generator. The third embedding representation set (output by the camera pose generator) obtained in step 825 for predicting the corresponding conditional probability distribution of the camera pose, and the multiple probabilities of the keyframes used to mark shot boundaries can be represented as follows: and ,like Figure 5 As shown in the image.

[0107] Evaluate This section discusses a user study conducted on representative end-users from diverse backgrounds, as well as another expert-rated study designed to evaluate user design outcomes in the initial study. The aim is to assess whether and how Cinemassist can facilitate the user's design process and influence design outcomes.

[0108] Evaluation of users from different backgrounds The usability of Cinemassist may vary depending on the end-user's background. In particular, novice users lacking experience in 3D digital cinematography may struggle to master game engine-based film production platforms, making it difficult to create cinematic content. Conversely, experienced users, while frequently using game engine-based film production tools in their work or studies, may face different design challenges. Considering these potential differences among the target end-users, 18 participants with diverse backgrounds were recruited for the study. All participants had experience using mainstream game engine-based film production platforms such as Unity Timeline, Unreal Sequencer, and CryENGINE Trackview. Nine of the 18 participants had more than three years of digital cinematography experience, while the remaining participants had between one and three years of relevant experience. Furthermore, it was noted that the participants had diverse backgrounds; almost half were animation professionals (majoring in animation or working in the digital game or film industry), while the rest were studying or had graduated from design school without animation expertise.

[0109] Participants were required to use Unity Timeline to design cinematic compositions for a 72-second 3D animation, which was built in a Unity virtual scene using open-source assets from the Unity Asset Store by a digital game expert. The animation’s story followed three phases of a “hero’s journey”

[11] , a classic story pattern widely used in digital game design and film production.

[0110] process Before the formal research begins, a research guide will be read to each participant, introducing the research process. In addition, a tutorial session will be arranged, where each participant will watch a short tutorial video on how to use Unity for cinematic composition design and complete a warm-up exercise with a demo animation to familiarize themselves with cinematic composition design in the Unity Timeline.

[0111] Task PhaseParticipants receive an animation and are tasked with completing two cinematic composition design tasks on the Unity Timeline, one without Cinemassist and the other with Cinemassist. The order of the two tasks is randomly assigned to each participant. There are no time limits for either task, allowing participants to complete them to the best of their ability. Furthermore, the themes and expected emotional states that participants need to express in their composition designs are not pre-defined, allowing them to freely decide the specific messages they want to convey based on their understanding of the animation's storyline. This provides participants with more room to unleash their creativity. While designing using Cinemassist, participants are free to experiment with different input semantic settings and explore different camera behaviors. As a deliverable for each task, participants need to save the final composition sequence as a "storyboard," each storyboard containing a series of different animated camera perspectives, and export the corresponding video. To generate the video of the composition sequence, linear interpolation is used to interpolate camera movement between two consecutive keyframes, and the video is rendered using Unity Recorder at a frame rate of 24 FPS. Additionally, the design process for both tasks is screen-recorded. The time it took each participant to complete the task was also recorded.

[0112] Post-task questionnaire survey After completing the task, participants were required to rate their confidence in the task design outcome using a 5-point Likert scale and explain their reasoning. Specifically, these confidence questions were designed to assess the overall quality of the design outcome, as well as its novelty and appropriateness, which are widely considered to be the essence of “creativity”

[46] . In addition, each participant was required to rate the functionality and ease of use of Cinemassist using a 5-point Likert scale (1: strongly agree, 5: strongly disagree) and explain their reasoning. Importantly, questions 1 through 4 assessed two functions of Cinemassist in promoting participants’ divergent and convergent thinking, which are considered to be two important components of the creative process [18, 24, 45]. Participants further reported their satisfaction with the ease of use of Cinemassist through questions 5 through 7, which followed the “System Usability Scale (SUS) criteria” [7].

[0113] Reflective interview session After completing two tasks and their corresponding questionnaires, semi-structured reflective interviews were conducted with the participants. Interview questions were developed based on observations of the participants' design process and aimed to gather feedback to improve the system's functionality and usability. With each participant's consent, the interviews were recorded for subsequent analysis.

[0114] result Figure 9 The data shows participants' self-assessment of confidence in the quality of their design tasks during the questionnaire survey. The data reveals a clear trend in participants' perception of "overall quality." Without Cinemassist, over 75% of participants were confident in their designs, but none expressed "very much confidence." After using Cinemassist, a similar percentage of participants maintained confidence, and three participants upgraded their self-assessment to "very much confidence." P11 explained, "In framing the shots, I used more professional angles, enhancing the narrative effect." P13 also agreed, stating, "AI seems to have guaranteed quality, making my design look closer to a real film." Conversely, participants' confidence scores regarding the "novelty" of their design outcomes showed a significant change. After completing the task without Cinemassist, less than 25% of participants felt "confident" or "very confident," with most expressing only "neutral" confidence. Furthermore, 25% of participants expressed a lack of confidence. In contrast, after completing the task using Cinemassist, the percentage of participants feeling "confident" or "very confident" rose to over 50%, with no one expressing a lack of confidence. Page 13 highlights, "I made some breakthroughs in expressing the scene's content." Page 16 comments, "Some unexpected shots made the final result more creative." Confidence in the "appropriateness" of the design varied minimally across different tasks, indicating that it had little impact on participants' self-perception in this regard.

[0115] Figure 10 The participants' evaluations of Cinemassist's two features and their ease of use were summarized. Specifically, over 80% of the participants agreed or strongly agreed that the frame-level and scene-level features helped them generate different potential design options. In particular, regarding the frame-level feature, P2 added, "After learning about alternatives to frame-level cinematic composition design, my original 'serial' ideation process transformed into a 'parallel' workflow." P4 added, "These suggestions freed me from the constraints of traditional paradigms." Regarding the scene-level feature, P5 assessed, "The suggested sequences are of above-average quality, with no obvious errors, and even those that exist are easy to correct." Similarly, P7 commented, "The coherence of these sequences already looks very good, requiring almost no manual adjustments."

[0116] In contrast, while all participants agreed or strongly agreed that frame-level features helped them achieve their intended design, less than 40% still agreed with scene-level features. For example, regarding frame-level features, participant P10 commented, “The system consistently recommends alternatives that are close to my envisioned concept, which I can immediately apply or fine-tune to achieve the desired effect.” However, when discussing scene-level features, P10 noted, “The recommended sequences don’t entirely match my initial concept overall.” P5 also acknowledged, “While some sequences may be close to my vision, I still need to spend time modifying the remaining frames to align with my initial concept.” Regarding ease of use, all participants indicated they would like to use Cinemassist frequently in their work. Over 80% of participants disagreed that Cinemassist was too complex, and over 90% confirmed that Cinemassist was easy to use.

[0117] It is noteworthy that some participants were observed to spend more time completing tasks using Cinemassist than without it. During the reflective interview, participants were asked to explain the reasons for this phenomenon. P11 explained, “Two of Cinemassist’s features saved me time manually adjusting the camera position, allowing me to spend more time exploring and evaluating which composition worked best and how to utilize these options in different ways to better tell the story.” When asked for suggestions on Cinemassist’s iterations and ease-of-use improvements, participants offered three points.

[0118] First, Cinemassist's scene-level functionality should recommend more diverse and easily editable composition sequences. P10 suggested, "Could the scene-level functionality allow users to directly edit the camera trajectory of recommended composition sequences, much like adjusting the 'inspiration bulb' at the frame level?" Second, semantic input needs enhancement beyond the existing "subject matter" and "expected sentiment" parameters to include more multimodal semantic specifications, such as expected directorial style, screenplay, and accompanying background music or narration. Finally, Cinemassist should continuously iterate to simplify or automate the task of adjusting the "target object" over time to match the narrative development in 3D animation. P8 complained, "As the story unfolds, I often forget to switch target objects, causing the system to still focus on the wrong objects when making recommendations. Could the system remind me to do so, or help me find the correct target?" P8 also suggested, "Sometimes I only want to focus on capturing one object, and sometimes I need to capture multiple objects simultaneously. Could the system expand its functionality in this area, moving beyond the current fixed 'dual-target' setting?" Expert rating study Two cinematography experts were invited to evaluate the quality of the composition design results from previous user research and the composition results automatically generated by Cinemassist. Both experts are professors from a digital film production academy with over 10 years of experience in digital animated film production and teaching. The automatic results were generated by selecting keyframes at uniform time intervals and randomly selecting target objects and input semantics into the proposed model.

[0119] process First, the experts will receive the story script for the animation and will need to evaluate the design work of 18 participants. When evaluating the participants’ performance, three design works will be randomly presented to the experts: two film composition sequences designed by the participants with and without Cinemassist, and an automatic sequence generated by Cinemassist. During the evaluation process, the experts will need to rank the sequences according to their quality, and rate them as “1st”, “2nd” and “3rd” (corresponding to the order presented on the slides). The ranking is based on two evaluation criteria: coherence and originality, which are considered to be two key aspects of storyboard design

[30] . In addition, the experts will need to rank the three sequences according to their overall quality.

[0120] result Based on the assessments of the two experts, there was a clear difference in evaluations between participants with and without an animation background. Specifically, participants who were professionals or students in the field of animated film production were categorized as having an animation background. Therefore, participants P1 through P9 were categorized as having an animation background, while participants P10 through P18 were categorized as not having an animation background.

[0121] The scoring results are as follows Figure 11 As shown, using Cinemassist significantly improves the performance of participants with an animation background. In nearly 75% of the three evaluation criteria, the two experts rated task designs using Cinemassist (i.e., human plus Cinemassist) as first. In contrast, task designs without Cinemassist (human) mostly ranked second. Furthermore, automatically generated results from Cinemassist were generally ranked last. The two experts agreed that the main flaw in automatically generated sequences was "focusing on the wrong characters" during the development of the animated story. This was primarily due to the irrelevant and randomly set input content for Cinemassist. It is expected that the automated method can achieve better results by using more sophisticated keyframe selection algorithms and a powerful sentiment predictor from the input 3D animation.

[0122] Figure 12 The presentation shows two works created by P7, one with and without Cinemassist, and a visual comparison of an automatically generated work. In the comparison, both experts rated the work created using Cinemassist (human-generated plus Cinemassist) as first, citing its coherence, novelty, and overall quality. Notably, one expert commented, "The composition of this work (human-generated plus Cinemassist) is very novel." Furthermore, another expert commented, "The other work (human-generated) looks somewhat stiff and lacks creativity." Additionally, both experts ranked the automatically generated work (generated by Cinemassist) last in terms of coherence and overall quality. One expert further explained, "I didn't see the two zombies appear as expected," indicating that the storyboard failed to focus on the intended goal.

[0123] Observations revealed that Cinemassist did not improve the performance of participants without an animation background. Specifically, most designs completed without Cinemassist ranked first, while designs using Cinemassist frequently ranked second, with Cinemassist's automatically generated results consistently ranking third. One expert commented on a sequence designed by P13 after completing the task using Cinemassist: "This work does not clearly demonstrate the audiovisual language of filmmaking, thus exhibiting a misuse of shot types." The expert speculated that this phenomenon might occur because users lacking an animation background and basic knowledge of film editing struggle to correctly integrate suggested compositions into a coherent sequence, thus failing to express novel and accurate content.

[0124] Therefore, the study found that while most participants in the initial user study were more confident in the task design outcomes supported by Cinemassist, subsequent expert evaluation studies yielded two drastically different results. The first result supported Cinemassist's ability to improve design outcomes for users with animated backgrounds, while the second result failed to provide similar evidence for Cinemassist's usability for non-animation users. Interestingly, among all the rating results, such as Figure 11 As shown, two outstanding automatically generated results rank first in terms of compositional novelty.

[0125] discuss Data-driven systems vs. rule-based systems Cinemassist was proposed and implemented as a ready-to-use creativity support tool to enhance cinematic composition on game engine-based filmmaking platforms. User research results demonstrate that this tool, employing a data-driven approach, provides multiple design options at both the keyframe and scene levels, significantly improving the novelty and overall quality of user-generated designs. In contrast, traditional game engine-based filmmaking systems rely on predefined rules and often only offer standardized design suggestions. These systems lack awareness of design context, severely limiting their ability to enhance the creativity or efficiency of the design process. While some systems attempt to introduce variability and novelty by adding random noise to rule-based suggestions, this often results in a loss of initial appropriateness, thus reducing their suitability as design alternatives. Cinemassist's data-driven approach effectively addresses these limitations. It leverages a large number of creative examples to learn a deep generative model capable of simultaneously considering spatial and semantic context and synthesizing novel options that go beyond the training examples. This approach produces more context-aware and innovative design options. However, the effectiveness of the data-driven approach is limited by the size of the training dataset. Introducing additional rules can help enhance the accuracy and relevance of model recommendations, thereby ensuring the robustness and appropriateness of the model across various scenarios.

[0126] Traditional design process vs. AI-assisted design process User research results indicate that, compared to traditional methods, Cinemassist can significantly improve the design process by presenting users with a variety of potential camera pose schemes at the keyframe and scene levels. User feedback suggests that this AI-assisted process not only shortens the time required to implement design options but may also provide users with more opportunities to explore and select the optimal design. Therefore, it is an interesting direction to study whether and how AI-suggested design options affect the original cognitive models of game engine-based filmmakers

[13] , especially in the areas of conception, implementation, exploration, and evaluation, which could be used for future iterations.

[0127] The results also show that the effectiveness of Cinemassist depends on the user's background. In particular, for users with an animation background, Cinemassist can significantly improve their design output. However, for users without an animation background, although Cinemassist can enhance their confidence in their own output, it cannot help them improve their performance according to expert ratings. This may be related to the mismatch between confidence and correctness in AI-assisted decision-making scenarios

[37] . One possible reason is that, in order to better support creativity, the model's output is diversified through random sampling, resulting in some suggestions with lower probability. These suggestions, while inspiring, may be difficult for beginners with limited knowledge of film composition to apply correctly. To help beginners more effectively, Cinemassist can be customized to provide only a small number of high-probability options to increase their chances of obtaining high-quality results.

[0128] Design inspiration for creativity support tools Based on the research results, four design insights were derived to provide a reference for the design of creative support tools for design tasks in 3D real-time environments. First, displaying the results of recommended design solutions can enable immediate evaluation, thereby promoting the designer's "divergent thinking" (I1). Second, presenting recommended design solutions as "hints" in the form of game engine-based film production design space (i.e., 3D scene) can promote the designer's "convergent thinking", thereby encouraging them to conduct more in-depth "in-situ" exploration and adjustment (I2). Third, enabling recommended design solutions to be instantly switched between the above two presentation modes can promote the designer's switching between divergent thinking and convergent thinking (the two ends of the cognitive continuum)

[18] , thereby jointly promoting their creative process (I3). Finally, given the two types of user needs observed in user research, the importance of considering the "level of control" is emphasized (I4).

[0129] Propose the advantages of the CINEMASSIST system Design suggestions are becoming more diverse. Unlike traditional systems that typically rely on basic cinematography rules, the proposed deep generative machine learning model is trained on a diverse dataset containing numerous design paradigms from master filmmakers. This approach enables the model to provide creative and diverse design suggestions, significantly expanding the traditional design space. Consequently, filmmakers gain access to a wider range of cinematic compositions, pushing the boundaries of standard filmmaking practices.

[0130] Enhanced coherence and context adaptability Cinemassist's advantage lies in its ability to provide frame-level context-sensitive cinematic composition suggestions while maintaining scene-level coherence. This is thanks to its integration of current 3D scene context and semantic factors, such as subject matter and anticipated emotional expression, during model training and real-time suggestion phases. Therefore, these suggestions are highly appropriate and coherent, enhancing the narrative effect and visual continuity of the animation.

[0131] User-centered design and deployment Cinemassist was conceived and implemented based on a comprehensive understanding of its target users, including both professional and amateur 3D animated filmmakers. This user-centric approach creates a "ready-to-use" solution that integrates seamlessly into filmmakers' existing workflows. Cinemassist is available as a plugin for popular game engine-based filmmaking platforms such as Unity Timeline, Unreal Sequencer, and CryENGINE Trackview. This compatibility allows filmmakers to easily adopt and benefit from these innovative features without disrupting their standard production processes. In short, Cinemassist not only expands the innovative possibilities for 3D animated filmmakers but also ensures that these new features are easy to use and practical, thereby improving the efficiency and quality of the filmmaking process.

[0132] Applications of CINEMASSIST Beyond its initial application scenario (3D digital filmmaking), Cinemassist can be deployed in other potential technology areas: intelligent cinematography in physical environments and virtual reality (VR) spaces. Specifically, in the physical space domain, Cinemassist's cinematic composition suggestions can be transmitted to other autonomous cinematography equipment (such as drones) to explore alternatives in real time. In the VR space domain, a potential variant of Cinemassist could automatically suggest different camera paths for viewers / players located in different positions within a virtual scene / space, enabling them to watch and enjoy adaptive, coherent storytelling / game content.

[0133] Figure 13 This is a block diagram of an exemplary Cinemassist system implementation based on some aspects of this disclosure. Figure 13 In this example, Cinemassist is designated by reference numeral 1305. Based on the above discussion, Cinemassist 1305 can be implemented in software form (e.g., as a computer program). Cinemassist 1305 can be implemented by computer devices 1600, 1700 (e.g.,...). Figure 16-17(as shown) or its components are executed. Cinemassist 1305 may include a keyframe feature extractor 1310 and a camera pose generator 1315. Each of these components 1310 and 1315 may communicate with each other directly or indirectly (e.g., via one or more buses) 1320.

[0134] Alternatively, each of the keyframe feature extractor 1310 and the camera pose generator 1315 can be implemented as a specific hardware module (e.g., an ASIC) to perform the same operation. Nevertheless, the implementation of components 1310 and 1315 can also be chosen to employ a combination of hardware and software modules as needed.

[0135] Figure 14 This is a block diagram of a keyframe feature extractor 1310 according to some aspects of this disclosure. The keyframe feature extractor 1310 can be implemented in software, at least steps 805, 810, which are related to the method 800 described herein. Figure 8 Correspondingly, the keyframe feature extractor 1310 can be generated by computer devices 1600, 1700 (such as...). Figure 16-17 (as shown) or its components are executed. The keyframe feature extractor 1310 may include a determination component 1405 and a timing processing component 1410. Each of these components 1405 and 1410 may communicate with each other directly or indirectly (e.g., via one or more buses) 1415.

[0136] Component 1405 determines feature vectors for each frame of a 3D animated video, semantic information about the subject matter of the 3D animation, semantic information about the estimated emotional state conveyed by the 3D animation, and embedded representations of camera poses in the 3D animation at first multiple time steps. The camera poses are associated with at least two characters in the corresponding frames of the video, and each feature vector represents the 3D spatial relationship between the two characters. (See reference...) Figure 8 As shown, the embedded representation of camera pose refers to the toroidal coordinate representation of camera pose.

[0137] The timing processing component 1410 can perform timing processing on keyframes of a 3D animation based on feature vectors to obtain keyframe feature vectors for each keyframe. Each keyframe defines a key moment in the 3D animation and is configured to be located at the center of multiple frames within a local window in the 3D animation. In addition, each keyframe feature vector is also configured to include context information of the local window.

[0138] Alternatively, each component 1405 and 1410 in the keyframe feature extractor 1310 (including the keyframe feature extractor 1310 itself) can be implemented as a specific hardware module (e.g., an ASIC) to perform the same operation. Nevertheless, the implementation of components 1405 and 1410 can also be chosen to employ a combination of hardware and software modules as needed.

[0139] Figure 15 This is a block diagram of a camera pose generator 1315 according to some aspects of this disclosure. The camera pose generator 1315 can be implemented in software at least steps 815, 820, and 825, which correspond to the method 800 described herein. Figure 8 The camera pose generator 1315 can be generated by computer devices 1600, 1700 (such as...). Figure 16-17 (as shown) or its components are executed. The camera pose generator 1315 may include a first processing component 1505, a second processing component 1510, a third processing component 1515, and a generation component 1520. Each of these components 1505, 1510, 1515, and 1520 may communicate with each other directly or indirectly (e.g., via one or more buses) 1525.

[0140] The first processing component 1505 can process the semantic information of the subject matter of the 3D animation and estimate the semantic information of the emotional state through multiple first feedforward neural networks to obtain a first embedding representation set.

[0141] The second processing component 1510 can process the one-hot vector embedding representation of the camera pose and the keyframe feature vector through multiple second feedforward neural networks to obtain a second set of embedding representations associated with keyframes at a second plurality of time steps. The camera pose generator 1315 is configured to communicate with the keyframe feature extractor 1310 to obtain the keyframe feature vectors therefrom. (See reference...) Figure 8 As shown, the one-hot vector embedding representation of the camera pose is obtained based on the toroidal coordinate representation of the camera pose.

[0142] The third processing component 1515 can process the first and second embedded representation sets via an autoregressive transform to obtain a third embedded representation set of predicted camera poses associated with the two characters at a second plurality of time steps. It should be understood that the first plurality of time steps precede the second plurality of time steps in temporal order.

[0143] The generation component 1520 can generate a corresponding conditional probability distribution for predicting camera pose, as well as multiple probabilities for keyframes marking lens boundaries, based on a third embedded representation set and multiple decoders. Each conditional probability distribution can be configured to reference multiple quantization categories.

[0144] Alternatively, each component 1505, 1510, 1515, and 1520 in the camera pose generator 1315 (including the camera pose generator 1315 itself) can be implemented as a specific hardware module (e.g., an ASIC) to perform the same operation. Nevertheless, the implementation of components 1505, 1510, 1515, and 1520 can also be chosen to employ a combination of hardware and software modules as needed.

[0145] Figure 16 It is used for implementation based on some aspects of this disclosure. Figure 8 A schematic diagram of an exemplary (first) computing device 1600 of method 800.

[0146] The computing device 1600 may include a keyboard 1602, a touchscreen 1604, a microphone 1606, a speaker 1608, and an antenna 1610. Users can operate the computing device 1600 to perform various functions / tasks, such as making phone calls, sending text messages, browsing the Internet, sending emails, and providing satellite navigation.

[0147] Computing device 1600 may include hardware for performing communication functions (such as telephone or data communication), as well as an application processor and corresponding supporting hardware to enable computing device 1600 to establish other functions, such as messaging, internet browsing, email functionality, etc. The communication hardware may include a radio frequency (RF) processor 1612 that provides RF signals to antenna 1610 for transmitting data signals and receives RF signals from antenna 1610. A baseband processor 1614 may be provided that provides signals to and receives signals from the RF processor 1612. As is known in the art, baseband processor 1614 may also interact with a Subscriber Identity Module (SIM) 1616. The communication subsystem enables computing device 1600 to communicate via a variety of different communication protocols, including 3G, 4G, 5G, New Radio (NR), GSM, WiFi, Bluetooth™, and / or CDMA. The communication subsystem of computing device 3400 is beyond the scope of this disclosure.

[0148] The keyboard 1602 and touchscreen 1604 are controlled by the application processor 1618. The power and audio controller 1620 supplies power from the battery 1622 to the communication subsystem, application processor 1618, and other hardware. The power and audio controller 1620 also controls input from the microphone 1606 and audio output through the speaker 1608. Additionally, a Global Positioning System (GPS) antenna and associated receiver element 1624, controlled by the application processor 1618, can be configured to receive GPS signals for satellite navigation functionality of the computing device 1600.

[0149] The computing device 1600 may provide various types of memory to complement the operation of the application processor 1618. The computing device 1600 may include random access memory (RAM) 1626 coupled to the application processor 1618, which data and program code can be written to and read from. Code stored in RAM 1626 can be executed from RAM 1626 by the application processor 1618. RAM 1626 is a type of volatile memory in the computing device 1600.

[0150] The computing device 1600 may also be equipped with non-volatile (long-term) memory 1628 coupled to the application processor 1618. Memory 1628 may be logically divided into three partitions: operating system (OS) partition 1630, system partition 1632, and user partition 1634. Memory 1628 may represent the non-volatile memory of the computing device 1600.

[0151] In this example, OS partition 1630 may contain the firmware of computing device 1600, including the operating system. Other computer programs, such as applications (also known as apps), may also be stored in memory 1628. In particular, applications critical to the operation of computing device 1600, such as communication applications in a smartphone, are typically stored in system partition 1632. Applications stored on system partition 1632 are typically pre-programmed in the factory settings of computing device 1600.

[0152] The user then adds and installs applications to computing device 1600, which are typically stored in user partition 1634.

[0153] Figure 16 The various functional components shown can also be integrated into a single component. For example, memory 1628 may include NAND flash memory, NOR flash memory, hard disk drive, or a combination thereof.

[0154] Figure 17 It is based on some aspects of this disclosure for implementation and execution. Figure 8 A schematic diagram of an exemplary (second) computing device 1700 of method 800. The following description of the computing device 1700 is provided as an example only and is not intended to be limiting.

[0155] like Figure 17As shown, the exemplary computing device 1700 may include a processor 1704 for executing software routines. For simplicity, only a single processor is shown in the figure, but the computing device 1700 may also be configured as a multiprocessor system (i.e., including multiple processors). The processor 1704 is coupled to a communication infrastructure 1706 to communicate with other components of the computing device 1700. The communication infrastructure 1706 may include, for example, a communication bus, a crossbar network, or a network.

[0156] The computing device 1700 also includes a main memory 1708, such as random access memory (RAM), and a secondary memory 1710. The secondary memory 1710 may include, for example, a hard disk drive 1712 and / or a removable storage drive 1714, which may include a floppy disk drive, magnetic tape drive, optical disk drive, etc. As known in the art, the removable storage drive 1714 reads and / or writes data from the removable storage unit 1718. The removable storage unit 1718 may include a floppy disk, magnetic tape, optical disk, etc., and is read and / or written data by the removable storage drive 1714. As those skilled in the art will understand, the removable storage unit 1718 may also include a computer-readable storage medium storing computer-executable program code instructions and / or data.

[0157] In other respects, the auxiliary storage 1710 may also or optionally include other similar means for loading computer programs or other instructions into the computing device 1700 for execution. Such means may include, for example, a removable storage unit 1722 and an associated interface 1720. Examples of removable storage units 1722 and interfaces 1720 may include program cartridges and cartridge interfaces (e.g., cartridge interfaces in video game console devices), removable storage chips (e.g., EPROM or PROM) and associated slots, as well as other exemplary removable storage units 1722 and interfaces 1720 that enable the transfer of software programs and / or data between the removable storage unit 1722 and the computing device 1700.

[0158] The computing device 1700 also includes at least one communication interface 1724. The communication interface 1724 allows software programs and data to be transferred between the computing device 1700 and external devices via a communication path 1726. In several ways, the communication interface 1724 allows data transfer between the computing device 1700 and a data communication network (e.g., a public data network or a private data communication network). The communication interface 1724 can be used to exchange data between different computing devices 1700 that may collectively form part of an interconnected computer network. Examples of the communication interface 1724 may include a modem, a network interface (e.g., an Ethernet card), a communication port, an antenna with associated circuitry, etc. The communication interface 1724 can be configured for wired or wireless connections. The software and data transmitted through the communication interface 1724 are in the form of signals, which may be electronic signals, electromagnetic signals, optical signals, or other signals that can be received by the communication interface 1724. These signals are provided to the communication interface via the communication path 1726.

[0159] The computing device 1700 may also include a display interface 1702 configured to perform operations of rendering images to an associated display 1730, and an audio interface 1732 configured to perform operations of playing audio content through an associated speaker 1734.

[0160] As used herein, the term "computer program product" may refer in part to removable storage unit 1718, removable storage unit 1722, hard disk installed in hard disk drive 1712, or a carrier wave that transmits software to communication interface 1724 via communication path 1726 (e.g., via a wireless link or cable). Computer-readable storage medium means any non-transitory tangible storage medium that provides recorded instructions and / or data to computing device 1700 for execution and / or processing. Examples of such storage media include floppy disks, magnetic tapes, CD-ROMs, DVDs, Blu-ray discs, hard disk drives, ROMs or integrated circuits, USB storage devices, magneto-optical discs, or computer-readable cards (such as PCMCIA cards), whether such devices are located internally or externally to computing device 1700. Examples of transient or non-tangible computer-readable transmission media that may also be involved in providing software, applications, instructions, and / or data to computing device 1700 include radio or infrared transmission channels, network connections to another computer or networked device, and the Internet or intranet (including email transmissions and information recorded on websites).

[0161] Computer programs (also referred to as computer program code / instructions) are stored in main memory 1708 and / or auxiliary memory 1710. Computer programs can also be received via communication interface 1724. When executed, these computer programs enable computing device 1700 to perform one or more aspects of this disclosure discussed above. In various aspects of this disclosure, the computer programs, when executed, enable processor 1704 to perform certain aspects of this disclosure. Therefore, such computer programs can act as controllers for computer device 1700.

[0162] The software can be stored in a computer program product and loaded onto the computing device 1700 using a removable storage drive 1714, a hard disk drive 1712, or an interface 1720. Alternatively, the computer program product can be downloaded directly to the computer device 1700 via communication path 1726. When the processor 1704 executes the software, the computing device 1700 will perform certain aspects of this disclosure.

[0163] It should be understood that, Figure 17 The computing device 1700 shown is for illustrative purposes only. Therefore, in some aspects, one or more features of the computing device 1700 may be omitted. Similarly, in other aspects, one or more features of the computing device 1700 may be combined or juxtaposed. Furthermore, in some aspects, one or more features of the computing device 1700 may be divided into one or more components.

[0164] It should be understood that, Figure 17 The components shown can also be used to provide execution. Figure 8 The various functional devices of the method 800 disclosed herein are as described in various aspects of this disclosure. Furthermore, the terms "computing device" 1600 and 1700 may include or refer to mobile devices, wireless devices, remote devices, handheld devices, tablet computers, laptop computers, computer servers, computer terminals, blade servers, etc. The computing devices 1600 and 1700 described herein can communicate with various types of devices, such as other computing devices 1600 and 1700, which may sometimes act as repeaters or may be configured to work together as a computer cluster to perform high-performance computing.

[0165] All methods described herein are presented with possible implementations. Operations and steps can be rearranged or otherwise modified, and other implementations may exist. Furthermore, if applicable, aspects of two or more methods can be appropriately combined.

[0166] The information and signals described herein can be represented using a variety of different techniques and methods. For example, data, instructions, commands, information, signals, bits, symbols, and chips mentioned throughout the description can be represented by voltage, current, electromagnetic waves, magnetic fields or particles, light fields or particles, or any combination thereof.

[0167] The various exemplary modules and components described herein can be implemented using general-purpose processors, DSPs, ASICs, CPUs, FPGAs or other programmable logic devices, discrete gate or transistor logic, discrete hardware components, or any combination thereof. A general-purpose processor can be a microprocessor, or any processor, controller, microcontroller, or state machine. Processors can also be implemented using combinations of computing devices (e.g., a combination of a DSP and a microprocessor, a combination of multiple microprocessors, a combination of one or more microprocessors with a DSP core, or any other such configuration).

[0168] The functions described herein can be implemented using hardware, processor-executed software, firmware, or any combination thereof. If implemented using processor-executed software, these functions can be stored on a computer-readable medium or transmitted as one or more instructions or code. Other examples and implementations are within the scope of this disclosure and the appended claims. For example, due to the nature of software, the functions described herein can be implemented using processor-executed software, hardware, firmware, hardwiring, or any combination thereof. Features implementing the functions can also be physically located in different locations, including being distributed in different physical locations such that some functions are implemented in different physical locations.

[0169] Computer-readable media include non-transitory computer storage media and communication media, the latter including any medium capable of facilitating the transfer of a computer program from one place to another. Non-transitory storage media can be any available medium accessible to a general-purpose computer or a special-purpose computer. By way of example and not limitation, non-transitory computer-readable media can include RAM, ROM, electrically erasable programmable ROM (EEPROM), flash memory, optical disc (CD) ROM or other optical disc storage devices, magnetic disk storage devices or other magnetic storage devices, or any other non-transitory medium that can be used to carry or store required program code in the form of instructions or data structures, and such media can be accessed by a general-purpose computer or a special-purpose computer, or a general-purpose processor or a special-purpose processor. Furthermore, any connection can be appropriately referred to as computer-readable media. For example, if software is transmitted from a website, server, or other remote source via coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technology (such as infrared, radio, and microwave), then that coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technology (such as infrared, radio, and microwave) is included within the definition of computer-readable media. The disks and optical discs used in this article include CDs, laser discs, optical discs, DVDs, floppy disks, and Blu-ray discs. Disks typically copy data magnetically, while optical discs use lasers to copy data optically. Combinations of these media are also included within the scope of computer-readable media.

[0170] As used herein (including the claims), “or” when used to enumerate items (e.g., a list of items beginning with phrases such as “at least one” or “one or more”) indicates an inclusive list; for example, enumerating at least one of A, B, or C means A or B or C or AB or AC or BC or ABC (i.e., A and B and C). Furthermore, as used herein, the term “based on” should not be construed as a reference to a closed set of conditions. For example, an exemplary step described as “based on condition A” may be based on both condition A and condition B simultaneously without departing from the scope of this disclosure. In other words, as used herein, the term “based on” should be interpreted in the same manner as the term “at least partially based on”.

[0171] In the accompanying drawings, similar parts or features may have the same reference numerals. Furthermore, different components of the same type can be distinguished by adding a hyphen and a second reference numeral after the reference numeral. If only the first reference numeral is used in the description, the description applies to any similar components having the same first reference numeral, regardless of the second or other subsequent reference numerals.

[0172] This document describes exemplary configurations in conjunction with the accompanying drawings, but does not represent all implementable examples or all examples within the scope of the claims. The term "example" as used herein means "as an example, instance, or illustration," and not "preferred" or "superior to other examples." The detailed description includes specific details intended to aid in understanding the described techniques. However, these techniques can be implemented even without these specific details. In some cases, known structures and apparatuses are shown in block diagram form to avoid obscuring the concept of the described examples.

[0173] The description herein is intended to enable those skilled in the art to make or use the contents of this disclosure. Various modifications can be made to the contents of this disclosure by those skilled in the art, and the general principles defined herein can be applied to other variations without departing from the scope of this disclosure. Therefore, the contents of this disclosure are not limited to the examples and designs described herein, but should have the broadest scope consistent with the principles and novel features disclosed herein.

[0174] Example The following examples are disclosed according to various aspects of the present invention.

[0175] Example 1: A computer implementation method for generating movie videos using a game engine, the method comprising: determining feature vectors of each frame of a three-dimensional (3D) animation video, thematic semantic information of the 3D animation, semantic information of an estimated emotional state conveyed by the 3D animation, and an embedded representation of camera poses in the 3D animation at a first plurality of time steps, wherein the camera poses are associated with at least two characters in corresponding frames of the video, and each feature vector characterizes the 3D spatial relationship between the two characters, and wherein the embedded representation is a toroidal coordinate representation of the camera pose; based on the feature vectors, performing temporal processing on keyframes of the 3D animation to obtain keyframe feature vectors for each keyframe, wherein each keyframe defines a key moment in the 3D animation and is configured to be located at the center of a plurality of frames within a local window in the 3D animation, and wherein each keyframe feature vector is configured to contain... The process includes: including the contextual information of the local window; processing the semantic information of the subject matter of the 3D animation and the semantic information of the estimated emotional state through a first feedforward neural network to obtain a first embedding representation set; and processing the one-hot vector embedding representation of the camera pose and the keyframe feature vector through a second feedforward neural network to obtain a second embedding representation set associated with the keyframe at a second plurality of time steps, wherein the one-hot vector embedding representation is an embedding representation based on the camera pose; processing the first and second embedding representation sets through an autoregressive transformer to obtain a third embedding representation set of the predicted camera pose associated with the two characters at a second plurality of time steps, wherein the first plurality of time steps precede the second plurality of time steps in temporal order; and generating the corresponding conditional probability distribution of the predicted camera pose and multiple probabilities of the keyframes marking the shot boundaries based on the third embedding representation set through multiple decoders.

[0176] Example 2: The method is the same as in Example 1, where the local window contains multiple frames, including 5 frames.

[0177] Example 3: The method of any of Examples 1 to 2, wherein each keyframe feature vector is configured as an 1123-dimensional vector.

[0178] Example 4: The method of any of Examples 1 to 3, wherein temporal processing of keyframes to obtain keyframe feature vectors includes: processing the determined feature vectors using a set of 1D convolutions along the time axis.

[0179] Example 5: The method of any of Examples 1 to 4, where the two characters are selected from multiple characters in a 3D animated video.

[0180] Example 6: The method of any of Examples 1 to 4, wherein the two characters include a character in a 3D animated video and a virtual character introduced as a reference character.

[0181] Example 7: The method, as in any of Examples 1 to 6, also includes receiving a 3D animated video as input.

[0182] Example 8: The method of any of Examples 1 to 7, wherein multiple first feedforward neural networks and multiple second feedforward neural networks are implemented using a multilayer perceptron (MLP).

[0183] Example 9: A method as in any of Examples 1 through 8, wherein multiple decoders include a multilayer perceptron (MLP) configured with a softmax output layer for class probability prediction, and an MLP configured with a sigmoid output layer.

[0184] Example 10: A method as described in any of Examples 1 through 9, where the camera pose is constrained by an embedding representation defined as follows: , in, For camera posture, and For these two roles The 2D screen position of the head, and is the angle in the toroidal coordinate system.

[0185] Example 11: The method of any of Examples 1 to 10 further includes: referring to the conditional probability distribution of the predicted camera pose obtained at the time step, providing a subset of the predicted camera pose at the time step based on the order likelihood value of a subset of the predicted camera pose.

[0186] Example 12: The method of any of Examples 1 to 11, wherein the semantic information of the subject matter of the 3D animation and the semantic information of the estimated emotional state are encoded into their respective one-hot vectors before being processed by the first feedforward neural network.

[0187] Example 13: The method of any one of Examples 1 to 12, wherein processing the one-hot vector embedding representation of the camera pose and the keyframe feature vector by the second feedforward neural network to obtain the second embedding representation set includes, for each time step: processing the one-hot vector of the corresponding camera pose by the third feedforward neural network; concatenating the processed one-hot vector of the corresponding camera pose with the keyframe feature vector at that time step to obtain the concatenated embedding representation; and providing the concatenated embedding representation to the feedforward neural network in the second feedforward neural network for processing.

[0188] Example 14: The method as in any of Examples 1 through 13, where each feature vector is defined as ,in, These two characters 3D distance between heads , It is a connecting role and roles The angle formed by the first line of the head and the second line connecting the shoulders.

[0189] Example 15: A method similar to any of Examples 1 through 14, where each conditional probability distribution is configured to be referenced across multiple quantization categories.

[0190] Example 16: A computing device for generating movie videos using a game engine, comprising: one or more memories storing executable code; and one or more processors coupled to the one or more memories and configured to execute the code to cause the device to: determine feature vectors of individual frames of a three-dimensional (3D) animated video, semantic information of the subject matter of the 3D animation and semantic information of an estimated emotional state conveyed by the 3D animation, and an embedded representation of camera poses in the 3D animation at first plurality of time steps, wherein the camera poses are associated with at least two characters in corresponding frames of the video, and each feature vector characterizes a 3D spatial relationship between the two characters, and wherein the embedded representation is a toroidal coordinate representation of the camera poses; and based on the feature vectors, perform temporal processing on keyframes of the 3D animation to obtain keyframe feature vectors for each keyframe, wherein each keyframe defines a key moment in the 3D animation and is configured to be located in a local window in the 3D animation. The process involves: determining the center of multiple frames, wherein each keyframe feature vector is configured to include contextual information of a local window; processing semantic information of the subject matter of the 3D animation and semantic information of the estimated emotional state through a first feedforward neural network to obtain a first embedding representation set; processing one-hot vector embedding representations of camera pose and keyframe feature vectors through a second feedforward neural network to obtain a second embedding representation set associated with keyframes at a second plurality of time steps, wherein the one-hot vector embedding representations are embedding representations based on camera pose; processing the first and second embedding representation sets through an autoregressive transformer to obtain a third embedding representation set of predicted camera poses associated with the two characters at a second plurality of time steps, wherein the first plurality of time steps precede the second plurality of time steps in temporal order; and generating a corresponding conditional probability distribution of the predicted camera pose based on the third embedding representation set through multiple decoders, as well as multiple probabilities of keyframes marking shot boundaries.

[0191] Example 17: A computing device for generating movie videos using a game engine, comprising: means for determining feature vectors of individual frames of a three-dimensional (3D) animation video, semantic information of the subject matter of the 3D animation and semantic information of an estimated emotional state conveyed by the 3D animation, and an embedded representation of camera poses in the 3D animation at first plurality of time steps, wherein the camera poses are associated with at least two characters in corresponding frames of the video, and each feature vector characterizes a 3D spatial relationship between the two characters, and wherein the embedded representation is a toroidal coordinate representation of the camera pose; means for performing temporal processing on keyframes of the 3D animation based on the feature vectors to obtain keyframe feature vectors for each keyframe, wherein each keyframe defines a key moment in the 3D animation and is configured to be located at the center of a plurality of frames within a local window in the 3D animation, and wherein each keyframe feature vector is configured to include a local window. The apparatus includes: contextual information of the mouth; means for processing semantic information of the subject matter of the 3D animation and semantic information of the estimated emotional state through a first feedforward neural network to obtain a first embedding representation set; and means for processing one-hot vector embedding representations of camera pose and keyframe feature vectors through a second feedforward neural network to obtain a second embedding representation set associated with keyframes at a second plurality of time steps, wherein the one-hot vector embedding representations are embedding representations based on camera pose; means for processing the first and second embedding representation sets through an autoregressive transformer to obtain a third embedding representation set of predicted camera poses associated with the two characters at a second plurality of time steps, wherein the first plurality of time steps precede the second plurality of time steps in temporal order; and means for generating a corresponding conditional probability distribution of predicted camera poses and a plurality of probabilities of keyframes marking shot boundaries based on the third embedding representation set through a plurality of decoders.

[0192] Example 18: A non-transitory computer-readable medium comprising executable code that, when executed by a processor of a computing device, causes the device to perform a method as described in any one of Examples 1 to 15.

[0193] References [1] H Porter Abbott. 2002. The Cambridge introduction to narrative.Cambridge University Press. [2] Shivam Sharma Karan Bilakhiya Aman Gupta, Amey Band. 2020.text2emotion. https: / / pypi.org / project / text2emotion / [3] Hideaki Anno. 2021. Evangelion: 3.0+1.01 Thrice Upon Time[Film]. IMDb (2021). [4] Ido Arev, Hyun Soo Park, Yaser Sheikh, Jessica Hodgins, and Ariel Shamir. 2014. Automatic Editing of Footage from Multiple Social Cameras. ACMTrans. Graph. 33, 4, Article 81 (July 2014), 11 pages. https: / / doi.org / 10.1145 / 2601097.2601198 [5] Daniel Arizona. 1991. Grammar of the film language. Silman-JamesPress. [6] J. Aumont, University of Texas Press, A. Bergala, M. Marie, R. Neupert, and M. Vernet. 1992. The Aesthetics of Film. University of Texas Press.https: / / books.google.com.hk / books?id=nmpZAAAAMAAJ [7] Aaron Bangor, Philip T Kortum, and James T Miller. 2008. An empirical evaluation of the system usability scale. Intl. Journal of Human-Computer Interaction 24, 6 (2008), 574-594. [8] Marcel Berger. 2009. Geometry i. Springer Science & BusinessMedia. [9] Nathalie Bonnardel. 2000. Towards understanding and supportingcreativity in design: analogies in a constrained cognitive environment.Knowledge-Based Systems 13, 7-8 (2000), 505-513.

[10] Christopher Bowen. 2013. Grammar of the Shot. Routledge.

[11] Joseph Campbell. 2008. The hero with a thousand faces. Vol. 17.New World Library.

[12] Marc Christie, Patrick Olivier, and Jean-Marie Normand. 2008.Camera control in computer graphics. In Computer Graphics Forum, Vol. 27.Wiley Online Library, 2197-2218.

[13] Nicholas Davis, Boyang Li, Brian O’Neill, Mark Riedl, andMichael Nitsche. 2011. Distributed Creative Cognition in Digital Filmmaking.In Proceedings of the 8th ACM Conference on Creativity and Cognition(Atlanta, Georgia, USA) (C&C ’11). Association for Computing Machinery, NewYork, NY, USA, 207-216. https: / / doi.org / 10.1145 / 2069618.2069654

[14] Nicholas Davis, Alexander Zook, Brian O’Neill, Brandon Headrick,Mark Riedl, Ashton Grosz, and Michael Nitsche. 2013. Creativity support fornovice digital filmmaking. In Proceedings of the SIGCHI Conference on HumanFactors in Computing Systems. 651-660.

[15] Edirlei E. S. de Lima, Cesar T. Pozzer, Marcos C. d’Ornellas,Angelo E. M. Ciarlini, Bruno Feijó, and Antonio L. Furtado. 2009. VirtualCinematography Director for Interactive Storytelling. In Proceedings of theInternational Conference on Advances in Computer Entertainment Technology(Athens, Greece) (ACE ’09). Association for Computing Machinery, New York,NY, USA, 263-270. https: / / doi.org / 10.1145 / 1690388.1690432

[16] Will Eisner. 2008. Graphic storytelling and visual narrative:Principles and practices from the legendary cartoonist (Rev. ed). NY: Norton.(Original work published 1996) (2008).

[17] Inan Evin, Perttu Hämäläinen, and Christian Guckelsberger. 2022.Cine-AI: Generating Video Game Cutscenes in the Style of Human Directors.Proceedings of the ACM on Human-Computer Interaction 6, CHI PLAY (2022), 1-23.

[18] HJ Eysenck. 2003. Creativity, personality and the convergent-divergent continuum. (2003).

[19] Gerhard Fischer. 2004. Social creativity: turning barriers intoopportunities for collaborative design. In Proceedings of the eighthconference on Participatory design: Artful integration: interweaving media,materials and practices-Volume 1. 152-161.

[20] Jonas Frich, Lindsay MacDonald Vermeulen, Christian Remy,Michael Mose Biskjaer, and Peter Dalsgaard. 2019. Mapping the landscape ofcreativity support tools in HCI. In Proceedings of the 2019 CHI Conference onHuman Factors in Computing Systems. 1-18.

[21] Jonas Frich, Michael Mose Biskjaer, and Peter Dalsgaard. 2018.Twenty years of creativity research in human-computer interaction: Currentstate and future directions. In Proceedings of the 2018 Designing InteractiveSystems Conference. 1235-1257.

[22] Quentin Galvane and Rémi Ronfard. 2017. Implementing hitchcock-the role of focalization and viewpoint. In Eurographics Workshop onIntelligent Cinematography and Editing. The Eurographics Association.

[23] Francis Glebas. 2012. Directing the story: professionalstorytelling and storyboarding techniques for live action and animation.Routledge.

[24] Joy Paul Guilford. 1967. The nature of human intelligence.(1967).

[25] Li-wei He, Michael F Cohen, and David H Salesin. 1996. Thevirtual cinematographer: A paradigm for automatic real-time camera controland directing. In Proceedings of the 23rd annual conference on Computergraphics and interactive techniques. 217-224.

[26] Qingqiu Huang, Yu Xiong, Anyi Rao, JiazeWang, and Dahua Lin.2020. Movienet: A holistic dataset for movie understanding. In ComputerVision-ECCV 2020: 16th European Conference, Glasgow, UK, August 23-28, 2020,Proceedings, Part IV 16. Springer, 709-727.

[27] David G Jansson and Steven M Smith. 1991. Design fxation. Designstudies 12, 1 (1991), 3-11.

[28] Arnav Jhala and R Michael Young. 2011. Intelligent machinimageneration for visual storytelling. In Artifcial Intelligence for ComputerGames. Springer, 151-170.

[29] Hongda Jiang, Bin Wang, Xi Wang, Marc Christie, and BaoquanChen. 2020. Example-driven virtual cinematography by learning camerabehaviors. ACM Transactions on Graphics (TOG) 39, 4 (2020), 45-1.

[30] Steven Douglas Katz. 1991. Film directing shot by shot:visualizing from concept to screen. Gulf Professional Publishing.

[31] Arthur Koestler. 1964. The act of creation. (1964).

[32] Aki Kubota. 2021. Hideaki Anno: The Final Challenge ofEvangelion [Film]. NHK (2021).

[33] Nicole E Lemon. 2012. Previsualization in Computer AnimatedFilmmaking. Ph.D. Dissertation. The Ohio State University.

[34] Christophe Lino and Marc Christie. 2015. Intuitive and efcientcamera control with the toric space. ACM Transactions on Graphics (TOG) 34, 4(2015), 1-12.

[35] Christophe Lino, Marc Christie, Fabrice Lamarche, Schofeld Guy,and Patrick Olivier. 2010. A real-time cinematography system for interactive3d environments. In SCA’10 Proceedings of the 2010 ACM SIGGRAPH / EurographicsSymposium on Computer Animation. 139-148.

[36] Henry Lowood and Michael Nitsche. 2011. The machinima reader.MIT Press.

[37] Shuai Ma, Xinru Wang, Ying Lei, Chuhan Shi, Ming Yin, andXiaojuan Ma. 2024. “Are You Really Sure?” Understanding the Efects of HumanSelf-Confdence Calibration in AI-Assisted Decision Making. In Proceedings ofthe CHI Conference on Human Factors in Computing Systems (Honolulu, HI, USA)(CHI ’24). Association for Computing Machinery, New York, NY, USA, Article840, 20 pages. https: / / doi.org / 10.1145 / 3613904.3642671

[38] James Mairata, Mairata, and Aboujieb. 2018. Steven Spielberg’sStyle by Stealth. Springer.

[39] Paul Marino. 2004. 3D game-based flmmaking: The art ofmachinima. Paraglyph Press.

[40] Marcos Mateu-Mestre and Jefrey Katzenberg. 2010. Framed ink:Drawing and composition for visual storytellers. Design Studio Press.

[41] R. McKee. 1997. Story: Substance, Structure, Style and thePrinciples of Screenwriting. HarperCollins Publishers.

[42] Manshad Abbasi Mohsin and Anatoly Beltiukov. 2019. Summarizingemotions from text using Plutchik’s wheel of emotions. In 7th ScientifcConference on Information Technologies for Intelligent Decision MakingSupport (ITIDS 2019). Atlantis Press, 291-294.

[43] K. L. Bhanu Moorthy, Moneish Kumar, Ramanathan Subramanian, andVineet Gandhi. 2020. GAZED-Gaze-Guided Cinematic Editing of Wide-AngleMonocular Video Recordings. In Proceedings of the 2020 CHI Conference onHuman Factors in Computing Systems (Honolulu, HI, USA) (CHI ’20). Associationfor Computing Machinery, New York, NY, USA, 1-11. https: / / doi.org / 10.1145 / 3313831.3376544

[44] Kumiyo Nakakoji. 2006. Meanings of tools, support, and uses forcreative design processes. In International design research symposium, Vol.6. 156-165.

[45] Mark A Runco. 2014. Creativity theories and themes: research,development, and practice. (2014).

[46] Mark A Runco and Garrett J Jaeger. 2012. The standard defnitionof creativity. Creativity research journal 24, 1 (2012), 92-96.

[47] István Sárándi, Timm Linder, Kai Oliver Arras, and BastianLeibe. 2020. Metrabs: metric-scale truncation-robust heatmaps for absolute 3dhuman pose estimation. IEEE Transactions on Biometrics, Behavior, andIdentity Science 3, 1 (2020), 16-30.

[48] Ben Shneiderman. 2002. Creativity support tools. Commun. ACM 45,10 (2002), 116-120.

[49] T. Sobchack and V.C. Sobchack. 1987. An Introduction to Film.Little, Brown. https: / / books.google.com.hk / books?id=SwMbAQAAIAAJ

[50] EDP Symons. 1986. Edison’s electric light. Biography of aninvention.

Claims

1. A computer-based method for generating movie videos using a game engine, the method comprising: The method determines feature vectors of each frame of a three-dimensional (3D) animation video, semantic information of the subject matter of the 3D animation and semantic information of the estimated emotional state conveyed by the 3D animation, and an embedded representation of the camera pose in the 3D animation at a first plurality of time steps, wherein the camera pose is associated with at least two characters in the corresponding frame of the video, and each feature vector characterizes the 3D spatial relationship between the two characters, and wherein the embedded representation is a toroidal coordinate representation of the camera pose; Based on the feature vector, time processing is performed on the keyframes of the 3D animation to obtain the keyframe feature vectors of each keyframe. Each keyframe defines a key moment in the 3D animation and is configured to be located at the center of multiple frames within a local window in the 3D animation. Each keyframe feature vector is configured to include the context information of the local window. The semantic information of the subject matter of the 3D animation and the semantic information of the estimated emotional state are processed by a first feedforward neural network to obtain a first embedding representation set. The one-hot vector embedding representation of the camera pose and the key frame feature vector are processed by a second feedforward neural network to obtain a second embedding representation set associated with the key frame at a second plurality of time steps. The one-hot vector embedding representation is an embedding representation based on the camera pose. The first and second embedded representation sets are processed by an autoregressive transform to obtain a third embedded representation set of predicted camera poses associated with the two roles at a second plurality of time steps, wherein the first plurality of time steps precede the second plurality of time steps in temporal order; and The corresponding conditional probability distribution of the predicted camera pose, as well as multiple probabilities of the keyframes marking the lens boundaries, are generated by multiple decoders based on a third embedding representation set.

2. The method according to claim 1, wherein, The local window contains multiple frames, including 5 frames.

3. The method according to any one of claims 1 to 2, wherein, Each keyframe feature vector is configured as an 1123-dimensional vector.

4. The method according to any one of claims 1 to 3, wherein, Temporal processing of the keyframes to obtain the keyframe feature vectors includes processing the determined feature vectors using a set of 1D convolutions along the time axis.

5. The method according to any one of claims 1 to 4, wherein, The two characters are selected from a plurality of characters in the 3D animated video.

6. The method according to any one of claims 1 to 4, wherein, The two roles include the characters in the 3D animated video, and the virtual characters introduced as reference characters.

7. The method according to any one of claims 1 to 6, further comprising: The video of the 3D animation is received as input.

8. The method according to any one of claims 1 to 7, wherein, Multiple first feedforward neural networks and multiple second feedforward neural networks are implemented using a multilayer perceptron (MLP).

9. The method according to any one of claims 1 to 8, wherein, The plurality of decoders include a multilayer perceptron (MLP) configured with a softmax output layer for class probability prediction, and an MLP configured with a sigmoid output layer.

10. The method according to any one of claims 1 to 9, wherein, The camera pose is defined by the following embedding representation: , in, The camera pose, and For the two roles The 2D screen position of the head, and Let be the angle in the toroidal coordinate system.

11. The method according to any one of claims 1 to 10, further comprising: Referring to the conditional probability distribution of the predicted camera pose obtained at a time step, the subset of the predicted camera pose at the time step is provided based on the ranking likelihood value of the subset of the predicted camera pose.

12. The method according to any one of claims 1 to 11, wherein, Before being processed by the first feedforward neural network, the semantic information of the subject matter of the 3D animation and the semantic information of the estimated emotional state are encoded into their respective one-hot vectors.

13. The method according to any one of claims 1 to 12, wherein, The second embedding representation set is obtained by processing the one-hot vector embedding representation of the camera pose and the keyframe feature vectors through a second feedforward neural network, including, for each time step: The one-hot vector corresponding to the camera pose is processed by a third feedforward neural network; The processed one-hot vector of the corresponding camera pose is concatenated with the keyframe feature vector at the time step to obtain a concatenated embedded representation; and The concatenated embedding representation is provided to the feedforward neural network in the second feedforward neural network for processing.

14. The method according to any one of claims 1 to 13, wherein, Each feature vector is defined as ,in, The two roles mentioned 3D distance between heads , It is a connecting role and roles The angle formed by the first line of the head and the second line connecting the shoulders.

15. The method according to any one of claims 1 to 14, wherein, Each conditional probability distribution is configured to be referenced across multiple quantization categories.

16. A computing device for generating movie videos using a game engine, comprising: One or more memories that store executable code; as well as One or more processors, coupled to the one or more memories and configured to execute the code to enable the device: The method determines feature vectors of each frame of a three-dimensional (3D) animation video, semantic information of the subject matter of the 3D animation and semantic information of the estimated emotional state conveyed by the 3D animation, and an embedded representation of the camera pose in the 3D animation at a first plurality of time steps, wherein the camera pose is associated with at least two characters in the corresponding frame of the video, and each feature vector characterizes the 3D spatial relationship between the two characters, and wherein the embedded representation is a toroidal coordinate representation of the camera pose; Based on the feature vector, time processing is performed on the keyframes of the 3D animation to obtain the keyframe feature vectors of each keyframe. Each keyframe defines a key moment in the 3D animation and is configured to be located at the center of multiple frames within a local window in the 3D animation. Each keyframe feature vector is configured to include the context information of the local window. The semantic information of the subject matter of the 3D animation and the semantic information of the estimated emotional state are processed by a first feedforward neural network to obtain a first embedding representation set. The one-hot vector embedding representation of the camera pose and the key frame feature vector are processed by a second feedforward neural network to obtain a second embedding representation set associated with the key frame at a second plurality of time steps. The one-hot vector embedding representation is an embedding representation based on the camera pose. The first and second embedded representation sets are processed by an autoregressive transform to obtain a third embedded representation set of predicted camera poses associated with the two roles at a second plurality of time steps, wherein the first plurality of time steps precede the second plurality of time steps in temporal order; and The corresponding conditional probability distribution of the predicted camera pose, as well as multiple probabilities of the keyframes marking the lens boundaries, are generated by multiple decoders based on a third embedding representation set.

17. A computing device for generating movie videos using a game engine, comprising: A means for determining feature vectors of individual frames of a video of a three-dimensional (3D) animation, semantic information of the subject matter of the 3D animation and semantic information of the estimated emotional state conveyed by the 3D animation, and an embedded representation of camera pose in the 3D animation at a first plurality of time steps, wherein the camera pose is associated with at least two characters in corresponding frames of the video, and each feature vector characterizes a 3D spatial relationship between the two characters, and wherein the embedded representation is a toroidal coordinate representation of the camera pose; An apparatus for performing time processing on keyframes of a 3D animation based on the feature vectors to obtain keyframe feature vectors for each keyframe, wherein each keyframe defines a key moment in the 3D animation and is configured to be located at the center of multiple frames within a local window in the 3D animation, and wherein each keyframe feature vector is configured to include context information of the local window. An apparatus for processing semantic information of the subject matter of the 3D animation and semantic information of the estimated emotional state through a first feedforward neural network to obtain a first embedding representation set, and for processing one-hot vector embedding representation of the camera pose and keyframe feature vectors through a second feedforward neural network to obtain a second embedding representation set associated with the keyframe at a second plurality of time steps, wherein the one-hot vector embedding representation is an embedding representation based on the camera pose. A means for processing a first and a second set of embedded representations via an autoregressive transform to obtain a third set of embedded representations of predicted camera poses associated with the two roles at a second plurality of time steps, wherein the first plurality of time steps precede the second plurality of time steps in temporal order; and An apparatus for generating, via multiple decoders, a corresponding conditional probability distribution of the predicted camera pose and multiple probabilities of the keyframes marking the lens boundaries, based on a third embedding representation set.

18. A non-transitory computer-readable medium comprising executable code that, when executed by a processor of a computing device, causes the device to perform the method of any one of claims 1 to 15.