Video generation method, computer readable medium, electronic device and program product
By obtaining video materials and determining rendering parameters, calling video template files to render video content on the server side, solving the problem of quickly generating high-quality cross-platform videos, improving generation efficiency and saving client resources.
Patent Information
- Application Number
- CN202510830601.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-19
- Publication Date
- 2025-08-29
AI Technical Summary
How to quickly generate high-quality videos and improve video generation efficiency, especially maintain consistent video effects across platforms while avoiding consuming client device resources.
By obtaining video materials, determining the parameters required for video rendering, calling the video template file, rendering the video frames of the video content display process based on these parameters, generating the target video, and rendering it on the server side to output the video.
The consistency of cross-platform video effects is achieved, the video generation efficiency is significantly improved, and the use of client resources is avoided.
Smart Images

Figure CN120568099A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer technology, and in particular, to a video generation method, a computer-readable medium, an electronic device, and a program product. Background Art
[0002] With the development of Internet technology, especially video content platforms, users' demand for sharing personalized content in the form of videos is also increasing. How to quickly generate high-quality videos has become a technical problem that needs to be solved urgently. Summary of the Invention
[0003] This summary is provided to briefly introduce concepts that will be described in detail in the detailed description below. This summary is not intended to identify key features or essential features of the claimed technical solution, nor is it intended to limit the scope of the claimed technical solution.
[0004] In a first aspect, the present disclosure provides a video generation method, comprising:
[0005] Get video material;
[0006] Determining video rendering parameters required for video rendering based on the video material;
[0007] Calling a video template file, and generating a target video for displaying the video material according to the video rendering parameters, wherein the video template file is used to define a video content display process, and the video template file renders the video content of the video frame corresponding to each video content display process according to the video rendering parameters;
[0008] The target video is output.
[0009] In a second aspect, the present disclosure provides a video generation method, which is executed by a client, and the method includes:
[0010] In response to a material selection operation, determining the video material indicated by the material selection operation;
[0011] Display the target video, wherein the target video is generated according to the video rendering parameters required for video rendering based on the video material, and a video template file is called, wherein the video template file is used to define the video content display process, and the video template file renders the video content of the video frames corresponding to each video content display process according to the video rendering parameters.
[0012] In a third aspect, the present disclosure provides a computer-readable medium having a computer program stored thereon, which, when executed by a processing device, implements the steps of the method described in the first aspect, or implements the steps of the method described in the second aspect.
[0013] In a fourth aspect, the present disclosure provides an electronic device, comprising:
[0014] a storage device having a computer program stored thereon;
[0015] A processing device is used to execute the computer program in the storage device to implement the steps of the method described in the first aspect, or to implement the steps of the method described in the second aspect.
[0016] In a fifth aspect, the present disclosure provides a computer program product, comprising a computer program, which, when executed by a processor, implements the steps of the method described in the first aspect, or implements the steps of the method described in the second aspect.
[0017] Based on the above technical solution, by obtaining video material, determining the video rendering parameters required for video rendering based on the video material, and then calling the video template file used to define the video content display process, rendering the video content of the video frames corresponding to each video content display process according to the video rendering parameters, generating a target video for displaying the video material, and then outputting the target video, it is possible to achieve batch generation of videos with consistent effects across multiple terminals by encoding the video template file once, significantly improving video generation efficiency. Moreover, video rendering through the server does not occupy client device resources.
[0018] Other features and advantages of the present disclosure will be described in detail in the following detailed description. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] The above and other features, advantages and aspects of the various embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. Throughout the drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic and that the originals and elements are not necessarily drawn to scale. In the drawings:
[0020] Figure 1 The figure is a flowchart of a video generating method according to an exemplary embodiment.
[0021] Figure 2 The figure is a flowchart of generating a target video according to an exemplary embodiment.
[0022] Figure 3 The diagram is an architecture diagram of a video rendering service according to an exemplary embodiment.
[0023] Figure 4The figure is a timing diagram showing the operation of a video rendering service according to an exemplary embodiment.
[0024] Figure 5 The diagram is a structural diagram of a video template file according to an exemplary embodiment.
[0025] Figure 6 The figure is a schematic diagram showing data flow according to an exemplary embodiment.
[0026] Figure 7 FIG. 4 is a flowchart of generating a target video according to another exemplary embodiment.
[0027] Figure 8 is a flowchart of a video generating method according to another exemplary embodiment.
[0028] Figure 9 FIG. 4 is an architecture diagram of a video generation system according to an exemplary embodiment.
[0029] Figure 10 It is a logic timing diagram of a video generating method according to an exemplary embodiment.
[0030] Figure 11 The diagram is an architecture diagram of a video rendering service instance according to an exemplary embodiment.
[0031] Figure 12 is a flowchart of a video generating method according to yet another exemplary embodiment.
[0032] Figure 13 is a flowchart of a video generating method according to another exemplary embodiment.
[0033] Figure 14 is a schematic diagram of an editing interface according to an exemplary embodiment.
[0034] Figure 15 The figure is a schematic structural diagram of a video generating apparatus according to an exemplary embodiment.
[0035] Figure 16 FIG. 4 is a schematic structural diagram of a video generating apparatus according to another exemplary embodiment.
[0036] Figure 17 FIG. 4 is a schematic structural diagram of a video generating apparatus according to yet another exemplary embodiment.
[0037] Figure 18 It is a schematic structural diagram of an electronic device according to an exemplary embodiment. DETAILED DESCRIPTION
[0038] The following describes embodiments of the present disclosure in more detail with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments described herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are for illustrative purposes only and are not intended to limit the scope of protection of the present disclosure.
[0039] It should be understood that the various steps described in the method embodiments of the present disclosure may be performed in different orders and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this respect.
[0040] As used herein, the term "including" and its variations are open-ended, i.e., "including but not limited to." The term "based on" means "based, at least in part, on." The term "one embodiment" means "at least one embodiment," the term "another embodiment" means "at least one additional embodiment," and the term "some embodiments" means "at least some embodiments." Other terms are defined in the following description.
[0041] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.
[0042] It should be noted that the modifications of "one" and "multiple" mentioned in the present disclosure are illustrative rather than restrictive, and those skilled in the art should understand that unless otherwise clearly indicated in the context, they should be understood as "one or more".
[0043] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only used for illustrative purposes and are not used to limit the scope of these messages or information.
[0044] It is understandable that before using the technical solutions disclosed in the various embodiments of this disclosure, the type, scope of use, usage scenarios, etc. of the personal information involved in this disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.
[0045] For example, in response to a user's active request, a prompt message is sent to the user to clearly inform the user that the operation requested will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the electronic device, application, server, storage medium, or other software or hardware that performs the operations of the disclosed technical solution based on the prompt message.
[0046] As an optional but non-limiting implementation, in response to receiving a user's active request, the prompt information may be sent to the user in the form of a pop-up window, in which the prompt information may be presented in text form. Furthermore, the pop-up window may also contain a selection control for the user to select "agree" or "disagree" to provide personal information to the electronic device.
[0047] It is understandable that the above notification and user authorization process are merely illustrative and do not limit the implementation of the present disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of the present disclosure.
[0048] At the same time, it is understandable that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) must comply with the requirements of relevant laws, regulations and relevant provisions.
[0049] Figure 1 FIG. 1 is a flow chart of a video generation method according to an exemplary embodiment. Figure 1 As shown, the embodiment of the present disclosure provides a video generation method, which can be executed by a server. The method can be specifically executed by a video generation device, which can be implemented by software and / or hardware, and the device can be configured in the server. Figure 1 As shown, the method may include the following steps.
[0050] In step 110, video material is obtained.
[0051] Here, the video material can be sent by the client. Exemplarily, the video material includes at least the conversation content between the user and the virtual object. Among them, the virtual object can be a virtual character driven by artificial intelligence (AI). For example, the virtual object can be an AI virtual character with a specific role setting, and the user can communicate with the AI virtual character through text, voice, and image, and experience the interactive process from unfamiliar to familiar. Accordingly, the conversation content between the user and the virtual object may include but is not limited to text content, image content, voice content, etc. generated by the user and the virtual object. Among them, the image content can be pictures and / or video clips.
[0052] Of course, in other implementations, the video material may also include image information of the virtual object, role information of the virtual object, creator information of the virtual object, and the like.
[0053] In some embodiments, the user may select a video material in a user interface for communicating with a virtual object and send the video material to a server.
[0054] For example, in a scenario of having a conversation with an AI virtual character, the user can select a wonderful conversation with the AI virtual character as video material to instruct the server to generate a video for showing the wonderful conversation between the user and the AI virtual character, so that the user can share the interaction between the user and the AI virtual character through the video.
[0055] In step 120 , video rendering parameters required for video rendering are determined based on the video material.
[0056] Here, the video rendering parameters may include video material and pre-processing parameters obtained based on the video material. Pre-processing parameters may include frame rate (FPS), video encoding format, audio data converted from text content in the video material, audio duration corresponding to the audio data, video duration corresponding to the video data included in the video material, and conversation length corresponding to the conversation content, etc. It should be noted that the video rendering parameters include both all video material required for video rendering and pre-processing parameters calculated from the video material.
[0057] In step 130, the video template file is called to generate a target video for displaying the video material according to the video rendering parameters, wherein the video template file is used to define the video content display process, and the video template file renders the video content of the video frames corresponding to each video content display process according to the video rendering parameters.
[0058] Here, the video content presentation process can refer to the entire process from the beginning to the end of a video. The video template file provides a clear and complete framework for presenting video content, defining the video assets to be displayed during each time period of the video content presentation process, as well as the presentation styles used for these assets. For example, the video content presentation process defined by the video template file may include the video opening, introduction, main content, climax, and ending. The video opening presents a compelling opening image, such as a beautiful title animation or brand logo. The introduction introduces the video's theme or background information, setting the stage for subsequent content, such as displaying relevant images, text, or interview clips. The main content details product features, storylines, and teaching points. The climax brings the video to a climax, emphasizing the most important information or highlights, such as the product's key advantages, a turning point in the story, or the core conclusion of the teaching content. The ending summarizes the video content, such as displaying the brand logo or slogan. Furthermore, the same or different animation effects can be used during each time period of the video content presentation process, all of which can be defined in the video template file.
[0059] It should be noted that the video template file can be understood as a video template code developed through React (a JavaScript library for building user interfaces) components and Remotion (a video creation framework based on React components). Each video template file contains a complete video implementation logic, which defines the video content corresponding to the video to be rendered corresponding to the video template file in each video content display process in the form of code. Each video template file can be reused, and the same video template file can be used to generate videos with different video content but similar video content display processes. Each time the video template file is used, the corresponding video rendering parameters can be replaced, which greatly improves the efficiency of video production. For example, assuming that there is a video content display process for displaying conversation content in the video template file, the video length corresponding to the video content display process will be different for conversation content of different lengths.
[0060] After obtaining the video rendering parameters required for video rendering, the corresponding video template file is called. The video template file renders the video content of the video frames corresponding to each video content display process according to the video rendering parameters, and generates a target video containing video materials.
[0061] For example, in a scenario of having a conversation with an AI virtual character, the user can select a wonderful conversation with the AI virtual character as video material. The server determines the video rendering parameters required for video rendering based on the video material, and calls the video template file to generate a video based on the video rendering parameters to show the wonderful conversation between the user and the AI virtual character, so that the user can share the interaction between the user and the AI virtual character through the video.
[0062] It should be understood that in the embodiments of the present disclosure, the video template file can be pre-created by the developer. The server can pre-store video template files corresponding to different video types, and the user can select the required video template file by selecting the corresponding video type. Of course, the server can also select a video template file corresponding to the video type suitable for the video material sent by the client.
[0063] In other embodiments, the video template file can also be constructed by the user. The user can create a user-specific video template file based on their requirements for the video style. For example, candidate video implementation logic and candidate functional components can be displayed, and then in response to a selection operation, the video template file is constructed based on the video implementation logic and functional components indicated by the selection operation. It should be understood that how to create a user-specific video template file will be described in detail in subsequent embodiments.
[0064] Here, the server can return the target video to the client that sent the video material so that the client can play the target video and / or share the target video. It should be understood that there are various ways for the server to return the target video to the client that sent the video material, and this is not specifically limited in the embodiments of the present disclosure. For example, the server can directly send the target video to the client that sent the video material. For another example, the server can return a download link of the target video to the client that sent the video material so that the client can obtain the target video according to the download link.
[0065] Thus, by obtaining video material, determining the video rendering parameters required for video rendering based on the video material, and then calling the video template file used to define the video content display process, rendering the video content of the video frames corresponding to each video content display process according to the video rendering parameters, generating the target video for displaying the video material, and then outputting the target video, it is possible to achieve batch generation of videos with consistent effects across multiple terminals by encoding the video template file once, significantly improving video generation efficiency. Moreover, video rendering through the server does not occupy client device resources.
[0066] In some feasible implementations, in step 130 , a target video may be generated according to video rendering parameters through a video rendering service, where the video rendering service is configured to generate the target video according to a video template file and video rendering parameters.
[0067] Here, the video rendering service can provide an external interface for external calls to the video rendering service to render the video, where the interface can be an HTTP (Hypertext Transfer Protocol) interface. Within the video rendering service, the video rendering service uses a pre-developed video template file and video rendering parameters to generate the target video, and then returns the generated target video to the caller through the interface.
[0068] In some embodiments, the video rendering service can be deployed in a Function as a Service (FaaS) container.
[0069] The video rendering service can be understood as a service based on Node.js (a cross-platform JavaScript runtime environment). The video rendering service is deployed in the FaaS container and provides an HTTP interface for external calls.
[0070] It should be understood that because the video rendering service is deployed in a FaaS environment and is dedicated to video rendering, without interfacing with other business components, it can be reused across scenarios. Furthermore, deploying the video rendering service in a FaaS environment allows for automatic capacity expansion based on traffic conditions, achieving the most efficient resource utilization. For example, computing resources can be dynamically scaled based on the real-time load of the video rendering service, effectively responding to traffic peaks while also avoiding resource waste during low traffic periods, significantly improving the system's scalability and cost-effectiveness.
[0071] Figure 2 FIG. 1 is a flow chart showing a method for generating a target video according to an exemplary embodiment. Figure 2 As shown, in some possible implementations, the video rendering service can generate a target video through the following steps.
[0072] In step 210, video rendering parameters are obtained.
[0073] The video rendering service includes an interface layer, which includes an HTTP interface. This HTTP interface is used to receive rendering requests containing video rendering parameters and return the rendered target video. It's important to note that the interface layer serves as the entry point for the video rendering service, receiving video rendering parameters and returning the target video.
[0074] In step 220 , metadata information required for video rendering is determined based on the video rendering parameters and the video template file.
[0075] Here, after receiving the video rendering parameters, the video rendering service calculates metadata information required for video rendering based on the received video rendering parameters using a video template file pre-configured by the video rendering service.
[0076] The metadata information includes time metadata corresponding to each video content display process and content data corresponding to the time metadata.
[0077] Temporal metadata can include the start time and duration of each video content presentation process. Temporal metadata actually describes the timeline of the video frame sequence. Through temporal metadata, it can provide global time control data for video rendering and provide precise time references for components. It is important to note that temporal metadata is dynamically determined by the video material included in the video rendering parameters, allowing video template files to adapt to different video materials. For example, different dialogue content and different audio durations will affect the calculated temporal metadata.
[0078] The content data may refer to the video material used by the corresponding time metadata. Through the time metadata and content data, the video frames included in each video content display process and the position of the video material in the video frame can be accurately determined, thereby rendering the video content of each video frame in the target video.
[0079] In step 230 , context information required for video rendering is constructed based on the metadata information and the video rendering parameters.
[0080] Here, after the video rendering service calculates the metadata information, it constructs the context information required for video rendering based on the metadata information and the video rendering parameters. For example, the metadata information and the video rendering parameters can be directly used as the context information required for video rendering.
[0081] It is worth noting that the context information includes both the raw data of the video rendering parameters and the derived data of the calculated metadata information. The context information can be globally accessible so that functional components can access consistent information and reduce data conversion costs.
[0082] In step 240, a rendering engine is called to render the target video according to the functional components and context information indicated by the video template file.
[0083] Here, the video rendering service can use a video generation module to perform video rendering, where the video generation module can be a Remotion service. Specifically, the Remotion service can call a rendering engine, which renders the target video based on the functional components indicated by the video template file and the previously constructed context information. The rendering engine can be a Remotion rendering engine. It should be noted that the video template file includes the functional components used by the video template file, which are React components.
[0084] Figure 3 FIG. 1 is an architecture diagram of a video rendering service according to an exemplary embodiment. Figure 3As shown, the video rendering service includes an interface layer, a controller layer, a service layer, a rendering engine, a component library and a storage layer. An Http interface is provided in the interface layer, and the Http interface is used to obtain video rendering parameters. A rendering controller (RenderController) is provided in the controller layer. The control logic of the video rendering service is in the rendering controller, and the rendering controller is used to coordinate the entire video rendering process. For example, the rendering controller can be used to verify the legitimacy of the rendering request received by the interface, the rendering controller is used to create and manage context information, the rendering controller is used to control the start and cancellation of the video rendering process, the rendering controller is used to feedback the rendering progress to the client, and the rendering controller is used to return the rendered target video to the client and perform error handling.
[0085] The service layer is equipped with a Remotion service, which encapsulates the video rendering logic and is responsible for calling the Remotion interface to perform the actual video rendering work. The rendering controller calls the Remotion service, which receives the video rendering parameters passed in via the HTTP interface. The Remotion service determines the metadata information required for video rendering based on the video rendering parameters and the video template file, and constructs the context information required for video rendering based on the metadata information and the video rendering parameters. In addition, the Remotion service obtains the React components required for video rendering from the component library based on the functional components indicated by the video template file. Specifically, the corresponding React component can be found from the component registry of the component library through the component identifier (ID). The Remotion service then calls the rendering engine and renders the target video based on the functional components and context information indicated by the video template file.
[0086] The rendering engine includes Remotion Core (Remotion's core functionality, used to define video compositions and manage animation states). Remotion Core launches a browser instance, selects the functional components in CompositionSelection (which selects or specifies the container for the video to be rendered), and uses MediaRendering (the process of rendering the defined video content into the final media file) to render the video content corresponding to each frame based on the browser instance and Composition Selection. It then stitches all the video frames together to form the final target video. Of course, the rendering controller also returns the target video to the client through an interface.
[0087] Therefore, by pre-calculating metadata information before video rendering, precise content control at the frame level can be achieved, thereby ensuring the quality of the video image.
[0088] In some feasible implementations, the target video can be uploaded to the video cloud platform through the video rendering service, and a download link for the target video can be generated. Then, the download link can be returned to the client corresponding to the video material through the video rendering service, so that the client can obtain the target video from the video cloud platform according to the download link.
[0089] Here, the video cloud platform can be a cloud storage service for storing the rendered target video. After the video rendering service renders the target video, the video rendering service stores the target video in the video cloud platform.
[0090] like Figure 3 As shown, the service layer of the video rendering service includes a storage service that encapsulates the video storage logic. The rendering controller can call the storage service to store the target video in the object-based video cloud platform within the storage layer. The storage service returns a download link for the target video in the video cloud platform to the rendering controller. The rendering controller then returns the download link to the client through the interface layer, allowing the client to retrieve the target video from the video cloud platform based on the download link.
[0091] Therefore, through the above implementation, bandwidth can be saved by returning the download link to the client.
[0092] The following is combined with Figure 4 Describes the video rendering service in detail.
[0093] Figure 4 FIG. 1 is a timing diagram of a video rendering service according to an exemplary embodiment. Figure 4 As shown, the rendering controller receives rendering requests from the client via the HTTP interface. It verifies the legitimacy of the rendering request and, upon successful verification, sets an SSE (Server-Sent Event) response header for subsequent event notifications to the client. The rendering controller then sends a Started event to the client, informing it that video rendering has begun. The rendering controller then creates a RenderContext. Note that the context created by the rendering controller is blank at this point, preparing the context for video rendering.
[0094] The rendering controller then calls the Remotion service to obtain metadata. The rendering controller can call the Remotion service's getMetadata() method to retrieve metadata. The Remotion service selects a functional component. The Remotion service can call the selectComposition() method to select the appropriate functional component based on the rendering request. The Remotion service returns the metadata to the rendering controller, which calls the Remotion service's setMetadata() and setRenderOptions() methods, respectively, to set the metadata and video rendering parameters in the context.
[0095] Next, the rendering controller instructs the Remotion service to start video rendering. The Remotion service creates a browser instance and uses the browser instance to render the target video.
[0096] After rendering the target video, the rendering controller calls the storage service to upload the target video to the object storage. The object storage returns the target video's key to the storage service. The storage service generates a download link based on the target video's key and returns the download link to the rendering controller. The rendering controller returns the download link to the client via the HTTP interface. The client uses the download link to access the object storage and obtain the corresponding target video, which can then be played on the client or shared with other users.
[0097] It is worth noting that during the video rendering process, the Remotion service will continuously update the progress data of the video rendering. The Remotion service passes the progress data to the rendering controller through the addProgress() method. The rendering controller calculates the video rendering progress based on the progress data and pushes the video rendering progress to the client by sending progress events to display the video rendering progress on the client.
[0098] In some possible implementations, the video template file includes:
[0099] The video implementation logic for defining a video content presentation process is configured to determine metadata information required for video rendering based on the video rendering parameters, and to construct context information required for video rendering based on the metadata information and the video rendering parameters, wherein the metadata information includes time metadata corresponding to each video content presentation process and content data corresponding to the time metadata;
[0100] The functional component is used to render the video content corresponding to each video frame according to the context information to obtain the target video.
[0101] In some possible implementations, the video template file further includes:
[0102] A root component, used to receive the video rendering parameters;
[0103] The registration component nested in the root component is used to register the combined component and pass the video rendering parameters to the combined component. The video implementation logic and the functional component are nested in the combined component.
[0104] Here, the root component serves as the entry point for the entire video rendering and is used to receive the video rendering parameters passed by the video rendering service. The registered component (ArtifactCompositions) is nested inside the root component, that is, the registered component is a child component of the root component, and the root component is the parent component of the registered component. The registered component is used to contain and register the composite component so that all components in the composite component can be used in the entire architecture. The registered component can pass the video rendering parameters from the root component to the composite component. The composite component is used to nest the video implementation logic, that is, the video implementation logic is a child component of the composite component, and the composite component is the parent component of the video implementation logic. By nesting different video implementation logic in the composite component, different video content can be achieved.
[0105] The video implementation logic receives the video rendering parameters passed by the composition component, determines the metadata information required for video rendering based on the video rendering parameters, and constructs the context information required for video rendering based on the metadata information and the video rendering parameters. It should be noted that the video implementation logic can be a component that implements specific video rendering logic. The video implementation logic can include a context component to provide context information to the underlying functional components.
[0106] The functional component is a child component of the video implementation logic, which is then its parent component. The functional component is called by the video implementation logic and, based on the context provided by the video implementation logic, renders the video content corresponding to each video frame to obtain the target video.
[0107] In the video template file, the root component, registration component, combination component, video implementation logic, and functional component adopt a multi-layer nested component architecture, which allows each component to be reused, not only in the same video template file, but also in different video template files.
[0108] Figure 5 FIG. 1 is a diagram showing the architecture of a video template file according to an exemplary embodiment. Figure 5As shown, the video template file includes a root component layer, a composite component layer, a video template layer, and a functional component layer. A root component (RemotionRoot) is provided in the root component layer for receiving video rendering parameters. A registration component (ArtifactCompositions) and a composition component (Composition) are provided in the composite component layer. The registration component is included in the root component and is used to register the composition component. For example, the composition component in the composition component layer can be used to generate a video showing the content of the conversation.
[0109] The video template layer can include video implementation logic, which defines the video content display process and provides a complete implementation for rendering specific video content. Of course, the video implementation logic is primarily used to provide context information to functional components so that functional components can render the target video based on this context information. The video implementation logic can also include context components (such as DebuggingProvider and TimeMetadataProvider) that provide context information.
[0110] The video template layer includes functional components, which are called by the video implementation logic. Based on the context information, the functional components render the video content corresponding to each video frame to obtain the target video. The functional components include at least one component for implementing a specific video rendering function. Functional components can be composite components composed of multiple components.
[0111] like Figure 5 As shown, the functional components may include a Dialogs (interface element for creating a modal dialog box) component located in the sub-component layer for rendering the dialog content, a Background component for rendering the background image, and an Ending component for rendering the end of the video.
[0112] Of course, each component in the sub-component layer can also be composed of one or more atomic components. For example, the Dialogs component can include a BubbleList component (a component for displaying a list of dialogue bubbles with automatic scrolling function). The BubbleList component can include a single robot dialogue bubble (SingleBotBubble) component for displaying the dialogue content of the virtual object, a player dialogue bubble (PlayerBubble) component for displaying the dialogue content of the user, a story dialogue bubble (BubbleStory) component for displaying story-related dialogues, a narration dialogue bubble (NarratorBubble) component for displaying narration content, and an audio (OptionalAudio) component for presenting audio.
[0113] Therefore, through the above implementation method, the video template file is designed with a multi-layer nested component architecture, so that the components of each layer can be reused, greatly improving development efficiency.
[0114] In some possible implementations, the video implementation logic includes:
[0115] The data source layer includes a data model and a metadata calculation function. The data model is used to store video rendering parameters through a preset data structure. The metadata calculation function is used to determine the metadata information required for video rendering based on the video rendering parameters stored in the data model.
[0116] The context component is used to construct the context information required for video rendering based on metadata information and video rendering parameters;
[0117] The animation calculation layer is used to determine the frame number corresponding to each video frame being rendered;
[0118] The functional component is used to render the video content corresponding to each video frame according to the context information and frame number to obtain the target video.
[0119] Here, the data source layer serves as the input starting point for receiving video rendering parameters. The data source layer includes a data model (VideoRenderingResources), which is used to store the received video rendering parameters using a preset data structure. This unified data structure ensures that each component can access consistent information, reducing data conversion costs. The metadata calculation function (calculateMetadata) included in the data source layer receives the data model as input and calculates the metadata information required for video rendering. The data source layer allows the separation of the calculation of original video rendering parameters and derived metadata information, improving the flexibility and maintainability of the system.
[0120] The context component receives the metadata information calculated by the metadata calculation function and the video rendering parameters in the data model, constructs the context information required for video rendering, and provides this context information to the functional component, providing it with an accurate time reference for frame-level rendering. It should be understood that the context component can be implemented through React's context interface, allowing downstream functional components to easily obtain context information without the cumbersome parameter passing between components, greatly simplifying data communication between components.
[0121] The animation calculation layer is responsible for determining the frame number of each video frame being rendered. This frame number is the basis for all animation calculations and can be used to calculate the state and style changes corresponding to the video frame being rendered. For example, the animation calculation layer can obtain the frame number of the currently rendered video frame by calling useCurrentFrame (a core hook provided by Remotion). When rendering each video frame, useCurrentFrame is called and returns the frame number corresponding to the current video frame.
[0122] Correspondingly, the functional component can receive the context information and render the video content corresponding to each video frame according to the frame number and context information corresponding to the current video frame.
[0123] In some embodiments, the animation calculation layer is further configured to determine, based on the frame number, an elastic animation value used for the video frame corresponding to the frame number. Accordingly, the functional component is configured to render the video content corresponding to each video frame based on the context information, the frame number, and the elastic animation value to obtain a target video.
[0124] Among them, the animation calculation layer can receive the frame number corresponding to the current video frame as a parameter through the spring function (a tool for creating smooth animations), calculate the physically based elastic animation value, so that the functional component can generate a smooth and natural transition effect according to the elastic animation value, making the appearance and movement of elements more vivid.
[0125] In some possible implementations, the context component is further used to provide a debugging status during the video rendering process; and the function component is further used to display the debugging status.
[0126] Here, the context component can use DebuggingProvide (a debugging tool provided by Remotion) to collect debugging information during the video rendering process. This information is then used to output debugging status during development and debugging, helping developers understand the system's operating status and troubleshoot problems. Through DebuggingProvider, the display of debugging status can be centrally controlled. During the development and testing phases, debug mode is enabled to display detailed debugging status; in a production environment, debug mode can be turned off to prevent the debugging status from interfering with the user experience. During the development process, functional components can decide whether to display debugging status, such as logs, warnings, and error messages, based on the received debugging status, to help developers understand the functional component's operating status and data flow.
[0127] Figure 6 FIG. 1 is a schematic diagram showing the data flow according to an exemplary embodiment. Figure 6As shown, the data source layer includes a data model and a metadata calculation function. The metadata calculation function calculates metadata information based on the video rendering parameters in the data model and passes the metadata information to the TimingMetadataProvider in the context component. Similarly, the original video rendering parameters in the data model are also passed to the TimingMetadataProvider to form the context information.
[0128] The animation calculation layer obtains the frame number corresponding to the video frame being rendered through useCurrrentFrame, and passes the frame number corresponding to the video frame being rendered to the spring function and functional component respectively. The spring function calculates the corresponding elastic animation value based on the frame number and passes the calculated elastic animation value to the functional component.
[0129] like Figure 6 As shown, functional components can include basic components (such as Dialogs components, Background components, Ending components, etc.), Dialogs components and BubbleList components. Among them, the basic components are the carriers of business logic and are responsible for organizing the overall structure of video content, such as background, title, ending, etc. The Dialogs component is used to process the rendering logic related to the dialogue content, filter and process the dialogue content, control the animation and transition effects, and manage the display logic of the dialogue bubbles. The Dialogs component mainly uses time metadata to determine when to display which dialogue content. The BubbleList component is used to process bubble rendering and animation, calculate style changes based on frame numbers, and manage the scrolling effects of multiple bubbles. The BubbleList component receives the dialogue content that has been filtered and batched from the Dialogs component to implement the specific interactive presentation of the bubble interface.
[0130] like Figure 6 As shown, the DebuggingProvider in the context component collects the debugging status and provides it to the functional component so that the functional component displays the debugging status.
[0131] Therefore, through the architecture of the above data source layer, context component, animation calculation layer and functional component, the context information calculation and functional component are separated, the context and functional components can be decoupled and the functional components can be reused.
[0132] Figure 7 FIG. 1 is a flow chart showing a method of generating a target video according to another exemplary embodiment. Figure 7 As shown, in some possible implementations, the target video can be generated through the following steps.
[0133] In step 710, the total number of video frames corresponding to the video rendering parameters is determined based on the video rendering parameters and the duration corresponding to the video content display process included in the video template file.
[0134] Here, the total number of video frames refers to the total number of video frames contained in the video to be rendered using the video rendering parameters and the video template file. Since the video template file defines all video content display processes, and the video rendering parameters affect the corresponding video duration of different video content display processes, the video rendering parameters and the video template file can be used to calculate the total number of video frames in the target video to be rendered.
[0135] For example, the total number of video frames can be calculated using a first calculation formula, which is:
[0136] totalFrames=getDialogueFrames(dialogues)+conditional(intro+ending)*fps+(startDelay+endingDelay+sloganDelay)*fps+sloganFrames
[0137] Where totalFrames is the total number of video frames, fps is the frame rate, dialogues is the dialogue content, startDelay is the delay time of the beginning of the video, endingDelay is the delay time of the ending part, sloganDelay is the delay time of the brand slogan, sloganFrames is the length of the video clip corresponding to the display process of the brand slogan, getDialogueFrames(dialogues) is the length of the audio obtained based on the dialogue content, ending is the length of the ending part, and intro is the lead delay of the work introduction video clip.
[0138] Note that startDelay, endingDelay, sloganDelay, and sloganFrames can be fixed values based on business needs. In other words, they are fixed values in the video template file. Intro, ending, and audio duration are dynamic parameters. For example, the audio duration depends on the length of the dialogue in the video clip.
[0139] It is worth noting that the duration corresponding to the video content display process included in the video template file used in the above first calculation formula is only used as an example. Different video template files may have different video content display processes and the duration corresponding to the video content display process. Therefore, for different video template files, even if the video rendering parameters are the same, the calculated total number of video frames may also be different.
[0140] In step 720, the total number of video frames is divided into multiple video slices according to the preset video segment duration, and a video rendering task corresponding to each video slice is created.
[0141] Here, the preset video clip length is the optimal length for video rendering by a single video rendering service instance calculated through big data.
[0142] In some embodiments, the duration of the video clip can be determined based on the resource information of the service cluster where the video rendering service is located, the maximum load of the service cluster, the average rendering speed of the video rendering service in historical time, and the average duration of historical video materials processed by the video rendering service.
[0143] Among them, the resource information of the service cluster may include the total number of CPU (central processing unit) cores used by the service cluster. The maximum load of the service cluster may refer to the maximum concurrent number of rendering requests sent by the client received by the service cluster. It should be understood that the maximum concurrent number can be obtained through big data measurement. The average duration may be the average duration of historical video materials obtained by statistically analyzing all historical video materials processed by the video rendering service. It should be understood that if the video material includes conversation content, the average duration may refer to the average length of the conversation content. The average rendering speed of the video rendering service in historical time may be obtained by statistically analyzing the speed at which the video rendering service generates video. The average rendering speed of the video rendering service in historical time may include the average rendering speed of text screens and the average rendering speed of video screens.
[0144] For example, the total number of video frames can be calculated using a second calculation formula, which is:
[0145]
[0146] Among them, ClipDuration is the length of the video clip, TotalCpuCores is the total number of cores, qpsMax is the maximum concurrency, dialoguesAug is the average length of the dialogue content, renderSpeed is the average rendering speed of the video rendering service in historical time, textRenderSpeed is the average rendering speed of text images, livePhoto is the length of the dynamic video, and videoRenderSpeed is the average rendering speed of the video images.
[0147] The calculated total number of video frames is split into multiple video slices based on the preset video segment duration, with each video slice including the number of video frames corresponding to the video segment duration. It should be noted that the total number of video frames can be split into multiple video slices in chronological order, with each video slice including the number of video frames corresponding to the video segment duration.
[0148] For each video slice, a video rendering task corresponding to the video slice is generated. The video rendering task is used to instruct the video rendering service to render a number of video frames corresponding to the video slice.
[0149] In step 730, the video rendering service is called to execute multiple video rendering tasks in parallel, and the video clips corresponding to the respective video rendering tasks are obtained according to the video rendering parameters and the video template file.
[0150] The video rendering service can create multiple instances, each of which performs a single video rendering task. This allows all video rendering tasks to be executed simultaneously through multiple instances. Each instance renders all the frames included in the corresponding video slice based on the received video rendering parameters and the used video template file, forming the video clip for the corresponding video rendering task.
[0151] It is worth noting that, regarding how the video rendering service renders the video clips according to the received video rendering parameters and the used video template file, reference can be made to the relevant description of the above embodiment, which will not be repeated here.
[0152] In step 740 , multiple video clips are spliced together to obtain a target video.
[0153] Here, after obtaining a plurality of video segments, the plurality of video segments may be spliced together in a time sequence of the video segments to obtain a corresponding target video.
[0154] Therefore, by splitting a video rendering task into multiple video rendering tasks and having the video rendering service execute the multiple video rendering tasks in parallel, the time consumed in video generation can be significantly shortened.
[0155] In some achievable embodiments, the video slices are video frames without audio, and accordingly, the video segments rendered by the video rendering service are video segments without audio. The total audio duration corresponding to the video rendering parameters can be determined based on the video rendering parameters and the duration corresponding to the video content display process included in the video template file. The total audio duration can be divided into multiple audio slices based on the preset audio segment durations, and an audio rendering task corresponding to each audio slice can be constructed. The video rendering service can then be called to execute the multiple audio rendering tasks in parallel, and the audio segments corresponding to each audio rendering task can be obtained based on the video rendering parameters and the video template file.
[0156] Here, if the video and audio are rendered together, due to the difference between the video encoding rate (30FPS) and the audio sampling rate (44kHZ), there will be a sound pause of millimeter length at the splicing of the video clips. In order to ensure the losslessness of the audio, the length of the single audio rendering needs to be as long as possible, and it is best to render the audio independently at one time. However, if the audio is rendered at one time, in long video scenes, audio rendering will become a performance shortcoming. In a distributed rendering system, all audio and video images are started to render at the same time. The longer the video image, the longer the audio. Since in the above embodiment, the video frame will be split into multiple rendering tasks for parallel rendering, the time-consuming audio rendering will become a performance shortcoming.
[0157] To ensure that all video segments render in parallel with each other, minimizing the total rendering time, it's necessary to independently set the lengths of video frames and audio segments. Therefore, in this disclosed embodiment, the total audio duration is divided into multiple audio slices based on the preset audio segment duration. Each audio slice includes a number of audio frames corresponding to the audio segment duration. Then, an audio rendering task is constructed for each audio slice.
[0158] It's important to note that the total audio duration corresponding to the video rendering parameters refers to the duration of the audio in the target video that will be rendered using these parameters. The total audio duration is determined by the video rendering parameters and the duration of the video content display process included in the video template file. For details, refer to the description of the total video frame count in the above embodiment.
[0159] In this embodiment, the duration of the audio segment is in a preset ratio to the duration of the video segment. For example, the duration of the audio segment and the duration of the video segment can be 1:4. For example, if a video segment is 3 seconds, then an audio segment can be 12 seconds.
[0160] After determining the audio rendering tasks corresponding to each audio slice, the video rendering service can be scheduled to execute multiple audio rendering tasks in parallel. According to the video rendering parameters and the video template file, the audio clips corresponding to each audio rendering task are obtained. It should be understood that the audio rendering tasks and the video rendering tasks can be performed synchronously.
[0161] Accordingly, after all video rendering tasks and audio rendering tasks are completed, multiple video clips can be spliced to obtain an initial video without audio, multiple audio clips can be spliced to obtain the target audio, and then the initial video and the target audio can be spliced to obtain the target video.
[0162] Since the audio and video images are rendered independently during the rendering stage, after receiving the audio clips and video clips, all video clips without audio can be spliced into an initial video without audio, all audio clips can be spliced into target audio, and then the initial video without audio and the target audio can be spliced to obtain the target audio containing audio and video images.
[0163] For example, the audio and video clips may be joined using FFmpeg (a set of open source computer programs that can be used to record, convert, and stream digital audio and video).
[0164] You can use the command "ffmpeg -f concat -safe 0 -i audioFileList.txt -c copy audio.mp3" to concatenate audio clips. ffmpeg calls FFmpeg, -f concat specifies the input format as "concat" and uses FFmpeg's concatenation function, -safe 0 allows the use of file lists containing special characters or relative paths, -i audioFileList.txt specifies a text file containing the list of audio clips to be concatenated, -c copy copies the audio data without re-encoding, and audio.mp3 is the output of the concatenated target audio.
[0165] You can use "ffmpeg -f concat -safe 0 -i videoFileList.txt -c copy video.mp4" to join video clips. ffmpeg calls FFmpeg, -f concat specifies the input format as "concat" and uses FFmpeg's concatenation function, -safe 0 allows file lists containing special characters or relative paths, -ivideoFileList.txt specifies a text file containing the list of video clips to be joined, -c copy copies the video data without re-encoding, and video.mp4 is the output of the merged video.
[0166] You can use "ffmpeg -i video.mp4 -i audio.mp3 -map 0:v -map 1:ac:v copyoutput.mp4" to merge the original video and the target audio. Among them, ffmpeg means calling FFmpeg, -i video.mp4 means specifying the original video as the first input file, -i audio.mp3 means specifying the target audio as the second input file, -map 0:v means selecting the video stream from the first input file, -map 1:a means selecting the audio stream from the second input file, -c:v copy means copying the video stream and audio stream without re-encoding, and output.mp4 means the input merged target video.
[0167] It should be noted that the more audio and video clips there are, the longer the splicing time will be. To reduce the splicing time, the splicing operation can be performed in a subprocess to increase the splicing speed.
[0168] After the target video is obtained by splicing, the target video can be returned to the client. Of course, continuing with the above embodiment, a download link of the target video can be returned to the client.
[0169] Figure 8 FIG. 1 is a flow chart of a video generation method according to another exemplary embodiment. Figure 8 As shown, in some possible implementations, the video generation method may include the following steps.
[0170] In step 810, the video material sent by the client is obtained through the scheduling service.
[0171] Here, the scheduling service may receive a rendering request from a client, where the rendering request carries video materials required for video rendering.
[0172] In some embodiments, a scheduling service is used to receive video material sent by a server corresponding to a client; the client is used to send a unique identifier corresponding to the video material to the server, and the server is used to obtain the video material from a database according to the unique identifier.
[0173] It's important to note that the rendering request sent by the client to the server can include a unique identifier corresponding to the video material. Upon receiving the rendering request, the server uses the unique identifier to query the server's database, retrieve the required video material from the database, and pass it to the scheduling service. For example, in a scenario where a conversation is displayed, the user can send a rendering request to the server with the unique identifier of the conversation content, and the server will retrieve the corresponding conversation content from the database.
[0174] In step 820, the video rendering parameters required for video rendering are determined based on the video material through the scheduling service.
[0175] Here, after the scheduling service obtains the video material, it can process the video material to obtain the video rendering parameters. It should be understood that the detailed description of the video rendering parameters can refer to the relevant description of the above embodiment.
[0176] In some embodiments, the video material includes text content, and accordingly, the video rendering parameters include audio data corresponding to the text content. In response to the video material including text content, a text-to-speech (TTS) service can be called through a scheduling service to convert the text content into audio data.
[0177] It should be noted that the scheduling service can call the TTS service to convert text content into audio data.
[0178] In step 830, a plurality of video rendering tasks are determined according to the video template file and the video rendering parameters through the scheduling service.
[0179] Here, the detailed description of the video rendering task can refer to the relevant description of the above embodiment, which will not be repeated here. Continuing with the above implementation, both audio and video images can be rendered separately and in parallel using the above implementation.
[0180] In step 840 , a plurality of video rendering tasks are executed in parallel by the video rendering service to obtain a video segment corresponding to each video rendering task.
[0181] Here, the scheduling service can call the video rendering service to execute multiple video rendering tasks in parallel and obtain the video clips corresponding to each video rendering task. Continuing with the above implementation, the video rendering service can be deployed in a function-as-a-service container specifically for video rendering.
[0182] In step 850, multiple video clips are spliced together through the scheduling service to obtain a target video.
[0183] Here, after the video rendering service obtains the video clip, it returns the video clip to the scheduling service, and then the scheduling service splices multiple video clips to obtain the target video.
[0184] In step 860 , the target video is returned to the client through the scheduling service.
[0185] Here, after the scheduling service obtains the target video, it can return the target video to the client. For example, the scheduling service can upload the target video to the video cloud platform and generate a download link for the target video. The scheduling service then returns the download link to the client corresponding to the video material, so that the client can obtain the target video from the video cloud platform according to the download link.
[0186] Therefore, through multiple microservices such as scheduling service, video rendering service, TTS service, video cloud platform, etc., video generation can be completed in collaboration, greatly improving the efficiency and quality of video generation.
[0187] Figure 9 FIG. 1 is an architecture diagram of a video generation system according to an exemplary embodiment. Figure 9 As shown, the client can select a video material and send a rendering request to the server via the HTTP interface. This rendering request carries the unique identifier corresponding to the video material. The server responds to the rendering request, queries the database based on the unique identifier, performs a cache hit check, determines the video material, and then constructs a rendering task based on the video material and submits it to the task queue. The rendering task then carries the video material to the scheduling service.
[0188] In response to the received video material, the scheduling service processes the video material into video rendering parameters and creates multiple video rendering tasks. The scheduling service then calls the video rendering service, which receives the video rendering parameters passed by the scheduling service and executes the multiple video rendering tasks in parallel through multiple video rendering service instances to obtain the video clips corresponding to the video rendering tasks. The video rendering service can communicate with external parties through the FaaS gateway. While processing the video material, the scheduling service can call the text-to-speech service to convert the text content into audio data.
[0189] The video rendering service returns video clips to the scheduling service. The scheduling service uses the video stitching service to stitch the video clips into the target video and returns the target video to the server. The server returns the target video to the client via a Websocket (a full-duplex communication protocol). Of course, the scheduling service can also store the target video in the video cloud and return a download link to the client. The client can then download the target video from the video cloud and play / share the target video.
[0190] In addition, during the video rendering process, the scheduling service can feedback the rendering progress to the client through the server, so that the rendering progress of the video can be displayed in real time while the client is waiting for the video to be generated.
[0191] Figure 10 FIG. 1 is a logic sequence diagram of a video generation method according to an exemplary embodiment. Figure 10 As shown, the user can select a video material in the client and then send the video material to the server by clicking a button. The server submits a rendering task including the video material to the scheduling service. The scheduling service can call the TTS service to convert the text content into audio data, and the TSS service returns the audio data to the scheduling service. The scheduling service uploads the audio data to the temporary storage, and the temporary storage returns the audio data link to the scheduling service. The scheduling service constructs multiple video rendering tasks and instructs the video rendering service to execute multiple video rendering tasks in parallel. During the execution of the video rendering task, the video rendering task can obtain the required audio data from the temporary storage through the audio data link, and the temporary storage returns the audio data to the video rendering service. During the execution of the video rendering task, the video rendering service feeds back the rendering progress of the video rendering task to the scheduling service, and the scheduling service returns the rendering progress to the client through the server. The video rendering service uploads the rendered video clips to the temporary storage.
[0192] After the video rendering is complete, the scheduling service downloads the video clips from temporary storage and stitches them together to create the target video. The scheduling service uploads the target video to the video cloud platform, which returns the target video key to the scheduling service. The scheduling service generates a download link based on the target video key and returns the download link to the client through the server. Users can access the video cloud platform through the download link to view the target video.
[0193] Figure 11 FIG. 1 is an architecture diagram of a video rendering service instance according to an exemplary embodiment. Figure 11As shown, in the pre-rendering preparation stage, the video rendering service instance performs operations to initialize Stream (streaming of data during video rendering), initialize AbortControl (an interface for controlling asynchronous operations), and configure video rendering parameters to initialize the context. In the video rendering stage, the video rendering service instance performs operations to select a video template file, start the rendering process, and start rendering to start video rendering. In the output result stage, the video rendering service instance performs operations to output rendering progress, input rendered video clips, and construct synthesis parameters. Among them, the synthesis parameters are used for the scheduling service to synthesize the target video. In the rendering process, the timeline of the video frame sequence is provided as a time reference for the functional components, and the corresponding video frames are rendered.
[0194] Figure 12 FIG. 1 is a flow chart of a video generation method according to another exemplary embodiment. Figure 12 As shown, the embodiment of the present disclosure provides a video generation method, which can be executed by a client. The method can be specifically executed by a video generation device, which can be implemented by software and / or hardware, and the device can be configured in the client. Figure 12 As shown, the method may include the following steps.
[0195] S1201, in response to a material selection operation, determining a video material indicated by the material selection operation;
[0196] S1202, display the target video. The target video is generated based on the video material, the video rendering parameters required for video rendering are determined, and the video template file is called and generated according to the video rendering parameters. The video template file is used to define the video content display process. The video template file renders the video content of the video frames corresponding to each video content display process according to the video rendering parameters.
[0197] Here, the user can select video material in the user interface by triggering a material selection operation in the user interface.
[0198] In some embodiments, the video material includes at least the content of the conversation between the user and the virtual object. Accordingly, the conversation content can be displayed in the user interface of the conversation between the user and the virtual object, and in response to the user selecting a material for the conversation content, the selected conversation content is determined as the video material.
[0199] The target video is displayed in the user interface. The target video is generated based on the video material, by determining the video rendering parameters required for video rendering, and calling a video template file to generate a video for displaying the video material based on the video rendering parameters. It should be understood that the generation of the target video can be referred to the relevant description of the above embodiment and will not be repeated here.
[0200] In some feasible implementations, candidate video implementation logic and candidate functional components can be displayed, the video implementation logic is used to define the video content display process, and the functional components are used to render video materials. Then, in response to a selection operation, a video template file is constructed according to the video implementation logic and functional components indicated by the selection operation.
[0201] In some possible implementations, the video template file includes:
[0202] The video implementation logic used to define the video content display process is used to determine the metadata information required for video rendering based on the video rendering parameters, and to construct the context information required for video rendering based on the metadata information and the video rendering parameters. The metadata information includes the time metadata corresponding to each video content display process and the content data corresponding to the time metadata;
[0203] The functional component is used to render the video content corresponding to each video frame according to the context information to obtain the target video.
[0204] In some possible implementations, the video template file further includes:
[0205] Root component, used to receive video rendering parameters;
[0206] The registration component nested in the root component is used to register the composite component and pass the video rendering parameters to the composite component. The video implementation logic and functional components are nested in the composite component.
[0207] In some possible implementations, the video implementation logic includes:
[0208] The data source layer includes a data model and a metadata calculation function. The data model is used to store video rendering parameters through a preset data structure. The metadata calculation function is used to determine the metadata information required for video rendering based on the video rendering parameters stored in the data model.
[0209] The context component is used to construct the context information required for video rendering based on metadata information and video rendering parameters;
[0210] The animation calculation layer is used to determine the frame number corresponding to each video frame being rendered;
[0211] The functional component is used to render the video content corresponding to each video frame according to the context information and frame number to obtain the target video.
[0212] In some possible implementations, the animation calculation layer is further used to determine, based on the frame number, an elastic animation value used by the video frame corresponding to the frame number;
[0213] The functional component is used to render the video content corresponding to each video frame according to the context information, frame number and elastic animation value to obtain the target video.
[0214] In some feasible implementations, in step S1202, video material can be sent to the server, and the target video sent by the server can be received and displayed. The server is used to determine the video rendering parameters required for video rendering based on the video material, and generate the target video according to the video rendering parameters through the video rendering service. The video rendering service is used to generate the target video according to the video template file and the video rendering parameters.
[0215] In some feasible implementations, in step S1202, video material can be sent to the server, and the target video sent by the server is received and displayed. The server is used to determine the video rendering parameters required for video rendering based on the video material, and determine the total number of video frames corresponding to the video rendering parameters based on the video rendering parameters and the duration corresponding to the video content display process included in the video template file. According to the preset video clip duration, the total number of video frames is split into multiple video slices, and a video rendering task corresponding to each video slice is created, and the video rendering service is called to execute multiple video rendering tasks in parallel. According to the video rendering parameters and the video template file, the video clip corresponding to each video rendering task is obtained, and the multiple video clips are spliced to obtain the target video.
[0216] Figure 13 FIG. 1 is a flow chart of a video generation method according to another exemplary embodiment. Figure 13 As shown, the embodiment of the present disclosure provides a video generation method, which can be executed by a client. The method can be specifically executed by a video generation device, which can be implemented by software and / or hardware, and the device can be configured in the client. Figure 13 As shown, the method may include the following steps.
[0217] S1301: Display candidate video implementation logic and candidate functional components. The video implementation logic is used to define the video content display process, and the functional components are used to render video materials.
[0218] S1302, in response to the selection operation, construct a video template file according to the video implementation logic and functional components indicated by the selection operation; the video template file is used to render the video content of the video frames corresponding to each video content display process according to the video rendering parameters corresponding to the video material, and obtain the target video.
[0219] Here, candidate video implementation logic and candidate functional components are displayed in the editing interface. Users can then select the desired video implementation logic and functional components in the editing interface to build the desired video template file. Each candidate video implementation logic is used to define a different video content display process. Each different functional component is used to render the video material with different effects.
[0220] Figure 14 FIG is a schematic diagram of an editing interface according to an exemplary embodiment. Figure 14 As shown, a video implementation logic list can be displayed in the editing interface. The video implementation logic list displays candidate video implementation logics for the user to select. It is important to note that each video implementation logic can define a different video content display process. For example, different video implementation logics for displaying conversation content can use different video content display processes to display conversation content. In the user's visual perception, the video effects generated by using different video implementation logics are different.
[0221] The editing interface displays a list of functional components, showing candidate functional components for the user to select. For example, the functional component list may display Dialogs, Background, Ending, and other components that implement different video effects. Of course, the Dialogs component can also provide a BubbleList component composed of different subcomponents. For example, a BubbleList component can be composed of a SingleBotBubble component, a PlayerBubble component, a BubbleStory component, a NarratorBubble component, and an OptionalAudio component. Users can freely select different functional components based on the desired video effect.
[0222] The user can select different video implementation logics and functional components by triggering a selection operation in the editing interface, so that the client can construct a video template file based on the video implementation logic and functional components selected by the selection operation. Figure 4 As shown, the user can select the desired video implementation logic and functional components by dragging and dropping them to the editing area of the editing interface. Of course, the editing interface can have a preview area. After the user completes the selection of the video implementation logic and functional components, the video template file constructed according to the selected video implementation logic and functional components can be displayed in the preview area. The video effect preview corresponding to the video template file is displayed.
[0223] It should be understood that the constructed video template file can be sent to the server to generate a video according to the video template file created by the user in a subsequent process.
[0224] Therefore, through the above implementation method, users can select different video implementation logic and functional components to build the required video template file according to the required video effect, allowing users to personalize the video effect, greatly improving the user's video production experience.
[0225] Figure 15 FIG. 1 is a schematic diagram showing the structure of a video generating device according to an exemplary embodiment. Figure 15 As shown, the embodiment of the present disclosure provides a video generating device 1500, which is configured in a server. The video generating device 1500 includes:
[0226] The acquisition module 1501 is configured to acquire video material;
[0227] The determination module 1502 is configured to determine video rendering parameters required for video rendering according to the video material;
[0228] The generating module 1503 is configured to call a video template file and generate a target video for displaying the video material according to the video rendering parameters, wherein the video template file is used to define a video content display process, and the video template file renders the video content of the video frame corresponding to each video content display process according to the video rendering parameters;
[0229] The output module 1504 is configured to output the target video.
[0230] Optionally, the generating module 1503 is specifically configured to:
[0231] The target video is generated according to the video rendering parameters through a video rendering service, and the video rendering service is used to generate the target video according to the video template file and the video rendering parameters.
[0232] Optionally, the video rendering service performs the following steps:
[0233] Obtaining the video rendering parameters;
[0234] Determining metadata information required for video rendering according to the video rendering parameters and the video template file, the metadata information including time metadata corresponding to each video content display process and content data corresponding to the time metadata;
[0235] Constructing context information required for video rendering according to the metadata information and the video rendering parameters;
[0236] A rendering engine is called to render the target video according to the functional components indicated by the video template file and the context information.
[0237] Optionally, the video rendering service further performs the following steps:
[0238] Upload the target video to the video cloud platform and generate a download link for the target video;
[0239] The download link is returned to the client corresponding to the video material, so that the client obtains the target video from the video cloud platform according to the download link.
[0240] Optionally, the video rendering service is deployed in a function-as-a-service container.
[0241] Optionally, the video template file includes:
[0242] The video implementation logic for defining a video content presentation process is configured to determine metadata information required for video rendering based on the video rendering parameters, and to construct context information required for video rendering based on the metadata information and the video rendering parameters, wherein the metadata information includes time metadata corresponding to each video content presentation process and content data corresponding to the time metadata;
[0243] The functional component is used to render the video content corresponding to each video frame according to the context information to obtain the target video.
[0244] Optionally, the video template file further includes:
[0245] A root component, used to receive the video rendering parameters;
[0246] The registration component nested in the root component is used to register the combined component and pass the video rendering parameters to the combined component. The video implementation logic and the functional component are nested in the combined component.
[0247] Optionally, the video implementation logic includes:
[0248] A data source layer, comprising a data model and a metadata calculation function, wherein the data model is used to store the video rendering parameters using a preset data structure, and the metadata calculation function is used to determine metadata information required for video rendering based on the video rendering parameters stored in the data model;
[0249] A context component, configured to construct context information required for video rendering based on the metadata information and the video rendering parameters;
[0250] The animation calculation layer is used to determine the frame number corresponding to each video frame being rendered;
[0251] The functional component is used to render the video content corresponding to each video frame according to the context information and the frame number to obtain the target video.
[0252] Optionally, the animation calculation layer is further configured to determine, based on the frame number, an elastic animation value used by the video frame corresponding to the frame number;
[0253] The functional component is used to render the video content corresponding to each video frame according to the context information, the frame number and the elastic animation value to obtain the target video.
[0254] Optionally, the context component is further used to provide a debugging status during the video rendering process; and the functional component is further used to display the debugging status.
[0255] Optionally, the generating module 1503 is specifically configured to:
[0256] Determining the total number of video frames corresponding to the video rendering parameters according to the video rendering parameters and the duration corresponding to the video content display process included in the video template file;
[0257] Splitting the total number of video frames into multiple video slices according to the preset video segment length, and creating a video rendering task corresponding to each video slice;
[0258] Calling a video rendering service to execute a plurality of the video rendering tasks in parallel, and obtaining a video clip corresponding to each of the video rendering tasks according to the video rendering parameters and the video template file;
[0259] The plurality of video clips are spliced together to obtain the target video.
[0260] Optionally, the generating module 1503 is specifically configured to:
[0261] The duration of the video clip is determined based on the resource information of the service cluster where the video rendering service is located, the maximum load of the service cluster, the average rendering speed of the video rendering service in historical time, and the average duration of historical video materials processed by the video rendering service.
[0262] Optionally, the video slice is a video picture without audio; the generating module 1503 is further configured to:
[0263] Determine the total audio duration corresponding to the video rendering parameters according to the video rendering parameters and the duration corresponding to the video content display process included in the video template file;
[0264] Divide the total audio duration into multiple audio slices according to a preset audio segment duration, and construct an audio rendering task corresponding to each audio slice, wherein the audio segment duration is in a preset ratio to the video segment duration;
[0265] Calling a video rendering service to execute the plurality of audio rendering tasks in parallel, and obtaining an audio clip corresponding to each audio rendering task according to the video rendering parameters and the video template file;
[0266] The generating module 1503 is specifically configured to:
[0267] splicing a plurality of the video clips to obtain an initial video without audio;
[0268] Splicing multiple audio clips to obtain target audio;
[0269] The initial video and the target audio are spliced together to obtain the target video.
[0270] Optionally, the acquisition module 1501 is specifically configured to:
[0271] Get the video material sent by the client through the scheduling service;
[0272] The determining module 1502 is specifically configured to:
[0273] Determining, through the scheduling service, video rendering parameters required for video rendering based on the video material;
[0274] The generating module 1503 is specifically configured to:
[0275] Determining, by the scheduling service, a plurality of video rendering tasks according to the video template file and the video rendering parameters;
[0276] By means of a video rendering service, a plurality of the video rendering tasks are executed in parallel to obtain a video clip corresponding to each of the video rendering tasks; the video rendering service is used to generate a video clip corresponding to the video rendering task according to the video template file and the video rendering parameters corresponding to the video rendering task;
[0277] splicing the plurality of video clips through the scheduling service to obtain the target video;
[0278] The output module 1504 is specifically configured to:
[0279] The target video is returned to the client through the scheduling service.
[0280] Optionally, the acquisition module 1501 is specifically configured to:
[0281] The scheduling service receives video material sent by the server corresponding to the client; the client is used to send a unique identifier corresponding to the video material to the server, and the server is used to obtain the video material from a database according to the unique identifier.
[0282] Optionally, the video material includes text content, and the video rendering parameters include audio data corresponding to the text content; the determining module 1502 is specifically configured to:
[0283] In response to the video material including the text content, the scheduling service calls a text-to-speech service to convert the text content into audio data.
[0284] Regarding the video generating device 1500 in the above embodiment, the method logic executed by each functional module has been described in detail in the part about the method, and will not be repeated here.
[0285] Figure 16 FIG. 1 is a schematic diagram showing the structure of a video generating device according to another exemplary embodiment. Figure 16 As shown, the embodiment of the present disclosure provides a video generating device 1600, which is configured in a client. The video generating device 1600 includes:
[0286] The selection module 1601 is configured to, in response to a material selection operation, determine the video material indicated by the material selection operation;
[0287] The display module 1602 is configured to display the target video, which is generated based on the video material, the video rendering parameters required for video rendering, and calling the video template file according to the video rendering parameters. The video template file is used to define the video content display process, and the video template file renders the video content of the video frames corresponding to each video content display process according to the video rendering parameters.
[0288] Optionally, the video generating device 1600 further includes:
[0289] A display unit configured to display candidate video implementation logic and candidate functional components, wherein the video implementation logic is used to define a video content display process, and the functional components are used to render video materials;
[0290] The construction unit is configured to construct a video template file in response to a selection operation according to the video implementation logic and the functional components indicated by the selection operation.
[0291] Optionally, the video template file includes:
[0292] The video implementation logic for defining a video content presentation process is configured to determine metadata information required for video rendering based on the video rendering parameters, and to construct context information required for video rendering based on the metadata information and the video rendering parameters, wherein the metadata information includes time metadata corresponding to each video content presentation process and content data corresponding to the time metadata;
[0293] The functional component is used to render the video content corresponding to each video frame according to the context information to obtain the target video.
[0294] Optionally, the video template file further includes:
[0295] A root component, used to receive the video rendering parameters;
[0296] The registration component nested in the root component is used to register the combined component and pass the video rendering parameters to the combined component. The video implementation logic and the functional component are nested in the combined component.
[0297] Optionally, the video implementation logic includes:
[0298] A data source layer, comprising a data model and a metadata calculation function, wherein the data model is used to store the video rendering parameters using a preset data structure, and the metadata calculation function is used to determine metadata information required for video rendering based on the video rendering parameters stored in the data model;
[0299] A context component, configured to construct context information required for video rendering based on the metadata information and the video rendering parameters;
[0300] The animation calculation layer is used to determine the frame number corresponding to each video frame being rendered;
[0301] The functional component is used to render the video content corresponding to each video frame according to the context information and the frame number to obtain the target video.
[0302] Optionally, the animation calculation layer is further configured to determine, based on the frame number, an elastic animation value used by the video frame corresponding to the frame number;
[0303] The functional component is used to render the video content corresponding to each video frame according to the context information, the frame number and the elastic animation value to obtain the target video.
[0304] Optionally, the display module 1602 is specifically configured to:
[0305] Sending the video material to a server;
[0306] Receive and display the target video sent by the server, the server is used to determine the video rendering parameters required for video rendering based on the video material, and generate the target video according to the video rendering parameters through the video rendering service, and the video rendering service is used to generate the target video according to the video template file and the video rendering parameters.
[0307] Optionally, the display module 1602 is specifically configured to:
[0308] Sending the video material to a server;
[0309] Receive and display the target video sent by the server, the server is used to determine the video rendering parameters required for video rendering based on the video material, determine the total number of video frames corresponding to the video rendering parameters based on the video rendering parameters and the duration corresponding to the video content display process included in the video template file, split the total number of video frames into multiple video slices based on the preset video segment duration, create a video rendering task corresponding to each of the video slices, call the video rendering service, execute multiple video rendering tasks in parallel, obtain the video segment corresponding to each of the video rendering tasks based on the video rendering parameters and the video template file, splice the multiple video segments, and obtain the target video.
[0310] Optionally, the video material at least includes the conversation content between the user and the virtual object.
[0311] Regarding the video generating device 1600 in the above embodiment, the method logic executed by each functional module has been described in detail in the part about the method, and will not be repeated here.
[0312] Figure 17 FIG. 1 is a structural diagram of a video generating device according to another exemplary embodiment. Figure 17 As shown, the embodiment of the present disclosure provides a video generating device 1700, which is configured in a client. The video generating device 1700 includes:
[0313] Display module 1701 is configured to display candidate video implementation logic and candidate functional components, wherein the video implementation logic is used to define the video content display process, and the functional components are used to render video materials;
[0314] Construction module 1702 is configured to construct a video template file in response to a selection operation according to the video implementation logic and the functional components indicated by the selection operation; the video template file is used to render the video content of the video frames corresponding to each video content display process according to the video rendering parameters corresponding to the video material to obtain the target video.
[0315] Regarding the video generating device 1700 in the above embodiment, the method logic executed by each functional module has been described in detail in the part about the method, which will not be repeated here.
[0316] Reference below Figure 18 , which shows an electronic device (eg Figure 1 The client in the embodiment of the present disclosure may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), in-vehicle terminals (such as in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 18 The electronic device shown is only an example and should not limit the functions and scope of use of the embodiments of the present disclosure.
[0317] like Figure 18 As shown, the electronic device 1800 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 1801, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 1802 or a program loaded from a storage device 1808 into a random access memory (RAM) 1803. Various programs and data required for the operation of the electronic device 1800 are also stored in the RAM 1803. The processing device 1801, the ROM 1802, and the RAM 1803 are connected to each other via a bus 1804. An input / output (I / O) interface 1805 is also connected to the bus 1804.
[0318] Typically, the following devices may be connected to the I / O interface 1805: an input device 1806 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 1807 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 1808 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 1809. The communication device 1809 may allow the electronic device 1800 to communicate with other devices wirelessly or by wire to exchange data. Although Figure 18 The electronic device 1800 is shown with various devices, but it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed instead.
[0319] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program includes a program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication device 1809, or installed from the storage device 1808, or installed from the ROM 1802. When the computer program is executed by the processing device 1801, the above-mentioned functions defined in the method of the embodiment of the present disclosure are performed.
[0320] It should be noted that the computer-readable medium mentioned above in the present disclosure may be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two. A computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or component, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, a computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, device, or component. In the present disclosure, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any suitable medium, including but not limited to wires, optical cables, RF (radio frequency), etc., or any suitable combination thereof.
[0321] In some embodiments, the client and server can communicate using any currently known or future developed network protocol, such as HTTP (HyperText Transfer Protocol), and can be interconnected with any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network ("LAN"), a wide area network ("WAN"), an internet (e.g., the Internet), and a peer-to-peer network (e.g., an ad hoc peer-to-peer network), as well as any currently known or future developed network.
[0322] The computer-readable medium may be included in the electronic device, or may exist independently without being incorporated into the electronic device.
[0323] The above-mentioned computer-readable medium carries one or more programs. When the above-mentioned one or more programs are executed by the electronic device, the electronic device is enabled to: obtain video material; determine the video rendering parameters required for video rendering based on the video material; call the video template file, and generate a target video for displaying the video material based on the video rendering parameters, wherein the video template file is used to define the video content display process, and the video template file renders the video content of the video frames corresponding to each video content display process based on the video rendering parameters; and output the target video.
[0324] Alternatively, the computer-readable medium carries one or more programs, and when the one or more programs are executed by the electronic device, the electronic device: sends video material to a server; the server is used to determine the video rendering parameters required for video rendering based on the video material, call a video template file, and generate a target video for displaying the video material based on the video rendering parameters, wherein the video template file is used to define the video content display process, and the video template file renders the video content of the video frames corresponding to each video content display process based on the video rendering parameters; receives the target video sent by the server, and displays the target video.
[0325] Alternatively, the computer-readable medium carries one or more programs, and when the one or more programs are executed by the electronic device, the electronic device: displays candidate video implementation logic and candidate functional components, the video implementation logic is used to define the video content display process, and the functional components are used to render video materials; in response to a selection operation, constructs a video template file according to the video implementation logic and the functional components indicated by the selection operation; the video template file is used to render the video content of the video frames corresponding to each of the video content display processes according to the video rendering parameters corresponding to the video material, and obtain the target video.
[0326] Computer program code for performing the operations of the present disclosure may be written in one or more programming languages, or a combination thereof, including, but not limited to, object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on the remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).
[0327] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the module, program segment, or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of the boxes in the block diagram and / or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0328] The modules described in the embodiments of the present disclosure may be implemented in software or hardware, wherein the name of a module does not necessarily limit the module itself.
[0329] The functions described above herein may be performed, at least in part, by one or more hardware logic components. For example, and without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chip (SOCs), complex programmable logic devices (CPLDs), and the like.
[0330] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0331] The above description is merely a preferred embodiment of the present disclosure and an illustration of the technical principles employed. Those skilled in the art should understand that the scope of disclosure involved in the present disclosure is not limited to the technical solutions formed by the specific combination of the above-mentioned technical features, but also includes other technical solutions formed by any combination of the above-mentioned technical features or their equivalents without departing from the above-mentioned disclosed concepts. For example, a technical solution formed by replacing the above-mentioned features with (but not limited to) technical features with similar functions disclosed in this disclosure.
[0332] In addition, although each operation is described in a specific order, this should not be understood as requiring these operations to be performed in the specific order shown or in a sequential order. Under certain circumstances, multitasking and parallel processing may be advantageous. Similarly, although some specific implementation details have been included in the above discussion, these should not be interpreted as limiting the scope of the present disclosure. Some features described in the context of a separate embodiment can also be implemented in a single embodiment in combination. On the contrary, the various features described in the context of a single embodiment can also be implemented in multiple embodiments individually or in any suitable sub-combination mode.
[0333] Although the subject matter has been described using language specific to structural features and / or methodological logical acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are merely example forms of implementing the claims. Regarding the apparatus in the above-described embodiments, the specific manner in which each module performs operations has been described in detail in the embodiments related to the method and will not be elaborated upon here.
Claims
1. A video generation method, characterized in that: include: Get video material; Determining video rendering parameters required for video rendering based on the video material; Calling a video template file, and generating a target video for displaying the video material according to the video rendering parameters, wherein the video template file is used to define a video content display process, and the video template file renders the video content of the video frame corresponding to each video content display process according to the video rendering parameters; The target video is output.
2. The method according to claim 1, characterized in that The calling of the video template file and generating a target video for displaying the video material according to the video rendering parameters include: The target video is generated according to the video rendering parameters through a video rendering service, and the video rendering service is used to generate the target video according to the video template file and the video rendering parameters.
3. The method according to claim 2, characterized in that The video rendering service generates the target video through the following steps: Obtaining the video rendering parameters; Determining metadata information required for video rendering according to the video rendering parameters and the video template file, the metadata information including time metadata corresponding to each video content display process and content data corresponding to the time metadata; Constructing context information required for video rendering according to the metadata information and the video rendering parameters; A rendering engine is called to render the target video according to the functional components indicated by the video template file and the context information.
4. The method according to any one of claims 1 to 3, characterized in that The video template file includes: The video implementation logic for defining a video content presentation process is configured to determine metadata information required for video rendering based on the video rendering parameters, and to construct context information required for video rendering based on the metadata information and the video rendering parameters, wherein the metadata information includes time metadata corresponding to each video content presentation process and content data corresponding to the time metadata; The functional component is used to render the video content corresponding to each video frame according to the context information to obtain the target video.
5. The method according to claim 4, characterized in that The video template file also includes: A root component, used to receive the video rendering parameters; The registration component nested in the root component is used to register the combined component and pass the video rendering parameters to the combined component. The video implementation logic and the functional component are nested in the combined component.
6. The method according to claim 4, characterized in that The video implementation logic includes: A data source layer, comprising a data model and a metadata calculation function, wherein the data model is used to store the video rendering parameters using a preset data structure, and the metadata calculation function is used to determine metadata information required for video rendering based on the video rendering parameters stored in the data model; A context component, configured to construct context information required for video rendering based on the metadata information and the video rendering parameters; The animation calculation layer is used to determine the frame number corresponding to each video frame being rendered; The functional component is used to render the video content corresponding to each video frame according to the context information and the frame number to obtain the target video.
7. The method according to claim 6, characterized in that The animation calculation layer is further used to determine, according to the frame number, the elastic animation value used by the video frame corresponding to the frame number; The functional component is used to render the video content corresponding to each video frame according to the context information, the frame number and the elastic animation value to obtain the target video.
8. The method according to any one of claims 1 to 3, characterized in that The calling of the video template file and generating a target video for displaying the video material according to the video rendering parameters include: Determining the total number of video frames corresponding to the video rendering parameters according to the video rendering parameters and the duration corresponding to the video content display process included in the video template file; Splitting the total number of video frames into multiple video slices according to the preset video segment length, and creating a video rendering task corresponding to each video slice; Calling a video rendering service to execute a plurality of the video rendering tasks in parallel, and obtaining a video clip corresponding to each of the video rendering tasks according to the video rendering parameters and the video template file; The plurality of video clips are spliced together to obtain the target video.
9. The method according to claim 8, characterized in that The preset video clip duration is determined by the following steps: The duration of the video clip is determined based on the resource information of the service cluster where the video rendering service is located, the maximum load of the service cluster, the average rendering speed of the video rendering service in historical time, and the average duration of historical video materials processed by the video rendering service.
10. The method according to claim 8, characterized in that The video slice is a video picture without audio; the method further includes: Determine the total audio duration corresponding to the video rendering parameters according to the video rendering parameters and the duration corresponding to the video content display process included in the video template file; Divide the total audio duration into multiple audio slices according to a preset audio segment duration, and construct an audio rendering task corresponding to each audio slice, wherein the audio segment duration is in a preset ratio to the video segment duration; Calling a video rendering service to execute the plurality of audio rendering tasks in parallel, and obtaining an audio clip corresponding to each audio rendering task according to the video rendering parameters and the video template file; The step of splicing the plurality of video clips to obtain the target video includes: splicing a plurality of the video clips to obtain an initial video without audio; Splicing multiple audio clips to obtain target audio; The initial video and the target audio are spliced together to obtain the target video.
11. The method according to any one of claims 1 to 3, characterized in that The obtaining of video material includes: Get the video material sent by the client through the scheduling service; The step of determining video rendering parameters required for video rendering based on the video material includes: Determining, through the scheduling service, video rendering parameters required for video rendering based on the video material; The calling of the video template file and generating a target video for displaying the video material according to the video rendering parameters include: Determining, by the scheduling service, a plurality of video rendering tasks according to the video template file and the video rendering parameters; By means of a video rendering service, a plurality of the video rendering tasks are executed in parallel to obtain a video clip corresponding to each of the video rendering tasks; the video rendering service is used to generate a video clip corresponding to the video rendering task according to the video template file and the video rendering parameters corresponding to the video rendering task; splicing the plurality of video clips through the scheduling service to obtain the target video; Outputting the target video includes: The target video is returned to the client through the scheduling service.
12. A video generation method, characterized in that: Executed by a client, the method includes: In response to a material selection operation, determining the video material indicated by the material selection operation; Display the target video, wherein the target video is generated according to the video rendering parameters required for video rendering based on the video material, and a video template file is called, wherein the video template file is used to define the video content display process, and the video template file renders the video content of the video frames corresponding to each video content display process according to the video rendering parameters.
13. The method according to claim 12, characterized in that The video template file is determined by the following steps: Display candidate video implementation logic and candidate functional components, wherein the video implementation logic is used to define the video content display process, and the functional components are used to render video materials; In response to a selection operation, a video template file is constructed according to the video implementation logic and the functional components indicated by the selection operation.
14. The method according to claim 13, characterized in that The video template file includes: The video implementation logic for defining a video content presentation process is configured to determine metadata information required for video rendering based on the video rendering parameters, and to construct context information required for video rendering based on the metadata information and the video rendering parameters, wherein the metadata information includes time metadata corresponding to each video content presentation process and content data corresponding to the time metadata; The functional component is used to render the video content corresponding to each video frame according to the context information to obtain the target video.
15. The method according to claim 14, characterized in that The video implementation logic includes: A data source layer, comprising a data model and a metadata calculation function, wherein the data model is used to store the video rendering parameters using a preset data structure, and the metadata calculation function is used to determine metadata information required for video rendering based on the video rendering parameters stored in the data model; A context component, configured to construct context information required for video rendering based on the metadata information and the video rendering parameters; The animation calculation layer is used to determine the frame number corresponding to each video frame being rendered; The functional component is used to render the video content corresponding to each video frame according to the context information and the frame number to obtain the target video.
16. The method according to any one of claims 12 to 15, characterized in that The target video to be displayed includes: Sending the video material to a server; Receive and display the target video sent by the server, the server is used to determine the video rendering parameters required for video rendering based on the video material, and generate the target video according to the video rendering parameters through the video rendering service, and the video rendering service is used to generate the target video according to the video template file and the video rendering parameters.
17. The method according to any one of claims 12 to 15, characterized in that The target video to be displayed includes: Sending the video material to a server; Receive and display the target video sent by the server, the server is used to determine the video rendering parameters required for video rendering based on the video material, determine the total number of video frames corresponding to the video rendering parameters based on the video rendering parameters and the duration corresponding to the video content display process included in the video template file, split the total number of video frames into multiple video slices based on the preset video segment duration, create a video rendering task corresponding to each of the video slices, call the video rendering service, execute multiple video rendering tasks in parallel, obtain the video segment corresponding to each of the video rendering tasks based on the video rendering parameters and the video template file, splice the multiple video segments, and obtain the target video.
18. A computer readable medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processing device, the computer program implements the steps of the method according to any one of claims 1 to 11, or implements the steps of the method according to any one of claims 12 to 17.
19. An electronic device, characterized in that: include: a storage device having a computer program stored thereon; A processing device, configured to execute the computer program in the storage device to implement the steps of the method according to any one of claims 1 to 11, or to implement the steps of the method according to any one of claims 12 to 17.
20. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the computer program implements the steps of the method according to any one of claims 1 to 11, or the computer program implements the steps of the method according to any one of claims 12 to 17.