A video generation method and device, computer equipment and storage medium

By generating event streams in a virtual 3D scene and controlling target objects to perform event actions, the cumbersome 3D animation production process caused by the involvement of multiple software programs in existing technologies is solved, thus achieving a simplified 3D animation production effect.

CN115761064BActive Publication Date: 2026-04-21DOUYIN VISION CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
DOUYIN VISION CO LTD
Filing Date
2022-11-10
Publication Date
2026-04-21

AI Technical Summary

Technical Problem

Current technology requires the use of multiple specialized software programs to create 3D animations, resulting in a cumbersome process, high learning costs for users, and difficulty in completing the task independently.

Method used

By generating a virtual 3D scene and responding to event editing operations of the target object, an event flow is constructed, and the target object is automatically controlled to perform event actions to generate the first target video.

Benefits of technology

It reduces the difficulty of 3D animation production, enabling users to complete 3D animation production in one stop, simplifying the process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115761064B_ABST
    Figure CN115761064B_ABST
Patent Text Reader

Abstract

This disclosure provides a video generation method, apparatus, computer device, and storage medium. The method includes: generating a virtual three-dimensional scene; the virtual three-dimensional scene includes at least one target object; in response to performing an event editing operation on the target object, generating an event stream corresponding to the target object; the event stream includes event information corresponding to multiple nodes; wherein the event information is determined based on the event editing operation and is used to describe the event action performed by the target object; based on the event information corresponding to the multiple nodes in the event stream, controlling the target object to perform the event action corresponding to the event information in the virtual three-dimensional scene, and acquiring a first target video of the target object while controlling the target object to perform the event action.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of computer image processing technology, and more specifically, to a video generation method, apparatus, computer device, and storage medium. Background Technology

[0002] Currently, creating a 3D animation often requires the use of multiple professional software programs. For example, 3D model animation uses Maya and 3ds Max, lighting and rendering use Katana and Maya, and compositing and editing use Nuke and Premiere Pro. This makes the process of creating a 3D animation too complicated, and users need to spend a lot of time learning these professional software programs, making it difficult for users to complete a 3D animation on their own. Summary of the Invention

[0003] This disclosure provides at least one video generation method, apparatus, computer device, and storage medium.

[0004] In a first aspect, embodiments of this disclosure provide a video generation method, including:

[0005] Generate a virtual 3D scene; the virtual 3D scene includes at least one target object;

[0006] In response to performing an event editing operation on the target object, an event stream corresponding to the target object is generated; the event stream includes event information corresponding to multiple nodes respectively; wherein, the event information is determined based on the event editing operation and is used to describe the event action performed by the target object;

[0007] Based on the event information corresponding to multiple nodes in the event stream, the target object is controlled to perform event actions corresponding to the event information in the virtual 3D scene, and when the target object is controlled to perform the event actions, the first target video of the target object is obtained.

[0008] In one optional implementation, it further includes:

[0009] Acquire a second target video of a real-world scene;

[0010] The first target video and the second target video are fused together to obtain a target video that includes the target object and the real scene.

[0011] In one optional implementation, generating the virtual 3D scene includes:

[0012] Generate a virtual three-dimensional space corresponding to the virtual three-dimensional scene;

[0013] Determine the coordinates of at least one target object in the virtual three-dimensional space;

[0014] Based on the coordinates of the target object in the virtual three-dimensional space, the three-dimensional model corresponding to the target object is added to the virtual three-dimensional space to form the virtual three-dimensional scene;

[0015] In one optional implementation, it further includes:

[0016] In response to an addition operation that adds the target object to the virtual 3D scene, the initial features of the target object in the virtual 3D scene are determined; the initial features include at least one of the following: initial pose, initial animation, initial lighting type, and initial camera angle;

[0017] Based on the initial features, the three-dimensional model corresponding to the target object is added to the virtual three-dimensional scene.

[0018] In one optional implementation, the event editing operation on the target object includes:

[0019] Generate the node and receive the basic event materials corresponding to the generated node;

[0020] Based on the aforementioned basic event materials, event information corresponding to the generated nodes is generated.

[0021] In one optional implementation, the node includes: a time node, and generating the node includes:

[0022] Determine the event execution time on the timeline, and generate the corresponding time node based on the event execution time.

[0023] In one optional implementation, controlling the target object to perform corresponding event actions in the virtual 3D scene based on the event flow includes:

[0024] For each time node in the event stream, the target object is controlled to perform the event action corresponding to that time node in the virtual 3D scene;

[0025] If the event action corresponding to a given time node is completed but the next time node corresponding to that time node has not yet arrived, the event action corresponding to that time node is repeated.

[0026] In one optional implementation, the node includes: an event node; the step of controlling the target object to perform corresponding event actions in the virtual 3D scene based on the event flow includes:

[0027] For each event node in the event stream, the target object is controlled to perform the event action corresponding to that event node in the virtual 3D scene, and after the event action corresponding to that event node is completed, the event action corresponding to the next event node is executed.

[0028] In one optional implementation, there are multiple target objects; controlling the target objects to perform event actions corresponding to the event information in the virtual 3D scene based on the event information corresponding to the multiple nodes in the event stream includes:

[0029] Based on the event execution time of multiple nodes in the event streams associated with multiple target objects, the event information in the event streams associated with the multiple target objects is merged to obtain the event execution script;

[0030] Based on the event execution script, multiple target objects are controlled to perform event actions corresponding to each of the multiple target objects in the virtual 3D scene.

[0031] Secondly, embodiments of this disclosure also provide a video generation apparatus, comprising:

[0032] The first generation module is used to generate a virtual 3D scene; the virtual 3D scene includes at least one target object;

[0033] The second generation module is used to generate an event stream corresponding to the target object in response to an event editing operation on the target object; the event stream includes event information corresponding to multiple nodes respectively; wherein the event information is determined based on the event editing operation and is used to describe the event action performed by the target object;

[0034] The first acquisition module is used to control the target object to perform an event action corresponding to the event information in the virtual three-dimensional scene based on the event information corresponding to multiple nodes in the event stream, and to acquire the first target video of the target object when controlling the target object to perform the event action.

[0035] In one optional embodiment, the apparatus further includes a second acquisition module, configured to:

[0036] Acquire a second target video of a real-world scene;

[0037] The first target video and the second target video are fused together to obtain a target video that includes the target object and the real scene.

[0038] In one optional implementation, the first generation module is further configured to:

[0039] Generate a virtual three-dimensional space corresponding to the virtual three-dimensional scene;

[0040] Determine the coordinates of at least one target object in the virtual three-dimensional space;

[0041] Based on the coordinates of the target object in the virtual three-dimensional space, the three-dimensional model corresponding to the target object is added to the virtual three-dimensional space to form the virtual three-dimensional scene;

[0042] In one optional implementation, the apparatus further includes a third generation module, configured to:

[0043] In response to an addition operation that adds the target object to the virtual 3D scene, the initial features of the target object in the virtual 3D scene are determined; the initial features include at least one of the following: initial pose, initial animation, initial lighting type, and initial camera angle;

[0044] Based on the initial features, the three-dimensional model corresponding to the target object is added to the virtual three-dimensional scene.

[0045] In one optional implementation, the second generation module, when performing editing operations on the target object, is further configured to:

[0046] Generate the node and receive the basic event materials corresponding to the generated node;

[0047] Based on the aforementioned basic event materials, event information corresponding to the generated nodes is generated.

[0048] In one optional implementation, the second generation module is further configured to:

[0049] Determine the event execution time on the timeline, and generate the corresponding time node based on the event execution time.

[0050] In one optional implementation, when the second generation module controls the target object to perform corresponding event actions in the virtual 3D scene based on the event stream, it is used to:

[0051] For each time node in the event stream, the target object is controlled to perform the event action corresponding to that time node in the virtual 3D scene;

[0052] If the event action corresponding to a given time node is completed but the next time node corresponding to that time node has not yet arrived, the event action corresponding to that time node is repeated.

[0053] In an optional implementation, when the second generation module controls the target object to perform corresponding event actions in the virtual 3D scene based on the event stream, it is further configured to:

[0054] For each event node in the event stream, the target object is controlled to perform the event action corresponding to that event node in the virtual 3D scene, and after the event action corresponding to that event node is completed, the event action corresponding to the next event node is executed.

[0055] In one optional implementation, when the second generation module controls the target object to perform an event action corresponding to the event information in the virtual 3D scene based on event information corresponding to multiple nodes in the event stream, it is used to:

[0056] Based on the event execution time of multiple nodes in the event streams associated with multiple target objects, the event information in the event streams associated with the multiple target objects is merged to obtain the event execution script;

[0057] Based on the event execution script, multiple target objects are controlled to perform event actions corresponding to each of the multiple target objects in the virtual 3D scene.

[0058] Thirdly, an optional implementation of this disclosure also provides a computer device, a processor, and a memory, wherein the memory stores machine-readable instructions executable by the processor, and the processor is configured to execute the machine-readable instructions stored in the memory. When the machine-readable instructions are executed by the processor, the steps of the first aspect above, or any possible implementation of the first aspect, are performed.

[0059] Fourthly, an optional implementation of this disclosure also provides a computer-readable storage medium storing a computer program that, when run, performs the steps of the first aspect or any possible implementation of the first aspect.

[0060] The video generation method provided in this disclosure reduces the difficulty of video generation by responding to event editing operations on a target object, determining an event stream composed of event information corresponding to multiple nodes, and using the event stream to automatically control the target object to perform event actions.

[0061] Furthermore, during event editing operations, the position of nodes in the event flow can be changed by controlling each node, thereby adjusting the timing of the target object executing the event action corresponding to the event information, and realizing the editing of the execution order and timing of the event actions; through different event editing operations, event information corresponding to the event editing operation is generated, realizing the editing of the execution content of the event action; and by obtaining the first target video of the target object when the target object executes the event action, the output of 3D animation is realized, thus completing the production of 3D animation in one stop and reducing the difficulty of 3D animation production.

[0062] To make the above-mentioned objects, features and advantages of this disclosure more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. Attached Figure Description

[0063] To more clearly illustrate the technical solutions of the embodiments of this disclosure, the accompanying drawings used in the embodiments will be briefly described below. These drawings are incorporated in and constitute a part of this specification. They illustrate embodiments conforming to this disclosure and, together with the specification, serve to explain the technical solutions of this disclosure. It should be understood that the following drawings only show some embodiments of this disclosure and should not be considered as limiting the scope. Those skilled in the art can obtain other related drawings based on these drawings without creative effort.

[0064] Figure 1 A flowchart of a video generation method provided by some embodiments of this disclosure is shown;

[0065] Figure 2 An example diagram illustrating an animation provided by some embodiments of this disclosure is shown;

[0066] Figure 3a This illustration shows one of the example diagrams illustrating an event editing operation on a target object provided by some embodiments of this disclosure;

[0067] Figure 3b This is illustrated as a second example diagram of an event editing operation on a target object provided by some embodiments of this disclosure;

[0068] Figure 4 An example diagram is shown illustrating a merging operation on the overall timeline provided by some embodiments of this disclosure;

[0069] Figure 5 A flowchart is shown below illustrating another video generation method provided by some embodiments of this disclosure;

[0070] Figure 6 A schematic diagram of a video generation apparatus provided by some embodiments of the present disclosure is shown;

[0071] Figure 7 A schematic diagram of a computer device provided by some embodiments of the present disclosure is shown. Detailed Implementation

[0072] To make the objectives, technical solutions, and advantages of the embodiments of this disclosure clearer, the technical solutions of the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this disclosure, and not all of them. The components of the embodiments of this disclosure described and shown herein can generally be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of this disclosure is not intended to limit the scope of the claimed disclosure, but merely represents selected embodiments of this disclosure. All other embodiments obtained by those skilled in the art based on the embodiments of this disclosure without inventive effort are within the scope of protection of this disclosure.

[0073] Research has shown that the production of a 3D animation can be simply divided into three steps: Step 1: 3D model creation, texturing, and material processing; Step 2: Adding narrative and interactive elements, such as the 3D model's movements, speech, scene lighting, and camera effects; Step 3: Adjusting the coordination between the 3D model's movements, speech, scene lighting, and camera effects; and finally, generating the 3D animation.

[0074] Currently, creating a 3D animation often requires the use of multiple professional software programs. For example, 3D model animation uses Maya and 3ds Max, lighting and rendering use Katana and Maya, and compositing and editing use Nuke and Premiere Pro. This makes the process of creating a 3D animation too complicated, and users need to spend a lot of time learning these software programs, making it difficult for users to complete a 3D animation on their own.

[0075] Based on the above research, this disclosure provides a video generation method that determines an event stream composed of event information corresponding to multiple nodes by responding to event editing operations on a target object, and uses the event stream to automatically control the target object to perform event actions to obtain a first target video, thereby reducing the difficulty of video generation.

[0076] The shortcomings of the above solutions are the result of the inventor's practical experience and careful research. Therefore, the discovery process of the above problems and the solutions proposed in this disclosure below should be considered as the inventor's contribution to this disclosure.

[0077] It should be noted that similar labels and letters in the following figures indicate similar items. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures.

[0078] To facilitate understanding of this embodiment, a video generation method disclosed in this disclosure will first be described in detail. The execution entity of the video generation method provided in this disclosure is generally a computer device with certain computing capabilities. This computer device may include, for example, a terminal device, a server, or other processing devices. The terminal device may be a user equipment (UE), mobile device, user terminal, terminal, cellular phone, cordless phone, personal digital assistant (PDA), handheld device, computing device, in-vehicle device, wearable device, etc. In some possible implementations, this video generation method can be implemented by a processor calling computer-readable instructions stored in memory.

[0079] The following describes a video generation method provided by an embodiment of this disclosure.

[0080] See Figure 1 The diagram shows a flowchart of a video generation method provided in this embodiment of the present disclosure. The method includes steps S101 to S103, wherein:

[0081] S101: Generate a virtual 3D scene; the virtual 3D scene includes at least one target object;

[0082] S102: In response to performing an event editing operation on the target object, an event stream corresponding to the target object is generated; the event stream includes event information corresponding to multiple nodes respectively; wherein, the event information is determined based on the event editing operation and is used to describe the event action performed by the target object;

[0083] S103: Based on the event information corresponding to multiple nodes in the event stream, control the target object to perform the event action corresponding to the event information in the virtual three-dimensional scene, and obtain the first target video of the target object when controlling the target object to perform the event action.

[0084] The steps described above in this disclosure will be explained in detail below.

[0085] Regarding step S101 above, the virtual three-dimensional scene can be, for example, a virtual scene generated by computer technology; the virtual three-dimensional scene is presented on the computer screen based on the shooting perspective of a virtual camera, and different perspectives of the virtual three-dimensional scene can be obtained by changing the shooting perspective of the virtual camera; the virtual three-dimensional scene includes at least one target object, which can be, for example, a virtual character, animal or other virtual role controlled by the user and presented in the virtual three-dimensional scene, or a virtual item such as a hat, weapon, scroll, tree, or vegetation.

[0086] In addition, these can also include virtual light sources, virtual cameras, and other elements that display scene effects within a virtual 3D scene. Generally, virtual light sources and virtual cameras are visible to the user during the event flow editing process, but become invisible to the user when entering the effects preview stage after editing is complete.

[0087] In one embodiment provided in this disclosure, generating a virtual three-dimensional scene may include: generating a virtual three-dimensional space corresponding to the virtual three-dimensional scene; determining the coordinate values ​​of at least one target object in the virtual three-dimensional space; and adding a three-dimensional model corresponding to the target object to the virtual three-dimensional space based on the coordinate values ​​of the target object in the virtual three-dimensional space to constitute the virtual three-dimensional scene.

[0088] Here, the virtual 3D space includes a 3D coordinate system. Any coordinate position in the 3D coordinate system corresponds to its coordinate value, and any coordinate value in the 3D coordinate system can be mapped to the corresponding spatial position in the virtual 3D scene.

[0089] For example, the coordinates of a target object in the virtual 3D space can be determined based on the final position of the target object within the virtual 3D space controlled by the user. For instance, if a user selects a target object from the resource library outside the virtual 3D space and drags it into the virtual 3D space, the coordinates of the target object in the virtual 3D space can be determined based on its final position. Here, the final position refers to the location of the target object in the virtual 3D space when the user triggers the release operation.

[0090] In another embodiment provided in this disclosure, in response to an addition operation that adds the target object to the virtual three-dimensional scene, initial features of the target object in the virtual three-dimensional scene are determined; based on the initial features, a three-dimensional model corresponding to the target object is added to the virtual three-dimensional scene.

[0091] For example, when a target object is added to a virtual 3D scene, an initial feature option can pop up on the computer's display interface. The initial feature option includes function options such as adjusting the target object's initial pose, initial animation, initial lighting and shadow type, and initial camera angle. These function options can all have default settings, which users can adjust or accept.

[0092] For example, when a user places a target object at any location in a virtual 3D scene, its initial position can be determined. At this time, the user can adjust the initial orientation of the target object in the initial pose function option, thereby determining the initial pose of the target object in the virtual scene.

[0093] Here, when the target object is a virtual character, the orientation of the virtual character's face can be specified as the basis for orientation adjustment to determine the initial orientation of the virtual character; when the target object is a virtual item, the front of the virtual item can be specified as the basis for orientation adjustment to determine the initial orientation of the virtual item.

[0094] In addition, when the user chooses to accept the default settings, the target object will be added to the virtual 3D scene with a preset initial pose.

[0095] For example, you can define the initial animation of the target object in the initial animation function options.

[0096] Here, when the target object is a virtual character, you can add initial animations such as waving, smiling, and blinking; when the target object is virtual vegetation, you can add initial animations of swaying in the wind.

[0097] In addition, when the user chooses to accept the default settings, the target object will be added to the virtual 3D scene with a preset initial animation. The preset initial animation may be set to no animation, that is, the target object does not perform any action, depending on the actual situation.

[0098] For example, when the target object is a virtual light source, the initial lighting type of the virtual 3D scene can be determined in the initial lighting type function option when adding the virtual light source to the virtual 3D scene. The initial lighting type includes, for example, scene light, ambient light, etc.

[0099] For example, when the target object is a virtual character, the initial camera view can be determined in the initial camera view function options. The initial camera view includes, for example, a close-up of the face, a close-up of the whole body, and a panoramic view.

[0100] Regarding step S102 above, the event editing operation on the target object can include editing at least one of the target object's actions, animations, language, lighting effects, and camera movement effects.

[0101] In one embodiment provided in this disclosure, the node is generated by performing an event editing operation on the target object, and basic event materials corresponding to the generated node are received; based on the basic event materials, event information corresponding to the generated node is generated.

[0102] Here, a node can be represented as a marker of an event message. The event flow of the target object consists of event messages corresponding to multiple nodes. Basic event material can be represented as the content of an event message, such as at least one of the following: a set of actions, a piece of audio, a piece of lighting and shadow effects, a piece of camera movement effects, etc.

[0103] For example, you can add target objects, as well as a set of actions, a voice clip, a lighting effect, etc., to the resource library. The resource library can be preset by software developers, or it can be created by resource publishers using other modeling software, such as 3D MAX, and uploaded to the target application that uses this method. The resource library includes resources such as actions, expressions, character models, lighting effects, and sound effects.

[0104] Resource publishers include the user themselves and other users who upload their created sound, lighting effects, models, animations, and other resources to the resource server. Users then retrieve the resource list and download the relevant resources from the resource server. If the resource was created by the user, they can directly import it into the target application for their own use; if they wish to share it with other users, they can choose to upload it to the resource server.

[0105] For example, a virtual character is added to the character resource library of the virtual 3D scene, and event editing operations are performed on the virtual character to generate nodes; the event editing stage is entered, in which the event editing operation controls the virtual character to move a certain distance in the virtual 3D scene. Specifically, a target position is determined in the virtual 3D scene. After the target position is determined, the special effects preview stage is entered, and the virtual character moves from the current position to the target position. At this time, the virtual character moves from the current position to the target position and basic event materials are generated.

[0106] For example, see Figure 2 The image shown is an example of adding animation. Figure 2 In the event editing operation, when adding an action and facial expression to the virtual character, the action resource library A21 and the expression resource library A22 are loaded in the virtual 3D scene. The user can select the corresponding action in the action resource library A21 and the expression for the changes in the virtual character's facial features in the expression resource library A22. After the virtual character's action and expression are determined, the special effects preview node is entered, and the virtual character performs the action and expression. At this time, the changes in the virtual character's action and expression are received and basic event materials are generated.

[0107] For example, the event editing operation involves controlling a virtual camera to shoot around a virtual character. Specifically, a virtual camera is generated in a virtual 3D scene, and the user can drag the virtual camera to determine its shooting trajectory. When entering the effects preview stage, the virtual 3D scene switches to the perspective of the virtual camera generated in the virtual 3D scene and moves along the shooting trajectory. At this time, the display screen of the virtual 3D scene captured by the virtual camera according to the shooting trajectory is received and basic event materials are generated.

[0108] It should be noted that during the event editing phase, the virtual camera is positioned outside the virtual 3D scene to capture the virtual 3D scene. Users can control the display of the virtual 3D scene using control buttons or by swiping the screen. When users need to add camera movement effects, a virtual camera is added to the virtual 3D scene. Users can drag the virtual camera and use the control buttons to control the display of the virtual 3D scene and determine the shooting trajectory of the virtual camera. When entering the effects preview phase, the display of the virtual 3D scene is shown according to the shooting trajectory of the virtual camera added to the virtual 3D scene.

[0109] In one possible implementation, see Figure 3a One of the example diagrams shown illustrates an event editing operation on a target object. Figure 3a The system includes a virtual scene and multiple controls. The virtual scene includes a virtual character S. The controls include joysticks A31 and A32 for controlling the virtual camera's viewpoint during the event editing phase; a preview control B31 for entering the effects preview phase; an output video control B32 for generating the first target video; an option control B33 for setting, for example, the event execution time and execution order for each node; and an event editing control C31 for generating nodes and performing event editing operations on the virtual character S during the event editing phase. When the user clicks the event editing control C31, three sub-controls pop up: a movement sub-control C311 for editing the movement of the virtual character S; an animation sub-control C312 for editing the animation of the virtual character S; and a lens sub-control C313 for editing the camera angle. Clicking any sub-control generates the node corresponding to that sub-control and enters the event editing phase, such as... Figure 3bThe second example diagram illustrates an event editing operation on a target object. The user clicks the move sub-control C311 to generate a move node and enter the event editing stage. At this time, the preview control B31, output video control B32, and option control B33 in the virtual 3D scene are hidden. The user can drag the joystick A31 to zoom in, zoom out, pan left, and pan right, and drag the joystick A32 to rotate up, down, left, and right, thus changing the display of the virtual 3D scene. The user clicks to determine a target position D1 in the virtual 3D scene, which indicates the position the virtual character S will move to when entering the effects preview stage. Here, after determining one or more target positions, the user clicks the event editing control C31 again to end the event editing stage. The user can then click the preview control B31 to enter the effects preview stage. At this time, the virtual character S will receive and generate basic event materials based on the target position D1 determined in the event editing stage. Event information is generated based on the basic event materials and nodes, and the virtual character S is controlled to move from its current position to the target position D1 based on the event information.

[0110] Here, the generated basic event material can be saved by the user, or it can be saved automatically when the user triggers the next event editing operation.

[0111] In addition, users can click on option control B33 to set the movement speed, motion range, camera movement speed, and number of motion loops for the virtual character S. The specific settings depend on the actual situation, and this disclosure does not impose any restrictions.

[0112] Regarding S103 above, there may be multiple nodes in the event stream of a target object, each corresponding to event information. After performing event editing operations on the target object, event information is generated, which can control the target object to execute the event actions corresponding to the event information.

[0113] Here, nodes include time nodes and event nodes. Depending on the node, the target object is controlled to perform corresponding event actions in the virtual 3D scene, including at least one of the following M1 and M2:

[0114] M1: For each time node, determine the event execution time on the timeline, and generate a time node corresponding to that event execution time based on the event execution time.

[0115] For example, when a user triggers an event editing operation, a timeline corresponding to a target object can be added to the virtual 3D scene. Dragging the cursor on the timeline allows the addition of a time node at any point in time, and event editing can begin at that point. When the event editing ends and the target object is previewed with special effects, basic event materials will be generated starting from the beginning of the timeline until the basic event materials included in the event information corresponding to the last time node are executed.

[0116] Furthermore, in one embodiment provided in this disclosure, for each time node in the event stream, the target object is controlled to perform an event action corresponding to that time node in the virtual three-dimensional scene; in response to the completion of the event action corresponding to that time node and the fact that the next time node corresponding to that time node has not yet arrived, the event action corresponding to that time node is repeatedly executed.

[0117] For example, taking a virtual character as the target object, when a virtual character has multiple time nodes, an event flow can be generated in chronological order based on the event information corresponding to the multiple time nodes. For example, the event flow of a virtual character includes event information corresponding to three time nodes, namely event A corresponding to the virtual character's "waving" animation, event B corresponding to the virtual character's "moving from point D1 to point D2" animation, and event C corresponding to the virtual character's "smiling" animation. Here, the three events are generated into an event flow in the node order of event A, event B, and event C, where the time node of event A is 0s on the timeline and the event execution time is 1s; the time node of event B is 3s on the timeline and the event execution time is 2s; and the time node of event C is 5s and the event execution time is 1s. At this point, when entering the special effects preview stage, the virtual character is controlled to execute corresponding event actions in the order of each time node on the timeline. For example, at the beginning, the virtual character is controlled to perform an event A, waving. Based on the event execution time of event A and the time node of the next event B, the waving action is repeated 3 times. Then, event B is executed to control the virtual character to move from point D1 to point D2. Here, based on the event execution time of event B and the time node of the next event C, only one movement event action is needed. If multiple movements are required, the movement can be stopped and maintained when the character first moves to point D2, waiting for the time node of event C; or the movement can be stopped when the character first moves to point D2, waiting for the time node of event C; or, through some detection methods, if a situation occurs where the character moves to the target position but the next event has not been triggered, it is considered that there is a lack of smoothness in the action, and a prompt message is displayed to remind the user to adjust the position, or the action can be adjusted automatically.

[0118] In another conceivable embodiment, to achieve the "wave while moving" action, the event information corresponding to two or more time points needs to be overlapped on the timeline. For example, the virtual character's event stream includes event information corresponding to two time points: event A corresponding to the virtual character's "wave" animation, and event B corresponding to the virtual character "moving from point D1 to point D2". This allows these two events to overlap on the timeline, that is, the time point of event B is inserted into the execution time interval of event A. Here, event A and event B can be strictly synchronized; for example, the time points of event A and event B are the same, and the event execution times are also the same. In this way, an effect of "wave while moving" can be obtained, where waving begins when movement starts and stopping when movement stops. Alternatively, events can be executed at different times. For example, event A's time point is 0s on the timeline, and the event execution time is 3s; event B's time point is 1s on the timeline, and the event execution time is 2s. This creates the effect of a virtual character waving at one location for one second, then waving while moving to the next location, and stopping waving after reaching the next location.

[0119] M2: For each event node in the event stream, control the target object to perform the event action corresponding to that event node in the virtual 3D scene, and after the event action corresponding to that event node is completed, execute the event action corresponding to the next event node.

[0120] For example, an event node only determines its order with other event nodes, starting from the first event on the event axis. For instance, an event flow executes events A, B, and C in the order of the event nodes. The user can change the event actions performed by the target object by adjusting the position of the event node relative to other event nodes in the event flow. For example, taking a virtual character as the target object, event A is an animation of "waving," event B is an animation of "clapping," and event C is an animation of "bowing." According to the event flow before adjustment, the virtual character performs the event actions of waving, clapping, and bowing. Now, if the user adjusts event B before event A, the adjusted event flow becomes event B, event A, and event C. Now, according to the adjusted event flow, the virtual character performs the event actions of clapping, waving, and bowing.

[0121] Here, multiple event axes can also be divided in the event flow according to the event type, and multiple event axes can be executed simultaneously when an event action is performed.

[0122] In one embodiment provided in this disclosure, there are multiple target objects; based on the event execution time of multiple nodes in the event streams associated with the multiple target objects respectively, the event information in the event streams associated with the multiple target objects is merged to obtain an event execution script; based on the event execution script, the multiple target objects are controlled to perform event actions corresponding to the multiple target objects respectively in the virtual three-dimensional scene.

[0123] For example, in a virtual 3D scene, multiple target objects can be added. For instance, in a virtual 3D scene, a target object A1 corresponding to a virtual character, a target object A2 corresponding to a virtual pet that flies with the virtual character, a target object A3 corresponding to a virtual light source that produces a "stage light" effect following the movement of the virtual character, a target object A4 corresponding to a virtual stage, and a target object A5 corresponding to a virtual camera that follows the movement of the virtual character. Each of these target objects can be independently edited to generate its own event stream. These event streams can then be merged on a single timeline to obtain the event execution script.

[0124] For example, target object A41 and target object A42 are placed in a virtual 3D scene. The event stream of target object A41 includes event information a411, event information a412, and event information a413, while the event stream of target object A42 includes event information a421 and event information a422. When merging the event streams of target object A41 and target object A42, the merging can be performed based on the time nodes and event execution events included in the event information of each event stream. See also Figure 4 The diagram shown is an example of a merging operation performed on the overall timeline. Figure 4 In the event stream, event information a411 has a time node of 0s and an execution time of 1s; event information a412 has a time node of 1s and an execution time of 3s; event information a413 has a time node of 4s and an execution time of 2s; event information a421 has a time node of 0s and an execution time of 2s; and event information a422 has a time node of 3s and an execution time of 3s. Based on the time nodes and execution times included in the event information in each event stream, the execution order of the events on the total time axis T is: [(a411, a421), a412, a422, a413], where (a411, a421) indicates simultaneous execution.

[0125] Here, the execution order of events in the event execution script can be changed by adjusting the time nodes and event execution times in the event information. The specific order depends on the actual situation, and this disclosure does not impose any restrictions.

[0126] In addition, when controlling the target object to perform the event action, the first target video of the target object is acquired.

[0127] In one embodiment provided in this disclosure, a second target video of a real-world scene is acquired; the first target video and the second target video are fused together to obtain a target video including the target object and the real-world scene.

[0128] For example, target effects are generated based on the content of the first target video and added to the target application. Before the user shoots a real-world scene using the target application's camera function, they can choose to add these target effects. At this time, target effects composed of virtual characters, virtual objects, and virtual effects included in the first target video are generated in the shooting screen. Outside the area of ​​these target effects, the real-world scene captured by the user's terminal device's camera is generated. When the user starts recording, the first target video is played synchronously. The virtual characters, virtual objects, and virtual effects in the first target video begin to execute event actions, i.e., target effects, according to the event execution script. The user can make corresponding synchronized actions based on the target effects to obtain the target video.

[0129] In another example, the second target video includes a video already shot by the user. This video can be fused with the target effects presented in the first target video and the real-world scene in the second target video. For example, the speed, model size, and model position of the first and / or second target videos can be adjusted. These adjustments can be made manually by the user, or by using a neural network algorithm to identify the positions of various models in the first target video and the target objects in the second target video that are in sync with the models. A preset algorithm is then used to adjust the size and relative position of the models and the target objects in sync, ultimately resulting in the target video.

[0130] Additionally, see Figure 5 As shown, this disclosure also provides a flowchart of another video generation method, including steps S501 to S505.

[0131] Step S501: Add the target object.

[0132] Add target object A, target object B, and target object C to the virtual 3D scene. Target objects can include virtual characters, virtual animals, virtual items, etc.

[0133] Step S502: Generate an event stream.

[0134] Click the "Add Event" button in the virtual 3D scene, and select "Add Movement Event," "Animation Event," "Light and Shadow Effect Event," or "Camera Effect Event" from the pop-up event types. This will generate event flow A for target object A, event flow B for target object B, and event flow C for target object C. Each event flow can include multiple event information.

[0135] Step S503: Generate green screen video.

[0136] Event streams A, B, and C are merged to obtain the green screen video, also known as the first target video. The green screen area here refers to the area in the virtual 3D scene excluding the area occupied by the target object.

[0137] Step S504: Take a video with the user.

[0138] The user video and the green screen video are combined and processed, so that the target object in the green screen video appears on the user video, and the green screen area is filled with the content of the user video.

[0139] Here, the light and shadow effects determine whether the user's video can overlap with the area where the light and shadow effects are applied based on the transparency of the light and shadow effects. The lower the transparency, the more difficult it is for the content in the user's video to be displayed in the area where the light and shadow effects are applied.

[0140] Those skilled in the art will understand that, in the above-described method of the specific implementation, the order in which each step is written does not imply a strict execution order and does not constitute any limitation on the implementation process. The specific execution order of each step should be determined by its function and possible internal logic.

[0141] Based on the same inventive concept, this disclosure also provides a video generation device corresponding to the video generation method. Since the principle of the device in this disclosure for solving the problem is similar to that of the video generation method described above, the implementation of the device can refer to the implementation of the method, and the repeated parts will not be described again.

[0142] Reference Figure 6 The diagram shown is a schematic representation of a video generation apparatus provided in an embodiment of this disclosure. The apparatus includes: a first generation module 61, a second generation module 62, and a first acquisition module 63; wherein,

[0143] This disclosure also provides a video generation apparatus, including:

[0144] The first generation module 61 is used to generate a virtual three-dimensional scene; the virtual three-dimensional scene includes at least one target object;

[0145] The second generation module 62 is used to generate an event stream corresponding to the target object in response to an event editing operation on the target object; the event stream includes event information corresponding to multiple nodes respectively; wherein the event information is determined based on the event editing operation and is used to describe the event action performed by the target object;

[0146] The first acquisition module 63 is used to control the target object to perform an event action corresponding to the event information in the virtual three-dimensional scene based on the event information corresponding to multiple nodes in the event stream, and to acquire the first target video of the target object when controlling the target object to perform the event action.

[0147] In an optional embodiment, the device further includes a second acquisition module 64, configured to:

[0148] Acquire a second target video of a real-world scene;

[0149] The first target video and the second target video are fused together to obtain a target video that includes the target object and the real scene.

[0150] In one optional implementation, the first generation module 61 is further configured to:

[0151] Generate a virtual three-dimensional space corresponding to the virtual three-dimensional scene;

[0152] Determine the coordinates of at least one target object in the virtual three-dimensional space;

[0153] Based on the coordinates of the target object in the virtual three-dimensional space, the three-dimensional model corresponding to the target object is added to the virtual three-dimensional space to form the virtual three-dimensional scene;

[0154] In an optional embodiment, the apparatus further includes a third generation module 65, configured to:

[0155] In response to an addition operation that adds the target object to the virtual 3D scene, the initial features of the target object in the virtual 3D scene are determined; the initial features include at least one of the following: initial pose, initial animation, initial lighting type, and initial camera angle;

[0156] Based on the initial features, the three-dimensional model corresponding to the target object is added to the virtual three-dimensional scene.

[0157] In an optional implementation, the second generation module 62, when performing editing operations on the target object, is further configured to:

[0158] Generate the node and receive the basic event materials corresponding to the generated node;

[0159] Based on the aforementioned basic event materials, event information corresponding to the generated nodes is generated.

[0160] In an optional implementation, the second generation module 62 is further configured to:

[0161] Determine the event execution time on the timeline, and generate the corresponding time node based on the event execution time.

[0162] In one optional implementation, when the second generation module 62 controls the target object to perform corresponding event actions in the virtual 3D scene based on the event stream, it is used to:

[0163] For each time node in the event stream, the target object is controlled to perform the event action corresponding to that time node in the virtual 3D scene;

[0164] If the event action corresponding to a given time node is completed but the next time node corresponding to that time node has not yet arrived, the event action corresponding to that time node is repeated.

[0165] In an optional implementation, when the second generation module 62 controls the target object to perform corresponding event actions in the virtual 3D scene based on the event stream, it is further configured to:

[0166] For each event node in the event stream, the target object is controlled to perform the event action corresponding to that event node in the virtual 3D scene, and after the event action corresponding to that event node is completed, the event action corresponding to the next event node is executed.

[0167] In one optional implementation, when the second generation module 62 controls the target object to perform an event action corresponding to the event information in the virtual 3D scene based on event information corresponding to multiple nodes in the event stream, it is used to:

[0168] Based on the event execution time of multiple nodes in the event streams associated with multiple target objects, the event information in the event streams associated with the multiple target objects is merged to obtain the event execution script;

[0169] Based on the event execution script, multiple target objects are controlled to perform event actions corresponding to each of the multiple target objects in the virtual 3D scene.

[0170] The processing flow of each module in the device and the interaction flow between each module can be referred to the relevant descriptions in the above method embodiments, and will not be detailed here.

[0171] This disclosure also provides a computer device, such as... Figure 7 The diagram shown is a schematic representation of a computer device structure provided in an embodiment of this disclosure, including:

[0172] A processor 71 and a memory 72; the memory 72 stores machine-readable instructions executable by the processor 71, and the processor 71 executes the machine-readable instructions stored in the memory 72. When the machine-readable instructions are executed by the processor 71, the processor 71 performs the following steps:

[0173] Generate a virtual 3D scene; the virtual 3D scene includes at least one target object;

[0174] In response to performing an event editing operation on the target object, an event stream corresponding to the target object is generated; the event stream includes event information corresponding to multiple nodes respectively; wherein, the event information is determined based on the event editing operation and is used to describe the event action performed by the target object;

[0175] Based on the event information corresponding to multiple nodes in the event stream, the target object is controlled to perform event actions corresponding to the event information in the virtual 3D scene, and when the target object is controlled to perform the event actions, the first target video of the target object is obtained.

[0176] The aforementioned memory 72 includes a main memory 721 and an external memory 722; the main memory 721, also known as internal memory, is used to temporarily store the computational data in the processor 71, as well as the data exchanged with external memory 722 such as a hard disk. The processor 71 exchanges data with the external memory 722 through the main memory 721.

[0177] The specific execution process of the above instructions can be referred to the steps of the video generation method described in the embodiments of this disclosure, and will not be repeated here.

[0178] This disclosure also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, performs the steps of the video generation method described in the above-described method embodiments. The storage medium may be a volatile or non-volatile computer-readable storage medium.

[0179] This disclosure also provides a computer program product carrying program code. The program code includes instructions that can be used to execute the steps of the video generation method described in the above method embodiments. For details, please refer to the above method embodiments, which will not be repeated here.

[0180] The aforementioned computer program product can be implemented through hardware, software, or a combination thereof. In one optional embodiment, the computer program product is specifically embodied in a computer storage medium; in another optional embodiment, the computer program product is specifically embodied in a software product, such as a software development kit (SDK), etc.

[0181] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems and devices described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here. In the several embodiments provided in this disclosure, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of units is only a logical functional division; in actual implementation, there may be other division methods. Furthermore, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Another point is that the displayed or discussed mutual coupling or direct coupling or communication connection may be through some communication interfaces; the indirect coupling or communication connection of devices or units may be electrical, mechanical, or other forms.

[0182] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0183] In addition, the functional units in the various embodiments of this disclosure can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.

[0184] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a processor-executable, non-volatile, computer-readable storage medium. Based on this understanding, the technical solution of this disclosure, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this disclosure. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0185] Finally, it should be noted that the above-described embodiments are merely specific implementations of this disclosure, used to illustrate the technical solutions of this disclosure, and not to limit it. The protection scope of this disclosure is not limited thereto. Although this disclosure has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that any person skilled in the art can still modify or easily conceive of changes to the technical solutions described in the foregoing embodiments, or make equivalent substitutions for some of the technical features, within the scope of the technology disclosed in this disclosure. Such modifications, changes, or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this disclosure, and should all be covered within the protection scope of this disclosure. Therefore, the protection scope of this disclosure should be determined by the protection scope of the claims.

Claims

1. A video generation method, characterized in that, include: Generate virtual 3D scenes; The virtual 3D scene includes at least one target object; In response to performing an event editing operation on the target object, an event stream corresponding to the target object is generated; the event stream includes event information corresponding to multiple nodes respectively; wherein, the event information is determined based on the event editing operation and is used to describe the event action performed by the target object, the multiple nodes include time nodes on the time axis and event nodes on the event axis, and the event stream is generated according to the event information corresponding to the multiple nodes in the order of the multiple nodes; Based on the event information corresponding to multiple nodes in the event stream, the target object is controlled to perform an event action corresponding to the event information in the virtual 3D scene, and while the target object is controlled to perform the event action, the first target video of the target object is obtained by shooting the target object; The process of generating an event flow corresponding to the target object includes: triggering an event addition control in the virtual 3D scene, selecting an event of the target type from the pop-up event types, and generating an event flow for the target object in the order of addition.

2. The method according to claim 1, characterized in that, The method further includes: Acquire a second target video of a real-world scene; The first target video and the second target video are fused together to obtain a target video that includes the target object and the real scene.

3. The method according to claim 1, characterized in that, The generation of the virtual 3D scene includes: Generate a virtual three-dimensional space corresponding to the virtual three-dimensional scene; Determine the coordinates of at least one target object in the virtual three-dimensional space; Based on the coordinates of the target object in the virtual three-dimensional space, the three-dimensional model corresponding to the target object is added to the virtual three-dimensional space to form the virtual three-dimensional scene.

4. The method according to any one of claims 1-3, characterized in that, Also includes: In response to an addition operation that adds the target object to the virtual 3D scene, the initial features of the target object in the virtual 3D scene are determined; The initial features include at least one of the following: initial pose, initial animation, initial lighting type, and initial camera angle; Based on the initial features, the three-dimensional model corresponding to the target object is added to the virtual three-dimensional scene.

5. The method according to claim 1, characterized in that, The event editing operation on the target object includes: Generate the node and receive the basic event materials corresponding to the generated node; Based on the aforementioned basic event materials, event information corresponding to the generated nodes is generated.

6. The method according to claim 5, characterized in that, The node includes the time node, and generating the node includes: The event execution time is determined on the timeline, and a time node corresponding to the event execution time is generated based on the event execution time.

7. The method according to claim 6, characterized in that, The step of controlling the target object to perform event actions corresponding to the event information in the virtual 3D scene based on the event information corresponding to multiple nodes in the event stream includes: For each time node in the event stream, the target object is controlled to perform the event action corresponding to the time node in the virtual 3D scene; If the event action corresponding to the time node is completed but the next time node corresponding to the time node has not arrived, the event action corresponding to the time node is repeated.

8. The method according to claim 5, characterized in that, The node includes the event node; controlling the target object to perform an event action corresponding to the event information in the virtual 3D scene based on the event information corresponding to multiple nodes in the event stream includes: For each event node in the event stream, the target object is controlled to perform the event action corresponding to the event node in the virtual 3D scene, and after the event action corresponding to the event node is completed, the event action corresponding to the next event node is executed.

9. The method according to claim 1, characterized in that, There are multiple target objects; the step of controlling the target objects to perform event actions corresponding to the event information in the virtual 3D scene based on the event information corresponding to the multiple nodes in the event stream includes: Based on the event execution time of multiple nodes in the event streams associated with multiple target objects, the event information in the event streams associated with the multiple target objects is merged to obtain the event execution script; Based on the event execution script, multiple target objects are controlled to perform event actions corresponding to each of the multiple target objects in the virtual 3D scene.

10. A video generation apparatus, characterized in that, The device includes: The first generation module is used to generate a virtual 3D scene; the virtual 3D scene includes at least one target object; The second generation module is used to generate an event stream corresponding to the target object in response to an event editing operation on the target object. The event stream includes event information corresponding to multiple nodes. The event information is determined based on the event editing operation and is used to describe the event actions performed by the target object. The multiple nodes include time nodes on the time axis and event nodes on the event axis. The event stream is generated according to the event information corresponding to the multiple nodes in the order of the multiple nodes. The first acquisition module is used to control the target object to perform an event action corresponding to the event information in the virtual three-dimensional scene based on the event information corresponding to multiple nodes in the event stream, and to obtain a first target video of the target object by shooting the target object while controlling the target object to perform the event action; When generating an event stream corresponding to the target object, the second generation module is used to: trigger the event addition control in the virtual 3D scene, select the target type of event from the pop-up event types, and generate an event stream for the target object according to the addition order.

11. A computer device, characterized in that, include: The processor and the memory, the memory storing machine-readable instructions executable by the processor, the processor executing the machine-readable instructions stored in the memory, wherein when the machine-readable instructions are executed by the processor, the processor performs the steps of the video generation method as described in any one of claims 1 to 9.

12. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, which, when executed by a computer device, performs the steps of the video generation method as described in any one of claims 1 to 9.

Citation Information

Patent Citations

  • Video generation method and device, electronic equipment and storage medium

    CN114245099A

  • Method for processing a computer-animated scene and corresponding device

    US20140253561A1