Video generation method, electronic equipment and storage medium

By creating a virtual 3D environment for the experience cabin and using multi-view camera synchronous shooting technology, the problem of spatial mismatch in multi-screen video production was solved, achieving improved image consistency and immersion.

CN121603744APending Publication Date: 2026-03-03CHANGZHOU CHINA DINOSAUR LAND CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511674148.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-14
Publication Date
2026-03-03

AI Technical Summary

Technical Problem

Existing technologies lack virtual environment modeling for multi-display screens when creating videos for experience cabins, resulting in a mismatch between the spatial position of the images and the virtual 3D environment, which affects the user experience.

Method used

By acquiring scene files and special effects files, a virtual 3D environment corresponding to the experience cabin is created. Multi-view cameras are set up, and their movement trajectory and rotation posture are controlled to shoot synchronously, generating video files for synchronous playback on multiple displays.

Benefits of technology

Ensure that the display effect on the screen is consistent with the user's viewing angle, reduce jarring sensations, enhance visual immersion, avoid content gaps and disproportion, and improve the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121603744A_ABST
    Figure CN121603744A_ABST
Patent Text Reader

Abstract

The invention provides a video generation method, electronic equipment and a storage medium. The method comprises the following steps: integrating a scene file and a special effect file through third three-dimensional software to obtain a three-dimensional main scene with a special effect; creating a virtual three-dimensional environment corresponding to the experience cabin; setting a virtual camera corresponding to each display area in the virtual three-dimensional environment, and taking the virtual three-dimensional environment with the virtual cameras as a multi-view camera; creating a motion track and a rotation posture of the multi-view camera; and in the three-dimensional main scene, controlling the multi-view camera to move based on the movement track and the rotation attitude, and performing synchronous shooting through each virtual camera in the movement process to obtain a video file for synchronous playing of the M display screens. Therefore, the problem that the picture is not matched with the spatial position of the virtual three-dimensional environment easily during the production of the video played by the multiple display screens of the experience cabin can be improved, and the user experience can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of video production technology, and more specifically, to a video generation method, an electronic device, and a storage medium. Background Technology

[0002] The experience cabin provides an immersive experience by combining video content played on multiple displays. Multiple displays installed on the interior walls of the cabin allow for multi-view video playback, enhancing the immersive experience. However, for visual simulations of entering a virtual 3D environment (such as a Jurassic World setting), using traditional panoramic screens instead of multiple displays for video playback would require rendering the entire environment during video production, processing massive amounts of pixel and sampling data, resulting in high rendering costs and long processing times per session.

[0003] If multiple displays are used for video playback, each display needs to serve as a visual window for the user in a virtual 3D environment. The video displayed on each display simulates the various scenes in the virtual 3D environment as seen by the user through these windows. Due to the unique physical location (such as installation angle and spacing), size, and spatial layout of each display in the experience cabin, there is currently a lack of virtual environment modeling for such multi-distributed displays during video production. This may lead to defects in the produced video content. For example, the images displayed on the displays from different perspectives may not match the spatial location of the virtual 3D environment, resulting in noticeable content breaks, repetitions, or disproportionate effects that negatively impact the user experience. Summary of the Invention

[0004] In view of this, the purpose of this application is to provide a video generation method, electronic device and storage medium that can improve the problem that the video played on multiple screens in the production experience cabin is prone to mismatch between the spatial position of the screen and the virtual three-dimensional environment, which affects the user experience.

[0005] To achieve the above technical objectives, the technical solution adopted in this application is as follows:

[0006] In a first aspect, embodiments of this application provide a video generation method, the method comprising:

[0007] Acquire scene files obtained using the first 3D software and special effects files obtained using the second 3D software;

[0008] By integrating the scene files and special effects files using third-party 3D software, a 3D main scene with special effects is obtained.

[0009] In the main three-dimensional scene, a virtual three-dimensional environment corresponding to the experience cabin is created. The virtual three-dimensional environment includes a display area corresponding to M displays in the experience cabin. The M displays are distributed on at least three inner walls of the experience cabin, where M is an integer greater than or equal to 3.

[0010] Within the virtual 3D environment, a virtual camera corresponding to each display area is set up, and the virtual 3D environment with the virtual camera is used as a multi-view camera. The field of view of the virtual camera is obtained based on the position and field of view of the display screen in the experience cabin.

[0011] In the main 3D scene, create the motion trajectory and rotation posture of the multi-view camera;

[0012] In the three-dimensional main scene, the multi-view camera is controlled to move based on the motion trajectory and rotation posture, and during the movement, each virtual camera is used to shoot synchronously to obtain a video file for synchronous playback on the M displays.

[0013] Secondly, embodiments of this application also provide an electronic device, which includes a processor and a memory coupled to each other. The memory stores a computer program, and when the computer program is executed by the processor, the electronic device performs the above-described method.

[0014] Thirdly, embodiments of this application also provide a computer-readable storage medium storing a computer program that, when run on a computer, causes the computer to perform the above-described method.

[0015] The invention employing the above technical solution has the following advantages:

[0016] In the technical solution provided in this application, by creating a virtual 3D environment corresponding to the motion cabin, the actual distribution of the displays inside the motion cabin can be replicated, ensuring that the layout of the virtual environment is completely aligned with the physical space. This eliminates potential problems caused by deviations in the position and size of the display area, ensuring spatial consistency of the displayed image. The shooting field of view of the virtual camera is obtained based on the position and field of view of the displays inside the experience cabin. This helps ensure that the final display effect on the screen meets the user's visual expectations, reducing "jumping" and enhancing visual immersion. By controlling the multi-view camera to move based on the motion trajectory and rotation posture and to shoot synchronously, multiple virtual cameras share the same motion logic and timeline. This eliminates phenomena such as motion trajectory misalignment and shooting delay, ensuring that the videos played on the M displays are completely synchronized in terms of content progress and motion rhythm. This avoids obvious content breaks, repetitions, or proportional imbalances at the edges of adjacent screens, ensuring that the multi-display screens form a coherent overall visual space, thereby improving the user experience. Attached Figure Description

[0017] This application can be further illustrated by the non-limiting embodiments given in the accompanying drawings. It should be understood that the following drawings only illustrate some embodiments of this application and should not be considered as limiting the scope. For those skilled in the art, other related drawings can be obtained from these drawings without any inventive effort.

[0018] Figure 1 This is a flowchart illustrating the video generation method provided in an embodiment of this application.

[0019] Figure 2 This is a schematic diagram illustrating the positioning of a virtual camera in a virtual three-dimensional environment, as provided in an embodiment of this application. Detailed Implementation

[0020] The present application will be described in detail below with reference to the accompanying drawings and specific embodiments. It should be noted that similar or identical parts are referred to by the same reference numerals in the drawings or description. Implementations not shown or described in the drawings are forms known to those skilled in the art. In the description of this application, terms such as "first" and "second" are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.

[0021] Please refer to Figure 1 This application provides a video generation method, which can be applied to an electronic device and can be executed or implemented by the electronic device. The electronic device can be, but is not limited to, a personal computer, a server, or a combination thereof; no specific limitation is made here. The video generation method may include the following steps:

[0022] Step 110: Obtain the scene file obtained through the first 3D software and the special effects file obtained through the second 3D software;

[0023] Step 120: Using third-party 3D software, integrate the scene files and special effects files to obtain a 3D main scene with special effects;

[0024] Step 130: In the main three-dimensional scene, create a virtual three-dimensional environment corresponding to the experience cabin. The virtual three-dimensional environment includes display areas corresponding to M displays in the experience cabin. The M displays are distributed on at least three inner walls of the experience cabin, where M is an integer greater than or equal to 3.

[0025] Step 140: In the virtual three-dimensional environment, set up a virtual camera corresponding to each of the display areas, and use the virtual three-dimensional environment with the virtual camera as a multi-view camera. The shooting field of view of the virtual camera is obtained based on the position and field of view of the display screen in the experience cabin.

[0026] Step 150: In the three-dimensional main scene, create the motion trajectory and rotation posture of the multi-view camera;

[0027] Step 160: In the three-dimensional main scene, control the multi-view camera to move based on the motion trajectory and rotation posture, and during the movement, synchronously shoot through each of the virtual cameras to obtain a video file for synchronous playback on the M displays.

[0028] The following is a detailed explanation of each step in the video generation method:

[0029] Prior to step 110, the method may further include steps of creating / generating scene files and effects files. For example, prior to step 110, the method may further include:

[0030] Step 010: In the first 3D software, based on the pre-created import plugin, the path of the target file is obtained through clipboard operation. The target file stores a variety of pre-created first asset files for scene construction. The first asset files include object models, textures, hierarchical structures and material binding relationships.

[0031] Step 020: Based on the path, import all first asset files that meet the preset conditions into the first 3D software to form an initial 3D main scene. In the initial 3D main scene, through the format conversion plugin in the first 3D software, convert the images representing materials or textures in all the first asset files into image files of a specified format, including TX format.

[0032] Step 030: Export the initial 3D main scene in USD (Universal Scene Description) format to obtain the scene file.

[0033] In this embodiment, the first 3D software, the second 3D software, and the third 3D software are all conventional software used for 3D animation production, and each has its own advantages.

[0034] The first 3D software is software suitable for creating scenes, such as, but not limited to, Clarisse.

[0035] The second 3D software is software suitable for creating special effects, such as, but not limited to, Houdini.

[0036] Third-dimensional software refers to software suitable for creating character movements and skeletal rigging animations, such as, but not limited to, Maya.

[0037] Clarisse, Houdini, and Maya are all commonly used software for 3D animation production. Combining the strengths of multiple 3D software programs in video production can improve video quality and enhance the user's visual experience. The following will use these three software programs as examples to illustrate the implementation process of this method.

[0038] In this embodiment, step 010, in Clarisse, the construction of a complex virtual 3D scene is typically composed of various asset files, each used to form objects within the virtual 3D scene. For example, for a "Jurassic World" scene, different objects in the scene, such as different types of trees and rocks, can be created separately, forming their own asset files. Each asset file may include, but is not limited to, object models (e.g., the 3D structure of rocks), textures (e.g., the external texture of rocks), hierarchical structures (e.g., trees containing trunks, leaves, fruits / flowers, etc.), and material binding relationships (e.g., whether it is a reflective material), all of which can be flexibly set according to the actual situation.

[0039] After creating each asset file, it needs to be imported into Clarisse for integration. During asset file import, rendering layers and object properties are automatically assigned based on predefined templates, thereby improving asset loading and scene building efficiency. Currently, importing asset files into Clarisse is inefficient, typically involving manual import of single asset files, and cannot handle batch imports. Traditional manual import methods are not only inefficient but also prone to problems such as lost paths, mismatched materials, and inconsistent file versions, severely impacting production schedules and rendering quality. Therefore, this embodiment utilizes a dedicated semi-automated import plugin to deeply optimize the asset loading process.

[0040] In this embodiment, the import plugin is used to obtain the path of the target file through clipboard operations. For example, the system clipboard is accessed through win32clipboard, and the uniformly encoded text is read first with CF_UNICODETEXT. When file drag-and-drop data is detected, HDROP (file path array) is parsed to support batch import of multiple asset files, and self-adaptive parsing is performed on the content of the target file to import asset files that meet preset conditions.

[0041] Step 020, based on the path, import all first asset files that meet preset conditions into the first 3D software to form an initial 3D main scene, which may include:

[0042] Based on the path, a first initial asset file representing the target asset is determined in the target file, and a second initial asset file representing the material or texture is determined in the target file. The first initial asset file and the second initial asset file are used as the first asset file that meets the preset conditions. The target asset includes any one of the asset files in USD format, ABC format, and OBJ format.

[0043] Understandably, the conditions for satisfying the preset conditions can be one or more combinations of the following:

[0044] The target file contains USD / ABC / OBJ asset paths (including relative / absolute / UNC), that is, it includes any asset file in USD format, ABC format, or OBJ format (the first initial asset file). It should be noted that USD format, ABC format, and OBJ format are all common data formats for outputting asset files during the 3D animation production process.

[0045] The target file contains material / texture paths, meaning there is a second initial asset file representing the material or texture.

[0046] Importing plugins can automatically discard non-compliant paths and restore the default calling mode.

[0047] In this embodiment, the image format representing the material or texture may include, but is not limited to, PNG, JPEG, EXR, and other formats.

[0048] After importing the various asset files into Clarisse, designers can integrate them according to their needs, such as adjusting the position and size of the models in the 3D scene, to form the initial 3D main scene.

[0049] In Clarisse, the format conversion plugin allows for batch conversion of images representing materials or textures in the first asset file into image files of a specified format, thus improving the efficiency of image format conversion. The specified format can be TX format.

[0050] The TX format uses block compression technology, which can compress texture image data to 1 / 4 to 1 / 8 of its original size. In addition, the TX format supports on-demand tile loading (Auto-tile). In Clarisse, for images representing materials or textures, the format conversion plugin can convert formats such as PNG, JPEG, and EXR to TX format, which helps optimize rendering performance, save memory resources, and improve the processing efficiency of complex scenes.

[0051] USD format is an open-source 3D scene description format with cross-platform compatibility, supported by mainstream software such as Clarisse, Houdini, Maya, and UE4. In Clarisse, the initial 3D main scene can be exported in USD format to form a scene file, which facilitates subsequent integration of scene files and effects files in Maya.

[0052] When exporting the initial 3D main scene in USD format, the process also includes layered compression and on-demand loading of the USD file: the USD file of the initial 3D main scene is divided into multiple layers according to element type (object model, texture, material, animation data); each layer file is compressed using the LZ4 compression algorithm (compression ratio 1:3-1:5), and a layer index table is generated (recording the storage path, size, and loading priority of each layer file); when loading in a third-party 3D software, based on the current operation requirements (such as only editing the object model), only the corresponding layer file is loaded through the index table, reducing the amount of data loaded.

[0053] Prior to step 010, the method may also include steps to create an import plugin and / or a format conversion plugin. Plugins may be created using C++ development (suitable for deep integration) or Python extensions (suitable for lightweight options).

[0054] As an example, the process of creating a format conversion plugin using a Python extension can be as follows:

[0055] Before acquiring the scene file obtained through the first 3D software and the special effects file obtained through the second 3D software, the method further includes:

[0056] Configure the environment based on the format conversion plugin to be created. The environment configuration includes installing the txmake tool for TX format conversion and obtaining the path of the txmake tool.

[0057] Based on the aforementioned environment configuration and format conversion functional requirements, an executable Python script is created. These functional requirements include, but are not limited to, support for multiple file selection, error messages, and Clarisse console log output.

[0058] Associate the Python script with Clarisse's Shelf toolbar to add function buttons for format conversion to the Shelf toolbar, thus creating the import plugin.

[0059] This document describes how to create a format conversion plugin / script using Python. Its functional requirements include converting non-specified formats such as PNG / JPEG / EXR to TX format, supporting batch processing and triggering within Clarisse. Specifically, the creation process for this format conversion plugin can be as follows:

[0060] First, configure the environment. For example, prepare the txmake tool. txmake can be installed with Autodesk software or downloaded separately, and it must be compatible with the operating system, such as Windows. Next, record the (installation) path of the txmake tool to obtain its absolute path. Then, obtain the Clarisse scripts directory path and the test file path. The test file is the folder containing the images that need to be converted.

[0061] Next, a Python script is created. During the script creation process, a batch file selection function and a txmake call function are created. Error handling and logging are added. In addition, the os.path module is used to handle path separators (such as os.path.splitext to separate filenames and suffixes) to ensure compatibility with Windows (\) and Unix ( / ) systems.

[0062] Next, the Python script is associated with Clarisse's Shelf (toolbar), creating a Shelf button as a function button for format conversion. This function button is linked to the Python script, thus forming a format conversion plugin in Clarisse. When a user clicks this function button, the format conversion logic is triggered, achieving a "one-click trigger" function, eliminating the need for the user to manually run the script.

[0063] As another example, if using the C++ language, the process of creating a format conversion plugin can be as follows:

[0064] The first 3D software is Clarisse. Before acquiring the scene file obtained through the first 3D software and the special effects file obtained based on the second 3D software, the method further includes:

[0065] Obtain and configure the path to the SDK (Software Development Kit) that matches the Clarisse;

[0066] Establish the dependency library path, which includes the OpenImageIO path for reading input images and the txmake tool path for TX format conversion;

[0067] CMake generates the corresponding build configuration file for Clarisse. CMake is a cross-platform installation (build) tool.

[0068] Based on the SDK corresponding to Clarisse, the source code file is generated according to the functional requirements of the format conversion plugin to be created;

[0069] The format conversion plugin is compiled and generated based on the SDK path, the dependency library path, the compilation configuration file, and the source code file.

[0070] Create a subdirectory corresponding to the format conversion plugin in the Clarisse plugin directory, and copy the format conversion plugin to the subdirectory to obtain Clarisse with the format conversion plugin.

[0071] As an example, the process of creating a format conversion plugin can be as follows:

[0072] First, prepare the environment:

[0073] (1.1) Prepare the Clarisse SDK installation package, OpenImageIO (OIIO) source code / pre-compiled package, txmake tool, and compilation toolchain.

[0074] (1.2) Unzip the Clarisse SDK to the specified directory (e.g., D: / Clarisse-SDK-5.0), compile OpenImageIO (OIIO) into a static library (enable STATIC_LINK=ON), and output it to D: / OIIO-static (including include / and lib / OpenImageIO.lib); extract the txmake tool to D: / txmake / . In this way, you can get the Clarisse SDK directory, the OIIO static library directory, and the txmake tool path.

[0075] (1.3). Create CMakeLists.txt in the project root directory (e.g., D: / tx_converter_project), fill in the dependency paths from step (1.2) (including the OIIO static library directory and the txmake tool path), and then execute the CMake command to generate the compilation configuration file corresponding to the platform.

[0076] (1.4). Based on the functional requirements of the desired format conversion plugin, define the interface specifications for plugin registration, import hooks, and command callbacks through the Clarisse SDK interface documentation, forming the source code file. Functional requirements may include: batch automatic format conversion (import hook), manual format conversion (menu command), error handling, and log output. Automatic format conversion refers to: when importing textures into Clarisse (e.g., via File > Import or loading textures via material nodes), the plugin intercepts the file path, checks if the format is PNG / JPEG / EXR, and if so, automatically converts it to TX and replaces the reference. This requires utilizing Clarisse's resource import hook. Manual format conversion refers to: adding custom commands to the Clarisse menu (e.g., Edit > Convert to TX) or right-click menu, allowing users to manually select files to trigger conversion. This requires registering menu extensions and command callback functions through the Clarisse SDK.

[0077] (1.5). Using the compilation toolchain (VS2019 tool can be used on Windows system), import the compilation configuration file and source code file for compilation. During the compilation process, link the Clarisse SDK library and OIIO static library to form a plugin file. Use the plugin file and the supporting tool (txmake) as a format conversion plugin.

[0078] (1.6). Create the corresponding subdirectory in the Clarisse plugin directory and copy the format conversion plugin to the corresponding subdirectory to enable Clarisse to perform batch conversion of TX format.

[0079] In this embodiment, the creation process of the import plugin is similar to that of the format conversion plugin. As an example, the principle behind creating the import plugin is as follows:

[0080] Based on the "list of available environments (e.g., Clarisse version, operating system, and required asset formats (USD / ABC / OBJ))" and the functional requirement of "retrieving paths from the clipboard and batch importing assets," a Python script is created. This script reads file paths from the clipboard and automatically calls Clarisse's import tool according to the specified format. Next, the script is validated. For example, the script is run within Clarisse to test various scenarios (e.g., single file, multiple mixed-format files, invalid files) to see if it imports correctly or generates errors. After validation, the Python script is placed in a folder accessible to Clarisse. A new button is created in Clarisse's Shelf panel, associated with this script, and its name and icon are set, forming an import plugin. This import plugin can then use Clarisse's built-in Qt library (PySide2) to read clipboard content, extract file paths, and match corresponding import nodes based on file extensions (.usd / .abc / .obj, etc.) to automatically import the corresponding asset files using the appropriate import tool.

[0081] Prior to step 110, the method further includes:

[0082] Step 040: Using the second 3D software, create corresponding special effects based on the 3D main scene to obtain a special effects file. The position and timeline of the special effects are matched with the 3D main scene.

[0083] Step 050: Using the second 3D software, volumetric effects are converted into first-type effect files in OpenVDB format, and mesh-type dynamic object effects are converted into second-type effect files in USD format. The purpose is to reduce the resource consumption of effect production, improve workflow efficiency, and ensure data consistency throughout the entire process from creation to delivery.

[0084] In this embodiment, volumetric effects may include, but are not limited to, smoke, flames, explosions, and clouds; mesh-type dynamic object effects may include, but are not limited to, broken fragments, fluid surfaces, soft body deformation, and particle geometry.

[0085] In this embodiment, most regions in the voxel data of volumetric effects (such as smoke and fog) are "null" (lacking density / temperature information). OpenVDB uses a sparse raster structure (storing only non-null voxels), and with an efficient compression algorithm, it can compress the data volume to 1 / 10 or even 1 / 100 of the original dense storage. Therefore, using the OpenVDB format for volumetric effects can reduce memory usage. Furthermore, OpenVDB smoke files exported from Houdini can be directly imported into Nuke for density adjustment, or used in Arnold for light sampling to calculate volumetric shadows, avoiding detail loss due to format incompatibility.

[0086] The core of mesh-based dynamic objects is the "time-varying geometric topology / vertex position" (such as the trajectory of broken fragments or the ripple deformation of a fluid surface). Its characteristics include "a large number of elements (potentially tens of thousands of fragments), complex hierarchical relationships (such as parent-child constraints and instantiation), and dense temporal sampling (24 / 30 frames per second)." Mesh-based dynamic object effects use the USD format, which allows for structured management of complex hierarchies and relationships. The hierarchical scene descriptions of USD (Layer, Prim, Attribute structure) can structurally store these relationships: for example, broken "main fragments," "splashed fragments," and "dust particles" can be grouped as different Prims, and their generation logic can be recorded through "relationship attributes." This facilitates quick location of related elements during later adjustments in Houdini or reuse of the hierarchical structure in other software.

[0087] Special effects production is a pipeline process: Houdini is responsible for generating dynamic meshes, Maya may be used to add materials, Nuke is used for compositing, and the renderer is responsible for the final output. USD, as a "universal scene description" standard, can transfer information across the entire chain, including geometry, materials, animation, and lighting, without loss, avoiding information loss caused by traditional formats (such as OBJ which only contains static meshes, and FBX which has poor compatibility).

[0088] In step 110, the method of obtaining the files can be flexibly selected according to the actual situation. For example, the files can be obtained locally, or from a server, external hard drive, etc., to obtain pre-prepared scene files and special effects files.

[0089] In step 120, the method of integrating scene files and special effects files may include: adjusting the position and time of the special effects corresponding to the special effects files in the 3D main scene so that the position and time period of the special effects appearing in the 3D main scene meet the design requirements.

[0090] In step 130, the relative position and size of each display area in the virtual 3D environment are the same as the position and size of the corresponding display screen in the physical experience cabin. There is a one-to-one correspondence between display areas and display screens. The number M of display screens in the experience cabin can be flexibly adjusted according to actual conditions. Display screens are installed on at least three inner walls to form semi-enclosed or enclosed display areas, which helps to improve the immersive effect. For example, a curved display screen can be installed at the front and rear of the experience cabin; one or more flat display screens can be installed at the top of the experience cabin.

[0091] As an example, please refer to Figure 2 This is a schematic diagram of the layout of the display area and virtual camera in a virtual 3D environment from a frontal viewpoint. The virtual 3D environment can include four display areas, corresponding to four viewing directions (front, back, top back, and top front). The front and back areas are curved display areas, while the top back and top front areas are flat display areas.

[0092] In step 140, the field of view is calculated based on the display screen size and the user's viewing distance. The method for obtaining the field of view is conventional and will not be elaborated here. There is a one-to-one correspondence between the display area and the virtual camera. The field of view can include the horizontal field of view (FOV) formed by the left and right edges of the display screen. h The vertical field of view (FOV) formed by the top and bottom edges v "Physical Field of View" (FOV) h FOV v This can be converted into the field of view parameters of a virtual camera. For example, the physical field of view can be scaled proportionally to serve as the virtual camera's field of view parameters. The scaling ratio can be flexibly adjusted according to the actual situation to ensure that the horizontal range of the captured image covers the physical width of the display screen (no black borders, no overflow); and that the vertical range of the captured image covers the physical height of the display screen. In this way, the virtual camera's field of view can reproduce the user's real perspective when viewing the display screen inside the motion pod, ensuring that when the subsequently captured video footage is played on the display screen, the content range, perspective effect, and proportion relationship conform to physical spatial perception, eliminating the problem of "image and space misalignment" from the bottom up.

[0093] The virtual camera's field of view is based on the position and field of view of the display screen in the experience cabin, replacing the extensive settings that rely on experience in existing technologies. By establishing a precise mapping relationship between the virtual camera's field of view and the physical parameters of the display screen (installation position, field of view), problems such as image edge distortion, insufficient or excessive field of view can be avoided, ensuring that the images played on each display screen meet the user's visual expectations, reducing the "jumping feeling" and improving visual immersion.

[0094] In step 150, the motion trajectory and rotation posture can be flexibly created as needed to simulate the motion trajectory and rotation posture of the experience cabin in the 3D main scene, providing a stable and immersive motion foundation for subsequent synchronous shooting. The content displayed on the experience cabin screen is typically used to simulate the user driving, roaming, or performing other activities in the experience cabin.

[0095] The motion trajectories of multi-view cameras must meet the principle of "unified base and synchronous offset," with all virtual cameras sharing the same core trajectory (ensuring synchronization). For example, the relative positions and orientations of each virtual camera are fixed in the corresponding virtual 3D environment of the experience cabin, and all virtual cameras share the same timeline (consistent frame rate, such as 30fps), ensuring that the motion states (position, angle, speed) of each camera at the same moment are completely matched (error ≤ 1ms).

[0096] As an example, the motion trajectory can be generated as follows: On the timeline of the main 3D scene (e.g., 0-60 seconds), keyframes are set according to the scene's rhythm (e.g., 0 seconds for "takeoff," 10 seconds for "turn," 30 seconds for "dive"). Each keyframe contains the camera's 3D position coordinates (X...). k ,Y k Z k ) and timestamp t k Interpolate the trajectory between keyframes using cubic Bézier curves or spline curves to ensure continuous and smooth position changes. Simultaneously, adjust the trajectory's velocity curve according to scene dynamics (such as deceleration when approaching obstacles) (e.g., using the Sigmoid function to achieve a smooth transition between acceleration and deceleration).

[0097] Before shooting with virtual cameras, the method may also include: analyzing the scene complexity within the field of view of each virtual camera in the 3D main scene, including the number of polygons (unit: 10,000) and the number of effect layers (such as the number of smoke and particle layers); classifying virtual cameras into high priority (e.g., complexity ≥ 5 million polygons, or the number of effect layers ≥ 8), medium priority (2-5 million polygons, or 4-8 effect layers), and low priority (< 2 million polygons, and < 4 effect layers) according to scene complexity; allocating 60% of CPU / GPU rendering resources to high-priority virtual cameras, 30% to medium-priority cameras, and 10% to low-priority cameras, and adjusting in real time during the rendering process (e.g., releasing resources to low-priority cameras after high-priority scene rendering is complete), thereby improving overall rendering efficiency.

[0098] In step 160, synchronous shooting is performed using each of the virtual cameras to obtain a video file for synchronous playback on the M displays, which may include:

[0099] For each virtual camera's field of view, continuous images within the field of view are rendered to obtain image frames, wherein the image frames of all virtual cameras are time-synchronized.

[0100] For each virtual camera, the image frames obtained from multiple consecutive shots are synthesized to obtain a video file for each virtual camera. All video files of the virtual cameras correspond one-to-one with the corresponding display screens of the M display screens, so as to serve as video files for synchronous playback on the M display screens.

[0101] Specifically, video files can be generated in the following ways:

[0102] Step 1: Synchronous image rendering based on a unified time reference. That is, for M virtual cameras, the timeline of the 3D main scene is used as the sole reference, and the synchronous generation of image frames is achieved through "timestamp binding and parallel rendering control".

[0103] (1) Pre-configure rendering parameters (matching display characteristics): Configure rendering parameters for each virtual camera that are strictly consistent with the corresponding display screen to ensure that the physical properties of the image frame are adapted to the display screen / playback terminal.

[0104] (2) Rigid timestamp binding (ensuring frame synchronization): Based on the global timeline of the 3D main scene (denoted as T, unit: seconds), a unique timestamp is generated for each time point t (e.g., t=0, 1 / 30, 2 / 30, ..., T total). All virtual cameras trigger rendering synchronously at the same timestamp. Each virtual camera only renders the 3D main scene content within its field of view, ensuring that the image content matches the visual range of the corresponding display screen.

[0105] (3) Image frame verification and correction (to ensure frame integrity): After each frame is rendered, the timestamps of all virtual camera image frames are checked for consistency to ensure time synchronization. Image quality consistency is also checked by comparing the average brightness and color deviation of frames from different cameras (e.g., calculated by PSNR peak signal-to-noise ratio). If the difference is greater than the threshold (e.g., PSNR < 30dB), the rendering parameters of the corresponding camera (e.g., light source intensity, exposure value) are adjusted and the image is re-rendered.

[0106] Step 2: Video file synthesis based on frame sequence, that is, combining the continuous image frames of each virtual camera into a video file in chronological order, while embedding synchronization markers to support coordinated playback on the display screen.

[0107] (1) Frame sequence sorting and alignment: For a single virtual camera, all image frames within the time axis T are collected and sorted in ascending order of timestamp t to form an ordered frame sequence, ensuring that the frame order is completely consistent with the time flow of the 3D main scene. If a frame at a certain timestamp is missing due to rendering failure, a complete frame is generated by interpolating the preceding and following frames (such as pixel offset prediction based on motion trajectory) to avoid video stuttering.

[0108] (2) Video encoding and format adaptation: The ordered frame sequence is synthesized into a video file. The encoding parameters must match the hardware decoding capability of the display screen. The encoding parameters include the encoding format (H.265 / HEVC), bit rate (e.g., 4K+30fps), and container format (e.g., MP4).

[0109] (3) Embedding synchronization markers: embed time synchronization markers for each video file to ensure frame-level synchronization when playing on M displays. For example, embed electrical signal markers in the starting frame (t=0) and key frames (such as scene transition times) of the video file. When the display detects the signal, it forces the alignment of the playback progress of all videos (error ≤ 1ms).

[0110] Step 3: Multi-video collaborative verification. After the video is synthesized, the M video files are verified as a whole to check their synchronization and content coherence. For example, the continuity of image stitching is checked.

[0111] In this embodiment, multi-camera parallel shooting simulation can be achieved in virtual space without the need for overall rendering of the panoramic environment. This not only reduces the number of pixels and sampling overhead required for a single rendering, but also allows for independent adjustment and optimization of the content on each screen.

[0112] It should be noted that after rendering, the footage output from each camera position can be imported into Nuke for spatial geometry correction and stitching. During this stage, various tools in Nuke are used to spatially register the four feeds, correcting edge deviations caused by camera parallax or projection distortion, thereby ensuring visual continuity, consistent brightness, and no obvious seams at the joints between screens.

[0113] In this embodiment, the technical approach of "split-screen multi-camera rendering and post-processing geometric correction" can significantly shorten the overall rendering time and reduce rendering frame overhead. At the same time, it can make each shot more focused on the effective field of view, which is conducive to improving the sense of immersion and realistic surround effect of the immersive scene.

[0114] This application also provides an electronic device that may include a processor and a memory. The memory stores a computer program, which, when executed by the processor, enables the electronic device to perform corresponding steps in the video generation method described below.

[0115] In this embodiment, the processor can be an integrated circuit chip with signal processing capabilities. For example, the processor can be a central processing unit (CPU), a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components, capable of implementing or executing the methods, steps, and logic block diagrams disclosed in the embodiments of this application.

[0116] The memory can be, but is not limited to, random access memory, read-only memory, programmable read-only memory, erasable programmable read-only memory, electrically erasable programmable read-only memory, etc. In this embodiment, the memory can be used to store scene files, special effects files, video files, etc. Of course, the memory can also be used to store programs, which the processor executes after receiving execution instructions.

[0117] It should be noted that those skilled in the art will understand that, for the sake of convenience and brevity, the specific working process of the electronic device described above can be referred to the corresponding steps in the aforementioned method, and will not be elaborated further here.

[0118] This application also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program that, when run on a computer, causes the computer to perform the video generation method as described in the above embodiments.

[0119] Based on the above description of the embodiments, those skilled in the art can clearly understand that this application can be implemented by hardware or by using software plus necessary general-purpose hardware platforms. Based on this understanding, the technical solution of this application can be embodied in the form of a software product. This software product can be stored in a non-volatile storage medium (such as a CD-ROM, USB flash drive, mobile hard drive, etc.) and includes several instructions to cause a computer device (such as a personal computer, electronic device, or network device, etc.) to execute the methods described in the various implementation scenarios of this application.

[0120] In the embodiments provided in this application, it should be understood that the disclosed devices and methods can also be implemented in other ways. The device and method embodiments described above are merely illustrative. For example, the flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of devices, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code, which includes one or more executable instructions for implementing a specified logical function. It should also be noted that each block in a block diagram and / or flowchart, and combinations of blocks in block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions. Furthermore, the functional modules in the various embodiments of this application can be integrated together to form an independent part, or each module can exist independently, or two or more modules can be integrated to form an independent part.

[0121] The above description is merely an embodiment of this application and is not intended to limit the scope of protection of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of protection of this application.

Claims

1. A video generation method, characterized in that, The method includes: Acquire scene files obtained using the first 3D software and special effects files obtained using the second 3D software; By integrating the scene files and special effects files using third-party 3D software, a 3D main scene with special effects is obtained. In the main three-dimensional scene, a virtual three-dimensional environment corresponding to the experience cabin is created. The virtual three-dimensional environment includes a display area corresponding to M displays in the experience cabin. The M displays are distributed on at least three inner walls of the experience cabin, where M is an integer greater than or equal to 3. Within the virtual 3D environment, a virtual camera corresponding to each display area is set up, and the virtual 3D environment with the virtual camera is used as a multi-view camera. The field of view of the virtual camera is obtained based on the position and field of view of the display screen in the experience cabin. In the main 3D scene, create the motion trajectory and rotation posture of the multi-view camera; In the three-dimensional main scene, the multi-view camera is controlled to move based on the motion trajectory and rotation posture, and during the movement, each virtual camera is used to shoot synchronously to obtain a video file for synchronous playback on the M displays.

2. The method according to claim 1, characterized in that, Before acquiring the scene file obtained through the first 3D software and the special effects file obtained based on the second 3D software, the method further includes: In the first 3D software, based on a pre-created import plugin, the path of the target file is obtained through clipboard operation. The target file stores a variety of pre-created first asset files for scene construction. The first asset file includes object models, textures, hierarchical structures and material binding relationships. Based on the path, all first asset files that meet the preset conditions are imported into the first 3D software to form an initial 3D main scene. In the initial 3D main scene, the images representing materials or textures in all the first asset files are converted into image files of a specified format, including TX format, through the format conversion plugin in the first 3D software. The initial 3D main scene is exported in USD format to obtain the scene file.

3. The method according to claim 2, characterized in that, Based on the path, all first asset files that meet preset conditions are imported into the first 3D software to form an initial 3D main scene, including: Based on the path, a first initial asset file representing the target asset is determined in the target file, and a second initial asset file representing the material or texture is determined in the target file. The first initial asset file and the second initial asset file are used as the first asset file that meets the preset conditions. The target asset includes any one of the asset files in USD format, ABC format, and OBJ format.

4. The method according to claim 2, characterized in that, The first 3D software is Clarisse. Before acquiring the scene file obtained through the first 3D software and the special effects file obtained based on the second 3D software, the method further includes: Configure the environment based on the format conversion plugin to be created. The environment configuration includes installing the txmake tool for TX format conversion and obtaining the path of the txmake tool. Based on the aforementioned environment configuration and format conversion functional requirements, an executable Python script is created, including support for multiple file selection. Associate the Python script with Clarisse's Shelf toolbar to add function buttons for format conversion to the Shelf toolbar, thus creating the import plugin.

5. The method according to claim 2, characterized in that, The first 3D software is Clarisse. Before acquiring the scene file obtained through the first 3D software and the special effects file obtained based on the second 3D software, the method further includes: Obtain and configure the path to the SDK that matches the Clarisse; Establish the dependency library path, which includes the OpenImageIO path for reading input images and the txmake tool path for TX format conversion; Generate the Clarisse build configuration file using CMake; Based on the SDK corresponding to Clarisse, the source code file is generated according to the functional requirements of the format conversion plugin to be created; The format conversion plugin is compiled and generated based on the SDK path, the dependency library path, the compilation configuration file, and the source code file. Create a subdirectory corresponding to the format conversion plugin in the Clarisse plugin directory, and copy the format conversion plugin to the subdirectory to obtain Clarisse with the format conversion plugin.

6. The method according to claim 1, characterized in that, Before acquiring the scene file obtained through the first 3D software and the special effects file obtained based on the second 3D software, the method further includes: Using the second 3D software, corresponding special effects are created based on the 3D main scene to obtain special effects files. The position and timeline of the special effects are matched with the 3D main scene. Using the second 3D software, volumetric effects are converted into first-type effect files in OpenVDB format, and mesh-type dynamic object effects are converted into second-type effect files in USD format.

7. The method according to claim 6, characterized in that, The volumetric effects include at least one of smoke, flame, explosion, and cloud; the mesh-type dynamic object effects include at least one of broken fragments, fluid surfaces, soft deformation, and particle geometry.

8. The method according to claim 1, characterized in that, By synchronously capturing images using each of the virtual cameras, video files are obtained for synchronous playback on the M displays, including: For each virtual camera's field of view, continuous images within the field of view are rendered to obtain image frames, wherein the image frames of all virtual cameras are time-synchronized. For each virtual camera, the image frames obtained from multiple consecutive shots are synthesized to obtain a video file for each virtual camera. All video files of the virtual cameras correspond one-to-one with the corresponding displays in the M displays, so as to serve as video files for synchronous playback on the M displays.

9. An electronic device, characterized in that, The electronic device includes a processor and a memory coupled together, the memory storing a computer program that, when executed by the processor, causes the electronic device to perform the method as described in any one of claims 1 to 8.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when run on a computer, causes the computer to perform the method as described in any one of claims 1 to 8.