Video generation method and device in three-dimensional scene, equipment and storage medium

By introducing motion mask calculation, clustering and change prediction steps in the three-dimensional Gaussian splattering technology, the flickering artifact problem of three-dimensional Gaussian splattering during stream free view synthesis is solved, and the video quality is significantly improved.

CN120201258APending Publication Date: 2025-06-24PENG CHENG LAB
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510382331.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-28
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

When performing stream free view synthesis, the synthesized free view video often has obvious flickering artifacts and the video quality is not high.

Method used

By acquiring scene images from multiple perspectives at each collected frame, the initial three-dimensional Gaussian primitives are generated, and the three-dimensional Gaussian primitives are updated through iteratively, the motion mask is calculated, the motion primitives are clustered, the change prediction is performed, and the change residual data is generated to reduce flicker artifacts and improve the video quality.

Benefits of technology

It effectively reduces flicker artifacts in free-view synthetic videos, improves the quality of video generation, and improves the accuracy and stability of synthetic videos.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120201258A_ABST
    Figure CN120201258A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a video generation method and device in a three-dimensional scene, equipment and a storage medium, and relates to the technical field of computer vision. According to the method, a motion mask is obtained through calculation of two adjacent frames of images, and a dynamic area in a scene is accurately identified. Then determining a surface motion element and a motion element based on the motion mask, further distinguishing a dynamic region from a static region, and avoiding the interference of the dynamic region on the static region; and then, updating the three-dimensional Gaussian primitive through iteration to ensure that the change characteristics of the dynamic region are kept consistent between frames, generating change residual data between the frames, accurately describing the change of the dynamic region between adjacent frames, and avoiding flicker artifacts caused by inconsistent change of the dynamic region. The video generation quality can be effectively improved, the application range is wide, and the accuracy and stability of the synthesized video are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer vision technology, and in particular to a method, apparatus, device, and storage medium for video generation in a three-dimensional scene. Background Art

[0002] Dynamic Novel View Synthesis (DNVS) uses a multi-view video sequence to achieve free viewing of a dynamic scene from any perspective. With the rapid development of virtual reality (VR) and augmented reality (AR) technologies, the application demand for novel view synthesis of dynamic scenes in real scenarios is increasing day by day. Its core goal is to generate continuous frames from any perspective in a dynamic scene through multi-view video data, providing users with an immersive viewing experience. Due to its high efficiency and real-time performance, it is widely used in multiple downstream scenarios. For example: online live broadcast, sports production, VR navigation, stage reconstruction, or sports event broadcasting, etc.

[0003] In related technologies, a three-dimensional Gaussian splatting method is used to achieve dynamic novel view synthesis. Based on explicit Gaussian ellipsoidal primitives, while maintaining high synthesis quality, it can also achieve real-time rendering speed. However, when performing free view synthesis using three-dimensional Gaussian splatting, the synthesized free view video often has obvious flickering artifacts and the video quality is not high. Summary of the Invention

[0004] The main purpose of the embodiments of this application is to propose a method, apparatus, device, and storage medium for video generation in a three-dimensional scene, reducing the flickering artifacts in the video obtained by free view synthesis and improving the video generation quality.

[0005] To achieve the above object, the first aspect of the embodiments of this application proposes a method for video generation in a three-dimensional scene, including:

[0006] Obtaining the scene images corresponding to each acquisition frame from multiple perspectives, and generating a plurality of initial three-dimensional Gaussian primitives according to the scene images of the first frame of each of the perspectives;

[0007] Obtain the three-dimensional Gaussian basis elements of the previous frame, as well as the current frame image and the previous frame image under each of the perspectives. Calculate a motion mask based on the current frame image and the previous frame image. Obtain surface motion basis elements according to the three-dimensional Gaussian basis elements and all the motion masks. Cluster the surface motion basis elements to obtain at least one cluster. Obtain motion basis elements according to the cluster and the surface motion basis elements. Perform change prediction based on the motion basis elements to obtain change residual data corresponding to the current frame. Obtain the current frame Gaussian basis elements according to the three-dimensional Gaussian basis elements and the change residual data. Use the current frame Gaussian basis elements as the three-dimensional Gaussian basis elements of the previous frame and perform iteration. The initial value of the three-dimensional Gaussian basis elements is the initial three-dimensional Gaussian basis elements, and the initial value of the previous frame image is the scene image of the first frame;

[0008] Obtain rendering Gaussian basis elements according to the initial three-dimensional Gaussian basis elements of the first frame and the current frame Gaussian basis elements of other frames. The rendering Gaussian basis elements are used for rendering to generate a target video.

[0009] In some embodiments, the calculating a motion mask based on the current frame image and the previous frame image includes:

[0010] Perform optical flow motion estimation on the current frame image and the previous frame image to obtain the optical flow difference at each pixel position. Compare the optical flow difference with a preset optical flow value to generate a significant motion mask;

[0011] Perform frame difference motion estimation on the current frame image and the previous frame image to obtain the frame difference at each pixel position. Compare the frame difference with a preset frame difference value to generate a tiny motion mask;

[0012] Fuse the significant motion mask and the tiny motion mask to obtain the motion mask.

[0013] In some embodiments, the obtaining surface motion basis elements according to the three-dimensional Gaussian basis elements and all the motion masks includes:

[0014] For each pixel position, calculate the contribution value of each three-dimensional Gaussian basis element to the pixel position, and select the three-dimensional Gaussian basis element corresponding to the maximum value of the contribution value as the target Gaussian basis element of the pixel position;

[0015] Under each perspective, form an initial basis element sequence according to the target Gaussian basis elements, multiply the initial basis element sequence by the motion mask to obtain a target basis element sequence, and remove duplicates of the three-dimensional Gaussian basis elements in the target basis element sequences of all perspectives to obtain the surface motion basis elements.

[0016] In some embodiments, the calculating the contribution value of each three-dimensional Gaussian basis element to the pixel position includes:

[0017] Obtain the opacity of each of the three-dimensional Gaussian basis elements;

[0018] Take each of the three-dimensional Gaussian basis elements as a to-be-tested Gaussian basis element one by one, and obtain the three-dimensional Gaussian basis elements located in front of the to-be-tested Gaussian basis element as occluding Gaussian basis elements in sequence;

[0019] Cumulatively multiply the difference between 1 and the opacity of the occluding Gaussian basis element to obtain an occlusion quantization value;

[0020] Calculate the product between the opacity of the to-be-tested Gaussian basis element and the occlusion quantization value as the contribution value of the to-be-tested Gaussian basis element.

[0021] In some embodiments, the obtaining the motion basis element according to the clustering cluster and the surface motion basis element includes:

[0022] Perform a convex hull operation on each of the clustering clusters to obtain a convex boundary corresponding to the clustering cluster;

[0023] Take all the three-dimensional Gaussian basis elements located within the convex boundary as internal Gaussian basis elements;

[0024] Obtain the motion basis element according to the surface motion basis element and all the internal Gaussian basis elements.

[0025] In some embodiments, the obtaining the change residual data corresponding to the current frame by performing change prediction according to the motion basis element includes:

[0026] Obtain the first position data of each of the motion basis elements, input the first position data into a pre-trained rigid change network for displacement prediction to obtain the position offset and rotation offset corresponding to each of the motion basis elements in the current frame;

[0027] Determine a changed motion basis element based on the motion basis element, the position offset, the rotation offset, the current frame image, and the motion basis element;

[0028] Obtain the second position data of each of the changed motion basis elements, input the second position data into a pre-trained optimization network for color prediction to obtain the color offset corresponding to each of the changed motion basis elements in the current frame;

[0029] Obtain the change residual data according to the position offset, the rotation offset, and the color offset.

[0030] In some embodiments, the determining a changed motion basis element based on the motion basis element, the position offset, the rotation offset, the current frame image, and the motion basis element includes:

[0031] Displace on the motion primitive based on the corresponding position offset and rotation offset to obtain the changed motion primitive of the current frame;

[0032] Perform image rendering based on all the changed motion primitives to obtain the rendered image of the current frame;

[0033] For each pixel position, calculate the difference value according to the rendered image and the current frame image, and obtain the attention value according to the difference value;

[0034] Obtain the attention map according to the attention values of all the pixel positions, and determine the changed motion primitive based on the attention map and the motion primitive.

[0035] In some embodiments, the determining the changed motion primitive based on the attention map and the motion primitive includes:

[0036] For each pixel position, select the changed motion primitive corresponding to the maximum value of the contribution value as the attention motion primitive of the pixel position according to the contribution value of each changed motion primitive to the pixel position;

[0037] Construct an attention primitive sequence according to all the attention motion primitives, multiply the attention primitive sequence by the attention map to obtain a changed primitive sequence, and take the intersection of the changed primitive sequence and the motion primitive to obtain the changed motion primitive.

[0038] In some embodiments, the obtaining the current frame Gaussian primitive according to the three-dimensional Gaussian primitive and the changed residual data includes:

[0039] Remove the motion primitive from the three-dimensional Gaussian primitive to obtain a stationary Gaussian primitive;

[0040] Perform parameter variation on the motion primitive based on at least one of the position offset, the rotation offset or the color offset to obtain the changed primitive of the current frame;

[0041] Obtain the current frame Gaussian primitive according to the changed primitive of the current frame and the stationary Gaussian primitive.

[0042] To achieve the above object, a second aspect of the embodiments of the present application proposes a video generation device in a three-dimensional scene, including:

[0043] Initialization module: configured to obtain the scene images corresponding to each acquisition frame of multiple perspectives, and generate a plurality of initial three-dimensional Gaussian primitives according to the scene images of the first frame of each perspective;

[0044] Iterative module: It is used to obtain the three-dimensional Gaussian basis elements of the previous frame, as well as the current frame image and the previous frame image under each perspective. A motion mask is calculated based on the current frame image and the previous frame image. Surface motion basis elements are obtained according to the three-dimensional Gaussian basis elements and all the motion masks. At least one clustering cluster is obtained by clustering the surface motion basis elements. Motion basis elements are obtained according to the clustering cluster and the surface motion basis elements. Change prediction is performed based on the motion basis elements to obtain the change residual data corresponding to the current frame. The current frame Gaussian basis elements are obtained according to the three-dimensional Gaussian basis elements and the change residual data. The current frame Gaussian basis elements are used as the three-dimensional Gaussian basis elements of the previous frame for iteration. The initial value of the three-dimensional Gaussian basis elements is the initial three-dimensional Gaussian basis elements, and the initial value of the previous frame image is the scene image of the first frame;

[0045] Rendering module: It is used to obtain rendering Gaussian basis elements according to the initial three-dimensional Gaussian basis elements of the first frame and the current frame Gaussian basis elements of other frames. The rendering Gaussian basis elements are used for rendering to generate a target video.

[0046] To achieve the above object, a third aspect of the embodiments of the present application proposes an electronic device, which includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the method described in the first aspect above is implemented.

[0047] To achieve the above object, a fourth aspect of the embodiments of the present application proposes a storage medium, which is a storage medium that stores a computer program. When the computer program is executed by a processor, the method described in the first aspect above is implemented.

[0048] The video generation method, device, equipment, and storage medium proposed in the embodiments of the present application obtain the scene images corresponding to each acquisition frame from multiple perspectives, and generate multiple initial three-dimensional Gaussian basis elements based on the scene images of the first frame of each perspective. Next, the three-dimensional Gaussian basis elements of the previous frame, the current frame image, and the previous frame image of each perspective are obtained. A motion mask is calculated based on the current frame image and the previous frame image. Surface motion basis elements are obtained based on the three-dimensional Gaussian basis elements and all motion masks. The surface motion basis elements are clustered to obtain at least one cluster. Motion basis elements are obtained based on the cluster and the surface motion basis elements. Change prediction is performed based on the motion basis elements to obtain the change residual data corresponding to the current frame. The current frame Gaussian basis element is obtained based on the three-dimensional Gaussian basis element and the change residual data. The current frame Gaussian basis element is used as the three-dimensional Gaussian basis element of the next frame for iteration, where the initial value of the three-dimensional Gaussian basis element is the initial three-dimensional Gaussian basis element, and the initial value of the previous frame image is the scene image of the first frame. Finally, the rendering Gaussian basis element is obtained based on the initial three-dimensional Gaussian basis element of the first frame and the current frame Gaussian basis elements of other frames. In the embodiments of the present application, first, a motion mask is calculated through the images of two adjacent frames to accurately identify the dynamic regions in the scene. Then, based on the motion mask, the surface motion basis elements and the motion basis elements are determined to further distinguish the dynamic and static regions and avoid the interference of the dynamic regions on the static regions. Then, the three-dimensional Gaussian basis element is iteratively updated to ensure that the change characteristics of the dynamic regions are consistent between frames, and at the same time, the change residual data between frames is generated to accurately describe the changes of the dynamic regions between adjacent frames and avoid the flicker artifacts caused by inconsistent changes of the dynamic regions. It can effectively improve the quality of video generation, has a wide range of applications, and significantly improves the accuracy and stability of the synthesized video. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] Figure 1 is a flowchart of the video generation method in the three-dimensional scene provided by the embodiments of the present application.

[0050] Figure 2 is a flowchart of calculating the motion mask based on the current frame image and the previous frame image provided by the embodiments of the present application.

[0051] Figure 3 is a flowchart of obtaining the surface motion basis elements based on the three-dimensional Gaussian basis elements and all motion masks provided by the embodiments of the present application.

[0052] Figure 4 is a flowchart of calculating the contribution value of each three-dimensional Gaussian basis element to the pixel position provided by the embodiments of the present application.

[0053] Figure 5 is a flowchart of obtaining the motion basis elements based on the cluster and the surface motion basis elements provided by the embodiments of the present application.

[0054] Figure 6It is a flowchart for obtaining the change residual data corresponding to the current frame by predicting changes based on motion primitives provided in an embodiment of the present application.

[0055] Figure 7 It is a flowchart for determining the change motion primitives based on motion primitives, position offsets, rotation offsets, the current frame image, and motion primitives provided in an embodiment of the present application.

[0056] Figure 8 It is a flowchart for determining the change motion primitives based on the attention map and motion primitives provided in an embodiment of the present application.

[0057] Figure 9 It is a flowchart for obtaining the current frame Gaussian basis by using three-dimensional Gaussian basis elements and change residual data provided in an embodiment of the present application.

[0058] Figure 10 It is the overall flowchart of the video generation method in a three-dimensional scene provided in an embodiment of the present application.

[0059] Figure 11 It is the flowchart of the motion positioning part and the change and optimization part of the video generation method in a three-dimensional scene provided in an embodiment of the present application.

[0060] Figure 12 It is the structural block diagram of the video generation device in a three-dimensional scene provided in another embodiment of the present application.

[0061] Figure 13 It is the schematic diagram of the hardware structure of the electronic device provided in an embodiment of the present application. Detailed implementation manners

[0062] In order to make the objectives, technical solutions, and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0063] It should be noted that although the functional modules are divided in the device schematic diagram and the logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order from the module division in the device or the order in the flowchart.

[0064] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which the present application belongs. The terms used herein are only for the purpose of describing the embodiments of the present application and are not intended to limit the present application.

[0065] First, several nouns involved in the present application are analyzed:

[0066] Artificial Intelligence (AI): It is a new technical science that studies and develops theories, methods, technologies, and application systems for simulating, extending, and expanding human intelligence. Artificial intelligence is a branch of computer science. It attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. The research in this field includes robots, speech recognition, image recognition, natural language processing, and expert systems, etc. Artificial intelligence can simulate the information process of human consciousness and thinking. It also uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, sense the environment, acquire knowledge, and use knowledge to obtain the best results in terms of theories, methods, technologies, and application systems.

[0067] Dynamic Novel View Synthesis (DNVS) uses multi-view video sequences to enable free viewing of dynamic scenes from any perspective. With the rapid development of virtual reality (VR) and augmented reality (AR) technologies, the application demand for novel view synthesis of dynamic scenes in real scenes is increasing day by day. Its core goal is to generate continuous frames from any perspective in a dynamic scene through multi-view video data, providing users with an immersive viewing experience. Due to its high efficiency and real-time performance, it is widely used in multiple downstream scenarios. For example: online live broadcast, sports production, VR navigation, stage reconstruction, or sports event broadcast, etc.

[0068] In related technologies, the three-dimensional Gaussian splatting method is used to achieve dynamic novel view synthesis. Based on explicit Gaussian ellipsoidal primitives, while maintaining high synthesis quality, it can also achieve real-time rendering speed. However, when performing free-viewpoint synthesis of flow, the synthesized free-viewpoint video often has obvious flicker artifacts and the video quality is not high.

[0069] Based on this, the embodiments of this application provide a video generation method, device, equipment, and storage medium in a three-dimensional scene. First, a motion mask is calculated from two adjacent frames of images to accurately identify the dynamic regions in the scene. Then, based on the motion mask, surface motion primitives and motion primitives are determined to further distinguish between dynamic and static regions and avoid interference of dynamic regions on static regions. Then, by iteratively updating the three-dimensional Gaussian primitives, it is ensured that the change characteristics of the dynamic regions are consistent between frames, and at the same time, the change residual data between frames is generated to accurately describe the changes of the dynamic regions between adjacent frames, avoiding flicker artifacts caused by inconsistent changes of the dynamic regions. It can effectively improve the generation quality of the video, has a wide range of applications, and significantly improves the accuracy and stability of the synthesized video.

[0070] The embodiments of the present application provide a method, apparatus, device, and storage medium for video generation in a three-dimensional scene, which will be specifically described through the following embodiments. First, the method for video generation in a three-dimensional scene in the embodiments of the present application will be described.

[0071] The embodiments of the present application can acquire and process relevant data based on artificial intelligence technology. Among them, Artificial Intelligence (AI) is a theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a similar way to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines, enabling the machines to have the functions of perception, reasoning, and decision-making.

[0072] Artificial intelligence technology is a comprehensive discipline that involves a wide range of fields, including both hardware-level technologies and software-level technologies. The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0073] The method for video generation in a three-dimensional scene provided by the embodiments of the present application relates to the field of computer vision technology. The method for video generation in a three-dimensional scene provided by the embodiments of the present application can be applied to a terminal, or to a server, or can be a computer program running on a terminal or a server. For example, the computer program can be a native program or software module in an operating system; it can be a local (Native) application (Application, APP), that is, a program that needs to be installed in an operating system to run, such as a client that supports video generation in a three-dimensional scene, that is, a program that only needs to be downloaded to a browser environment to run; it can also be a small program that can be embedded in any APP. In short, the above computer program can be any form of application program, module, or plug-in. Among them, the terminal communicates with the server through a network. The method for video generation in a three-dimensional scene can be executed by the terminal or the server, or by the terminal and the server in cooperation.

[0074] In some embodiments, the terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart watch, etc. In addition, the terminal can also be an intelligent vehicle-mounted device. The intelligent vehicle-mounted device applies the video generation method in a three-dimensional scene of this embodiment to provide relevant services and enhance the driving experience. The server can be an independent server or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms; it can also be a service node in a blockchain system, and the service nodes in the blockchain system form a Peer To Peer (P2P) network, and the P2P protocol is an application layer protocol running on top of the Transmission Control Protocol (TCP) protocol. The terminal and the server can be connected through communication connection methods such as Bluetooth, Universal Serial Bus (USB), or network, and this embodiment does not limit this here.

[0075] This application can be used in many general or special computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multi-processor systems, microprocessor-based systems, set-top boxes, programmable consumer electronic devices, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and so on. This application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. This application can also be practiced in a distributed computing environment where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.

[0076] The video generation method in a three-dimensional scene in the embodiments of this application will be described below.

[0077] Figure 1 It is an optional flowchart of the video generation method in a three-dimensional scene provided by the embodiments of this application. Figure 1 The method in can include but is not limited to steps 110 to 130. At the same time, it can be understood that this embodiment does not specifically limit the order of steps 110 to 130 in, and the order of steps can be adjusted according to actual needs, or some steps can be reduced or added. Figure 1 The order of steps 110 to 130 in is not specifically limited, and the order of steps can be adjusted according to actual needs, or some steps can be reduced or added.

[0078] Step 110: Obtain the scene images corresponding to multiple perspectives in each acquisition frame, and generate multiple initial 3D Gaussian primitives based on the scene images of the first frame of each perspective.

[0079] In one embodiment, for a target scene, image acquisition devices are arranged at multiple perspectives. At a certain perspective, the target scene is imaged at different times using the corresponding image acquisition device, and the time is defined as the acquisition frame. Among them, there are scene images corresponding to multiple acquisition frames for each perspective, and these scene images contain the geometric, texture, and illumination information of the target scene from different perspectives.

[0080] Next, according to the scene images from different perspectives, use the COLMAP algorithm to calibrate all image acquisition devices, estimate the camera internal parameters (such as focal length, principal point coordinates, and distortion coefficients), camera external parameters (such as rotation matrix and translation vector), and sparse point cloud for each perspective. Subsequently, select the first acquisition frame as the first frame, and obtain the scene images of the first frame for each perspective. Input the scene images of the first frame, camera internal parameters, camera external parameters, and sparse point cloud into the 3D Gaussian splashing algorithm to generate multiple initial 3D Gaussian primitives corresponding to the first frame of the target scene.

[0081] Among them, each initial 3D Gaussian primitive can be regarded as a 3D ellipsoid, or a 3D Gaussian distribution, which is used to describe the geometric shape, color, density, and other information of a local area in the target scene. Through the combination of multiple initial 3D Gaussian primitives, a 3D model of the target scene can be constructed.

[0082] Step 120: Obtain the 3D Gaussian primitives of the previous frame, the current frame image and the previous frame image for each perspective, calculate the motion mask based on the current frame image and the previous frame image, obtain the surface motion primitives according to the 3D Gaussian primitives and all motion masks, cluster the surface motion primitives to obtain at least one cluster, obtain the motion primitives according to the cluster and the surface motion primitives, perform change prediction based on the motion primitives to obtain the change residual data corresponding to the current frame, obtain the current frame Gaussian primitives for each perspective according to the 3D Gaussian primitives and the change residual data, and use the current frame Gaussian primitives as the 3D Gaussian primitives of the previous frame for iteration.

[0083] In one embodiment, each acquisition frame is defined to correspond to a preset number of three-dimensional Gaussian basis elements. During the rendering process of each frame, the same number of three-dimensional Gaussian basis elements are used for splashing rendering according to the viewing angle. Each three-dimensional Gaussian basis element is distinguished by an index, and the preset number is set according to the actual situation. For the first frame, the three-dimensional Gaussian basis elements are the initial three-dimensional Gaussian basis elements. Except for the first frame, other acquisition frames do not need to obtain the corresponding three-dimensional Gaussian basis elements in the same way as the first frame, but are obtained through the feature changes between adjacent frames. Since the number of acquisition frames is uncertain, for each acquisition frame, a cyclic iteration method is adopted to determine the three-dimensional Gaussian basis elements of other acquisition frames except the first frame.

[0084] Among them, the three-dimensional Gaussian basis elements include the following attributes: position, occupancy, ellipsoid axis length, rotation angle, and color characterized by spherical harmonic coefficients. The position defines the coordinates of the three-dimensional Gaussian basis element in three-dimensional space, describing the specific position of the three-dimensional Gaussian basis element in the target scene. The occupancy represents the visibility or importance of the three-dimensional Gaussian basis element in the target scene, reflecting whether the area is occupied by an object in the target scene or whether it contributes to the visibility of the target scene. The ellipsoid axis length defines the geometric shape of the three-dimensional Gaussian basis element, describing the dimensions of the three-dimensional Gaussian basis element in the three principal axis directions, thereby determining its ellipsoid shape. The rotation angle defines the rotation posture of the three-dimensional Gaussian basis element in three-dimensional space, describing the rotation state of the three-dimensional Gaussian basis element relative to the global coordinate system, thereby determining the direction of its ellipsoid. The color can be characterized by spherical harmonic coefficients, which is used to describe the color distribution of the three-dimensional Gaussian basis element from different viewing angles. The spherical harmonic coefficients can represent complex color changes, such as lighting and shadow effects, with a relatively low storage cost.

[0085] The following takes a random acquisition frame as an example to describe the iteration process.

[0086] In any iteration, the adjacent frames are respectively called the current frame and the previous frame. First, it is necessary to obtain the three-dimensional Gaussian basis elements of the previous frame, as well as the current frame image and the previous frame image at each viewing angle. Taking the first iteration as an example, the adjacent frames are the first frame and the second frame. The second frame is used as the current frame, and the first frame is used as the previous frame. The current frame image is the scene image of the second frame, and the previous frame image is the scene image of the first frame. The three-dimensional Gaussian basis elements of the previous frame are the initial three-dimensional Gaussian basis elements. Therefore, the initial value of the three-dimensional Gaussian basis elements is the initial three-dimensional Gaussian basis elements, and the initial value of the previous frame image is the scene image of the first frame.

[0087] By analyzing the feature changes between the previous frame and the current frame, the three-dimensional Gaussian basis elements corresponding to the current frame can be updated. This process continues in subsequent iterations, that is, in each iteration, the three-dimensional Gaussian basis elements of the previous frame are used as input, combined with the scene image information of the current frame, to optimize and generate the three-dimensional Gaussian basis elements of the current frame. This iterative way ensures the continuity and consistency of the three-dimensional Gaussian basis elements between frames. Through continuous iterative updates, the three-dimensional Gaussian basis elements can accurately reflect the change characteristics of dynamic regions while maintaining the stability of static regions, thus effectively improving the visual quality and realism of the finally generated video.

[0088] In one embodiment, referring to Figure 2 , Figure 2 is a flowchart of calculating a motion mask based on the current frame image and the previous frame image provided by an embodiment of the present application, which specifically includes the following steps:

[0089] Step 210: Perform optical flow motion estimation based on the current frame image and the previous frame image to obtain the optical flow difference at each pixel position, compare the optical flow difference with a preset optical flow value, and generate a significant motion mask.

[0090] In one embodiment, the current frame image I t and the previous frame image I t-1 can be jointly input into a trained optical flow model, and the optical flow information corresponding to the two scene images of adjacent frames is obtained by using the optical flow model. Among them, the optical flow model can be used to estimate the pixel motion in the image sequence, and the optical flow information is obtained by analyzing the pixel changes between adjacent frames. Here, the optical flow model can adopt relevant models in related technologies.

[0091] Then, the pixel positions are selected one by one, and the optical flow difference at each pixel position is obtained according to the difference between the optical flow information of the current frame and the optical flow information of the previous frame, that is, the motion vector of each pixel between the two frames. Then, the optical flow difference is compared with a preset optical flow value. If the optical flow difference is greater than the preset optical flow value, the significant mask value at the corresponding pixel position is 1, otherwise it is 0. All the significant mask values are statistically obtained to generate a significant motion mask. For example, the preset optical flow value can be set to 1, so as to shield the noise of unnecessary static regions and highlight the significantly moving regions in the target scene.

[0092] It can be seen that the significant motion mask is a binary image generated by comparing the optical flow difference with the preset optical flow value, and can identify the significantly moving regions in the target scene.

[0093] Step 220: Perform frame difference motion estimation based on the current frame image and the previous frame image to obtain the frame difference at each pixel position, compare the frame difference with a preset frame difference value, and generate a micro motion mask.

[0094] In one embodiment, considering that there may still be motion that is not obvious or the motion of semi-transparent objects in the target scene, frame difference motion estimation is introduced to extract the motion features of this part. Among them, frame difference motion estimation detects the motion area by calculating the pixel intensity difference between the current frame image and the previous frame image. The frame difference is the intensity difference between the current frame image and the previous frame image at the same pixel position.

[0095] In addition, after obtaining the frame difference, it is compared with a preset frame difference value, and the preset frame difference value here can be 10. If the frame difference is greater than the preset frame difference value, the micro mask value at the corresponding pixel position is 1, indicating that the pixel position belongs to the micro motion area. If the frame difference is less than or equal to the preset frame difference value, the micro mask value at the corresponding pixel position is 0, indicating that the pixel position belongs to the static area or noise. By counting the micro mask values of all pixels, a micro motion mask is generated. The micro motion mask is a binary image generated by comparing the frame difference with the preset frame difference value, and is used to identify the areas of micro motion in the scene, such as slowly moving objects or subtle changes.

[0096] Step 230: Fuse the significant motion mask and the micro motion mask to obtain a motion mask.

[0097] In one embodiment, the significant motion mask and the micro motion mask are superimposed to obtain a motion mask. Among them, when superimposing, if at the same pixel position, the value of either the significant motion mask or the micro motion mask is 1, the value of the motion mask is 1, indicating that the pixel position belongs to the motion area. If the values of both the significant motion mask and the micro motion mask are 0, the value of the motion mask is 0, indicating that the pixel position belongs to the static area or noise. By superimposing the significant motion mask and the micro motion mask, the motion mask can capture both fast motion and slow motion in the scene, providing a more comprehensive motion detection result.

[0098] Motion mask The calculation process can be expressed as:

[0099]

[0100] where δ represents the step function, which is used to binarize the optical flow difference, represents the optical flow model, τ represents the preset optical flow value, and They respectively represent morphological dilation and erosion operations, abs represents taking the absolute value, and · represents dot product. Among them, morphological erosion is used to remove noise and small regions, and morphological dilation is used to fill holes and connect broken regions. Through morphological dilation and erosion, the continuity and robustness of the moving region can be enhanced. The dot product combines the significant motion mask and the micro-motion mask to generate the final motion mask. Additionally, a motion mask can also be calculated using a motion segmentation model, which is not limited in the embodiments of this application.

[0101] For example, in a target scene, there is a person and a cat moving. In the current frame image, the person is walking fast, which belongs to significant motion, while the cat is moving slowly, which belongs to micro-motion. At this time, a significant motion mask can be generated through an optical flow model to identify the area where the person is walking; a micro-motion mask can be generated through frame difference motion estimation to identify the area where the cat is moving.

[0102] When the significant motion mask and the micro-motion mask are superimposed, if the value of the area where the person is walking is 1 in the significant motion mask and also 1 in the motion mask. The value of the area where the cat is moving is 1 in the micro-motion mask and also 1 in the motion mask. The value of the background area is 0 in both the significant motion mask and the micro-motion mask, and also 0 in the motion mask. In the finally generated motion mask, the area where the person is walking and the area where the cat is moving are both marked as 1, and the background area is marked as 0. It can be seen that the motion mask in the embodiments of this application can detect both fast motion and slow motion in the target scene simultaneously.

[0103] In one embodiment, next, it is necessary to use the motion mask to distinguish between the static region and the moving region. Refer to Figure 3 , Figure 3 is a flowchart of obtaining surface motion primitives according to three-dimensional Gaussian primitives and all motion masks provided by the embodiments of this application, which specifically includes the following steps:

[0104] Step 310: For each pixel position, calculate the contribution value of each three-dimensional Gaussian primitive to the pixel position, and select the three-dimensional Gaussian primitive corresponding to the maximum contribution value as the target Gaussian primitive of the pixel position.

[0105] In one embodiment, since the positions of each three-dimensional Gaussian basis element in the target scene are different, their distances to the image plane are different. During the rendering process, first, the three-dimensional Gaussian basis elements are sorted according to their depth, that is, the distance from the image acquisition device. At this time, the closer three-dimensional Gaussian basis elements may occlude the farther ones. Therefore, for each pixel position, if multiple three-dimensional Gaussian basis elements are projected onto this pixel position, the closer three-dimensional Gaussian basis element will occlude the farther one. And the transparency of each three-dimensional Gaussian basis element is different, and the transparency can be calculated based on the occupancy. The three-dimensional Gaussian basis element with a higher transparency allows part of the light to pass through, so that the three-dimensional Gaussian basis element behind it is still visible. For each pixel position, if multiple three-dimensional Gaussian basis elements are projected onto this pixel position, the three-dimensional Gaussian basis element with a lower transparency will significantly occlude the three-dimensional Gaussian basis element behind it, while the three-dimensional Gaussian basis element with a higher transparency has a weaker occlusion effect. Therefore, for each pixel position, the embodiment of the present application calculates the contribution degree of the three-dimensional Gaussian basis element to this pixel position, and selects the three-dimensional Gaussian basis element that has the greatest impact on it as the target Gaussian basis element.

[0106] In one embodiment, referring to Figure 4 , Figure 4 is a flowchart for calculating the contribution value of each three-dimensional Gaussian basis element to the pixel position provided by the embodiment of the present application, which specifically includes the following steps:

[0107] Step 410: Obtain the opacity of each three-dimensional Gaussian basis element.

[0108] In one embodiment, the opacity is obtained by multiplying the occupancy by the corresponding Gaussian value.

[0109] Step 420: Take each three-dimensional Gaussian basis element as the to-be-tested Gaussian basis element in turn, and obtain the three-dimensional Gaussian basis elements located in front of the to-be-tested Gaussian basis element as the occluding Gaussian basis elements in order.

[0110] In one embodiment, since the three-dimensional Gaussian basis elements have been sorted according to the depth, it is assumed that the to-be-tested Gaussian basis element is the i-th three-dimensional Gaussian basis element, then the occluding Gaussian basis elements are the first i - 1 three-dimensional Gaussian basis elements.

[0111] Step 430: Cumulatively multiply the difference between one and the opacity of the occluding Gaussian basis element to obtain the occlusion quantization value.

[0112] In one embodiment, the occlusion quantization value is expressed as:

[0113]

[0114] where α j represents the opacity of the j-th occluding Gaussian basis element.

[0115] Step 440: Calculate the product of the opacity of the to-be-measured Gaussian basis element and the occlusion quantization value as the contribution value of the to-be-measured Gaussian basis element.

[0116] In one embodiment, the contribution value of the to-be-measured Gaussian basis element is expressed as:

[0117]

[0118] where α i represents the opacity of the i-th three-dimensional Gaussian basis element.

[0119] With the contribution degree of each three-dimensional Gaussian basis element to the pixel position, select the three-dimensional Gaussian basis element corresponding to the maximum contribution value as the target Gaussian basis element of the pixel position. Therefore, the target Gaussian basis element is expressed as:

[0120]

[0121] where GIM n represents the target Gaussian basis element of the n-th pixel position.

[0122] Then calculate the target Gaussian basis element of each pixel position in the above manner.

[0123] Step 320: At each viewing angle, construct an initial basis element sequence according to the target Gaussian basis element, multiply the initial basis element sequence by the motion mask to obtain a target basis element sequence, and remove duplicates of the three-dimensional Gaussian basis elements in the target basis element sequences of all viewing angles to obtain the surface motion basis element.

[0124] In one embodiment, at each viewing angle, the initial basis element sequence can be constructed by the target Gaussian basis elements of all pixel positions, and the motion mask is a set composed of the mask values of all pixel positions. Since only the mask value of the motion area in the motion mask is 1 and the others are 0, multiplying the initial basis element sequence by the motion mask can filter out the target Gaussian basis elements related to the motion area at this viewing angle.

[0125] Specifically, for each target Gaussian basis element, check whether the area projected onto the image plane overlaps with the pixels with a value of 1 in the motion mask. If there is an overlap, retain the target Gaussian basis element; otherwise, delete it. The finally obtained target basis element sequence only contains the target Gaussian basis elements related to the moving target.

[0126] Since the target Gaussian basis elements may be the same for adjacent pixel positions or the same pixel position under different viewing angles, there may be duplicate 3D Gaussian basis elements in the target basis element sequence. At this time, the target basis element sequences of all viewing angles can be de-duplicated according to the index, and the obtained set of Gaussian basis elements is called the surface motion basis elements corresponding to the acquisition frame. Using the surface motion basis elements, the moving targets in the target scene under all viewing angles can be accurately obtained, reducing the interference of static areas.

[0127] As can be seen from the above process, the surface motion basis elements are expressed as:

[0128]

[0129] Among them, G o represents the surface motion basis elements, V represents the total number of viewing angles, GIM i represents the initial basis element sequence of the i-th viewing angle, M i represents the motion mask of the i-th viewing angle, GIM i ·M i represents the target basis element sequence of the i-th viewing angle obtained through dot product operation.

[0130] In one embodiment, next, the surface motion basis elements are clustered into different clusters, and each cluster can represent a different object surface. The clustering can be DBSCAN clustering based on position, and the clustering process is expressed as:

[0131]

[0132] Among them, represents the i-th surface motion basis element, L represents a certain cluster of clustering, N represents the total number of clusters, represents the position of the i-th 3D Gaussian basis element in the i-th surface motion basis element, represents the position of the i-th 3D Gaussian basis element in the j-th surface motion basis element, ∈ represents the threshold, which can be set to 2 according to experience, and δ() is an indicator function used to judge and Whether the position difference between is less than or equal to the threshold ∈. If the condition is met, δ() returns 1, otherwise it returns 0. is used to judge the entire product result. If the result is 0, then belongs to the cluster L.

[0133] In one embodiment, next, for each cluster, a convex hull operation is performed to locate the 3D Gaussian basis elements related to motion inside the object based on the surface motion basis elements. Refer to Figure 5 , Figure 5It is a flowchart for obtaining motion primitives based on clustering clusters and surface motion primitives provided by an embodiment of the present application, specifically including the following steps:

[0134] Step 510: Perform a convex hull operation on each clustering cluster to obtain the convex boundary corresponding to the clustering cluster.

[0135] In one embodiment, the Delaunay convex hull operation can be used to obtain the convex boundary corresponding to each clustering cluster. For a certain clustering cluster, obtain the positions of each three-dimensional Gaussian primitive in the clustering cluster, and these position points will be used as the input for Delaunay triangulation and convex hull calculation. Among them, Delaunay triangulation is a method of dividing a point set into triangles, ensuring that no other points are contained within the circumcircles of all triangles. Perform Delaunay triangulation on the position point set of the clustering cluster, connect the point set into a continuous surface, and generate a triangular mesh. On the basis of Delaunay triangulation, calculate the convex hull of the point set to generate the convex boundary. The convex hull defines the boundary of the object surface, separating the inside and outside of the object. The convex boundary can be understood as enclosing the positions of each three-dimensional Gaussian primitive in the clustering cluster with a minimum polygon mesh.

[0136] Step 520: Take all three-dimensional Gaussian primitives located within the convex boundary as internal Gaussian primitives.

[0137] In one embodiment, since the convex boundary is equivalent to enclosing the three-dimensional Gaussian primitives corresponding to a clustering cluster with a mesh, and this mesh is obtained based on the three-dimensional Gaussian primitives on the object surface, any point located inside the convex hull is considered to be a point inside the object, and all three-dimensional Gaussian primitives located within the convex boundary are taken as internal Gaussian primitives.

[0138] Step 530: Obtain motion primitives based on the surface motion primitives and all internal Gaussian primitives.

[0139] In one embodiment, summing up the surface motion primitives and all internal Gaussian primitives can obtain all motion primitives including the object surface and interior, expressed as:

[0140]

[0141] Among them, G L represents the internal Gaussian primitive corresponding to the L-th clustering cluster, G o represents the surface motion primitive, and G m represents the motion primitive.

[0142] Next, motion optimization needs to be performed based on the motion primitives to achieve the modeling of each frame in a dynamic scenario.

[0143] In one embodiment, refer to Figure 6 , Figure 6It is a flowchart of obtaining the change residual data corresponding to the current frame by predicting changes based on motion primitives provided by an embodiment of the present application, specifically including the following steps:

[0144] Step 610: Obtain the first position data of each motion primitive, and input the first position data into a pre-trained rigid change network for displacement prediction to obtain the position offset and rotation offset corresponding to each motion primitive in the current frame.

[0145] In one embodiment, the motion primitive is obtained based on the three-dimensional Gaussian primitive of the previous frame. At this time, the position of the motion primitive is obtained as the first position data, and then the first position data is input into a pre-trained rigid change network to predict its position change and rotation change. The input of the rigid change network is the first position data of the motion primitive, and the output is the position offset and rotation offset of each motion primitive in the current frame.

[0146] In one embodiment, the rigid change network can be implemented by combining hash network encoding with a lightweight MLP. Among them, hash network encoding is used to efficiently map high-dimensional first position data to a low-dimensional feature space, reducing computational complexity; the lightweight MLP is used to learn the non-linear relationship between position and rotation changes to generate accurate prediction results. In this way, the rigid change network can quickly and accurately predict the motion state of the motion primitive in the current frame, and obtain the position offset and rotation offset of each motion primitive in the current frame. By combining hash network encoding and lightweight MLP, the rigid change network significantly improves computational efficiency while ensuring prediction accuracy, and can adapt to dynamic changes in complex scenarios, such as rapid movement or rotation of objects. It can be understood that the above is only a schematic illustration of the structure of the rigid change network, and does not represent a limitation on it. For example, the hash grid can be replaced by a tri-planar voxel grid or Plenoxels, etc., or only using a grid or only using an MLP to model is also possible.

[0147] For example, in a dynamic indoor scene, the motion primitives of the previous frame describe the local features of a person's arms, legs, and torso. By predicting the position offset and rotation offset of these motion primitives in the current frame through the rigid change network, the motion state of the person can be updated in real time. If a person's arm changes from a vertical state to a horizontal state, the rigid change network will predict the position offset and rotation offset of the arm, thereby accurately describing its motion trajectory.

[0148] In one embodiment, the above displacement prediction process is expressed as:

[0149]

[0150] Among them, represents the first position data of the motion primitive in the previous frame, represents a rigid transformation network, Δu t and Δq t respectively represent the position offset and rotation offset of the motion primitive compared to the previous frame.

[0151] Step 620: Determine the changing motion primitive based on the motion primitive, position offset, rotation offset, current frame image, and motion primitive.

[0152] In the dynamic analysis of the target scene, the motion between two consecutive frames can be divided into two categories: motion continuation and sudden appearance. For the case of motion continuation, that is, the same object appears in both the previous and current frames but with a change in the motion state, it can be characterized by the above displacement prediction process. Specifically, using the motion primitive of the previous frame as input, the rigid transformation network predicts its position offset and rotation offset in the current frame, thereby describing the change in the motion state of the object. This prediction process can accurately capture the motion trajectory of the object. For example, if the coffee pot was picked up in the previous frame and its position and attitude change in the current frame, the rigid transformation network can predict its displacement and rotation, providing a continuous and consistent motion trajectory for dynamic scene analysis.

[0153] For the case of sudden appearance, that is, an object that did not appear in the previous frame but suddenly appears in the current frame, it is necessary to further determine which existing three-dimensional Gaussian primitives are needed to simulate the "suddenly appeared" target in the current frame. These three-dimensional Gaussian primitives used to simulate the new target are called changing motion primitives. The selection of changing motion primitives can be based on the context information of the scene and the geometric features of the object. For example, if the coffee pot was picked up in the previous frame to pour coffee and coffee suddenly appears in the current frame, by analyzing the position and attitude of the coffee pot, the relevant three-dimensional Gaussian primitives can be selected to simulate the appearance of the coffee. These changing motion primitives can accurately describe the geometric shape, color, and density information of the new target and can adapt to the dynamic changes in complex scenes.

[0154] In one embodiment, referring to Figure 7 , Figure 7 is a flowchart for determining the changing motion primitive based on the motion primitive, position offset, rotation offset, current frame image, and motion primitive provided by the embodiment of the present application, specifically including the following steps:

[0155] Step 710: Displace the motion primitive based on the corresponding position offset and rotation offset to obtain the changing motion primitive of the current frame.

[0156] In one embodiment, after obtaining the position offset and rotation offset, the motion primitive can be displaced according to the position offset and rotation offset to obtain the changing motion primitive corresponding to each motion unit in the current frame, expressed as:

[0157] ut = u t-1 + Δu t

[0158] q t = q t-1 × Δq t

[0159] Wherein, u t represents the position of the changing motion primitive of the current frame, and q t represents the rotation angle of the changing motion primitive of the current frame.

[0160] Step 720: Perform image rendering based on all changing motion primitives to obtain the rendered image of the current frame.

[0161] In one embodiment, replace the corresponding three-dimensional Gaussian primitives in the three-dimensional Gaussian primitives of the previous frame with the changing primitives, and perform image rendering using all the replaced three-dimensional Gaussian primitives, then the rendered image of the current frame from each perspective can be obtained.

[0162] Step 730: For each pixel position, calculate the difference value according to the rendered image and the current frame image, and obtain the attention value according to the difference value.

[0163] In one embodiment, taking a certain perspective as an example, for each pixel position, according to the rendered image I d and the current frame image I t calculate the difference value, and the difference value can be the L1 loss value, expressed as:

[0164]

[0165] Wherein, i and j respectively represent the indexes of the pixel positions on the rendered image and the current frame image.

[0166] Next, according to the comparison between the difference value and the threshold, determine the attention value of each pixel position, expressed as:

[0167]

[0168] Wherein, τ represents the threshold, which can be set to 99 according to experience, represents the attention value obtained by comparing the i-th pixel position of the rendered image and the j-th pixel position of the current frame image.

[0169] Step 740: Obtain the attention map according to the attention values of all pixel positions, and determine the changing motion primitives based on the attention map and the motion primitives.

[0170] In one embodiment, when i = j, an attention map corresponding to each perspective is obtained according to the attention value, and then the attention map is multiplied by the motion primitive to determine the changing motion primitive of the perspective. Then, all perspectives are aggregated and duplicate entries are removed to obtain the changing motion primitive corresponding to the acquisition frame.

[0171] In one embodiment, referring to Figure 8 , Figure 8 is a flowchart for determining the changing motion primitive based on the attention map and the motion primitive provided by an embodiment of the present application, which specifically includes the following steps:

[0172] Step 810: For each pixel position, according to the contribution value of each changing motion primitive to the pixel position, the changing motion primitive corresponding to the maximum contribution value is selected as the attention motion primitive of the pixel position.

[0173] In one embodiment, for each pixel position, the contribution value of each changing motion primitive is also calculated according to the above calculation method of the contribution value to evaluate the contribution degree of each changing motion primitive to the suddenly changing object, and the changing motion primitive corresponding to the maximum contribution value is selected as the attention motion primitive of each pixel position.

[0174] Step 820: An attention primitive sequence is formed according to all the attention motion primitives, the attention primitive sequence is multiplied by the attention map to obtain a changing primitive sequence, and the intersection of the changing primitive sequence and the motion primitive is taken to obtain the changing motion primitive.

[0175] In one embodiment, an attention primitive sequence is formed according to all the attention motion primitives, the attention primitive sequence is multiplied by the attention map to obtain a changing primitive sequence, and then the changing primitive sequences of all perspectives are aggregated. This process is expressed as:

[0176]

[0177] Wherein, represents the attention primitive sequence of the i-th perspective, M ai represents the attention map of the i-th perspective, represents the changing primitive sequence of the i-th perspective, and V represents the number of perspectives.

[0178] Since the three-dimensional Gaussian primitive related to the suddenly appearing object belongs to a subset of the motion primitive, the changing motion primitive G new is obtained by taking the intersection, which is expressed as:

[0179]

[0180] Step 630: Obtain the second position data of each varying motion primitive, and input the second position data into a pre-trained optimization network for color prediction to obtain the color offset corresponding to each varying motion primitive in the current frame.

[0181] In one embodiment, after determining that the three-dimensional Gaussian primitives representing sudden appearance are varying motion primitives, it is necessary to determine the colors of these three-dimensional Gaussian primitives for better simulation. Therefore, first obtain the positions of the varying motion primitives as the second position data, and then input the second position data into a pre-trained optimization network for color prediction to obtain the color offset corresponding to each varying motion primitive in the current frame. The optimization network here can adopt the same network structure as the previous rigid change network, but since the prediction tasks of the two are different, their model parameters are different.

[0182] This process is expressed as:

[0183]

[0184] sh t = sh t-1 + Δsh t

[0185] Wherein, represents the second position data, represents the optimization network, and Δsh t represents the color offset, sh t-1 represents the color value of the three-dimensional Gaussian primitive in the previous frame, and sh t represents the color value of the three-dimensional Gaussian primitive corresponding to the current frame.

[0186] Step 640: Obtain the varying residual data according to the position offset, rotation offset, and color offset.

[0187] In one embodiment, the parts of motion continuation are characterized by motion primitives, and the parts of sudden appearance are characterized by varying motion units. Therefore, for motion primitives, they all include the position offset and rotation offset of the current frame relative to the previous frame. For varying motion units, on the basis of the position offset and rotation offset, they also include the color offset. Summarize these position offsets, rotation offsets, and color offsets to obtain the varying residual data.

[0188] It can be understood that for the objects of "sudden appearance" above, only the color offset is calculated, and it is also possible to perform quantitative calculations on the offsets of parameters such as position, occupancy, ellipsoid axis length, and rotation angle. This embodiment does not limit it.

[0189] In one embodiment, next, parameter prediction of the current frame can be performed based on the three-dimensional Gaussian basis elements of the previous frame and the change residual data corresponding to the current frame. Refer to Figure 9 , Figure 9 is a flowchart for obtaining the Gaussian basis elements of the current frame according to the three-dimensional Gaussian basis elements and the change residual data provided by the embodiments of the present application, specifically including the following steps:

[0190] Step 910: Remove the motion basis elements from the three-dimensional Gaussian basis elements to obtain the static Gaussian basis elements.

[0191] In one embodiment, the changing motion basis elements are a subset of the motion basis elements. Therefore, by removing the motion basis elements from the three-dimensional Gaussian basis elements, the remaining three-dimensional motion basis elements are used as the static Gaussian basis elements. The static Gaussian basis elements can be regarded as the three-dimensional Gaussian basis elements that do not change in terms of displacement, appearance, etc. from the previous frame to the current frame.

[0192] Step 920: Perform parameter changes on the motion basis elements based on at least one of the position offset, rotation offset, or color offset to obtain the changing basis elements of the current frame.

[0193] In one embodiment, for each three-dimensional Gaussian basis element, as long as it has any one of the position offset, rotation offset, or color offset, parameter superposition will be performed according to the included position offset, rotation offset, or color offset to obtain the changing basis elements of the current frame.

[0194] Step 930: Obtain the Gaussian basis elements of the current frame according to the changing basis elements of the current frame and the static Gaussian basis elements.

[0195] In one embodiment, the changing basis elements of the current frame and the static Gaussian basis elements of the previous frame are summarized to obtain the Gaussian basis elements of the current frame corresponding to the current frame. Then, the next frame is selected as the current frame, the current frame is used as the previous frame, and the Gaussian basis elements of the current frame are used as the three-dimensional Gaussian basis elements of the previous frame at the start of iteration, and iteration is performed until the Gaussian basis elements of the current frame corresponding to each acquired frame are obtained.

[0196] In one embodiment, considering the storage and transmission costs, except for saving the initial three-dimensional Gaussian basis elements in the first frame, only the relevant parameters of the three-dimensional Gaussian basis elements related to motion are stored in other subsequent frames. Specifically, the subsequent frames only save: the index of the motion basis elements, the index of the changing motion basis elements, and the change residual data. Compared with the full amount of data, especially in the case of more static regions, this method can greatly improve the storage efficiency and reduce the storage cost.

[0197] Step 130: Obtain the rendering Gaussian basis elements according to the initial three-dimensional Gaussian basis elements of the first frame and the Gaussian basis elements of the current frame of other frames.

[0198] In one embodiment, for the rendering Gaussian basis elements used in the rendering process, the rendering Gaussian basis elements are used for rendering to generate a target video. When rendering the first frame, the rendering Gaussian basis elements are the initial three-dimensional Gaussian basis elements. For other frames, the current frame Gaussian basis elements obtained according to the above process are used as the rendering Gaussian basis elements. The rasterization rendering process of three-dimensional Gaussian splashing is directly used to render the corresponding rendering Gaussian basis elements to obtain three-dimensional data corresponding to each view, and a corresponding target video is generated according to the three-dimensional data and the actually selected view. It can be understood that the generation and rendering processes of the rendering Gaussian basis elements can be separated and carried out on different devices or on the same device, and the main body is determined according to actual needs, which is not limited in this embodiment.

[0199] In one embodiment, the process of generating the initial three-dimensional Gaussian basis elements can be generated by using a neural network model related to three-dimensional Gaussian splashing. Therefore, at least three models including this neural network model, the rigid transformation network, and the optimization network are included in the embodiments of the present application. When training the models, the photometric loss between images and the D-SSIM loss can be used to constrain the parameter terms at each stage.

[0200] The loss function is calculated as follows:

[0201]

[0202] where represents the value of the loss function, represents the photometric loss between images, represents the D-SSIM loss, and I respectively represent the image rendered from the training view and the corresponding ground truth image, and λ represents the loss coefficient, which can be set to 0.2 according to empirical values.

[0203] When training the neural network model corresponding to the three-dimensional Gaussian basis elements, the loss function is used to optimize the calculation processes of the position, occupancy, ellipsoidal axis length, rotation angle, and color of the three-dimensional Gaussian basis elements. When training the rigid transformation network, the loss function is used to optimize the calculation processes of the position offset and rotation offset corresponding to the motion area. When training the optimization network, the loss function is used to optimize the calculation process of the color offset corresponding to the motion area. It can be understood that the true value images are set according to the corresponding prediction tasks in the three stages.

[0204] In one embodiment, referring to Figure 10 , Figure 10 is the overall flowchart of the video generation method in a three-dimensional scene provided by the embodiments of the present application. Referring to Figure 10 , the video generation method in a three-dimensional scene of the embodiments of the present application can be divided into three parts, namely the initialization part, the motion positioning part, and the change and optimization part.

[0205] At the beginning of initialization, starting from the acquisition frame T = 0, multiple image acquisition devices are used to capture videos from different perspectives. After obtaining the scene image of the first frame, the initialization process of the three-dimensional Gaussian basis element is carried out accordingly to generate multiple initial three-dimensional Gaussian basis elements.

[0206] Next, enter the motion localization part. Taking the T-th frame as the current frame, obtain the three-dimensional Gaussian basis elements of the previous frame, as well as the current frame image and the previous frame image in each perspective. Calculate the motion mask based on the current frame image and the previous frame image. Obtain the surface motion basis elements according to the three-dimensional Gaussian basis elements and all motion masks. Cluster the surface motion basis elements to obtain at least one cluster. Obtain the motion basis elements according to the cluster and the surface motion basis elements. Then enter the change and optimization part. Predict the change based on the motion basis elements to obtain the change residual data corresponding to the current frame. Obtain the current frame Gaussian basis elements in each perspective according to the three-dimensional Gaussian basis elements and the change residual data. Next, for each subsequent frame, the motion localization part uses the three-dimensional Gaussian basis elements trained in the previous frame as initialization and locates the motion basis elements related to motion in the next frame. The change and optimization part will train the change motion basis elements located in the previous part. After that, obtain the current frame Gaussian basis elements of each frame, store them, and transmit them to the web page or the player end of the local area according to actual needs. It can be understood that during the training process, the loss function can be calculated using the ground truth image to update the models of the rigid change network and the optimization network.

[0207] In one embodiment, refer to Figure 11 , Figure 11 is the flowchart of the motion localization part and the change and optimization part of the video generation method in the three-dimensional scene provided by the embodiment of the present application.

[0208] In the motion localization part, first, for each perspective, the optical flow difference at each pixel position is obtained by performing optical flow motion estimation based on the current frame image and the previous frame image. The optical flow difference is compared with a preset optical flow value to generate a significant motion mask. The frame difference at each pixel position is obtained by performing frame difference motion estimation based on the current frame image and the previous frame image. The frame difference is compared with a preset frame difference value to generate a minor motion mask. Then, the significant motion mask and the minor motion mask are fused to obtain a motion mask. As can be seen from the schematic diagram of the motion mask, the static area is the black part in the motion mask, and it can also be considered that the mask value is 0, while the mask value of the motion area is 1. Next, surface motion primitives are obtained based on three-dimensional Gaussian basis elements and all motion masks. The surface motion primitives are clustered into different clusters, and each cluster can represent a different object surface. Then, a convex hull operation is performed on each cluster to obtain the convex boundary corresponding to the cluster. The cluster L1 and the cluster L2 are schematically shown in the figure. All three-dimensional Gaussian basis elements located within the convex boundary are used as internal Gaussian basis elements, and all motion primitives are obtained based on the surface motion primitives and all internal Gaussian basis elements. During the training process, the rigid transformation network needs to be updated using the optimized partial rendering graph and the corresponding ground truth image.

[0209] In the deformation and optimization part, based on the motion primitives, first, the sequence of change primitives for all perspectives is determined based on the motion primitives, position offset, rotation offset, the current frame image, and the motion primitives. Then, the intersection of the sequence of change primitives and the motion primitives is calculated to determine the changed motion primitives. In this process, the corresponding rendering image needs to be obtained using the three-dimensional Gaussian splashing technique. During the training process, the optimization network needs to be updated using the deformed partial rendering graph and the corresponding ground truth image.

[0210] An object of an embodiment of the present application is to be able to achieve high rendering quality, low storage cost, and stable temporal consistency for the flowing dynamic free-viewpoint synthesis given the video streams captured by multiple camera positions. Compared with the flowing dynamic free-viewpoint synthesis in the related art, the embodiment of the present application explicitly separates the dynamic and static regions in the scene and only trains, optimizes, and stores the part of the parameters related to the dynamic in the scene, achieving high-fidelity rendering quality while maintaining a high temporal consistency for free-viewpoint viewing. In addition, in the optimization stage, the embodiment of the present application locates and optimizes the newly emerged objects using the maximum pixel error of the rendering result, which can more accurately and efficiently find the pixels of the "suddenly emerged" objects, thereby maintaining a high temporal stability of the static region in the scene and avoiding the waste of unnecessary static region training, transmission, and storage costs.

[0211] The technical solution provided by the embodiments of this application is as follows: obtain the scene images corresponding to each acquisition frame from multiple perspectives, and generate multiple initial three-dimensional Gaussian primitives based on the scene images of the first frame of each perspective. Next, obtain the three-dimensional Gaussian primitives of the previous frame, as well as the current frame image and the previous frame image under each perspective, calculate the motion mask based on the current frame image and the previous frame image, obtain the surface motion primitives based on the three-dimensional Gaussian primitives and all motion masks, cluster the surface motion primitives to obtain at least one cluster, obtain the motion primitives based on the cluster and the surface motion primitives, perform change prediction based on the motion primitives to obtain the change residual data corresponding to the current frame, obtain the current frame Gaussian primitives based on the three-dimensional Gaussian primitives and the change residual data, use the current frame Gaussian primitives as the three-dimensional Gaussian primitives of the next frame, and perform iteration. Among them, the initial value of the three-dimensional Gaussian primitives is the initial three-dimensional Gaussian primitives, and the initial value of the previous frame image is the scene image of the first frame. Finally, render based on the initial three-dimensional Gaussian primitives of the first frame and the current frame Gaussian primitives of other frames to generate the target video. In the embodiments of this application, first calculate the motion mask through the images of two adjacent frames to accurately identify the dynamic areas in the scene. Then, based on the motion mask, determine the surface motion primitives and the motion primitives to further distinguish the dynamic and static areas and avoid the interference of the dynamic areas on the static areas. Then, update the three-dimensional Gaussian primitives iteratively to ensure that the change characteristics of the dynamic areas are consistent between frames. At the same time, generate the change residual data between frames to accurately describe the changes of the dynamic areas between adjacent frames and avoid the flicker artifacts caused by inconsistent changes of the dynamic areas. It can effectively improve the generation quality of the video, has a wide range of applications, and significantly improves the accuracy and stability of the synthesized video.

[0212] The embodiments of this application also provide a video generation device in a three-dimensional scene, which can implement the above video generation method in a three-dimensional scene. Refer to Figure 12 and this device includes:

[0213] Initialization module 1210: used to obtain the scene images corresponding to each acquisition frame from multiple perspectives, and generate multiple initial three-dimensional Gaussian primitives based on the scene images of the first frame of each perspective.

[0214] Iterative module 1220: It is used to obtain the three-dimensional Gaussian basis elements of the previous frame, as well as the current frame image and the previous frame image in each viewing angle. Calculate the motion mask based on the current frame image and the previous frame image. Obtain the surface motion basis elements according to the three-dimensional Gaussian basis elements and all the motion masks. Cluster the surface motion basis elements to obtain at least one cluster. Obtain the motion basis elements according to the cluster and the surface motion basis elements. Perform change prediction according to the motion basis elements to obtain the change residual data corresponding to the current frame. Obtain the current frame Gaussian basis elements according to the three-dimensional Gaussian basis elements and the change residual data. Use the current frame Gaussian basis elements as the three-dimensional Gaussian basis elements of the previous frame and perform iteration. The initial value of the three-dimensional Gaussian basis elements is the initial three-dimensional Gaussian basis elements, and the initial value of the previous frame image is the scene image of the first frame.

[0215] Rendering module 1230: It is used to obtain the rendering Gaussian basis elements according to the initial three-dimensional Gaussian basis elements of the first frame and the current frame Gaussian basis elements of other frames. The rendering Gaussian basis elements are used for rendering to generate the target video.

[0216] The specific implementation of the three-dimensional scene video generation device in this embodiment is basically the same as that of the above-mentioned three-dimensional scene video generation method, and will not be elaborated here.

[0217] This application embodiment also provides an electronic device, including:

[0218] At least one memory;

[0219] At least one processor;

[0220] At least one program;

[0221] The program is stored in the memory, and the processor executes the at least one program to implement the three-dimensional scene video generation method described above in this application. This electronic device can be any intelligent terminal including mobile phones, tablet computers, personal digital assistants (Personal Digital Assistant, PDA), in-vehicle computers, etc.

[0222] Please refer to Figure 13 , Figure 13 which schematically shows the hardware structure of an electronic device in another embodiment. The electronic device includes:

[0223] Processor 1301, which can be implemented in ways such as a general-purpose central processing unit (Central Processing Unit, CPU), a microprocessor, an application-specific integrated circuit (Application Specific Integrated Circuit, ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in this application embodiment;

[0224] The memory 1302 can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM), etc. The memory 1302 can store an operating system and other application programs. When implementing the technical solutions provided in the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 1302 and are called by the processor 1301 to execute the method for generating a video in a three-dimensional scene according to the embodiments of this application;

[0225] The input / output interface 1303 is used to implement information input and output;

[0226] The communication interface 1304 is used to implement communication interaction between this device and other devices. Communication can be achieved through wired means (such as USB, network cable, etc.) or through wireless means (such as mobile network, WIFI, Bluetooth, etc.);

[0227] The bus 1305 transmits information between various components of the device (such as the processor 1301, the memory 1302, the input / output interface 1303, and the communication interface 1304);

[0228] Among them, the processor 1301, the memory 1302, the input / output interface 1303, and the communication interface 1304 achieve communication connections with each other inside the device through the bus 1305.

[0229] The embodiments of this application also provide a storage medium. The storage medium is a storage medium that stores a computer program. When the computer program is executed by a processor, the above method for generating a video in a three-dimensional scene is implemented.

[0230] As a non-transitory storage medium, the memory can be used to store non-transitory software programs and non-transitory computer executable programs. In addition, the memory can include a high-speed random access memory, and can also include non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory optionally includes a memory remotely set relative to the processor, and these remote memories can be connected to the processor through a network. Examples of the above networks include, but are not limited to, the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.

[0231] The video generation method, device, equipment and storage medium proposed in the embodiments of the present application obtain the scene images corresponding to each acquisition frame from multiple perspectives, and generate multiple initial three-dimensional Gaussian primitives based on the scene images of the first frame of each perspective. Next, obtain the three-dimensional Gaussian primitives of the previous frame, as well as the current frame image and the previous frame image of each perspective, calculate the motion mask based on the current frame image and the previous frame image, obtain the surface motion primitives according to the three-dimensional Gaussian primitives and all motion masks, cluster the surface motion primitives to obtain at least one cluster, obtain the motion primitives according to the cluster and the surface motion primitives, perform change prediction based on the motion primitives to obtain the change residual data corresponding to the current frame, obtain the current frame Gaussian primitives according to the three-dimensional Gaussian primitives and the change residual data, use the current frame Gaussian primitives as the three-dimensional Gaussian primitives of the next frame, and perform iteration, where the initial value of the three-dimensional Gaussian primitives is the initial three-dimensional Gaussian primitives, and the initial value of the previous frame image is the scene image of the first frame. Finally, obtain the rendering Gaussian primitives according to the initial three-dimensional Gaussian primitives of the first frame and the current frame Gaussian primitives of other frames. In the embodiments of the present application, first calculate the motion mask through the images of two adjacent frames to accurately identify the dynamic area in the scene. Then, determine the surface motion primitives and motion primitives based on the motion mask to further distinguish the dynamic and static areas and avoid the interference of the dynamic area on the static area. Then, update the three-dimensional Gaussian primitives iteratively to ensure that the change characteristics of the dynamic area are consistent between frames, and at the same time generate the change residual data between frames to accurately describe the change of the dynamic area between adjacent frames and avoid the flicker artifacts caused by inconsistent changes in the dynamic area. It can effectively improve the quality of video generation, has a wide range of applications, and significantly improves the accuracy and stability of the synthesized video.

[0232] The embodiments described in the embodiments of the present application are for more clearly explaining the technical solutions of the embodiments of the present application, and do not constitute a limitation to the technical solutions provided by the embodiments of the present application. Those skilled in the art know that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of the present application are also applicable to similar technical problems.

[0233] Those skilled in the art can understand that the technical solutions shown in the figure do not constitute a limitation to the embodiments of the present application, and may include more or fewer steps than those shown in the figure, or combine some steps, or different steps.

[0234] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0235] Those of ordinary skill in the art will understand that all or some of the steps in the methods disclosed above, and the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, or a suitable combination thereof.

[0236] As used in the description of the present application and the above drawings, the terms "first", "second", "third", "fourth", etc. (if any) are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data may be interchanged under appropriate circumstances so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that comprises a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.

[0237] It should be understood that in the present application, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated objects and indicates that there can be three relationships. For example, "A and / or B" can mean: only A exists, only B exists, and both A and B exist simultaneously. Here, A and B can be singular or plural. The character " / " generally means that the associated objects before and after are in an "or" relationship. "At least one (one)" or a similar expression thereof refers to any combination of these items, including any combination of single items (ones) or plural items (ones). For example, at least one (one) of a, b, or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.

[0238] In several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the above division of units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces, and the indirect coupling or communication connection of devices or units can be in an electrical, mechanical, or other form.

[0239] The units described above as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0240] In addition, each functional unit in various embodiments of the present application may be integrated into a processing unit, or each unit may exist physically alone, or two or more units may be integrated into one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.

[0241] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in various embodiments of the present application. The foregoing storage medium includes: various media that can store programs such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs.

[0242] The preferred embodiments of the embodiments of the present application have been described above with reference to the accompanying drawings, and thus do not limit the scope of rights of the embodiments of the present application. Any modification, equivalent replacement, and improvement made by those skilled in the art without departing from the scope and essence of the embodiments of the present application shall be within the scope of rights of the embodiments of the present application.

Claims

1. A method for generating a video in a three-dimensional scene, characterized in that: include: Acquire scene images corresponding to multiple viewing angles in each acquisition frame, and generate multiple initial three-dimensional Gaussian primitives according to the scene image of the first frame of each viewing angle; Acquire a three-dimensional Gaussian primitive of a previous frame and a current frame image and a previous frame image under each of the viewing angles, calculate a motion mask based on the current frame image and the previous frame image, obtain a surface motion primitive according to the three-dimensional Gaussian primitive and all the motion masks, cluster the surface motion primitives to obtain at least one cluster cluster, obtain a motion primitive according to the cluster cluster and the surface motion primitive, perform change prediction according to the motion primitive to obtain change residual data corresponding to the current frame, obtain a current frame Gaussian primitive according to the three-dimensional Gaussian primitive and the change residual data, use the current frame Gaussian primitive as the three-dimensional Gaussian primitive of the previous frame, perform iteration, the initial value of the three-dimensional Gaussian primitive is the initial three-dimensional Gaussian primitive, and the initial value of the previous frame image is the scene image of the first frame; A rendering Gaussian primitive is obtained according to the initial three-dimensional Gaussian primitive of the first frame and the current frame Gaussian primitives of other frames, and the rendering Gaussian primitive is used for rendering to generate a target video.

2. The method for generating a video in a three-dimensional scene according to claim 1, characterized in that: The calculating and obtaining a motion mask based on the current frame image and the previous frame image includes: Performing optical flow motion estimation based on the current frame image and the previous frame image to obtain an optical flow difference value at each pixel position, comparing the optical flow difference value with a preset optical flow value, and generating a significant motion mask; Performing frame difference motion estimation based on the current frame image and the previous frame image to obtain a frame difference at each pixel position, comparing the frame difference with a preset frame difference value, and generating a slight motion mask; The motion mask is obtained by fusing the significant motion mask and the subtle motion mask.

3. The method for generating a video in a three-dimensional scene according to claim 1, characterized in that: The step of obtaining a surface motion primitive according to the three-dimensional Gaussian primitive and all the motion masks comprises: For each pixel position, calculating the contribution value of each of the three-dimensional Gaussian primitives to the pixel position, and selecting the three-dimensional Gaussian primitive corresponding to the maximum value of the contribution value as the target Gaussian primitive of the pixel position; At each viewing angle, an initial primitive sequence is constructed according to the target Gaussian primitives, the target primitive sequence is obtained by multiplying the initial primitive sequence with the motion mask, and the three-dimensional Gaussian primitives in the target primitive sequence of all viewing angles are deduplicated to obtain the surface motion primitives.

4. The method for generating a video in a three-dimensional scene according to claim 3, characterized in that: The calculating the contribution value of each of the three-dimensional Gaussian primitives to the pixel position comprises: Obtaining the opacity of each of the three-dimensional Gaussian primitives; The three-dimensional Gaussian basis elements are used one by one as Gaussian basis elements to be tested, and the three-dimensional Gaussian basis elements located in front of the Gaussian basis elements to be tested are obtained in sequence as blocking Gaussian basis elements; Multiplying the difference between one and the opacity of the occluded Gaussian primitive to obtain an occlusion quantization value; The product of the opacity of the Gaussian primitive to be tested and the occlusion quantization value is calculated as the contribution value of the Gaussian primitive to be tested.

5. The method for generating a video in a three-dimensional scene according to claim 1, characterized in that: The obtaining of motion primitives according to the clusters and the surface motion primitives comprises: Performing a convex hull operation on each of the clusters to obtain a convex boundary corresponding to the cluster; All three-dimensional Gaussian primitives within the convex boundary are regarded as internal Gaussian primitives; The motion primitive is obtained according to the surface motion primitive and all the internal Gaussian primitives.

6. The method for generating a video in a three-dimensional scene according to claim 3, characterized in that: The step of performing change prediction according to the motion primitive to obtain change residual data corresponding to the current frame includes: Acquire first position data of each of the motion primitives, input the first position data into a pre-trained rigid change network for displacement prediction, and obtain a position offset and a rotation offset corresponding to each of the motion primitives in a current frame; Determine a change motion primitive based on the motion primitive, the position offset, the rotation offset, the current frame image and the motion primitive; Acquire second position data of each of the change motion primitives, input the second position data into a pre-trained optimization network for color prediction, and obtain a color offset corresponding to each of the change motion primitives in the current frame; The change residual data is obtained according to the position offset, the rotation offset and the color offset.

7. The method for generating a video in a three-dimensional scene according to claim 6, characterized in that: The step of determining a change motion primitive based on the motion primitive, the position offset, the rotation offset, the current frame image and the motion primitive comprises: Performing displacement on the motion primitive based on the corresponding position offset and the rotation offset to obtain a changed motion primitive of the current frame; Performing image rendering based on all the changing motion primitives to obtain a rendered image of the current frame; For each pixel position, a difference value is calculated according to the rendered image and the current frame image, and an attention value is obtained according to the difference value; An attention map is obtained according to the attention values ​​of all the pixel positions, and a change motion primitive is determined based on the attention map and the motion primitive.

8. The method for generating a video in a three-dimensional scene according to claim 7, characterized in that: The determining of the change motion primitive based on the attention map and the motion primitive comprises: For each pixel position, according to the contribution value of each of the change motion primitives to the pixel position, the change motion primitive corresponding to the maximum contribution value is selected as the attention motion primitive of the pixel position; An attention primitive sequence is formed according to all the attention motion primitives, the attention primitive sequence is multiplied with the attention map to obtain a change primitive sequence, and the change primitive sequence and the motion primitive are intersected to obtain the change motion primitive.

9. The method for generating a video in a three-dimensional scene according to claim 6, characterized in that: The step of obtaining the current frame Gaussian primitives according to the three-dimensional Gaussian primitives and the change residual data includes: Removing the moving primitive from the three-dimensional Gaussian primitive to obtain a stationary Gaussian primitive; Perform parameter change on the motion primitive based on at least one of the position offset, the rotation offset or the color offset to obtain a current frame change primitive; The current frame Gaussian primitive is obtained according to the current frame change primitive and the static Gaussian primitive.

10. A video generation device in a three-dimensional scene, characterized in that: include: Initialization module: used for acquiring scene images corresponding to multiple viewing angles in each acquisition frame, and generating multiple initial three-dimensional Gaussian primitives according to the scene image of the first frame of each viewing angle; Iteration module: used for obtaining the three-dimensional Gaussian primitive of the previous frame and the current frame image and the previous frame image under each of the viewing angles, calculating a motion mask based on the current frame image and the previous frame image, obtaining a surface motion primitive according to the three-dimensional Gaussian primitive and all the motion masks, clustering the surface motion primitives to obtain at least one cluster cluster, obtaining a motion primitive according to the cluster cluster and the surface motion primitive, performing change prediction according to the motion primitive to obtain change residual data corresponding to the current frame, obtaining a current frame Gaussian primitive according to the three-dimensional Gaussian primitive and the change residual data, taking the current frame Gaussian primitive as the three-dimensional Gaussian primitive of the previous frame, and performing iteration, wherein the initial value of the three-dimensional Gaussian primitive is the initial three-dimensional Gaussian primitive, and the initial value of the previous frame image is the scene image of the first frame; Rendering module: used for obtaining rendering Gaussian primitives according to the initial three-dimensional Gaussian primitives of the first frame and the current frame Gaussian primitives of other frames, wherein the rendering Gaussian primitives are used for rendering to generate a target video.

11. An electronic device, characterized in that: The electronic device includes a memory and a processor, the memory stores a computer program, and the processor implements the video generation method in a three-dimensional scene as described in any one of claims 1 to 9 when executing the computer program.

12. A storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method for generating a video in a three-dimensional scene according to any one of claims 1 to 9 is implemented.

Citation Information

Cited By

  • Video fusion method and system based on large model

    CN120568159A

  • Video processing method and device, electronic equipment and storage medium

    CN121585886A

  • Method, apparatus, electronic device and storage medium for video processing

    CN121585886B