Three-dimensional video processing method and device, equipment and storage medium
By generating and quantizing the encoded keyframe representation model and transformation parameters on the encoding end, the problem of time domain error accumulation in three-dimensional video processing is solved, and the reconstruction performance of three-dimensional video is improved.
Patent Information
- Application Number
- CN202510012414.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-02
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-01-02
AI Technical Summary
There is a problem of time domain error accumulation in three-dimensional video processing, which affects the reconstruction performance of three-dimensional video.
A three-dimensional video processing method is proposed, by generating keyframe representation models on the encoding end and performing quantization codes, generating transformation parameters and performing quantization codes, and sending them to the decoding end for decoding and residual fusion to offset the time domain error.
It effectively reduces the accumulation of time domain errors, improves the reconstruction performance of three-dimensional videos, and ensures the accuracy and reliability of the video data processing process.
Smart Images

Figure CN119996694A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer vision technology, and in particular to a three-dimensional video processing method, device, equipment and storage medium. Background Art
[0002] 3D Gaussian Splatting (3DGS) technology has surpassed early 3D scene representation technologies such as point cloud, mesh, neural radiance field (NeRF) in terms of 3D scene reconstruction quality, reconstruction speed, interactive freedom and rendering speed, and has become a widely used 3D scene representation method today.
[0003] In the related art, the 3D Gaussian splashing technology uses a neural network to model the rotation and translation changes of Gaussian points at adjacent moments, and uses the inter-frame prediction results as a reference for the next frame for 3D reconstruction. However, this method has the problem of temporal error accumulation, which affects the reconstruction performance of 3D video. Summary of the invention
[0004] The main purpose of the embodiments of the present application is to propose a three-dimensional video processing method, device, equipment and storage medium to reduce the accumulation of time domain errors and improve the reconstruction performance of three-dimensional video.
[0005] To achieve the above-mentioned purpose, a first aspect of an embodiment of the present application proposes a three-dimensional video processing method, which is applied at an encoding end. The method includes:
[0006] At multiple acquisition moments, video frames corresponding to the target scene at different viewing angles are obtained respectively, and the video frames corresponding to the initial acquisition moment are used as key frame sequences, and the video frames corresponding to other acquisition moments are used as non-key frame sequences;
[0007] Generate a key frame representation model based on the key frame sequence, perform a quantization encoding operation on the key frame representation model to obtain a key frame encoding model, and decode the key frame encoding model to obtain a key frame decoding model;
[0008] The non-key frame sequences are selected one by one as processing frame sequences, a reference frame sequence is obtained according to a previous acquisition moment of the processing frame sequence, transformation parameters are generated according to a decoding model corresponding to the reference frame sequence, a quantization encoding operation is performed on the transformation parameters to obtain coded transformation parameters, and the coded transformation parameters are sent to a decoding end. If a residual flag exists, a residual representation model corresponding to the processing frame sequence is generated based on the transformation parameters, a quantization encoding operation is performed on the residual representation model to obtain a residual coding model, and the residual coding model is sent to a decoding end. The decoding model corresponding to the first processing frame sequence is the key frame decoding model.
[0009] In one embodiment, the step of performing a quantization encoding operation on the key frame representation model to obtain a key frame encoding model includes:
[0010] Acquire a quantization step length, and perform a quantization operation on a first model parameter of the key frame representation model according to the quantization step length to obtain a quantized representation model;
[0011] The quantized representation model is entropy encoded to obtain a key frame encoding model.
[0012] In one embodiment, obtaining the quantization step size includes:
[0013] Obtaining first center positions of multiple three-dimensional Gaussian distributions in the key frame representation model;
[0014] The first center position is input into a hash list to obtain a first context feature, and the first context feature is input into a quantization prediction model for data processing to obtain the quantization step size.
[0015] In one embodiment, generating transformation parameters according to the decoding model includes:
[0016] Obtaining second center positions of multiple three-dimensional Gaussian distributions in the decoding model;
[0017] The second center position is input into a hash list to obtain a second context feature, and the second context feature is input into a motion transformation prediction model for data processing to obtain the transformation parameters, wherein the transformation parameters include the predicted parameter change of the second model parameter corresponding to the three-dimensional Gaussian distribution in the decoding model at the acquisition moment corresponding to the processing frame sequence.
[0018] In one embodiment, generating the residual representation model corresponding to the processed frame sequence based on the transformation parameters includes:
[0019] Performing inter-frame transformation on the decoding model according to the transformation parameters to obtain a reference decoding model;
[0020] generating the residual identifier according to the reference decoding model;
[0021] The residual representation model is generated based on the residual identifier.
[0022] In one embodiment, performing inter-frame transformation on the decoding model according to the transformation parameters to obtain a reference decoding model includes:
[0023] At least obtaining a center position parameter, a covariance parameter, a color parameter, and an opacity parameter corresponding to each three-dimensional Gaussian distribution of the decoding model;
[0024] According to the transformation parameters, at the acquisition time corresponding to the processing frame sequence, a center position update parameter corresponding to the center position parameter, a covariance update parameter corresponding to the covariance parameter, a color update parameter corresponding to the color parameter, and an opacity update parameter corresponding to the opacity parameter are obtained;
[0025] The three-dimensional Gaussian distribution is updated based on the center position update parameter, the covariance update parameter, the color update parameter, and the opacity update parameter to obtain the reference decoding model.
[0026] In one embodiment, generating the residual identifier according to the reference decoding model includes:
[0027] Rendering the reference decoding model to obtain a reference rendered image corresponding to the processed frame sequence;
[0028] The residual identifier is generated according to a distortion parameter between the reference rendered image and a video frame corresponding to the processed frame sequence.
[0029] In one embodiment, generating the residual representation model includes:
[0030] Determining a residual representation region of the processed frame sequence based on the residual identifier;
[0031] At least one three-dimensional Gaussian distribution is generated for the residual representation area to form the residual representation model.
[0032] In one embodiment, the method further comprises:
[0033] Obtaining a rendering distortion loss item corresponding to the key frame decoding model or the non-key frame decoding model;
[0034] Obtaining a bit rate loss term, where the bit rate loss term is calculated by the decoding model or the residual representation model during a quantization encoding process;
[0035] A rate-distortion joint loss term is calculated according to the rendering distortion loss term and the bit rate loss term, and the rate-distortion joint loss term is used to optimize at least the decoding model, the residual representation model and the transformation parameters.
[0036] To achieve the above-mentioned purpose, a second aspect of an embodiment of the present application proposes a three-dimensional video processing method, which is applied at a decoding end. The method includes:
[0037] Decoding the acquired key frame encoding model to obtain a key frame decoding model;
[0038] For each processed frame sequence, the corresponding encoding transformation parameters are obtained, and the decoding transformation parameters are obtained according to the encoding transformation parameters. The decoding transformation parameters and the decoding model corresponding to the reference frame sequence are used to perform inter-frame transformation to obtain a reference decoding model, and the reference decoding model is used as a non-key frame decoding model. If the processed frame sequence includes a residual coding model, the residual coding model is decoded to obtain a residual decoding model, and the residual decoding model and the reference decoding model are residually fused to obtain the non-key frame decoding model. The decoding model of the first processed frame sequence is the key frame decoding model.
[0039] In one embodiment, decoding the acquired key frame encoding model to obtain the key frame decoding model includes:
[0040] Performing entropy decoding on the key frame coding model to obtain a quantized decoding model;
[0041] A quantization step size is obtained, and a dequantization operation is performed on the parameters of the quantization decoding model according to the quantization step size to obtain the key frame decoding model.
[0042] In one embodiment, the method further comprises:
[0043] Performing two-dimensional projection on a plurality of three-dimensional Gaussian distributions included in the key frame decoding model and the non-key frame decoding model at each viewing angle to obtain a two-dimensional Gaussian distribution at a corresponding viewing angle;
[0044] Performing image rendering of corresponding viewing angles on the two-dimensional Gaussian distribution one by one to obtain target rendered images corresponding to all viewing angles at each acquisition moment of the target scene;
[0045] A three-dimensional scene corresponding to the target scene is generated according to the target rendered image.
[0046] To achieve the above-mentioned purpose, a third aspect of an embodiment of the present application provides a three-dimensional video processing device, which is applied to an encoding end. The device includes:
[0047] Frame division module: used to obtain the corresponding video frames of the target scene at different viewing angles at multiple acquisition moments, and use the video frames corresponding to the initial acquisition moment as the key frame sequence, and the video frames corresponding to other acquisition moments as the non-key frame sequence;
[0048] A key frame encoding module: used for generating a key frame representation model based on the key frame sequence, performing a quantization encoding operation on the key frame representation model to obtain a key frame encoding model, and decoding the key frame encoding model to obtain a key frame decoding model;
[0049] Non-key frame encoding module: used to select the non-key frame sequence as the processing frame sequence one by one, obtain the reference frame sequence according to the previous acquisition moment of the processing frame sequence, generate transformation parameters according to the decoding model corresponding to the reference frame sequence, perform quantization encoding operation on the transformation parameters to obtain the encoding transformation parameters, send the encoding transformation parameters to the decoding end, if there is a residual flag, generate the residual representation model corresponding to the processing frame sequence based on the transformation parameters, perform quantization encoding operation on the residual representation model to obtain the residual encoding model, send the residual encoding model to the decoding end, the decoding model corresponding to the first processing frame sequence is the key frame decoding model.
[0050] To achieve the above-mentioned purpose, the fourth aspect of an embodiment of the present application proposes a storage medium, which is a storage medium. The storage medium stores a computer program, and when the computer program is executed by a processor, the method described in the first aspect is implemented.
[0051] The three-dimensional video processing method, device, equipment and storage medium proposed in the embodiment of the present application are to obtain the video frames corresponding to the target scene at different viewing angles at multiple acquisition moments, respectively, take the video frames corresponding to the initial acquisition moment as the key frame sequence, and take the video frames corresponding to other acquisition moments as the non-key frame sequence, then generate the key frame representation model based on the key frame sequence, perform quantization coding operation on the key frame representation model, obtain the key frame coding model, decode the key frame coding model, and obtain the key frame decoding model, then select the non-key frame sequence one by one as the processing frame sequence, obtain the reference frame sequence according to the previous acquisition moment of the processing frame sequence, generate the transformation parameters according to the decoding model corresponding to the reference frame sequence, perform quantization coding operation on the transformation parameters, obtain the coding transformation parameters and send them to the decoding end, if there is a residual mark, generate the residual representation model corresponding to the processing frame sequence based on the transformation parameters, perform quantization coding operation on the residual representation model, obtain the residual coding model, and send the residual coding model to the decoding end. The embodiment of the present application uses the decoding model of a previous acquisition moment as a reference, generates the transformation parameters according to the decoding model, and uses these transformation parameters to indicate the data change trend and law of the previous and next acquisition moments. Subsequently, the residual representation model corresponding to the acquisition moment is generated according to the transformation parameters, and the residual representation model is used to fit the actual dynamic change trend of the video frame, thereby offsetting the time domain error to the greatest extent, thereby avoiding the continuous accumulation of errors over the time series, thereby ensuring the accuracy and reliability of the entire video data processing process and improving the performance of subsequent three-dimensional video reconstruction. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] Figure 1 It is a flow chart of a three-dimensional video processing method provided in an embodiment of the present application.
[0053] Figure 2This is a flowchart of performing quantization encoding operations on a key frame representation model to obtain a key frame encoding model provided by an embodiment of the present application.
[0054] Figure 3 This is a flowchart for obtaining a quantization step size provided in an embodiment of the present application.
[0055] Figure 4 This is a flowchart of generating transformation parameters according to a decoding model provided in an embodiment of the present application.
[0056] Figure 5 It is a flowchart of a residual representation model corresponding to a frame sequence generated based on transformation parameters provided in an embodiment of the present application.
[0057] Figure 6 It is a flowchart of performing inter-frame transformation on a decoding model according to transformation parameters to obtain a reference decoding model provided by an embodiment of the present application.
[0058] Figure 7 This is a flowchart of generating a residual representation model provided in an embodiment of the present application.
[0059] Figure 8 It is a flow chart of the rate-distortion joint optimization process provided in an embodiment of the present application.
[0060] Fig. 9 It is a schematic diagram of the overall flow of three-dimensional video processing at the encoding end provided in an embodiment of the present application.
[0061] Fig.10 This is an optional flowchart of the three-dimensional video processing method provided in the embodiment of the present application.
[0062] Fig.11 It is a flowchart of decoding the acquired key frame encoding model to obtain the key frame decoding model provided by an embodiment of the present application.
[0063] Fig.12 This is a structural block diagram of a three-dimensional video processing device provided in another embodiment of the present application.
[0064] Fig.13 It is a schematic diagram of the hardware structure of the electronic device provided in the embodiment of the present application. DETAILED DESCRIPTION
[0065] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0066] It should be noted that although the functional modules are divided in the device schematic and the logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than the module division in the device or the order in the flowchart.
[0067] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.
[0068] After obtaining multi-viewpoint video sequences collected synchronously from different perspectives, immersive holographic video technology focuses on completing the high-fidelity reconstruction of three-dimensional dynamic scenes, and must be able to support photo-realistic real-time rendering from any perspective. Immersive communication covers new media business types such as XR communication and holographic communication. Looking back over the past few decades, video coding and decoding technology and corresponding standards have always been built and developed around planar video. However, this can no longer meet the application needs of immersive holographic communication in the future 6G era. With the rapid growth of immersive holographic video data, it poses extremely severe challenges to storage space and network bandwidth. How to achieve immersive holographic video encoding efficiently and quickly has become an important technical problem that needs to be solved urgently.
[0069] The 3D Gaussian Splatting (3DGS) technology has surpassed early 3D scene representation technologies such as point cloud, mesh, neural radiance field (NeRF) in terms of 3D scene reconstruction quality, reconstruction speed, interactive freedom and rendering speed, and has become a widely used 3D scene representation method today, especially in the reconstruction scenes of immersive holographic videos.
[0070] In related technologies, the 3D Gaussian splashing technology uses a neural network to model the rotation and translation changes of Gaussian points at adjacent moments, and uses the inter-frame prediction results as a reference for the next frame to carry out 3D reconstruction. However, when using this method, there is a problem of continuous accumulation of time domain errors, which has a negative impact on the reconstruction performance of 3D video. In addition, most of the 3D Gaussian reconstruction technologies in related technologies lack a compression coding process. On the one hand, this will lead to increased transmission costs and processing time; on the other hand, since it is impossible to jointly optimize the encoding bit rate and distortion, it also has an adverse effect on the reconstruction performance.
[0071] Based on this, the embodiments of the present application provide a three-dimensional video processing method, device, equipment and storage medium, which uses the decoding model of a previous acquisition moment as a reference, generates transformation parameters based on the decoding model, and uses these transformation parameters to indicate the data change trends and laws of the previous and subsequent acquisition moments. Subsequently, a residual representation model corresponding to the acquisition moment is generated according to the transformation parameters, and the residual representation model is used to fit the actual dynamic change trend of the video frame, to offset the time domain error to the greatest extent, thereby avoiding the continuous accumulation of errors with the time series, thereby ensuring the accuracy and reliability of the entire video data processing process and improving the performance of subsequent three-dimensional video reconstruction.
[0072] The embodiments of the present application provide a three-dimensional video processing method, apparatus, device and storage medium, which are specifically described through the following embodiments. First, the three-dimensional video processing method in the embodiments of the present application is described.
[0073] The embodiments of the present application can acquire and process relevant data based on artificial intelligence technology. Among them, artificial intelligence (AI) is the theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a similar way to human intelligence. Artificial intelligence is to study the design principles and implementation methods of various intelligent machines so that the machines have the functions of perception, reasoning and decision-making.
[0074] Artificial intelligence technology is a comprehensive discipline that covers a wide range of fields, including both hardware-level and software-level technologies. The basic technologies of artificial intelligence generally include sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, mechatronics and other technologies. Artificial intelligence software technology mainly includes computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0075] The three-dimensional video processing method provided in the embodiment of the present application relates to the field of computer vision technology. The three-dimensional video processing method provided in the embodiment of the present application can be applied to a terminal, can be applied to a server side, or can be a computer program running in a terminal or a server side. For example, a computer program can be a native program or software module in an operating system; it can be a local (Native) application (Application, APP), that is, a program that needs to be installed in the operating system to run, such as a client that supports three-dimensional video processing, that is, a program that can be run only by downloading it to a browser environment; it can also be a small program that can be embedded in any APP. In short, the above-mentioned computer program can be an application, module or plug-in in any form. Among them, the terminal communicates with the server via a network. The three-dimensional video processing method can be executed by a terminal or a server, or by a terminal and a server in collaboration.
[0076] In some embodiments, the terminal may be a smart phone, a tablet computer, a laptop computer, a desktop computer or a smart watch, etc. In addition, the terminal may also be an intelligent vehicle-mounted device. The intelligent vehicle-mounted device applies the three-dimensional video processing method of this embodiment to provide related services and enhance the driving experience. The server may be an independent server, or it may be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (Content Delivery Network, CDN), and big data and artificial intelligence platforms; it may also be a service node in a blockchain system, and each service node in the blockchain system forms a peer-to-peer (Peer To Peer, P2P) network, and the P2P protocol is an application layer protocol running on the Transmission Control Protocol (Transmission Control Protocol, TCP) protocol. The terminal and the server may be connected via Bluetooth, Universal Serial Bus (Universal Serial Bus, USB) or network and other communication connection methods, which are not limited in this embodiment.
[0077] The present application can be used in many general or special computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, etc. The present application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The present application can also be practiced in distributed computing environments, in which tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.
[0078] It should be noted that in each specific implementation of the present application, when it comes to the need to perform relevant processing based on data related to user identity or characteristics such as user information, user behavior data, user historical data, and user location information, the user's permission or consent will be obtained first, and the collection, use, and processing of these data will comply with relevant laws, regulations, and standards. In addition, when the embodiment of the present application needs to obtain the user's sensitive personal information, the user's separate permission or consent will be obtained through a pop-up window or by jumping to a confirmation page. After clearly obtaining the user's separate permission or consent, the necessary user-related data for the normal operation of the embodiment of the present application will be obtained.
[0079] The following describes a three-dimensional video processing method in an embodiment of the present application.
[0080] In one embodiment, the three-dimensional video processing method involved in the embodiment of the present application needs to be completed by the encoding end and the decoding end in collaboration.
[0081] On the one hand, the encoding end and the decoding end can be deployed in the same processing device. If the processing device resources are limited, then choosing to perform encoding and decoding operations in the device and using encoding operations to compress and integrate data can effectively reduce the amount of data to be processed and reduce the burden on the device so that the device can complete subsequent processing work with limited resources. In addition, if there is a need to complete the reconstruction process as quickly as possible, such as in application scenarios with high real-time requirements, performing encoding and decoding operations in the same device can reduce the time spent on intermediate links such as data transmission, and quickly achieve the conversion from original data to the final reconstruction result, thereby ensuring processing efficiency.
[0082] On the other hand, the encoding and decoding ends can also be deployed in a distributed manner. The encoding operation is first performed on the encoding end, and then the corresponding encoded bitstream is transmitted, and then the decoding operation is performed on the corresponding processing device on the decoding end. In the case of limited transmission bandwidth, transmitting the encoded bitstream can effectively reduce the amount of transmitted data and improve transmission efficiency. In this way, the entire reconstruction process will not be affected by delays in the transmission link, thereby improving the overall reconstruction efficiency.
[0083] It is understandable that this embodiment has no specific restrictions on the deployment forms of the encoding end and the decoding end. The following describes in detail the processing flow of the encoding end and the processing flow of the decoding end respectively.
[0084] In one embodiment, Figure 1 is an optional flowchart of the three-dimensional video processing method provided in an embodiment of the present application. Figure 1 The method in the embodiment is applied to the encoding end, and may include but is not limited to steps 110 to 130. At the same time, it can be understood that this embodiment Figure 1 The order of step 110 to step 130 is not specifically limited, and the order of steps can be adjusted or some steps can be reduced or added according to actual needs.
[0085] Step 110: at multiple acquisition moments, respectively obtain video frames corresponding to the target scene at different viewing angles, use the video frames corresponding to the initial acquisition moment as a key frame sequence, and use the video frames corresponding to other acquisition moments as a non-key frame sequence.
[0086] In one embodiment, the target scene refers to those scenes that need to be three-dimensionally modeled. As far as the target scene is concerned, a 360-degree panoramic surround layout can be used to deploy image acquisition devices, which can achieve synchronous acquisition. This surround layout has the characteristics of all-round coverage of the target scene, which can ensure that information from any angle is not missed, and thus data of the target scene can be collected from all directions. Different image acquisition devices correspond to different viewing angles, which is like observing the target scene from different positions. The pictures seen at different positions are different.
[0087] In one embodiment, at multiple acquisition moments, image acquisition is performed from different viewing angles, so that a captured video can be obtained from each viewing angle. The captured video includes video frames corresponding to the corresponding acquisition moments, and these video frames are the basic units that constitute the captured video. The frames are arranged in sequence according to the time sequence of the acquisition moments, thereby forming a continuous captured video.
[0088] In one embodiment, in view of the fact that, except for the initial video frame, the remaining video frames in the captured video have a temporal reference relationship, which means that there is a motion association between the current video frame and the video frame corresponding to the previous capture moment, in the embodiment of the present application, the initial capture moment is marked as the key frame capture moment, and the other capture moments are marked as non-key frame capture moments. Accordingly, the video frames corresponding to different viewing angles at the initial capture moment are regarded as key frame sequences, and the video frames corresponding to different viewing angles at other capture moments are regarded as non-key frame sequences.
[0089] Step 120: Generate a key frame representation model based on the key frame sequence, perform quantization encoding operation on the key frame representation model to obtain a key frame encoding model, and decode the key frame encoding model to obtain a key frame decoding model.
[0090] In one embodiment, for a key frame sequence, the Colmap calibration method can be used in 3DGS to perform camera calibration on video frames of different viewing angles corresponding to the key frame sequence, and obtain a sparse point cloud representing the target scene, as well as internal and external parameters corresponding to each image acquisition device. The sparse point cloud is then initialized and represented as a key frame representation model composed of multiple three-dimensional Gaussian distributions, each of which can be regarded as a three-dimensional ellipsoid, and multiple three-dimensional ellipsoids are used to fit the sparse point cloud.
[0091] Specifically, in this embodiment, the Colmap calibration method is used to analyze the feature information contained in each video frame. Such feature information includes corner points, edges and other significant visual features in the video frame. These feature points are then accurately identified with the help of a feature extraction algorithm. Matching operations are then performed based on the viewpoint correspondence between these feature points in different video frames. Next, a sparse point cloud for characterizing the target scene is obtained using the feature point information that has been successfully matched. The sparse point cloud is a discretized representation of the target scene in three-dimensional space, consisting of numerous feature points that are relatively sparsely distributed, and they outline key features such as the general outlines of the main objects in the target scene and their spatial position relationships.
[0092] After the sparse point cloud is obtained, it is converted into a form aggregated from multiple different three-dimensional Gaussian distributions. Each three-dimensional Gaussian distribution can be regarded as a three-dimensional ellipsoid. Intuitively speaking, the three-dimensional ellipsoid has parameters such as the center position, the major and minor semi-axes, and can cover a certain spatial area in a relatively flexible manner that conforms to the actual spatial distribution law. Using multiple such three-dimensional ellipsoids to fit the sparse point cloud means that the distribution and shape of these three-dimensional ellipsoids are as close as possible to the spatial form of the actual objects represented by the sparse point cloud. For example, in a target scene containing multiple irregularly shaped objects, by adjusting the center position, covariance and other parameters of each three-dimensional ellipsoid, it can well wrap the sparse point cloud part corresponding to the object, and then accurately simulate the spatial structure of the entire target scene through the combination of multiple three-dimensional ellipsoids.
[0093] In addition, after completing the feature matching, the internal and external parameters corresponding to each image acquisition device are calculated. Among them, the internal parameters reflect the imaging characteristics of the camera itself, such as the focal length of the camera, the position of the principal point, and possible lens distortion parameters; the external parameters reflect the position and posture information of the camera in space, such as the rotation angle and translation vector of the camera relative to the target scene. The internal and external parameters corresponding to the image acquisition device will be used in the subsequent rendering and reconstruction process.
[0094] It is understandable that the key frame representation model generated by 3DGS is only for illustration, and the key frame representation model can also be generated by the neural radiation field NeRF, or directly a point cloud model, a voxel grid model, etc.
[0095] In one embodiment, after obtaining the key frame representation model, it is necessary to enter the encoding stage. Figure 2 , Figure 2 The flowchart of performing quantization encoding operation on the key frame representation model to obtain the key frame encoding model provided by the embodiment of the present application specifically includes the following steps:
[0096] Step 210: Obtain a quantization step length, and perform a quantization operation on a first model parameter of the key frame representation model according to the quantization step length to obtain a quantized representation model.
[0097] In one embodiment, referring to Figure 3 , Figure 3 : is a flowchart of obtaining a quantization step size provided in an embodiment of the present application, which specifically includes the following steps:
[0098] Step 310: Obtain the first center position of multiple three-dimensional Gaussian distributions in the key frame representation model.
[0099] In one embodiment, for the key frame representation model, it includes multiple three-dimensional Gaussian distributions, and each three-dimensional Gaussian distribution has at least a center position, a covariance matrix, a color parameter, and an opacity parameter. The center position can determine the core of the three-dimensional Gaussian distribution in the three-dimensional space, just like a "positioning point", clarifying the approximate coordinate position of the spatial area mainly described by the Gaussian distribution. The covariance matrix reflects the data discreteness of the Gaussian distribution in each dimension in the three-dimensional space and the correlation between different dimensions, which is similar to the size information and rotation information of the Gaussian ellipsoid corresponding to the three-dimensional Gaussian distribution. The color parameter determines the color condition of the area corresponding to the three-dimensional Gaussian distribution in application scenarios such as visualization, which is fitted using spherical harmonic functions. The opacity parameter reflects the transparency of the area. In this embodiment, the center position of all three-dimensional Gaussian distributions in the key frame representation model is obtained, and then the first center position is formed.
[0100] Step 320: Input the first center position into the hash list to obtain the first context feature, and input the first context feature into the quantization prediction model for data processing to obtain the quantization step size.
[0101] In one embodiment, a binary hash table is pre-constructed to store first context features corresponding to different positions. For example, the first context feature can characterize the distribution of objects in the target scene, the spatial relationship with adjacent feature points, and the semantic information that may be involved. By constructing the binary hash table in advance, these rich and diverse context information can be stored in a structured form, which is convenient for subsequent retrieval and use. In addition, the binary hash table can be optimized during the application process. The way of storing information in the hash table, association rules, and related parameters will be continuously adjusted according to the processing frame sequence. In this embodiment, the first center position is input into the binary hash table, and the coordinate values therein are mapped to obtain the corresponding first context feature. Among them, the position information is used as input, just like providing an "index" to the binary hash table to clarify the information to be found.
[0102] Next, the first context feature is input into the quantization prediction model to carry out data processing, thereby obtaining the quantization step size. The quantization prediction model here can use a multi-layer perceptron. After the data is predicted by the quantization prediction model, the quantization step size can be obtained. Based on the obtained quantization step size, the first model parameter of the key frame representation model is quantized, and finally the quantization representation model is obtained.
[0103] In one embodiment, the first model parameters of the key frame representation model include the center position of each three-dimensional Gaussian distribution mentioned above, the covariance matrix, the color parameter, the opacity parameter, and the like. The quantization operation is to discretize the first model parameters that originally take continuous values according to the obtained quantization step length, so as to transform them into data that take values at specific intervals within a specific range. For example, with respect to the color parameter, it can originally take values within a wider continuous color value interval, but through the quantization operation, it is limited to a number of discrete color value points according to the quantization step length. Such an approach helps to reduce the amount of data, simplify the model representation, and to a certain extent retain the key model features. Finally, after such a quantization operation, a quantized representation model can be obtained. Compared with the key frame representation model, the quantized representation model is more compact and concise in terms of data form. The embodiment of the present application does not limit the specific quantization calculation process.
[0104] Step 220: Perform entropy coding on the quantized representation model to obtain a key frame coding model.
[0105] In one embodiment, there are many encoding methods for entropy coding, such as arithmetic coding, Huffman tree coding, etc. The embodiment of the present application does not limit the specific method of entropy coding.
[0106] Taking arithmetic coding as an example, arithmetic coding is an entropy coding technology based on probability statistics, which treats the data of the entire quantization representation model as a whole symbol sequence for processing. When arithmetic coding is performed on the quantization representation model, the probability distribution corresponding to each parameter and its value in the quantization representation model is analyzed first. For example, for different color parameter values in the quantization representation model, their frequencies of occurrence in the entire model are counted to determine their corresponding coding intervals. Then, based on these probability information, all data in the quantization representation model are gradually mapped to the corresponding coding intervals, and by continuously subdividing and determining the coding intervals, a unique corresponding key frame coding model is finally generated. Among them, the key frame coding model further reduces the amount of data compared to the quantization representation model, occupies less storage space, and presents the information in the original quantization representation model in a more compact and efficient coding form.
[0107] In one embodiment, after the key frame coding model is obtained, the encoding end needs to perform related decoding operations, and the key frame coding model is entropy decoded according to the corresponding entropy decoding method of entropy coding to obtain a quantized decoding model. Similarly, according to the inverse quantization process of the quantization method, the parameters of the quantized decoding model are inversely quantized according to the quantization step size to obtain the key frame decoding model. After the key frame decoding model is obtained, it is stored in a cache for subsequent use.
[0108] Step 130: Select non-key frame sequences one by one as processing frame sequences, obtain a reference frame sequence according to the previous acquisition moment of the processing frame sequence, generate transformation parameters according to the decoding model corresponding to the reference frame sequence, perform quantization encoding operation on the transformation parameters to obtain coded transformation parameters, and send the coded transformation parameters to the decoding end. If there is a residual flag, generate a residual representation model corresponding to the processing frame sequence based on the transformation parameters, perform quantization encoding operation on the residual representation model to obtain a residual coding model, and send the residual coding model to the decoding end. The decoding model corresponding to the first processing frame sequence is the key frame decoding model.
[0109] In one embodiment, after the key frame sequence is processed, the non-key frame sequence is processed. Since the non-key frame sequence contains time domain information, the non-key frame sequences are selected one by one as the processing frame sequence, and the frame sequence corresponding to the previous acquisition moment of the processing frame sequence is used as the reference frame sequence, so as to more effectively utilize the time correlation between the non-key frame sequences and between them and the key frame sequence. Specifically, if the first non-key frame sequence is selected, then its corresponding reference frame sequence is the key frame sequence; and for other non-key frame sequences, their corresponding reference frame sequence is the non-key frame sequence corresponding to the previous acquisition moment.
[0110] It is understandable that the decoding model corresponding to the reference frame sequence can be sent back to the encoding end for storage after being decoded by the decoding end, or can be obtained by decoding by the encoding end according to the same decoding method as the decoding end, and this embodiment does not limit this. The following takes the first non-key frame sequence as an example for processing the frame sequence.
[0111] In one embodiment, a reference frame sequence corresponding to a processing frame sequence is first obtained, and then a decoding model corresponding to the reference frame sequence is obtained from a decoding end. For the first processing frame sequence, its decoding model is a key frame decoding model corresponding to a key frame encoding model. After obtaining the decoding model, data processing work is performed on the processing frame sequence according to the corresponding decoding model.
[0112] In one embodiment, referring to Figure 4 , Figure 4 This is a flowchart of generating transformation parameters according to a decoding model provided by an embodiment of the present application, which specifically includes the following steps:
[0113] Step 410: Obtain second center positions of multiple three-dimensional Gaussian distributions in the decoding model.
[0114] In one embodiment, the decoding model also includes multiple three-dimensional Gaussian distributions, and these three-dimensional Gaussian distributions at least have a center position, a covariance matrix, a color parameter, and an opacity parameter. The center position of the three-dimensional Gaussian distribution in the decoding model is obtained to obtain a second center position.
[0115] Step 420: Input the second center position into the hash list to obtain a second context feature, and input the second context feature into the motion transformation prediction model for data processing to obtain transformation parameters.
[0116] In one embodiment, referring to step 320, the second center position is first input into the hash list to obtain the second context feature. Then, the second context feature is input into the motion transformation prediction model to carry out data processing to obtain the transformation parameter. The motion transformation prediction model can also be constructed by a multi-layer perceptron.
[0117] In view of the fact that the parameters of the three-dimensional Gaussian distribution often change at adjacent acquisition moments due to the movement of objects, changes in scene lighting, or other dynamic factors, in an embodiment of the present application, a motion transformation prediction model is used to analyze the input second context features, so as to accurately capture the changes of these parameters over time, and express the changes as transformation parameters. Therefore, the transformation parameters can indicate the change in the predicted parameters of the second model parameters corresponding to the three-dimensional Gaussian distribution in the decoding model at the acquisition moment corresponding to the processing frame sequence. For example, for the center position of a three-dimensional Gaussian distribution, the transformation parameters will clearly indicate how much its coordinate position in the three-dimensional space has changed from the reference frame sequence to the acquisition moment corresponding to the processing frame sequence.
[0118] Next, in one embodiment, refer to Figure 5 , Figure 5 This is a flowchart of a residual representation model corresponding to a frame sequence generated based on transformation parameters provided by an embodiment of the present application, which specifically includes the following steps:
[0119] Step 510: Perform inter-frame transformation on the decoding model according to the transformation parameters to obtain a reference decoding model.
[0120] In one embodiment, referring to Figure 6 , Figure 6 The flowchart of performing inter-frame transformation on the decoding model according to the transformation parameters to obtain the reference decoding model provided by the embodiment of the present application specifically includes the following steps:
[0121] Step 610: Obtain at least the center position parameter, covariance parameter, color parameter and opacity parameter corresponding to each three-dimensional Gaussian distribution of the decoding model.
[0122] Step 620: Obtain, based on the transformation parameters, center position update parameters corresponding to the center position parameters, covariance update parameters corresponding to the covariance parameters, color update parameters corresponding to the color parameters, and opacity update parameters corresponding to the opacity parameters at the acquisition moment corresponding to the processing frame sequence.
[0123] In one embodiment, according to the change in the predicted parameters at the acquisition moment corresponding to the processing frame sequence indicated by the transformation parameter, the corresponding change is added to the center position parameter to obtain the center position update parameter, the corresponding change is added to the covariance parameter to obtain the covariance update parameter, the corresponding change is added to the color parameter to obtain the color update parameter, and the corresponding change is added to the opacity parameter to obtain the opacity update parameter.
[0124] Step 630: Update the three-dimensional Gaussian distribution based on the center position update parameter, the covariance update parameter, the color update parameter, and the opacity update parameter to obtain a reference decoding model.
[0125] In one embodiment, the three-dimensional Gaussian distribution is updated according to the center position update parameter, the covariance update parameter, the color update parameter and the opacity update parameter to obtain a reference decoding model, and then the reference decoding model is sent to the decoding end. It can be understood that the reference decoding model is a plurality of three-dimensional Gaussian distributions corresponding to the predicted processing frame sequence.
[0126] In one embodiment, rendering is performed according to a reference decoding model to obtain a reference rendered image, and a distortion comparison is performed between the reference rendered image and the corresponding video frame. If the degree of distortion is within an acceptable threshold, there is no need to generate a residual flag. Otherwise, a residual flag is generated for the encoding end, so that the encoding end can update the change parameters according to the residual flag, and then update the reference decoding model, so that the rendering effect meets the actual needs.
[0127] Step 520: Generate a residual identifier according to the reference decoding model.
[0128] In one embodiment, the specific process of generating a residual identifier based on a reference decoding model includes: rendering the reference decoding model to obtain a reference rendered image corresponding to a processing frame sequence, generating a residual identifier based on distortion parameters between the reference rendered image and a video frame corresponding to the processing frame sequence, and when the residual identifier indicates updating transformation parameters, sending the residual identifier to the encoding end.
[0129] In one embodiment, first, a rendering operation is performed on the reference decoding model to obtain a reference rendered image corresponding to the processed frame sequence. Taking 3DGS as an example to introduce the rendering process, at each viewing angle, the three-dimensional Gaussian distribution in the reference decoding model is first projected into a two-dimensional space to obtain a two-dimensional Gaussian distribution, and then rendering is performed based on the two-dimensional Gaussian distribution to obtain a reference rendered image corresponding to each viewing angle.
[0130] Subsequently, a residual identifier is generated based on the distortion parameters between the reference rendered image and the video frame corresponding to the processed frame sequence. Specifically, the reference rendered image of different viewing angles is compared with the corresponding video frame in terms of pixel difference, and the area where the pixel distortion is greater than the preset threshold is determined as the residual representation area. Finally, all the residual representation areas are summarized to generate a residual identifier for indicating the update of the transformation parameters, and the residual identifier is sent to the encoder. If the pixel distortion is less than the preset threshold, no residual identifier will be generated.
[0131] Step 530: Generate a residual representation model based on the residual identifier.
[0132] In one embodiment, if there is a residual identifier, it indicates that a residual representation model needs to be generated. Figure 7 , Figure 7 This is a flowchart of generating a residual representation model provided by an embodiment of the present application, which specifically includes the following steps:
[0133] Step 710: Determine a residual representation region of the processed frame sequence based on the residual identifier.
[0134] In one embodiment, the decoding end can perform a region comparison between the reference rendered image and the corresponding video frame, and analyze the differences between the two in terms of image presentation region by region. For example, the color deviation, the clarity of the object outline, and the consistency of the layout of the elements in the scene are compared, so as to accurately determine the areas with poor rendering effect.
[0135] Then, the area with poor rendering effect is set as the residual representation area. The residual representation area indicates that under the current encoding and rendering mechanism, the picture in this area has a significant gap with the actual expected presentation effect.
[0136] Step 720: Generate at least one three-dimensional Gaussian distribution for the residual representation area to form a residual representation model.
[0137] In one embodiment, the encoding end needs to add some three-dimensional Gaussian distributions to the residual representation area or segment the three-dimensional Gaussian distributions.
[0138] For example, for an area containing objects with complex textures, the original three-dimensional Gaussian distribution may not be able to accurately present its texture details, while adding a new three-dimensional Gaussian distribution can better capture these details. Alternatively, although some larger three-dimensional Gaussian distributions can cover the corresponding areas, due to their relatively "extensive" coverage method, it is difficult to fit the complex object shapes, spatial layouts, etc. in the area. Therefore, after dividing it into several small three-dimensional Gaussian distributions, each small three-dimensional Gaussian distribution can describe the local spatial characteristics more specifically and better fit the actual form of the target scene in the area, thereby achieving a more accurate and delicate fitting effect for the entire target scene, further improving the quality of subsequent rendering.
[0139] At this time, a residual representation model generated based on the residual representation area is obtained, and the residual representation model is also composed of multiple three-dimensional Gaussian distributions.
[0140] It is understandable that the above-mentioned update process of transformation parameters can be performed multiple times until no residual mark is received. For example, judging by the gradient, if there is always a relatively large gradient in the area, or the gradient exceeds a certain threshold, a new three-dimensional Gaussian distribution can be generated by random initialization or splitting from the existing Gaussians in the surrounding area. The random initialization here can also be completed by re-labeling the processed frame sequence based on COLMAP. Use COLMAP to obtain a sparse point cloud, and the position of the point cloud is used as the position of the Gaussian point, while other parameters are randomly initialized.
[0141] In addition, in order to further improve the reconstruction accuracy, the above-mentioned quantization encoding process can be performed only on one or more of the transformation parameters and the residual representation model. Similarly, the entropy decoding and inverse quantization processes at the decoding end are set accordingly.
[0142] In addition, the embodiment of the present application also performs a rate-distortion joint optimization process on both key frame sequences and non-key frame sequences, and jointly optimizes the transformation parameters of the encoding end, the update process of the residual representation model, and the decoding model of the decoding end, which can prevent the accumulation of time domain errors, improve prediction performance, and thereby improve video coding efficiency.
[0143] In one embodiment, referring to Figure 8 , Figure 8 : is a flow chart of the rate-distortion joint optimization process provided by an embodiment of the present application, which specifically includes the following steps:
[0144] Step 810: Obtain a rendering distortion loss item corresponding to a key frame decoding model or a non-key frame decoding model.
[0145] In one embodiment, the decoding model includes a key frame decoding model or a non-key frame decoding model. At this time, for the decoding model, the structural loss value of its rendered image and video frame is calculated using calculation methods such as D-SSIM, PSNR, LPIPS, etc. as the rendering distortion loss item D.
[0146] Step 820: Obtain a bit rate loss item.
[0147] In one embodiment, the rate loss term is calculated by the decoding model or the residual representation model during the quantization coding process. For example, the rate required for coding according to the quantized parameters is R2, and the rate loss term is R2, where the rate can be calculated based on information entropy.
[0148] Step 830: Calculate a rate-distortion joint loss term according to the rendering distortion loss term and the bit rate loss term.
[0149] In one embodiment, the rate-distortion joint loss term is used to optimize at least the decoding model, the residual representation model and the transformation parameters, where the rate-distortion joint loss term L is expressed as:
[0150] L=D+λR
[0151] Among them, λ represents the balancing factor, which is determined according to actual needs.
[0152] Through the above process, for the key frame sequence, the encoder generates a key frame coding model for the key frame sequence and sends it to the decoder. For the non-key frame sequence, if the residual mark is not received, it means that the rendering effect according to the current coding transformation parameters is satisfactory, and it is only necessary to perform quantization coding operations on the transformation parameters according to the generation method of the key frame coding model to obtain the coding transformation parameters, and send the coding transformation parameters to the decoder to participate in the subsequent rendering process. If the residual mark is received, a residual representation model is generated, and the residual representation model is quantized and coded according to the same method to obtain a residual coding model, and the coding transformation parameters are updated at the same time, and the last updated coding transformation parameters and residual coding model are sent to the decoder.
[0153] From the above process, we can see that Fig. 9 , Fig. 9 It is a schematic diagram of the overall flow of three-dimensional video processing at the encoding end provided in an embodiment of the present application.
[0154] Firstly, the multi-perspective videos are divided into key frame sequences and non-key frame sequences according to the acquisition time. The purpose is to eliminate the time domain redundancy in non-key frames, thereby significantly improving the compression efficiency of immersive holographic videos.
[0155] For the key frame sequence at the initial acquisition moment, the key frame rate distortion optimization is performed, and then a compact key frame representation model is obtained at the encoding end. Next, the key frame representation model is quantized and entropy encoded to obtain a key frame encoding model, which is stored or transmitted in the form of a binary bit stream. At the decoding end, the key frame encoding model needs to be entropy decoded and dequantized to obtain a key frame decoding model, which is then stored in the decoding buffer for interaction with the encoding end and rendering at the decoding end.
[0156] Next, for the non-key frame sequence at the non-initial acquisition moment in the multi-view video, the decoding model of the previous acquisition moment needs to be extracted from the decoding buffer at the encoding end, and it is used as a time domain reference, and then the non-key frame rate distortion optimization is performed on the non-key frame sequence at the current acquisition moment, thereby obtaining the residual representation model and transformation parameters. After that, the residual representation model and transformation parameters are quantized and entropy coded respectively, and then the residual coding model and coding transformation parameters corresponding to the non-key frame sequence are obtained.
[0157] Then, at the decoding end, the residual coding model and the coding transformation parameters are entropy decoded and dequantized in turn to obtain the decoding transformation parameters and the residual decoding model. On this basis, the inter-frame transformation operation is performed on the decoding model of the time domain reference based on the decoding transformation parameters to obtain the reference decoding model. In addition, the residual fusion operation is performed on the reference decoding model and the residual decoding model to finally obtain the non-key frame decoding model, which is then stored in the decoding buffer to serve as the decoding model of the reference frame sequence at the next acquisition moment. This cycle is repeated until all acquisition moments are processed.
[0158] The processing flow of the decoding end is described in detail below.
[0159] In one embodiment, Fig.10 is an optional flowchart of the three-dimensional video processing method provided in an embodiment of the present application. Fig.10 The method in the embodiment is applied to the encoding end, and may include but is not limited to steps 1010 to 1020. At the same time, it can be understood that this embodiment Fig.10 The order of step 1010 to step 1020 is not specifically limited, and the order of steps can be adjusted or some steps can be reduced or added according to actual needs.
[0160] Step 1010: Decode the acquired key frame encoding model to obtain a key frame decoding model.
[0161] In one embodiment, referring to Fig.11 , Fig.11The flowchart of decoding the acquired key frame encoding model to obtain the key frame decoding model provided by the embodiment of the present application specifically includes the following steps:
[0162] Step 1110: Perform entropy decoding on the key frame coding model to obtain a quantized decoding model.
[0163] In one embodiment, the key frame coding model is entropy decoded in a corresponding entropy decoding manner of the entropy coding at the coding end to obtain a quantized decoding model, which will not be described in detail here.
[0164] Step 1120: Obtain a quantization step size, and perform a dequantization operation on the parameters of the quantization decoding model according to the quantization step size to obtain a key frame decoding model.
[0165] In one embodiment, similarly, according to the inverse quantization process of the encoding end quantization method, the parameters of the quantized decoding model are inversely quantized according to the quantization step size to obtain the key frame decoding model.
[0166] Step 1020: For each processed frame sequence, obtain the corresponding encoding transformation parameters, obtain the decoding transformation parameters according to the encoding transformation parameters, perform inter-frame transformation using the decoding transformation parameters and the decoding model corresponding to the reference frame sequence to obtain a reference decoding model, and use the reference decoding model as a non-key frame decoding model. If the processed frame sequence includes a residual coding model, decode the residual coding model to obtain a residual decoding model, perform residual fusion on the residual decoding model and the reference decoding model to obtain a non-key frame decoding model.
[0167] In one embodiment, the definition of the processing frame sequence is consistent with that of the encoding end, so each processing frame sequence includes a corresponding reference frame sequence of the previous acquisition moment. Since not all processing frame sequences need to update the transformation parameters, not every processing frame sequence includes a residual coding model.
[0168] For a processing frame sequence that only contains coding transformation parameters, the obtained coding transformation parameters are processed according to the same entropy decoding and inverse quantization process as the above-mentioned key frame coding model to obtain decoding transformation parameters. The decoding transformation parameters can be used to indicate the parameter changes between the current processing frame sequence and the reference frame sequence. Subsequently, the decoding transformation parameters and the decoding model corresponding to the reference frame sequence are used to perform the above-mentioned similar inter-frame transformation operation, thereby obtaining a reference decoding model, and the reference decoding model is used as the non-key frame decoding model of this processing frame sequence. It should be noted that the decoding model of the first processing frame sequence is a key frame decoding model.
[0169] For the processing frame sequence including the coding transformation parameters and the residual coding model, the obtained coding transformation parameters and the residual coding model are processed respectively according to the same entropy decoding and inverse quantization process as the key frame coding model, thereby obtaining the decoding transformation parameters and the residual decoding model. Next, the residual fusion operation is performed on the residual decoding model and the reference decoding model, and the corresponding three-dimensional Gaussian distribution is added to the residual representation area by using the residual decoding model, or the existing three-dimensional Gaussian distribution is updated, and finally the corresponding non-key frame decoding model is obtained.
[0170] It can be seen from the above process that a corresponding decoding model is generated at each acquisition moment, including a key frame decoding model and at least one non-key frame decoding model. In this process, the residual decoding model usually increases the number of three-dimensional Gaussian distributions. At this time, if the number of three-dimensional Gaussian distributions in a decoding model is too large, it can be pruned by the opacity threshold judgment method, which is not limited in this embodiment.
[0171] After the corresponding decoding model is generated at each acquisition moment, three-dimensional reconstruction can be performed. Specifically, multiple three-dimensional Gaussian distributions contained in the key frame decoding model and the non-key frame decoding model are projected two-dimensionally at each viewing angle to obtain the two-dimensional Gaussian distribution of the corresponding viewing angle, and the two-dimensional Gaussian distribution is rendered one by one at the corresponding viewing angle to obtain the target rendering image corresponding to all viewing angles at each acquisition moment of the target scene, and finally the three-dimensional scene corresponding to the target scene is generated according to the target rendering image.
[0172] The technical solution provided in the embodiment of the present application can be applied to immersive holographic video. By acquiring the video frames corresponding to the target scene at different viewing angles at multiple acquisition moments, the video frames corresponding to the initial acquisition moment are used as key frame sequences, and the video frames corresponding to other acquisition moments are used as non-key frame sequences, and then a key frame representation model is generated based on the key frame sequence, and a quantization encoding operation is performed on the key frame representation model to obtain a key frame encoding model, and the key frame encoding model is decoded to obtain a key frame decoding model. Next, non-key frame sequences are selected one by one as processing frame sequences, and a reference frame sequence is obtained according to the previous acquisition moment of the processing frame sequence. Transformation parameters are generated according to the decoding model corresponding to the reference frame sequence, and the transformation parameters are quantized and encoded. After the encoding transformation parameters are obtained, they are sent to the decoding end. If there is a residual flag, a residual representation model corresponding to the processing frame sequence is generated based on the transformation parameters, and a quantization encoding operation is performed on the residual representation model to obtain a residual encoding model, and the residual encoding model is sent to the decoding end. The embodiment of the present application uses the decoding model of a previous acquisition moment as a reference, generates transformation parameters according to the decoding model, and uses these transformation parameters to indicate the data change trend and law of the previous and subsequent acquisition moments. Subsequently, the residual representation model corresponding to the acquisition moment is generated according to the transformation parameters, and the residual representation model is used to fit the actual dynamic change trend of the video frame, thereby offsetting the time domain error to the greatest extent, thereby avoiding the continuous accumulation of errors over the time series, thereby ensuring the accuracy and reliability of the entire video data processing process and improving the performance of subsequent three-dimensional video reconstruction.
[0173] The embodiment of the present application also provides a three-dimensional video processing device, which can implement the three-dimensional video processing method applied to the encoding end. Fig.12 , the device comprises:
[0174] Frame division module 1210: used to obtain video frames corresponding to the target scene at different viewing angles at multiple acquisition moments, and use the video frames corresponding to the initial acquisition moment as a key frame sequence and the video frames corresponding to other acquisition moments as a non-key frame sequence.
[0175] The key frame encoding module 1220 is used to generate a key frame representation model based on a key frame sequence, perform quantization encoding operations on the key frame representation model to obtain a key frame encoding model, and decode the key frame encoding model to obtain a key frame decoding model.
[0176] Non-key frame encoding module 1230: used to select non-key frame sequences one by one as processing frame sequences, obtain a reference frame sequence according to the previous acquisition moment of the processing frame sequence, generate transformation parameters according to the decoding model corresponding to the reference frame sequence, perform quantization encoding operation on the transformation parameters to obtain the coded transformation parameters, and send the coded transformation parameters to the decoding end. If there is a residual flag, a residual representation model corresponding to the processing frame sequence is generated based on the transformation parameters, and a quantization encoding operation is performed on the residual representation model to obtain a residual coding model. The residual coding model is sent to the decoding end. The decoding model corresponding to the first processing frame sequence is the key frame decoding model.
[0177] The specific implementation of the 3D video processing device of this embodiment is basically the same as the specific implementation of the 3D video processing method described above, and will not be described in detail herein.
[0178] The present application also provides an electronic device, including:
[0179] at least one memory;
[0180] at least one processor;
[0181] at least one program;
[0182] The program is stored in the memory, and the processor executes the at least one program to implement the above-mentioned three-dimensional video processing method of the present application. The electronic device can be any intelligent terminal including a mobile phone, a tablet computer, a personal digital assistant (PDA), a car computer, etc.
[0183] See also Fig.13 , Fig.13 The hardware structure of an electronic device of another embodiment is illustrated, and the electronic device includes:
[0184] The processor 1301 may be implemented by a general-purpose central processing unit (CPU), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of the present application;
[0185] The memory 1302 may be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 1302 may store an operating system and other application programs. When the technical solution provided in the embodiment of this specification is implemented by software or firmware, the relevant program code is stored in the memory 1302, and the processor 1301 is used to call and execute the three-dimensional video processing method of the embodiment of this application;
[0186] Input / output interface 1303, used to implement information input and output;
[0187] Communication interface 1304, used to realize communication interaction between the device and other devices, which can be realized by wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.); and
[0188] A bus 1305 that transmits information between the various components of the device (e.g., the processor 1301, the memory 1302, the input / output interface 1303, and the communication interface 1304);
[0189] The processor 1301 , the memory 1302 , the input / output interface 1303 and the communication interface 1304 are connected to each other in communication within the device via a bus 1305 .
[0190] The embodiment of the present application further provides a storage medium, which is a storage medium storing a computer program. When the computer program is executed by a processor, the above-mentioned three-dimensional video processing method is implemented.
[0191] As a non-transitory storage medium, the memory can be used to store non-transitory software programs and non-transitory computer executable programs. In addition, the memory may include a high-speed random access memory, and may also include a non-transitory memory, such as at least one disk storage device, a flash memory device, or other non-transitory solid-state storage device. In some embodiments, the memory may optionally include a memory remotely arranged relative to the processor, and these remote memories may be connected to the processor via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0192] The three-dimensional video processing method, device, equipment and storage medium proposed in the embodiment of the present application are to obtain the video frames corresponding to the target scene at different viewing angles at multiple acquisition moments, respectively, take the video frames corresponding to the initial acquisition moment as the key frame sequence, and take the video frames corresponding to other acquisition moments as the non-key frame sequence, then generate the key frame representation model based on the key frame sequence, perform quantization coding operation on the key frame representation model, obtain the key frame coding model, decode the key frame coding model, and obtain the key frame decoding model, then select the non-key frame sequence one by one as the processing frame sequence, obtain the reference frame sequence according to the previous acquisition moment of the processing frame sequence, generate the transformation parameters according to the decoding model corresponding to the reference frame sequence, perform quantization coding operation on the transformation parameters, obtain the coding transformation parameters and send them to the decoding end, if there is a residual mark, generate the residual representation model corresponding to the processing frame sequence based on the transformation parameters, perform quantization coding operation on the residual representation model, obtain the residual coding model, and send the residual coding model to the decoding end. The embodiment of the present application uses the decoding model of a previous acquisition moment as a reference, generates the transformation parameters according to the decoding model, and uses these transformation parameters to indicate the data change trend and law of the previous and next acquisition moments. Subsequently, the residual representation model corresponding to the acquisition moment is generated according to the transformation parameters, and the residual representation model is used to fit the actual dynamic change trend of the video frame, thereby offsetting the time domain error to the greatest extent, thereby avoiding the continuous accumulation of errors over the time series, thereby ensuring the accuracy and reliability of the entire video data processing process and improving the performance of subsequent three-dimensional video reconstruction.
[0193] The embodiments described in the embodiments of the present application are intended to more clearly illustrate the technical solutions of the embodiments of the present application and do not constitute a limitation on the technical solutions provided in the embodiments of the present application. Those skilled in the art will appreciate that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided in the embodiments of the present application are also applicable to similar technical problems.
[0194] Those skilled in the art will appreciate that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present application, and may include more or fewer steps than shown in the figures, or a combination of certain steps, or different steps.
[0195] The device embodiments described above are merely illustrative, and the units described as separate components may or may not be physically separated, that is, they may be located in one place or distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0196] Those skilled in the art will appreciate that all or some of the steps in the methods disclosed above, and the functional modules / units in the systems and devices may be implemented as software, firmware, hardware, or a suitable combination thereof.
[0197] The terms "first", "second", "third", "fourth", etc. (if any) in the specification of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0198] It should be understood that in the present application, "at least one (item)" means one or more, and "plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships may exist. For example, "A and / or B" can mean: only A exists, only B exists, and A and B exist at the same time, where A and B can be singular or plural. The character " / " generally indicates that the objects associated before and after are in an "or" relationship. "At least one of the following" or similar expressions refers to any combination of these items, including any combination of single or plural items. For example, at least one of a, b or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.
[0199] In the several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the above units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0200] The units described above as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0201] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.
[0202] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including multiple instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of various embodiments of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), disk or optical disk and other media that can store programs.
[0203] The preferred embodiments of the present application are described above with reference to the accompanying drawings, but the scope of the rights of the present application is not limited thereto. Any modification, equivalent substitution and improvement made by a person skilled in the art without departing from the scope and essence of the present application should be within the scope of the rights of the present application.
Claims
1. A three-dimensional video processing method, characterized in that: Applied at the encoding end, the method includes: At multiple acquisition moments, video frames corresponding to the target scene at different viewing angles are obtained respectively, and the video frames corresponding to the initial acquisition moment are used as key frame sequences, and the video frames corresponding to other acquisition moments are used as non-key frame sequences; Generate a key frame representation model based on the key frame sequence, perform a quantization encoding operation on the key frame representation model to obtain a key frame encoding model, and decode the key frame encoding model to obtain a key frame decoding model; The non-key frame sequences are selected one by one as processing frame sequences, a reference frame sequence is obtained according to a previous acquisition moment of the processing frame sequence, transformation parameters are generated according to a decoding model corresponding to the reference frame sequence, a quantization encoding operation is performed on the transformation parameters to obtain coded transformation parameters, and the coded transformation parameters are sent to a decoding end. If a residual flag exists, a residual representation model corresponding to the processing frame sequence is generated based on the transformation parameters, a quantization encoding operation is performed on the residual representation model to obtain a residual coding model, and the residual coding model is sent to a decoding end. The decoding model corresponding to the first processing frame sequence is the key frame decoding model.
2. The three-dimensional video processing method according to claim 1, characterized in that: The step of performing a quantization encoding operation on the key frame representation model to obtain a key frame encoding model includes: Acquire a quantization step length, and perform a quantization operation on a first model parameter of the key frame representation model according to the quantization step length to obtain a quantized representation model; The quantized representation model is entropy encoded to obtain a key frame encoding model.
3. The three-dimensional video processing method according to claim 2, characterized in that: The obtaining of the quantization step size comprises: Obtaining first center positions of multiple three-dimensional Gaussian distributions in the key frame representation model; The first center position is input into a hash list to obtain a first context feature, and the first context feature is input into a quantization prediction model for data processing to obtain the quantization step size.
4. The three-dimensional video processing method according to claim 1, characterized in that: The generating of transformation parameters according to the decoding model comprises: Obtaining second center positions of multiple three-dimensional Gaussian distributions in the decoding model; The second center position is input into a hash list to obtain a second context feature, and the second context feature is input into a motion transformation prediction model for data processing to obtain the transformation parameters, wherein the transformation parameters include the predicted parameter change of the second model parameter corresponding to the three-dimensional Gaussian distribution in the decoding model at the acquisition moment corresponding to the processing frame sequence.
5. The three-dimensional video processing method according to claim 1, characterized in that: The generating the residual representation model corresponding to the processed frame sequence based on the transformation parameters comprises: Performing inter-frame transformation on the decoding model according to the transformation parameters to obtain a reference decoding model; generating the residual identifier according to the reference decoding model; The residual representation model is generated based on the residual identifier.
6. The three-dimensional video processing method according to claim 5, characterized in that: The inter-frame transformation of the decoding model according to the transformation parameters to obtain a reference decoding model includes: At least obtaining a center position parameter, a covariance parameter, a color parameter, and an opacity parameter corresponding to each three-dimensional Gaussian distribution of the decoding model; According to the transformation parameters, at the acquisition time corresponding to the processing frame sequence, a center position update parameter corresponding to the center position parameter, a covariance update parameter corresponding to the covariance parameter, a color update parameter corresponding to the color parameter, and an opacity update parameter corresponding to the opacity parameter are obtained; The three-dimensional Gaussian distribution is updated based on the center position update parameter, the covariance update parameter, the color update parameter, and the opacity update parameter to obtain the reference decoding model.
7. The three-dimensional video processing method according to claim 5, characterized in that: The generating the residual identifier according to the reference decoding model includes: Rendering the reference decoding model to obtain a reference rendered image corresponding to the processed frame sequence; The residual identifier is generated according to a distortion parameter between the reference rendered image and a video frame corresponding to the processed frame sequence.
8. The three-dimensional video processing method according to claim 5, characterized in that: The generating the residual representation model comprises: Determining a residual representation region of the processed frame sequence based on the residual identifier; At least one three-dimensional Gaussian distribution is generated for the residual representation area to form the residual representation model.
9. The three-dimensional video processing method according to claim 1, characterized in that: The method further comprises: Obtaining a rendering distortion loss item corresponding to the key frame decoding model or the non-key frame decoding model; Obtaining a bit rate loss term, where the bit rate loss term is calculated by the decoding model or the residual representation model during a quantization encoding process; A rate-distortion joint loss term is calculated according to the rendering distortion loss term and the bit rate loss term, and the rate-distortion joint loss term is used to optimize at least the decoding model, the residual representation model and the transformation parameters.
10. A three-dimensional video processing method, characterized in that: Applied at the decoding end, the method includes: Decoding the acquired key frame encoding model to obtain a key frame decoding model; For each processed frame sequence, the corresponding encoding transformation parameters are obtained, and the decoding transformation parameters are obtained according to the encoding transformation parameters. The decoding transformation parameters and the decoding model corresponding to the reference frame sequence are used to perform inter-frame transformation to obtain a reference decoding model, and the reference decoding model is used as a non-key frame decoding model. If the processed frame sequence includes a residual coding model, the residual coding model is decoded to obtain a residual decoding model, and the residual decoding model and the reference decoding model are residually fused to obtain the non-key frame decoding model. The decoding model of the first processed frame sequence is the key frame decoding model.
11. The three-dimensional video processing method according to claim 10, characterized in that: The step of decoding the acquired key frame encoding model to obtain a key frame decoding model includes: Performing entropy decoding on the key frame coding model to obtain a quantized decoding model; A quantization step size is obtained, and a dequantization operation is performed on the parameters of the quantization decoding model according to the quantization step size to obtain the key frame decoding model.
12. The three-dimensional video processing method according to claim 10, characterized in that: The method further comprises: Performing two-dimensional projection on a plurality of three-dimensional Gaussian distributions included in the key frame decoding model and the non-key frame decoding model at each viewing angle to obtain a two-dimensional Gaussian distribution at a corresponding viewing angle; Performing image rendering of corresponding viewing angles on the two-dimensional Gaussian distribution one by one to obtain target rendered images corresponding to all viewing angles at each acquisition moment of the target scene; A three-dimensional scene corresponding to the target scene is generated according to the target rendered image.
13. A three-dimensional video processing device, characterized in that: Applied to the encoding end, the device includes: Frame division module: used to obtain the corresponding video frames of the target scene at different viewing angles at multiple acquisition moments, and use the video frames corresponding to the initial acquisition moment as the key frame sequence, and the video frames corresponding to other acquisition moments as the non-key frame sequence; A key frame encoding module: used for generating a key frame representation model based on the key frame sequence, performing a quantization encoding operation on the key frame representation model to obtain a key frame encoding model, and decoding the key frame encoding model to obtain a key frame decoding model; Non-key frame encoding module: used to select the non-key frame sequence as the processing frame sequence one by one, obtain the reference frame sequence according to the previous acquisition moment of the processing frame sequence, generate transformation parameters according to the decoding model corresponding to the reference frame sequence, perform quantization encoding operation on the transformation parameters to obtain the encoding transformation parameters, send the encoding transformation parameters to the decoding end, if there is a residual flag, generate the residual representation model corresponding to the processing frame sequence based on the transformation parameters, perform quantization encoding operation on the residual representation model to obtain the residual encoding model, send the residual encoding model to the decoding end, the decoding model corresponding to the first processing frame sequence is the key frame decoding model.
14. An electronic device, characterized in that: The electronic device comprises a memory and a processor, the memory stores a computer program, and the processor implements the three-dimensional video processing method according to any one of claims 1 to 12 when executing the computer program.
15. A storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the three-dimensional video processing method according to any one of claims 1 to 12 is implemented.
Citation Information
Patent Citations
Multi-reference inter-frame prediction method and system, equipment and storage medium
CN113938687A
Video coding and decoding processing method and device, computer equipment and storage medium
CN116233445A
Point cloud coding processing method, point cloud decoding processing method and related equipment
CN118827998A
Three-dimensional volume video coding compression method
CN119052510A
Residual encoding method and apparatus, video encoding method and device, and system
WO2022261838A1