Holographic video coding method and device, computer equipment and storage medium

By encoding the relevant parameters in the keyframes and non-keyframes of holographic videos, a holographic video file in the form of binary code streams is solved, and the multi-platform and multi-device compatibility of holographic videos is achieved.

CN120091120APending Publication Date: 2025-06-03PENG CHENG LAB
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510016478.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-06
Publication Date
2025-06-03

AI Technical Summary

Technical Problem

Holographic videos generated based on three-dimensional Gaussian splattering technology have compatibility issues on different devices and platforms, which lead to unplayable.

Method used

By encoding neural network parameters, position parameters, attribute parameters and hash features in keyframes and non-keyframes of holographic videos, a holographic video file in the form of binary code streams is generated to ensure the compatibility of the file on multiple devices and platforms.

Benefits of technology

The multi-platform and multi-device compatibility of holographic video files is realized, ensuring that holographic videos can be played smoothly on different devices and platforms.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120091120A_ABST
    Figure CN120091120A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a holographic video coding method and device, computer equipment and a storage medium. Determining a current frame to be coded by determining a video frame coding sequence corresponding to the holographic video; when the current frame is a key frame, obtaining a neural network parameter, a position parameter, an attribute parameter and a hash feature corresponding to a target anchor point in the key frame, and respectively encoding to obtain key frame binary code stream data corresponding to the key frame; when the current frame is a non-key frame, obtaining a neural network parameter and a cache hash feature corresponding to a target anchor point in the non-key frame, and encoding the neural network parameter and the cache hash feature respectively to obtain non-key frame binary code stream data corresponding to the non-key frame; key frame binary code stream data of each key frame and non-key frame binary code stream data of each non-key frame data in the holographic video are determined, and a holographic video file corresponding to the holographic video is generated according to the key frame binary code stream data of each key frame and the non-key frame binary code stream data of each non-key frame data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and particularly to a holographic video encoding method, apparatus, computer device, and storage medium. Background Art

[0002] With the development of computer technology, videos currently evolve from two-dimensional to three-dimensional holographic videos. 3D Gaussian Splatting (3DGS) has surpassed technologies such as point clouds, meshes, and Neural Radiance Fields (NeRF) in aspects such as three-dimensional scene reconstruction quality, reconstruction speed, interaction freedom, and rendering speed, becoming the preferred technology for constructing three-dimensional holographic videos. For the holographic videos generated based on the 3D Gaussian Splatting technology, since they contain information such as positions, attributes, neural networks, and hash features, the data volume is large, and it is necessary to encode them to achieve data compression for the transmission of holographic videos.

[0003] In related technologies, only holistic data compression can be performed on the holographic video to generate a corresponding three-dimensional video file. However, this three-dimensional video file has compatibility issues on different devices and platforms. For example, some devices and platforms cannot open the three-dimensional video file, resulting in the inability to play the holographic video. Therefore, the three-dimensional video files generated by data compression of holographic videos in related technologies have poor compatibility.

[0004] Therefore, encoding the holographic videos generated based on the 3D Gaussian Splatting technology to form holographic video files that are generally compatible on multiple platforms and devices has become an urgent technical problem to be solved currently. Summary of the Invention

[0005] Embodiments of this application provide a holographic video encoding method, apparatus, computer device, and storage medium, which can encode relevant parameters of a holographic video to generate a holographic video file in the form of a binary bitstream. This holographic video file has good compatibility and can be applicable to multiple devices and platforms.

[0006] To achieve the above objective, embodiments of this application provide a holographic video encoding method, including:

[0007] Determine the video frame encoding order corresponding to the holographic video, and determine the current frame to be encoded in the holographic video according to the video frame encoding order;

[0008] When the current frame is a key frame, obtain the neural network parameters, position parameters, attribute parameters, and hash features corresponding to the target anchor points in the key frame;

[0009] Encode the neural network parameters, position parameters, attribute parameters, and hash features corresponding to the target anchor points in the key frame respectively to obtain the key frame binary bitstream data corresponding to the key frame;

[0010] When the current frame is a non-key frame, obtain the neural network parameters and cached hash features corresponding to the target anchor points in the non-key frame, and encode the neural network parameters and cached hash features corresponding to the target anchor points in the non-key frame respectively to obtain the non-key frame binary bitstream data corresponding to the non-key frame;

[0011] Determine the key frame binary bitstream data corresponding to each key frame and the non-key frame binary bitstream data corresponding to each non-key frame data in the holographic video, and generate the holographic video file corresponding to the holographic video according to the key frame binary bitstream data corresponding to each key frame and the non-key frame binary bitstream data corresponding to each non-key frame data.

[0012] To achieve the above object, an embodiment of the present application provides a holographic video encoding device, including:

[0013] A determination module, configured to determine the video frame encoding order corresponding to the holographic video, and determine the current frame to be encoded in the holographic video according to the video frame encoding order;

[0014] An acquisition module, configured to, when the current frame is a key frame, acquire the neural network parameters, position parameters, attribute parameters, and hash features corresponding to the target anchor points in the key frame;

[0015] A first encoding module, configured to encode the neural network parameters, position parameters, attribute parameters, and hash features corresponding to the target anchor points in the key frame respectively to obtain the key frame binary bitstream data corresponding to the key frame;

[0016] A second encoding module, configured to, when the current frame is a non-key frame, acquire the neural network parameters and cached hash features corresponding to the target anchor points in the non-key frame, and encode the neural network parameters and cached hash features corresponding to the target anchor points in the non-key frame respectively to obtain the non-key frame binary bitstream data corresponding to the non-key frame;

[0017] A generation module, configured to determine the key frame binary bitstream data corresponding to each key frame and the non-key frame binary bitstream data corresponding to each non-key frame data in the holographic video, and generate the holographic video file corresponding to the holographic video according to the key frame binary bitstream data corresponding to each key frame and the non-key frame binary bitstream data corresponding to each non-key frame data.

[0018] In some embodiments, the first encoding module is configured to:

[0019] Obtain the range of attribute parameters and the probability of attribute parameters corresponding to the target anchor point in the key frame;

[0020] Construct an entropy model corresponding to the attribute parameters according to the range of the attribute parameters and the probability of the attribute parameters;

[0021] Encode the attribute parameters corresponding to the target anchor point in the key frame according to the entropy model to obtain sub-data of the attribute binary code stream, and the key frame binary code stream data includes the sub-data of the attribute binary code stream.

[0022] In some embodiments, the first encoding module is configured to:

[0023] Determine the grid network corresponding to the target anchor point in the key frame;

[0024] Input the hash feature corresponding to the target anchor point in the key frame into the grid network, and output the probability of the attribute parameters corresponding to the target anchor point in the key frame.

[0025] In some embodiments, the second encoding module is configured to:

[0026] After obtaining the neural network parameters and the cached hash feature corresponding to the target anchor point in the non-key frame, obtain the historical position parameters and historical attribute parameters corresponding to the target anchor point in the previous frame of the non-key frame;

[0027] Determine the neural transformation cache network corresponding to the non-key frame according to the neural network parameters corresponding to the target anchor point in the non-key frame;

[0028] Input the cached hash feature, the historical position parameters and the historical attribute parameters into the neural transformation cache network, and output the position residual and the attribute residual corresponding to the target anchor point in the non-key frame;

[0029] Encode the position residual and the attribute residual corresponding to the target anchor point in the non-key frame to obtain sub-data of the residual binary code stream, and the non-key frame binary code stream data includes the sub-data of the residual binary code stream, and the non-key frame binary code stream data contains the sub-data of the residual binary code stream.

[0030] In some embodiments, the holographic video encoding device further includes a position determination module, configured to:

[0031] After inputting the cache hash feature, the historical position parameter, and the historical attribute parameter into the neural transformation cache network and outputting the position residual and the attribute residual corresponding to the target anchor point in the non-key frame, generate the target position parameter corresponding to the target anchor point in the non-key frame according to the position residual and the historical position parameter, and generate the target attribute parameter corresponding to the target anchor point in the non-key frame according to the attribute residual and the historical attribute parameter;

[0032] Determine the next frame to be encoded of the holographic video. When the next frame to be encoded is a non-key frame, determine the cache hash feature of the next frame to be encoded;

[0033] Input the cache hash feature of the next frame to be encoded, the target position parameter and the target attribute parameter corresponding to the target anchor point in the non-key frame into the neural transformation cache network corresponding to the next frame to be encoded, and output the position residual and the attribute residual corresponding to the target anchor point in the next frame to be encoded.

[0034] In some embodiments, the determining module is configured to:

[0035] Obtain the video frame sequence header corresponding to the holographic video;

[0036] Determine the video frame encoding order corresponding to the holographic video according to the video frame sequence header.

[0037] In some embodiments, the determining module is further configured to:

[0038] After obtaining the video frame sequence header corresponding to the holographic video, encode the video frame sequence header to obtain the sequence header binary bitstream data corresponding to the video frame sequence header.

[0039] In some embodiments, the generating module is configured to:

[0040] Determine the key frame header of each key frame and the non-key frame header of each non-key frame in the holographic video;

[0041] Encode the key frame header to generate the key frame header binary bitstream data, and encode the non-key frame header to generate the non-key frame header binary bitstream data;

[0042] Generate the holographic video file corresponding to the holographic video according to the sequence header binary bitstream data, the key frame header binary bitstream data, the non-key frame header binary bitstream data, the key frame binary bitstream data corresponding to each key frame, and the non-key frame binary bitstream data corresponding to each non-key frame data.

[0043] To achieve the above object, an embodiment of the present application provides a computer-readable storage medium storing multiple instructions adapted to be loaded by a processor to execute the holographic video encoding method provided by the embodiment of the present application.

[0044] To achieve the above object, an embodiment of the present application provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the holographic video encoding method provided by the embodiment of the present application is implemented.

[0045] In the embodiment of the present application, by determining the video frame encoding order corresponding to the holographic video and determining the current frame to be encoded in the holographic video according to the video frame encoding order; when the current frame is a key frame, obtaining the neural network parameters, position parameters, attribute parameters, and hash features corresponding to the target anchor points in the key frame; encoding the neural network parameters, position parameters, attribute parameters, and hash features corresponding to the target anchor points in the key frame respectively to obtain the key frame binary bitstream data corresponding to the key frame; when the current frame is a non-key frame, obtaining the neural network parameters and cached hash features corresponding to the target anchor points in the non-key frame, and encoding the neural network parameters and cached hash features corresponding to the target anchor points in the non-key frame respectively to obtain the non-key frame binary bitstream data corresponding to the non-key frame; determining the key frame binary bitstream data corresponding to each key frame in the holographic video and the non-key frame binary bitstream data corresponding to each non-key frame data, and generating the holographic video file corresponding to the holographic video according to the key frame binary bitstream data corresponding to each key frame and the non-key frame binary bitstream data corresponding to each non-key frame data.

[0046] Therefore, in the holographic video, the neural network parameters, position parameters, attribute parameters, and hash features corresponding to the target anchor points in the key frames of the holographic video are encoded according to the video frame encoding order to obtain the key frame binary bitstream data corresponding to the key frames. The neural network parameters and cached hash features corresponding to the target anchor points in the non-key frames are encoded to obtain the non-key frame binary bitstream data corresponding to the non-key frames. Finally, a holographic video file is generated based on the key frame binary bitstream data of each key frame and the non-key frame binary bitstream data of each non-key frame. Since the holographic video file is composed of data in binary form, the holographic video file has stronger compatibility than the three-dimensional video files in the related art. The playback device only needs to decode according to the corresponding decoding rules to obtain the information such as neural network parameters, position parameters, attribute parameters, and hash features contained in the key frame binary bitstream data. Subsequently, the rendering information of the key frames of the holographic video, such as color information and motion information, can be determined through these. The information such as neural network parameters and cached hash features contained in each non-key frame binary bitstream data can be obtained. Subsequently, the rendering information of the non-key frames of the holographic video, such as color information and motion information, can be determined through this information and in combination with the relevant information of the reference frames. The playback device can play the holographic video according to the rendering information. Therefore, in this application, it is possible to encode the relevant parameters of the holographic video to generate a holographic video file in the form of a binary bitstream. This holographic video file has good compatibility and can be applied to various devices and platforms.

[0047] Other features and advantages of the present application will be described in the subsequent specification, and, in part, will be obvious from the specification, or will be understood by implementing the present application. The objectives and other advantages of the present application can be realized and obtained through the structures specifically pointed out in the specification, claims, and drawings. Brief Description of the Drawings

[0048] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those skilled in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0049] Figure 1 It is a schematic diagram of the system framework corresponding to the holographic video encoding method provided by the embodiment of the present application;

[0050] Figure 2 It is a schematic diagram of the scenario of the holographic video encoding method provided by the embodiment of the present application;

[0051] Figure 3It is a schematic flowchart of the holographic video encoding method provided by an embodiment of the present application;

[0052] Figure 4 It is another schematic flowchart of the holographic video encoding method provided by an embodiment of the present application;

[0053] Figure 5 Schematic structural diagram of the holographic video encoding device provided by an embodiment of the present application;

[0054] Figure 6 It is a schematic structural diagram of a computer device provided by an embodiment of the present application. Detailed implementation manners

[0055] In order to enable those skilled in the art to better understand the solution of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative efforts shall fall within the protection scope of the present application.

[0056] It can be understood that in the specific implementation manners of the present application, when it comes to data related to holographic videos, when the above embodiments of the present application are applied to specific products or technologies, user permission or consent needs to be obtained, and the collection, use, and processing of relevant data need to comply with relevant laws, regulations, and standards.

[0057] It should be noted that in some processes described in the specification, claims, and the above-mentioned drawings, there are multiple steps that appear in a specific order. However, it should be clearly understood that these steps may not be executed in the order in which they appear in this document or may be executed in parallel. The step numbers are only used to distinguish different steps, and the numbers themselves do not represent any execution order. In addition, descriptions such as "first", "second", or "target" in this document are used to distinguish similar objects and do not necessarily describe a specific order or sequence.

[0058] This application can be used in numerous general-purpose or special-purpose computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer computing devices, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and so on. This application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. This application can also be practiced in a distributed computing environment where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.

[0059] Before further elaborating on the embodiments of this application, the nouns and terms involved in the embodiments of this application are explained. The nouns and terms involved in the embodiments of this application are applicable to the following explanations:

[0060] Three-dimensional Gaussian Splatting (3DGS): It is an emerging three-dimensional scene representation and rendering technology that uses 3D Gaussian functions to represent points in a scene. By optimizing the parameters of these Gaussian functions, it realizes efficient and realistic rendering from images to 3D objects.

[0061] Hash feature (hash_embedding): A hash feature is a form of feature representation obtained by processing and encoding the relevant information of 3D Gaussian points through a specific hash function. It maps the information such as the position and attributes of Gaussian points in 3D space into a low-dimensional feature vector to represent the specific attributes of the point and its relative relationship in space.

[0062] Cached hash feature (ntc_embeddings): The full name is neural transformation cached hash feature. The neural transformation cache network (NTC) is an important part of the 3D Gaussian Splatting technology. It is mainly used to process and predict the transformation information of 3D Gaussian points between different frames to achieve more efficient holographic video rendering. And the neural transformation cached hash feature generally refers to a representation method that maps high-dimensional data to a low-dimensional vector space. It can capture the features and semantic information of the data. In 3D Gaussian Splatting, the neural transformation cached hash feature is a low-dimensional vector representation output by the NTC network, used to describe the relevant features of 3D Gaussian points in non-key frames, such as features like position and attributes.

[0063] Neural Transformation Cache-Multi-Layer Perceptron (ntc_mlp): It is mainly used to predict the transformation information of 3D Gaussian points, including changes in position, rotation, scaling, etc., so as to determine the state of Gaussian points in the next frame, in order to generate continuous holographic video frames. It is responsible for performing non-linear transformation and processing on the input feature data, such as the cached hash features of anchor points, the positions and attributes corresponding to the anchor points in the adjacent previous frame, etc., and outputs coordinate residuals and attribute residuals for updating the positions and attributes of the anchor points in the current non-key frame.

[0064] MLP Grid: It is used to input the hash features of anchor points and determine the probability of attribute parameters corresponding to the anchor points according to the hash features. The mlp_grid can extract features from the input image data or scene information from multiple angles and positions at the same time. Each MLP is responsible for processing a small part of the data and generating corresponding feature representations, and then fusing these local features through a grid structure to obtain more comprehensive and representative global features.

[0065] Multi-Layer Perceptron (MLP) is a feedforward artificial neural network that contains multiple neurons arranged in layers, including an input layer, one or more hidden layers, and an output layer. It processes the input data through the connections and weights between neurons to learn the mapping relationship between the input and the output.

[0066] First, describe the technical problems existing in the related technologies:

[0067] With the development of computer technology, currently videos have evolved from two-dimensional to three-dimensional holographic videos. 3D Gaussian Splatting (3DGS) has surpassed technologies such as point clouds, meshes, and Neural Radiance Fields (NeRF) in terms of three-dimensional scene reconstruction quality, reconstruction speed, interaction freedom, and rendering speed, and has become the preferred technology for constructing three-dimensional holographic videos. The holographic videos generated based on the 3D Gaussian Splatting technology have a large amount of data because they contain information such as position, attributes, neural networks, and hash features, so they need to be encoded to achieve data compression for the transmission of holographic videos.

[0068] In the related art, only overall data compression can be performed on the holographic video to generate a corresponding three-dimensional video file. However, the three-dimensional video file may have compatibility issues on different devices and platforms. For example, some devices and platforms cannot open the three-dimensional video file, resulting in the inability to play the holographic video. Therefore, the three-dimensional video file generated by data compression of the holographic video in the related art has poor compatibility.

[0069] Therefore, data encoding of holographic videos generated based on three-dimensional Gaussian splashing technology to form universally compatible holographic video files on multiple platforms and multiple devices has become a technical problem that needs to be solved urgently.

[0070] In order to solve the above technical problems, the embodiments of the present application provide a solution for obtaining a binary code stream of relevant information by binary encoding the relevant information in the holographic video, and finally generating a holographic video file with strong compatibility based on the binary code stream data corresponding to the holographic video. Specifically, the embodiments of the present application provide a holographic video encoding method, device, computer equipment and storage medium, which will be described in detail later.

[0071] See also Figure 1 , Figure 1 Schematic diagram of the system framework corresponding to the holographic video encoding method provided in the embodiment of the present application. The holographic video encoding method provided in the embodiment of the present application can be applied to the system framework.

[0072] It includes a terminal 140, the Internet 130, a gateway 120, a server 110, and the like.

[0073] The terminal 140 or the server 110 may be a device that performs the holographic video encoding method.

[0074] The terminal 140 includes but is not limited to mobile phones, computers, intelligent voice interaction devices, smart home appliances, vehicle-mounted terminals, aircraft, etc. The embodiments of the present application can be applied to various scenarios, including but not limited to cloud office, enterprise management, etc. In addition, it can be a single device or a collection of multiple devices. For example, multiple desktop computers are interconnected through a local area network, share a display, etc. to work together, and together constitute a terminal 140. The terminal 140 can communicate with the Internet 130 in a wired or wireless manner to exchange data.

[0075] Server 110 refers to a computer system that can provide certain services to terminal 140. Compared with ordinary terminal 140, server 110 has higher requirements in terms of stability, security, performance, etc. Server 110 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, as well as big data and artificial intelligence platforms.

[0076] The gateway 120 is also called an internetwork connector and protocol converter. The gateway realizes network interconnection at the transport layer and is a computer system or device that acts as a converter. Between two systems that use different communication protocols, data formats, or languages, or even have completely different architectures, the gateway is a translator. At the same time, the gateway can also provide filtering and security functions. The message sent by terminal 140 to server 110 needs to be sent to the corresponding server 110 through gateway 120. The message sent by server 110 to terminal 140 also needs to be sent to the corresponding terminal 140 through gateway 120.

[0077] The holographic video encoding method in the embodiments of this application can be applied to a variety of scenarios, such as virtual reality, multimedia playback, and other scenarios. The scenarios to which the holographic video encoding method in this application is applied are not limited herein.

[0078] Please refer to Figure 2 , Figure 2 which is a schematic diagram of the scenario of the holographic video encoding method provided by the embodiments of this application.

[0079] Among them, the holographic video described in this application can be a three-dimensional video generated based on the three-dimensional Gaussian splash technology. The holographic video contains multiple video frames, and these multiple video frames form a video frame sequence. By playing the video frames in sequence according to the video frame sequence, the playback of the holographic video can be realized.

[0080] Among them, the video frame sequence contains key frames and non-key frames. The key frames refer to the frames that are of special importance in the holographic video. They usually can represent the key scenes, key actions, or key time points in the video, etc., and are the iconic frames in the entire video sequence. These frames play a key role in understanding the main content and structure of the video, similar to the key scenes in a movie or the key poses in an animation. During the processing of the three-dimensional Gaussian splash technology, the information of the key frames is often used as a reference to process other frames. For example, when performing 3D reconstruction and rendering, the positions, colors, shapes, and other attributes of the 3D Gaussian points in the key frames and the relationships between them can provide important bases for the calculation and adjustment of the corresponding attributes of the non-key frames, helping to ensure the coherence and consistency of the entire video in space and time.

[0081] Non-key frames are frames that are of secondary importance in a video relative to key frames. They are used for the transition and supplementation between key frames to more delicately represent the continuity and dynamic changes of the video. One of the main functions of non-key frames is to transition between adjacent key frames, making the action and scene changes in the video more natural and smooth. By inserting an appropriate number of non-key frames between key frames, the sense of jump in the video can be avoided, allowing the audience to more naturally feel the dynamic processes such as the movement of objects and the transition of scenes.

[0082] During the process of encoding a holographic video, the video frame sequence header can be encoded to generate the video frame sequence header binary bitstream data, that is, data in binary form composed of the characters '0' and '1'. The video frame sequence header is the header file of the entire video frame sequence, which contains the basic information of the video frame sequence, such as the number of frames in the sequence.

[0083] The key frame header can also be encoded to obtain the key frame header binary bitstream data, that is, data in binary form composed of the characters '0' and '1'. The key frame header is the header file corresponding to the key frame, which contains the basic information of the key frame, such as the number of anchor points.

[0084] Encode the key frame to generate the key frame binary bitstream data, that is, data in binary form composed of the characters '0' and '1'. Specifically, the neural network parameters, position parameters, attribute parameters, and hash features corresponding to the target anchor points in the key frame can be obtained; the neural network parameters, position parameters, attribute parameters, and hash features corresponding to the target anchor points in the key frame are encoded respectively to obtain the key frame binary bitstream data corresponding to the key frame. Among them, the target anchor points in the key frame can be each anchor point in the key frame. The neural network parameters include the relevant parameters of various neural networks. The position parameters include the three-dimensional space coordinates of the target anchor points in the key frame. The attribute parameters include the anchor point features, such as the color and opacity corresponding to the anchor point, and the attribute parameters also include parameters such as the scale and offset corresponding to the anchor point. The hash feature can be used as a concise and representative way to represent the anchor point. Since the hash function can map complex anchor point information (such as position, attribute, etc.) to a hash code of a fixed length, this enables each anchor point to be described in a compact form, thus reducing the burden of data storage and processing. The hash feature can be used as an index of the anchor point to facilitate the quick query and access of specific anchor points and their related position, attribute, etc. data.

[0085] For non-key frames, the non-key frame header can be encoded to obtain non-key frame binary bitstream data, that is, data in binary form consisting of the characters '0' and '1'. The non-key frame header is the header file corresponding to the non-key frame and contains the corresponding basic information of the non-key frame, such as the non-key frame start code.

[0086] Encode the non-key frame to generate non-key frame binary bitstream data, that is, data in binary form consisting of the characters '0' and '1'. Specifically, obtain the neural network parameters and cache hash features corresponding to the target anchor points in the non-key frame, and encode the neural network parameters and cache hash features corresponding to the target anchor points in the non-key frame respectively to obtain the non-key frame binary bitstream data corresponding to the non-key frame. Among them, the target anchor points in the non-key frame can be each anchor point in the non-key frame. The neural network parameters include the parameters corresponding to the neural transformation cache network. The NTC network first extracts features from the input non-key frame data, which may include information such as the positions, colors, and normals of 3D Gaussian points, as well as the relationship with the surrounding environment. Through a series of neural network layers, these features are encoded and transformed, mapped into a low-dimensional vector space to obtain cache hash features. The cache hash features can be used to identify the static features of the current non-key frame and can also reflect the dynamic changes of the 3D scene in the time dimension.

[0087] Finally, a holographic video file corresponding to the holographic video can be generated according to the sequence header binary bitstream data, the key frame header binary bitstream data corresponding to each key frame, the non-key frame header binary bitstream data corresponding to each non-key frame data, the key frame binary bitstream data corresponding to each key frame, and the non-key frame binary bitstream data corresponding to each non-key frame data.

[0088] For example, the sequence header of the holographic video can be encoded first to obtain the sequence header binary bitstream data. When a key frame is read, the key frame header is encoded first to obtain the key frame header binary bitstream data, and then the key frame data corresponding to the key frame is encoded to obtain the key frame binary bitstream data. When a non-key frame is read, the non-key frame header is encoded first to obtain the non-key frame header binary bitstream data, and then the non-key frame data corresponding to the non-key frame is encoded to obtain the non-key frame binary bitstream data. After the holographic video is encoded, a holographic video file corresponding to the holographic video can be generated according to the sequence header binary bitstream data, the key frame header binary bitstream data and the key frame binary bitstream data corresponding to each key frame, and the non-key frame header binary bitstream data and the non-key frame binary bitstream data corresponding to each non-key frame.

[0089] Since the holographic video file is composed of data in binary form, the holographic video file has stronger compatibility compared to the 3D video files in the related art. The playback device only needs to decode according to the corresponding decoding rules to obtain the information such as neural network parameters, position parameters, attribute parameters, and hash features contained in the key frame binary code stream data. Subsequently, the rendering information of the holographic video key frame, such as color information and motion information, can be determined through these. For each non-key frame binary code stream data, the information such as neural network parameters and cached hash features is obtained. Subsequently, the rendering information of the holographic video non-key frame, such as color information and motion information, can be determined through this information and in combination with the relevant information of the reference frame. The playback device can play the holographic video according to the rendering information. Therefore, in this application, it is possible to encode the relevant parameters of the holographic video to generate a holographic video file in the form of a binary code stream. This holographic video file has good compatibility and can be applied to multiple devices and platforms.

[0090] To understand the holographic video encoding method provided by the embodiments of this application in more detail, please continue to refer to Figure 3 , Figure 3 which is a schematic flowchart of the holographic video encoding method provided by the embodiments of this application. The holographic video encoding method may include the following steps:

[0091] Step 210: Determine the video frame encoding order corresponding to the holographic video, and determine the current frame to be encoded in the holographic video according to the video frame encoding order;

[0092] Step 220: When the current frame is a key frame, obtain the neural network parameters, position parameters, attribute parameters, and hash features corresponding to the target anchor points in the key frame;

[0093] Step 230: Encode the neural network parameters, position parameters, attribute parameters, and hash features corresponding to the target anchor points in the key frame respectively to obtain the key frame binary code stream data corresponding to the key frame;

[0094] Step 240: When the current frame is a non-key frame, obtain the neural network parameters and cached hash features corresponding to the target anchor points in the non-key frame, and encode the neural network parameters and cached hash features corresponding to the target anchor points in the non-key frame respectively to obtain the non-key frame binary code stream data corresponding to the non-key frame;

[0095] Step 250: Determine the key frame binary code stream data corresponding to each key frame and the non-key frame binary code stream data corresponding to each non-key frame data in the holographic video, and generate the holographic video file corresponding to the holographic video according to the key frame binary code stream data corresponding to each key frame and the non-key frame binary code stream data corresponding to each non-key frame data.

[0096] The following will describe steps 210 to 250 in detail.

[0097] In step 210, determine the encoding order of the video frames corresponding to the holographic video, and determine the current frame to be encoded in the holographic video according to the video frame encoding order.

[0098] It can be understood that in the holographic video, there is a sequence of video frames arranged in order, and the relevant information of the non-key frame target anchor points needs to be determined based on the relevant information of the previous frame anchor points as reference information. For example, the position and attributes of the target anchor points in the non-key frames need to be determined based on the position and attributes of the target anchor points in the previous video frame as a reference. Therefore, it is necessary to encode each video frame in the holographic video based on the encoding order of the video frames.

[0099] In some embodiments, determining the encoding order of the video frames corresponding to the holographic video includes:

[0100] (1.1) Obtain the video frame sequence header corresponding to the holographic video;

[0101] (1.2) Determine the encoding order of the video frames corresponding to the holographic video according to the video frame sequence header.

[0102] Among them, the video frame sequence header corresponding to the holographic video includes: video sequence start code (video_sequence_start_code), represented by the string '0x000001B0', used to identify the start of the video sequence; number of frames in the sequence (num_frames), used to identify the total number of frames in the video sequence; group of pictures size (gop_size), used to identify the size of the encoded group of pictures.

[0103] The encoding order of the video frames corresponding to the holographic video can be determined according to the video frame sequence header, that is, the encoding order from the first video frame to the last video frame. Then, determine the current frame to be encoded in the holographic video according to the video frame encoding order.

[0104] In some embodiments, after obtaining the video frame sequence header corresponding to the holographic video, the video frame sequence header can also be encoded to obtain the sequence header binary bitstream data corresponding to the video frame sequence header. For example, the video sequence start code, number of frames in the sequence, and group of pictures size can be encoded to generate the sequence header binary bitstream data.

[0105] In step 220, when the current frame is a key frame, obtain the neural network parameters, position parameters, attribute parameters, and hash features corresponding to the target anchor points in the key frame.

[0106] If it is determined that the current frame is a key frame, the neural network parameters, position parameters, attribute parameters, and hash features corresponding to the target anchor points in the key frame can be obtained. Among them, the target anchor points in the key frame can be each anchor point in the key frame. The neural network parameters include the relevant parameters of various neural networks. For example, the opacity network (mlp_opacity) is a multi-layer perceptron network used to derive the opacity of 3D Gaussian points for the target anchor points; the covariance network (mlp_cov) is a multi-layer perceptron network used to derive the covariance of 3D Gaussian points for the target anchor points; the color network (mlp_color) is a multi-layer perceptron network used to derive the color of 3D Gaussian points for the target anchor points; the grid network (mlp_grid) is a multi-layer perceptron network used to derive the probability of the attribute parameters of the target anchor points from the hash features.

[0107] The position parameters include the three-dimensional spatial coordinates of the target anchor points in the key frame.

[0108] The attribute parameters include the anchor point feature, anchor point scaling, and anchor point offset. The anchor point feature is used to identify the feature attributes of the target anchor points, such as attributes like opacity, covariance, and color used to derive 3D Gaussian points. The anchor point scaling is used to identify the scaling attributes of the target anchor points and is used for the position attributes of deriving 3D Gaussian points. The anchor point offset is used to identify the offset attributes of the target anchor points and is used for the position attributes of deriving 3D Gaussian points.

[0109] The hash feature can be used as a concise and representative way to represent the target anchor points. Since the hash function can map complex anchor point information (such as position, attributes, etc.) to a hash code of a fixed length, this enables each target anchor point to be described in a compact form, thereby reducing the burden of data storage and processing. The hash feature can be used as an index for the anchor points, facilitating quick query and access to the target anchor points and their related position, attribute, and other data.

[0110] In some embodiments, before encoding the neural network parameters, position parameters, attribute parameters, and hash features corresponding to the target anchor points in the key frame respectively to obtain the key frame binary code stream data corresponding to the key frame, it is also necessary to encode the key frame header to generate the key frame header binary code stream data.

[0111] Among them, the key frame header includes: the intra-frame start code, such as represented by the string '0x000001B2', which is used to identify the start of the key frame;

[0112] The number of anchor points (num_anchors), which is used to identify the number of anchor points in the three-dimensional scene;

[0113] The maximum anchor batch size (max_batch_size) is used to identify the maximum number of anchor batches during 3D scene encoding;

[0114] The minimum anchor coordinate value (x_bound_min) is used to identify the minimum value of the 3D coordinates of the anchors in the 3D scene;

[0115] The maximum anchor coordinate value (x_bound_max) is used to identify the maximum value of the 3D coordinates of the anchors in the 3D scene;

[0116] The list of minimum values of anchor features per batch (min_feat_list) is used to identify the list of minimum values of the quantized features of the anchors per batch. This list is the index of the anchors per batch, and the corresponding value of the index is the minimum value of the quantized features of the anchors in this batch;

[0117] The list of maximum values of anchor features per batch (max_feat_list) is used to identify the list of maximum values of the quantized features of the anchors per batch. This list is the batch index of the anchors per batch, and the corresponding value of the index is the maximum value of the quantized features of the anchors in this batch;

[0118] The list of minimum values of anchor scales per batch (min_scaling_list) is used to identify the list of minimum values of the quantized scales of the anchors per batch. This list is the batch index of the anchors per batch, and the corresponding value of the index is the minimum value of the quantized scale of the anchors in this batch;

[0119] The list of maximum values of anchor scales per batch (max_scaling_list) is used to identify the list of maximum values of the quantized scales of the anchors per batch. This list is the batch index of the anchors per batch, and the corresponding value of the index is the maximum value of the quantized scale of the anchors in this batch;

[0120] The list of minimum values of anchor offsets per batch (min_offsets_list) is used to identify the list of minimum values of the quantized offsets of the anchors per batch. This list is the batch index of the anchors per batch, and the corresponding value of the index is the minimum value of the quantized offset of the anchors in this batch;

[0121] The list of maximum values of anchor offsets per batch (max_offsets_list) is used to identify the list of maximum values of the quantized offsets of the anchors per batch. This list is the batch index of the anchors per batch, and the corresponding value of the index is the maximum value of the quantized offset of the anchors in this batch;

[0122] The binary hash grid probability distribution (prob_hash) is used to identify the probability of "+1" in the binary hash grid, and entropy encoding is performed on the features in the binary hash grid based on the probability.

[0123] It is possible to encode the content included in the above key frame header to generate the key pillow binary code stream data in binary form.

[0124] In step 230, the neural network parameters, position parameters, attribute parameters, and hash features corresponding to the target anchor points in the key frame are respectively encoded to obtain the key frame binary bitstream data corresponding to the key frame.

[0125] Among them, the neural network parameters can be encoded, and the binary bitstream data generated during the encoding process can be obtained. For example, the opaque network, covariance network, color network, and grid network are encoded to obtain the binary bitstream data corresponding to the neural network parameters. The position parameters can be encoded, and the binary bitstream data generated during the encoding process can be obtained. The hash features can be encoded, and the binary bitstream data generated during the encoding process can be obtained. The attribute parameters can be encoded, and the binary bitstream data generated during the encoding process can be obtained.

[0126] The hash features corresponding to each anchor point in each video frame of the holographic video can be stored in the hash table corresponding to each video frame. This hash table can be understood as a set of hash features and has a mapping relationship between the hash features and the three-dimensional spatial coordinates of the anchor points. The hash features corresponding to the anchor points can be found in the hash table according to this mapping relationship and the three-dimensional spatial coordinates of the anchor points.

[0127] During the encoding process of the key frame, the hash table corresponding to the key frame can be encoded. Since the hash table contains the hash features corresponding to the target anchor points in the key frame, during the encoding process of the hash table, the hash features of the target anchor points in the key frame are also encoded, thereby obtaining the binary bitstream data generated during the encoding process of the hash features of the target anchor points in the key frame.

[0128] These binary bitstream data can be understood as components of the key frame binary bitstream data, and these binary bitstream data can be determined as sub-data of the key frame binary bitstream data.

[0129] It should be noted that when encoding the neural network parameters, position parameters, attribute parameters, and hash features of the anchor points in the key frame, all the anchor points in the key frame can be divided into multiple batches of anchor points. Specifically, it is the ceiling of the result of dividing the total number of anchor points by the maximum batch size (max_batch_size), thereby obtaining the number of anchor point batches. The relevant information of the anchor points corresponding to each batch can be encoded until the relevant information of all the anchor points in all batches is encoded, then the encoding of the key frame is completed, and the key frame binary bitstream data corresponding to the key frame is obtained.

[0130] In some embodiments, encoding the attribute parameters corresponding to the target anchor points in the key frame includes:

[0131] (1.1) Obtain the range and probability of the attribute parameters corresponding to the target anchor point in the key frame;

[0132] (1.2) Construct an entropy model corresponding to the attribute parameters according to the range and probability of the attribute parameters;

[0133] (1.3) Encode the attribute parameters corresponding to the target anchor point in the key frame according to the entropy model to obtain sub-data of the attribute binary code stream, and the key frame binary code stream data includes the sub-data of the attribute binary code stream.

[0134] Among them, the corresponding hash feature of the target anchor point can be found according to the three-dimensional spatial coordinates of the target anchor point. For example, the hash feature corresponding to the target anchor point is found through the index relationship between the hash feature and the three-dimensional spatial coordinates, and then the range of the attribute parameters corresponding to the target anchor point is searched according to the hash feature of the target anchor point, that is, through the mapping relationship between the hash feature and the anchor point-related information, the maximum and minimum values of the attribute parameters corresponding to the target anchor point are found, so as to determine the range of the attribute parameters.

[0135] For example, find the minimum value of the anchor point feature corresponding to the target anchor point in the list of minimum values of each batch of anchor point features corresponding to the target anchor point, find the maximum value of the anchor point feature corresponding to the target anchor point in the list of maximum values of each batch of anchor point features corresponding to the target anchor point, find the minimum value of the anchor point scale corresponding to the target anchor point in the list of minimum values of each batch of anchor point scales, find the maximum value of the anchor point scale corresponding to the target anchor point in the list of maximum values of each batch of anchor point scales, find the minimum value of the anchor point offset corresponding to the target anchor point in the list of minimum values of each batch of anchor point offsets, and find the maximum value of the anchor point offset corresponding to the target anchor point in the list of maximum values of each batch of anchor point offsets. Then, through these maximum and minimum values, the range of the attribute parameters corresponding to the target anchor point can be determined.

[0136] Then, an entropy model corresponding to the attribute parameters is constructed according to the range and probability of the attribute parameters, and the attribute parameters corresponding to the target anchor point in the key frame are encoded according to the entropy model to obtain sub-data of the attribute binary code stream, and the key frame binary code stream data includes the sub-data of the attribute binary code stream.

[0137] Among them, the entropy model can be a model determined based on Shannon entropy. Shannon entropy measures the uncertainty of the information source. According to the principle of Shannon entropy, the goal of information encoding is to represent information with as few encoding symbols as possible. To achieve this goal, for symbols with high probability, shorter encodings should be used; for symbols with low probability, longer encodings should be used. Therefore, in this application, the minimum and maximum numbers of the encoding symbols taken by the entropy model for encoding the attribute parameters can be determined through the range of the attribute parameters, and the number of encoding symbols corresponding to the attribute parameters of the target anchor point can be determined within the range between the maximum number and the minimum number. The probability of the encoding symbols corresponding to the attribute parameters can be determined through the attribute parameter probability.

[0138] In the present application, an entropy model can encode the attribute parameters corresponding to the target anchor points in the key frames, so as to obtain the binary bitstream data corresponding to the attribute parameters.

[0139] In some embodiments, obtaining the probability of the attribute parameters corresponding to the target anchor points in the key frames includes:

[0140] (1.1.1) Determine the grid network corresponding to the target anchor points in the key frames;

[0141] (1.1.2) Input the hash features corresponding to the target anchor points in the key frames into the grid network, and output the probability of the attribute parameters corresponding to the target anchor points in the key frames.

[0142] Among them, the grid network corresponding to the target anchor points in the key frames can be determined. For example, the grid network can be determined through the neural network parameters corresponding to the target anchor points in the key frames. The grid network can be a trained neural network, and then the hash features corresponding to the target anchor points in the key frames are input into the grid network, so as to output the probability of the attribute parameters corresponding to the target anchor points in the key frames.

[0143] In step 240, when the current frame is a non-key frame, obtain the neural network parameters and cached hash features corresponding to the target anchor points in the non-key frame, and encode the neural network parameters and cached hash features corresponding to the target anchor points in the non-key frame respectively to obtain the non-key frame binary bitstream data corresponding to the non-key frame.

[0144] In some embodiments, if the current frame is a non-key frame, before encoding the non-key frame, the non-key frame header can be encoded to obtain the non-key frame header binary bitstream data corresponding to the non-key frame header.

[0145] Among them, the non-key frame header includes:

[0146] The non-key frame start code (predictive_frame_start_code), represented by the bit string '0x000001B3', is used to identify the start of the non-key frame;

[0147] The binary hash grid probability distribution (prob_ntc), which is used to identify the probability of '+1' in the binary hash grid, and entropy encodes the features in the binary hash grid based on the probability. The binary hash grid is used to implement the neural transformation cache of the non-key frame.

[0148] If the current frame is a non-key frame, the neural network parameters and cached hash features corresponding to the target anchor points in the non-key frame can be obtained. Among them, the neural network parameters include the relevant parameters of various neural networks. For example, the opacity network (mlp_opacity) is a multi-layer perceptron network used when deriving the opacity of 3D Gaussian points for the target anchor points; the covariance network (mlp_cov) is a multi-layer perceptron network used when deriving the covariance of 3D Gaussian points for the target anchor points; the color network (mlp_color) is a multi-layer perceptron network used when deriving the color of 3D Gaussian points for the target anchor points; the grid network (mlp_grid) is a multi-layer perceptron network used when deriving the probability of the target anchor point attribute parameters from the hash features; the neural transformation cache network (ntc_mlp) is a multi-layer perceptron network used when caching the time-domain transformation parameters of the anchor point offset derived from the hash features.

[0149] The cached hash feature is actually a neural network hash cached feature, which can be understood as a vector corresponding to the target anchor points in the non-key frame. The cached hash feature is a compact feature vector representation obtained by comprehensively encoding various attributes of the 3D Gaussian points. Based on specific mathematical transformations and encoding methods, it maps the attribute information of the 3D Gaussian points in the holographic video, such as position, color, shape, transparency, etc., as well as its dynamic change information in the video sequence, such as motion trajectory, speed, etc., into a vector in a high-dimensional vector space. The cached hash feature can capture the semantic information of the 3D Gaussian points and their state features in the neural transformation to a certain extent.

[0150] In the process of generating and rendering holographic videos, the computational cost of neural transformation is usually very high. By caching the already calculated neural transformation results, the cached hash feature avoids repeated calculations, thus greatly improving the processing efficiency. For example, for adjacent frames or similar scene parts in the video, their neural transformation results may be similar. By caching the cached hash features corresponding to these results, they can be directly called in subsequent processing, reducing unnecessary computational overhead.

[0151] In holographic videos, it is necessary to ensure the transformation consistency and coherence between different frames to avoid unnatural visual effects such as flickering and jumping. The cached hash feature helps to maintain this consistency and coherence by caching and reusing similar features between adjacent frames. For example, during the movement of an object, the position and attribute changes of the Gaussian points in adjacent frames are usually continuous. By using the cached hash features, this continuity can be better maintained, making the generated holographic video more smooth and natural.

[0152] The neural network parameters corresponding to the target anchor points in the non-key frame can be encoded to obtain the binary bitstream data corresponding to the neural network parameters.

[0153] The cache hash features corresponding to the target anchor points in non-key frames can be encoded to obtain the binary bitstream data corresponding to the cache hash features. The cache hash features corresponding to each anchor point in the non-key frame can be stored in the hash table corresponding to the non-key frame. This hash table can be understood as a set of cache hash features and has a mapping relationship between the cache hash features and the three-dimensional spatial coordinates of the target anchor points in the non-key frame. The cache hash features corresponding to the target anchor points can be found in the hash table according to this mapping relationship and the three-dimensional spatial coordinates of the target anchor points in the non-key frame.

[0154] During the encoding process of the non-key frame, the hash table corresponding to the non-key frame can be encoded. Since the hash table contains the cache hash features corresponding to the target anchor points in the non-key frame, during the encoding process of the hash table, the encoding of the cache hash features of the target anchor points in the non-key frame is also realized, so as to obtain the binary bitstream data generated during the encoding process of the cache hash features of the target anchor points in the non-key frame.

[0155] Among them, the binary bitstream data corresponding to the non-key frame includes the binary bitstream data corresponding to the neural network parameters and the binary bitstream data corresponding to the cache hash features. The target anchor points in the non-key frame can be each anchor point. After the encoding of the neural network parameters and cache hash features corresponding to all the anchor points in the non-key frame is completed, it is considered that the encoding of the non-key frame is completed, and the binary bitstream data corresponding to the non-key frame is obtained.

[0156] In some embodiments, after obtaining the neural network parameters and cache hash features corresponding to the target anchor points in the non-key frame, it further includes:

[0157] (1.1) Obtain the historical position parameters and historical attribute parameters corresponding to the target anchor points in the non-key frame in the previous frame;

[0158] (1.2) Determine the neural transformation cache network corresponding to the non-key frame according to the neural network parameters corresponding to the target anchor points in the non-key frame;

[0159] (1.3) Input the cache hash features, historical position parameters and historical attribute parameters into the neural transformation cache network, and output the position residuals and attribute residuals corresponding to the target anchor points in the non-key frame;

[0160] (1.4) Encode the position residuals and attribute residuals corresponding to the target anchor points in the non-key frame to obtain the residual binary bitstream sub-data, and the binary bitstream data corresponding to the non-key frame includes the residual binary bitstream sub-data.

[0161] Among them, non-key frames serve as supplementary frames between key frames. The relevant information of the anchor points in non-key frames is determined based on the relevant information of the anchor points in the previous non-key frame or the previous key frame. For example, the position parameters and attribute parameters of the anchor points in the previous key frame can be used as reference information for the target anchor points in the current non-key frame, so as to update the position parameters and attribute parameters of the target anchor points in the non-key frame.

[0162] According to the neural network parameters corresponding to the target anchor points in the non-key frame, the neural transformation cache network corresponding to the non-key frame is determined. Then, the cached hash feature, historical position parameters, and historical attribute parameters are input into the neural transformation cache network, and the position residuals and attribute residuals corresponding to the target anchor points in the non-key frame are output. The position residual can be understood as the position change value of the target anchor point in the non-key frame relative to its position in the previous frame, and the attribute residual can be understood as the attribute change value of the target anchor point in the non-key frame relative to its attribute in the previous frame.

[0163] Therefore, the position residuals and attribute residuals corresponding to the target anchor points in the non-key frame can be encoded to obtain residual binary code stream sub-data. The non-key frame binary code stream data includes the residual binary code stream sub-data, and the non-key frame binary code stream data contains the residual binary code stream sub-data. In this way, the positions and attributes corresponding to the target anchor points in the non-key frame can be updated according to these position residuals and attribute residuals respectively later.

[0164] In some embodiments, after inputting the cached hash feature, historical position parameters, and historical attribute parameters into the neural transformation cache network and outputting the position residuals and attribute residuals corresponding to the target anchor points in the non-key frame, it further includes:

[0165] (2.1) Generate the target position parameters corresponding to the target anchor points in the non-key frame according to the position residuals and historical position parameters, and generate the target attribute parameters corresponding to the target anchor points in the non-key frame according to the attribute residuals and historical attribute parameters;

[0166] (2.2) Determine the next frame to be encoded in the holographic video. When the next frame to be encoded is a non-key frame, determine the cached hash feature of the next frame to be encoded;

[0167] (2.3) Input the cached hash feature of the next frame to be encoded, the target position parameters and target attribute parameters corresponding to the target anchor points in the non-key frame into the neural transformation cache network corresponding to the next frame to be encoded, and output the position residuals and attribute residuals corresponding to the target anchor points in the next frame to be encoded.

[0168] For example, add the position residual and the historical position parameter to obtain the target position parameter corresponding to the target anchor point in the current non-key frame, and add the attribute residual and the historical attribute parameter to obtain the target attribute parameter corresponding to the target anchor point in the current non-key frame.

[0169] Then, determine the next frame to be encoded in the holographic video. When the next frame to be encoded is a non-key frame, determine the cached hash feature of the next frame to be encoded, and input the cached hash feature of the next frame to be encoded, the target position parameter and the target attribute parameter corresponding to the target anchor point in the non-key frame into the neural transformation cache network corresponding to the next frame to be encoded, and output the position residual and the attribute residual corresponding to the target anchor point in the next frame to be encoded.

[0170] That is to say, the position residual and the attribute residual corresponding to the target anchor point in the next frame to be encoded are determined based on the target position parameter and the target attribute parameter corresponding to the target anchor point in the current non-key frame. In this way, when encoding the next frame to be encoded, the position residual and the attribute residual corresponding to the target anchor point in the next frame to be encoded can be encoded, so as to obtain the residual binary code stream sub-data corresponding to the target anchor point in the next frame to be encoded.

[0171] In some embodiments, after obtaining the neural network parameters and the cached hash feature corresponding to the target anchor point in the non-key frame, it further includes:

[0172] (3.1) Determine the neural transformation cache network corresponding to the non-key frame according to the neural network parameters corresponding to the target anchor point in the non-key frame;

[0173] (3.2) Input the cached hash feature corresponding to the target anchor point in the non-key frame into the neural transformation cache network, and output the position residual and the attribute residual corresponding to the target anchor point in the non-key frame;

[0174] (3.3) Encode the position residual and the attribute residual corresponding to the target anchor point in the non-key frame to obtain the residual binary code stream sub-data, and the non-key frame binary code stream data includes the residual binary code stream sub-data.

[0175] Determine the neural transformation cache network corresponding to the non-key frame according to the neural network parameters corresponding to the target anchor point in the non-key frame, and then input the cached hash feature into the neural transformation cache network, and output the position residual and the attribute residual corresponding to the target anchor point in the non-key frame. The position residual can be understood as the position change value of the target anchor point in the non-key frame relative to its position in the previous frame, and the attribute residual can be understood as the attribute change value of the target anchor point in the non-key frame relative to its attribute in the previous frame.

[0176] Therefore, the position residuals and attribute residuals corresponding to the target anchor points in the non-key frames can be encoded to obtain sub-data of the residual binary bitstream. The non-key frame binary bitstream data includes the sub-data of the residual binary bitstream, and the non-key frame binary bitstream data contains the sub-data of the residual binary bitstream. In this way, during the subsequent decoding process of the holographic video file, the sub-data of the residual binary bitstream can be decoded to provide the position residuals and attribute residuals corresponding to the target anchor points in the non-key frames, so as to update the positions and attributes of the target anchor points according to the position residuals and attribute residuals corresponding to the target anchor points in the non-key frames.

[0177] It should be noted that after obtaining the position residuals and attribute residuals corresponding to the target anchor points in the non-key frames, the position residuals and attribute residuals corresponding to the target anchor points in the non-key frames may not be encoded. Instead, the positions and attributes of the target anchor points are updated according to the position residuals and attribute residuals corresponding to the target anchor points in the non-key frames, so as to obtain the target position parameters and target attribute parameters corresponding to the target anchor points in the non-key frames. When the next frame to be encoded is a non-key frame, the target position parameters and target attribute parameters corresponding to the target anchor points in the current non-key frame actually serve as reference information for the next frame to be encoded, to help determine the positions and attributes corresponding to the target anchor points in the next frame to be encoded, so as to generate the cached hash features corresponding to the target anchor points in the next frame to be encoded. In this way, in the non-key frame binary bitstream data corresponding to the non-key frame, it actually only needs to contain the cached hash features corresponding to the target anchor points in the non-key frame and the binary bitstream data corresponding to the neural network parameters.

[0178] Step 250: Determine the key frame binary bitstream data corresponding to each key frame and the non-key frame binary bitstream data corresponding to each non-key frame data in the holographic video, and generate a holographic video file corresponding to the holographic video according to the key frame binary bitstream data corresponding to each key frame and the non-key frame binary bitstream data corresponding to each non-key frame data.

[0179] In this application, each key frame header, each key frame, each non-key frame header, and each non-key frame in the holographic video can be encoded to obtain the key frame header binary bitstream data corresponding to each key frame header, the key frame binary bitstream data corresponding to each key frame, the non-key frame header binary bitstream data corresponding to each non-key frame header, and the non-key frame binary bitstream data corresponding to each non-key frame data.

[0180] Finally, a holographic video file corresponding to the holographic video is generated based on the sequence header binary bitstream data, the key frame header binary bitstream data and the key frame binary bitstream data corresponding to each key frame, and the non-key frame header binary bitstream data and the non-key frame binary bitstream data corresponding to each non-key frame data. For example, the sequence header of the holographic video can be encoded first to obtain the sequence header binary bitstream data. When a key frame is read, the key frame header is encoded first to obtain the key frame header binary bitstream data, and then the key frame data corresponding to the key frame is encoded to obtain the key frame binary bitstream data. When a non-key frame is read, the non-key frame is encoded first to obtain the non-key frame header binary bitstream data, and then the non-key frame data corresponding to the non-key frame is encoded to obtain the non-key frame binary bitstream data. After the holographic video is encoded, a holographic video file corresponding to the holographic video can be generated based on the sequence header binary bitstream data, the key frame header binary bitstream data and the key frame binary bitstream data corresponding to each key frame, and the non-key frame header binary bitstream data and the non-key frame binary bitstream data corresponding to each non-key frame.

[0181] In the embodiment of the present application, by determining the video frame encoding order corresponding to the holographic video and determining the current frame to be encoded in the holographic video according to the video frame encoding order; when the current frame is a key frame, obtaining the neural network parameters, position parameters, attribute parameters and hash features corresponding to the target anchor points in the key frame; encoding the neural network parameters, position parameters, attribute parameters and hash features corresponding to the target anchor points in the key frame respectively to obtain the key frame binary bitstream data corresponding to the key frame; when the current frame is a non-key frame, obtaining the neural network parameters and cached hash features corresponding to the target anchor points in the non-key frame, and encoding the neural network parameters and cached hash features corresponding to the target anchor points in the non-key frame respectively to obtain the non-key frame binary bitstream data corresponding to the non-key frame; determining the key frame binary bitstream data corresponding to each key frame and the non-key frame binary bitstream data corresponding to each non-key frame data in the holographic video, and generating a holographic video file corresponding to the holographic video according to the key frame binary bitstream data corresponding to each key frame and the non-key frame binary bitstream data corresponding to each non-key frame data.

[0182] Therefore, in the holographic video, according to the encoding order of video frames, the neural network parameters, position parameters, attribute parameters, and hash features corresponding to the target anchor points in the key frames of the holographic video are encoded respectively to obtain the key frame binary bitstream data corresponding to the key frames. The neural network parameters and cached hash features corresponding to the target anchor points in the non-key frames are encoded respectively to obtain the non-key frame binary bitstream data corresponding to the non-key frames. Finally, a holographic video file is generated according to the key frame binary bitstream data of each key frame and the non-key frame binary bitstream data of each non-key frame. Since the holographic video file is composed of data in binary form, the holographic video file has stronger compatibility than the three-dimensional video files in the related art. The playback device only needs to decode according to the corresponding decoding rules to obtain the information such as neural network parameters, position parameters, attribute parameters, and hash features contained in the key frame binary bitstream data. Subsequently, the rendering information of the key frames of the holographic video, such as color information and motion information, can be determined through these. The information such as neural network parameters and cached hash features contained in each non-key frame binary bitstream data can be obtained. Subsequently, the rendering information of the non-key frames of the holographic video, such as color information and motion information, can be determined through this information and in combination with the relevant information of the reference frames. The playback device can play the holographic video according to the rendering information. Therefore, in this application, it is possible to encode the relevant parameters of the holographic video to generate a holographic video file in the form of a binary bitstream. The holographic video file has good compatibility and can be applied to various devices and platforms.

[0183] Please refer to Figure 4 , Figure 4 which is another schematic flowchart of the holographic video encoding method provided by the embodiments of this application. The holographic video encoding method may include the following steps:

[0184] Step 301, obtain the video frame sequence header corresponding to the holographic video;

[0185] Step 302, determine the video frame encoding order corresponding to the holographic video according to the video frame sequence header;

[0186] Step 303, encode the video frame sequence header to obtain the sequence header binary bitstream data corresponding to the video frame sequence header;

[0187] Step 304, when the current frame is a key frame, encode the key frame header to generate the key frame header binary bitstream data;

[0188] Step 305, obtain the neural network parameters, position parameters, attribute parameters, and hash features corresponding to the target anchor points in the key frame;

[0189] Step 306: Encode the neural network parameters, position parameters, attribute parameters, and hash features corresponding to the target anchor points in the key frame respectively to obtain the key frame binary bitstream data corresponding to the key frame;

[0190] Step 307: When the current frame is a non-key frame, encode the non-key frame header to generate non-key frame header binary bitstream data;

[0191] Step 308: Obtain the neural network parameters and cached hash features corresponding to the target anchor points in the non-key frame, and encode the neural network parameters and cached hash features corresponding to the target anchor points in the non-key frame respectively to obtain the non-key frame binary bitstream data corresponding to the non-key frame;

[0192] Step 309: Determine the key frame binary bitstream data corresponding to each key frame and the non-key frame binary bitstream data corresponding to each non-key frame data in the holographic video;

[0193] Step 310: Generate a holographic video file corresponding to the holographic video according to the sequence header binary bitstream data, the key frame header binary bitstream data and the key frame binary bitstream data corresponding to each key frame, and the non-key frame header binary bitstream data and the non-key frame binary bitstream data corresponding to each non-key frame data.

[0194] In the above embodiments, the descriptions of the various embodiments have their own emphases. For the parts not detailed in a certain embodiment, reference may be made to the detailed description of the holographic video encoding method above, and details will not be repeated here.

[0195] Please refer to Figure 5 , Figure 5 which is a schematic structural diagram of a holographic video encoding device provided by an embodiment of the present application. This holographic video encoding device is used to execute the above holographic video encoding method.

[0196] In an embodiment of the present application, the term "module" or "unit" refers to a computer program or a part of a computer program with a predetermined function, which works together with other related parts to achieve a predetermined goal, and can be fully or partially implemented by using software, hardware (such as a processing circuit or a memory), or a combination thereof. Similarly, one processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of the overall module or unit including the function of the module or unit.

[0197] The holographic video encoding device 400 includes:

[0198] A determination module 410, configured to determine the video frame encoding order corresponding to the holographic video, and determine the current frame to be encoded in the holographic video according to the video frame encoding order;

[0199] An acquisition module 420, configured to acquire neural network parameters, position parameters, attribute parameters, and hash features corresponding to a target anchor point in a key frame when the current frame is a key frame;

[0200] A first encoding module 430, configured to encode the neural network parameters, position parameters, attribute parameters, and hash features corresponding to the target anchor point in the key frame respectively to obtain key frame binary bitstream data corresponding to the key frame;

[0201] A second encoding module 440, configured to acquire neural network parameters and cached hash features corresponding to the target anchor point in a non-key frame when the current frame is a non-key frame, and encode the neural network parameters and cached hash features corresponding to the target anchor point in the non-key frame respectively to obtain non-key frame binary bitstream data corresponding to the non-key frame;

[0202] A generation module 450, configured to determine key frame binary bitstream data corresponding to each key frame in the holographic video and non-key frame binary bitstream data corresponding to each non-key frame data, and generate a holographic video file corresponding to the holographic video according to the key frame binary bitstream data corresponding to each key frame and the non-key frame binary bitstream data corresponding to each non-key frame data.

[0203] In some embodiments, the first encoding module 430 is configured to:

[0204] Acquire the range and probability of attribute parameters corresponding to the target anchor point in the key frame;

[0205] Construct an entropy model corresponding to the attribute parameters according to the range and probability of the attribute parameters;

[0206] Encode the attribute parameters corresponding to the target anchor point in the key frame according to the entropy model to obtain attribute binary bitstream sub-data, and the key frame binary bitstream data includes the attribute binary bitstream sub-data.

[0207] In some embodiments, the first encoding module 430 is configured to:

[0208] Determine the grid network corresponding to the target anchor point in the key frame;

[0209] Input the hash features corresponding to the target anchor point in the key frame into the grid network, and output the probability of the attribute parameters corresponding to the target anchor point in the key frame.

[0210] In some embodiments, the second encoding module 440 is configured to:

[0211] After acquiring the neural network parameters and cached hash features corresponding to the target anchor point in the non-key frame, acquire the historical position parameters and historical attribute parameters corresponding to the target anchor point in the previous frame of the non-key frame;

[0212] Determine the neural transformation cache network corresponding to the non-key frame according to the neural network parameters corresponding to the target anchor points in the non-key frame;

[0213] Input the cached hash feature, historical position parameter, and historical attribute parameter into the neural transformation cache network, and output the position residual and attribute residual corresponding to the target anchor point in the non-key frame;

[0214] Encode the position residual and attribute residual corresponding to the target anchor point in the non-key frame to obtain the residual binary code stream sub-data. The non-key frame binary code stream data includes the residual binary code stream sub-data, and the non-key frame binary code stream data contains the residual binary code stream sub-data.

[0215] In some embodiments, the holographic video encoding device further includes a position determination module 410, which is used for:

[0216] After inputting the cached hash feature, historical position parameter, and historical attribute parameter into the neural transformation cache network and outputting the position residual and attribute residual corresponding to the target anchor point in the non-key frame, generate the target position parameter corresponding to the target anchor point in the non-key frame according to the position residual and historical position parameter, and generate the target attribute parameter corresponding to the target anchor point in the non-key frame according to the attribute residual and historical attribute parameter;

[0217] Determine the next frame to be encoded of the holographic video. When the next frame to be encoded is a non-key frame, determine the cached hash feature of the next frame to be encoded;

[0218] Input the cached hash feature of the next frame to be encoded, the target position parameter and target attribute parameter corresponding to the target anchor point in the non-key frame into the neural transformation cache network corresponding to the next frame to be encoded, and output the position residual and attribute residual corresponding to the target anchor point in the next frame to be encoded.

[0219] In some embodiments, the determination module 410 is used for:

[0220] Obtain the video frame sequence header corresponding to the holographic video;

[0221] Determine the video frame encoding order corresponding to the holographic video according to the video frame sequence header.

[0222] In some embodiments, the determination module 410 is further used for:

[0223] After obtaining the video frame sequence header corresponding to the holographic video, encode the video frame sequence header to obtain the sequence header binary code stream data corresponding to the video frame sequence header.

[0224] In some embodiments, the generation module 450 is used for:

[0225] Determine the key frame headers of each key frame and the non-key frame headers of each non-key frame in the holographic video;

[0226] Encode the key frame headers to generate key frame header binary bitstream data, and encode the non-key frame headers to generate non-key frame header binary bitstream data;

[0227] Generate a holographic video file corresponding to the holographic video according to the sequence header binary bitstream data, the key frame header binary bitstream data, the non-key frame header binary bitstream data, the key frame binary bitstream data corresponding to each key frame, and the non-key frame binary bitstream data corresponding to each non-key frame data.

[0228] In the embodiment of the present application, the determination module 410 determines the video frame encoding order corresponding to the holographic video, and determines the current frame to be encoded in the holographic video according to the video frame encoding order; when the current frame is a key frame, the acquisition module 420 acquires the neural network parameters, position parameters, attribute parameters, and hash features corresponding to the target anchor points in the key frame; the first encoding module 430 encodes the neural network parameters, position parameters, attribute parameters, and hash features corresponding to the target anchor points in the key frame respectively to obtain the key frame binary bitstream data corresponding to the key frame; when the current frame is a non-key frame, the second encoding module 440 acquires the neural network parameters and cached hash features corresponding to the target anchor points in the non-key frame, and encodes the neural network parameters and cached hash features corresponding to the target anchor points in the non-key frame respectively to obtain the non-key frame binary bitstream data corresponding to the non-key frame; the generation module 450 determines the key frame binary bitstream data corresponding to each key frame and the non-key frame binary bitstream data corresponding to each non-key frame data in the holographic video, and generates a holographic video file corresponding to the holographic video according to the key frame binary bitstream data corresponding to each key frame and the non-key frame binary bitstream data corresponding to each non-key frame data.

[0229] Therefore, in the holographic video, according to the encoding order of video frames, the neural network parameters, position parameters, attribute parameters, and hash features corresponding to the target anchor points in the key frames of the holographic video are encoded respectively to obtain the key frame binary bitstream data corresponding to the key frames. The neural network parameters and cached hash features corresponding to the target anchor points in the non-key frames are encoded respectively to obtain the non-key frame binary bitstream data corresponding to the non-key frames. Finally, a holographic video file is generated according to the key frame binary bitstream data of each key frame and the non-key frame binary bitstream data of each non-key frame. Since the holographic video file is composed of data in binary form, the holographic video file has stronger compatibility than the three-dimensional video files in the related art. The playback device only needs to decode according to the corresponding decoding rules to obtain the information such as neural network parameters, position parameters, attribute parameters, and hash features contained in the key frame binary bitstream data. Subsequently, the rendering information of the key frames of the holographic video, such as color information and motion information, can be determined through these. The information such as neural network parameters and cached hash features contained in each non-key frame binary bitstream data can be obtained. Subsequently, the rendering information of the non-key frames of the holographic video, such as color information and motion information, can be determined through these information and in combination with the relevant information of the reference frames. The playback device can play the holographic video according to the rendering information. Therefore, in this application, it is possible to encode the relevant parameters of the holographic video to generate a holographic video file in the form of a binary bitstream. The holographic video file has good compatibility and can be applied to various devices and platforms.

[0230] An embodiment of the present application also provides a computer device, which includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the above-mentioned holographic video encoding method is implemented. The computer device can be any intelligent terminal including a tablet computer, an in-vehicle computer, etc.

[0231] Please refer to Figure 6 , Figure 6 which schematically shows the hardware structure of a computer device in another embodiment. The computer device includes:

[0232] A processor 501, which can be implemented in ways such as a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided by the embodiments of the present application;

[0233] The memory 502 can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM), etc. The memory 502 can store an operating system and other application programs. When implementing the technical solutions provided in the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 502 and are called by the processor 501 to execute the holographic video encoding method of the embodiments of this application;

[0234] The input / output interface 503 is used to implement information input and output;

[0235] The communication interface 504 is used to implement communication and interaction between this device and other devices. Communication can be achieved through wired means (such as USB, network cable, etc.) or through wireless means (such as mobile network, WIFI, Bluetooth, etc.);

[0236] The bus 505 transmits information between various components of the device (such as the processor 501, the memory 502, the input / output interface 503, and the communication interface 504);

[0237] Among them, the processor 501, the memory 502, the input / output interface 503, and the communication interface 504 are communicatively connected to each other inside the device through the bus 505.

[0238] The embodiments of this application also provide a computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the above-mentioned holographic video encoding method is implemented.

[0239] As a non-transitory computer-readable storage medium, the memory can be used to store non-transitory software programs and non-transitory computer-executable programs. In addition, the memory can include high-speed random access memory and can also include non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory can optionally include a memory remotely located relative to the processor, and these remote memories can be connected to the processor through a network. Examples of the above-mentioned network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.

[0240] The embodiments of the present application provide a holographic video encoding method, apparatus, computer device, and storage medium. The method determines the encoding order of video frames corresponding to a holographic video, and determines the current frame to be encoded in the holographic video according to the video frame encoding order; when the current frame is a key frame, obtains the neural network parameters, position parameters, attribute parameters, and hash features corresponding to the target anchor points in the key frame; encodes the neural network parameters, position parameters, attribute parameters, and hash features corresponding to the target anchor points in the key frame respectively to obtain the key frame binary bitstream data corresponding to the key frame; when the current frame is a non-key frame, obtains the neural network parameters and cached hash features corresponding to the target anchor points in the non-key frame, and encodes the neural network parameters and cached hash features corresponding to the target anchor points in the non-key frame respectively to obtain the non-key frame binary bitstream data corresponding to the non-key frame; determines the key frame binary bitstream data corresponding to each key frame and the non-key frame binary bitstream data corresponding to each non-key frame data in the holographic video, and generates a holographic video file corresponding to the holographic video according to the key frame binary bitstream data corresponding to each key frame and the non-key frame binary bitstream data corresponding to each non-key frame data.

[0241] Thus, in a holographic video, the neural network parameters, position parameters, attribute parameters, and hash features corresponding to the target anchor points in the key frame of the holographic video are encoded respectively according to the video frame encoding order to obtain the key frame binary bitstream data corresponding to the key frame, and the neural network parameters and cached hash features corresponding to the target anchor points in the non-key frame are encoded respectively to obtain the non-key frame binary bitstream data corresponding to the non-key frame. Finally, a holographic video file is generated according to the key frame binary bitstream data of each key frame and the non-key frame binary bitstream data of each non-key frame. Since the holographic video file is composed of data in binary form, the holographic video file has stronger compatibility than the three-dimensional video files in the related art. The playback device only needs to decode according to the corresponding decoding rules to obtain the information such as neural network parameters, position parameters, attribute parameters, and hash features contained in the key frame binary bitstream data. Subsequently, the rendering information of the key frames of the holographic video, such as color information and motion information, can be determined through these. The information such as neural network parameters and cached hash features is obtained from each non-key frame binary bitstream data. Subsequently, the rendering information of the non-key frames of the holographic video, such as color information and motion information, can be determined through this information and in combination with the relevant information of the reference frames. The playback device can play the holographic video according to the rendering information. Therefore, in the present application, it is possible to encode the relevant parameters of the holographic video to generate a holographic video file in the form of a binary bitstream. The holographic video file has good compatibility and can be applied to various devices and platforms.

[0242] The embodiments described in the embodiments of the present application are for more clearly illustrating the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions provided by the embodiments of the present application. Those skilled in the art can know that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of the present application are equally applicable to similar technical problems.

[0243] Those skilled in the art can understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present application, and may include more or fewer steps than shown in the figures, or combine certain steps, or different steps.

[0244] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0245] Those of ordinary skill in the art can understand that all or some of the steps in the methods disclosed above, and the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, and their appropriate combinations.

[0246] The terms "first", "second", "third", "fourth", etc. (if any) in the specification of the present application and the above drawings are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present application described here can be implemented in an order other than those illustrated or described here. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.

[0247] It should be understood that in this application, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated objects and indicates that three relationships can exist. For example, "A and / or B" can mean: only A exists, only B exists, and both A and B exist at the same time. Here, A and B can be singular or plural. The character " / " generally indicates an "or" relationship between the associated objects before and after. "At least one (one)" or a similar expression below refers to any combination of these items, including any combination of single items (one) or plural items (ones). For example, at least one (one) of a, b, or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.

[0248] In several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the above division of units is only a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of devices or units can be in electrical, mechanical, or other forms.

[0249] The units described above as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0250] In addition, each functional unit in various embodiments of this application can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.

[0251] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods of various embodiments of this application. The aforementioned storage medium includes: various media that can store programs, such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs.

[0252] The preferred embodiments of the embodiments of this application have been described above with reference to the accompanying drawings, which does not limit the scope of rights of the embodiments of this application. Any modifications, equivalent replacements, and improvements made by those skilled in the art without departing from the scope and essence of the embodiments of this application shall be within the scope of rights of the embodiments of this application.

Claims

1. A holographic video encoding method, characterized in that: include: Determine a video frame coding sequence corresponding to the holographic video, and determine a current frame to be encoded in the holographic video according to the video frame coding sequence; When the current frame is a key frame, obtaining neural network parameters, position parameters, attribute parameters and hash features corresponding to the target anchor point in the key frame; Encoding the neural network parameters, position parameters, attribute parameters and hash features corresponding to the target anchor point in the key frame respectively to obtain key frame binary code stream data corresponding to the key frame; When the current frame is a non-key frame, obtaining a neural network parameter and a cache hash feature corresponding to a target anchor point in the non-key frame, and respectively encoding the neural network parameter and the cache hash feature corresponding to the target anchor point in the non-key frame to obtain non-key frame binary code stream data corresponding to the non-key frame; Determine the key frame binary code stream data corresponding to each key frame in the holographic video and the non-key frame binary code stream data corresponding to each non-key frame data, and generate a holographic video file corresponding to the holographic video according to the key frame binary code stream data corresponding to each key frame and the non-key frame binary code stream data corresponding to each non-key frame data.

2. The holographic video encoding method according to claim 1, characterized in that: The encoding of the attribute parameters corresponding to the target anchor point in the key frame includes: Obtaining the attribute parameter range and attribute parameter probability corresponding to the target anchor point in the key frame; Constructing an entropy model corresponding to the attribute parameter according to the attribute parameter range and the attribute parameter probability; The attribute parameters corresponding to the target anchor point in the key frame are encoded according to the entropy model to obtain attribute binary code stream sub-data, and the key frame binary code stream data includes the attribute binary code stream sub-data.

3. The holographic video encoding method according to claim 2, characterized in that: The obtaining the probability of the attribute parameter corresponding to the target anchor point in the key frame includes: Determine a mesh network corresponding to a target anchor point in the key frame; The hash feature corresponding to the target anchor point in the key frame is input into the grid network, and the attribute parameter probability corresponding to the target anchor point in the key frame is output.

4. The holographic video encoding method according to claim 1, characterized in that: After obtaining the neural network parameters and cached hash features corresponding to the target anchor point in the non-key frame, the method further includes: Obtaining historical position parameters and historical attribute parameters corresponding to the target anchor point in the non-key frame in the previous frame; Determine a neural transform cache network corresponding to the non-key frame according to a neural network parameter corresponding to a target anchor point in the non-key frame; Inputting the cache hash feature, the historical position parameter and the historical attribute parameter into the neural transform cache network, and outputting the position residual and attribute residual corresponding to the target anchor point in the non-key frame; The position residual and the attribute residual corresponding to the target anchor point in the non-key frame are encoded to obtain residual binary code stream sub-data, wherein the non-key frame binary code stream data includes the residual binary code stream sub-data, and the non-key frame binary code stream data contains the residual binary code stream sub-data.

5. The holographic video encoding method according to claim 4, characterized in that: After inputting the cache hash feature, the historical position parameter and the historical attribute parameter into the neural transform cache network and outputting the position residual and attribute residual corresponding to the target anchor point in the non-key frame, the method further includes: Generating a target position parameter corresponding to the target anchor point in the non-key frame according to the position residual and the historical position parameter, and generating a target attribute parameter corresponding to the target anchor point in the non-key frame according to the attribute residual and the historical attribute parameter; Determine a next frame to be encoded of the holographic video, and when the next frame to be encoded is a non-key frame, determine a cache hash feature of the next frame to be encoded; The cached hash features of the next frame to be encoded, the target position parameters and target attribute parameters corresponding to the target anchor point in the non-key frame are input into the neural transformation cache network corresponding to the next frame to be encoded, and the position residual and attribute residual corresponding to the target anchor point in the next frame to be encoded are output.

6. The holographic video encoding method according to claim 1, characterized in that: The step of determining the video frame coding sequence corresponding to the holographic video includes: Obtaining a video frame sequence header corresponding to the holographic video; The video frame encoding order corresponding to the holographic video is determined according to the video frame sequence header.

7. The holographic video encoding method according to claim 6, characterized in that: After obtaining the video frame sequence header corresponding to the holographic video, the method further includes: Encoding the video frame sequence header to obtain sequence header binary code stream data corresponding to the video frame sequence header; The generating of the holographic video file corresponding to the holographic video according to the key frame binary code stream data corresponding to each key frame and the non-key frame binary code stream data corresponding to each non-key frame data comprises: Determine a key frame header of each key frame and a non-key frame header of each non-key frame in the holographic video; Encoding the key frame header to generate key frame header binary code stream data, and encoding the non-key frame header to generate non-key frame header binary code stream data; A holographic video file corresponding to the holographic video is generated according to the sequence header binary code stream data, the key frame header binary code stream data, the non-key frame header binary code stream data, the key frame binary code stream data corresponding to each key frame, and the non-key frame binary code stream data corresponding to each non-key frame data.

8. A holographic video encoding device, characterized in that: include: A determination module, used to determine a video frame coding sequence corresponding to a holographic video, and determine a current frame to be encoded in the holographic video according to the video frame coding sequence; An acquisition module, used for acquiring neural network parameters, position parameters, attribute parameters and hash features corresponding to a target anchor point in the key frame when the current frame is a key frame; A first encoding module is used to encode the neural network parameters, position parameters, attribute parameters and hash features corresponding to the target anchor point in the key frame, respectively, to obtain key frame binary code stream data corresponding to the key frame; A second encoding module is used for, when the current frame is a non-key frame, obtaining the neural network parameters and cache hash features corresponding to the target anchor point in the non-key frame, and encoding the neural network parameters and cache hash features corresponding to the target anchor point in the non-key frame respectively to obtain non-key frame binary code stream data corresponding to the non-key frame; A generation module is used to determine the key frame binary code stream data corresponding to each key frame in the holographic video and the non-key frame binary code stream data corresponding to each non-key frame data, and generate a holographic video file corresponding to the holographic video according to the key frame binary code stream data corresponding to each key frame and the non-key frame binary code stream data corresponding to each non-key frame data.

9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a plurality of instructions, and the instructions are suitable for being loaded by a processor to execute the holographic video encoding method according to any one of claims 1 to 7.

10. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the holographic video encoding method according to any one of claims 1 to 7 is implemented.

Citation Information

Cited By

  • Holographic video playback method and apparatus, computer device, and storage medium

    WO2026144325A1

  • Holographic video coding method and apparatus, and computer device and storage medium

    WO2026144401A1