A three-dimensional gaussian splash multi-view video joint semantic coding method
By employing a 3D Gaussian splash multi-view video joint semantic coding method, cross-view parallel semantic feature autoencoder is used to reduce redundancy between viewpoints. Combined with multi-view video image coding and feedforward 3DGS technology, real-time efficient transmission and high-quality rendering of immersive semantic communication are achieved.
Patent Information
- Application Number
- CN202511141476.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-15
- Publication Date
- 2025-11-11
- Estimated Expiration
- 2045-08-15
AI Technical Summary
Existing traditional coding methods struggle to support real-time transmission of multi-view high-definition video images in immersive semantic communication scenarios, especially in balancing coding complexity and compression rate. Furthermore, insufficient redundant compression between multiple views affects the real-time performance of communication.
A three-dimensional Gaussian splash multi-view video joint semantic coding method is adopted. The semantic features of the left and right views are encoded by a cross-view parallel semantic feature autoencoder to transmit cross-view context information, reduce redundancy between views, and perform lightweight prediction Gaussian parameter rendering of user view image at the receiving end.
It improves the real-time performance of video image encoding and reduces its complexity, meeting the real-time requirements of immersive semantic communication and enabling efficient transmission and high-quality rendering of multi-view video images.
Smart Images

Figure CN120769035B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer vision processing technology, and in particular to a joint semantic coding method for three-dimensional Gaussian splash multi-view video. Background Technology
[0002] Immersive semantic communication technology is a core application scenario of 6G, which will completely change the way people work, entertain themselves and communicate through realistic three-dimensional visual presentation and interaction. However, the requirements of ultra-high data transmission rate and ultra-low latency of immersive semantic communication have brought challenges to 6G network and data processing technology. This has prompted researchers to develop corresponding high-performance and low-latency video image coding technology for immersive semantic communication.
[0003] The relevant encoding and decoding methods are usually traditional encoding frameworks based on search, transform and entropy coding. The improvement of the encoding performance of traditional encoding methods depends on the continuous increase of encoding modes. However, the more encoding modes there are, the more complex the rate-distortion optimization search will be. The benefit ratio of the increase in compression ratio to the increase in encoding complexity becomes smaller and smaller. Especially in the immersive semantic communication scenario, the current traditional encoding methods are difficult to support the real-time transmission of panoramic or multi-view high-definition video images with huge data volume. Summary of the Invention
[0004] This application provides a joint semantic coding method for three-dimensional Gaussian splash multi-view video, which reduces redundancy between viewpoints by transmitting cross-view context information, and helps to achieve immersive semantic communication with low bandwidth requirements.
[0005] The first aspect of this application provides a joint semantic coding method for three-dimensional Gaussian splash multi-view video, the method comprising:
[0006] In the semantic communication system, the transmitting end performs left and right view semantic feature encoding in parallel for the left and right view images through a cross-view parallel semantic feature autoencoder to obtain the left and right view latent representations. The cross-view parallel semantic feature autoencoder reduces redundancy between viewpoints by transmitting cross-view context information.
[0007] The sending end sends the potential representations of the left and right perspectives to the receiving end in the semantic communication system;
[0008] The receiving end decodes the received left and right view latent representations to obtain decoded semantic features and reconstructed left and right view images. Based on the decoded semantic features and reconstructed left and right view images, it predicts Gaussian parameters and renders the user's view image based on the predicted Gaussian parameters.
[0009] Optionally, the cross-view parallel semantic feature autoencoder includes at least: a cross-view parallel disparity semantic feature autoencoder;
[0010] In a semantic communication system, the transmitting end uses a cross-view parallel semantic feature autoencoder to encode left and right view semantic features in parallel for both left and right view images, obtaining left and right view latent representations, which include at least:
[0011] The transmitting end uses a disparity estimation module to estimate the disparity of the left and right view images of the current frame, thereby obtaining the left and right disparity of the current frame and the semantic features of the left and right disparity of the current frame.
[0012] The transmitting end reads the left and right parallax semantic features decoded by the transmitting end of the previous frame from the buffer;
[0013] The transmitting end uses the cross-view parallel disparity semantic feature autoencoder to encode the left and right view disparity semantic features of the current frame in parallel, based on the left and right disparity semantic features of the current frame, the left and right disparity semantic features decoded by the transmitting end of the previous frame, and the left and right disparity context of the current frame, to obtain the potential representation of the left and right view disparity of the current frame.
[0014] The method further includes:
[0015] The sending end decodes the left and right view disparity latent representation of the current frame through the first cross-view parallel disparity semantic feature self-decoder to obtain the sending end decoded left and right disparities of the current frame and the sending end decoded left and right disparity semantic features of the current frame.
[0016] The transmitting end caches the left and right disparity semantic features decoded by the transmitting end of the current frame into the buffer for use in encoding the left and right view disparity semantic features of the next frame. The transmitting end also caches the left and right disparity of the current frame decoded by the transmitting end into the buffer for use in encoding the left and right view image semantic features of the next frame.
[0017] Optionally, the left and right disparity contexts of the current frame include: the left disparity context of the current frame and the right disparity context of the current frame;
[0018] The method further includes:
[0019] The transmitting end performs disparity compensation and semantic aggregation based on the left disparity of the current frame, the right disparity semantic features of the current frame, the left disparity semantic features of the current frame, and the semantic correlation of the transmitting end's left viewpoint of the current frame to obtain the left disparity context of the current frame; the semantic correlation of the transmitting end's left viewpoint of the current frame is determined based on the left disparity of the current frame and the transmitting end disparity from the right viewpoint to the left viewpoint of the current frame.
[0020] The transmitting end performs disparity compensation and semantic aggregation based on the right disparity of the current frame, the left disparity semantic features of the current frame, the right disparity semantic features of the current frame, and the right view semantic correlation of the transmitting end in the current frame to obtain the right disparity context of the current frame; the right view semantic correlation of the transmitting end in the current frame is determined based on the right disparity of the current frame and the transmitting end disparity from the left view to the right view of the current frame.
[0021] Optionally, the cross-view parallel semantic feature autoencoder further includes: a cross-view parallel image semantic feature autoencoder;
[0022] In the semantic communication system, the transmitting end uses a cross-view parallel semantic feature autoencoder to encode left and right view semantic features in parallel for both left and right view images, obtaining the left and right view latent representations. This also includes:
[0023] The sending end extracts features from the left and right view images of the current frame through the image feature extraction module to obtain the semantic features of the left and right view images of the current frame.
[0024] The transmitting end reads the semantic features of the left and right view images of the previous frame from the buffer;
[0025] The transmitting end uses the cross-view parallel image semantic feature autoencoder to encode the left and right view image semantic features of the current frame in parallel, based on the left and right view image semantic features of the current frame, the left and right view image semantic features decoded by the transmitting end of the previous frame, and the left and right view image context of the transmitting end of the current frame, to obtain the potential representation of the left and right view image of the current frame.
[0026] The method further includes:
[0027] The sending end decodes the latent representation of the left and right view images of the current frame through the first cross-view parallel image semantic feature self-decoder to obtain the sending end decoded left and right view images of the current frame and the sending end decoded left and right view image semantic features of the current frame.
[0028] The sending end caches the semantic features of the left and right view images decoded by the sending end in the current frame into the buffer for use in encoding the semantic features of the left and right view images in the next frame.
[0029] Optionally, the left and right view image contexts of the current frame sender include: the left view image context of the current frame sender and the right view image context of the current frame sender.
[0030] The method further includes:
[0031] The transmitting end performs disparity compensation and semantic aggregation based on the left disparity decoded by the transmitting end of the current frame, the semantic features of the right view image of the current frame, the semantic features of the left view image of the current frame, and the semantic correlation of the left view image of the transmitting end of the current frame, to obtain the left view image context of the transmitting end of the current frame.
[0032] The transmitting end performs disparity compensation and semantic aggregation based on the right disparity decoded by the transmitting end of the current frame, the semantic features of the left-view image of the current frame, the semantic features of the right-view image of the current frame, and the semantic correlation of the right-view image of the transmitting end of the current frame, to obtain the right-view image context of the transmitting end of the current frame.
[0033] Optionally, the receiving end decodes the received left and right view latent representations to obtain decoded semantic features and reconstructed left and right view images, including at least:
[0034] The receiving end decodes the received left and right view disparity latent representations through the second cross-view parallel disparity semantic feature self-decoder to obtain the receiver-decoded left and right disparities of the current frame and the receiver-decoded left and right disparity semantic features of the current frame.
[0035] The receiving end caches the left and right parallax of the current frame into the buffer for use in reconstructing the left and right view images of the next frame.
[0036] Optionally, the receiving end decodes the received left and right view latent representations to obtain decoded semantic features and reconstructed left and right view images, and further includes:
[0037] The receiving end decodes the latent representation of the received left and right view images to obtain the receiver-decoded left and right view images of the current frame and the semantic features of the receiver-decoded left and right view images of the current frame.
[0038] The receiving end reads the semantic features of the left and right view images of the previous frame from the buffer;
[0039] The receiving end uses the second cross-view parallel image semantic feature self-decoder to perform parallel reconstruction of the left and right view images of the current frame based on the left and right view image semantic features decoded by the receiving end of the current frame, the left and right view image semantic features decoded by the receiving end of the previous frame, and the left and right view image context of the receiving end of the current frame, to obtain the reconstructed left and right view images and the semantic features of the reconstructed left and right view images of the current frame.
[0040] The method further includes:
[0041] The receiving end caches the semantic features of the left and right view images decoded by the receiving end in the current frame into the buffer for use in the reconstruction of the left and right view images in the next frame.
[0042] Optionally, the receiving end left and right view image context of the current frame includes: the receiving end left view image context of the current frame and the receiving end right view image context of the current frame.
[0043] The method further includes:
[0044] The receiving end performs disparity compensation and semantic aggregation based on the left disparity decoded by the receiving end of the current frame, the semantic features of the right-view image decoded by the receiving end of the current frame, the semantic features of the left-view image decoded by the receiving end of the current frame, and the semantic correlation of the left-view image of the receiving end of the current frame, to obtain the left-view image context of the receiving end of the current frame; the semantic correlation of the left-view image of the receiving end of the current frame is determined based on the left disparity decoded by the receiving end of the current frame and the disparity from the right view to the left view of the current frame.
[0045] The receiving end performs disparity compensation and semantic aggregation based on the right disparity decoded by the receiving end of the current frame, the semantic features of the left-view image decoded by the receiving end of the current frame, the semantic features of the right-view image decoded by the receiving end of the current frame, and the semantic correlation of the right-view image of the receiving end of the current frame, to obtain the right-view image context of the receiving end of the current frame; the semantic correlation of the right-view image of the receiving end of the current frame is determined based on the right disparity decoded by the receiving end of the current frame and the disparity from the left view to the right view of the current frame.
[0046] Optionally, the receiving end predicts Gaussian parameters based on the decoded semantic features and the reconstructed left and right view images, including:
[0047] The receiving end decodes the left and right parallax of the current frame and predicts the center position.
[0048] The receiving end reconstructs the image based on the left and right perspectives of the current frame and predicts the color of the Gaussian ellipsoid.
[0049] The receiving end decodes the left and right parallax semantic features of the current frame and reconstructs the image semantic features from the left and right perspectives of the current frame to predict the covariance matrix and opacity.
[0050] Optionally, the sending end sends the left and right view latent representations to the receiving end in the semantic communication system, including:
[0051] The transmitting end quantizes the potential representation of the left and right views to obtain the quantized potential representation of the left and right views.
[0052] The transmitting end estimates the distribution parameters of the left and right view potential representations through a super-prior network based on the left and right view potential representations.
[0053] The transmitting end performs entropy encoding on the quantized left and right view potential representations according to the distribution parameters to obtain a bit stream;
[0054] The sending end sends the bit stream and the distribution parameters to the receiving end;
[0055] The receiving end processes the received bit stream according to the distribution parameters to obtain the received left and right view potential representations.
[0056] Based on the same inventive concept, a second aspect of this application provides a three-dimensional Gaussian splash multi-view video joint semantic coding apparatus, the apparatus comprising:
[0057] The compression module is used in the semantic communication system to perform left and right view semantic feature encoding on the left and right view images in parallel through a cross-view parallel semantic feature autoencoder to obtain the left and right view potential representations. The cross-view parallel semantic feature autoencoder reduces the redundancy between views by transmitting cross-view context information.
[0058] A transmission module is used by the sending end to send the potential representations of the left and right perspectives to the receiving end in the semantic communication system;
[0059] The reconstruction module is used by the receiving end to decode the received left and right view latent representations to obtain decoded semantic features and reconstructed left and right view images. Based on the decoded semantic features and reconstructed left and right view images, Gaussian parameters are predicted, and the user's view image is rendered based on the predicted Gaussian parameters.
[0060] Based on the same inventive concept, a third aspect of the present application provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing, implements the three-dimensional Gaussian splash multi-view video joint semantic coding method as proposed in the first aspect of the present application.
[0061] Compared with the prior art, this application has the following advantages:
[0062] This application provides a three-dimensional Gaussian splash multi-view video joint semantic coding method. In the semantic communication system, the transmitting end performs left and right view semantic feature encoding in parallel for the left and right view images through a cross-view parallel semantic feature autoencoder to obtain the left and right view latent representations. The cross-view parallel semantic feature autoencoder reduces the redundancy between views by transmitting cross-view context information. The transmitting end sends the left and right view latent representations to the receiving end in the semantic communication system. The receiving end decodes the received left and right view latent representations to obtain decoded semantic features and reconstructed left and right view images. Based on the decoded semantic features and reconstructed left and right view images, Gaussian parameters are predicted, and the user's view image is rendered based on the predicted Gaussian parameters.
[0063] Therefore, this scheme combines multi-view video image encoding and feedforward 3DGS technology, employing a parallel architecture of dual-view encoding and Gaussian parameter prediction, with shared parameters between the left and right view branch models, improving real-time processing and reducing complexity. Simultaneously, a cross-view parallel semantic feature autoencoder is designed to achieve information interaction between the left and right view branches by transmitting cross-view context information, thereby reducing redundancy between viewpoints. Finally, at the receiving end, lightweight 3DGS prediction is performed on the decoded features to render the video image from the user's perspective, meeting the real-time requirements of immersive semantic communication. Attached Figure Description
[0064] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the description of the embodiments of the present invention will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0065] Figure 1 This is a flowchart of a three-dimensional Gaussian splash multi-view video joint semantic coding method in one embodiment of this application;
[0066] Figure 2 This is a schematic diagram of the parallax coding architecture of the transmitting end in one embodiment of this application;
[0067] Figure 3 This is a schematic diagram of the image encoding architecture of the transmitting end in one embodiment of this application;
[0068] Figure 4 This is a schematic diagram of the architecture for transmitting the potential representation of left and right view disparity in one embodiment of this application;
[0069] Figure 5 This is a schematic diagram of the architecture for Gaussian parameter prediction in one embodiment of this application;
[0070] Figure 6This is a schematic diagram of the functional modules of a three-dimensional Gaussian splash multi-view video joint semantic coding device according to an embodiment of this application;
[0071] Figure 7 This is a schematic diagram of the structure of an electronic device according to one embodiment of this application. Detailed Implementation
[0072] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0073] Immersive semantic communication technology is a core application scenario of 6G, which will completely change the way people work, entertain themselves and communicate through realistic three-dimensional visual presentation and interaction. However, the requirements of ultra-high data transmission rate and ultra-low latency of immersive semantic communication have brought challenges to 6G network and data processing technology. This has prompted researchers to develop corresponding high-performance and low-latency video image coding technology for immersive semantic communication.
[0074] Deep learning-based video coding typically combines techniques such as motion estimation, motion compensation, nonlinear transformation, and entropy coding. For example, DVC utilizes these techniques to build the first end-to-end neural network-based video residual coding framework, predicting the current frame through optical flow estimation and motion compensation, and encoding optical flow and residuals. However, existing deep learning-based video coding models struggle to achieve a balance between coding speed and rate-distortion performance. Furthermore, immersive applications require constructing a 3D scene representation and rendering the image seen by the user from their perspective. Among related technologies, 3D reconstruction from multi-view images is employed, but due to imperfect image acquisition and noise affecting camera parameter estimation, traditional 3D reconstruction methods yield unsatisfactory results. 3DGS (3D Gaussian Splatting) technology proposes rasterizing a set of Gaussian ellipsoids to approximate the appearance of a 3D scene, achieving high-quality synthesis of new perspective views and allowing for fast convergence and real-time rendering at 1080p resolution (approximately 30 FPS), making low-cost 3D content creation and real-time applications possible.
[0075] However, in immersive semantic communication scenarios, video image encoding and reconstruction still face the following problems: 1. General 3DGS-based scene reconstruction requires separate training of 3D representations for different scenes and objects, resulting in insufficient real-time performance and generalization. Especially for dynamic 3D scenes, a single frame contains nearly a million Gaussian points, resulting in a massive amount of data and making real-time encoding and compression difficult. 2. Schemes that reconstruct and encode 3DGS at the sending end: After reconstructing 3DGS using a feedforward method, each pixel will have one or more Gaussian points, which increases redundancy compared to video. Furthermore, the structure of Gaussian points in 3D space is more complex, leading to compression difficulties. 3. Schemes that use video encoding and transmission followed by decoding and 3DGS reconstruction at the receiving end: The parallax of multi-view video images acquired at the sending end is large. Existing methods do not adequately compress the redundancy between multiple views, resulting in a large computational load at the receiving end and affecting the real-time performance of communication.
[0076] To address the aforementioned issues, this application proposes a joint semantic coding method for 3D Gaussian splash multi-view video. This method combines multi-view video image coding with feedforward 3DGS technology, employing a parallel architecture of dual-view coding and Gaussian parameter prediction. The left and right view branch models share parameters, improving real-time processing and reducing complexity. Simultaneously, a cross-view parallel semantic feature autoencoder is designed to facilitate information interaction between the left and right view branches by transmitting cross-view context information, thereby reducing redundancy between viewpoints. Finally, at the receiving end, lightweight 3DGS prediction is performed on the decoded features to render the video image from the user's perspective, meeting the real-time requirements of immersive semantic communication. The following detailed description, in conjunction with the accompanying drawings, through some embodiments and application scenarios, illustrates the joint semantic coding method for 3D Gaussian splash multi-view video provided by this application.
[0077] The first aspect of this application provides a joint semantic coding method for three-dimensional Gaussian splash multi-view video. The method is described below through Section 1.1 Method Overview, Section 1.2 Transmitter Encoding, Section 1.3 Transmission Process, Section 1.4 Receiver Decoding, and Section 1.5 3DGS Rendering Process.
[0078] 1.1 A brief overview of the joint semantic coding method for 3D Gaussian splash multi-view video: such as Figure 1 As shown, the method includes the following steps:
[0079] S101: In the semantic communication system, the sending end performs left and right view semantic feature encoding in parallel for the left and right view images through a cross-view parallel semantic feature autoencoder to obtain the left and right view potential representations. The cross-view parallel semantic feature autoencoder reduces the redundancy between views by transmitting cross-view context information.
[0080] In this embodiment, the semantic communication system includes a transmitter and a receiver. The transmitter extracts features from the input left and right dual-view images through a cross-view parallel semantic feature autoencoder to obtain the semantic content in the dual-view images, and encodes and expresses the semantic information (i.e., highly abstracts and compresses the semantic concepts in the dual-view images) to obtain the latent representations of the left and right views.
[0081] The receiving end is used to perform semantic decoding on the received left and right view potential representations, which is the reverse process of encoding. In this way, the semantic information of the dual-view images can be restored through the decoding process, and then the left and right view images can be reconstructed and rendered from the user's perspective to achieve immersive semantic communication.
[0082] The input data for the cross-view parallel semantic feature autoencoder at the transmitting end consists of left-view and right-view images, which together form a stereo vision pair, helping to obtain more comprehensive information about the user communication scene. During the encoding process, the cross-view parallel semantic feature autoencoder encodes the semantic features of the left and right view images in parallel to transmit cross-view contextual information and reduce redundancy between viewpoints.
[0083] For example, a cross-view parallel semantic feature autoencoder can employ a two-branch neural network (such as a dual-path CNN with shared weights) to process the left-view image separately. and right-angle image Each branch extracts hierarchical features (such as edges, textures, and object semantics) through multiple deep separable convolutions, and performs cross-view feature interaction and fusion through contextual features to ultimately generate latent representations for the left and right views. This preserves the unique information of each view while sharing semantics across views, reducing redundant information between the left and right views.
[0084] S102: The sender sends the potential representations of the left and right perspectives to the receiver in the semantic communication system.
[0085] S103: The receiver decodes the received left and right view latent representations to obtain decoded semantic features and reconstructed left and right view images. Based on the decoded semantic features and reconstructed left and right view images, it predicts Gaussian parameters and renders the user's view image based on the predicted Gaussian parameters.
[0086] In this embodiment, the sending end encodes the semantic features of the left and right perspectives through a cross-view parallel semantic feature autoencoder, obtains the latent representations of the left and right perspectives, and then sends the latent representations of the left and right perspectives to the receiving end in the semantic communication system.
[0087] After receiving the latent representations of the left and right perspectives, the receiving end performs decoding processing to obtain decoded semantic features, and further reconstructs the left and right perspective images based on the decoded semantic features (see Section 1.3 below for details). In this embodiment, it is assumed that the reconstructed left and right perspective images by the receiving end are approximately consistent with the original left and right perspective images by the sending end. However, it is easy to understand that in actual applications, due to the loss in the semantic feature extraction and compression encoding process, there may be slight differences between the reconstructed left and right perspective images by the receiving end and the original left and right perspective images by the sending end.
[0088] Next, 3DGS (3D Gaussian Splashing) is used for 3D scene rendering. 3DGS transforms each point or object in the scene into a Gaussian distribution with parameters such as position, shape, and color. It dynamically adjusts these parameters to fit the complex shapes of real-world objects (such as the curved surfaces of buildings or human hair). Utilizing GPU-accelerated differentiable rasterization technology, it employs a "snowball" algorithm to avoid jagged edges when projecting the Gaussian distribution onto a 2D screen, while also supporting dynamic optimization of the position, shape, and transparency of the Gaussian spots. Thus, complex 3D scenes are flexibly represented through discrete Gaussian distributions, achieving efficient real-time rendering.
[0089] In this embodiment, the receiving end first predicts Gaussian parameters, including the position, color, and opacity of the Gaussian ellipsoid, based on the decoded semantic features and the reconstructed left and right view images. Then, it renders the user's view image based on the predicted Gaussian parameters. That is, Gaussian points in three-dimensional space are projected onto a two-dimensional image plane, and these projected data points produce visual effects on the image in a certain way, thus appearing in the final rendered image. It is easy to understand that the user's view at the receiving end is often different from the viewpoint of the left and right view images acquired by the sending end. This embodiment performs left and right view semantic feature encoding in parallel at the sending end to transmit cross-view context information, thereby realizing image rendering of the user's new viewpoint at the receiving end.
[0090] For example, given an RGB video of a human-centered scene with sparse camera views, this scheme aims to compress two adjacent viewpoints, reduce transmission bandwidth requirements, and achieve real-time, high-quality rendering of new perspective views of the human body at the receiving end. Specifically, in the aforementioned semantic communication system, at each time point t, this scheme performs parallel encoding and decoding of left and right viewpoint video frames and prediction of Gaussian parameters to achieve immersive semantic communication.
[0091] This embodiment combines multi-view video image encoding and feedforward 3DGS technology, employing an architecture of parallel computation of dual-view encoding and Gaussian parameter prediction. The parameters of the left and right view branch models are fully shared to support parallel computing, improving real-time processing and reducing complexity. Simultaneously, a cross-view parallel semantic feature autoencoder is designed to achieve information interaction between the left and right view branches by transmitting cross-view context information, thereby reducing redundancy between views. Finally, at the receiving end, lightweight 3DGS prediction is performed on the decoded features to render the video image from the user's perspective, meeting the real-time requirements of immersive semantic communication.
[0092] It's easy to understand that the process of cross-view image encoding and compression at the sending end and image reconstruction at the receiving end in the aforementioned semantic communication system can be viewed as being completed by a single image reconstruction model. That is, the input to this image reconstruction model is the left and right view images, and the output is the reconstructed left and right view images. Some model parameters of this image reconstruction model, such as the network weights during feature extraction by the cross-view parallel semantic feature autoencoder, need to be obtained through training with a large number of training samples.
[0093] For example, training samples may include: original left and right view images and corresponding real reconstructed left and right view images. During training, the original left and right view images are input into an initial image reconstruction model for feature extraction and reconstruction prediction to obtain predicted reconstructed left and right view images. Then, based on the predicted and real reconstructed left and right view images, a loss function value is calculated, and the model parameters of the initial image reconstruction model are updated according to this loss function value. This process is repeated multiple times until the loss function converges or the preset number of training iterations is reached, at which point training ends, and the trained image reconstruction model is obtained.
[0094] 1.2 Encoding process at the sending end:
[0095] This scheme employs two cross-view parallel semantic feature autoencoders: a cross-view parallel disparity semantic feature autoencoder, used to encode the disparity of the left and right view images to obtain the left and right view disparity latent representations; and a cross-view parallel image semantic feature autoencoder, used to encode the left and right view images to obtain the left and right view image latent representations. The left and right view disparity latent representations and the left and right view image latent representations together constitute the left and right view latent representations, which are used for transmission. Internally, each of the two cross-view parallel semantic feature autoencoders contains two parallel autoencoders, each encoding either the disparity or the image of the left and right views respectively.
[0096] The following sections, 1.2.1 and 1.2.2, respectively, will explain the encoding process of disparity and images:
[0097] 1.2.1 Please refer to Figure 2 , Figure 2 This is a schematic diagram of the parallax coding architecture of the transmitting end in one embodiment of this application. Figure 2 As shown, the sending end includes a disparity estimation module, a cross-view parallel disparity semantic feature autoencoder, and a buffer. Specifically, the disparity encoding process at the sending end includes:
[0098] S201: The sending end uses the disparity estimation module to estimate the disparity of the left and right view images of the current frame, and obtains the left and right disparity and the semantic features of the left and right disparity of the current frame.
[0099] For horizontally aligned left and right views after stereo correction, the displacement of corresponding pixels between views is restricted to the horizontal direction. The predicted disparity map can be used to obtain the depth map of the view through linear transformation based on parameters such as camera focal length, baseline distance, and optical center deviation. The depth map can then be inversely projected onto the world coordinate system using pixel-by-pixel Gaussian points. Therefore, through explicit disparity estimation and disparity coding, the center position of the 3DGS can be directly obtained at the receiving end. Since this geometric property has a significant impact on rendering quality, transmitting disparity ensures that the accuracy of the 3DGS center position does not degrade with the decrease in decoded image quality.
[0100] Furthermore, disparity represents the pixel correspondence between views. Considering the large disparity of images in sparse views, explicit disparity estimation is beneficial for capturing the semantic correlation between viewpoints, thereby helping to compress redundancy between viewpoints. Therefore, in this embodiment, the disparity estimation module is based on the left and right view images of the current frame (i.e., Figure 2 In and The left and right disparities of the current frame are calculated by constructing a matching cost volume and updating it iteratively. Figure 2 In and This avoids slow 3D convolution computations. Then, features are extracted from the left and right disparities to obtain semantic features of the left and right disparities (i.e.,...). Figure 2 In and To further improve the real-time performance of the model, in this embodiment, the feature extraction of the current frame does not use the common method of downsampling through convolutional layers, but instead directly downsamples the current frame by 8 times and then uses a depthwise separable convolutional module to extract features.
[0101] Specifically, the process of the disparity estimation module can be represented as follows:
[0102]
[0103] in, Indicates the left disparity of the current frame. Indicates the right disparity of the current frame. This represents the feature extractor of the disparity estimation module. This indicates the disparity estimation module. This represents the left-view image of the current frame. This represents the right-view image of the current frame. This indicates the camera parameters from the left-hand perspective. This indicates the camera parameters from the right-hand perspective.
[0104] It should be noted that the left and right view images of the current frame refer to the left and right view images corresponding to the current time t, and the left and right view images of the previous frame refer to the left and right view images corresponding to the previous time t-1.
[0105] S202: The sending end reads the left and right parallax semantic features of the previous frame from the buffer.
[0106] like Figure 2 As shown, the buffer is mainly used to store the left and right disparity and left and right disparity semantic features of the video images processed during semantic communication.
[0107] In this embodiment, decoding the left and right disparity semantic features at the transmitting end means that the transmitting end decodes the potential representation of the left and right view disparity of the previous frame through the first cross-view parallel disparity semantic feature self-decoder to obtain the left and right disparity and left and right disparity semantic features of the previous frame.
[0108] Considering that scene changes between adjacent video frames are usually gradual (e.g., object movement, slow viewpoint changes), the disparity semantic features of the previous frame (e.g., object outlines, depth distribution) are strongly correlated with the current frame. Therefore, when encoding the left and right viewpoint disparity semantic features of the current frame, the transmitting end needs to read the decoded left and right disparity semantic features of the previous frame from the buffer (i.e., the features decoded by the transmitting end). Figure 2 In and This approach uses the disparity semantic features of the previous frame as contextual information for conditional coding, thereby leveraging the features of the previous frame as a priori and encoding only the residual (the changed part), reducing the bit rate. In other words, it improves coding efficiency and reconstruction quality through temporal coherence and disparity dynamic consistency.
[0109] S203: The transmitting end uses a cross-view parallel disparity semantic feature autoencoder to encode the left and right view disparity semantic features of the current frame in parallel, based on the left and right disparity semantic features of the current frame, the left and right disparity semantic features decoded by the transmitting end in the previous frame, and the left and right disparity context of the current frame, to obtain the potential representation of the left and right view disparity of the current frame.
[0110] In this embodiment, the left and right disparity context of the current frame refers to the disparity context information between the left and right view images generated through disparity compensation and semantic aggregation.
[0111] like Figure 2 As shown, when the sending end encodes the disparity of the current frame using a cross-view parallel disparity semantic feature autoencoder, it needs to combine three feature metrics: the left and right disparity semantic features of the current frame (i.e., Figure 2 In and The left and right parallax semantic features decoded at the sending end of the previous frame (i.e.) Figure 2 In and ), and the left and right disparity context of the current frame (i.e. Figure 2 In and Encode the left and right view disparity latent representations of the current frame to obtain the left and right view disparity latent representations (i.e., ... Figure 2 In and The data is transmitted to the receiving end for view reconstruction.
[0112] This embodiment encodes the left and right viewpoint parallax semantic features of the current frame in parallel based on the left and right parallax semantic features decoded by the sender of the previous frame and the left and right parallax context of the current frame, so as to realize parallax semantic interaction between the left and right views, and then effectively compress the redundancy between viewpoints through cross-view information flow.
[0113] In addition to deploying a cross-view parallel disparity semantic feature autoencoder, the transmitting end also deploys a first cross-view parallel disparity semantic feature autodecoder to decode the left and right view disparity latent representations of the current frame, thereby obtaining the transmitting end decoded left and right disparities of the current frame and the transmitting end decoded left and right disparity semantic features of the current frame.
[0114] Then, the sending end caches the left and right disparity semantic features decoded by the sending end in the current frame into a buffer for use in encoding the left and right view disparity semantic features of the next frame.
[0115] Considering that if the transmitting end only encodes the potential representation of left and right view disparity without local decoding, the decoding results between the receiving end and the transmitting end may differ due to quantization or noise (i.e., "drift error"), which will lead to a decline in reconstruction quality over a long period of time. Therefore, in this embodiment, a first cross-view parallel disparity semantic feature self-decoder is deployed at the transmitting end to decode the left and right disparities and left and right disparity semantic features of the current frame in real time, completely synchronized with the receiving end. Furthermore, the decoded left and right disparities and left and right disparity semantic features of the current frame are cached in a buffer area instead of caching the original encoded data, to ensure that the context used for encoding the next frame is consistent with that of the receiving end, forming a closed-loop feedback.
[0116] Specifically, the process of decoding the left and right disparity semantic features based on the current frame at the transmitting end and encoding the left and right view disparity semantic features for the next frame is similar to steps S201-S203 above, and will not be repeated here. For the process of decoding the left and right disparity based on the current frame at the transmitting end and encoding the left and right view image semantic features for the next frame, please refer to Section 1.2.2 below.
[0117] Furthermore, the left and right disparity contexts of the current frame used in disparity coding of the current frame in step S203 above include: the left disparity context and the right disparity context of the current frame. The calculation process of the left and right disparity contexts of the current frame is as follows:
[0118] S203-1: The transmitting end performs disparity compensation and semantic aggregation based on the left disparity of the current frame, the right disparity semantic features of the current frame, the left disparity semantic features of the current frame, and the semantic correlation of the transmitting end's left viewpoint of the current frame to obtain the left disparity context of the current frame; the semantic correlation of the transmitting end's left viewpoint of the current frame is determined based on the left disparity of the current frame and the transmitting end disparity from the right viewpoint to the left viewpoint of the current frame.
[0119] Similar to using motion compensation to model the correlation between consecutive frames in video coding, disparity reflects the pixel motion relationship between left and right views. Therefore, disparity compensation can be used to explicitly model the correlation between multi-view frames. For the left and right view images of the current frame, we have:
[0120]
[0121] in, This represents the prediction of the right-view image obtained by using the disparity compensation method on the left-view image. This represents a left-view image. This represents the right-view disparity. Considering the disparity from the left view to the right view and the disparity from the right view to the left view, they should be inversely related in the unobstructed condition. Therefore, disparity compensation can also be used to model the correlation between multi-view frames:
[0122]
[0123] in, This represents the predicted value of the right-view disparity generated based on the left-view disparity. Indicates left parallax. Indicates right parallax.
[0124] Furthermore, considering the occlusion problem, using only parallax compensation would introduce noise. Therefore, based on the physical meaning of parallax, the semantic correlation of parallax between viewpoints can be introduced:
[0125]
[0126] in, Indicates the semantic relevance from the left perspective of the sending end. This represents the learnable scaling factor. Indicates left parallax. Indicates right parallax.
[0127] Then, using semantic relevance Weighted semantic fusion is used to obtain the left disparity context of the current frame:
[0128]
[0129] in, The left disparity context of the current frame, Indicates the left disparity of the current frame. This represents the semantic features of the left disparity of the current frame. This represents the right disparity semantic features of the current frame. This indicates the semantic relevance from the left perspective of the sending end.
[0130] Therefore, the sending end performs disparity compensation and semantic aggregation based on the left disparity of the current frame, the right disparity semantic features of the current frame, the left disparity semantic features of the current frame, and the semantic correlation of the sending end's left viewpoint in the current frame, to obtain the left disparity context of the current frame.
[0131] S203-2: The transmitting end performs disparity compensation and semantic aggregation based on the right disparity of the current frame, the semantic features of the left disparity of the current frame, the semantic features of the right disparity of the current frame, and the semantic correlation of the right viewpoint of the transmitting end in the current frame, to obtain the right disparity context of the current frame; the semantic correlation of the right viewpoint of the transmitting end in the current frame is determined based on the right disparity of the current frame and the disparity of the transmitting end from the left viewpoint to the right viewpoint of the current frame.
[0132] In this embodiment, the calculation process of the right disparity context of the current frame is similar to that of the left disparity context. The calculation method of the right view semantic relevance at the sending end is as follows:
[0133]
[0134] in, Indicates the semantic relevance from the right perspective of the sender. This represents the learnable scaling factor. Indicates left parallax. Indicates right parallax.
[0135] Then, using semantic relevance Weighted semantic fusion is performed to obtain the right disparity context of the current frame:
[0136]
[0137] in, The right disparity context of the current frame, Indicates the right disparity of the current frame. This represents the semantic features of the left disparity of the current frame. This represents the right disparity semantic features of the current frame. This indicates the semantic relevance from the right perspective of the sending end.
[0138] To address parallax, this embodiment uses a feature-based parallax compensation and semantic aggregation method to generate contextual information between the left and right views, thereby enabling semantic interaction between views. This effectively eliminates noise introduced by occlusion in parallax compensation and further compresses redundancy between viewpoints through cross-view information flow, improving bandwidth transmission efficiency.
[0139] 1.2.2 Please refer to Figure 3 , Figure 3 This is a schematic diagram of the image encoding architecture of the transmitting end in one embodiment of this application. For example... Figure 3 As shown, the sending end includes an image feature extraction module, a cross-view parallel image semantic feature autoencoder, and a buffer. Specifically, the image encoding process at the sending end includes:
[0140] S301: The sending end extracts features from the left and right view images of the current frame through the image feature extraction module to obtain the semantic features of the left and right view images of the current frame.
[0141] In this embodiment, the sending end performs 8x downsampling on the left and right view images of the current frame through the image feature extraction module, and extracts image features through cascaded depthwise separable convolution to obtain the semantic features of the left and right view images of the current frame.
[0142] S302: The sending end reads the semantic features of the left and right view images from the buffer of the previous frame.
[0143] like Figure 3 As shown, in addition to storing the left and right disparity and semantic features of the video images processed during semantic communication, the buffer is also used to store the semantic features of the left and right view images.
[0144] In this embodiment, decoding the semantic features of the left and right view images at the transmitting end means that the transmitting end decodes the latent representation of the left and right view images of the previous frame through the first cross-view parallel image semantic feature self-decoder to obtain the semantic features of the left and right view images of the previous frame.
[0145] S303: The transmitting end uses a cross-view parallel image semantic feature autoencoder to encode the semantic features of the left and right view images of the current frame in parallel, based on the semantic features of the left and right view images of the current frame, the semantic features of the left and right view images decoded by the transmitting end of the previous frame, and the context of the left and right view images of the transmitting end of the current frame, to obtain the potential representation of the left and right view images of the current frame.
[0146] In this embodiment, the left and right view image context of the current frame refers to the image context information between the left and right view images generated through parallax compensation and semantic aggregation.
[0147] like Figure 3 As shown, when the sending end performs feature encoding on the current frame's image using a cross-view parallel image semantic feature autoencoder, it needs to combine three feature metrics: the semantic features of the left and right view images of the current frame (i.e., the semantic features of the current frame's left and right view images). Figure 3 In and The semantic features of the left and right view images decoded at the sending end of the previous frame (i.e., Figure 3 In and ), and the left and right view image context of the sending end of the current frame (i.e. Figure 3 In and Encode the images to obtain the left and right view latent representations of the current frame (i.e., ...). Figure 3 In and The data is transmitted to the receiving end for view reconstruction.
[0148] This embodiment encodes the semantic features of the left and right view images of the current frame in parallel based on the semantic features of the left and right view images decoded by the sender of the previous frame, and the context of the left and right view images of the sender of the current frame, so as to realize the semantic interaction of images between the left and right views, and then effectively compress the redundancy between views through cross-view information flow.
[0149] In addition to deploying a cross-view parallel image semantic feature autoencoder, the transmitting end also deploys a first cross-view parallel image semantic feature autodecoder to decode the latent representations of the left and right view images of the current frame, obtaining the transmitting end-decoded left and right view images and the transmitting end-decoded semantic features of the left and right view images of the current frame. Then, the transmitting end caches the transmitting end-decoded semantic features of the left and right view images of the current frame into a buffer for use in encoding the semantic features of the left and right view images of the next frame.
[0150] In this embodiment, a first cross-view parallel image semantic feature self-decoder is deployed at the transmitting end to decode the left and right view images and semantic features of the current frame in real time, achieving complete synchronization with the receiving end. Furthermore, the decoded left and right view images and semantic features of the current frame are cached in a buffer area instead of the original encoded data, ensuring that the context used for encoding the next frame is consistent with that of the receiving end, forming a closed-loop feedback.
[0151] Furthermore, the left and right image contexts of the current frame used in the current frame image encoding in step S303 above include: the sending end left-view image context of the current frame and the sending end right-view image context of the current frame. The calculation process of the sending end left and right-view image contexts of the current frame is as follows:
[0152] S303-1: The transmitting end performs disparity compensation and semantic aggregation based on the left disparity decoded by the transmitting end in the current frame, the semantic features of the right view image in the current frame, the semantic features of the left view image in the current frame, and the semantic correlation of the left view image of the transmitting end in the current frame, to obtain the context of the left view image of the transmitting end in the current frame.
[0153] S303-2: The transmitting end performs disparity compensation and semantic aggregation based on the right disparity decoded by the transmitting end in the current frame, the semantic features of the left view image in the current frame, the semantic features of the right view image in the current frame, and the semantic correlation of the right view image of the transmitting end in the current frame, to obtain the right view image context of the transmitting end in the current frame.
[0154] In this embodiment, the left-view image context of the transmitting end is calculated as follows:
[0155]
[0156] in, The left-view image context of the current frame. This indicates the left parallax decoded by the sender in the current frame. This represents the semantic features of the left-view image in the current frame. This represents the semantic features of the right-view image in the current frame. This indicates the semantic relevance of the sender's left perspective in the current frame.
[0157] Similarly, the calculation method for the right-view image context at the sending end is as follows:
[0158]
[0159] in, The right-view image context of the current frame. This indicates the right parallax decoded by the sender in the current frame. This represents the semantic features of the left-view image in the current frame. This represents the semantic features of the right-view image in the current frame. This indicates the semantic relevance of the right-hand view of the sender in the current frame.
[0160] For images, this embodiment uses feature-based disparity compensation and semantic aggregation methods to generate contextual information between the left and right views, so as to realize semantic interaction between views, effectively eliminate noise introduced by occlusion in disparity compensation, and then effectively compress redundancy between viewpoints through cross-view information flow, thereby improving bandwidth transmission efficiency.
[0161] It should be noted that the image encoding process at the transmitting end is similar to the disparity encoding process at the transmitting end. For the similarities, please refer to the relevant content on disparity encoding in Section 1.2.1 above.
[0162] 1.3 The process by which the sender transmits the latent representations of the left and right perspectives to the receiver in the semantic communication system specifically includes:
[0163] S103-1: The transmitting end quantizes the potential representations of the left and right views to obtain the quantized potential representations of the left and right views.
[0164] In this embodiment, the left and right view latent representations include the left and right view disparity latent representations and the left and right view image latent representations. The encoding, compression, and transmission processes of the two are similar. Here, the transmission process of the left and right view disparity latent representation is taken as an example to illustrate the entire transmission process of the left and right view latent representations.
[0165] Please refer to Figure 4 , Figure 4 This is a schematic diagram of the architecture for transmitting the potential representation of left and right view disparity in one embodiment of this application. Figure 4 As shown in this embodiment, the transmitting end quantizes the potential representation of the parallax of the left and right viewpoints, discretizes the continuous potential representation, and facilitates subsequent entropy coding.
[0166] S103-2: The transmitter estimates the distribution parameters of the left and right view latent representations through a super-prior network based on the left and right view latent representations.
[0167] In this embodiment, the Hyperprior Network includes an encoding part and a decoding part. The encoding part performs dimensionality reduction processing on the quantized left and right view latent representations of the input using a lightweight CNN, outputting the hyperprior latent representation. Then, through the decoding part... Decode into distribution parameters (e.g., mean) ,scale ), to describe the probability distribution of each position in the quantized left and right view potential representation.
[0168] S103-3: The transmitting end performs entropy encoding on the quantized left and right view potential representations according to the distribution parameters to obtain the bit stream.
[0169] In this embodiment, the advanced prior latent representation is first discussed. Perform independent quantization and entropy coding (such as arithmetic coding), then calculate the probability for each position of the quantized left and right view latent representations based on the distribution parameters, and use arithmetic coding (such as ANS) to compress the quantized left and right view latent representations into a bit stream, where symbols with higher probabilities are assigned shorter codewords.
[0170] S103-4: The sending end sends the bit stream and distribution parameters to the receiving end.
[0171] S103-5: The receiving end processes the received bit stream according to the distribution parameters to obtain the received left and right view potential representations.
[0172] In this embodiment, the transmitting end sends the bitstream and distribution parameters to the receiving end. The receiving end then processes the bitstream using arithmetic decoding based on the distribution parameters to recover the latent representations of the left and right perspectives. Finally, it performs inverse quantization to obtain the reconstructed latent representations of the left and right perspectives (e.g., ...). Figure 4 Left-view parallax latent representation And right-view parallax latent representation ), used for subsequent 3DGS reconstruction.
[0173] This embodiment uses a super-prior network to perform efficient entropy encoding on the latent representations of the left and right perspectives. By utilizing its statistical properties, the number of transmission bits is reduced. Thus, through quantization and entropy encoding, the left and right feature maps are compressed into low-dimensional vectors, reducing transmission bandwidth and promoting the realization of real-time semantic communication.
[0174] 1.4 Decoding process at the receiving end:
[0175] In this scheme, the receiving end deploys two decoders: a second cross-view parallel disparity semantic feature self-decoder, used to decode the received left and right view disparity latent representations, and a second cross-view parallel image semantic feature self-decoder, used to decode the received left and right view image latent representations and reconstruct the left and right view images. In one embodiment, the second cross-view parallel disparity semantic feature self-decoder is the same as the first cross-view parallel disparity semantic feature self-decoder in Section 1.2 above, and the second cross-view parallel image semantic feature self-decoder is the same as the first cross-view parallel image semantic feature self-decoder in Section 1.2 above.
[0176] Specifically, the process by which the receiving end decodes the received left and right view latent representations to obtain decoded semantic features and reconstructed left and right view images includes:
[0177] The receiver decodes the received left and right view disparity latent representations through the second cross-view parallel disparity semantic feature self-decoder to obtain the receiver-decoded left and right disparities of the current frame and the receiver-decoded left and right disparity semantic features of the current frame; the receiver caches the receiver-decoded left and right disparities of the current frame into a buffer for use in the reconstruction of the left and right view images of the next frame.
[0178] In this embodiment, the receiving end needs to perform disparity decoding. Specifically, the receiving end decodes the received left and right view disparity latent representations using a second cross-view parallel disparity semantic feature self-decoder to obtain the receiving end-decoded left and right disparities and the receiving end-decoded left and right disparity semantic features of the current frame. This process is similar to the process in Section 1.2 where the sending end decodes the left and right view disparity latent representations of the current frame using a first cross-view parallel disparity semantic feature self-decoder to obtain the sending end-decoded left and right disparities and the sending end-decoded left and right disparity semantic features of the current frame, and will not be repeated here. Furthermore, similarly, the receiving end buffers the receiving end-decoded left and right disparities of the current frame into a buffer for use in reconstructing the left and right view images of the next frame.
[0179] On the other hand, the receiving end also needs to perform image decoding and image reconstruction based on the decoded disparity semantic features and image semantic features. Specifically, this process mainly includes:
[0180] S401: The receiver decodes the latent representation of the received left and right view images to obtain the decoded left and right view images of the current frame and the semantic features of the decoded left and right view images of the current frame.
[0181] This process is similar to the process described in Section 1.2.2, where the sending end decodes the latent representations of the left and right view images of the current frame using the first cross-view parallel image semantic feature self-decoder, and obtains the process of the sending end decoding the left and right view images and the semantic features of the left and right view images of the current frame. Therefore, it will not be described again here.
[0182] In this embodiment, it is assumed that the receiver-decoded left and right view images of the current frame are the same as those of the sender-decoded left and right view images of the current frame. The semantic features of the receiver-decoded left and right view images of the current frame are the same as those of the sender-decoded left and right view images of the current frame.
[0183] S402: The receiver reads the semantic features of the left and right view images from the buffer of the previous frame.
[0184] Please refer to step S302 above for the process; it will not be repeated here.
[0185] S403: The receiving end uses the second cross-view parallel image semantic feature self-decoder to perform parallel reconstruction of the left and right view images of the current frame based on the semantic features of the left and right view images decoded by the receiving end of the current frame, the semantic features of the left and right view images decoded by the receiving end of the previous frame, and the context of the left and right view images of the receiving end of the current frame, to obtain the reconstructed left and right view images and semantic features of the reconstructed left and right view images of the current frame.
[0186] In this embodiment, the left and right view image context of the current frame refers to the image context information between the left and right view images of the receiver generated by the receiver through parallax compensation and semantic aggregation.
[0187] Specifically, when the receiving end reconstructs the left and right view images of the current frame using the second cross-view parallel image semantic feature self-decoder, it needs to combine three feature indicators: the semantic features of the left and right view images decoded by the receiving end of the current frame, the semantic features of the left and right view images decoded by the receiving end of the previous frame, and the context of the left and right view images of the receiving end of the current frame, so as to reconstruct higher quality left and right view images of the current frame through disparity compensation and semantic aggregation.
[0188] In addition, the receiving end will also cache the semantic features of the left and right view images decoded by the receiving end in the current frame into a buffer for use in the reconstruction of the left and right view images in the next frame.
[0189] Furthermore, the receiver's left and right view image contexts used in step S403 for reconstructing the current frame image include: the receiver's left view image context and the receiver's right view image context. The calculation process for the receiver's left and right view image contexts of the current frame is as follows:
[0190] S403-1: The receiver performs disparity compensation and semantic aggregation based on the receiver-decoded left disparity of the current frame, the receiver-decoded right-view image semantic features of the current frame, the receiver-decoded left-view image semantic features of the current frame, and the receiver-decoded left-view semantic correlation of the current frame, to obtain the receiver-decoded left-view image context of the current frame; the receiver-decoded left-view semantic correlation of the current frame is determined based on the receiver-decoded left disparity of the current frame and the receiver disparity from the right view to the left view of the current frame.
[0191] S403-2: The receiver performs disparity compensation and semantic aggregation based on the receiver-decoded right disparity of the current frame, the receiver-decoded left-view image semantic features of the current frame, the receiver-decoded right-view image semantic features of the current frame, and the receiver-decoded right-view semantic correlation of the current frame, to obtain the receiver-decoded right-view image context of the current frame; the receiver-decoded right-view semantic correlation of the current frame is determined based on the receiver-decoded right disparity of the current frame and the receiver disparity from the left view to the right view of the current frame.
[0192] The calculation method for the left and right view image context of the receiving end of the current frame is the same as the calculation method for the left and right view image context of the sending end of the current frame. Please refer to steps S303-1 and S303-2 above for details, which will not be repeated here.
[0193] 1.5 The process by which the receiver predicts Gaussian parameters based on the decoded semantic features and the reconstructed left and right view images includes:
[0194] S501: The receiver predicts the center position based on the left and right parallax of the current frame.
[0195] Please refer to Figure 5 , Figure 5 This is a schematic diagram of the Gaussian parameter prediction architecture in one embodiment of this application. Since the left and right views share parameters, and the Gaussian parameter prediction process is independent at each time step, for the sake of brevity, in... Figure 5 The time t and its subscripts and superscripts are omitted.
[0196] like Figure 5 As shown, the receiver first decodes the left and right parallax based on the current frame. Predict the center location Specifically, decoding yields the left and right parallax. Then, the disparity map can be transformed linearly according to the camera parameters to obtain the depth map of the view. The depth map can be inversely projected onto the world coordinate system by defining Gaussian points pixel by pixel, thereby calculating the center position of the Gaussian ellipsoid.
[0197] S502: The receiver reconstructs the image based on the left and right perspectives of the current frame and predicts the color of the Gaussian ellipsoid.
[0198] In this embodiment, considering that the light on the human body surface is mainly diffuse reflection, only the 0th-order spherical harmonic function is used to decode the reconstructed images of the left and right perspectives (i.e., Figure 5 In The RGB values of ) are used as the color of the Gaussian ellipsoid (i.e. Figure 5 In This reduces the complexity of the model.
[0199] S503: The receiver decodes the left and right parallax semantic features of the current frame and reconstructs the image semantic features from the left and right perspectives of the current frame, and predicts the covariance matrix and opacity.
[0200] In this embodiment, the left and right parallax semantic features (i.e., the decoded features obtained by the receiver) are used. Figure 5 In ) and left and right perspective reconstructed image semantic features (i.e. Figure 5 In The covariance matrix of a 3D Gaussian can be predicted using a small number of depthwise separable convolutional modules (decomposed into scale and rotation, i.e., ...). Figure 5 S and R in the text) and opacity (i.e. Figure 5 (O in)
[0201]
[0202]
[0203]
[0204] in, Used to represent the shape of a Gaussian ellipsoid. Used to indicate the direction of the Gaussian ellipsoid Used to represent the opacity of a Gaussian ellipsoid. This represents the maximum value of the Gaussian ellipsoid. This is a module for preprocessing semantic features of images and disparity in Gaussian parameter prediction. , and It is a head network used to obtain the parameters of the prediction, consisting of only a single depthwise separable convolution. This indicates that the receiver decodes the left and right parallax semantic features for the current frame. Including the left disparity semantic features decoded at the receiver. Decoding right disparity semantic features at the receiving end . This represents the semantic features of the reconstructed image from the left and right perspectives of the current frame. For the current frame, Including the left disparity semantic features decoded at the receiver. Decoding right disparity semantic features at the receiving end .
[0205] Finally, based on the value range of different attributes, the corresponding activation function or normalization method is used to obtain the corresponding Gaussian parameters. The prediction process is carried out in parallel by left and right view branches to reduce processing latency.
[0206] Based on the decoded semantic features and the reconstructed left and right view images, the Gaussian parameters are predicted, and the user's view image can be rendered according to the predicted Gaussian parameters. For specific rendering processes, please refer to existing technologies; this article does not impose any limitations on them.
[0207] This embodiment uses a cross-view parallel semantic feature autoencoder to encode the left and right view videos and predict Gaussian parameters in parallel, which reduces transmission bandwidth and improves the real-time performance of semantic communication.
[0208] Please refer to Figure 6Based on the same inventive concept, a second aspect of this application provides a three-dimensional Gaussian splash multi-view video joint semantic coding apparatus 600, which includes:
[0209] Compression module 601 is used in a semantic communication system for the sending end to encode left and right view semantic features in parallel for left and right view images through a cross-view parallel semantic feature autoencoder to obtain left and right view latent representations. The cross-view parallel semantic feature autoencoder reduces redundancy between viewpoints by transmitting cross-view context information.
[0210] Transmission module 602 is used by the sending end to send the left and right view potential representations to the receiving end in the semantic communication system;
[0211] The reconstruction module 603 is used by the receiving end to decode the received left and right view potential representations to obtain decoded semantic features and reconstructed left and right view images. Based on the decoded semantic features and reconstructed left and right view images, Gaussian parameters are predicted, and the user's view image is rendered based on the predicted Gaussian parameters.
[0212] As the device embodiment is basically similar to the method embodiment, the description is relatively simple, and relevant parts can be found in the description of the method embodiment.
[0213] Thirdly, based on the same inventive concept, and referring to... Figure 7 This application provides an electronic device 700, including a processor 701 and a memory 702; the memory 702 stores machine-executable instructions that can be executed by the processor 701, and the processor 701 is used to execute the machine-executable instructions to implement the three-dimensional Gaussian splash multi-view video joint semantic coding method as proposed in the first aspect of this application.
[0214] It should be noted that the specific implementation of the electronic device 700 in this application embodiment refers to the specific implementation of the three-dimensional Gaussian splash multi-view video joint semantic coding method proposed in the first aspect of the above-mentioned application embodiment, and will not be repeated here.
[0215] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The same or similar parts between the various embodiments can be referred to each other.
[0216] Those skilled in the art will understand that embodiments of this application can be provided as methods, apparatus, or computer program products. Therefore, embodiments of this application can take the form of entirely hardware embodiments, entirely software embodiments, or embodiments combining software and hardware aspects. Furthermore, embodiments of this application can take the form of computer program products implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0217] This application describes embodiments with reference to flowchart illustrations and / or block diagrams of methods, terminal devices (systems), and computer program products according to embodiments of this application. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing terminal device to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing terminal device, generate instructions for implementing the flowchart illustrations. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0218] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing terminal device to operate in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0219] These computer program instructions can also be loaded onto a computer or other programmable data processing terminal equipment, causing a series of operational steps to be performed on the computer or other programmable terminal equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable terminal equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0220] Although preferred embodiments of the present application have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments as well as all changes and modifications falling within the scope of the embodiments of the present application.
[0221] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or terminal device that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or terminal device. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or terminal device that includes said element.
[0222] The above provides a detailed description of a three-dimensional Gaussian splash multi-view video joint semantic coding method provided in this application. Specific examples have been used to illustrate the principles and implementation methods of this application. The description of the above embodiments is only for the purpose of helping to understand the method and its core ideas. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.
Claims
1. A joint semantic coding method for three-dimensional Gaussian splash multi-view video, characterized in that, The method includes: In the semantic communication system, the transmitting end performs left and right view semantic feature encoding in parallel for the left and right view images through a cross-view parallel semantic feature autoencoder to obtain the left and right view latent representations. The cross-view parallel semantic feature autoencoder reduces redundancy between viewpoints by transmitting cross-view context information. The sending end sends the potential representations of the left and right perspectives to the receiving end in the semantic communication system; The receiving end decodes the received left and right view latent representations to obtain decoded semantic features and reconstructed left and right view images. Based on the decoded semantic features and reconstructed left and right view images, it predicts Gaussian parameters and renders the user's view image based on the predicted Gaussian parameters.
2. The method according to claim 1, characterized in that, The cross-view parallel semantic feature autoencoder includes at least: a cross-view parallel disparity semantic feature autoencoder; In a semantic communication system, the transmitting end uses a cross-view parallel semantic feature autoencoder to encode left and right view semantic features in parallel for both left and right view images, obtaining left and right view latent representations, which include at least: The transmitting end uses a disparity estimation module to estimate the disparity of the left and right view images of the current frame, thereby obtaining the left and right disparity of the current frame and the semantic features of the left and right disparity of the current frame. The transmitting end reads the left and right parallax semantic features decoded by the transmitting end of the previous frame from the buffer; The transmitting end uses the cross-view parallel disparity semantic feature autoencoder to encode the left and right view disparity semantic features of the current frame in parallel, based on the left and right disparity semantic features of the current frame, the left and right disparity semantic features decoded by the transmitting end of the previous frame, and the left and right disparity context of the current frame, to obtain the potential representation of the left and right view disparity of the current frame. The method further includes: The sending end decodes the left and right view disparity latent representation of the current frame through the first cross-view parallel disparity semantic feature self-decoder to obtain the sending end decoded left and right disparities of the current frame and the sending end decoded left and right disparity semantic features of the current frame. The transmitting end caches the left and right disparity semantic features decoded by the transmitting end of the current frame into the buffer for use in encoding the left and right view disparity semantic features of the next frame. The transmitting end also caches the left and right disparity of the current frame decoded by the transmitting end into the buffer for use in encoding the left and right view image semantic features of the next frame.
3. The method according to claim 2, characterized in that, The left and right disparity contexts of the current frame include: the left disparity context of the current frame and the right disparity context of the current frame; The method further includes: The transmitting end performs disparity compensation and semantic aggregation based on the left disparity of the current frame, the right disparity semantic features of the current frame, the left disparity semantic features of the current frame, and the semantic correlation of the transmitting end's left viewpoint of the current frame to obtain the left disparity context of the current frame; the semantic correlation of the transmitting end's left viewpoint of the current frame is determined based on the left disparity of the current frame and the transmitting end disparity from the right viewpoint to the left viewpoint of the current frame. The transmitting end performs disparity compensation and semantic aggregation based on the right disparity of the current frame, the left disparity semantic features of the current frame, the right disparity semantic features of the current frame, and the right view semantic correlation of the transmitting end in the current frame to obtain the right disparity context of the current frame; the right view semantic correlation of the transmitting end in the current frame is determined based on the right disparity of the current frame and the transmitting end disparity from the left view to the right view of the current frame.
4. The method according to claim 2, characterized in that, The cross-view parallel semantic feature autoencoder further includes: a cross-view parallel image semantic feature autoencoder; In the semantic communication system, the transmitting end uses a cross-view parallel semantic feature autoencoder to encode left and right view semantic features in parallel for both left and right view images, obtaining the left and right view latent representations. This also includes: The sending end extracts features from the left and right view images of the current frame through the image feature extraction module to obtain the semantic features of the left and right view images of the current frame. The transmitting end reads the semantic features of the left and right view images of the previous frame from the buffer; The transmitting end uses the cross-view parallel image semantic feature autoencoder to encode the left and right view image semantic features of the current frame in parallel, based on the left and right view image semantic features of the current frame, the left and right view image semantic features decoded by the transmitting end of the previous frame, and the left and right view image context of the transmitting end of the current frame, to obtain the potential representation of the left and right view image of the current frame. The method further includes: The sending end decodes the latent representation of the left and right view images of the current frame through the first cross-view parallel image semantic feature self-decoder to obtain the sending end decoded left and right view images of the current frame and the sending end decoded left and right view image semantic features of the current frame. The sending end caches the semantic features of the left and right view images decoded by the sending end in the current frame into the buffer for use in encoding the semantic features of the left and right view images in the next frame.
5. The method according to claim 4, characterized in that, The left and right view image contexts of the current frame's sender include: the left view image context of the current frame's sender and the right view image context of the current frame's sender. The method further includes: The transmitting end performs disparity compensation and semantic aggregation based on the left disparity decoded by the transmitting end of the current frame, the semantic features of the right view image of the current frame, the semantic features of the left view image of the current frame, and the semantic correlation of the left view image of the transmitting end of the current frame, to obtain the left view image context of the transmitting end of the current frame. The transmitting end performs disparity compensation and semantic aggregation based on the right disparity decoded by the transmitting end of the current frame, the semantic features of the left-view image of the current frame, the semantic features of the right-view image of the current frame, and the semantic correlation of the right-view image of the transmitting end of the current frame, to obtain the right-view image context of the transmitting end of the current frame.
6. The method according to claim 2, characterized in that, The receiving end decodes the received left and right view latent representations to obtain decoded semantic features and reconstructed left and right view images, including at least: The receiving end decodes the received left and right view disparity latent representations through the second cross-view parallel disparity semantic feature self-decoder to obtain the receiver-decoded left and right disparities of the current frame and the receiver-decoded left and right disparity semantic features of the current frame. The receiving end caches the left and right parallax of the current frame into the buffer for use in reconstructing the left and right view images of the next frame.
7. The method according to claim 6, characterized in that, The receiving end decodes the received left and right view latent representations to obtain decoded semantic features and reconstructed left and right view images, and also includes: The receiving end decodes the latent representation of the received left and right view images to obtain the receiver-decoded left and right view images of the current frame and the semantic features of the receiver-decoded left and right view images of the current frame. The receiving end reads the semantic features of the left and right view images of the previous frame from the buffer; The receiving end uses a second cross-view parallel image semantic feature self-decoder to perform parallel reconstruction of the left and right view images of the current frame based on the left and right view image semantic features decoded by the receiving end of the current frame, the left and right view image semantic features decoded by the receiving end of the previous frame, and the left and right view image context of the receiving end of the current frame, to obtain the reconstructed left and right view images and the semantic features of the reconstructed left and right view images of the current frame. The method further includes: The receiving end caches the semantic features of the left and right view images decoded by the receiving end in the current frame into the buffer for use in the reconstruction of the left and right view images in the next frame.
8. The method according to claim 7, characterized in that, The receiver's left and right view image contexts for the current frame include: the receiver's left view image context and the receiver's right view image context for the current frame. The method further includes: The receiving end performs disparity compensation and semantic aggregation based on the left disparity decoded by the receiving end of the current frame, the semantic features of the right-view image decoded by the receiving end of the current frame, the semantic features of the left-view image decoded by the receiving end of the current frame, and the semantic correlation of the left-view image of the receiving end of the current frame, to obtain the left-view image context of the receiving end of the current frame; the semantic correlation of the left-view image of the receiving end of the current frame is determined based on the left disparity decoded by the receiving end of the current frame and the disparity from the right view to the left view of the current frame. The receiving end performs disparity compensation and semantic aggregation based on the right disparity decoded by the receiving end of the current frame, the semantic features of the left-view image decoded by the receiving end of the current frame, the semantic features of the right-view image decoded by the receiving end of the current frame, and the semantic correlation of the right-view image of the receiving end of the current frame, to obtain the right-view image context of the receiving end of the current frame; the semantic correlation of the right-view image of the receiving end of the current frame is determined based on the right disparity decoded by the receiving end of the current frame and the disparity from the left view to the right view of the current frame.
9. The method according to claim 7, characterized in that, The receiving end predicts Gaussian parameters based on the decoded semantic features and the reconstructed left and right view images, including: The receiving end decodes the left and right parallax of the current frame and predicts the center position. The receiving end reconstructs the image based on the left and right perspectives of the current frame and predicts the color of the Gaussian ellipsoid. The receiving end decodes the left and right parallax semantic features of the current frame and reconstructs the image semantic features from the left and right perspectives of the current frame to predict the covariance matrix and opacity.
10. The method according to any one of claims 1-9, characterized in that, The sending end transmits the left and right view latent representations to the receiving end in the semantic communication system, including: The transmitting end quantizes the potential representation of the left and right views to obtain the quantized potential representation of the left and right views. The transmitting end estimates the distribution parameters of the left and right view potential representations through a super-prior network based on the left and right view potential representations. The transmitting end performs entropy encoding on the quantized left and right view potential representations according to the distribution parameters to obtain a bit stream; The sending end sends the bit stream and the distribution parameters to the receiving end; The receiving end processes the received bit stream according to the distribution parameters to obtain the received left and right view potential representations.
Citation Information
Patent Citations
Three-dimensional image coding method for machine vision
CN119835395A
Video generation method and device, electronic equipment, storage medium and program product
CN120434373A