Three-dimensional Gaussian splash multi-view video joint semantic coding method

Through cross-view parallel semantic feature autoencoders and Gaussian parameter prediction, the redundancy of multi-view video images is reduced, real-time image transmission for immersive semantic communication is achieved, and the balance problem between coding complexity and compression rate in traditional coding methods is solved.

CN120769035AActive Publication Date: 2025-10-10TSINGHUA UNIVERSITY

Patent Information

Application Number
CN202511141476.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-15
Publication Date
2025-10-10
Estimated Expiration
2045-08-15

AI Technical Summary

Technical Problem

Existing traditional coding methods are difficult to support real-time transmission of multi-perspective high-definition video images in immersive semantic communication scenarios, especially when it is difficult to strike a balance between coding complexity and compression rate, and the redundant compression between multiple perspectives is insufficient, which affects the real-time performance of communication.

Method used

A three-dimensional Gaussian splattering multi-view video joint semantic coding method is adopted. The semantic features of the left and right views are encoded through a cross-view parallel semantic feature autoencoder, cross-view context information is transmitted, the redundancy between views is reduced, and lightweight Gaussian parameters are predicted at the receiving end to render the user's view image.

Benefits of technology

It improves the real-time performance and processing efficiency of video image coding, reduces coding complexity, and meets the real-time requirements of immersive semantic communication.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120769035A_ABST
    Figure CN120769035A_ABST
Patent Text Reader

Abstract

The invention provides a three-dimensional Gaussian splash multi-view video joint semantic coding method, and relates to the technical field of computer vision processing. According to the method, the multi-view video image coding and the feedforward 3DGS technology are combined, the architecture of double-view coding and Gaussian parameter prediction parallel operation is adopted, left and right view branch model parameters are shared, the processing real-time performance is improved, and the complexity is reduced. Meanwhile, a cross-view-angle parallel semantic feature auto-encoder is designed, and information interaction between left and right view angle branches is realized by transmitting cross-view-angle context information, so that redundancy between view angles is reduced. And finally, realizing 3DGS lightweight prediction on the decoding features at a receiving end, and rendering to obtain a video image of a user view angle so as to meet the real-time requirement of immersive semantic communication.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of computer vision processing, in particular to a three-dimensional Gaussian splash multi-view video joint semantic coding method. BACKGROUND

[0002] Immersive semantic communication technology is a core application scenario of 6G, which completely changes people's work, entertainment and communication methods through realistic three-dimensional visual presentation and interaction. However, the requirements of immersive semantic communication for ultra-high data transmission rate and ultra-low delay bring challenges to 6G network and data processing technology, which prompts researchers to develop corresponding high-performance low-delay video image coding technology for immersive semantic communication.

[0003] The related coding method is usually based on the traditional coding framework of search, transformation and entropy coding. The improvement of the coding performance of the traditional coding method depends on the increasing coding mode. However, the more the coding modes are, the more the complexity of the rate-distortion optimization search will increase, and the benefit ratio of the compression rate increase and the coding complexity improvement is smaller and smaller. Especially in the immersive semantic communication scenario, the current traditional coding method is difficult to support real-time transmission of panoramic or multi-view high-definition video images with huge data volume. SUMMARY

[0004] The present application provides a three-dimensional Gaussian splash multi-view video joint semantic coding method, which reduces the inter-view redundancy by transmitting cross-view context information, and helps to realize low-bandwidth immersive semantic communication.

[0005] The first aspect of the embodiment of the present application provides a three-dimensional Gaussian splash multi-view video joint semantic coding method, which comprises: The sending end in the semantic communication system performs left and right view semantic feature coding on the left and right view images in parallel through a cross-view parallel semantic feature self-encoder, and obtains left and right view latent representations, wherein the cross-view parallel semantic feature self-encoder reduces the inter-view redundancy by transmitting cross-view context information. The sending end sends the left and right view latent representations to the receiving end in the semantic communication system. The receiving end decodes the received left and right view latent representations to obtain decoded semantic features and reconstructed left and right view images, predicts Gaussian parameters according to the decoded semantic features and the reconstructed left and right view images, and renders images of a user's view according to the predicted Gaussian parameters.

[0006] Optionally, the cross-view parallel semantic feature self-encoder comprises at least a cross-view parallel disparity semantic feature self-encoder. The transmitter in the semantic communication system uses a cross-view parallel semantic feature autoencoder to encode the left and right view semantic features in parallel for the left and right view images, and obtains the left and right view potential representations, which at least include: The transmitting end performs disparity estimation on the left and right perspective images of the current frame through a disparity estimation module to obtain the left and right disparities of the current frame and the left and right disparity semantic features of the current frame; The sending end reads the sending end decoded left and right disparity semantic features of the previous frame from the buffer; The transmitter encodes the left and right perspective disparity semantic features of the current frame in parallel using the cross-perspective parallel disparity semantic feature autoencoder according to the left and right disparity semantic features of the current frame, the left and right disparity semantic features decoded by the transmitter of the previous frame, and the left and right disparity context of the current frame, to obtain a left and right perspective disparity potential representation of the current frame; The method further comprises: The transmitter decodes the left and right view disparity potential representations of the current frame through a first cross-view parallel disparity semantic feature self-decoder to obtain the transmitter-decoded left and right disparities of the current frame and the transmitter-decoded left and right disparity semantic features of the current frame; The sending end caches the left and right disparity semantic features decoded by the sending end of the current frame to the buffer zone for encoding the left and right perspective disparity semantic features of the next frame, and the sending end caches the left and right disparity semantic features decoded by the sending end of the current frame to the buffer zone for encoding the left and right perspective image semantic features of the next frame.

[0007] Optionally, the left and right disparity contexts of the current frame include: a left disparity context of the current frame and a right disparity context of the current frame; The method further comprises: The transmitter performs disparity compensation and semantic aggregation based on the left disparity of the current frame, the right disparity semantic feature of the current frame, the left disparity semantic feature of the current frame, and the transmitter left perspective semantic relevance of the current frame to obtain the left disparity context of the current frame; the transmitter left perspective semantic relevance of the current frame is determined based on the left disparity of the current frame and the transmitter disparity from the right perspective to the left perspective of the current frame; The sending end performs disparity compensation and semantic aggregation based on the right disparity of the current frame, the left disparity semantic features of the current frame, the right disparity semantic features of the current frame, and the sending end right perspective semantic correlation of the current frame to obtain the right disparity context of the current frame; the sending end right perspective semantic correlation of the current frame is determined based on the right disparity of the current frame and the sending end disparity from the left perspective to the right perspective of the current frame.

[0008] Optionally, the cross-view parallel semantic feature autoencoder further comprises: a cross-view parallel image semantic feature autoencoder; The transmitter in the semantic communication system uses a cross-view parallel semantic feature autoencoder to encode the left and right view semantic features in parallel for the left and right view images to obtain the left and right view potential representations. It also includes: The transmitting end extracts features from the left and right perspective images of the current frame through an image feature extraction module to obtain semantic features of the left and right perspective images of the current frame; The sending end reads the semantic features of the left and right viewing angle images decoded by the sending end of the previous frame from the buffer; The transmitting end uses the cross-view parallel image semantic feature autoencoder to encode the left and right view image semantic features of the current frame in parallel according to the left and right view image semantic features of the current frame, the left and right view image semantic features decoded by the transmitting end of the previous frame, and the left and right view image context of the transmitting end of the current frame, to obtain the left and right view image potential representations of the current frame; The method further comprises: The transmitter decodes the potential representations of the left and right view images of the current frame through the first cross-view parallel image semantic feature self-decoder to obtain the transmitter-decoded left and right view images of the current frame and the transmitter-decoded left and right view image semantic features of the current frame; The transmitting end caches the semantic features of the left and right viewing angle images decoded by the transmitting end of the current frame in the buffer zone for encoding the semantic features of the left and right viewing angle images of the next frame.

[0009] Optionally, the transmitting end left and right view image contexts of the current frame include: the transmitting end left view image context of the current frame and the transmitting end right view image context of the current frame; The method further comprises: The transmitter performs disparity compensation and semantic aggregation based on the transmitter-decoded left disparity of the current frame, the semantic features of the right view image of the current frame, the semantic features of the left view image of the current frame, and the transmitter-left view semantic relevance of the current frame to obtain the transmitter-left view image context of the current frame; The transmitter performs disparity compensation and semantic aggregation based on the transmitter-decoded right disparity of the current frame, the semantic features of the left-view image of the current frame, the semantic features of the right-view image of the current frame, and the transmitter-right-view semantic correlation of the current frame to obtain the transmitter-right-view image context of the current frame.

[0010] Optionally, the receiving end decodes the received left and right view potential representations to obtain decoded semantic features and reconstructed left and right view images, which at least includes: The receiving end decodes the received left and right view disparity potential representations through a second cross-view parallel disparity semantic feature self-decoder to obtain the receiving end decoded left and right disparities of the current frame and the receiving end decoded left and right disparity semantic features of the current frame; The receiving end caches the receiving end-decoded left and right disparities of the current frame in the buffer zone for use in reconstructing the left and right viewing angle images of the next frame.

[0011] Optionally, the receiving end decodes the received left and right view potential representations to obtain decoded semantic features and reconstructed left and right view images, further comprising: The receiving end decodes the received potential representations of the left and right view images to obtain the receiving end decoded left and right view images of the current frame and the receiving end decoded left and right view images semantic features of the current frame; The receiving end reads the semantic features of the left and right viewing angle images decoded by the receiving end of the previous frame from the buffer; The receiving end, through the second cross-view parallel image semantic feature self-decoder, performs, in parallel, left and right view image reconstruction of the current frame based on the semantic features of the left and right view images decoded by the receiving end of the current frame, the semantic features of the left and right view images decoded by the receiving end of the previous frame, and the left and right view image context of the current frame, to obtain left and right view reconstructed images of the current frame and semantic features of the left and right view reconstructed images of the current frame; The method further comprises: The receiving end caches the semantic features of the left and right viewing angle images decoded by the receiving end of the current frame in the buffer zone for use in reconstructing the left and right viewing angle images of the next frame.

[0012] Optionally, the receiving end left and right view image contexts of the current frame include: the receiving end left view image context of the current frame and the receiving end right view image context of the current frame; The method further comprises: The receiving end performs disparity compensation and semantic aggregation based on the receiving end decoded left disparity of the current frame, the receiving end decoded right view image semantic features of the current frame, the receiving end decoded left view image semantic features of the current frame, and the receiving end left view semantic relevance of the current frame to obtain the receiving end left view image context of the current frame; the receiving end left view semantic relevance of the current frame is determined based on the receiving end decoded left disparity of the current frame and the receiving end disparity from the right view to the left view of the current frame; The receiving end performs disparity compensation and semantic aggregation based on the receiving end decoded right disparity of the current frame, the receiving end decoded left perspective image semantic features of the current frame, the receiving end decoded right perspective image semantic features of the current frame, and the receiving end right perspective semantic correlation of the current frame to obtain the receiving end right perspective image context of the current frame; the receiving end right perspective semantic correlation of the current frame is determined based on the receiving end decoded right disparity of the current frame and the receiving end disparity from the left perspective to the right perspective of the current frame.

[0013] Optionally, the receiving end predicts Gaussian parameters based on the decoded semantic features and the reconstructed left and right view images, including: The receiving end predicts a center position according to left and right disparities decoded by the receiving end of the current frame; The receiving end reconstructs an image according to the left and right viewing angles of the current frame and predicts the color of the Gaussian ellipsoid; The receiving end predicts a covariance matrix and opacity according to the receiving end-decoded left and right disparity semantic features of the current frame and the left and right viewing angle-reconstructed image semantic features of the current frame.

[0014] Optionally, the sending end sending the left and right view potential representations to a receiving end in the semantic communication system includes: The transmitting end quantizes the left and right view potential representations to obtain quantized left and right view potential representations; The transmitting end estimates the distribution parameters of the left and right view potential representations through a super prior network according to the left and right view potential representations; The transmitting end performs entropy coding on the quantized left and right view potential representations according to the distribution parameters to obtain a bit stream; The transmitting end sends the bit stream and the distribution parameter to the receiving end; The receiving end processes the received bit stream according to the distribution parameters to obtain the received left and right perspective potential representations.

[0015] Based on the same inventive concept, a second aspect of an embodiment of the present application provides a three-dimensional Gaussian splatter multi-view video joint semantic coding device, the device comprising: A compression module is configured to encode left-view semantic features of a left-view image and a right-view image in parallel at a transmitting end in a semantic communication system using a cross-view parallel semantic feature autoencoder to obtain left-view latent representations. The cross-view parallel semantic feature autoencoder reduces inter-view redundancy by transmitting cross-view context information. A transmission module, configured for the transmitting end to transmit the left and right perspective potential representations to a receiving end in the semantic communication system; The reconstruction module is configured to decode the received left and right view latent representations to obtain decoded semantic features and reconstructed left and right view images, predict Gaussian parameters according to the decoded semantic features and the reconstructed left and right view images, and render an image of a user view according to the predicted Gaussian parameters.

[0016] Based on the same inventive concept, a third aspect of the embodiments of the present application provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, and the processor implements the three-dimensional Gaussian splash multi-view video joint semantic coding method according to the first aspect of the present application.

[0017] Compared with the prior art, the present application has the following advantages: The three-dimensional Gaussian splash multi-view video joint semantic coding method provided by the embodiments of the present application comprises the following steps: a sending end in a semantic communication system performs left and right view semantic feature coding on left and right view images in parallel through a cross-view parallel semantic feature self-encoder to obtain left and right view latent representations, wherein the cross-view parallel semantic feature self-encoder reduces inter-view redundancy by transmitting cross-view context information; the sending end sends the left and right view latent representations to a receiving end in the semantic communication system; and the receiving end decodes the received left and right view latent representations to obtain decoded semantic features and reconstructed left and right view images, predicts Gaussian parameters according to the decoded semantic features and the reconstructed left and right view images, and renders an image of a user view according to the predicted Gaussian parameters.

[0018] Therefore, the present scheme combines multi-view video image coding and feedforward 3DGS technology, adopts a dual-view coding and Gaussian parameter prediction parallel operation architecture, and shares left and right view branch model parameters, thereby improving real-time processing and reducing complexity. Meanwhile, a cross-view parallel semantic feature self-encoder is designed to realize information interaction between left and right view branches by transmitting cross-view context information to reduce inter-view redundancy. Finally, 3DGS lightweight prediction is performed on decoded features at the receiving end, and a video image of a user view is rendered to meet the real-time requirement of immersive semantic communication. BRIEF DESCRIPTION OF DRAWINGS

[0019] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings needed to be used in the description of the embodiments of the present application. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can also be obtained by those skilled in the art without any creative labor.

[0020] Figure 1 is a flowchart of a three-dimensional Gaussian splash multi-view video joint semantic coding method according to an embodiment of the present application; Figure 2 2 is a schematic diagram of a disparity coding architecture of a transmitting end in one embodiment of the present application; Figure 3 This is a schematic diagram of an image coding architecture at a transmitting end in one embodiment of the present application; Figure 4 2 is a schematic diagram of an architecture for transmitting potential representations of left and right perspective disparity in one embodiment of the present application; Figure 5 Schematic diagram of the architecture of Gaussian parameter prediction in one embodiment of the present application; Figure 6 This is a functional module diagram of a device for joint semantic coding of 3D Gaussian splattered multi-view videos in one embodiment of the present application; Figure 7 It is a structural diagram of an electronic device in one embodiment of the present application. DETAILED DESCRIPTION

[0021] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0022] Immersive semantic communication technology is a core application scenario of 6G, which will completely change the way people work, entertain and communicate through realistic three-dimensional visual presentation and interaction. However, the ultra-high data transmission rate and ultra-low latency requirements of immersive semantic communication have brought challenges to 6G networks and data processing technologies, prompting researchers to develop corresponding high-performance and low-latency video image coding technologies for immersive semantic communication.

[0023] Deep learning-based video coding typically combines techniques such as motion estimation, motion compensation, nonlinear transformation, and entropy coding. For example, DVC utilizes these techniques to build the first end-to-end neural network-based video residual coding framework. This framework uses optical flow estimation and motion compensation to predict the current frame and then encodes the optical flow and residual. However, existing deep learning-based video coding models struggle to strike a balance between encoding speed and rate-distortion performance. Furthermore, immersive services require constructing a 3D scene representation and rendering the user's image based on the user's perspective. Related technologies use multi-view images for 3D reconstruction, but due to imperfect image acquisition and noise-affected camera parameter estimation, traditional 3D reconstruction methods produce suboptimal results. 3DGS (3D Gaussian Splatting) rasterizes a set of Gaussian ellipsoids to approximate the appearance of a 3D scene. This technology achieves high-quality synthesis of new viewpoints and allows for fast convergence and real-time rendering (approximately 30 FPS) at 1080p resolution, making low-cost 3D content creation and real-time applications possible.

[0024] However, in immersive semantic communication scenarios, video image coding and reconstruction still face the following problems: 1. General 3DGS-based scene reconstruction requires separate training of three-dimensional representations for different scenes and objects, which lacks real-time and generalization capabilities. Especially for dynamic three-dimensional scenes, a single frame contains nearly one million Gaussian points, and the amount of data is huge, making real-time coding and compression difficult. 2. Adopting a solution to reconstruct and encode 3DGS at the sending end: After reconstructing 3DGS in a feedforward manner, each pixel will have one or more Gaussian points, which increases redundancy compared to video, and the Gaussian point structure in three-dimensional space is more complex, resulting in compression difficulties. 3. Adopting a solution to decode and reconstruct 3DGS at the receiving end after video encoding transmission: The parallax of the multi-view video images collected by the sending end is large, and the existing methods do not sufficiently compress the redundancy between multiple perspectives, resulting in a large computational load at the receiving end, affecting the real-time performance of communication.

[0025] In order to solve the above problems, the embodiment of the present application proposes a three-dimensional Gaussian splash multi-view video joint semantic coding method, which combines multi-view video image coding with feedforward 3DGS technology, adopts a dual-view coding and Gaussian parameter prediction parallel operation architecture, and shares left and right view branch model parameters, thereby improving processing real-time performance and reducing complexity. At the same time, a cross-view parallel semantic feature autoencoder is designed to achieve information interaction between left and right view branches by transmitting cross-view context information to reduce redundancy between views. Finally, 3DGS lightweight prediction is implemented on the decoding features at the receiving end, and the video image from the user's perspective is rendered to meet the real-time requirements of immersive semantic communication. Below, in conjunction with the accompanying drawings, a three-dimensional Gaussian splash multi-view video joint semantic coding method provided by the embodiment of the present application is described in detail through some embodiments and their application scenarios.

[0026] In a first aspect, embodiments of the present application provide a method for joint semantic coding of multi-view videos using three-dimensional Gaussian splatting. The following describes the first aspect of the method through Section 1.1 Method Overview, Section 1.2 Transmitter Encoding, Section 1.3 Transmission Process, Section 1.4 Receiver Decoding, and Section 1.5 3DGS Rendering Process.

[0027] 1.1 A brief overview of the joint semantic coding method for 3D Gaussian splatter multi-view video: Figure 1 As shown, the method includes the following steps: S101: The transmitter in the semantic communication system uses a cross-view parallel semantic feature autoencoder to encode left-view semantic features in parallel for the left-view image and the right-view image to obtain left-view potential representations. The cross-view parallel semantic feature autoencoder reduces redundancy between views by transmitting cross-view context information.

[0028] In this embodiment, the semantic communication system includes a transmitter and a receiver. The transmitter extracts features from the input left and right dual-view images using a cross-view parallel semantic feature autoencoder, obtains the semantic content in the dual-view images, and encodes the semantic information (i.e., highly abstracts and compresses the semantic concepts in the dual-view images) to obtain latent representations of the left and right views.

[0029] The receiving end is used to perform semantic decoding on the received left and right perspective potential representations, which is the inverse process of encoding. The semantic information of the dual-perspective image is restored through the decoding process, and then the left and right perspective images are reconstructed and rendered from the user perspective to achieve immersive semantic communication.

[0030] The input data for the transmitting end's cross-view parallel semantic feature autoencoder is the left and right view images, which form a stereoscopic visual pair and help obtain more comprehensive information about the user's communication scene. During the encoding process, the cross-view parallel semantic feature autoencoder encodes the left and right view semantic features in parallel for the left and right view images, conveying cross-view context information and reducing redundancy between views.

[0031] For example, the cross-view parallel semantic feature autoencoder can use a two-branch neural network (such as a two-way CNN with shared weights) to process the left view image separately and right view image Each branch extracts hierarchical features (such as edges, textures, and object semantics) through multiple layers of depthwise separable convolutions. Cross-view feature interaction and fusion are then performed using contextual features to ultimately generate latent representations for both left and right views. This preserves the unique information of each view while sharing the semantics of each view, reducing redundant information between the left and right views.

[0032] S102: The sending end sends the potential representations of the left and right perspectives to the receiving end in the semantic communication system.

[0033] S103: The receiving end decodes the received left and right view potential representations to obtain decoded semantic features and reconstructed left and right view images, predicts Gaussian parameters based on the decoded semantic features and the reconstructed left and right view images, and renders the image from the user's perspective based on the predicted Gaussian parameters.

[0034] In this embodiment, the transmitter encodes the semantic features of the left and right views through a cross-view parallel semantic feature autoencoder, obtains the potential representations of the left and right views, and then sends the potential representations of the left and right views to the receiver in the semantic communication system.

[0035] After receiving the left and right view latent representations, the receiver decodes them to obtain decoded semantic features, and then further reconstructs the left and right view images based on the decoded semantic features (see Section 1.3 for details). In this implementation, the left and right view images reconstructed by the receiver are assumed to be approximately the same as the original left and right view images from the transmitter. However, it is easy to understand that in actual applications, due to losses in the semantic feature extraction and compression encoding processes, the left and right view images reconstructed by the receiver may differ slightly from the original left and right view images from the transmitter.

[0036] Next, 3DGS (three-dimensional Gaussian splatting) is used to render the 3D scene. 3DGS transforms each point or object in the scene into a Gaussian distribution with parameters such as position, shape, and color. These parameters are dynamically adjusted to fit the complex shapes of real-world objects (such as the curved surfaces of buildings and the hair of people). Utilizing GPU-accelerated differentiable rasterization technology, 3DGS employs a "snowballing" algorithm to project the Gaussian distribution onto a 2D screen to avoid jagged edges. It also supports dynamic optimization of the Gaussian splatter's position, shape, and transparency. This allows for flexible representation of complex 3D scenes using discrete Gaussian distributions, enabling efficient real-time rendering.

[0037] In this embodiment, the receiving end first predicts the Gaussian parameters, including the position, color, and opacity of the Gaussian ellipsoid, based on the decoded semantic features and the reconstructed left and right perspective images. Then, the image from the user's perspective is rendered based on the predicted Gaussian parameters. That is, the Gaussian points in the three-dimensional space are projected onto the two-dimensional image plane, and these projected data points produce a visual effect on the image in a certain way, thereby appearing in the final rendered image. It is easy to understand that the user's perspective at the receiving end is often different from the perspective of the left and right perspective images collected by the sending end. This embodiment performs left and right perspective semantic feature encoding in parallel at the sending end, transmits cross-perspective context information, and then realizes image rendering of the user's new perspective at the receiving end.

[0038] Exemplarily, if given an RGB video of a human-centered scene with sparse camera views, the present scheme aims to compress the video of two adjacent views, reduce the transmission bandwidth requirement, and realize real-time human high-quality new view view rendering at the receiving end. Specifically, in the above semantic communication system, at each time point t, the present scheme parallelly performs the encoding and decoding of the left and right view video frames and the prediction of the Gaussian parameters, realizing immersive semantic communication.

[0039] The present embodiment combines multi-view video image encoding and feedforward 3DGS technology, adopts a dual-view encoding and Gaussian parameter prediction parallel operation architecture, and fully shares the parameters of the left and right view branch models to support parallel computing, thereby improving the real-time processing and reducing the complexity. Meanwhile, a cross-view parallel semantic feature autoencoder is designed to realize information interaction between the left and right view branches by delivering cross-view context information, thereby reducing the inter-view redundancy. Finally, at the receiving end, 3DGS lightweight prediction is performed on the decoded features, and the video image of the user's view is rendered to meet the real-time requirement of immersive semantic communication.

[0040] As can be easily understood, the process of cross-view image encoding compression by the cross-view parallel semantic feature autoencoder at the sending end and the image reconstruction at the receiving end in the above semantic communication system can be regarded as being completed by a complete image reconstruction model, i.e., the input of the image reconstruction model is the left and right view images, and the output is the reconstructed left and right view images. Some model parameters of the image reconstruction model, such as the network weights during feature extraction by the cross-view parallel semantic feature autoencoder, need to be obtained by training a large number of training samples.

[0041] Exemplarily, the training samples can include original left and right view images and corresponding real reconstructed left and right view images. During the training process, the original left and right view images are input into the initial image reconstruction model for feature extraction and reconstruction prediction to obtain predicted reconstructed left and right view images. Then, the loss function value is calculated according to the predicted reconstructed left and right view images and the real reconstructed left and right view images, and the model parameter of the initial image reconstruction model is updated according to the loss function value. This process is repeated for multiple rounds of model training until the loss function converges or a preset number of training times is reached, and the training is ended to obtain the trained image reconstruction model.

[0042] 1.2 Encoding process of the sending end: This scheme uses two parallel cross-view semantic feature autoencoders: a parallel cross-view disparity semantic feature autoencoder, which encodes the disparity between the left and right view images to obtain a latent representation of the disparity; and a parallel cross-view image semantic feature autoencoder, which encodes the left and right view images to obtain a latent representation of the left and right view images. The latent representations of the left and right view disparity and the left and right view images together form the latent representations of the left and right view for transmission. Within each of these parallel cross-view semantic feature autoencoders, two autoencoders run in parallel, encoding the disparity or image of the left and right view, respectively.

[0043] The following sections describe the disparity and image encoding processes, respectively, in Sections 1.2.1 and 1.2.2. 1.2.1 Please refer to Figure 2 , Figure 2 FIG. 1 is a schematic diagram of the disparity coding architecture of the transmitting end in one embodiment of the present application. Figure 2 As shown in Figure 1, the transmitter includes a disparity estimation module, a cross-view parallel disparity semantic feature autoencoder, and a buffer. Specifically, the disparity encoding process at the transmitter includes: S201: The sending end performs disparity estimation on the left and right perspective images of the current frame through a disparity estimation module to obtain the left and right disparities of the current frame and the left and right disparity semantic features of the current frame.

[0044] For the horizontally aligned left and right views after stereo rectification, the displacement of corresponding pixels between the views is limited to the horizontal direction. The predicted disparity map can be linearly transformed to obtain the depth map of the view based on parameters such as the camera's focal length, baseline distance, and optical center deviation. The depth map can be used to inversely project the Gaussian points defined pixel by pixel into the world coordinate system. Therefore, through explicit disparity estimation and disparity encoding, the center position of the 3DGS can be directly obtained at the receiver. Because this geometric property significantly affects rendering quality, transmitting disparity ensures that the accuracy of the 3DGS center position does not degrade with the degradation of decoded image quality.

[0045] In addition, disparity represents the pixel correspondence between views. Considering that the disparity of images is large under sparse views, explicit disparity is helpful in capturing the semantic correlation between views, thereby helping to compress redundancy between views. Therefore, in this embodiment, the disparity estimation module is based on the left and right view images of the current frame (i.e. Figure 2 in and ), by constructing the matching cost volume and calculating the left and right disparity of the current frame (i.e. Figure 2 in and ), avoiding slow 3D convolution calculations. Then, feature extraction is performed on the left and right disparities to obtain the semantic features of the left and right disparities (i.e. Figure 2 To further improve the real-time performance of the model, in the embodiment, the feature extraction of the current frame is not performed in a step-by-step down-sampling manner using a common convolutional layer, but the current frame is directly down-sampled by 8 times and then a deep separable convolution module is used to extract features.

[0046] Specifically, the process of the disparity estimation module can be represented as follows:

[0047] wherein, represents the left disparity of the current frame, represents the right disparity of the current frame, represents a feature extractor of the disparity estimation module, represents the disparity estimation module, represents the left perspective image of the current frame, represents the right perspective image of the current frame, represents the camera parameters under the left perspective, represents the camera parameters under the right perspective.

[0048] It should be noted that the left and right perspective images of the current frame refer to the left and right perspective images corresponding to the current time t, and the left and right perspective images of the previous frame refer to the left and right perspective images corresponding to the previous time t-1.

[0049] S202: The sending end reads the sending end decoded left and right disparity semantic features of the previous frame from the buffer.

[0050] As shown in FIG. 2, the buffer is mainly used to store the left and right disparities and the left and right disparity semantic features of the video images obtained in the semantic communication process. Figure 2 In the embodiment, the sending end decoded left and right disparity semantic features refer to that the sending end decodes the left and right perspective disparity latent representations of the previous frame through the first cross-perspective parallel disparity semantic feature self-decoder to obtain the left and right disparity semantic features of the previous frame.

[0051] Considering that the scene change between adjacent frames of a video is usually gradual (such as object movement and slow change of perspective). The disparity semantic features (such as object contour and depth distribution) of the previous frame have strong correlation with the current frame. Therefore, when the sending end encodes the left and right perspective disparity semantic features of the current frame, the sending end needs to read the sending end decoded left and right disparity semantic features of the previous frame from the buffer (i.e.

[0052] Figure 2 ​​​​​), using the previous frame's disparity semantic features as context for conditional encoding. This uses the previous frame's features as a priori information to encode only the residual (the changing portion), reducing the bitrate. In other words, temporal coherence and disparity dynamic consistency improve coding efficiency and reconstruction quality.

[0053] S203: The sending end uses a cross-view parallel disparity semantic feature autoencoder to encode the left and right perspective disparity semantic features of the current frame in parallel according to the left and right disparity semantic features of the current frame, the left and right disparity semantic features decoded by the sending end of the previous frame, and the left and right disparity context of the current frame, to obtain the potential representation of the left and right perspective disparity of the current frame.

[0054] In this embodiment, the left and right disparity context of the current frame refers to disparity context information between left and right perspective images generated by disparity compensation and semantic aggregation.

[0055] like Figure 2 As shown in the figure, when the transmitter encodes the disparity of the current frame through the cross-view parallel disparity semantic feature autoencoder, it needs to combine three feature indicators: the left and right disparity semantic features of the current frame (i.e. Figure 2 in and ), the transmitter decodes the left and right disparity semantic features of the previous frame (i.e. Figure 2 in and ), and the left and right disparity context of the current frame (i.e. Figure 2 in and ) is encoded to obtain the potential representation of the left and right perspective disparity of the current frame (i.e. Figure 2 in and ), transmitted to the receiving end for view reconstruction.

[0056] This embodiment encodes the left and right perspective disparity semantic features of the current frame in parallel based on the left and right disparity semantic features of the current frame, the left and right disparity semantic features decoded by the sending end of the previous frame, and the left and right disparity context of the current frame, so as to realize the disparity semantic interaction between the left and right views, and further effectively compress the redundancy between perspectives through the cross-view information flow.

[0057] In addition, in addition to the cross-view parallel disparity semantic feature autoencoder, the sending end also deploys a first cross-view parallel disparity semantic feature autodecoder to decode the potential representation of the left and right view disparity of the current frame, and obtain the sending end decoded left and right disparity of the current frame and the sending end decoded left and right disparity semantic features of the current frame.

[0058] Then, the sender caches the sender-decoded left and right disparity semantic features of the current frame to the buffer for encoding the left and right perspective disparity semantic features of the next frame, and the sender caches the sender-decoded left and right disparities of the current frame to the buffer for encoding the left and right perspective image semantic features of the next frame.

[0059] Considering that if the transmitter only encodes and transmits the potential representation of the left and right view disparity without local decoding, the decoding results between the receiver and the transmitter may differ due to quantization or noise (i.e., "drift error"), and long-term accumulation will lead to a decrease in reconstruction quality. Therefore, in this embodiment, by deploying a first cross-view parallel disparity semantic feature self-decoder at the transmitter, the left and right disparity and left and right disparity semantic features of the current frame are decoded in real time, completely synchronized with the receiver. The decoded left and right disparity and left and right disparity semantic features of the current frame are cached in a buffer area instead of caching the original encoded data, to ensure that the context used when encoding the next frame is consistent with that of the receiver, forming a closed-loop feedback loop.

[0060] Specifically, the process of encoding the left and right view disparity semantic features of the next frame based on the transmitter's decoding of the left and right disparity semantic features of the current frame is similar to steps S201-S203 above and will not be repeated here. For the process of encoding the left and right view image semantic features of the next frame based on the transmitter's decoding of the left and right disparity semantic features of the current frame, please refer to Section 1.2.2 below.

[0061] Furthermore, the left and right disparity contexts of the current frame used for disparity encoding of the current frame in step S203 include: the left disparity context of the current frame and the right disparity context of the current frame. The calculation process of the left and right disparity contexts of the current frame is as follows: S203-1: The sender performs disparity compensation and semantic aggregation based on the left disparity of the current frame, the right disparity semantic features of the current frame, the left disparity semantic features of the current frame, and the semantic relevance of the sender's left perspective of the current frame to obtain the left disparity context of the current frame; the semantic relevance of the sender's left perspective of the current frame is determined based on the left disparity of the current frame and the sender's disparity from the right perspective to the left perspective of the current frame.

[0062] Similar to the use of motion compensation to model the correlation between previous and next frames in video coding, disparity reflects the pixel motion relationship between the left and right views. Therefore, disparity compensation can be used to explicitly model the correlation between multi-view frames. For the left and right view images of the current frame, we have:

[0063] in, It indicates the prediction of the right view image obtained by using the parallax compensation method for the left view image. represents the left view image, Represents the right view disparity. For the left and right view disparity, considering that the disparity from the left view to the right view and the disparity from the right view to the left view should be in an inverse relationship in the absence of occlusion, disparity compensation can also be used to model the correlation between multi-view frames:

[0064] in, represents the predicted value of the right view disparity generated based on the left view disparity, represents the left disparity, Indicates right parallax.

[0065] In addition, considering the occlusion problem, using only parallax compensation will introduce noise. Therefore, we can introduce the semantic correlation of parallax between perspectives based on the physical meaning of parallax:

[0066] in, Indicates the semantic relevance of the left perspective of the sender, represents a learnable scaling factor, represents the left disparity, Indicates right parallax.

[0067] Then, using semantic relevance Weighted semantic fusion is performed to obtain the left disparity context of the current frame:

[0068] in, The left disparity context of the current frame, represents the left disparity of the current frame, represents the left disparity semantic feature of the current frame, represents the right disparity semantic feature of the current frame, Indicates the semantic relevance of the left perspective of the sender.

[0069] Therefore, the sender performs disparity compensation and semantic aggregation based on the left disparity of the current frame, the right disparity semantic features of the current frame, the left disparity semantic features of the current frame, and the semantic relevance of the sender's left perspective of the current frame to obtain the left disparity context of the current frame.

[0070] S203-2: The sender performs disparity compensation and semantic aggregation based on the right disparity of the current frame, the left disparity semantic features of the current frame, the right disparity semantic features of the current frame, and the sender's right perspective semantic relevance of the current frame to obtain the right disparity context of the current frame; the sender's right perspective semantic relevance of the current frame is determined based on the right disparity of the current frame and the sender's disparity from the left perspective to the right perspective of the current frame.

[0071] In this embodiment, the calculation process of the right disparity context of the current frame is similar to the calculation process of the left disparity context. The calculation method of the right view semantic relevance at the sending end is:

[0072] in, Indicates the semantic relevance of the right perspective of the sender, represents a learnable scaling factor, represents the left disparity, Indicates right parallax.

[0073] Then, using semantic relevance Weighted semantic fusion is performed to obtain the right disparity context of the current frame:

[0074] in, The right disparity context of the current frame, represents the right disparity of the current frame, represents the left disparity semantic feature of the current frame, represents the right disparity semantic feature of the current frame, Indicates the semantic relevance of the right perspective of the sender.

[0075] Regarding parallax, this embodiment uses feature-based parallax compensation and semantic aggregation methods to generate contextual information between the left and right views to achieve semantic interaction between views, effectively eliminating the noise introduced by parallax compensation due to occlusion, and then effectively compressing the redundancy between perspectives through cross-view information flow, thereby improving bandwidth transmission efficiency.

[0076] 1.2.2 Please refer to Figure 3 , Figure 3 Schematic diagram of the image coding architecture of the transmitting end in one embodiment of the present application. Figure 3 As shown in Figure 1, the transmitter includes an image feature extraction module, a cross-view parallel image semantic feature autoencoder, and a buffer. Specifically, the image encoding process at the transmitter includes: S301: The sending end extracts features from the left and right view images of the current frame through an image feature extraction module to obtain semantic features of the left and right view images of the current frame.

[0077] In this embodiment, the sending end downsamples the left and right view images of the current frame by 8 times through the image feature extraction module, and extracts image features through cascaded depthwise separable convolution to obtain semantic features of the left and right view images of the current frame.

[0078] S302: The sender reads the sender-decoded left and right view image semantic features of the previous frame from the buffer.

[0079] like Figure 3As shown, the buffer is used to store not only the left and right disparities and left and right disparity semantic features of the video image processed during the semantic communication process, but also the left and right perspective image semantic features.

[0080] In this embodiment, the sending end decodes the semantic features of the left and right view images, which means that the sending end decodes the potential representation of the left and right view images of the previous frame through the first cross-view parallel image semantic feature self-decoder to obtain the semantic features of the left and right view images of the previous frame.

[0081] S303: The transmitter uses a cross-view parallel image semantic feature autoencoder to encode the left and right view image semantic features of the current frame in parallel according to the left and right view image semantic features of the current frame, the left and right view image semantic features decoded by the transmitter of the previous frame, and the left and right view image context of the current frame, to obtain the potential representation of the left and right view images of the current frame.

[0082] In this embodiment, the left and right view image context of the current frame refers to image context information between the left and right view images generated by parallax compensation and semantic aggregation.

[0083] like Figure 3 As shown in the figure, when the transmitter encodes the image of the current frame through the cross-view parallel image semantic feature autoencoder, it needs to combine three feature indicators: the semantic features of the left and right view images of the current frame (i.e. Figure 3 in and ), the transmitter decodes the semantic features of the left and right view images of the previous frame (i.e. Figure 3 in and ), and the sending end left and right perspective image context of the current frame (i.e. Figure 3 in and ) is encoded to obtain the potential representation of the left and right view images of the current frame (i.e. Figure 3 in and ), transmitted to the receiving end for view reconstruction.

[0084] This embodiment encodes the semantic features of the left and right view images of the current frame in parallel based on the semantic features of the left and right view images of the current frame, the semantic features of the left and right view images decoded by the transmitter of the previous frame, and the left and right view image context of the current frame, so as to achieve image semantic interaction between the left and right views, and further effectively compress the redundancy between viewpoints through the cross-view information flow.

[0085] In addition to the cross-view parallel image semantic feature autoencoder, the transmitter also deploys a first cross-view parallel image semantic feature autodecoder to decode the latent representations of the left and right view images of the current frame, obtaining the transmitter-decoded left and right view images and the transmitter-decoded semantic features of the current frame. The transmitter then caches the transmitter-decoded semantic features of the current frame in a buffer for use in encoding the semantic features of the left and right view images of the next frame.

[0086] In this embodiment, a first cross-view parallel image semantic feature autodecoder is deployed at the transmitter to decode the left and right view images and their semantic features in real time, fully synchronized with the receiver. The decoded left and right view images and their semantic features are cached in a buffer, rather than the original encoded data. This ensures that the context used for encoding the next frame is consistent with that used by the receiver, forming a closed-loop feedback loop.

[0087] Furthermore, the left and right image contexts of the current frame used for encoding the current frame image in step S303 include: the sending end left view image context of the current frame and the sending end right view image context of the current frame. The calculation process of the sending end left and right view image contexts of the current frame is as follows: S303-1: The transmitter performs disparity compensation and semantic aggregation based on the transmitter-decoded left disparity of the current frame, the semantic features of the right-view image of the current frame, the semantic features of the left-view image of the current frame, and the transmitter-left-view semantic relevance of the current frame to obtain the transmitter-left-view image context of the current frame.

[0088] S303-2: The transmitter performs disparity compensation and semantic aggregation based on the transmitter-decoded right disparity of the current frame, the semantic features of the left-view image of the current frame, the semantic features of the right-view image of the current frame, and the transmitter-right-view semantic correlation of the current frame to obtain the transmitter-right-view image context of the current frame.

[0089] In this embodiment, the left view image context of the sending end is calculated as follows:

[0090] in, The left view image context of the current frame, Indicates the left disparity of the current frame decoded by the sender. Represents the semantic features of the left-view image of the current frame, Represents the semantic features of the right view image of the current frame, Indicates the semantic relevance of the sender's left view of the current frame.

[0091] Similarly, the sending end's right view image context is calculated as follows:

[0092] in, The right view image context of the current frame, Indicates the sender's decoded right disparity of the current frame. Represents the semantic features of the left-view image of the current frame, Represents the semantic features of the right view image of the current frame, Indicates the semantic relevance of the sender's right view of the current frame.

[0093] For images, this embodiment uses feature-based parallax compensation and semantic aggregation methods to generate contextual information between the left and right views to achieve semantic interaction between views, effectively eliminating the noise introduced by parallax compensation due to occlusion, and then effectively compressing the redundancy between perspectives through cross-view information flow, thereby improving bandwidth transmission efficiency.

[0094] It should be noted that the image encoding process at the sending end is similar to the disparity encoding process at the sending end. For the similarities, please refer to the relevant content of disparity encoding in Section 1.2.1 above.

[0095] 1.3 The process of the transmitter sending the potential representations of left and right views to the receiver in the semantic communication system specifically includes: S103-1: The sending end quantizes the left and right view potential representations to obtain quantized left and right view potential representations.

[0096] In this embodiment, the potential representation of left and right perspectives includes the potential representation of left and right perspective disparity and the potential representation of left and right perspective images. The encoding, compression and transmission processes of the two are similar. Here, the transmission process of the potential representation of left and right perspective disparity is taken as an example to illustrate the transmission process of the entire potential representation of left and right perspectives.

[0097] Please refer to Figure 4 , Figure 4 FIG. 1 is a schematic diagram of an architecture for transmitting potential representations of left and right perspective parallax in one embodiment of the present application. Figure 4 As shown, in this embodiment, the transmitter discretizes the continuous potential representation by quantizing the potential representation of the left and right view disparity, so as to facilitate subsequent entropy coding.

[0098] S103-2: The transmitter estimates the distribution parameters of the left and right view potential representations through a super prior network based on the left and right view potential representations.

[0099] In this embodiment, the Hyperprior Network includes an encoding part and a decoding part. The encoding part performs dimensionality reduction processing on the input quantized left and right view potential representations through a lightweight CNN and outputs a hyperprior potential representation. . Then through the decoding part Decoded into distribution parameters (e.g. mean ,scale ), to describe the probability distribution of each position of the quantized left and right view potential representation.

[0100] S103-3: The transmitter performs entropy coding on the quantized left and right view potential representations according to the distribution parameters to obtain a bitstream.

[0101] In this embodiment, the super prior potential representation is first Independent quantization and entropy coding (such as arithmetic coding) are performed, and then the probability of each position of the quantized left and right view potential representations is calculated according to the distribution parameters. Arithmetic coding (such as ANS) is used to compress the quantized left and right view potential representations into a bit stream, where symbols with higher probabilities are assigned shorter codewords.

[0102] S103-4: The transmitting end sends the bit stream and distribution parameters to the receiving end.

[0103] S103-5: The receiving end processes the received bit stream according to the distribution parameters to obtain the received left and right view potential representations.

[0104] In this embodiment, the transmitting end sends the bit stream and distribution parameters to the receiving end. The receiving end processes the bit stream using arithmetic decoding based on the distribution parameters to restore the left and right view potential representations. Then, the inverse quantization process is performed to obtain the reconstructed left and right view potential representations (e.g. Figure 4 The left view disparity latent representation in and right view disparity latent representation ), which is used for subsequent 3DGS reconstruction.

[0105] This embodiment uses a hyper-prior network to perform efficient entropy coding on the potential representations of the left and right perspectives, and uses its statistical characteristics to reduce the number of transmission bits. Thus, the left and right feature maps are compressed into low-dimensional vectors through quantization and entropy coding, thereby reducing the transmission bandwidth and promoting the realization of real-time semantic communication.

[0106] 1.4 Decoding process at the receiving end: In this solution, the receiving end deploys two decoders: a second cross-view parallel disparity semantic feature self-decoder for decoding the received left and right view disparity potential representations, and a second cross-view parallel image semantic feature self-decoder for decoding the received left and right view image potential representations and reconstructing the left and right view images. In one embodiment, the second cross-view parallel disparity semantic feature self-decoder is the same as the first cross-view parallel disparity semantic feature self-decoder in Section 1.2 above, and the second cross-view parallel image semantic feature self-decoder is the same as the first cross-view parallel image semantic feature self-decoder in Section 1.2 above.

[0107] Specifically, the receiving end decodes the received left and right view potential representations to obtain decoded semantic features and reconstructed left and right view images, including: The receiving end decodes the received left and right view disparity potential representations through the second cross-view parallel disparity semantic feature self-decoder to obtain the receiving end decoded left and right disparities of the current frame and the receiving end decoded left and right disparities semantic features of the current frame; the receiving end caches the receiving end decoded left and right disparities of the current frame to the buffer for left and right view image reconstruction of the next frame.

[0108] In this embodiment, on the one hand, the receiving end needs to perform disparity decoding. Specifically, the receiving end decodes the received left and right perspective disparity potential representations through the second cross-perspective parallel disparity semantic feature self-decoder, and obtains the receiving end-decoded left and right disparities of the current frame and the receiving end-decoded left and right disparity semantic features of the current frame. This process is similar to the process in Section 1.2 where the sending end decodes the left and right perspective disparity potential representations of the current frame through the first cross-perspective parallel disparity semantic feature self-decoder, and obtains the sending end-decoded left and right disparities of the current frame and the sending end-decoded left and right disparity semantic features of the current frame, so it will not be repeated here. And similarly, the receiving end caches the receiving end-decoded left and right disparities of the current frame into the buffer for left and right perspective image reconstruction of the next frame.

[0109] On the other hand, the receiving end also needs to decode the image and reconstruct the image based on the decoded disparity semantic features and image semantic features. Specifically, this process mainly includes: S401: The receiving end decodes the received latent representations of the left and right view images to obtain the receiving end decoded left and right view images of the current frame and the receiving end decoded left and right view images semantic features of the current frame.

[0110] This process is similar to the process in Section 1.2.2 where the transmitter decodes the potential representations of the left and right view images of the current frame through the first cross-view parallel image semantic feature self-decoder to obtain the transmitter-decoded left and right view images of the current frame and the transmitter-decoded semantic features of the left and right view images of the current frame. It will not be repeated here.

[0111] In this embodiment, the received left and right view images decoded by the receiving end are the same as the left and right view images decoded by the sending end, and the semantic features of the received left and right view images of the current frame are the same as the semantic features of the left and right view images decoded by the sending end.

[0112] S402: The receiving end reads the semantic features of the left and right view images of the previous frame from the buffer.

[0113] The process is described above in step S302, which will not be repeated here.

[0114] S403: The receiving end uses the second cross-view parallel image semantic feature self-decoder to perform parallel reconstruction of the left and right view images of the current frame according to the semantic features of the received left and right view images of the current frame, the semantic features of the received left and right view images of the previous frame, and the context of the received left and right view images of the current frame, to obtain the reconstructed left and right view images of the current frame and the semantic features of the reconstructed left and right view images of the current frame.

[0115] In this embodiment, the context of the received left and right view images of the current frame refers to the image context information between the received left and right view images generated by the receiving end through disparity compensation and semantic aggregation.

[0116] Specifically, when the receiving end reconstructs the left and right view images of the current frame using the second cross-view parallel image semantic feature self-decoder, it needs to combine three feature indicators: the semantic features of the received left and right view images of the current frame, the semantic features of the received left and right view images of the previous frame, and the context of the received left and right view images of the current frame, to reconstruct the left and right view images of the current frame through disparity compensation and semantic aggregation to obtain higher quality.

[0117] In addition, the receiving end also caches the semantic features of the received left and right view images of the current frame to the buffer for use in the reconstruction of the left and right view images of the next frame.

[0118] Further, the context of the received left and right view images of the current frame used in the reconstruction of the current frame in step S403 includes the context of the received left view image of the current frame and the context of the received right view image of the current frame. The calculation process of the context of the received left and right view images of the current frame is as follows: S403-1: The receiving end performs disparity compensation and semantic aggregation according to the receiving end decoded left disparity of the current frame, the receiving end decoded right view image semantic feature of the current frame, the receiving end decoded left view image semantic feature of the current frame, and the receiving end left view semantic correlation of the current frame to obtain the receiving end left view image context of the current frame. The receiving end left view semantic correlation of the current frame is determined according to the receiving end decoded left disparity of the current frame and the receiving end disparity from the right view to the left view of the current frame.

[0119] S403-2: The receiving end performs disparity compensation and semantic aggregation according to the receiving end decoded right disparity of the current frame, the receiving end decoded left view image semantic feature of the current frame, the receiving end decoded right view image semantic feature of the current frame, and the receiving end right view semantic correlation of the current frame to obtain the receiving end right view image context of the current frame. The receiving end right view semantic correlation of the current frame is determined according to the receiving end decoded right disparity of the current frame and the receiving end disparity from the left view to the right view of the current frame.

[0120] The calculation method of the receiving end left and right view image context of the current frame is the same as the calculation method of the left and right view image context of the current frame, and details are described above with reference to steps S303-1 and S303-2, which will not be repeated here.

[0121] 1.5 The process of predicting the Gaussian parameter by the receiving end according to the decoded semantic feature and the reconstructed left and right view images, comprising: S501: The receiving end predicts the center position according to the receiving end decoded left and right disparity of the current frame.

[0122] Please refer to Figure 5 , Figure 5 is the architecture schematic diagram of the Gaussian parameter prediction in an embodiment of the present application. Since the left and right views share the parameters, and the process of predicting the Gaussian parameter is independent at each time step, in order to express concisely, the time t and its upper and lower indexes are omitted in Figure 5 .

[0123] As shown in Figure 5 , the receiving end first predicts the center position according to the receiving end decoded left and right disparity of the current frame. Specifically, after the left and right disparity is decoded, the disparity map can be converted into the depth map of the view through linear transformation according to the camera parameters, and the depth map can inversely project the pixel-defined Gaussian point to the world coordinate system, so as to calculate the center position of the Gaussian ellipsoid.

[0124] S502: The receiving end predicts the color of the Gaussian ellipsoid according to the left and right view reconstructed images of the current frame.

[0125] In this embodiment, considering that the light on the human body surface is mainly diffuse reflection, only the 0th order spherical harmonic function, that is, the decoded left and right perspectives are used to reconstruct the image (i.e. Figure 5 in ) as the color of the Gaussian ellipsoid (i.e. Figure 5 in ) to reduce the complexity of the model.

[0126] S503: The receiving end predicts the covariance matrix and opacity according to the receiving end decoded left and right disparity semantic features of the current frame and the left and right viewing angles reconstructed image semantic features of the current frame.

[0127] In this embodiment, the decoded left and right disparity semantic features (i.e. Figure 5 in ) and left and right perspectives to reconstruct image semantic features (i.e. Figure 5 in ), a small number of depth-separable convolution modules can be used to predict the covariance matrix of the 3D Gaussian (decomposed into scale and rotation, i.e. Figure 5 S and R in ) and opacity (i.e. Figure 5 O in):

[0128]

[0129]

[0130] in, Used to represent the shape of the Gaussian ellipsoid, Used to represent the direction of the Gaussian ellipsoid, Used to represent the opacity of the Gaussian ellipsoid, represents the maximum size of the Gaussian ellipsoid, A module that preprocesses the semantic features of images and disparity in Gaussian parameter prediction. , and It is the head network used to obtain the predicted parameters, which consists of only one depth-wise separable convolution. Indicates that the receiving end decodes the left and right disparity semantic features. For the current frame, Including the receiving end decoding left disparity semantic features and the receiving end decodes the right disparity semantic features . Represents the semantic features of the left and right perspectives of the current frame. For the current frame, Including the receiving end decoding left disparity semantic features and the receiving end decodes the right disparity semantic features .

[0131] Finally, according to the value range of different attributes, the corresponding Gaussian parameters are obtained using the corresponding activation function or normalization method, and the prediction process is parallel processing of left and right view branches to reduce processing delay.

[0132] According to the decoded semantic features and the reconstructed left and right view images, the Gaussian parameters are predicted, and then the images of the user's view are rendered according to the predicted Gaussian parameters. For specific rendering process, please refer to the prior art, which is not limited herein.

[0133] The embodiment encodes the left and right view videos and predicts the Gaussian parameters in parallel through the cross-view parallel semantic feature autoencoder, reduces the transmission bandwidth, and improves the real-time performance of semantic communication.

[0134] Please refer to Figure 6 , based on the same inventive concept, the second aspect of the embodiment of the application provides a three-dimensional Gaussian splash multi-view video joint semantic encoding device, the three-dimensional Gaussian splash multi-view video joint semantic encoding device 600 comprises: The compression module 601 is configured to perform left and right view semantic feature encoding in parallel for left and right view images through a cross-view parallel semantic feature autoencoder at a sending end in a semantic communication system, to obtain left and right view latent representations, wherein the cross-view parallel semantic feature autoencoder reduces inter-view redundancy by transmitting cross-view context information. The transmission module 602 is configured to send the left and right view latent representations to a receiving end in the semantic communication system by the sending end. The reconstruction module 603 is configured to decode the received left and right view latent representations to obtain decoded semantic features and reconstructed left and right view images, predict Gaussian parameters according to the decoded semantic features and the reconstructed left and right view images, and render images of a user's view according to the predicted Gaussian parameters.

[0135] For the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the related parts are referred to the part of the method embodiment.

[0136] In the third aspect, based on the same inventive concept, please refer to Figure 7 The electronic device 700 comprises a processor 701 and a memory 702, the memory 702 stores machine executable instructions executable by the processor 701, and the processor 701 is configured to execute the machine executable instructions to implement the three-dimensional Gaussian splash multi-view video joint semantic encoding method according to the first aspect of the application.

[0137] It should be noted that the specific implementation of the electronic device 700 in the embodiment of the present application refers to the specific implementation of the three-dimensional Gaussian splashing multi-view video joint semantic coding method proposed in the first aspect of the embodiment of the present application, and will not be repeated here.

[0138] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can be referenced to each other.

[0139] Those skilled in the art will appreciate that the embodiments of the present application can be provided as methods, devices, or computer program products. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0140] The embodiments of the present application are described with reference to the flowcharts and / or block diagrams of the methods, terminal devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing terminal device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing terminal device generate instructions for implementing the steps in the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0141] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing terminal device to operate in a specific manner, so that the instructions stored in the computer readable memory produce a manufactured product including an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.

[0142] These computer program instructions can also be loaded onto a computer or other programmable data processing terminal device so that a series of operating steps are executed on the computer or other programmable terminal device to produce a computer-implemented process, thereby providing instructions for executing on the computer or other programmable terminal device to implement the process. Figure 1 a process or multiple processes and / or boxes Figure 1steps of the functions specified in the one or more blocks.

[0143] While the preferred embodiments of the application have been described above, it should be understood that many modifications and variations to these embodiments will be apparent to those skilled in the art once they learn of the basic inventive concepts. Therefore, the attached claims are intended to cover all such modifications and variations.

[0144] Finally, it is to be understood that the phraseology or terminology employed herein, such as "first" and "second", etc., are for descriptive purposes only and should not be construed to be indicative of a necessary order of occurrence, the relationships or sequences of elements or steps, etc., unless expressly so limited. Moreover, the terms "comprising", "including", or any other variation thereof, are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements does not include only those elements but can include other elements not expressly listed or inherent to such process, method, article, or apparatus. An element proceeded by "comprises a... " does not, without more constraints, exclude the existence of additional identical elements in the process, method, article, or apparatus that comprises the recited element.

[0145] The above describes in detail a three-dimensional Gaussian splash multi-view video joint semantic coding method provided by the application. The principles and implementation manners of the application are described by using specific examples. The above description of the embodiments is only used to help understand the method of the application and its core idea. Meanwhile, for those skilled in the art, the specific implementation manners and application ranges can be changed according to the idea of the application. In conclusion, the content of the specification should not be understood as a limitation of the application.

Claims

1. A three-dimensional Gaussian splatter multi-view video joint semantic coding method, characterized by: The method comprises: The transmitter in the semantic communication system uses a cross-view parallel semantic feature autoencoder to encode the left and right view semantic features in parallel for the left and right view images to obtain left and right view potential representations. The cross-view parallel semantic feature autoencoder reduces inter-view redundancy by transmitting cross-view context information. The sending end sends the left and right perspective potential representations to the receiving end in the semantic communication system; The receiving end decodes the received left and right perspective potential representations to obtain decoded semantic features and reconstructed left and right perspective images, predicts Gaussian parameters based on the decoded semantic features and the reconstructed left and right perspective images, and renders the image from the user's perspective based on the predicted Gaussian parameters.

2. The method according to claim 1, characterized in that The cross-view parallel semantic feature autoencoder at least includes: a cross-view parallel disparity semantic feature autoencoder; The transmitter in the semantic communication system uses a cross-view parallel semantic feature autoencoder to encode the left and right view semantic features in parallel for the left and right view images, and obtains the left and right view potential representations, which at least include: The transmitting end performs disparity estimation on the left and right perspective images of the current frame through a disparity estimation module to obtain the left and right disparities of the current frame and the left and right disparity semantic features of the current frame; The sending end reads the sending end decoded left and right disparity semantic features of the previous frame from the buffer; The transmitter encodes the left and right perspective disparity semantic features of the current frame in parallel using the cross-perspective parallel disparity semantic feature autoencoder according to the left and right disparity semantic features of the current frame, the left and right disparity semantic features decoded by the transmitter of the previous frame, and the left and right disparity context of the current frame, to obtain a left and right perspective disparity potential representation of the current frame; The method further comprises: The transmitter decodes the left and right view disparity potential representations of the current frame through a first cross-view parallel disparity semantic feature self-decoder to obtain the transmitter-decoded left and right disparities of the current frame and the transmitter-decoded left and right disparity semantic features of the current frame; The sending end caches the left and right disparity semantic features decoded by the sending end of the current frame to the buffer zone for encoding the left and right perspective disparity semantic features of the next frame, and the sending end caches the left and right disparity semantic features decoded by the sending end of the current frame to the buffer zone for encoding the left and right perspective image semantic features of the next frame.

3. The method according to claim 2, characterized in that The left and right disparity contexts of the current frame include: a left disparity context of the current frame and a right disparity context of the current frame; The method further comprises: The transmitter performs disparity compensation and semantic aggregation based on the left disparity of the current frame, the right disparity semantic feature of the current frame, the left disparity semantic feature of the current frame, and the transmitter left perspective semantic relevance of the current frame to obtain the left disparity context of the current frame; the transmitter left perspective semantic relevance of the current frame is determined based on the left disparity of the current frame and the transmitter disparity from the right perspective to the left perspective of the current frame; The sending end performs disparity compensation and semantic aggregation based on the right disparity of the current frame, the left disparity semantic features of the current frame, the right disparity semantic features of the current frame, and the sending end right perspective semantic correlation of the current frame to obtain the right disparity context of the current frame; the sending end right perspective semantic correlation of the current frame is determined based on the right disparity of the current frame and the sending end disparity from the left perspective to the right perspective of the current frame.

4. The method according to claim 2, characterized in that The cross-view parallel semantic feature autoencoder further comprises: a cross-view parallel image semantic feature autoencoder; The transmitter in the semantic communication system uses a cross-view parallel semantic feature autoencoder to encode the left and right view semantic features in parallel for the left and right view images to obtain the left and right view potential representations. It also includes: The transmitting end extracts features from the left and right perspective images of the current frame through an image feature extraction module to obtain semantic features of the left and right perspective images of the current frame; The sending end reads the semantic features of the left and right viewing angle images decoded by the sending end of the previous frame from the buffer; The transmitting end uses the cross-view parallel image semantic feature autoencoder to encode the left and right view image semantic features of the current frame in parallel according to the left and right view image semantic features of the current frame, the left and right view image semantic features decoded by the transmitting end of the previous frame, and the left and right view image context of the transmitting end of the current frame, to obtain the left and right view image potential representations of the current frame; The method further comprises: The transmitter decodes the potential representations of the left and right view images of the current frame through the first cross-view parallel image semantic feature self-decoder to obtain the transmitter-decoded left and right view images of the current frame and the transmitter-decoded left and right view image semantic features of the current frame; The transmitting end caches the semantic features of the left and right viewing angle images decoded by the transmitting end of the current frame in the buffer zone for encoding the semantic features of the left and right viewing angle images of the next frame.

5. The method according to claim 4, characterized in that The transmitting end left and right view image contexts of the current frame include: the transmitting end left view image context of the current frame and the transmitting end right view image context of the current frame; The method further comprises: The transmitter performs disparity compensation and semantic aggregation based on the transmitter-decoded left disparity of the current frame, the semantic features of the right view image of the current frame, the semantic features of the left view image of the current frame, and the transmitter-left view semantic relevance of the current frame to obtain the transmitter-left view image context of the current frame; The transmitter performs disparity compensation and semantic aggregation based on the transmitter-decoded right disparity of the current frame, the semantic features of the left-view image of the current frame, the semantic features of the right-view image of the current frame, and the transmitter-right-view semantic correlation of the current frame to obtain the transmitter-right-view image context of the current frame.

6. The method according to claim 2, characterized in that The receiving end decodes the received left and right view potential representations to obtain decoded semantic features and reconstructed left and right view images, which at least includes: The receiving end decodes the received left and right view disparity potential representations through a second cross-view parallel disparity semantic feature self-decoder to obtain the receiving end decoded left and right disparities of the current frame and the receiving end decoded left and right disparity semantic features of the current frame; The receiving end caches the receiving end-decoded left and right disparities of the current frame in the buffer zone for use in reconstructing the left and right viewing angle images of the next frame.

7. The method according to claim 6, characterized in that The receiving end decodes the received left and right view potential representations to obtain decoded semantic features and reconstructed left and right view images, further comprising: The receiving end decodes the received potential representations of the left and right view images to obtain the receiving end decoded left and right view images of the current frame and the receiving end decoded left and right view images semantic features of the current frame; The receiving end reads the semantic features of the left and right viewing angle images decoded by the receiving end of the previous frame from the buffer; The receiving end, through the second cross-view parallel image semantic feature self-decoder, performs, in parallel, left and right view image reconstruction of the current frame based on the semantic features of the left and right view images decoded by the receiving end of the current frame, the semantic features of the left and right view images decoded by the receiving end of the previous frame, and the left and right view image context of the current frame, to obtain left and right view reconstructed images of the current frame and semantic features of the left and right view reconstructed images of the current frame; The method further comprises: The receiving end caches the semantic features of the left and right viewing angle images decoded by the receiving end of the current frame in the buffer zone for use in reconstructing the left and right viewing angle images of the next frame.

8. The method according to claim 7, characterized in that The receiving end left and right perspective image contexts of the current frame include: the receiving end left perspective image context of the current frame and the receiving end right perspective image context of the current frame; The method further comprises: The receiving end performs disparity compensation and semantic aggregation based on the receiving end decoded left disparity of the current frame, the receiving end decoded right view image semantic features of the current frame, the receiving end decoded left view image semantic features of the current frame, and the receiving end left view semantic relevance of the current frame to obtain the receiving end left view image context of the current frame; the receiving end left view semantic relevance of the current frame is determined based on the receiving end decoded left disparity of the current frame and the receiving end disparity from the right view to the left view of the current frame; The receiving end performs disparity compensation and semantic aggregation based on the receiving end decoded right disparity of the current frame, the receiving end decoded left perspective image semantic features of the current frame, the receiving end decoded right perspective image semantic features of the current frame, and the receiving end right perspective semantic correlation of the current frame to obtain the receiving end right perspective image context of the current frame; the receiving end right perspective semantic correlation of the current frame is determined based on the receiving end decoded right disparity of the current frame and the receiving end disparity from the left perspective to the right perspective of the current frame.

9. The method according to claim 7, characterized in that The receiving end predicts Gaussian parameters based on the decoded semantic features and the reconstructed left and right view images, including: The receiving end predicts a center position according to left and right disparities decoded by the receiving end of the current frame; The receiving end reconstructs an image according to the left and right viewing angles of the current frame and predicts the color of the Gaussian ellipsoid; The receiving end predicts a covariance matrix and opacity according to the receiving end-decoded left and right disparity semantic features of the current frame and the left and right viewing angle-reconstructed image semantic features of the current frame.

10. The method according to any one of claims 1 to 9, characterized in that The sending end sends the left and right view potential representations to the receiving end in the semantic communication system, including: The transmitting end quantizes the left and right view potential representations to obtain quantized left and right view potential representations; The transmitting end estimates the distribution parameters of the left and right view potential representations through a super prior network according to the left and right view potential representations; The transmitting end performs entropy coding on the quantized left and right view potential representations according to the distribution parameters to obtain a bit stream; The transmitting end sends the bit stream and the distribution parameter to the receiving end; The receiving end processes the received bit stream according to the distribution parameters to obtain the received left and right perspective potential representations.

Citation Information

Patent Citations

  • Video coding method based on context inter-frame compression

    CN118474374A

  • Three-dimensional image coding method for machine vision

    CN119835395A

  • Gaussian point cloud rendering method and device, equipment and storage medium

    CN120166220A

  • Video generation method and device, electronic equipment, storage medium and program product

    CN120434373A

  • Image disparity estimation

    US20210142095A1

Cited By

  • 3D Gaussian Splitting compression method

    CN121193947A

  • 3d gaussian splash compression method

    CN121193947B