Data transmission method and related device

Through vector-matrix decomposition and neural data-related transformation based on neural radiation fields, the storage and rendering efficiency problems of existing systems when transmitting 3D data are solved, and efficient and robust 3D data transmission and view synthesis are achieved.

CN118842789BActive Publication Date: 2025-10-14BEIJING UNIV OF POSTS & TELECOMM
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410784881.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-06-18
Publication Date
2025-10-14
Estimated Expiration
2044-06-18

AI Technical Summary

Technical Problem

Existing data transmission systems have difficulty in effectively transmitting 3D data, especially in supporting diverse viewing angles and insufficient utilization of spatial correlation, resulting in transmission congestion.

Method used

A feature extraction method based on neural radiation field is adopted to decompose the three-dimensional image features into compact vector and matrix factors through vector-matrix decomposition and neural data correlation transformation, and then the source-channel joint encoder is used for encoding and decoding to achieve efficient 3D image transmission.

Benefits of technology

It reduces storage costs and memory consumption, improves rendering quality, and enables robust and efficient 3D scene transmission over variable wireless channels, supporting high-quality view synthesis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118842789B_ABST
    Figure CN118842789B_ABST
Patent Text Reader

Abstract

The application provides a data transmission method and related equipment. The method comprises the following steps: performing feature extraction on a three-dimensional image based on a pre-constructed neural radiation field to obtain a matrix feature and a vector feature; inputting the vector feature into a source channel joint encoder to obtain a vector symbol; inputting the matrix feature into the source channel joint encoder to obtain a matrix symbol; performing segmentation on the matrix symbol to obtain a main information symbol and a syntax information symbol; and performing decoding based on the syntax information symbol, the main information symbol and the vector symbol to obtain a recovered three-dimensional image, so as to complete transmission of the three-dimensional image. According to the embodiment of the application, the multi-dimensional voxel feature is reduced by performing vector decomposition on the features in the neural radiation field, and the low-dimensional features with insignificant spatial correlation are further compressed and transmitted by using a neural syntax after dimension reduction.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of data transmission, and in particular to a data transmission method and related equipment. BACKGROUND

[0002] In recent years, due to the demand for more immersive and interactive communication experience, the transmission of three dimensional (3D) data has become a new direction to promote the development of information and communication technology. The existing data transmission system is mainly customized for audio, image and video data, and it is difficult to realize the effective transmission of 3D data. Specifically, higher data dimension brings great challenges to the existing system. First, an effective 3D data representation technology is needed to support diversified observation angles. Second, the current feature extraction algorithm is not sufficient to utilize the spatial correlation in 3D scene data. Finally, due to the wide symbol demand of 3D data transmission, the retransmission mechanism used when errors are received can easily cause transmission congestion. SUMMARY

[0003] Therefore, the purpose of the present application is to provide a data transmission method and related equipment.

[0004] To achieve the above purpose, the present application provides a data transmission method, comprising:

[0005] performing feature extraction on a three-dimensional image based on a pre-constructed neural radiance field to obtain matrix features and vector features;

[0006] inputting the vector features into a source channel joint encoder to obtain vector symbols;

[0007] inputting the matrix features into the source channel joint encoder to obtain matrix symbols;

[0008] segmenting the matrix symbols to obtain primary information symbols and syntax information symbols;

[0009] decoding based on the syntax information symbols, the primary information symbols and the vector symbols to obtain a recovered three-dimensional image, so as to complete the transmission of the three-dimensional image.

[0010] In a possible implementation manner, the neural radiance field is constructed by the following method:

[0011] mapping spatial position coordinates of the three-dimensional image to corresponding three-dimensional density features by using a first function;

[0012] mapping observation directions of the three-dimensional image to three-dimensional appearance features by using a second function;

[0013] constructing the neural radiance field based on the first function and the second function.

[0014] In a possible implementation, the feature extraction on the three-dimensional image based on the pre-constructed neural radiance field includes:

[0015] The feature extraction on the three-dimensional image based on the pre-constructed neural radiance field includes:

[0016] The vector decomposition on the three-dimensional appearance feature includes:

[0017] The vector decomposition on the three-dimensional appearance feature includes:

[0018] In a possible implementation, the input of the vector feature into the source-channel joint encoder includes:

[0019] The input of the vector feature into the source-channel joint encoder includes:

[0020] In a possible implementation, the input of the matrix feature into the source-channel joint encoder includes:

[0021] The multiple down-sampling processing on the matrix feature in the source-channel joint encoder includes:

[0022] In a possible implementation, the segmentation of the matrix symbol includes:

[0023] The segmentation of the matrix symbol according to the channel of the matrix symbol includes:

[0024] The processing of the structure information by the syntax generator includes:

[0025] In a possible implementation, the decoding based on the syntax information symbol, the main information symbol and the vector symbol includes:

[0026] The adjustment of the matrix source-channel joint decoder based on the syntax information symbol includes:

[0027] The decoding of the vector symbol by the vector source-channel joint decoder includes:

[0028] restore the three-dimensional image based on the restored matrix feature and the restored vector feature.

[0029] Based on the same inventive concept, one or more embodiments of the present specification also provide a data transmission device, comprising:

[0030] a feature extraction module configured to extract features of a three-dimensional image based on a pre-constructed neural radiance field to obtain a matrix feature and a vector feature;

[0031] a first input module configured to input the vector feature into a source channel joint encoder to obtain a vector symbol;

[0032] a second input module configured to input the matrix feature into the source channel joint encoder to obtain a matrix symbol;

[0033] a segmentation module configured to segment the matrix symbol to obtain a main information symbol and a syntax information symbol;

[0034] a restoration module configured to decode based on the syntax information symbol, the main information symbol and the vector symbol to obtain a restored three-dimensional image, so as to complete transmission of the three-dimensional image.

[0035] Based on the same inventive concept, one or more embodiments of the present specification also provide an electronic device, comprising a memory, a processor and a computer program stored on the memory and executable on the processor, wherein the processor implements the data transmission method of any one of the above embodiments when executing the program.

[0036] Based on the same inventive concept, one or more embodiments of the present specification also provide a non-transitory computer readable storage medium, which stores computer instructions for causing the computer to execute the data transmission method of any one of the above embodiments.

[0037] From the above, it can be seen that the data transmission method and related equipment provided by the present application, based on the pre-constructed neural radiation field, perform feature extraction on the three-dimensional image to obtain matrix features and vector features; input the vector features into the source-channel joint encoder to obtain vector symbols; input the matrix features into the source-channel joint encoder to obtain matrix symbols; segment the matrix symbols to obtain main information symbols and grammatical information symbols; decode based on the grammatical information symbols, the main information symbols and the vector symbols to obtain the restored three-dimensional image to complete the transmission of the three-dimensional image. The embodiment of the present application can decompose the multi-dimensional features constructed by the neural radiation field into compact vectors and matrix factors through the vector-matrix decomposition (VM) technology, which can speed up the rendering time of NeRF and improve the rendering quality while reducing storage costs and memory consumption. In this way, the transmission problem of processing multi-dimensional data can be avoided. The feature matrix obtained by VM decomposition is compressed using a Neural Data-Dependent Transform, which constructs a data-dependent neural transform for it. By using an additional model stream to generate transformation parameters at the decoding end, the model can learn a more abstract neural grammar. The learned grammar helps to cluster the latent features of the matrix more compactly. This application can promote robust and efficient scene transmission over variable wireless channels and achieve high-quality view synthesis of 3D scenes at the receiving end. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] In order to more clearly illustrate the technical solutions in this application or related technologies, the following briefly introduces the drawings required for use in the embodiments or related technical descriptions. Obviously, the drawings described below are merely embodiments of this application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0039] Figure 1 This is a flowchart of a data transmission method according to an embodiment of the present application;

[0040] Figure 2 This is a schematic diagram of the overall flow of the data transmission method according to an embodiment of the present application;

[0041] Figure 3 This is a schematic diagram of an encoder according to an embodiment of the present application;

[0042] Figure 4 This is a schematic diagram of a matrix encoder according to an embodiment of the present application;

[0043] Figure 5 A schematic diagram comparing different data transmission schemes according to an embodiment of the present application;

[0044] Figure 6 FIG. 1 is a schematic diagram of a data transmission device according to an embodiment of the present application;

[0045] Figure 7 FIG. 2 is a schematic diagram of an electronic device according to an embodiment of the present application. DETAILED DESCRIPTION

[0046] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the embodiments and the accompanying drawings.

[0047] It should be noted that, unless otherwise defined, the technical terms or scientific terms used in the embodiments of the present application should be understood as the common meanings understood by those with ordinary skills in the art to which the present application belongs. The terms "first", "second" and similar terms used in the embodiments of the present application do not represent any order, number or importance, but are only used to distinguish different components. The terms "include" or "contain" and similar terms mean that the elements or objects before the terms encompass the elements or objects listed after the terms and their equivalents, and do not exclude other elements or objects. The terms "connect" or "connected" and similar terms are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect. The terms "upper", "lower", "left", "right" and the like are only used to represent relative positional relationships, and when the absolute positions of the described objects change, the relative positional relationships can also change accordingly.

[0048] As described in the background section, in recent years, due to the demand for more immersive and interactive communication experience, the transmission of 3D data has become a new direction to promote the development of information and communication technology. The existing data transmission system is mainly customized for audio, image and video data, and it is difficult to realize effective transmission of 3D data. Specifically, higher data dimension brings great challenges to the existing system. First, effective 3D data representation technology is needed to support diversified viewing angles. Second, the current feature extraction algorithm is not sufficient to utilize the spatial correlation in 3D scene data. Finally, due to the wide symbol demand of 3D data transmission, the retransmission mechanism used when errors are received is easy to cause transmission congestion.

[0049] In view of the above, the embodiment of the present application proposes a data transmission method, based on a pre-constructed neural radiance field, a three-dimensional image is feature extracted to obtain a matrix feature and a vector feature; the vector feature is input into a source channel joint encoder to obtain a vector symbol; the matrix feature is input into the source channel joint encoder to obtain a matrix symbol; the matrix symbol is segmented to obtain a main information symbol and a syntax information symbol; based on the syntax information symbol, the main information symbol and the vector symbol, decoding is performed to obtain a recovered three-dimensional image to complete transmission of the three-dimensional image. The vector-matrix decomposition technology can decompose the multi-dimensional features constructed by the neural radiance field into compact vector and matrix factors, which can reduce the storage cost and memory consumption, speed up the rendering time of NeRF and improve the rendering quality. In this way, the transmission problem of multi-dimensional data processing can be avoided. The neural data-dependent transformation is used to compress the feature matrix obtained by VM decomposition, a data-dependent neural transformation is constructed, a transformation parameter is generated at the decoding end by using an additional model stream, so that the model can learn more abstract neural syntax, and the learned syntax helps to more compactly cluster the hidden features of the matrix. The present application can promote robust and efficient scene transmission on variable wireless channels, and realize high-quality view synthesis of 3D scenes at the receiving end.

[0050] In the following, the technical solutions of the embodiments of the present application will be described in detail through specific embodiments.

[0051] Reference Figure 1 The data transmission method of the embodiment of the present application includes the following steps:

[0052] Step S101, based on a pre-constructed neural radiance field, a three-dimensional image is feature extracted to obtain a matrix feature and a vector feature;

[0053] Step S102, the vector feature is input into a source channel joint encoder to obtain a vector symbol;

[0054] Step S103, the matrix feature is input into the source channel joint encoder to obtain a matrix symbol;

[0055] Step S104, the matrix symbol is segmented to obtain a main information symbol and a syntax information symbol;

[0056] Step S105, based on the syntax information symbol, the main information symbol and the vector symbol, decoding is performed to obtain a recovered three-dimensional image to complete transmission of the three-dimensional image.

[0057] First, some technical terms of the present application are described.

[0058] Neural radiance field (NeRF) has gradually become the research center of 3D data representation in the field of artificial intelligence. The core idea of NeRF is to represent a continuous 3D scene as a 5D input neural radiance field. Specifically, NeRF maps spatial positions and viewing directions (i.e. 5D input, including three-dimensional coordinates (x, y, z) and observation direction ) to color and opacity (i.e. 4D output, including RGB color value and volume density) through a neural network. This mapping relationship allows NeRF to render new views in different ways using volume rendering techniques. In order to avoid the use of a large multi-layer perceptron (MLP) in classic NeRF, some subsequent researches use spatial data structures to store (multi-dimensional) neural features in NeRF, which are interpolated during the volume ray marching. These methods can use smaller MLPs or completely eliminate neural networks, improve rendering quality, and reduce the storage cost of 3D models.

[0059] Vector-matrix decomposition (VM) techniques can decompose the multi-dimensional features constructed by neural radiance field into compact vector and matrix factors, which can speed up the rendering time of NeRF and improve the rendering quality while reducing the storage cost and memory consumption. In this way, the transmission problem of processing multi-dimensional data can be avoided.

[0060] Neural Data-Dependent Transform is used to compress the feature matrix obtained by VM decomposition based on neural network, and a data-dependent neural transform is constructed for it. By using an additional model stream to generate transform parameters at the decoding end, the model can learn more abstract neural grammar, and the learned grammar can help to more compactly cluster the hidden features of the matrix.

[0061] Joint source-channel coding (JSCC) is a classic topic in information theory and coding theory. Traditional JSCC jointly designs source coding and channel coding, seeking to optimize and improve end-to-end, but for many years, due to the difficulty of manually designing coding schemes with compression and noise resistance, it has not been well developed. In recent years, with the development of artificial intelligence, neural networks can be used to replace part of the encoder and decoder. The sender encapsulates semantic feature extraction, source channel coding into an encoder module, and the receiver encapsulates source channel decoding and semantic feature fusion into a decoder module. At the same time, in the process of training the encoder and decoder, the distortion introduced by the channel transmission of the intermediate bottleneck layer data is considered, thereby increasing the ability of the encoder and decoder to resist channel noise, fading and other adverse factors.

[0062] In this application, there are two stages of offline training and online transmission. Offline training is mainly to build a neural radiation field and end-to-end train the neural radiation field transmission model. Online transmission is to use the trained transmission model to encode, simulate transmission and decode the neural features, and use the neural features at the receiving end to render the 3D scene in multiple views. The two stages are different in details. In the following, first introduce the construction of the neural radiation field offline and the end-to-end training of the neural feature transmission model: input the neural radiation field features constructed from 3D scene data into the 3DST model, output the reconstructed features, then calculate the mean square error between the original features and the reconstructed features as the loss function, and finally update the model parameters by back propagation.

[0063] Reference Figure 2 The data transmission method of the embodiment of the application is shown in the whole flowchart.

[0064] As Figure 2 shown, first, the volume density and color dispersion are decomposed to obtain vector features and matrix features. Then the vector features and the matrix features are input into the corresponding encoders. For the matrix features, segmentation processing is also needed to obtain main information symbols and structure information. Then the structure information is input into the semantic generator to obtain syntax information symbols. Then the vector symbols, the main information symbols and the syntax information symbols are all input into the wireless channel to obtain the corresponding symbols with noise. Then the syntax information symbols are used to train the convolution kernel parameter generator to adjust the parameters of the decoder corresponding to the matrix features. The vector features corresponding to the decoder are normally decoded to obtain the decoded results. Then the two are used to perform the outer product operation according to the functions involved in the construction of the neural radiation field in the foregoing steps to obtain the corresponding 3D images.

[0065] For step S101, first, the neural radiation field needs to be constructed.

[0066] In some embodiments, the neural radiance field is constructed by mapping spatial position coordinates of the three-dimensional image to corresponding three-dimensional density features using a first function, mapping viewing directions of the three-dimensional image to viewing-angle-dependent three-dimensional appearance features using a second function, and constructing the neural radiance field based on the first function and the second function.

[0067] In some embodiments, the feature extraction of the three-dimensional image based on the pre-constructed neural radiance field to obtain the matrix features and the vector features comprises: feature extraction of the three-dimensional image based on the pre-constructed neural radiance field to obtain the three-dimensional density features and the three-dimensional appearance features; vector decomposition of the three-dimensional density features to obtain density matrix features and density vector features; and vector decomposition of the three-dimensional appearance features to obtain appearance matrix features and appearance vector features.

[0068] Specifically, the present application uses a neural radiance field (NeRF) framework based on vector decomposition to construct the neural radiance field. The principle of this step is to find a function that maps any spatial position coordinates and viewing directions to their corresponding volume density σ and viewing-angle-dependent color scattering c. The present application conditions the neural field based on a 3D voxel grid, specifically, the present application uses a 3D voxel grid V σ to represent the volume density σ, and a 3D appearance feature V c with multiple channels to simulate the color scattering c that will change according to the viewing angle. Therefore, by applying a density activation function f σ and a color mapping function f c , a radiance field can be defined as: where f σ is usually a softplus activation function, f c can be represented by a spherical harmonics (SH) function, by sampling K samples from any camera viewing angle , the pixel point corresponding to the viewing angle can be rendered where is the density and color sampled by the K sample points, Δ i is the ray sampling interval, τ represents the cumulative transmittance, and c bg is the background color. By minimizing the error between the rendered pixel RGB value and the original pixel RGB value: the neural radiance field (hereinafter collectively referred to as ) can be optimized. Since Generally high-dimensional features, the present application utilizes vector decomposition to reduce the amount of data transmitted. Specifically, a 3D tensor can be represented as the sum of the outer products between R sets of mutually orthogonal matrices and vectors in the XYZ coordinate space (each set of matrices and vectors are perpendicular to each other), represented by the following equation:

[0069]

[0070] where, represents the vector parallel to the x-axis, represents the vector parallel to the y-axis, represents the vector parallel to the z-axis, represents the matrix in the yz plane, represents the matrix in the xz plane, represents the matrix in the xy plane.

[0071] Since the density activation function f σ and the color mapping function f c can be fixed functions, the transmitted symbols in the channel can be summarized as where, L RGB represents the error between the rendered pixel and the original pixel.

[0072] Next, the neural feature dataset is constructed. According to the above steps, the matrix features and the vector features derived from the neural radiance field of 25 different scenes are constructed and collected. Since are multi-channel data and the data between each channel does not have significant correlation, the present application divides the multi-channel matrix and vector features into single-channel data. Since each scene produces three matrix features and vector features corresponding to different coordinate planes, each feature consists of 64 channels, so there are a total of 4800 training matrix and training vector sets. During training, the matrix feature set will be randomly cropped and padded to 256x256 blocks, and the vector feature set will be randomly cropped and padded to 256 lengths. It should be noted that this is the processing of the neural feature dataset during the training process, and in the actual running process, only the three-dimensional density features and three-dimensional appearance features extracted from the image need to be subjected to corresponding vector decomposition.

[0073] Referring to Figure 3 , the encoder schematic diagram of the embodiment of the present application.

[0074] Further, the vector to be transmitted is input to the encoder represented by Figure 3 .

[0075] In some embodiments, the vector feature input source channel joint encoder to obtain the vector symbol, comprising: the vector feature input source channel joint encoder, the vector feature is processed by dimension, and the vector feature processed by dimension is passed through the signal-to-noise ratio attention module, and the vector symbol is obtained.

[0076] In the embodiment, since the vector feature space correlation is weak, in order to encode the symbol to have the ability to resist channel noise, the encoder adopts the design of dimensioning, that is, the one-dimensional convolutional neural network is used to convert the input vector of 1*256 into the output of 8*256, and the coded symbol has a certain robustness after passing through the SNR attention module.

[0077] Reference Figure 4 The matrix encoder of the embodiment of the application is shown in the schematic diagram.

[0078] In some embodiments, the matrix feature is input into the source channel joint encoder to obtain the matrix symbol, comprising: the matrix feature is processed by multiple downsampling in the source channel joint encoder to obtain the matrix symbol.

[0079] In some embodiments, the matrix symbol is segmented to obtain the main information symbol and the syntax information symbol, comprising: the matrix symbol is segmented according to the channel of the matrix symbol to obtain the main information symbol and the structure information; the structure information is processed by the syntax generator to obtain the syntax information symbol.

[0080] In the embodiment, the feature matrix to be transmitted is input into the matrix encoder Figure 4 The matrix encoder is shown in the figure, for a limited matrix feature data set, the application adopts a neural transform dependent manner to cluster the matrix feature. Specifically, through the matrix encoder, the multi-channel coding symbol is obtained by multiple downsampling After the channel segmentation The matrix content information symbol And the structure information, the structure information is generated by the syntax generator Here, the matrix encoder and the neural syntax generator are both implemented by using the convolutional neural network (CNN).

[0081] In some embodiments, the decoding based on the syntax information symbol, the main information symbol and the vector symbol to obtain the recovered three-dimensional image comprises: adjusting a matrix source channel joint decoder based on the syntax information symbol, and decoding the matrix symbol by using the adjusted matrix source channel joint decoder to obtain a recovered matrix feature; decoding the vector symbol by using a vector source channel joint decoder to obtain a recovered vector feature; and recovering an image based on the recovered matrix feature and the recovered vector feature to obtain the recovered three-dimensional image.

[0082] In the embodiment, the vector transmission symbol The matrix content information symbol And the syntax symbol After being transmitted through a wireless channel and being affected by the channel, the received symbol at the receiving end is And

[0083] Further, the 8xL vector received symbol is changed into a 1xL reconstructed vector by a vector decoder.

[0084] Further, the neural syntax received symbol First, it is changed into a set of 3x3 convolution kernels by a weight generator, which can be part of a matrix decoder, and an online mode decision mechanism is realized by syntax matching to jointly optimize the coding efficiency of each individual matrix. After that, the received matrix content information symbol is changed into a reconstructed matrix feature by its corresponding feature decoder

[0085] After that, the loss function of the end-to-end model is calculated, and the codec network parameters are adjusted by the Adam optimization algorithm. The above is an iterative optimization process of the neural radiance field-based end-to-end wireless 3D scene transmission framework (3DST) model during offline training. During training, the entire training set is traversed multiple times, and each traversal includes multiple iterative optimizations.

[0086] As for the actually running model, the reconstructed matrix and the reconstructed vector are used to perform an outer product operation to obtain a three-dimensional density feature and a three-dimensional appearance feature Each light sampling point is subjected to a density activation function and a spherical harmonic function to obtain the body density σ and the color dispersion c corresponding to the point. After that, the light ray is subjected to a volume rendering function to obtain the accumulated RGB value of the light ray, thereby realizing the recovery of the three-dimensional image.

[0087] Reference Figure 5 , for different data transmission schemes of the embodiments of the application.

[0088] It can be seen from the figure that 3DST is better than the JPEG scheme in all the range of AWGN signal-to-noise ratio under the given channel transmission bandwidth, better than the BPG scheme at low signal-to-noise ratio, and very close to the BPG scheme at high signal-to-noise ratio. BPG is currently the mainstream image and matrix compression performance of the first-class image encoding method, which can be used as the theoretical upper bound of the matrix feature transmission system. It can be seen that the 3DST scheme proposed in the application realizes excellent RD performance gain, and when the channel condition fluctuates sharply, 3DST can realize smooth reconstruction performance decline and will not cause decoding error, so 3DST will not cause transmission congestion due to retransmission mechanism caused by decoding error at the receiving end.

[0089] As can be seen from the above embodiments, the data transmission method described in the embodiments of the application extracts matrix features and vector features from a three-dimensional image based on a pre-constructed neural radiation field; inputs the vector features into a source channel joint encoder to obtain vector symbols; inputs the matrix features into the source channel joint encoder to obtain matrix symbols; splits the matrix symbols to obtain primary information symbols and syntax information symbols; decodes based on the syntax information symbols, the primary information symbols and the vector symbols to obtain a restored three-dimensional image, so as to complete transmission of the three-dimensional image. The vector-matrix decomposition technology can decompose the multi-dimensional features constructed by the neural radiation field into compact vector and matrix factors, which can reduce storage cost and memory consumption, speed up the rendering time of NeRF and improve the rendering quality. In this way, the transmission problem of processing multi-dimensional data can be avoided. The neural data-dependent transformation is used to compress the feature matrix obtained by VM decomposition, and a data-dependent neural transformation is constructed. By using an additional model stream to generate transformation parameters at the decoding end, the model can learn more abstract neural syntax, and the learned syntax can help to more compactly cluster the hidden features of the matrix. The application can promote robust and efficient scene transmission over variable wireless channels and realize high-quality view synthesis of 3D scenes at the receiving end.

[0090] It should be noted that the method of the embodiments of the application can be executed by a single device, such as a computer or a server. The method of the embodiments can also be applied to a distributed scenario, and completed by multiple devices cooperating with each other. In this distributed scenario, one of the multiple devices can only execute one or more steps in the method of the embodiments of the application, and the multiple devices will interact with each other to complete the method.

[0091] It is to be understood that the foregoing description is descriptive only. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in a different order and still achieve desirable results. Additionally, the processes depicted in the attached figures do not necessarily require the particular order shown or sequential order to achieve desirable results. In certain implementations, multitasking and parallel processing can be advantageous.

[0092] Based on the same inventive concept, the present application also provides a data transmission device corresponding to any of the above-mentioned embodiment methods.

[0093] Reference Figure 6 The data transmission device comprises:

[0094] The extraction module 61 is configured to perform feature extraction on the three-dimensional image based on a pre-constructed neural radiation field to obtain matrix features and vector features.

[0095] The first input module 62 is configured to input the vector features into a source channel joint encoder to obtain vector symbols.

[0096] The second input module 63 is configured to input the matrix features into the source channel joint encoder to obtain matrix symbols.

[0097] The segmentation module 64 is configured to segment the matrix symbols to obtain primary information symbols and syntax information symbols.

[0098] The recovery module 65 is configured to decode based on the syntax information symbols, the primary information symbols, and the vector symbols to obtain a recovered three-dimensional image, so as to complete transmission of the three-dimensional image.

[0099] For the convenience of description, the above device is described in various modules according to functions. Of course, in the implementation of the present application, the functions of each module can be implemented in one or more software and / or hardware.

[0100] The device of the above-mentioned embodiment is used to implement the corresponding data transmission method in any of the above-mentioned embodiments, and has the beneficial effects of the corresponding method embodiments, which are not described here.

[0101] Based on the same inventive concept, the present application also provides an electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the data transmission method of any one of the above-mentioned embodiments when executing the program.

[0102] Figure 7 A more specific electronic device hardware structure diagram provided by the embodiment is shown, which can include: a processor 1010, a memory 1020, an input / output interface 1030, a communication interface 1040, and a bus 1050. The processor 1010, the memory 1020, the input / output interface 1030, and the communication interface 1040 are connected to each other through the bus 1050 for internal communication.

[0103] The processor 1010 can be implemented by a general-purpose CPU (Central Processing Unit), a microprocessor, an ASIC (Application Specific Integrated Circuit), or one or more integrated circuits, etc., for executing related programs to implement the technical solutions provided by the embodiments of the present specification.

[0104] The memory 1020 can be implemented by a ROM (Read Only Memory), a RAM (Random Access Memory), a static storage device, a dynamic storage device, etc. The memory 1020 can store an operating system and other application programs, and when the technical solutions provided by the embodiments of the present specification are implemented by software or firmware, the related program codes are stored in the memory 1020 and called and executed by the processor 1010.

[0105] The input / output interface 1030 is used to connect input / output modules to realize information input and output. The input / output modules can be configured as components in the device (not shown in the figure) or externally connected to the device to provide corresponding functions. The input device can include a keyboard, a mouse, a touch screen, a microphone, various sensors, etc., and the output device can include a display, a speaker, a vibrator, an indicator light, etc.

[0106] The communication interface 1040 is used to connect a communication module (not shown in the figure) to realize the communication interaction between the device and other devices. The communication module can realize communication through a wired manner (such as USB, network cable, etc.) or through a wireless manner (such as mobile network, WIFI, Bluetooth, etc.).

[0107] The bus 1050 includes a channel for transmitting information between various components (such as the processor 1010, the memory 1020, the input / output interface 1030, and the communication interface 1040) of the device.

[0108] It should be noted that although the above device only shows the processor 1010, the memory 1020, the input / output interface 1030, the communication interface 1040 and the bus 1050, in the specific implementation process, the device can also include other components necessary for normal operation. In addition, those skilled in the art can understand that the above device can also only contain the components necessary to implement the embodiments of the present application, and does not necessarily contain all the components shown in the figure.

[0109] The electronic device of the above embodiment is used to implement the corresponding data transmission method in any of the preceding embodiments, and has the beneficial effects of the corresponding method embodiments, which are not described here.

[0110] Based on the same inventive concept, corresponding to any of the above embodiment methods, the present application also provides a non-transitory computer readable storage medium storing computer instructions for causing the computer to execute the data transmission method according to any of the above embodiments.

[0111] The computer readable medium of the present embodiment includes permanent and non-permanent, removable and non-removable media, which can be realized by any method or technology to store information. The information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transmission medium that can be used to store information that can be accessed by a computing device.

[0112] The computer instructions stored in the storage medium of the above embodiment are used to cause the computer to execute the data transmission method according to any of the above embodiments, and have the beneficial effects of the corresponding method embodiments, which are not described here.

[0113] Those skilled in the art should understand that the discussion of any of the above embodiments is only exemplary and is not intended to imply that the scope of the present application (including claims) is limited to these examples; under the idea of the present application, the above embodiments or technical features in different embodiments can also be combined, the steps can be implemented in any order, and there are many other changes of different aspects of the embodiments of the present application as described above. In order to be brief, they are not provided in detail.

[0114] Additionally, to simplify the description and discussion, and so as not to obscure the embodiments of the application being presented, the well-known functions or constructions of integrated circuit (IC) chips and other components can or can not be shown in the figures and will be omitted as not to unnecessarily obscure the embodiments of the application being presented. Moreover, the devices can be shown in block diagram form in order to avoid obscuring the embodiments of the application, and this also acknowledges the fact that the details in regard to the implementation of the block diagram devices are highly dependent on the platform within which the embodiments of the application are to be implemented (i.e., these details should be well within the purview of one of ordinary skill in the art). Where specific details are set forth in order to describe an illustrative embodiment of the application, it will be apparent to one of ordinary skill in the art that the embodiments of the application can be practiced without, or with variation of, these specific details. Thus, the description is to be considered as illustrative only and not restrictive in nature.

[0115] While the application has been described in connection with specific embodiments thereof, it will be understood that many modifications, substitutions and changes will be apparent to those of ordinary skill in the art once they have the benefit of the foregoing description. For example, other memory architectures (e.g., dynamic RAM (DRAM)) can use the embodiments discussed.

[0116] It is intended that the embodiments of the application encompass all such substitutions, modifications and variations as fall within the scope of the appended claims. Accordingly, any omission, modification, equivalent replacement, improvement, etc. made in the spirit and principle of the embodiments of the application should be included in the scope of protection of the application.

Claims

1. A data transmission method, characterized in that: include: Based on a pre-constructed neural radiation field, feature extraction is performed on a three-dimensional image to obtain matrix features and vector features, including: based on the pre-constructed neural radiation field, feature extraction is performed on the three-dimensional image to obtain three-dimensional density features and three-dimensional appearance features; vector decomposition is performed on the three-dimensional density features to obtain density matrix features and density vector features; vector decomposition is performed on the three-dimensional appearance features to obtain appearance matrix features and appearance vector features; Inputting the vector feature into a source-channel joint encoder to obtain a vector symbol; Inputting the matrix features into the source-channel joint encoder to obtain matrix symbols; Segmenting the matrix symbols to obtain main information symbols and grammatical information symbols; Decoding is performed based on the syntax information symbols, the main information symbols, and the vector symbols to obtain a restored three-dimensional image to complete the transmission of the three-dimensional image, including: adjusting a matrix source-channel joint decoder based on the syntax information symbols, and using the adjusted matrix source-channel joint decoder to decode the matrix symbols to obtain restored matrix features; using a vector source-channel joint decoder to decode the vector symbols to obtain restored vector features; and restoring the image based on the restored matrix features and the restored vector features to obtain the restored three-dimensional image.

2. The method according to claim 1, characterized in that The neural radiation field is constructed by the following method: Mapping the spatial position coordinates of the three-dimensional image to corresponding three-dimensional density features using a first function; mapping the observation direction of the three-dimensional image to a three-dimensional appearance feature using a second function; The neural radiation field is constructed based on the first function and the second function.

3. The method according to claim 1, characterized in that Inputting the vector feature into a source-channel joint encoder to obtain a vector symbol includes: The vector features are input into a source-channel joint encoder, the vector features are subjected to dimensionality increase processing, and the vector features after dimensionality increase processing are passed through a signal-to-noise ratio attention module to obtain the vector symbols.

4. The method according to claim 1, wherein Inputting the matrix features into the source-channel joint encoder to obtain matrix symbols includes: In the source-channel joint encoder, the matrix features are subjected to multiple downsampling processes to obtain the matrix symbols.

5. The method according to claim 1, wherein The segmenting of the matrix symbols to obtain main information symbols and grammatical information symbols includes: Segmenting the matrix symbol according to the channels of the matrix symbol to obtain the main information symbol and structural information; The structural information is processed by a grammar generator to obtain the grammar information symbol.

6. A data transmission device, characterized in that: include: The extraction module is configured to perform feature extraction on the three-dimensional image based on the pre-constructed neural radiation field to obtain matrix features and vector features, including: performing feature extraction on the three-dimensional image based on the pre-constructed neural radiation field to obtain three-dimensional density features and three-dimensional appearance features; performing vector decomposition on the three-dimensional density features to obtain density matrix features and density vector features; performing vector decomposition on the three-dimensional appearance features to obtain appearance matrix features and appearance vector features; A first input module is configured to input the vector feature into a source-channel joint encoder to obtain a vector symbol; A second input module is configured to input the matrix feature into the source-channel joint encoder to obtain a matrix symbol; a segmentation module configured to segment the matrix symbols to obtain main information symbols and grammatical information symbols; A recovery module is configured to decode based on the syntax information symbols, the main information symbols and the vector symbols to obtain a recovered three-dimensional image to complete the transmission of the three-dimensional image, including: adjusting a matrix source-channel joint decoder based on the syntax information symbols, and using the adjusted matrix source-channel joint decoder to decode the matrix symbols to obtain recovered matrix features; using a vector source-channel joint decoder to decode the vector symbols to obtain recovered vector features; and restoring the image based on the recovered matrix features and the recovered vector features to obtain the recovered three-dimensional image.

7. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the method according to any one of claims 1 to 5 is implemented.

8. A non-transitory computer-readable storage medium storing computer instructions, characterized in that: The computer instructions are used to enable a computer to execute the method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Neural radiation field model training method and device, human face generation method and device, and server

    CN113822969A

  • Bridge crack detection and crack three-dimensional visualization method and system

    CN116468683A