Data transmission method and device, electronic equipment and medium
Through the combination of transform encoding model and channel encoder, bandwidth resources are dynamically allocated, which solves the problem of insufficient data redundancy and robustness in traditional three-dimensional scene transmission, and realizes efficient and accurate three-dimensional scene reconstruction.
Patent Information
- Application Number
- CN202510342127.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-21
- Publication Date
- 2025-07-25
AI Technical Summary
Traditional three-dimensional scene transmission technology has high data redundancy and high modeling complexity. The existing solutions fail to effectively utilize the spatial correlation between perspectives, resulting in large amount of redundant data and low transmission efficiency, insufficient transmission robustness, and susceptible to channel noise.
The transform encoding model is used to compress the three-dimensional features, combine with the channel encoder to dynamically allocate bandwidth resources, reduce redundancy through the nonlinear transform encoding model, dynamically allocate channel resources, allocate long code words for important features, and use short code words for secondary features to optimize bandwidth occupation and improve transmission reliability.
It realizes efficient transmission and precise reconstruction of three-dimensional scenes, reduces data redundancy, improves the robustness of coding efficiency and reconstruction quality, and adapts to dynamic noise environments.
Smart Images

Figure CN120378594A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of communications, and in particular, to a data transmission method, apparatus, electronic device, and medium. Background Art
[0002] With the popularization of virtual reality, augmented reality, and digital twin technologies, the application demand for three-dimensional (3D) scene transmission has surged in fields such as smart cities and industrial simulation.
[0003] However, traditional 3D scene transmission technologies face two core challenges. One is high data redundancy and modeling complexity. Currently, it mostly relies on independent compression and transmission of raw data from multiple cameras (such as JPEG and H.264), ignoring the spatial correlation between viewpoints, resulting in a large amount of redundant data. Moreover, currently, manual intervention or complex processes are required during transmission, which is inefficient and costly. Summary of the Invention
[0004] In view of this, the purpose of the present disclosure is to propose a data transmission method, apparatus, electronic device, and medium to solve or partially solve the above problems.
[0005] Based on the above purpose, the first aspect of the present disclosure provides a data transmission method, which is applied to a sending end. The method includes:
[0006] Collect sensing data, construct a target scene according to the sensing data, and determine target 3D features corresponding to the target scene, where the sensing data is data from different viewpoints;
[0007] Input the target 3D features into a transform coding model, and after being processed by the transform coding model, output compact features corresponding to the target 3D features;
[0008] Determine the bandwidth resource allocation amount corresponding to the compact features, input the bandwidth resource allocation amount and the compact features into a channel encoder, and after being processed by the channel encoder, output a modulated symbol sequence;
[0009] Send the modulated symbol sequence to a physical channel and send it to a receiving end via the physical channel.
[0010] Based on the above purpose, the second aspect of the present disclosure provides a data transmission method, which is applied to a receiving end. The method includes:
[0011] Receive a target angle and the modulated symbol sequence sent by the sending end, perform decoding processing on the modulated symbol sequence to obtain reconstructed 3D features;
[0012] According to the target angle, perform rendering processing based on the reconstructed 3D features to obtain a target image corresponding to the target scene at the target angle.
[0013] Based on the same inventive concept, a third aspect of the present disclosure provides a data transmission device disposed at a sending end. The data transmission device includes a three-dimensional feature extraction module, a transform coding module, and a channel coding module. Among them,
[0014] The three-dimensional feature extraction module is configured to collect sensing data, construct a target scene according to the sensing data, and determine target three-dimensional features corresponding to the target scene, where the sensing data is data from different perspectives;
[0015] The transform coding module is configured to input the target three-dimensional features into a transform coding model, and output compact features corresponding to the target three-dimensional features after being processed by the transform coding model;
[0016] The channel coding module is configured to determine a bandwidth resource allocation amount corresponding to the compact features, input the bandwidth resource allocation amount and the compact features into a channel encoder, output a modulation symbol sequence after being processed by the channel encoder, send the modulation symbol sequence to a physical channel, and send it to a receiving end via the physical channel
[0017] Based on the same inventive concept, a fourth aspect of the present disclosure provides a data transmission device disposed at a receiving end. The data transmission device includes a data receiving module and a rendering module. Among them,
[0018] The data receiving module is configured to receive a target angle and a modulation symbol sequence sent by a sending end, and perform decoding processing on the modulation symbol sequence to obtain reconstructed three-dimensional features;
[0019] The rendering module is configured to perform rendering processing based on the reconstructed three-dimensional features according to the target angle to obtain a target image corresponding to the target scene at the target angle.
[0020] Based on the same inventive concept, a fifth aspect of the present disclosure provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable by the processor. When the processor executes the computer program, the data transmission method described above is implemented.
[0021] Based on the same inventive concept, a sixth aspect of the present disclosure provides a non-transitory computer-readable storage medium. The non-transitory computer-readable storage medium stores computer instructions for causing a computer to execute the data transmission method described above.
[0022] As can be seen from the above, the present disclosure proposes a data transmission method, apparatus, electronic device and medium, which collect sensing data, construct a target scene according to the sensing data, and determine target three-dimensional features corresponding to the target scene, wherein the sensing data are data from different perspectives; input the target three-dimensional features into a transform coding model, process them through the transform coding model, and output compact features corresponding to the target three-dimensional features. By adopting the transform coding model, redundancy is reduced, and many limitations of traditional linear transforms in processing high-dimensional data are overcome, significantly improving the coding efficiency. Determine the bandwidth resource allocation amount corresponding to the compact features, input the bandwidth resource allocation amount and the compact features into a channel encoder, process them through the channel encoder, and output a modulation symbol sequence. Through the channel resource allocation coefficient, channel resources are dynamically allocated during coding, long codewords are allocated to important features to ensure transmission reliability, and short codewords are used for secondary features to optimize bandwidth occupancy, effectively improving the robustness of the final scene rendering result. Send the modulation symbol sequence to a physical channel, and send it to a receiving end through the physical channel for the receiving end to receive and reconstruct three-dimensional features according to the modulation symbol sequence, and then generate image data corresponding to the target scene, realizing the efficient transmission and accurate reconstruction of three-dimensional scenes, and having broad application prospects and excellent system adaptability. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] In order to more clearly illustrate the technical solutions in the present disclosure or related technologies, the following will briefly introduce the drawings required for use in the embodiments or related technology descriptions. Obviously, the drawings in the following description are only embodiments of the present disclosure. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0024] Figure 1 It is a flowchart of the data transmission method according to an embodiment of the present disclosure;
[0025] Figure 2 It is a schematic structural diagram of the transform coding model according to an embodiment of the present disclosure;
[0026] Figure 3 It is a flowchart of the data transmission method according to another embodiment of the present disclosure;
[0027] Figure 4 It is a schematic structural diagram of the transform decoding model according to an embodiment of the present disclosure;
[0028] Figure 5 It is a schematic diagram of the experimental simulation results according to an embodiment of the present disclosure;
[0029] Figure 6 It is a structural block diagram of the data transmission apparatus according to an embodiment of the present disclosure;
[0030] Figure 7Block diagram of a data transmission device according to another embodiment of the present disclosure;
[0031] Figure 8 Schematic diagram of a data transmission system provided by another embodiment of the present disclosure;
[0032] Figure 9 Schematic structural diagram of an electronic device according to an embodiment of the present disclosure. Detailed implementation manners
[0033] To make the objectives, technical solutions, and advantages of the present disclosure more clear and understandable, the present disclosure will be further described in detail below with reference to specific embodiments and the accompanying drawings.
[0034] It should be noted that, unless otherwise defined, the technical terms or scientific terms used in the embodiments of the present disclosure should have the ordinary meanings understood by those of ordinary skill in the art to which the present disclosure belongs. The terms "first", "second", and similar terms used in the embodiments of the present disclosure do not indicate any order, quantity, or importance, but are only used to distinguish different components. The terms "including" or "comprising" and the like mean that the elements or objects appearing before this term cover the elements or objects listed after this term and their equivalents, without excluding other elements or objects. The terms "connected" or "coupled" and the like are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. The terms "upper", "lower", "left", "right", etc. are only used to represent relative positional relationships, and when the absolute position of the object being described changes, the relative positional relationship may also change accordingly.
[0035] The following are the explanations of the terms related to the present disclosure:
[0036] VR: Virtual Reality (VR) technology combines virtual and reality. It uses specific technologies to generate a virtual scenario, but gives people a real feeling. Theoretically speaking, virtual reality technology is a computer simulation system that can create and experience virtual worlds. It uses a computer to generate a simulated environment and enables users to immerse themselves in this environment.
[0037] AR: Augmented Reality (AR) technology integrates the virtual environment generated by a computer with the real environment around the user by means of optoelectronic display technology, interaction technology, various sensor technologies, and computer graphics and multimedia technologies, so that users are convinced from the sensory effects that the virtual environment is a part of the real environment around them.
[0038] JPEG: JPEG (Joint Photographic Experts Group) is an image compression technology standard developed by the expert group of the same name. This standard is formulated by the International Organization for Standardization (ISO) and is a compression standard for continuous-tone still images.
[0039] NeRF: Neural Radiance Fields (NeRF) is a computer vision technology used to generate high-quality 3D reconstruction models. It uses deep learning technology to extract the geometric shape and texture information of an object from images taken from multiple viewpoints, and then uses this information to generate a continuous 3D radiation field, enabling a highly realistic 3D model to be presented at any angle and distance.
[0040] Rate-Distortion: The rate distortion theory, or information rate-distortion theory, is a major branch of information theory. Its basic problem can be summarized as follows: For a given source (input signal) distribution and distortion measure, what is the minimum expected distortion that can be achieved at a specific bit rate; or what is the maximum allowable bit rate to meet a certain distortion limit.
[0041] With the popularization of virtual reality (VR), augmented reality (AR), and digital twin technologies, the application demand for 3D scene transmission in fields such as smart cities and industrial simulation has increased sharply.
[0042] However, traditional 3D scene transmission technologies face two core challenges: high data redundancy and high modeling complexity. Existing solutions mostly rely on the independent compression and transmission of raw data from multiple cameras (such as JPEG, H.264), ignoring the spatial correlation between viewpoints, resulting in a large amount of redundant data (usually reaching hundreds of megabytes to gigabytes). In addition, traditional 3D modeling requires manual intervention or complex processes, with low efficiency and high costs. The transmission robustness is insufficient. Existing methods adopt a separate source-channel coding design, such as feature compression schemes based on vector quantization or linear transformation. Although they can effectively compress data, they do not consider the impact of channel noise and rely on channel coding with a fixed bit rate (such as LDPC). Such schemes are prone to the "cliff effect" caused by fluctuations in channel conditions, resulting in a sharp decline in reconstruction quality.
[0043] In recent years, Neural Radiance Field (NeRF) technology has achieved high-quality 3D reconstruction through implicit neural representations, significantly reducing the complexity of data modeling. However, the implicit features of NeRF have high dimensions (usually reaching dozens of megabytes) and strong correlations, making it difficult for traditional linear compression methods to encode efficiently. In addition, existing NeRF transmission schemes mostly focus on source compression and are not jointly optimized with channel coding, making it difficult to adapt to the dynamic noise environment of wireless channels. Therefore, there is an urgent need for an innovative solution that integrates efficient compression of 3D features, semantic-driven resource allocation, and end-to-end noise-resistant transmission to break through the bottlenecks of existing technologies.
[0044] Based on the above description, to address the problems of high view redundancy, difficult-to-balance transmission efficiency and reconstruction quality in existing 3D scene transmission schemes, this embodiment proposes a data transmission method, as Figure 1 shown, which is applied to the sender side, and the method includes:
[0045] Step 101, collect perception data, construct a target scene according to the perception data, and determine the target 3D features corresponding to the target scene, where the perception data is data from different viewpoints;
[0046] Step 102, input the target 3D features into a transform coding model, and after being processed by the transform coding model, output the compact features corresponding to the target 3D features;
[0047] Step 103, determine the bandwidth resource allocation amount corresponding to the compact features, input the bandwidth resource allocation amount and the compact features into a channel encoder, and after being processed by the channel encoder, output a modulated symbol sequence;
[0048] Step 104, send the modulated symbol sequence to a physical channel and send it to the receiver side via the physical channel.
[0049] Specifically, the sender side collects perception data and constructs a target scene according to the perception data, where the perception data is data from different viewpoints of the same scene, that is, images obtained from different viewpoints, and then constructs a target scene according to the data from different viewpoints. At the same time, the perception data contains the spatial geometric structure and texture attribute information of the target scene, the perception data is multi-modal data, and the specific forms of the perception data include at least one of the following: images, videos, depth information, or point cloud data, etc.
[0050] The sender side determines the target 3D features corresponding to the target scene, and the 3D features are the global 3D structure of the target scene.
[0051] Since there is redundancy in the spatial structure of the target three-dimensional feature, a transformation coding model is obtained. The target three-dimensional feature is input into the transformation coding model and processed by the transformation coding model. The transformation coding model is used to compress the three-dimensional feature, reduce the dimension of the three-dimensional feature, and output a compact feature corresponding to the target three-dimensional feature, thereby reducing the spatial redundancy.
[0052] In this embodiment, the transformation coding model includes a linear transformation coding model or a non-linear transformation coding model. Preferably, the transformation coding model in this embodiment is a non-linear transformation coding model.
[0053] The contributions of features in different regions to the final reconstruction quality are different. Determine the bandwidth resource allocation amount corresponding to the compact feature. The bandwidth resource allocation amount indirectly represents the importance of features in different regions of the compact feature to the quality of the reconstructed three-dimensional scene.
[0054] Obtain a channel encoder. Input the bandwidth resource allocation amount and the compact feature into the channel encoder. After encoding processing by the channel encoder, output a modulation symbol sequence corresponding to the target scene. When the channel encoder performs encoding, dynamically allocate channel bandwidth resources according to the bandwidth resource allocation amount, allocate long codewords for important features to ensure transmission reliability, and use short codewords for secondary features to optimize bandwidth occupancy, effectively improving the robustness of the final scene rendering result.
[0055] Send the modulation symbol sequence to the physical channel and transmit it to the receiving end via the physical channel, so that the receiving end receives the modulation symbol sequence, performs decoding processing on the modulation symbol sequence, obtains the reconstructed three-dimensional feature corresponding to the modulation symbol sequence, and then performs three-dimensional scene reconstruction of the target scene based on the reconstructed three-dimensional feature.
[0056] In this embodiment, after the modulation symbol sequence is transmitted to the physical channel, noise can be added to the modulation symbol sequence in the physical channel to ensure the security of data transmission and avoid the problem of data leakage.
[0057] Through the above solution, perceptual data is collected, a target scene is constructed based on the perceptual data, and target three-dimensional features corresponding to the target scene are determined, where the perceptual data is data from different perspectives; the target three-dimensional features are input into a transform coding model, processed by the transform coding model, and compact features corresponding to the target three-dimensional features are output. By adopting the transform coding model, redundancy is reduced, many limitations of traditional linear transforms in processing high-dimensional data are broken through, and the coding efficiency is significantly improved. The bandwidth resource allocation amount corresponding to the compact features is determined, and the bandwidth resource allocation amount and the compact features are input into a channel encoder, and after being processed by the channel encoder, a modulated symbol sequence is output. Through the channel resource allocation coefficient, channel resources are dynamically allocated during coding, long codewords are allocated for important features to ensure transmission reliability, and short codewords are used for secondary features to optimize bandwidth occupancy, effectively improving the robustness of the final scene rendering result. The modulated symbol sequence is sent to a physical channel and then sent to a receiving end through the physical channel, so that the receiving end can receive and reconstruct three-dimensional features according to the modulated symbol sequence, and then generate image data corresponding to the target scene, realizing the efficient transmission and accurate reconstruction of three-dimensional scenes, and having broad application prospects and excellent system adaptability.
[0058] In some embodiments, in step 101, constructing a target scene according to the perceptual data and determining target three-dimensional features corresponding to the target scene specifically includes:
[0059] Step 1011, input the perceptual data into a neural radiance field model, and use the neural radiance field model to extract features to obtain target three-dimensional features corresponding to the perceptual data.
[0060] Specifically in implementation, a pre-trained neural radiance field model is obtained, the perceptual data is input into the neural radiance field model, and the neural radiance field model is used to extract features to obtain target three-dimensional features corresponding to the multiple images.
[0061] In some embodiments, taking the perceptual data as image data as an example, the training process of the neural radiance field model is specifically as follows:
[0062] Training image data and an initial neural radiance field model are obtained, where the training image data includes at least one training scene and sparse perspective images corresponding to the training scene, and the sparse perspective images include images obtained by observing the training scene from different perspectives.
[0063] The sparse perspective images are input into the initial neural radiance field model, and the initial neural radiance field model is trained using the sparse perspective images to obtain initial three-dimensional features. A training perspective is obtained, where the training perspective is one of the perspectives corresponding to the sparse perspective images.
[0064] Input the initial three-dimensional features and the training viewpoints into the initial neural radiance field model, and use the initial neural radiance field model for rendering to obtain training viewpoint images. Use the sparse viewpoint image corresponding to the training viewpoint as the actual viewpoint image, and determine the target loss function according to the training viewpoint image and the actual viewpoint image. The target loss function includes an image reconstruction loss and a voxel rendering loss. The image reconstruction loss measures the difference between the rendered color value and the true color value, and the voxel rendering loss measures the difference between the rendered depth value and the true depth value.
[0065] Train the initial neural radiance field model based on the target loss function until the preset training end condition is met, determine that the training of the initial neural radiance field model is completed, and obtain the neural radiance field model. Among them, the preset training end condition includes at least one of the following: the target loss function converges, the number of training times reaches the preset number threshold, or all the training image data is input into the initial neural radiance field model for training.
[0066] Exemplarily, the sparse viewpoint image is where each image u i represents a three-dimensional scene image observed from different viewpoints, I is the total number of sparse viewpoint images used for training, and the extracted initial three-dimensional features are represented as where D, H, W, and C correspond to the depth, height, width, and number of channels of the three-dimensional features respectively.
[0067] In this embodiment, the three-dimensional structure of the target scene is modeled by the neural radiance field model. The network samples each ray (the light ray starting from the viewpoint) and predicts the color and density information of the intersection point of the light ray and the scene, so as to generate the target viewpoint image for the subsequent receiving end.
[0068] In this embodiment, when modeling the three-dimensional scene features, methods such as tensor decomposition or voxel grid can be used. Usually, such explicit modeling methods have more explicit geometric meanings compared to the methods that only rely on neural networks, and can provide a more interpretable feature basis for the subsequent steps.
[0069] Through the above solution, by training the initial neural radiance field model, the trained neural radiance field model can learn the global three-dimensional structure of the target scene from limited viewpoint images, so as to generate high-quality target three-dimensional features.
[0070] In some embodiments, in step 102, input the target three-dimensional features into the transform coding model, and after being processed by the transform coding model, output the compact features corresponding to the target three-dimensional features, specifically including:
[0071] Step 1021: Obtain the target dimension, determine the current dimension corresponding to the target three-dimensional feature, and determine the target quantity according to the target dimension and the current dimension;
[0072] Step 1022: Input the target three-dimensional feature into the transform coding model, and use the transform coding model to perform dimensionality reduction and compression processing on the target three-dimensional feature to obtain the compact feature corresponding to the target three-dimensional feature, where the compact feature includes a plurality of data units, the number of data units is the same as the target quantity, and the dimension of the compact feature is the target dimension.
[0073] In specific implementation, obtain the target dimension, and the target dimension is the dimension of the set compact feature. Determine the current dimension corresponding to the target three-dimensional feature, and determine the target quantity according to the target dimension and the current dimension, where the target quantity is the number of data units included in the compact feature.
[0074] Specifically, the ratio of the current dimension to the target dimension can be processed, and the obtained ratio is the target quantity.
[0075] Exemplarily, the current dimension is W, the current dimension corresponding to the target three-dimensional feature is D, and the ratio of the current dimension to the target dimension is processed to obtain the target quantity H = D / W.
[0076] Taking a specific example to describe, the target dimension is 100*100*100, the current dimension corresponding to the target three-dimensional feature is 1000*1000*1000, then the ratio of the current dimension to the target dimension is processed, and the obtained ratio is 10*10*10, so the target quantity is 10*10*10.
[0077] Obtain the transform coding model, input the target three-dimensional feature into the transform coding model, and use the transform coding model to perform dimensionality reduction and compression processing on the target three-dimensional feature to reduce the dimension of the target three-dimensional feature and obtain the compact feature corresponding to the target three-dimensional feature.
[0078] Among them, the compact feature includes a plurality of data units, the number of data units is the same as the target quantity, and the dimension of the compact feature is the target dimension.
[0079] In this embodiment, the structural schematic diagram of the transform coding model is as Figure 2 shown. The transform coding model includes a three-dimensional convolutional layer, a normalization layer, an activation function, and a transformer encoder. The target three-dimensional feature is downsampled and the channel dimension is reduced through the transform coding model. The target three-dimensional feature is mapped to a compact feature, and the compact feature is expressed as v = {v1, v2, …, v i , …, vP}, where v i is the data unit to be transmitted. As the basic data unit for transmission processing, each data unit corresponds to a local area in the target three-dimensional feature. P is the total number of data units, and this total number is the target number.
[0080] Through the above solution, by using the cross-dimensional transform coding technology and the transform coding model with a multi-layer three-dimensional convolutional network, downsampling of the target three-dimensional feature and channel dimension compression are achieved, and the explicit target three-dimensional feature is converted into a more compact implicit compact feature. Compared with the traditional linear transform, this transform mapping can break through many limitations of the traditional linear transform in processing high-dimensional data and significantly improve the coding efficiency.
[0081] In some embodiments, determining the bandwidth resource allocation amount corresponding to the compact feature in step 103 specifically includes:
[0082] Step 1031, taking each data unit among the multiple data units included in the compact feature as the target data unit. For each target data unit:
[0083] Obtain the preset prior information and the preset conditional probability model, and determine the probability value corresponding to the target data unit according to the prior information and the conditional probability model;
[0084] Perform logarithmic processing on the probability value to obtain the information entropy value corresponding to the target data unit;
[0085] Obtain the preset quantization function, and map the information entropy value according to the quantization function to obtain the channel resource allocation coefficient corresponding to the target data unit;
[0086] Step 1032, statistically calculate all the channel resource allocation coefficients to obtain the channel resource allocation coefficient corresponding to the compact feature;
[0087] Step 1033, obtain the channel state information, determine the available bandwidth resource quantity according to the channel state information, and determine the bandwidth resource allocation amount corresponding to the compact feature according to the bandwidth resource quantity and the channel resource allocation coefficient.
[0088] During specific implementation, each data unit among the multiple data units included in the compact feature is respectively taken as the target data unit. For each target data unit:
[0089] Obtain the preset prior information and the preset conditional probability model. The conditional probability model is used to quantify the information entropy value of the target data unit. Determine the probability value corresponding to the target data unit according to the prior information and the conditional probability model.
[0090] In this embodiment, the preset conditional probability model is represented by the formula:
[0091]
[0092] where p v|z (v|z) is the conditional probability model, v i is the i-th target data unit, N is the Gaussian distribution, * is the convolution operation, and z is the preset prior information.
[0093] Taking the logarithm of the probability value to obtain the information entropy value corresponding to the target data unit. Based on the above example, the information entropy value is represented by the formula In this embodiment, the higher the information entropy value, the greater the contribution of this part of the data to the quality of the reconstructed three-dimensional scene.
[0094] Obtain a preset quantization function, and perform a mapping process on the information entropy value according to the quantization function to map it to the channel resource allocation coefficient corresponding to the target data unit, where the channel resource allocation coefficient is represented by the formula:
[0095]
[0096] where is the channel resource allocation coefficient, Q(·) is the quantization function, and η is the scaling factor used to adjust the loan allocation ratio. More channel symbols (long codewords) are allocated to high-entropy regions (such as regions with objects, texture details), and fewer symbols (short codewords) are allocated to low-entropy regions (such as background or blank regions).
[0097] Statistical all channel resource allocation coefficients to obtain the channel resource allocation coefficient corresponding to the compact feature.
[0098] Obtain the channel state information, where the channel state information is the current state of the physical channel sent by the physical channel, and the channel state information includes the total number of currently available bandwidth resources of the physical channel.
[0099] Determine the available bandwidth resource quantity according to the channel state information, and perform a multiplication process on the bandwidth resource quantity and the channel resource allocation coefficient to obtain the bandwidth resource allocation quantity corresponding to the compact feature, where the bandwidth resource allocation quantity includes the bandwidth resource quantity that each target data unit can be allocated.
[0100] With the above solution, since the contributions of the features in different regions to the final reconstruction quality vary, the channel resource allocation coefficient corresponding to each data unit is determined through a conditional probability model. Then, based on the channel state information and the channel resource allocation coefficient, considering the total available bandwidth resource quantity of the physical channel and the channel resource allocation ratio required for each data unit determined by the transmitting end, the bandwidth resource allocation amount corresponding to the compact feature is determined more accurately. Subsequently, during encoding, channel resources are allocated to each data unit in the compact feature according to the bandwidth resource allocation amount, improving the robustness of the final scene rendering.
[0101] In some embodiments, in step 103, the channel resource allocation coefficient and the compact feature are input into a channel encoder, and after being processed by the channel encoder, a modulated symbol sequence is output, including:
[0102] Step 103A, input the bandwidth resource allocation amount and the compact feature into the channel encoder, take each data unit in the compact feature as a data unit to be encoded, and for each data unit to be encoded:
[0103] Determine the target bandwidth resource allocation amount corresponding to the data unit to be encoded from the bandwidth resource allocation amount, and perform encoding processing on the data unit to be encoded according to the target bandwidth resource allocation amount to obtain the initial symbol data corresponding to the data unit to be encoded;
[0104] Step 103B, obtain the arrangement order of all data units to be encoded, count all the initial symbol data according to the arrangement order to obtain a modulated symbol sequence, and output the modulated symbol sequence.
[0105] During specific implementation, a channel encoder is obtained, and the input and output dimensions of the channel encoder are dynamically adjusted by the bandwidth resource allocation amount.
[0106] Input the bandwidth resource allocation amount and the compact feature into the channel encoder, take each data unit in the compact feature as a data unit to be encoded, and for each data unit to be encoded:
[0107] Determine the target bandwidth resource allocation amount corresponding to the data unit to be encoded from the bandwidth resource allocation amount, and perform encoding processing on the data unit to be encoded according to the target bandwidth resource allocation amount to obtain the initial symbol data corresponding to the data unit to be encoded, and the length of the initial symbol data is proportional to the target bandwidth resource allocation amount.
[0108] Obtain the arrangement order of all data units to be encoded, count all the initial symbol data according to the arrangement order to obtain a modulated symbol sequence, and output the modulated symbol sequence.
[0109] Exemplarily, the compact feature is v = {v1, v2, …, v i , …, v P}, and the channel encoder encodes the compact feature according to the bandwidth resource allocation amount to obtain a modulation symbol sequence, which is expressed by the formula as s = {s1, s2, …, s i , …, s P}, where the length of each initial symbol data s i is proportional to .
[0110] In this embodiment, the channel encoder at the transmitting end and the channel decoder at the receiving end can be jointly trained by optimizing the objective function to ensure that the generation of the modulation symbol sequence matches the robustness in the noise environment. Among them, the optimization objective function is the rate distortion criterion.
[0111] Another embodiment of the present disclosure proposes a data transmission method, as Figure 3 shown, which is applied to the receiving end, and the method includes:
[0112] Step 201, receiving the target angle and the modulation symbol sequence sent by the transmitting end, and decoding the modulation symbol sequence to obtain a reconstructed three-dimensional feature;
[0113] Step 202, based on the target angle, performing rendering processing on the reconstructed three-dimensional feature to obtain a target image corresponding to the target scene at the target angle.
[0114] Specifically, when implemented, the receiving end receives the target angle and the modulation symbol sequence sent by the transmitting end, and decodes the modulation symbol sequence to obtain a reconstructed three-dimensional feature.
[0115] It can be understood that if the received modulation symbol sequence is a noisy modulation symbol sequence, the noisy modulation symbol sequence can be denoised first to obtain a modulation symbol sequence.
[0116] Obtain a neural radiance field model, which is a model corresponding to the neural radiance field model used by the transmitting end to extract the target three-dimensional feature. Input the reconstructed three-dimensional feature and the target angle into the neural radiance field model, and after rendering processing by the neural radiance field model, output a target image corresponding to the target scene at the target angle.
[0117] Through the above solution, by using the neural radiance field model to sample the rays at any target viewing angle, predicting the color and density of the ray intersection points, and generating the target image, the obtained target image is more accurate. By jointly training the feature extraction, encoding, decoding, and rendering modules through the end-to-end optimization loss function, the reconstruction quality of the overall system is ensured.
[0118] In some embodiments, in step 202, receiving the target angle and the modulated symbol sequence sent by the sending end, and performing decoding processing on the modulated symbol sequence to obtain a reconstructed three-dimensional feature, including:
[0119] Step 2021, receiving the modulated symbol sequence and the bandwidth resource allocation amount sent by the sending end, inputting the bandwidth resource allocation amount and the modulated symbol sequence into a channel decoder, and after being processed by the channel decoder, outputting a reconstructed compact feature;
[0120] Step 2022, inputting the reconstructed compact feature into a transformation decoding model, and after being processed by the transformation decoding model, outputting the reconstructed three-dimensional feature corresponding to the reconstructed compact feature.
[0121] Specifically, when the sending end sends the modulated symbol sequence, it can also send the bandwidth resource allocation amount. After the receiving end receives the modulated symbol sequence and the bandwidth resource allocation amount sent by the sending end, it inputs the bandwidth resource allocation amount and the modulated symbol sequence into a channel decoder, and after being processed by the channel decoder, outputs a reconstructed compact feature.
[0122] In this embodiment, the channel decoder is the decoder obtained by being trained simultaneously with the channel encoder of the sending end in the above embodiment.
[0123] Inputting the reconstructed compact feature into a transformation decoding model, and after being processed by the transformation decoding model, outputting the reconstructed three-dimensional feature corresponding to the reconstructed compact feature.
[0124] In this embodiment, the transformation decoding model is the model obtained by being trained simultaneously with the transformation encoding model of the sending end in the above embodiment. The structural schematic diagram of the transformation decoding model is as Figure 4 shown. The transformation decoder includes a transformer decoder, a transposed three-dimensional convolutional layer, a normalization layer, and an activation function. At the same time, if the transformation encoding model of the sending end is a non-linear transformation encoding model, then the transformation decoding model of the receiving end is a non-linear transformation decoding model.
[0125] Based on the same inventive concept, another embodiment of the present disclosure provides an interaction process of a data transmission method. The interaction process involves a sending end and a receiving end, and the interaction process specifically includes:
[0126] Step 301, determining the target three-dimensional feature corresponding to the target scenario.
[0127] Specifically, training a neural radiance field model in advance, specifically including:
[0128] Extracting three-dimensional features from a set of sparse perspective images and where each image \(u\) i represents a 3D scene image observed from different perspectives, \(I\) is the total number of sparse perspective images for training, and \(D\), \(H\), \(W\), and \(C\) correspond to the depth, height, width, and number of channels of the 3D features respectively.
[0129] When it comes to complex scene modeling, to reduce the workload of manual modeling and improve the efficiency and accuracy of feature extraction, by combining NeRF technology, a deep neural network is used to model the 3D structure of the scene. The network samples each ray (the light ray starting from the perspective) and predicts the color and density information of the intersection of the ray and the scene. By training the neural network, NeRF can learn the global 3D structure of the scene from limited perspective images, thereby generating high-quality 3D feature representations Optionally, for 3D scene feature modeling, methods such as tensor decomposition or voxel grids can be used. Generally, such explicit modeling methods have more explicit geometric meanings compared to methods relying solely on neural networks and can provide a more interpretable feature basis for subsequent steps.
[0130] Obtain multiple images of the target scene, input the multiple images into the neural radiance field model, and use the neural radiance field model to extract features to obtain the target 3D features corresponding to the target scene.
[0131] Step 302, adopt a transform coding model to compress the target 3D features into compact features.
[0132] Specifically, obtain the transform coding model, input the target 3D features into the transform coding model, and use the transform coding model to perform downsampling on the target 3D features to reduce the dimension of the target 3D features and obtain the compact features corresponding to the target 3D features.
[0133] Among them, the compact features include multiple data units, the number of data units is the same as the target number, and the dimension of the compact features is the target dimension.
[0134] In this embodiment, the transform coding model includes a 3D convolutional layer, a normalization layer, an activation function, and a transformer encoder. The target 3D features are downsampled and the channel dimension is reduced through the transform coding model. The target 3D features are mapped to compact features, and the compact features are represented as \(v=\{v_1,v_2,\ldots,v i ,\ldots,v P \}\), where \(v i is the data unit for transmission and serves as the basic data unit for transmission processing. Each data unit corresponds to a local area in the target 3D features, and \(P\) is the total number of data units, and the total number is the target number.
[0135] Step 303, perform joint source-channel coding based on the importance difference.
[0136] Specifically, obtain the preset prior information and the preset conditional probability model, where the conditional probability model is used to quantify the information entropy value of the target data unit. Determine the probability value corresponding to the target data unit according to the prior information and the conditional probability model.
[0137] In this embodiment, the preset conditional probability model is expressed by the formula:
[0138]
[0139] where p v|z (v|z) is the conditional probability model, v i is the i-th target data unit, N is the Gaussian distribution, * is the convolution operation, and z is the preset prior information.
[0140] Perform logarithmic processing on the probability value to obtain the information entropy value corresponding to the target data unit. Based on the above example, the information entropy value is expressed by the formula In this embodiment, the higher the information entropy value, the greater the contribution of this part of the data to the quality of the reconstructed three-dimensional scene.
[0141] Obtain the preset quantization function, and perform mapping processing on the information entropy value according to the quantization function to map it to the channel resource allocation coefficient corresponding to the target data unit, where the channel resource allocation coefficient is expressed by the formula:
[0142]
[0143] where is the channel resource allocation coefficient, Q(·) is the quantization function, and η is the scaling factor used to adjust the loan allocation ratio. More channel symbols (long codewords) are allocated to high-entropy regions (such as regions with objects, texture details), and fewer symbols (short codewords) are allocated to low-entropy regions (such as background or blank regions).
[0144] Statistically calculate all the channel resource allocation coefficients to obtain the channel resource allocation coefficient corresponding to the compact feature.
[0145] Obtain the channel state information, determine the available bandwidth resource quantity according to the channel state information, and determine the bandwidth resource allocation amount corresponding to the compact feature according to the bandwidth resource quantity and the channel resource allocation coefficient.
[0146] Obtain a channel encoder, and the input and output dimensions of the channel encoder are dynamically adjusted by the bandwidth resource allocation amount.
[0147] Input the bandwidth resource allocation amount and the compact feature into a channel encoder. Take each data unit in the compact feature as a data unit to be encoded. For each data unit to be encoded:
[0148] Determine the target bandwidth resource allocation amount corresponding to the data unit to be encoded from the bandwidth resource allocation amount, and perform encoding processing on the data unit to be encoded according to the target bandwidth resource allocation amount to obtain the initial symbol data corresponding to the data unit to be encoded. The length of the initial symbol data is proportional to the target bandwidth resource allocation amount.
[0149] Obtain the arrangement order of all data units to be encoded, count all the initial symbol data according to the arrangement order to obtain a modulation symbol sequence, and output the modulation symbol sequence.
[0150] Exemplarily, the compact feature is v = {v1, v2, …, v i , …, v P}. Use the channel encoder to perform encoding processing on the compact feature according to the bandwidth resource allocation amount to obtain a modulation symbol sequence. The modulation symbol sequence is expressed by the formula s = {s1, s2, …, s i , …, s P}, where the length of each initial symbol data s i is proportional to .
[0151] Step 304, decoding and rendering at the receiving end.
[0152] Specifically, first perform channel symbol decoding. The receiving end obtains the noisy symbols, and sequentially uses the source-channel decoder and the transform decoding module to reconstruct the three-dimensional feature Through the optimization of the entire system, ensure that the restored three-dimensional feature is close to the original feature f.
[0153] Perform neural rendering and view synthesis. Based on the reconstructed three-dimensional feature Use the NeRF rendering network to sample the rays of any target view, predict the color and density of the ray intersection points, and generate the target view image Jointly train the feature extraction, encoding, decoding, and rendering modules through an end-to-end optimization loss function to ensure the reconstruction quality of the overall system.
[0154] In this embodiment, multiple groups of simulation experiments are carried out based on the Synthetic-NeRF dataset, and then compare the performance of the data transmission method, linear coding scheme, and vector quantization scheme described in this embodiment under the same transmission bandwidth resources. The experimental results are as Figure 5 shown Figure 5The horizontal axis is the channel bandwidth occupancy rate, and the vertical axis is the image quality evaluation index (peak signal-to-noise ratio, PSNR). As can be seen from Figure 5 it can be seen that for the method provided in this embodiment, on the basis of the same channel bandwidth occupancy rate, the image quality evaluation index is higher, that is, the image quality is better, and better performance gains are achieved under different bandwidth conditions.
[0155] It should be noted that the method of the embodiments of the present disclosure can be executed by a single device, such as a computer or a server. The method of this embodiment can also be applied to a distributed scenario and completed by multiple devices cooperating with each other. In such a distributed scenario, one of the multiple devices can only execute one or more steps of the method of the embodiments of the present disclosure, and these multiple devices will interact with each other to complete the described method.
[0156] It should be noted that some embodiments of the present disclosure have been described above. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in a different order than in the above embodiments and still achieve the desired results. Additionally, the processes depicted in the figures do not necessarily require the particular order or sequential order shown to achieve the desired results. In certain implementations, multitasking and parallel processing are also possible or may be advantageous.
[0157] Based on the same inventive concept, corresponding to the method of any of the above embodiments, the present disclosure further provides a data transmission device disposed at the sending end.
[0158] Referring to Figure 6 Figure 6 For the data transmission device of the embodiment, disposed at the sending end, the data transmission device includes a three-dimensional feature extraction module 601, a transform coding module 602, and a channel coding module 603, wherein,
[0159] The three-dimensional feature extraction module 601 is configured to collect sensing data, construct a target scene according to the sensing data, and determine a target three-dimensional feature corresponding to the target scene, wherein the sensing data is data from different perspectives;
[0160] The transform coding module 602 is configured to input the target three-dimensional feature into a transform coding model, and output a compact feature corresponding to the target three-dimensional feature after being processed by the transform coding model;
[0161] The channel coding module 603 is configured to determine the bandwidth resource allocation amount corresponding to the compact feature, input the bandwidth resource allocation amount and the compact feature into a channel encoder, process them via the channel encoder to output a modulated symbol sequence, and send the modulated symbol sequence to a physical channel and then to a receiving end via the physical channel.
[0162] In some embodiments, the transform coding module 602 is specifically configured to:
[0163] Obtain a target dimension, determine the current dimension corresponding to the target three-dimensional feature, and determine a target quantity according to the target dimension and the current dimension;
[0164] Input the target three-dimensional feature into a transform coding model, and use the transform coding model to perform dimensionality reduction and compression on the target three-dimensional feature to obtain a compact feature corresponding to the target three-dimensional feature, where the compact feature includes a plurality of data units, the number of data units is the same as the target quantity, and the dimension of the compact feature is the target dimension.
[0165] In some embodiments, the channel coding module 603 is specifically configured to:
[0166] Take each data unit in the plurality of data units included in the compact feature as a target data unit, and for each target data unit:
[0167] Obtain preset prior information and a preset conditional probability model, and determine a probability value corresponding to the target data unit according to the prior information and the conditional probability model;
[0168] Perform a logarithmic process on the probability value to obtain an information entropy value corresponding to the target data unit;
[0169] Obtain a preset quantization function, and map the information entropy value according to the quantization function to obtain a channel resource allocation coefficient corresponding to the target data unit;
[0170] Statistically calculate all the channel resource allocation coefficients to obtain a channel resource allocation coefficient corresponding to the compact feature;
[0171] Obtain channel state information, determine the available bandwidth resource quantity according to the channel state information, and determine the bandwidth resource allocation amount corresponding to the compact feature according to the bandwidth resource quantity and the channel resource allocation coefficient.
[0172] In some embodiments, the channel coding module 603 is specifically configured to:
[0173] Input the allocated bandwidth resource amount and the compact feature into a channel encoder. Treat each data unit in the compact feature as a data unit to be encoded. For each data unit to be encoded:
[0174] Determine the target bandwidth resource allocation amount corresponding to the data unit to be encoded from the allocated bandwidth resource amount, and perform encoding processing on the data unit to be encoded according to the target bandwidth resource allocation amount to obtain the initial symbol data corresponding to the data unit to be encoded;
[0175] Obtain the arrangement order of all data units to be encoded, count all the initial symbol data according to the arrangement order to obtain a modulation symbol sequence, and output the modulation symbol sequence.
[0176] Based on the same inventive concept, corresponding to the method in any of the above embodiments, the present disclosure further provides a data transmission device, which is arranged at the receiving end.
[0177] Refer to Figure 7 , Figure 7 For the data transmission device in the embodiment, which is arranged at the receiving end, the data transmission device includes a data receiving module 701 and a rendering module 702, wherein,
[0178] The data receiving module 701 is configured to receive the target angle and the modulation symbol sequence sent by the sending end, and perform decoding processing on the modulation symbol sequence to obtain a reconstructed three-dimensional feature;
[0179] The rendering module 702 is configured to perform rendering processing based on the reconstructed three-dimensional feature according to the target angle to obtain a target image corresponding to the target scene at the target angle.
[0180] In some embodiments, the data receiving module 701 is specifically configured to:
[0181] Receive the modulation symbol sequence and the allocated bandwidth resource amount sent by the sending end, input the allocated bandwidth resource amount and the modulation symbol sequence into a channel decoder, and output a reconstructed compact feature after being processed by the channel decoder;
[0182] Input the reconstructed compact feature into a transform decoding model, and output the reconstructed three-dimensional feature corresponding to the reconstructed compact feature after being processed by the transform decoding model.
[0183] Based on the same inventive concept, another embodiment of the present disclosure provides a data transmission system, as Figure 8 shown. The data transmission system specifically includes a three-dimensional feature extraction module, a transform encoding module, a channel encoding module, a data receiving module, a decoding module, and a rendering module, wherein,
[0184] The three-dimensional feature extraction module is configured to receive the target scene sent by the receiving end and determine the target three-dimensional feature corresponding to the target scene;
[0185] The transform coding module is configured to input the target three-dimensional feature into a transform coding model, and after being processed by the transform coding model, output a compact feature corresponding to the target three-dimensional feature;
[0186] The channel coding module is configured to determine a channel resource allocation coefficient corresponding to the compact feature, input the channel resource allocation coefficient and the compact feature into a channel encoder, and after being processed by the channel encoder, output a modulation symbol sequence, and send the modulation symbol sequence to a channel and send it to the receiving end via the channel;
[0187] The data receiving module is configured to receive the target scene and the target angle, and feedback the target scene to the sending end, so that after the sending end receives the target scene, it outputs a modulation symbol sequence corresponding to the target scene;
[0188] The decoding module is configured to receive the modulation symbol sequence sent by the sending end, perform decoding processing on the modulation symbol sequence, and obtain a reconstructed three-dimensional feature;
[0189] The rendering module is configured to input the reconstructed three-dimensional feature and the target angle into a neural radiance field model, and after being rendered by the neural radiance field model, output a target image corresponding to the target scene at the target angle.
[0190] For the convenience of description, when describing the above device, various modules are described separately according to their functions. Of course, when implementing the present disclosure, the functions of each module can be implemented in one or more software and / or hardware.
[0191] The device in the above embodiment is used to implement the corresponding data transmission method in any of the foregoing embodiments, and has the beneficial effects of the corresponding method embodiments, which will not be elaborated here.
[0192] Based on the same inventive concept, corresponding to the method of any of the above embodiments, the present disclosure also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the program, it implements the data transmission method described in any of the above embodiments.
[0193] Figure 9FIG. 0 shows a more specific schematic diagram of the hardware structure of the electronic device provided in this embodiment. The device may include: a processor 1010, a memory 1020, an input / output interface 1030, a communication interface 1040, and a bus 1050. Among them, the processor 1010, the memory 1020, the input / output interface 1030, and the communication interface 1040 are communicatively connected to each other inside the device through the bus 1050.
[0194] The processor 1010 may be implemented in the form of a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, etc., and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this specification.
[0195] The memory 1020 may be implemented in the form of a ROM (Read Only Memory), a RAM (Random Access Memory), a static storage device, a dynamic storage device, etc. The memory 1020 may store an operating system and other application programs. When implementing the technical solutions provided in the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 1020 and are called and executed by the processor 1010.
[0196] The input / output interface 1030 is used to connect to an input / output module to implement information input and output. The input / output module may be configured as a component in the device (not shown in the figure) or externally connected to the device to provide corresponding functions. Among them, the input device may include a keyboard, a mouse, a touch screen, a microphone, various sensors, etc., and the output device may include a display, a speaker, a vibrator, an indicator light, etc.
[0197] The communication interface 1040 is used to connect to a communication module (not shown in the figure) to implement communication and interaction between this device and other devices. Among them, the communication module may implement communication in a wired manner (such as USB, network cable, etc.) or in a wireless manner (such as mobile network, WIFI, Bluetooth, etc.).
[0198] The bus 1050 includes a path for transmitting information between various components of the device (such as the processor 1010, the memory 1020, the input / output interface 1030, and the communication interface 1040).
[0199] It should be noted that although the above device only shows the processor 1010, the memory 1020, the input / output interface 1030, the communication interface 1040, and the bus 1050, in the specific implementation process, the device may also include other components necessary for normal operation. In addition, those skilled in the art can understand that the above device may also only include the components necessary to implement the solution of the embodiments of this specification, and does not necessarily include all the components shown in the figure.
[0200] The electronic device in the above embodiment is used to implement the corresponding data transmission method in any of the foregoing embodiments, and has the beneficial effects of the corresponding method embodiments, which will not be elaborated here.
[0201] Based on the same inventive concept, corresponding to the method in any of the above embodiments, the present disclosure also provides a non-transitory computer-readable storage medium storing computer instructions for causing the computer to execute the data transmission method described in any of the foregoing embodiments.
[0202] The computer-readable medium of this embodiment includes permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, magnetic disk storage or other magnetic storage devices, or any other non-transmission medium that can be used to store information accessible by a computing device.
[0203] The computer instructions stored in the storage medium of the above embodiment are used to cause the computer to execute the data transmission method described in any of the foregoing embodiments, and have the beneficial effects of the corresponding method embodiments, which will not be elaborated here.
[0204] It can be understood that before using the technical solutions of the various embodiments of the present disclosure, the types, usage scopes, usage scenarios, etc. of the personal information involved will be informed to the user in an appropriate manner and the user's authorization will be obtained.
[0205] For example, when responding to receiving an active request from a user, a prompt message is sent to the user to clearly prompt the user that the operation requested by the user will require obtaining and using the user's personal information. Thus, the user can autonomously choose whether to provide personal information to software or hardware such as an electronic device, an application program, a server, or a storage medium that performs the operations of the present disclosure technical solution based on the prompt message.
[0206] As an optional but non-limiting implementation manner, when responding to receiving an active request from a user, the manner of sending a prompt message to the user may be, for example, in the form of a pop-up window, and the prompt message may be presented in text in the pop-up window. In addition, the pop-up window may also carry a selection control for the user to choose "agree" or "disagree" to provide personal information to the electronic device.
[0207] It can be understood that the above process of notifying and obtaining user authorization is only illustrative and does not limit the implementation manner of the present disclosure. Other manners that comply with relevant laws and regulations can also be applied to the implementation manner of the present disclosure.
[0208] Those of ordinary skill in the art should understand that: the discussion of any of the above embodiments is only exemplary and is not intended to imply that the scope of the present disclosure (including the claims) is limited to these examples; under the idea of the present disclosure, the technical features between the above embodiments or different embodiments can also be combined, the steps can be implemented in any order, and there are many other variations in different aspects of the embodiments of the present disclosure as described above, and they are not provided in detail for the sake of brevity.
[0209] In addition, for the sake of simplicity of description and discussion, and in order not to make the embodiments of the present disclosure difficult to understand, the well-known power / ground connections to integrated circuit (IC) chips and other components may or may not be shown in the provided drawings. In addition, the device may be shown in the form of a block diagram to avoid making the embodiments of the present disclosure difficult to understand, and this also takes into account the fact that the details of the implementation manner of these block diagram devices highly depend on the platform on which the embodiments of the present disclosure will be implemented (that is, these details should be fully within the understanding of those skilled in the art). In the case where specific details (such as circuits) are set forth to describe the exemplary embodiments of the present disclosure, it is obvious to those skilled in the art that the embodiments of the present disclosure can be implemented without these specific details or with variations of these specific details. Therefore, these descriptions should be considered illustrative rather than restrictive.
[0210] Although the present disclosure has been described in connection with specific embodiments thereof, many alternatives, modifications, and variations of these embodiments will be apparent to those of ordinary skill in the art in light of the foregoing description. For example, other memory architectures (e.g., dynamic RAM (DRAM)) may be used with the embodiments discussed.
[0211] Embodiments of the present disclosure are intended to cover all such alternatives, modifications, and variations that fall within the broad scope of the appended claims. Accordingly, any omissions, modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of the embodiments of the present disclosure shall be included within the protection scope of the present disclosure.
Claims
1. A data transmission method, characterized in that, Applied to the sending end, including: Collect perception data, construct a target scene according to the perception data, and determine the target three-dimensional features corresponding to the target scene, where the perception data is data from different perspectives; Input the target three-dimensional features into a transform coding model, process them through the transform coding model, and output the compact features corresponding to the target three-dimensional features; Determine the bandwidth resource allocation amount corresponding to the compact features, input the bandwidth resource allocation amount and the compact features into a channel encoder, process them through the channel encoder, and output a modulated symbol sequence; Send the modulated symbol sequence to a physical channel and send it to the receiving end through the physical channel.
2. The method according to claim 1, characterized in that The step of inputting the target three-dimensional features into a transform coding model, processing them through the transform coding model, and outputting the compact features corresponding to the target three-dimensional features includes: Obtain a target dimension, determine the current dimension corresponding to the target three-dimensional features, and determine a target quantity according to the target dimension and the current dimension; Input the target three-dimensional features into a transform coding model, use the transform coding model to perform dimensionality reduction and compression on the target three-dimensional features, and obtain the compact features corresponding to the target three-dimensional features, where the compact features include a plurality of data units, the number of data units is the same as the target quantity, and the dimension of the compact features is the target dimension.
3. The method according to claim 2, wherein The step of determining the bandwidth resource allocation amount corresponding to the compact features includes: Take each data unit in the plurality of data units included in the compact features as a target data unit. For each target data unit: Obtain preset prior information and a preset conditional probability model, and determine the probability value corresponding to the target data unit according to the prior information and the conditional probability model; Perform a logarithmic process on the probability value to obtain the information entropy value corresponding to the target data unit; Obtain a preset quantization function, map the information entropy value according to the quantization function, and obtain the channel resource allocation coefficient corresponding to the target data unit; Statistically calculate all the channel resource allocation coefficients to obtain the channel resource allocation coefficient corresponding to the compact features; Obtain channel state information, determine the available bandwidth resource quantity according to the channel state information, and determine the bandwidth resource allocation amount corresponding to the compact features according to the bandwidth resource quantity and the channel resource allocation coefficient.
4. The method according to claim 3, wherein The step of inputting the bandwidth resource allocation amount and the compact features into a channel encoder, processing them through the channel encoder, and outputting a modulated symbol sequence includes: Input the bandwidth resource allocation amount and the compact features into a channel encoder, take each data unit in the compact features as a data unit to be encoded. For each data unit to be encoded: Determine the target bandwidth resource allocation amount corresponding to the data unit to be encoded from the bandwidth resource allocation amount, and perform encoding processing on the data unit to be encoded according to the target bandwidth resource allocation amount to obtain the initial symbol data corresponding to the data unit to be encoded; Obtain the arrangement order of all data units to be encoded, count all the initial symbol data according to the arrangement order to obtain a modulation symbol sequence, and output the modulation symbol sequence.
5. A data transmission method, characterized in that, Applied to the receiving end, it includes: Receive the target angle and the modulation symbol sequence sent by the sending end, perform decoding processing on the modulation symbol sequence to obtain a reconstructed three-dimensional feature; Based on the target angle, perform rendering processing on the reconstructed three-dimensional feature to obtain the target image corresponding to the target scene at the target angle.
6. The method according to claim 5, characterized in that, The receiving the target angle and the modulation symbol sequence sent by the sending end, and performing decoding processing on the modulation symbol sequence to obtain a reconstructed three-dimensional feature includes: Receive the modulation symbol sequence sent by the sending end and the bandwidth resource allocation amount, input the bandwidth resource allocation amount and the modulation symbol sequence into a channel decoder, and after being processed by the channel decoder, output a reconstructed compact feature; Input the reconstructed compact feature into a transform decoding model, and after being processed by the transform decoding model, output the reconstructed three-dimensional feature corresponding to the reconstructed compact feature.
7. A data transmission device, characterized in that, Disposed at the sending end, the data transmission device includes a three-dimensional feature extraction module, a transform coding module, and a channel coding module, where The three-dimensional feature extraction module is configured to collect sensing data, construct a target scene according to the sensing data, and determine the target three-dimensional feature corresponding to the target scene, where the sensing data is data from different perspectives; The transform coding module is configured to input the target three-dimensional feature into a transform coding model, and after being processed by the transform coding model, output a compact feature corresponding to the target three-dimensional feature; The channel coding module is configured to determine the bandwidth resource allocation amount corresponding to the compact feature, input the bandwidth resource allocation amount and the compact feature into a channel encoder, and after being processed by the channel encoder, output a modulation symbol sequence, and send the modulation symbol sequence to a physical channel and send it to the receiving end via the physical channel.
8. A data transmission device, characterized in that, Disposed at the receiving end, the data transmission device includes a data receiving module and a rendering module, where The data receiving module is configured to receive the target angle and the modulation symbol sequence sent by the sending end, and perform decoding processing on the modulation symbol sequence to obtain a reconstructed three-dimensional feature; The rendering module is configured to perform rendering processing on the reconstructed three-dimensional feature based on the target angle to obtain the target image corresponding to the target scene at the target angle.
9. An electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, where when the processor executes the program, it implements the method according to any one of claims 1 to 6.
10. A non-transitory computer-readable storage medium, characterized in that, The non-transitory computer-readable storage medium stores computer instructions for causing a computer to execute the method according to any one of claims 1 to 6.