Image semantic transmission method containing depth data
Through semantic communication methods, feature extraction, combined source channel encoding and decoding are performed on RGBD images, solving the problems of image transmission and recovery in harsh channel environments, and achieving high-quality image recovery and compression effects.
Patent Information
- Application Number
- CN202510184653.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-19
- Publication Date
- 2025-06-03
AI Technical Summary
The prior art is difficult to effectively transmit and restore RGBD images in harsh channel environments, resulting in a ‘cliff effect’ and poor image recovery quality.
Using semantic communication method, RGBD images are converted into semantic features through feature extraction networks, encoded using a joint source channel encoder with a dual-branch structure, and decoded and restored through a joint source channel decoder and feature recovery network at the receiving end.
The image recovery quality is significantly improved at lower signal-to-noise ratio, overcomes the ‘cliff effect’, and improves the image compression performance.
Smart Images

Figure CN120091134A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of satellite communication, and relates to an image semantic transmission method containing depth data, which is applicable to image data transmission in harsh channel environments. Background Art
[0002] Compared with traditional RGB images, RGBD images add depth information and play an important role in the transmission of first-line information in modern battlefields. The addition of a dimension in RGBD images compared to RGB leads to a significant increase in data volume, posing higher requirements for the transmission method. Conventional communication technologies use the design idea of source-channel separation coding. In source coding, the image is compressed as much as possible so that each symbol can carry more information. In channel coding, redundancy is added so that the encoded data can resist channel interference and noise. This design idea of source-channel separation coding with independent module design will experience a cliff-like decline in performance when the signal-to-noise ratio is lower than the threshold of a certain module, showing the "cliff effect", and the communication ability fails here.
[0003] As a new communication paradigm, semantic communication realizes more efficient coding and transmission by extracting and utilizing the semantic content of information. Semantic communication uses source-channel joint coding. Through a deep learning model, the original information is converted into a semantic vector or a semantic tensor. After transmission, this tensor can retain certain errors and enter the decoding link in the case of certain errors. The original information is directly decoded and restored by the deep learning model. At present, the research on the semantic information transmission of RGB images is relatively sufficient, but there is no publicly reported research on the semantic extraction, transmission, and recovery technologies for images containing depth information such as RGBD. Summary of the Invention
[0004] The technical problem solved by the present invention is: overcoming the deficiencies of the prior art, and proposing a method for improving the bonding strength between rubber and fiber fabric, which can improve the data compression ability of RGBD images, improve the image recovery quality at a lower signal-to-noise ratio, and overcome the "cliff effect".
[0005] The solution to the technical problem of the present invention is: an image semantic transmission method containing depth data, including the following steps:
[0006] Step S1: At the sending end of image transmission, convert the images in the RGBD dataset into YCbCr color space data and divide them into four channels, namely the luminance y, chrominance u, v, and depth d channels;
[0007] Step S2: For the YCbCr color space data, construct 4 feature extraction networks to respectively perform semantic feature extraction on each input channel y, u, v, d, and after extraction, output the luminance feature yf 、Chromaticity feature u f 、v f and depth feature d f ,and for y f 、u f 、v f are merged to obtain the merged feature yuv f ;
[0008] Step S3, construct a joint source-channel encoder with a two-branch structure, where one branch receives the feature yuv f and the other branch receives the feature d f . After passing through the joint source-channel encoder, the encoded two-way features yuv j and d j are output;
[0009] Step S4, the two encoded features yuv j and d j are transmitted through a wireless channel to obtain the transmitted feature data and
[0010] Step S5, at the receiving end of the image transmission, construct a joint source-channel decoder, and . After decoding and processing by the joint source-channel decoder, the decoded features are output and
[0011] Step S6, construct a feature recovery network symmetric to the feature extraction network structure, and are used as the input of the feature recovery network to obtain the recovered features and
[0012] Step S7, convert the recovered features and into an RGBD image;
[0013] Step S8, use two metrics, the channel bandwidth ratio CBR occupied by the image and the multi-scale structural similarity MS-SSIM, to evaluate the image compression effect and the image recovery quality of the image semantic transmission method respectively.
[0014] Furthermore, the structures of each feature extraction network are the same, and they all use the first residual network and the Swin Transformer Blocks network in cascade for feature extraction;
[0015] Among them, the input information of the first residual network is high-dimensional data of 4 channels of y, u, v, and d. After feature extraction and downsampling processing, the output is low-dimensional 4-channel features;
[0016] The input information of the Swin Transformer Blocks network is low-dimensional 4-channel features. After re-feature extraction processing based on the self-attention mechanism, the output is the low-dimensional v f , u f , v f , d f 4-channel features.
[0017] Furthermore, the structure of the Swin Transformer Blocks network is designed as follows:
[0018] The features of a single channel enter the LayerNorm layer for normalization, and then through the key, query, value KQV mapping, multi-head self-attention processing is performed within the W-MSA local window. After processing, it enters the LayerNorm layer for normalization and the multi-layer perceptron MLP for feature extraction. The above network structure is performed 2 times to form the output features y f , u f , v f , d f .
[0019] Furthermore, the joint source-channel encoder is composed of multiple cascaded coding units. Each coding unit has the same structure and includes a second residual network and a cross-modal attention network;
[0020] Among them, the input of the second residual network is the feature data of 2 channels of yuv f and d f . After feature extraction processing, the output is the feature data of the same dimension of 2 channels;
[0021] The cross-modal attention network is used to eliminate modal redundancy and integrate information between different modalities. Its input is the feature data of the same dimension of 2 channels. After feature encoding processing of the cross-modal attention mechanism, the output is the feature of the same dimension of 2 channels of the encoded yuv j and d j .
[0022] Furthermore, the second residual network adopts the same structure design as the first residual network;
[0023] The structure of the cross-modal attention network is designed as follows:
[0024] The characteristic data of the two channels enter the LayerNorm layer for normalization, and then pass through the key, query, and value KQV mapping to perform information interaction and feature fusion in the non-overlapping local cross-attention window of W-MCA or the shifted cross-attention window of SW-MCA. After processing, it enters the LayerNorm layer for normalization and the multi-layer perceptron MLP for feature extraction. The above network structure is performed twice, and finally the encoded feature yuv is output. j and d j 。
[0025] Furthermore, the joint source-channel decoder is composed of multiple cascaded decoding units. Each decoding unit has the same structure and adopts the same structural design method as the encoding unit in the joint source-channel encoder.
[0026] The two-way feature data after transmission and are sent to the joint source-channel decoder for decoding. The decoding process includes multiple rounds of feature recovery and upsampling operations, and each round of feature recovery and upsampling operation is implemented by a decoding unit.
[0027] Furthermore, in the feature recovery network, first, the input is decomposed into intermediate features and according to luminance and chrominance. For each decomposed data path a first residual network and a Swin Transformer Blocks network are used for feature recovery to obtain the recovered luminance feature chrominance feature and depth feature
[0028] Furthermore, the calculation method of the channel bandwidth ratio CBR occupied by the image is: CBR = k / e, where e represents the source bandwidth and k is the bandwidth occupied by the compressed vector.
[0029] A computer-readable storage medium stores a computer program. When the computer program is executed by a processor, it implements the steps of the method for image semantic transmission with depth data.
[0030] An electronic device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the steps of the method for image semantic transmission with depth data.
[0031] The beneficial effects of the present invention compared with the prior art are:
[0032] (1) Compared with the traditional image transmission method, the present invention uses the RGBD image containing depth information as the information source, which contains richer information and helps to improve the perception ability of the battlefield environment.
[0033] (2) For the RGBD image, the present invention designs a semantic transmission depth model network. In the feature extraction stage, the residual network and Swin Transformer Blocks are adopted to fuse the RGB information and depth information. In the feature recovery stage, the RGB information and depth information are separated and recovered. Compared with the mainstream JPEG method, it has better image compression and recovery performance.
[0034] (3) The present invention uses source-channel joint coding as the coding and decoding method for RGBD information. Compared with the traditional source-channel separate coding method, it can still be decoded under a lower signal-to-noise ratio, overcomes the "cliff effect", and ensures the image recovery quality. Description of the Drawings
[0035] Figure 1 It is the image semantic transmission model containing depth data in the embodiment of the present invention;
[0036] Figure 2 It is the structural schematic diagram of Swin Transformer Blocks in the embodiment of the present invention;
[0037] Figure 3 It is the structural schematic diagram of the cross-modal attention network in the embodiment of the present invention;
[0038] Figure 4 It is the performance simulation diagram of the semantic transmission method of the present invention and the JPEG method under different CBRs;
[0039] Figure 5 It is the performance simulation diagram of the semantic method of the semantic transmission method of the present invention and the traditional method under different SNRs. Detailed Embodiment
[0040] The present invention will be further described below in conjunction with the drawings and embodiments.
[0041] The present invention proposes an image semantic transmission method containing depth data, including the following steps:
[0042] Step S1: At the sending end of the image transmission, convert the images in the RGBD dataset into YCbCr color space data, and divide them into four channels, namely the luminance y, chrominance u, v, and depth d channels, as the input data for step S2;
[0043] Step S2: For the YCbCr color space data, construct 4 feature extraction networks to perform semantic feature extraction on each input channel y, u, v, and d respectively, reduce data redundancy, and after extraction, output the luminance feature y f chrominance feature u f v f and depth feature d f , and for y f u f v f perform merging to obtain the merged feature yuv f ;
[0044] Step S3: Construct a joint source-channel encoder with a two-branch structure, where one branch receives the feature yuv generated in Step S2 f , and the other branch receives the feature d f . After passing through the joint source-channel encoder, output the two encoded features yuv j and d j ;
[0045] Step S4: The two encoded features yuv j and d j are transmitted through a wireless channel to obtain the transmitted feature data and
[0046] Step S5: At the receiving end of the image transmission, construct a joint source-channel decoder and . After decoding and processing by the joint source-channel decoder, output the decoded features and
[0047] Step S6: Construct a feature recovery network symmetric to the feature extraction network structure and as the input of the feature recovery network to obtain the recovered features and
[0048] Step S7: Convert the recovered features and into an RGBD image;
[0049] Step S8: Use two metrics, the channel bandwidth ratio CBR occupied by the image and the multi-scale structural similarity MS-SSIM, to evaluate the image compression effect and image recovery quality of the image semantic transmission method respectively.
[0050] The design solutions in the above method include:
[0051] In step S2, the structures of each feature extraction network are the same. Both use the cascading of the first residual network and the SwinTransformer Blocks network for feature extraction. The deep structure of the first residual network can gradually extract features from low-level to high-level. At the same time, to solve the problem of gradient disappearance or gradient explosion during the training of deep neural networks, skip connections (or called residual connections) are introduced into the residual network, thereby improving the training efficiency and performance of the network. The core of the SwinTransformer Blocks network lies in the design of the Swin Transformer Blocks. These blocks use the non-overlapping local window (W-MSA) self-attention mechanism for calculation, enabling the model to efficiently perform self-attention operations within the local window. At the same time, through the shifted window operation, the information flow between different windows is maintained, enhancing the feature extraction ability of the model. In addition, the Swin Transformer Blocks network adopts a hierarchical design strategy, gradually increasing the receptive field through downsampling layer by layer, which helps the model capture different-scale features from local details to global context.
[0052] Specifically, the input information of the first residual network is high-dimensional data of 4 channels of y, u, v, and d. After feature extraction and downsampling processing, the output is low-dimensional 4-channel features;
[0053] The input information of the Swin Transformer Blocks network is low-dimensional 4-channel features. After re-feature extraction processing based on the self-attention mechanism, the output is the low-dimensional y f 、u f 、v f 、d f 4-channel features.
[0054] Among them, the first residual network Using conventional residual network 。
[0055] Preferably, as Figure 2 shown, the structure design of the Swin Transformer Blocks network is as follows:
[0056] The features of a single channel enter the LayerNorm layer for normalization, and then through the KQV (key, query, value) mapping, multi-head self-attention processing is performed within the W-MSA local window. After processing, it enters the LayerNorm layer for normalization and the Multi-Layer Perceptron (MLP) for feature extraction. The above network structure is performed 2 times to form the output features y f 、u f 、v f 、df Finally, the feature extraction network processes the output features y f , u f , v f from the Swin Transformer Blocks network and combines them to obtain the combined feature yuv f .
[0057] In step S3, the joint source-channel encoder consists of multiple cascaded coding units, each with the same structure, which includes a second residual network and a cross-modal attention network. The input of the second residual network is the feature data of two channels, yuv f and d f . After feature extraction processing, it outputs feature data of the same dimension for the two channels;
[0058] The cross-modal attention network is used to eliminate modal redundancy and integrate information between different modalities. Its input is the feature data of the same dimension for two channels. After feature encoding processing using the cross-modal attention mechanism, it outputs the encoded feature data of the same dimension for the two channels, yuv j and d j .
[0059] Preferably, the structure of the second residual network is the same as that of the first residual network.
[0060] Preferably, as Figure 3 shown, the structure of the cross-modal attention network is designed as follows:
[0061] The feature data of the two channels enter the LayerNorm layer for normalization, and then through KQV (key, query, value) mapping, information interaction and feature fusion are performed in the non-overlapping local cross-attention window of W-MCA or the shifted cross-attention window of SW-MCA. After processing, they enter the LayerNorm layer for normalization and the Multi-Layer Perceptron (MLP) for feature extraction. The above network structure is repeated twice, and finally the encoded features yuv j and d j are output.
[0062] In step S5, the joint source-channel decoder consists of multiple cascaded decoding units, each with the same structure, and the same structural design method as the coding units in the joint source-channel encoder is adopted. The two-channel feature data and transmitted are sent to the joint source-channel decoder for decoding. The decoding process includes multiple rounds of feature recovery and upsampling operations, and each round of feature recovery and upsampling operation is implemented by a decoding unit.
[0063] In step S6, the structure of the feature recovery network is symmetric to that of the feature extraction network. As Figure 1 shown, in the feature recovery network, first, the input is decomposed into intermediate features and according to brightness and chrominance. For each path of the decomposed data a first residual network and Swin Transformer Blocks network are used for feature recovery to obtain the recovered brightness feature chrominance feature and depth feature
[0064] In step S8, two metrics are used to evaluate the image compression and recovery performance of the semantic transmission method. The first is the Channel Bandwidth Ratio (CBR) of the image occupancy as an index to measure the image compression effect. The calculation method designed in the present invention is CBR = k / e, where e represents the source bandwidth and k is the bandwidth occupied by the compressed vector. The second is to use the Multi-Scale Structural Similarity Index (MS-SSIM) as an index to measure the image recovery quality. MS-SSIM evaluates the image at different scales by iteratively applying a low-pass filter and downsampling the filtered image. This method allows combining image details at multiple resolution levels, thus more comprehensively evaluating the image quality. By calculating the comparison values of brightness, contrast, and structure at each scale and combining these values to form the final quality score.
[0065] Example 1
[0066] An image semantic transmission method with depth data describes the semantic feature extraction, transmission, and recovery of RGBD image data based on the semantic communication model framework. The semantic communication model framework includes original RGBD image data preparation, a feature extraction network, a joint source-channel encoder, a joint source-channel decoder, a feature recovery network, and RGBD image data reconstruction, as Figure 1 shown. The specific steps are as follows:
[0067] Step S1: At the sending end, the images in the RGBD dataset (initial dimension [4, 512, 512]) are converted into YCbCr color space data and divided into four channels, namely the brightness (y), chrominance (u, v), and depth (d) channels, as the input data of the model. The data dimension corresponding to each channel is [1, 512, 512], representing brightness, two chrominance channels, and depth information respectively, as the input of the model.
[0068] Step S2: For the YCbCr color space data, design 4 feature extraction networks to perform semantic feature extraction on each input channel (y, u, v, d) respectively, reducing data redundancy, as Figure 1 shown. Each feature extraction network has the same structure, and a residual network and Swin Transformer Blocks network are cascaded for feature extraction. Each input dimension is [1, 512, 512]. After being processed by the residual network and Swin Transformer Blocks, the feature dimension is gradually downsampled to adapt to multi-scale feature extraction. At this time, after downsampling, the feature dimension of each input drops to [3, 64, 64]. The deep structure of the residual network can gradually extract features from low-level to high-level. At the same time, to solve the problem of gradient disappearance or gradient explosion during the training process of deep neural networks, skip connections (or called residual connections) are introduced in the residual network, thereby improving the training efficiency and performance of the network. The core of the Swin Transformer Blocks network lies in the design of the Swin Transformer Blocks. These blocks use the non-overlapping local window (W-MSA) self-attention mechanism to calculate, enabling the model to efficiently perform self-attention operations within the local window. At the same time, through the shifted window operation, the information flow between different windows is maintained, enhancing the feature extraction ability of the model. In addition, Swin Transformer Blocks adopts a hierarchical design strategy, gradually increasing the receptive field through layer-by-layer downsampling, which helps the model capture different scale features from local details to global context.
[0069] Specifically, Swin Transformer Blocks is designed using a hierarchical strategy, as Figure 2 shown. The data first undergoes normalization processing through the LayerNorm (LN) layer. LayerNorm helps to stabilize the training process of the neural network, accelerate convergence, and improve the generalization ability of the model. Secondly, after passing through LayerNorm, the data undergoes the mapping of KQV (key, query, value), as well as the processing of multi-head self-attention (MHA) and multi-layer perceptron (MLP). In W-MSA, these operations are restricted to be carried out within the local window, thus achieving efficient and locally sensitive feature extraction. Finally, the output features y f 、u f 、v f and d f are formed, and y f 、u f 、v f are merged to obtain yuvf and d f , at this time, yuv f has a dimension of [9, 64, 64], and d f has a dimension of [3, 64, 64].
[0070] Step S3: Construct a joint source-channel encoder with a two-branch structure, as Figure 1 shown. One branch receives the yuv f generated in Step S2, and the other branch receives d f . The joint source-channel encoder consists of two parts, one is a residual network and the other is a cross-modal attention network. The cross-modal attention network is used to eliminate modal redundancy and integrate queries between different modalities, as Figure 3 shown. The framework of the cross-modal attention network is similar to that of Swin Transformer Blocks. The main difference is that the cross-modal attention network uses multi-head cross-attention instead of multi-head self-attention to achieve information interaction between modalities. This network mainly includes components such as the LN layer, W-MCA, SW-MCA, and MLP. LN is a normalization layer, W-MCA is a non-overlapping local cross-attention window, and SW-MCA is a shifted cross-attention window. Adding LN to W-MCA and SW-MCA means that information interaction will occur at the positions of the displacement windows, further enhancing the effectiveness of feature fusion. After passing through the joint source-channel encoder, two results are output, namely yuv j and d j , with dimensions of [9, 64, 64] and [3, 64, 64] respectively.
[0071] Step S4: The two encoded features yuv j and d j are transmitted through a wireless channel to obtain and as Figure 1 shown. The channel is set to an AWGN channel, and the signal-to-noise ratio SNR is set to [-2.5 4].
[0072] Step S5: At the receiving end, construct a joint source-channel decoder, as Figure 1 shown. The two received features and (with dimensions of [9, 64, 64] and [3, 64, 64]) are fed into the joint source-channel decoder for decoding. The decoding process includes multiple rounds of feature recovery and upsampling, and each round of operation is implemented through a residual network and a cross-modal attention network. After processing, the joint source-channel decoder outputs the recovered features and (with dimensions of [9, 64, 64] and [3, 64, 64]).
[0073] Step S6: Construct a feature recovery network, whose structure is symmetric to that of the feature extraction network. The and generated in Step S5 are used as the inputs of the feature recovery network. In the feature recovery network, first, is decomposed into and The dimension of each group of features is [3, 64, 64]. For each decomposed data stream residual networks and Swin Transformer Blocks are used for feature recovery, and finally, the recovered features and are obtained. The dimension of each group of features is [1, 512, 512].
[0074] Step S7: Convert the recovered features and into an RGBD image with a dimension of [4, 512, 512].
[0075] Step S8: Two metrics are used to evaluate the image compression and recovery performance of the semantic transmission method. The channel bandwidth ratio (CBR) of the image is used as a metric to measure the image compression effect. The calculation method is CBR = k / e, where e represents the source bandwidth and k is the bandwidth occupied by the compressed vector. The standard multi-scale structural similarity (MS-SSIM) is used as a metric to measure the image recovery quality.
[0076] The comparison method is the mainstream JPEG image compression and recovery method. The mainstream method decomposes the RGBD image into an RGB image and a D depth map, processes them separately, compresses the RGB image using the JPEG method, and compresses the depth map using arithmetic coding for transmission. The specific calculation of JPEG compression is as follows: the length of the original image is [4, 512, 512], the length after encoding is 15728, the range of each encoded data is 0 - 255, and the encoded data is modulated using 4QAM and transmitted through an AWGN channel.
[0077] When the signal-to-noise ratio SNR = 2dB, the calculation result of the CBR of the JPEG method is as follows:[[]]
[0078] k = 15728×8 / 0 = 62912,
[0079] e = 4×512×512 = 1048576,
[0080] CBR = k / e ≈ 0.06.
[0081] The calculation result of CBR for the semantic transmission method in this embodiment is as follows:
[0082] k = 9×64×64 + 3×64×64 = 49152,
[0083] e = 4×512×512 = 1048576,
[0084]
[0085] By comparison, the compression performance of the semantic transmission method is improved by (0.06 - 0.047) / 0.06 = 21.7% compared with the JPEG image method. When the signal-to-noise ratio is 2 dB, set CBR = [0.04 0.15], and the image recovery quality MS-SSIM under different CBRs is statistically analyzed. The results are as Figure 4 shown.
[0086] It can be seen that when CBR is 0.06, the MS-SSIM of the JPEG method is 0.72, while the MS-SSIM of the semantic method is 0.86. Compared with the JPEG method, the semantic method has higher image recovery quality.
[0087] Set CBR to 0.1 and the signal-to-noise ratio SNR to [-2.5 4], and statistically analyze the MS-SSIM performance of the semantic method and the JPEG method under different signal-to-noise ratios, as Figure 5 shown. It can be seen that when the signal-to-noise ratio is less than -0.5, the MS-SSIM of the JPEG method is 0.35, and the image recovery performance shows a "cliff-like" decline, that is, the "cliff effect", while the MS-SSIM of the semantic method is 0.87, and the recovery performance is good, overcoming the "cliff effect" in the traditional method.
[0088] It can be seen that a semantic transmission method for RGBD image data proposed by the present invention improves the efficient compression and high-quality image recovery of RGBD image data in a harsh channel environment through RGBD image data feature extraction, joint source-channel coding, joint source-channel decoding, feature recovery network, and RGBD image data reconstruction.
[0089] This application provides a computer-readable storage medium, and the computer-readable storage medium stores computer instructions. When the computer instructions run on a computer, the computer is made to execute Figure 1 the method described above.
[0090] Those skilled in the art should understand that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories and optical memories, etc.) that contain computer-usable program code.
[0091] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in Figure 1 one or more of the processes or multiple processes and / or blocks Figure 1 one or more of the blocks or multiple blocks.
[0092] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that implement the functions specified in Figure 1 one or more of the processes or multiple processes and / or blocks Figure 1 one or more of the blocks or multiple blocks.
[0093] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in Figure 1 one or more of the processes or multiple processes and / or blocks Figure 1 one or more of the blocks or multiple blocks.
[0094] Obviously, those skilled in the art can make various modifications and variations to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalent technologies, the present application also intends to include these modifications and variations.
[0095] The content not described in detail in the specification of the present invention belongs to the well-known technology of those skilled in the art.
Claims
1. A method for transmitting image semantics containing depth data, characterized in that: The following steps are involved: Step S1: at the transmitting end of the image transmission, the image in the RGBD data set is converted into YCbCr color space data and divided into four channels, namely, brightness y, chrominance u, v and depth d channels; Step S2: For YCbCr color space data, four feature extraction networks are constructed to extract semantic features from each input channel y, u, v, and d respectively, and the brightness feature y is output after extraction. f 、Chromaticity feature u f 、v f and the deep feature d f , and for y f 、u f 、v f Merge to obtain the merged feature yuv f ; Step S3: construct a dual-branch joint source channel encoder, where one branch receives the feature yuv f , the other branch receives feature d f , after the joint source channel encoder, the encoded two-way feature yuv is output j and d j ; Step S4, the encoded two-way feature yuv j and d j Transmit through wireless channels to obtain the transmitted characteristic data and Step S5: at the receiving end of the image transmission, construct a joint source-channel decoder. and After decoding by the joint source channel decoder, the decoded features are output and Step S6: construct a feature recovery network that is symmetrical to the feature extraction network structure. and As the input of the feature recovery network, the restored features are obtained and Step S7: restore the features and Convert to RGBD image; Step S8, using two indicators, channel bandwidth ratio (CBR) and multi-scale structural similarity (MS-SSIM), to respectively evaluate the image compression effect and image restoration quality of the image semantic transmission method.
2. The method for transmitting image semantics containing depth data according to claim 1, characterized in that: The structure of each feature extraction network is the same, and both use the first residual network and the Swin Transformer Blocks network cascade for feature extraction; The input information of the first residual network is high-dimensional data of y, u, v, and d4 channels, and after feature extraction and downsampling, the output is low-dimensional 4-channel features; The input information of the Swin Transformer Blocks network is a low-dimensional 4-channel feature. After re-feature extraction based on the self-attention mechanism, the output feature-extracted low-dimensional y f 、u f 、v f d f 4-channel features.
3. The method for transmitting image semantics containing depth data according to claim 2, characterized in that: The structure design of the SwinTransformer Blocks network is as follows: The features of a single channel enter the LayerNorm layer for normalization, and then through the key, query, value KQV mapping, multi-head self-attention processing is performed in the W-MSA local window. After processing, it enters the LayerNorm layer for normalization and multi-layer perceptron MLP for feature extraction. The above network structure is performed twice to form the output feature y f 、u f 、v f d f .
4. The method for transmitting image semantics containing depth data according to claim 3, characterized in that: The joint source channel encoder is composed of a plurality of cascaded encoding units, each encoding unit having the same structure, including a second residual network and a cross-modal attention network; Among them, the input of the second residual network is yuv f and d f After feature extraction processing is performed on the feature data of the two channels, feature data of the same dimension of the two channels are output; The cross-modal attention network is used to eliminate modal redundancy and integrate information between different modalities. Its input is the feature data of the same dimension of two channels. After the feature encoding processing of the cross-modal attention mechanism, the encoded two-channel feature data of the same dimension yuv is output. j and d j .
5. The method for transmitting image semantics containing depth data according to claim 4, characterized in that: The second residual network and the first residual network adopt the same structural design; The structure of the cross-modal attention network is designed as follows: The feature data of the two channels enter the LayerNorm layer for normalization, and then through the key, query, value KQV mapping, information interaction and feature fusion are performed in the W-MCA non-overlapping local cross attention window or the SW-MCA shifted cross attention window. After processing, it enters the LayerNorm layer for normalization and multi-layer perceptron MLP for feature extraction. The above network structure is performed twice, and finally the encoded feature yuv is output j and d j .
6. The method for transmitting image semantics containing depth data according to claim 5, characterized in that: The joint source channel decoder is composed of a plurality of cascaded decoding units, each decoding unit has the same structure and adopts the same structural design as the encoding unit in the joint source channel encoder; Two-way feature data after transmission and The data is sent to the joint source channel decoder for decoding. The decoding process includes multiple rounds of feature recovery and upsampling operations, and each round of feature recovery and upsampling operations is implemented by a decoding unit.
7. The method for transmitting image semantics containing depth data according to claim 3, characterized in that: In the feature recovery network, the input Decompose into intermediate features according to brightness and chromaticity and For each data after decomposition The first residual network and Swin Transformer Blocks network are used to restore features and obtain the restored brightness features. Chromaticity characteristics and deep features 8. The method for transmitting image semantics containing depth data according to claim 1, characterized in that: The channel bandwidth ratio CBR occupied by the image is calculated as follows: CBR=k / e, wherein e represents the information source bandwidth, and k represents the bandwidth occupied by the compressed vector.
9. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 8 are implemented.
10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 8 are implemented.