A channel environment adaptive sampling-semantic-channel coding joint optimization method and system

By adopting a channel environment-adaptive sampling-semantic-channel coding joint optimization method, the problem of communication performance degradation under low signal-to-noise ratio is solved, enabling image transmission in environments with varying signal-to-noise ratios, reducing sensing and computational overhead, and ensuring image reconstruction quality.

CN119561650BActive Publication Date: 2025-11-11TSINGHUA UNIVERSITY +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411509807.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-28
Publication Date
2025-11-11
Estimated Expiration
2044-10-28

AI Technical Summary

Technical Problem

Existing communication transmission methods suffer from drastic performance degradation at low signal-to-noise ratios. Furthermore, existing semantic communication systems are designed relatively independently from sensing systems, failing to consider the integration of sampling and communication at the semantic level, and thus cannot support image transmission in environments with varying signal-to-noise ratios.

Method used

A channel environment adaptive sampling-semantic-channel coding joint optimization method is adopted. Semantic sampling and coding are performed at the transmitting end, and preprocessing and decoding are performed at the receiving end. The encoder and decoder, which are alternately cascaded with multiple residual convolution modules and channel environment adaptive modules, realize adaptive feature extraction and reconstruction of the channel environment.

Benefits of technology

It supports image transmission in environments with varying signal-to-noise ratios, ensures image reconstruction quality, reduces the number of sensing operations and computational overhead, and adapts to the high requirements of dynamic transmission environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119561650B_ABST
    Figure CN119561650B_ABST
Patent Text Reader

Abstract

The application provides a channel environment adaptive sampling-semantic-channel coding joint optimization method and system, relates to the technical field of semantic communication, and performs semantic sampling on an image to be transmitted by a sending end to obtain a semantic sampling result; the sending end inputs the semantic sampling result into an encoder to obtain an image feature vector to be transmitted; a receiving end receives an image feature vector affected by physical channel noise and performs preprocessing; the receiving end inputs the preprocessing result into a decoder to obtain an image feature vector to be semantically reconstructed; and the receiving end performs semantic reconstruction according to the image feature vector to be semantically reconstructed to obtain a reconstructed image. In the method provided in the application, the characteristics of a signal source and a channel are combined, the semantic sampling and reconstruction are considered, the image transmission in a signal-to-noise ratio variable environment can be supported, and the reconstruction quality of the image is ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of semantic communication technology, and in particular to a sampling-semantic-channel coding joint optimization method and system that is adaptive to the channel environment. Background Technology

[0002] With increasing communication demands, existing communication transmission methods need to meet the requirements of low sensing costs and low computing overhead, and adapt to the high demands of dynamic transmission environments.

[0003] However, the proposed independent transmission scheme suffers from a cliff effect at low signal-to-noise ratios, which causes a sharp deterioration in performance. This results in a sudden and significant reduction in the performance and quality of the transmitted images at the receiving end. Furthermore, existing semantic communication systems and sensing systems are designed relatively independently, without considering the integrated design of sampling and communication at the semantic level.

[0004] Furthermore, the semantic communication joint encoding and decoding architecture based on single signal-to-noise ratio training cannot overcome the performance loss caused by the mismatch between training and testing channel environments, and is difficult to support image transmission in environments with varying signal-to-noise ratios.

[0005] Therefore, there is an urgent need for a new method that integrates sampling and communication. Summary of the Invention

[0006] This application provides a channel environment adaptive sampling-semantic-channel coding joint optimization method and system to solve the problem that existing transmission methods do not consider the integration of sampling and communication at the semantic level and are difficult to support image transmission in environments with varying signal-to-noise ratios.

[0007] In a first aspect of this application, a channel environment adaptive sampling-semantic-channel coding joint optimization method is proposed, the method comprising:

[0008] The sending end performs semantic sampling on the image to be transmitted to obtain the semantic sampling result;

[0009] The transmitting end inputs the semantic sampling result into the encoder to obtain the image feature vector to be transmitted. The encoder is composed of multiple residual convolution modules and multiple channel environment adaptive modules cascaded alternately. The input of each channel environment adaptive module in the encoder is the signal-to-noise ratio and the intermediate feature output by the previous residual convolution module, and the output is the intermediate feature that is adapted to the signal-to-noise ratio of the current channel environment.

[0010] The receiving end receives image feature vectors affected by physical channel noise and performs preprocessing.

[0011] The receiving end inputs the preprocessing result into the decoder to obtain the image feature vector that needs semantic reconstruction. The decoder is composed of multiple residual transposed convolutional modules and multiple channel environment adaptive modules cascaded alternately. The input of each channel environment adaptive module in the decoder is the signal-to-noise ratio and the intermediate feature output by the previous residual transposed convolutional module, and the output is the intermediate feature that is adapted to the signal-to-noise ratio of the current channel environment.

[0012] The receiving end performs semantic reconstruction based on the image feature vector that needs semantic reconstruction, and obtains the reconstructed image.

[0013] Optionally, the sending end performs semantic sampling on the image to be transmitted to obtain semantic sampling results, including:

[0014] After normalizing the image to be transmitted, the sending end divides it into blocks to obtain multiple image blocks;

[0015] The transmitting end inputs the image to be transmitted and the signal-to-noise ratio into the semantic scanning network to obtain a semantic information distribution map;

[0016] The transmitting end uses a block sampling ratio aggregation correction algorithm to obtain a semantic sampling ratio allocation map based on the semantic information distribution map and the signal-to-noise ratio;

[0017] The sending end performs singular value decomposition on the training image set to obtain the initial matrix;

[0018] For the i-th image block among multiple image blocks, the semantic sampling ratio of the i-th image block is obtained based on the given overall sampling rate and the semantic information distribution map;

[0019] Based on the semantic sampling ratio of the i-th image block and the initial matrix, the semantic sampling matrix of the i-th image block is obtained;

[0020] Based on the i-th image patch and its semantic sampling matrix, semantic sampling is performed to obtain the i-th initial semantic sampling result;

[0021] Based on the i-th initial semantic sampling result and the semantic sampling matrix of the transposed i-th image block, block initialization sampling rate alignment is performed to obtain the final semantic sampling result.

[0022] Optionally, the semantic scanning network consists of a first convolutional layer, a first channel environment adaptive module, multiple simple residual blocks, a second channel environment adaptive module, and convolutional layers cascaded sequentially; the semantic scanning network is used to evaluate the saliency of semantic information at each location on the image to be transmitted.

[0023] Optionally, the sending end employs a block sampling ratio aggregation correction algorithm to obtain a semantic sampling ratio allocation map based on the semantic information distribution map and the signal-to-noise ratio, including:

[0024] The semantic information distribution map is divided into sections of size [size missing] in a left-to-right and top-to-bottom order. B × B of l Image blocks;

[0025] right l The image blocks are aggregated to obtain the semantic sampling ratio allocation map, which contains the block sampling rate of all image block allocations. 。

[0026] Optionally, the receiving end performs semantic reconstruction based on the image feature vector requiring semantic reconstruction to obtain a reconstructed image, including:

[0027] The receiving end expands the image feature vector that needs semantic reconstruction into l For each image block, a random transformation enhancement is performed, followed by block gradient descent to obtain the intermediate results for each image block. The intermediate results for all image blocks are then aggregated.

[0028] The semantic sampling ratio allocation map sent by the receiving end is used as the corresponding sampling ratio in the semantic sampling ratio allocation map as the filling element to obtain the extended semantic sampling ratio allocation map;

[0029] The features of the extended semantic sampling ratio allocation map are embedded into the feature space through a semantic extraction network oriented towards semantic sampling ratio allocation map and signal-to-noise ratio, and a feature map extracted based on semantics and signal-to-noise ratio is output.

[0030] The intermediate results of each image patch are concatenated with the feature maps extracted based on semantics and signal-to-noise ratio, and the concatenation results are processed based on a deep learning near-end mapping network.

[0031] The processing result is added to the intermediate result of each image block to obtain the reconstructed image data, and then inverse normalization is performed to obtain the reconstructed image.

[0032] Optionally, the channel environment adaptive module includes a sequentially connected average pooling layer, a first fully connected layer, a PReLU layer, a second fully connected layer, and a sigmoid layer; the channel environment adaptive module reprocesses the input features based on a known signal-to-noise ratio, as follows:

[0033] The intermediate features output by the previous module connected to the channel environment adaptive module are subjected to average pooling to obtain the average pooling result.

[0034] The average pooling result is concatenated with the signal-to-noise ratio to obtain the concatenated feature;

[0035] The splicing features are processed by a first fully connected layer, a PReLU layer, a second fully connected layer, and a sigmoid layer to obtain an adaptive scaling factor for the channel environment.

[0036] Based on the channel environment adaptive scaling factor and the intermediate features output by the previous module connected to the channel environment adaptive module, intermediate features that adapt to the signal-to-noise ratio of the current channel environment are obtained.

[0037] Optionally, the channel environment-adaptive sampling-semantic-channel coding joint optimization method is implemented through a channel environment-adaptive sampling-semantic-channel coding joint model, and the training process includes:

[0038] Calculate the loss between the reconstructed image and the corresponding image in the training image set;

[0039] Based on the loss, the parameters in the semantic sampling module, encoder, decoder, and semantic reconstruction module are optimized to obtain the channel environment-adaptive sampling-semantic-channel coding joint model.

[0040] A second aspect of this application proposes a channel environment adaptive sampling-semantic-channel coding joint optimization system, the system comprising:

[0041] The semantic sampling module is used to perform semantic sampling on the image to be transmitted and obtain the semantic sampling results.

[0042] The encoding module is used to input the semantic sampling result into the encoder to obtain the image feature vector to be transmitted. The encoder is composed of multiple residual convolution modules and multiple channel environment adaptive modules cascaded alternately. The input of each channel environment adaptive module in the encoder is the signal-to-noise ratio and the intermediate feature output by the previous residual convolution module, and the output is the intermediate feature that is adaptive to the signal-to-noise ratio of the current channel environment.

[0043] The preprocessing module is used to receive image feature vectors affected by physical channel noise and perform preprocessing.

[0044] The decoding module is used to input the preprocessing results into the decoder to obtain the image feature vector that needs to be semantically reconstructed. The decoder is composed of multiple residual transposed convolutional modules and multiple channel environment adaptive modules cascaded alternately. The input of each channel environment adaptive module in the decoder is the signal-to-noise ratio and the intermediate features output by the previous residual transposed convolutional module, and the output is the intermediate features that are adapted to the signal-to-noise ratio of the current channel environment.

[0045] The semantic reconstruction module is used to perform semantic reconstruction based on the image feature vector that needs to be semantically reconstructed, and obtain the reconstructed image.

[0046] In a third aspect of this application, an electronic device is provided, including a memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement the channel environment adaptive sampling-semantic-channel coding joint optimization method described in any one of the first aspects above.

[0047] In a fourth aspect of this application, a computer-readable storage medium is provided, on which a computer program / instruction is stored, which, when executed by a processor, implements the channel environment adaptive sampling-semantic-channel coding joint optimization method described in any one of the first aspects above.

[0048] This application includes the following advantages: It provides a channel environment adaptive sampling-semantic-channel coding joint optimization method and system. The transmitting end performs semantic sampling on the image to be transmitted to obtain semantic sampling results. The transmitting end inputs the semantic sampling results into an encoder to obtain the image feature vector to be transmitted. The encoder is composed of multiple residual convolutional modules and multiple channel environment adaptive modules alternately cascaded. The input of each channel environment adaptive module in the encoder is the signal-to-noise ratio (SNR) and the intermediate feature output by the previous residual convolutional module, and the output is an intermediate feature adaptive to the SNR of the current channel environment. The receiving end receives the image feature vector affected by physical channel noise and performs preprocessing. The receiving end inputs the preprocessing result into a decoder to obtain the image feature vector to be semantically reconstructed. The decoder is composed of multiple residual transposed convolutional modules and multiple channel environment adaptive modules alternately cascaded. The input of each channel environment adaptive module in the decoder is the SNR and the intermediate feature output by the previous residual transposed convolutional module, and the output is an intermediate feature adaptive to the SNR of the current channel environment. The receiving end performs semantic reconstruction based on the image feature vector to be semantically reconstructed to obtain the reconstructed image. The method proposed in this application incorporates an attention-based channel environment adaptation module into the semantic-based sampling and reconstruction, encoding-decoding end-to-end joint design, further combining the characteristics of the source and the channel. While considering semantic sampling and reconstruction, it can support image transmission in environments with varying signal-to-noise ratios and ensure the quality of image reconstruction. Attached Figure Description

[0049] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the description of the embodiments of this application will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0050] Figure 1This is a flowchart illustrating the steps of a channel environment adaptive sampling-semantic-channel coding joint optimization method provided in an embodiment of this application;

[0051] Figure 2 A flowchart illustrating a channel environment adaptive sampling-semantic-channel coding joint optimization method proposed in this application embodiment;

[0052] Figure 3 This is a schematic diagram illustrating the process of semantic sampling of an image by a semantic sampling module according to an embodiment of this application;

[0053] Figure 4 This is a schematic diagram of a semantic-channel-encoder-decoder encoding and decoding process proposed in an embodiment of this application;

[0054] Figure 5 This is a schematic diagram illustrating the process of semantic reconstruction of an image by a semantic reconstruction module provided in an embodiment of this application;

[0055] Figure 6 This is a schematic diagram of the architecture of a channel environment adaptive module provided in an embodiment of this application;

[0056] Figure 7 This is an architecture diagram of a channel environment adaptive sampling-semantic-channel coding joint optimization system proposed in an embodiment of this application;

[0057] Figure 8 This is a schematic diagram of an electronic device provided in an embodiment of this application. Detailed Implementation

[0058] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0059] With increasing communication demands, existing communication transmission methods need to meet the requirements of low sensing costs and low computing overhead, and adapt to the high demands of dynamic transmission environments.

[0060] However, the proposed independent transmission scheme suffers from a cliff effect at low signal-to-noise ratios, which causes a sharp deterioration in performance. That is, below a certain signal-to-noise ratio, the number of erroneous codewords in the transmitted encoded image far exceeds the error correction capability of the error control encoding and decoding. The performance and quality of the transmitted image at the receiving end will suddenly drop significantly. Furthermore, the existing semantic communication system and sensing system are designed relatively independently, without considering the integrated design of sampling and communication at the semantic level.

[0061] Furthermore, the semantic communication joint encoding and decoding architecture based on single signal-to-noise ratio training cannot overcome the performance loss caused by the mismatch between training and testing channel environments, and is difficult to support image transmission in environments with varying signal-to-noise ratios.

[0062] Based on this, this application proposes a transmission method that adapts to the channel environment and integrates sampling and communication.

[0063] In the first aspect of this application, a channel environment adaptive sampling-semantic-channel coding joint optimization method is provided, see reference. Figure 1 , Figure 1 This is a flowchart illustrating the steps of a channel environment adaptive sampling-semantic-channel coding joint optimization method provided in this application embodiment. The method includes the following steps:

[0064] Step 101: The sending end performs semantic sampling on the image to be transmitted to obtain the semantic sampling result;

[0065] Step 102: The transmitting end inputs the semantic sampling result into the encoder to obtain the image feature vector to be transmitted. The encoder is composed of multiple residual convolution modules and multiple channel environment adaptive modules cascaded alternately. The input of each channel environment adaptive module in the encoder is the signal-to-noise ratio and the intermediate feature output by the previous residual convolution module, and the output is the intermediate feature that is adapted to the signal-to-noise ratio of the current channel environment.

[0066] Step 103: The receiving end receives the image feature vector affected by physical channel noise and performs preprocessing;

[0067] Step 104: The receiving end inputs the preprocessing result into the decoder to obtain the image feature vector that needs semantic reconstruction. The decoder is composed of multiple residual transposed convolutional modules and multiple channel environment adaptive modules cascaded alternately. The input of each channel environment adaptive module in the decoder is the signal-to-noise ratio and the intermediate feature output by the previous residual transposed convolutional module, and the output is the intermediate feature that is adapted to the signal-to-noise ratio of the current channel environment.

[0068] Step 105: The receiving end performs semantic reconstruction based on the image feature vector that needs semantic reconstruction to obtain the reconstructed image.

[0069] To clearly describe the channel environment-adaptive sampling-semantic-channel coding joint optimization method proposed in this application, combined with Figure 2 This method will be explained in detail. Figure 2 This is a flowchart illustrating a channel environment adaptive sampling-semantic-channel coding joint optimization method proposed in an embodiment of this application.

[0070] In specific implementation step 101, the above method is applied to an end-to-end image transmission system, that is, transmission from the sending end to the receiving end. The sending end can be the source end, and the receiving end can be the sink end. Considering that the semantic communication system with joint source-channel coding and decoding currently proposed does not have semantic sampling and semantic reconstruction modules, downsampling processing of the object to be transmitted based on a certain strategy is very important in some scenarios where it is necessary to reduce the number of sensing operations and computational overhead. For example, multiple medical imaging is harmful to the patient's health, some sensors such as satellite imaging are expensive to operate, and excessively large amounts of data and models will increase the network storage and computational burden. To avoid the above situations, this application sets up a semantic sampling module based on downsampling at the sending end to perform semantic sampling on the image to be transmitted and obtain semantic sampling results. Semantic downsampling processing helps to ensure that image content with high semantic importance or rich semantic information is accurately extracted and encoded in subsequent processing, while redundant other parts are reduced in attention or even ignored, thereby reducing the source size and further reducing the transmission overhead of communication.

[0071] In one optional embodiment of this application, see [reference] Figure 3 , Figure 3 This is a schematic diagram of the semantic sampling module for semantic sampling of an image proposed in an embodiment of this application. When performing semantic sampling on an image, given that the overall average sampling rate of the entire image is determined, the image is divided into multiple image blocks. The sampling rate allocation for each image block in each image is accomplished by semantic information and signal-to-noise ratio. That is, the number of samples for each block is determined based on the given overall average sampling number and is jointly determined by semantic information scanning and block sampling rate aggregation correction.

[0072] In one optional embodiment of this application, the semantic scanning network is composed of a first convolutional layer, a first channel environment adaptive module, multiple simple residual blocks, a second channel environment adaptive module, and convolutional layers sequentially cascaded. This semantic scanning network is used to evaluate the saliency of semantic information at each location in the image to be transmitted, highlighting the importance of different regions in the image. Because the channel environment adaptive module focuses attention on the channel signal-to-noise ratio (SNR), the semantic scanning network comprehensively considers the channel conditions and can adapt to the SNR of the current channel environment, achieving more accurate allocation by combining the characteristics of the source channel. Furthermore, to achieve accurate sampling ratio allocation for each block, this application proposes a block sampling ratio aggregation correction algorithm to process the semantic information distribution map and SNR, obtaining a semantic sampling ratio allocation map, which is shared with the reconstruction end. Simultaneously, based on a given overall sampling rate, the number of samples for each block after image segmentation is adjusted and allocated.

[0073] In one optional embodiment of this application, the sending end performs semantic sampling on the image to be transmitted to obtain semantic sampling results, specifically including the following process: First, the sending end normalizes the image to be transmitted, and then divides the normalized image S into blocks to obtain multiple image blocks. The image to be transmitted is... n is the image size. Let H be the set of real numbers, W be the image height, and C be the image width, i.e., n = H × W × C. The set of all image patches can be represented as: 'l' represents the number of image patches, and its value is an integer greater than or equal to 2. Then, the transmitting end sends the image to be transmitted along with the signal-to-noise ratio... μ Input to semantic scanning network D In the process, a semantic information distribution map is obtained. M The sending end employs a block sampling ratio aggregation correction algorithm based on the semantic information distribution map. M and the signal-to-noise ratio μ The semantic sampling ratio allocation map is obtained. R The sending end sends the training image set. Perform singular value decomposition (SVD) To determine the number of images in the training image set, an initial matrix A is obtained. The initial matrix A has a size of N×N, where N is the dimension of its feature basis vectors. The features of the initial matrix A are based on decreasing importance from the first row to the last row. This means the generator matrix can be intentionally cropped to obtain sampling matrices with arbitrary sampling rates and can be learned from the data. For example, if image patch A has more details than image patch B, then the image patch is more important, and the subsequent semantic sampling matrix will retain more rows from the initial matrix A, meaning it will collect more information.

[0074] Furthermore, based on the given overall sampling rate r and the aforementioned semantic information distribution map M, the semantic sampling ratio of the i-th image patch is obtained. Based on the semantic sampling ratio of the i-th image patch Together with the initial matrix A mentioned above, we obtain the semantic sampling matrix of the i-th image patch. The sending end uses the semantic sampling matrix of the i-th image block and the i-th image block as its basis. Perform semantic sampling to obtain the i-th initial semantic sampling result. The initial semantic sampling result set for each image patch can be represented as: Based on the i-th initial semantic sampling result The semantic sampling matrix of the i-th image patch after transposition Perform block initialization sampling rate alignment to obtain the final semantic sampling result. The set of semantic sampling results for each image patch can be represented as: .

[0075] In the process of sampling an image using a semantic sampling module, the parameter of the semantic sampling module is ξ, and the semantic sampling process can be represented as:

[0076] in, (·) represents the semantic sampling module, and s represents the image to be sampled. μ denoted as signal-to-noise ratio, and r as the given overall sampling rate.

[0077] In one optional embodiment of this application, the transmitting end employs a block sampling ratio aggregation correction algorithm based on the semantic information distribution map M and the signal-to-noise ratio. μ The semantic sampling ratio allocation map R is obtained, specifically including the following method: The semantic information distribution map M is divided into sections of size M in a left-to-right and top-to-bottom order. B × B of l There are 1 image block, where the number of image blocks l is the same as the number of image blocks l obtained by dividing the normalized image S into blocks; for l The image blocks are aggregated to obtain the block sampling rate that includes all block allocations. The semantic sampling ratio allocation diagram R above.

[0078] In one optional embodiment of this application, during the actual semantic sampling process, due to the limited size of the semantic sampling matrix, the available sampling rate for each block can only be from... Choose from, among which q i Represented as the first i The precise sampling size of the block is given by N, which is the dimension of the eigenvectors of the initial matrix A. The block sampling ratio aggregation correction algorithm can be summarized as softmax normalization, sumpooling aggregation, and error correction. The semantic sampling ratio allocation graph obtained using the block sampling ratio aggregation correction algorithm is as follows: First, softmax normalization is applied to the semantic information distribution graph M. B × B Heap pooling yields an aggregated weight map, which is then compared with the target measurement size and... ql ( ql Multiplying the sample sizes of l image patches yields the measurement size mapping. Q Then to Q Perform iterative verification and correction. In each correction iteration, [the following is done:] Q Perform shearing and discretization so that all its elements are [0, ..., ... N The integers within ] specify the upper bound of the image patch measurement size, and finally the semantic sampling ratio allocation map R is obtained.

[0079] In specific implementation step 102, the semantic sampling result obtained by the semantic sampling module is input into the encoder in the semantic-channel encoder-decoder to obtain the image feature vector to be transmitted. The process can be found in [reference needed]. Figure 4 , Figure 4 This is a schematic diagram illustrating the encoding and decoding process of a semantic-channel-encoder-decoder according to an embodiment of this application. The parameter of the semantic channel-encoder is θ, and the encoding process can be represented as follows:

[0080] in, (·) represents the semantic channel encoder, and x represents the semantic sampling result.

[0081] After the semantic sampling result x is encoded by the encoder, it undergoes shape reshaping to generate a vector of size 1×k. Before outputting the obtained image feature vector to be transmitted, in order to satisfy the average power constraint, that is, to satisfy the following condition:

[0082]

[0083] The encoded symbol z (corresponding to the image feature vector) is obtained by applying the power normalization formula to the vector. The calculation formula is as follows:

[0084] ,

[0085] in, for The complex conjugate, For the output of the last submodule of the encoder before power normalization, this application sets the normalized power. =1. Then the encoded symbol z is input into the physical channel for transmission.

[0086] In one optional embodiment of this application, the encoder is composed of multiple residual convolutional modules (RCMs) and multiple adaptive channel environments (ACAMs) cascaded alternately, and the input of each ACAM in the encoder is the signal-to-noise ratio. μ (Corresponding to SNR in the figure) and the intermediate features output by the previous residual convolutional module RCM, the output is an intermediate feature that adaptively adjusts the signal-to-noise ratio for the current channel environment. In an optional embodiment of this application, the output channel size of the last convolutional layer of the last residual convolutional module RCM in the encoder is changed. k This allows for the creation of models with different bandwidth ratios (CBRs), enabling the encoder to better adapt to changes in the channel environment.

[0087] In specific implementation step 103, the encoded symbol z is determined by the function... (·) represents the noisy physical channel transmission. Let the channel gain be 1. This channel type is either an AWGN channel or a Rayleigh fading channel. After being affected by noise, the channel output is represented as follows: The AWGN channel transmission model can be represented as:

[0088] in, n It is noise.

[0089] In other words, the image feature vector received by the receiving end is an image feature vector affected by physical channel noise. The received image feature vector needs to be preprocessed before it is transmitted to the decoder for decoding. The preprocessing includes padding and reshaping.

[0090] In specific implementation step 104, the sending end inputs the preprocessing result into the decoder, and performs decoding based on the signal-to-noise ratio to obtain the image feature vector that needs semantic reconstruction. The decoder parameters are: The decoding process can be represented as:

[0091] in, (·) represents the semantic channel-decoder.

[0092] See Figure 4 The decoder described above is composed of multiple Residual Transposed Convolutional Modules (RTCMs) and multiple Adaptive Channel Modules (ACAMs) cascaded alternately. The input to each ACAM in the decoder is the signal-to-noise ratio. μ The intermediate features, combined with the intermediate features output from the previous residual transposed convolutional module (RTCM), are used to output intermediate features that are adaptive to the signal-to-noise ratio (SNR) of the current channel environment. Figure 4 The ↓↑ symbols on each residual convolutional module or residual transposed convolutional module represent upsampling and downsampling operations, respectively, and the | symbol represents the interval. Taking the numbers on the first residual convolutional module (RCM) in the encoder as an example, "9×9×256" represents the width W, height H, and number of output channels C of the convolutional kernel used in that module, respectively.

[0093] In one optional embodiment of this application, the encoder in the semantic-channel-encoder-decoder is composed of five residual convolutional modules (RCM) and five channel environment adaptive modules (ACAM) cascaded alternately, and the decoder is composed of five residual transposed convolutional modules (RTCM) and five channel environment adaptive modules (ACAM) cascaded alternately.

[0094] In specific implementation step 105, the receiving end performs semantic reconstruction based on the image feature vector requiring semantic reconstruction to obtain the reconstructed image. In this embodiment, semantic reconstruction is implemented through a semantic reconstruction module, which performs semantic reconstruction based on the decoding result, signal-to-noise ratio, and the semantic sampling ratio allocation map R shared by the transmitting and receiving ends, followed by inverse normalization processing to obtain the reconstructed image. The above semantic reconstruction process can be represented as follows:

[0095] in, (·) represents the semantic reconstruction module, and the parameter of the semantic reconstruction module is ζ.

[0096] The semantic reconstruction process described above can be found in [reference needed]. Figure 5 , Figure 5 This is a schematic diagram illustrating the process of semantic reconstruction of an image by a semantic reconstruction module according to an embodiment of this application. Specifically, the receiving end performs semantic reconstruction based on the image feature vector to be semantically reconstructed, obtaining a reconstructed image, including the following process: First, the receiving end converts the aforementioned image feature vector to be semantically reconstructed... Expand as l Given a set of image patches, the set of all expanded image patches can be represented as: , This refers to processing image patches that are the results of the previous iteration. Each image patch undergoes random transformation enhancement, followed by block gradient descent, to obtain an intermediate result for each image patch, which can be represented as: And summarizing the intermediate results of all image patches, it can be represented as The intermediate results of all image blocks are sent to the next stage for joint restoration. Random transformation enhancement of each image block refers to randomly performing enhancements such as vertical flipping, horizontal flipping, scaling, random brightness / darkening, contrast adjustment, hue adjustment, and multi-image stitching according to a set ratio. Using random transformation enhancement can better address the bottleneck between two adjacent iterations. The following formula is used when performing block gradient descent processing on the randomly transformed enhanced image blocks:

[0097]

[0098] Where k represents the current iteration round, and the total number of iteration rounds is... N p , Let be the iteration step size for the k-th round. Let be the semantic sampling matrix of the i-th image patch, and T denote the transpose of the semantic sampling matrix of the i-th image patch. For the i-th initial semantic sampling result , This refers to the i-th image patch expanded in the (k-1)-th round.

[0099] Intermediate results obtained by folding the i-th image patch The i-th image patch expanded in the k-th round is obtained according to the following formula. :

[0100] .

[0101] This formula represents the process of approximate point mapping, where, The semantic sampling result for the i-th image patch. Here, is the regularization term and its scaling factor, and argmin represents the parameter value at which a function reaches its minimum value in its domain.

[0102] Then, the semantic sampling ratio allocation map R sent by the receiving end is used as the filling element to obtain the extended semantic sampling ratio allocation map. Furthermore, a semantic extraction network oriented towards semantic sampling ratio allocation graphs and signal-to-noise ratio (i.e., the SNR of the input in the graph) is used. The extended semantic sampling ratio allocation map The features are embedded into the feature space, and the output is a feature map extracted based on semantics and signal-to-noise ratio. ; the intermediate results of each of the above image patches The feature maps extracted based on semantics and signal-to-noise ratio mentioned above Connectivity, and a deep learning-based proximal mapping network The connection results are processed; finally, the processed result is compared with the intermediate results of each image patch. The data are added together to obtain the reconstructed image data, which is then denormalized to obtain the reconstructed image.

[0103] The aforementioned deep learning-based proximal mapping network This is used to guide the approximate point mapping reconstruction process based on semantics and signal-to-noise ratio, by taking the connection result as input and providing the recovered residual content. Specifically, it is achieved through the following formula:

[0104]

[0105] in, For the image data reconstructed in the k-th round, `concat` represents the connection processing; the meanings of the remaining letters are the same as those described above. An iterative proximal mapping network is used. Np The iteration step size is... ρ The optimized result was obtained.

[0106] In an optional embodiment of this application, the above-mentioned semantic extraction network It consists of a concatenated convolutional layer, a channel environment adaptive module (ACAM), three residual blocks, another channel environment adaptive module (ACAM), and a 1×1 kernel convolutional layer.

[0107] In an optional embodiment of this application, the architecture of the Channel Environment Adaptive Module (ACAM) used in the semantic sampling module, encoder, decoder, and semantic reconstruction module described above can be found in [reference needed]. Figure 6 , Figure 6 This is a schematic diagram of the architecture of a channel environment adaptive module provided in an embodiment of this application. The aforementioned channel environment adaptive module (ACAM) includes an average pooling layer, a first fully connected layer, a PReLU layer, a second fully connected layer, and a sigmoid layer connected in sequence. The function of the channel environment adaptive module ACAM is to further process the input features based on the signal-to-noise ratio (SNR) in the known environment. The specific process is as follows: First, an average pooling operation is performed on the intermediate features output by the previous module connected to the channel environment adaptive module ACAM to obtain an average pooling result. Then, the average pooling result is concatenated with the SNR to obtain a concatenated feature. The concatenated feature is processed through the first fully connected layer, the PReLU layer, the second fully connected layer, and the sigmoid layer to obtain a channel environment adaptive scaling factor. Based on the channel environment adaptive scaling factor and the intermediate features output by the previous module connected to the channel environment adaptive module ACAM, intermediate features that are adaptive to the SNR of the current channel environment are obtained.

[0108] For example, the intermediate feature output by the previous module connected to the aforementioned Channel Environment Adaptive Module (ACAM) is: Each feature has a size of h×w, and c is the number of features. After performing average pooling on this intermediate feature as input, the average pooling result is obtained. , means as follows:

[0109] ,

[0110] Wherein, `gloalAvergePooling` is the average pooling operation. intermediate features of the input In location [ j , k Elements on ] It is the set of real numbers.

[0111] Average pooling results With signal-to-noise ratio μ The splicing features are obtained by splicing. I Among them, splicing operation The concatenation of two features along a specific dimension can be represented as:

[0112]

[0113] Then, a nonlinear network consisting of fully connected layers and activation functions is used to predict the scaling parameter S for each feature. The first fully connected layer... FC The activation function after 1 is the PReLU layer, the second fully connected layer. FC The activation function after step 2 is a sigmoid layer, used to map the output to the range [0, 1], acting as a distribution ratio, as shown below:

[0114]

[0115] Finally, element-to-element multiplication is performed to obtain the convergent input features and the output features based on the signal-to-noise ratio (SNR). , Its size is related to the input features Same, as shown below: .

[0116] In an optional embodiment of this application, the channel environment-adaptive sampling-semantic-channel coding joint optimization method proposed in this application is implemented through a channel environment-adaptive sampling-semantic-channel coding joint model. The training process is as follows: calculate the loss between the reconstructed image and the corresponding image in the training image set; optimize the parameters in the semantic sampling module, encoder, decoder, and semantic reconstruction module based on the loss to obtain the channel environment-adaptive sampling-semantic-channel coding joint model. The set of all end-to-end trainable parameters is referred to as... This indicates that the semantic scanning network, including the semantic sampling module, is included. Generate matrix A, encoder parameters decoder parameters Gradient descent step size in semantic reconstruction module Semantic extraction network Proximal mapping network , is represented as:

[0117]

[0118] In an optional embodiment of this application, for having N b Zhang Dawei training set of images Given the average size of the block samples q i (i.e., overall sampling ratio) During each training round, the channel signal-to-noise ratio μ ∈[0,20]dB and uniformly distributed, i.e. μ from Medium probability random selection, and q i from Medium probability random selection, of which N = B × B For the size of the block, l Using a loss based on mean squared error (mean squared error is used to measure the distortion between the input image and the reconstructed image) for a given number of blocks can achieve stable convergence and higher reconstruction accuracy. (·) indicates an end-to-end processing procedure. Let represent the image reconstructed after transmission through the system. Then, the loss function for end-to-end training can be expressed as follows:

[0119]

[0120] This application proposes a channel environment adaptive sampling-semantic-channel coding joint optimization method. The transmitting end performs semantic sampling on the image to be transmitted, obtaining a semantic sampling result. The transmitting end inputs the semantic sampling result into an encoder to obtain the image feature vector to be transmitted. The encoder is composed of multiple residual convolutional modules and multiple channel environment adaptive modules alternately cascaded. The input of each channel environment adaptive module in the encoder is the signal-to-noise ratio (SNR) and the intermediate feature output by the previous residual convolutional module, and the output is an intermediate feature adaptive to the SNR of the current channel environment. The receiving end receives the image feature vector affected by physical channel noise and performs preprocessing. The receiving end inputs the preprocessing result into a decoder to obtain the image feature vector to be semantically reconstructed. The decoder is composed of multiple residual transposed convolutional modules and multiple channel environment adaptive modules alternately cascaded. The input of each channel environment adaptive module in the decoder is the SNR and the intermediate feature output by the previous residual transposed convolutional module, and the output is an intermediate feature adaptive to the SNR of the current channel environment. The receiving end performs semantic reconstruction based on the image feature vector to be semantically reconstructed to obtain the reconstructed image. The method proposed in this application incorporates an attention-based channel environment adaptation module into the semantic-based sampling and reconstruction, encoding-decoding end-to-end joint design, further combining the characteristics of the source and the channel. While considering semantic sampling and reconstruction, it can support image transmission in environments with varying signal-to-noise ratios and ensure the quality of image reconstruction.

[0121] In a second aspect of this application, a channel environment-adaptive sampling-semantic-channel coding joint optimization system is proposed, see [link to relevant documentation]. Figure 7 , Figure 7 This is an architecture diagram of a channel environment adaptive sampling-semantic-channel coding joint optimization system proposed in an embodiment of this application. The system includes:

[0122] The semantic sampling module 701 is used to perform semantic sampling on the image to be transmitted and obtain the semantic sampling result.

[0123] The encoding module 702 is used to input the semantic sampling result into the encoder to obtain the image feature vector to be transmitted. The encoder is composed of multiple residual convolution modules and multiple channel environment adaptive modules cascaded alternately. The input of each channel environment adaptive module in the encoder is the signal-to-noise ratio and the intermediate feature output by the previous residual convolution module, and the output is the intermediate feature that is adaptive to the signal-to-noise ratio of the current channel environment.

[0124] Preprocessing module 703 is used to receive image feature vectors affected by physical channel noise and perform preprocessing;

[0125] The decoding module 704 is used to input the preprocessing result into the decoder to obtain the image feature vector that needs to be semantically reconstructed. The decoder is composed of multiple residual transposed convolutional modules and multiple channel environment adaptive modules cascaded alternately. The input of each channel environment adaptive module in the decoder is the signal-to-noise ratio and the intermediate feature output by the previous residual transposed convolutional module, and the output is the intermediate feature that is adapted to the signal-to-noise ratio of the current channel environment.

[0126] The semantic reconstruction module 705 is used to perform semantic reconstruction based on the image feature vector that needs to be semantically reconstructed, and obtain the reconstructed image.

[0127] The semantic sampling module includes:

[0128] The segmentation submodule is used to normalize the image to be transmitted and then segment it into multiple image blocks.

[0129] The semantic scanning submodule is used to input the image to be transmitted and the signal-to-noise ratio into the semantic scanning network to obtain a semantic information distribution map;

[0130] The aggregation correction submodule is used to obtain a semantic sampling ratio allocation map based on the semantic information distribution map and the signal-to-noise ratio using a block sampling ratio aggregation correction algorithm.

[0131] The decomposition submodule is used to perform singular value decomposition on the training image set to obtain the initial matrix;

[0132] The semantic sampling ratio acquisition submodule is used to obtain the semantic sampling ratio of the i-th image block among multiple image blocks, based on a given overall sampling rate and the semantic information distribution map;

[0133] The semantic sampling matrix acquisition submodule is used to obtain the semantic sampling matrix of the i-th image block based on the semantic sampling ratio of the i-th image block and the initial matrix;

[0134] The semantic sampling result acquisition submodule is used to perform semantic sampling based on the i-th image block and the semantic sampling matrix of the i-th image block to obtain the i-th initial semantic sampling result;

[0135] The sampling rate alignment submodule is used to perform block initialization sampling rate alignment based on the i-th initial semantic sampling result and the semantic sampling matrix of the transposed i-th image block to obtain the final semantic sampling result.

[0136] The semantic scanning network in the semantic scanning submodule consists of a first convolutional layer, a first channel environment adaptive module, multiple simple residual blocks, a second channel environment adaptive module, and a convolutional layer cascaded sequentially. The semantic scanning network is used to evaluate the saliency of semantic information at each location on the image to be transmitted.

[0137] The aggregation correction submodule includes:

[0138] A partitioning unit is used to divide the semantic information distribution map into units of size [size missing] in a left-to-right and top-to-bottom order. B × B of l Image blocks;

[0139] Block aggregation unit, used for l The image blocks are aggregated to obtain the semantic sampling ratio allocation map, which contains the block sampling rate of all image block allocations. 。

[0140] The semantic reconstruction module includes:

[0141] The expansion submodule is used to expand the image feature vector that needs semantic reconstruction into... l For each image block, a random transformation enhancement is performed, followed by block gradient descent to obtain the intermediate results for each image block. The intermediate results for all image blocks are then aggregated.

[0142] The filling extension submodule is used to receive the semantic sampling ratio allocation map sent by the sending end, and use the corresponding sampling ratio in the semantic sampling ratio allocation map as filling elements to obtain the extended semantic sampling ratio allocation map;

[0143] The feature embedding submodule is used to embed the features of the extended semantic sampling ratio allocation map into the feature space through a semantic extraction network oriented towards semantic sampling ratio allocation map and signal-to-noise ratio, and output a feature map extracted based on semantics and signal-to-noise ratio.

[0144] The connection submodule is used to connect the intermediate results of each image patch with the feature map extracted based on semantics and signal-to-noise ratio, and to process the connection results based on a deep learning near-end mapping network.

[0145] The reconstruction submodule is used to add the processing result to the intermediate result of each image block to obtain the reconstructed image data, and then perform inverse normalization to obtain the reconstructed image.

[0146] The channel environment adaptive module in the channel environment adaptive sampling-semantic-channel coding joint optimization system includes a sequentially connected average pooling layer, a first fully connected layer, a PReLU layer, a second fully connected layer, and a sigmoid layer. The channel environment adaptive module reprocesses the input features based on a known signal-to-noise ratio. The channel environment adaptive module includes:

[0147] The average pooling submodule is used to perform an average pooling operation on the intermediate features output by the previous module connected to the channel environment adaptive module to obtain the average pooling result.

[0148] The splicing submodule is used to splice the average pooling result with the signal-to-noise ratio to obtain the spliced ​​feature;

[0149] The feature prediction scaling submodule is used to process the spliced ​​features through a first fully connected layer, a PReLU layer, a second fully connected layer, and a sigmoid layer to obtain an adaptive scaling factor for the channel environment.

[0150] The feature aggregation submodule is used to obtain intermediate features that are adaptive to the signal-to-noise ratio of the current channel environment based on the channel environment adaptive scaling factor and the intermediate features output by the previous module connected to the channel environment adaptive module.

[0151] The system further includes a training module, which comprises:

[0152] The loss calculation submodule is used to calculate the loss between the reconstructed image and the corresponding image in the training image set;

[0153] The optimization submodule is used to optimize the parameters in the semantic sampling module, encoder, decoder and semantic reconstruction module based on the loss, so as to obtain the channel environment adaptive sampling-semantic-channel coding joint model.

[0154] Based on the same concept, this application discloses an electronic device in a third aspect. Figure 8 A schematic diagram of an electronic device disclosed in an embodiment of this application is shown, such as... Figure 8As shown, the electronic device 100 includes a memory 110 and a processor 120. The memory of the electronic device is not less than 12G, and the main frequency of the processor is not less than 2.4GHz. The memory 110 and the processor 120 are connected by a bus communication. The memory 110 stores a computer program, which can run on the processor 120 to implement a channel environment adaptive sampling-semantic-channel coding joint optimization method disclosed in the embodiments of this application.

[0155] Based on the same concept, this application discloses a computer-readable storage medium storing a computer program / instruction thereon in a fourth aspect. When the computer program / instruction is executed by a processor, it implements a channel environment adaptive sampling-semantic-channel coding joint optimization method disclosed in this application.

[0156] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The same or similar parts between the various embodiments can be referred to each other.

[0157] This application describes embodiments with reference to flowchart illustrations and / or block diagrams of methods, apparatuses, electronic devices, and computer program products according to embodiments of this application. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing terminal device to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing terminal device, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0158] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing terminal device to operate in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0159] These computer program instructions can also be loaded onto a computer or other programmable data processing terminal equipment, causing a series of operational steps to be performed on the computer or other programmable terminal equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable terminal equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0160] Although preferred embodiments of the present application have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments as well as all changes and modifications falling within the scope of the embodiments of the present application.

[0161] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or terminal device that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or terminal device. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or terminal device that includes said element.

[0162] The above provides a detailed description of the channel environment adaptive sampling-semantic-channel coding joint optimization method and system provided in this application. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the method and its core ideas. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.

Claims

1. A channel environment adaptive sampling-semantic-channel coding joint optimization method, characterized in that, The method includes: The sending end performs semantic sampling on the image to be transmitted to obtain the semantic sampling result; The transmitting end inputs the semantic sampling result into the encoder to obtain the image feature vector to be transmitted. The encoder is composed of multiple residual convolution modules and multiple channel environment adaptive modules cascaded alternately. The input of each channel environment adaptive module in the encoder is the signal-to-noise ratio and the intermediate feature output by the previous residual convolution module, and the output is the intermediate feature that is adapted to the signal-to-noise ratio of the current channel environment. The receiving end receives image feature vectors affected by physical channel noise and performs preprocessing. The receiving end inputs the preprocessing result into the decoder to obtain the image feature vector that needs semantic reconstruction. The decoder is composed of multiple residual transposed convolutional modules and multiple channel environment adaptive modules cascaded alternately. The input of each channel environment adaptive module in the decoder is the signal-to-noise ratio and the intermediate feature output by the previous residual transposed convolutional module, and the output is the intermediate feature that is adapted to the signal-to-noise ratio of the current channel environment. The receiving end performs semantic reconstruction based on the image feature vector that needs semantic reconstruction to obtain the reconstructed image; The sending end performs semantic sampling on the image to be transmitted, obtaining the semantic sampling results, including: After normalizing the image to be transmitted, the sending end divides it into blocks to obtain multiple image blocks; The transmitting end inputs the image to be transmitted and the signal-to-noise ratio into the semantic scanning network to obtain a semantic information distribution map; The transmitting end uses a block sampling ratio aggregation correction algorithm to obtain a semantic sampling ratio allocation map based on the semantic information distribution map and the signal-to-noise ratio; The sending end performs singular value decomposition on the training image set to obtain the initial matrix; For the i-th image block among multiple image blocks, the semantic sampling ratio of the i-th image block is obtained based on the given overall sampling rate and the semantic information distribution map; Based on the semantic sampling ratio of the i-th image block and the initial matrix, the semantic sampling matrix of the i-th image block is obtained; Based on the i-th image patch and its semantic sampling matrix, semantic sampling is performed to obtain the i-th initial semantic sampling result; Based on the i-th initial semantic sampling result and the semantic sampling matrix of the transposed i-th image block, block initialization sampling rate alignment is performed to obtain the final semantic sampling result.

2. The channel environment adaptive sampling-semantic-channel coding joint optimization method according to claim 1, characterized in that, The semantic scanning network consists of a first convolutional layer, a first channel environment adaptive module, multiple simple residual blocks, a second channel environment adaptive module, and convolutional layers cascaded sequentially; the semantic scanning network is used to evaluate the saliency of semantic information at each location on the image to be transmitted.

3. The channel environment adaptive sampling-semantic-channel coding joint optimization method according to claim 1, characterized in that, The transmitting end employs a block sampling ratio aggregation correction algorithm to obtain a semantic sampling ratio allocation map based on the semantic information distribution map and the signal-to-noise ratio, including: The semantic information distribution map is divided into l image blocks of size B×B in order from left to right and from top to bottom; Block aggregation is performed on l image blocks to obtain the semantic sampling ratio allocation map, which includes the block sampling rate of all image block allocations.

4. The channel environment adaptive sampling-semantic-channel coding joint optimization method according to claim 1, characterized in that, The receiving end performs semantic reconstruction based on the image feature vector requiring semantic reconstruction to obtain a reconstructed image, including: The receiving end expands the image feature vector that needs semantic reconstruction into l image blocks, performs random transformation enhancement on each image block, and then performs block gradient descent to obtain the intermediate result of each image block, and summarizes the intermediate results of all image blocks. The semantic sampling ratio allocation map sent by the receiving end is used as the corresponding sampling ratio in the semantic sampling ratio allocation map as the filling element to obtain the extended semantic sampling ratio allocation map; The features of the extended semantic sampling ratio allocation map are embedded into the feature space through a semantic extraction network oriented towards semantic sampling ratio allocation map and signal-to-noise ratio, and a feature map extracted based on semantics and signal-to-noise ratio is output. The intermediate results of each image patch are concatenated with the feature maps extracted based on semantics and signal-to-noise ratio, and the concatenation results are processed based on a deep learning near-end mapping network. The processing result is added to the intermediate result of each image block to obtain the reconstructed image data, and then inverse normalization is performed to obtain the reconstructed image.

5. The channel environment adaptive sampling-semantic-channel coding joint optimization method according to claim 1, characterized in that, The channel environment adaptive module includes sequentially connected average pooling layer, first fully connected layer, PReLU layer, second fully connected layer, and sigmoid layer; the channel environment adaptive module reprocesses the input features based on the known signal-to-noise ratio, and the specific process is as follows: The intermediate features output by the previous module connected to the channel environment adaptive module are subjected to average pooling to obtain the average pooling result. The average pooling result is concatenated with the signal-to-noise ratio to obtain the concatenated feature; The splicing features are processed by a first fully connected layer, a PReLU layer, a second fully connected layer, and a sigmoid layer to obtain an adaptive scaling factor for the channel environment. Based on the channel environment adaptive scaling factor and the intermediate features output by the previous module connected to the channel environment adaptive module, intermediate features that adapt to the signal-to-noise ratio of the current channel environment are obtained.

6. The channel environment adaptive sampling-semantic-channel coding joint optimization method according to claim 1, characterized in that, The channel environment-adaptive sampling-semantic-channel coding joint optimization method is implemented through a channel environment-adaptive sampling-semantic-channel coding joint model. The training process includes: Calculate the loss between the reconstructed image and the corresponding image in the training image set; Based on the loss, the parameters in the semantic sampling module, encoder, decoder, and semantic reconstruction module are optimized to obtain the channel environment-adaptive sampling-semantic-channel coding joint model.

7. A channel environment adaptive sampling-semantic-channel coding joint optimization system, characterized in that, The system includes: The semantic sampling module is used to perform semantic sampling on the image to be transmitted and obtain the semantic sampling results. The encoding module is used to input the semantic sampling result into the encoder to obtain the image feature vector to be transmitted. The encoder is composed of multiple residual convolution modules and multiple channel environment adaptive modules cascaded alternately. The input of each channel environment adaptive module in the encoder is the signal-to-noise ratio and the intermediate feature output by the previous residual convolution module, and the output is the intermediate feature that is adaptive to the signal-to-noise ratio of the current channel environment. The preprocessing module is used to receive image feature vectors affected by physical channel noise and perform preprocessing. The decoding module is used to input the preprocessing results into the decoder to obtain the image feature vector that needs to be semantically reconstructed. The decoder is composed of multiple residual transposed convolutional modules and multiple channel environment adaptive modules cascaded alternately. The input of each channel environment adaptive module in the decoder is the signal-to-noise ratio and the intermediate features output by the previous residual transposed convolutional module, and the output is the intermediate features that are adapted to the signal-to-noise ratio of the current channel environment. The semantic reconstruction module is used to perform semantic reconstruction based on the image feature vectors that need to be semantically reconstructed, and obtain the reconstructed image. The semantic sampling module includes: The segmentation submodule is used to normalize the image to be transmitted and then segment it into multiple image blocks. The semantic scanning submodule is used to input the image to be transmitted and the signal-to-noise ratio into the semantic scanning network to obtain a semantic information distribution map. The aggregation correction submodule is used to obtain a semantic sampling ratio allocation map based on the semantic information distribution map and the signal-to-noise ratio using a block sampling ratio aggregation correction algorithm. The decomposition submodule is used to perform singular value decomposition on the training image set to obtain the initial matrix; The semantic sampling ratio acquisition submodule is used to obtain the semantic sampling ratio of the i-th image block among multiple image blocks, based on a given overall sampling rate and the semantic information distribution map; The semantic sampling matrix acquisition submodule is used to obtain the semantic sampling matrix of the i-th image block based on the semantic sampling ratio of the i-th image block and the initial matrix; The semantic sampling result acquisition submodule is used to perform semantic sampling based on the i-th image block and the semantic sampling matrix of the i-th image block to obtain the i-th initial semantic sampling result; The sampling rate alignment submodule is used to perform block initialization sampling rate alignment based on the i-th initial semantic sampling result and the semantic sampling matrix of the transposed i-th image block to obtain the final semantic sampling result.

8. An electronic device comprising a memory, a processor, and a computer program stored in the memory, characterized in that, The processor executes the computer program to implement the channel environment adaptive sampling-semantic-channel coding joint optimization method according to any one of claims 1-6.

9. A computer-readable storage medium, characterized in that, It stores a computer program / instruction that, when executed by a processor, implements the channel environment adaptive sampling-semantic-channel coding joint optimization method as described in any one of claims 1-6.

Citation Information

Patent Citations

  • Image restoration method based on global texture and structure

    CN115035170A

  • Image encoding and decoding, video encoding and decoding: methods, systems, and training methods

    CN116584098A