Video coding and decoding transmission system construction method and system, equipment and medium
The proposed video encoding and decoding system, utilizing OFDM and noise reduction techniques, addresses the limitations of DeepJSCC in multiple-path fading environments, enhancing video transmission quality and stability.
Patent Information
- Application Number
- CN202510787968.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-13
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2045-06-13
AI Technical Summary
The existing deep joint source channel encoding (DeepJSCC) has degraded performance in actual wireless multipath fading environments and is unable to effectively combat frequency selective fading, resulting in a sharp drop in video quality or inability to decode.
The OFDM video transmission system is built, combining video codecs and denoising networks, and video frames are processed through multi-scale conditional context features and optical flow networks, and the multi-path fading is used to combat multi-path fading, and channel noise filtering is performed on the receiving end to simplify the decoding task.
High-quality video transmission and reconstruction under complex channel conditions are realized, improving the robustness and transmission performance of video data, and reducing the complexity of decoding.
Smart Images

Figure CN120321401A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of information and communication technologies, and particularly to a method and system, device, and medium for constructing a video codec transmission system based on deep joint source-channel coding for wireless video transmission in a multipath fading channel environment. Background Art
[0002] With the rapid development of the sixth-generation mobile communication (6G), new intelligent video applications such as ultra-high-definition video communication, virtual reality (VR), video conferencing, and telemedicine have put forward higher requirements for the bandwidth, latency, and reliability of wireless communication systems. Traditional video transmission methods generally adopt a separately designed source coding and channel coding (SSCC) scheme to separately complete video data compression and error control in the communication channel.
[0003] However, in scenarios with limited bandwidth, complex and dynamically changing channel conditions, this traditional scheme has obvious performance bottlenecks and limitations: once the channel condition deteriorates beyond the error correction ability, the video quality will drop sharply or even be completely undecodable, that is, the "cliff effect". In recent years, new communication paradigms represented by semantic information transmission have gradually attracted attention. Among them, deep joint source-channel coding (DeepJSCC) realizes end-to-end video compression and channel adaptation through a deep neural network, showing better robustness and transmission performance in complex channel environments.
[0004] The inventors found that most existing DeepJSCC research is limited to idealized simple channel environments (such as additive white Gaussian noise channels), and its performance will drop significantly under the conditions of multipath propagation and frequency-selective fading widely existing in the actual wireless communication environment.
[0005] Based on the above background situation, the inventors realized that there is an urgent need for a DeepJSCC system for video transmission facing the actual wireless multipath fading environment, which can effectively combat frequency-selective fading to meet the strict requirements of future wireless video communication. Summary of the Invention
[0006] To solve the problems, this application proposes a method for constructing a video codec transmission system, an application system, a receiver, a device, and a medium. The technical solutions are as follows: On the one hand, a method for constructing a video codec transmission system is provided. The video codec transmission system includes a video codec. The method for constructing the video codec transmission system includes the steps of: Construct an OFDM video transmission system; Construct a video codec, which includes a video encoder and a video decoder. The video encoder is arranged at the sending end of the OFDM video transmission system, and the video decoder is arranged at the receiving end of the OFDM video transmission system; The video codec is trained via an OFDM video transmission system using the original training data set until convergence.
[0007] Among them, the OFDM video transmission system includes an OFDM transmitter, a wireless multipath fading channel, and an OFDM receiver; Among them, the OFDM transmitter converts the compressed video output by the video encoder into frequency-domain symbols, applies the inverse discrete Fourier transform, and adds a cyclic prefix to obtain the OFDM transmission signal. The compressed video is obtained after the video encoder processes the original training data set; The OFDM transmission signal is transmitted through the wireless multipath fading channel and arrives at the OFDM receiver; The OFDM receiver removes the cyclic prefix of the received signal and applies the discrete Fourier transform to finally obtain the received frequency-domain symbols; the received frequency-domain symbols are used to train the video decoder.
[0008] Among them, the frame sequence of the video is divided into key frames and interpolation frames; The constructed video encoder includes a key-frame encoder and an interpolation-frame encoder, and the video decoder includes a key-frame decoder and an interpolation-frame decoder; Among them, the key-frame encoder is used to perform compression encoding on the images of the key frames, and the key-frame decoder is used to decode and reconstruct the images of the key frames; The interpolation-frame encoder uses the multi-scale conditional context features of the reference frames to perform conditional compression encoding on the images of the interpolation frames, and the interpolation-frame decoder is used to reconstruct the images of the interpolation frames.
[0009] Among them, the reference frames include the two adjacent front and back frames of the interpolation frame, and the multi-scale conditional context features are obtained through the following methods: Construct a feature extraction network, use the feature extraction network to map the reference frames into high-dimensional features, and further perform Gaussian blur through a set of Gaussian smoothing kernels with different scales to generate the scale space volume in the feature domain; Construct an optical flow network, use the optical flow network to estimate the scale space flow between the interpolation frame and the reference frames to obtain the motion information between the interpolation frame and the reference frames; Perform a feature space warping operation on the scale space flow and the scale space volume to obtain the multi-scale conditional context features.
[0010] Among them, the video codec transmission system further includes a denoising network, and the method for constructing the video codec transmission system further includes the steps of: Construct a denoising network. The denoising network is set at the receiving end of the OFDM video transmission system and is used to receive the noisy pilot symbols and data symbols obtained after OFDM demodulation through the multipath fading channel; The denoising network and the video codec are jointly trained using the original training dataset via an OFDM video transmission system until convergence.
[0011] Among them, the denoising network receives the OFDM demodulated signal after passing through a multipath fading channel, including the received frequency-domain symbols and received pilot symbols after channel fading and noise interference , as well as the original pilot symbols , and uses the channel state information implicit in the pilot symbols for adaptive feature weighting to efficiently filter out the noise and interference introduced by the channel from the received noisy data symbols, so as to finally obtain a clean compressed representation , that is: ; The output symbols processed by the denoising network form a clean compressed representation, which reduces the learning complexity of the subsequent decoder and improves the reconstruction quality of video frames. The key frame and interpolation frame decoders no longer need to receive pilot inputs to implicitly complete channel estimation and equalization, and the decoding task is simplified, that is: , .
[0012] Among them, the step of jointly training the denoising network and the video codec using the original training dataset via an OFDM video transmission system until convergence includes: First, fix the interpolation frame codec, and separately train the key frame codec and the denoising network. Using the OFDM system model, under the simulated multipath fading channel conditions, optimize the peak signal-to-noise ratio and multi-scale structural similarity of key frame reconstruction; When the key frame codec training converges, then introduce the relevant networks of the interpolation frame codec, including the scale space flow estimation network, the conditional context feature extraction network, and the interpolation frame decoding network, and perform end-to-end training of the interpolation frame path; The training objectives include the denoising loss of the compressed representation and the video frame reconstruction loss. The loss function is constructed using the weighted mean square error, and the final model convergence is achieved through joint optimization. For the training loss of the video frame sequence , that is , Among them, is the weighting coefficient of the denoising loss.
[0013] On the other hand, a video codec transmission system constructed by any of the above methods is provided. The video codec transmission system is used to achieve accurate encoding and decoding of video transmission for a wireless multipath fading environment.
[0014] In another aspect, there is provided an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of any one of the above methods are implemented.
[0015] In another aspect, there is provided a readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the method for constructing a video codec transmission system of any one of the above is implemented.
[0016] The beneficial effects of this application are as follows: A method and system, device, and medium for constructing a video codec transmission system are proposed. The video codec transmission system includes a video codec. The method for constructing the video codec transmission system includes the steps of: constructing an OFDM video transmission system; constructing a video codec, which includes a video encoder and a video decoder. The video encoder is arranged at the sending end of the OFDM video transmission system, and the video decoder is arranged at the receiving end of the OFDM video transmission system; training the video codec via the OFDM video transmission system using the original training data set until convergence. By introducing OFDM modulation in the training system, the frequency-selective fading in the multipath fading channel is effectively countered, and high-quality transmission and reconstruction of video data under complex channel conditions are achieved. Description of the Drawings
[0017] Figure 1 It is a schematic flowchart of an embodiment of the method for constructing a video codec transmission system of this application; Figure 2 It is a schematic flowchart of another embodiment of the method for constructing a video codec transmission system of this application; Figure 3 It is a schematic architecture diagram of an embodiment of a video transmission system based on OFDM technology of this application; Figure 4 It is a schematic flowchart of an embodiment of interpolation frame coding based on a conditional context coding mechanism of this application; Figure 5 It is a schematic structural diagram of an embodiment of a lightweight denoising network of this application; Figure 6 It is a schematic diagram of an embodiment of the influence comparison of the convergence performance by applying a lightweight denoising module of this application; Figure 7 It is a schematic diagram of an embodiment of the performance comparison of video reconstruction PSNR by applying different technologies of this application; Figure 8 It is a schematic diagram of an embodiment of the performance comparison of video reconstruction MS-SSIM by applying different technologies of this application; Figure 9 It is a schematic framework diagram of an embodiment of an electronic device provided by this application; Figure 10It is a framework schematic diagram of an embodiment of the computer-readable storage medium provided by this application. Detailed implementation manners
[0018] To facilitate the understanding of this application, the following will combine the accompanying drawings and specific embodiments to describe this application in more detail. The accompanying drawings show preferred embodiments of this application. However, this application can be implemented in many different forms and is not limited to the embodiments described in this specification. On the contrary, the purpose of providing these embodiments is to make the understanding of the disclosed content of this application more thorough and comprehensive.
[0019] It should be noted that unless otherwise defined, all technical and scientific terms used in this specification have the same meaning as commonly understood by those skilled in the technical field to which this application belongs. The terms used in the specification of this application are only for the purpose of describing specific embodiments and are not used to limit this application. For example, the term "plurality" includes two or more.
[0020] The following are some term explanations of this application: OFDM: Orthogonal Frequency Division Multiplexing; 6G: The sixth generation of mobile communication; VR: Virtual Reality; SSCC: Separate Source Channel Coding, source-channel separation coding; DeepJSCC: Deep Joint Source-Channel Coding.
[0021] Traditional video transmission methods generally adopt a separately designed source coding and channel coding (SSCC) scheme to separately complete video data compression and error control in the communication channel.
[0022] SSCC is the core design paradigm of classical communication systems. Its characteristic is to optimize the source coding (removing data redundancy) and channel coding (anti-interference error correction) as separate modules respectively; although SSCC has the advantages of simple design and modularity, its performance is limited by the assumptions in Shannon's separation theorem (such as infinite code length, ideal channel conditions, etc.). In practical applications, the separation of modules will lead to problems such as the "cliff effect" (channel decoding errors causing source decoding to collapse) and low spectral efficiency.
[0023] However, in scenarios with limited bandwidth, complex and dynamically changing channel conditions, this traditional solution has obvious performance bottlenecks and limitations: once the channel condition deteriorates beyond the error correction ability, the video quality will drop sharply or even be completely undecodable, i.e., the "cliff effect". In recent years, new communication paradigms represented by semantic information transmission have gradually attracted attention. Among them, Deep Joint Source-Channel Coding (DeepJSCC) realizes end-to-end video compression and channel adaptation through deep neural networks, showing better robustness and transmission performance in complex channel environments.
[0024] The inventors found that most of the existing DeepJSCC research is limited to idealized simple channel environments (such as additive white Gaussian noise channels), and its performance will drop significantly under the conditions of multipath propagation and frequency selective fading that widely exist in actual wireless communication environments.
[0025] Based on the above background, this application innovatively proposes a method and system, device, and medium for constructing a video codec transmission system. The video codec transmission system includes a video codec. For details, please refer to Figure 1 , the method for constructing a video codec transmission system includes the steps of: constructing an OFDM video transmission system; constructing a video codec, which includes a video encoder and a video decoder. The video encoder is set at the sending end of the OFDM video transmission system, and the video decoder is set at the receiving end of the OFDM video transmission system; training the video codec via the OFDM video transmission system using the original training data set until convergence. By introducing OFDM modulation in the training system, the frequency selective fading in the multipath fading channel is effectively countered, realizing high-quality transmission and reconstruction of video data under complex channel conditions.
[0026] In addition, due to the large inter-frame redundancy in video data, how to effectively model and compress this redundancy is also an important factor affecting video transmission efficiency.
[0027] Therefore, further, please refer to Figure 3 、 4 As shown in
[0028] Meanwhile, in the traditional DeepJSCC scheme, a single decoder needs to handle multiple tasks such as channel estimation, signal recovery, and semantic reconstruction simultaneously. The decoding chain is too long and complex, resulting in difficult effective training and convergence of the model, which limits the performance of actual deployment.
[0029] Therefore, further, please refer to Figure 2 , 5 As shown, the video codec transmission system in the present application further includes a denoising network. The method for constructing the video codec transmission system further includes the steps of: constructing a denoising network, which is arranged at the receiving end of the OFDM video transmission system and is used to receive the noisy pilot symbols and data symbols obtained after OFDM demodulation through a multipath fading channel; jointly training the denoising network and the video codec via the OFDM video transmission system using the original training dataset until convergence.
[0030] The following specifically elaborates on the embodiments in detail to more fully explain the solution of the present application.
[0031] On the one hand, referring to Figure 1 , 3 , and Figure 4, a method for constructing a video codec transmission system is provided. The video codec transmission system includes a video codec. The method for constructing the video codec transmission system includes the steps of: Constructing an OFDM video transmission system; Constructing a video codec, which includes a video encoder and a video decoder. The video encoder is arranged at the sending end of the OFDM video transmission system, and the video decoder is arranged at the receiving end of the OFDM video transmission system; Training the video codec via the OFDM video transmission system using the original training dataset until convergence.
[0032] Among them, referring to Figure 3 , the OFDM video transmission system includes an OFDM transmitter (the part within the dashed box of the OFDM transmitter in the figure), a wireless multipath fading channel, and an OFDM receiver (the part within the dashed box of the OFDM receiver); Among them, the OFDM transmitter converts the compressed video output by the video encoder into frequency-domain symbols, applies the inverse discrete Fourier transform (IDFT), and adds a cyclic prefix (Add CP) to obtain the OFDM transmission signal. The compressed video is obtained after the video encoder processes the original training dataset; The OFDM transmission signal is transmitted through the wireless multipath fading channel and reaches the OFDM receiver; The OFDM receiver removes the cyclic prefix (CP) of the received signal and applies the discrete Fourier transform (DFT) to finally obtain the received frequency-domain symbols; the received frequency-domain symbols are used to train the video decoder.
[0033] Specifically, please refer to Figure 3 shown in the architecture schematic diagram of the video transmission system based on OFDM technology. Among them, The multipath fading channel is a wireless multipath fading channel, with the transmitting end on its left and the receiving end on its right. The transmitting end includes a key frame encoder (Key Encoder) and an interpolation frame encoder (Interp. Encoder). The key frame encoder independently encodes and compresses the key frames in the video sequence, and the interpolation frame encoder uses the multi-scale conditional context features extracted from the adjacent front and back key frames to assist in conditionally encoding the interpolation frames to effectively compress the inter-frame redundancy. Then, the deeply compressed video data symbols are mapped to multiple OFDM subcarriers, and pilot symbols (Pilot) are inserted for assisting channel estimation. Subsequently, IDFT is performed and CP is added, and the data is transmitted through the wireless multipath fading channel.
[0034] At the receiving end, first the CP is removed, and then the DFT is performed on the received signal to obtain the noisy pilot symbols and the frequency-domain data signal. Finally, these features are reconstructed by the key frame decoder (Key Decoder) and the interpolation frame decoder (Interp. Decoder) respectively, so as to restore high-quality video frames and achieve the robust transmission of video data.
[0035] Among them, the OFDM video transmission system is specifically set as: The video sequence to be encoded and transmitted consists of groups of pictures (GoP), and each GoP consists of frames of video frames of images. Each video frame is an 8-bit RGB three-color channel image; The encoder maps the video frame to the compressed representation , is the length of the compressed representation, that is, , The OFDM transmitter reshapes (Reshape) the compressed representation into the frequency-domain symbols , and applies the inverse discrete Fourier transform (IDFT) to the frequency-domain symbols and the pilot symbols , and adds the cyclic prefix (CP) to obtain the OFDM transmitted signal ; Among them, is the number of OFDM packets, and each OFDM packet consists of data symbols and pilot symbols, and represent the number of subcarriers corresponding to each symbol and the cyclic prefix length respectively; The OFDM transmitted signal is transmitted through a wireless multipath fading channel to reach the receiver, that is , where represents the received signal, * represents the convolution operation, represents the sampled space channel impulse response, is the multi- path number, represents complex-valued additive white Gaussian noise (Gaussian channel) with variance , is the identity matrix; Each multipath channel path experiences independent Rayleigh fading satisfying , is the path index, and the variance of each path follows exponential decay, that is , where represents the delay spread, is the normalization coefficient, and the sum of the variances of all paths is 1, that is ; The OFDM receiver removes the cyclic prefix of the received signal and applies the discrete Fourier transform (DFT) to finally obtain the received frequency-domain symbols and the received pilot symbols , that is , , where represents the channel frequency response of the m-th subcarrier, and both represent noise samples; The decoder reconstructs the video frame using the received signal and the original pilot, that is , The data symbols at the transmitter are subject to an average power constraint, that is , The channel bandwidth constraint of the OFDM system is defined as .
[0036] Among them, the frame sequence of the video is divided into key frames and interpolated frames; The constructed video encoder includes a key frame encoder and an interpolated frame encoder, and the video decoder includes a key frame decoder and an interpolated frame decoder; Among them, the key frame encoder is used to perform compression encoding on the images of key frames, and the key frame decoder is used to decode and reconstruct the images of key frames; The interpolated frame encoder performs conditional compression encoding on the images of interpolated frames using the multi-scale conditional context features of reference frames, and the interpolated frame decoder is used to reconstruct the images of interpolated frames.
[0037] Among them, the reference frames include the two adjacent front and rear frames of the interpolated frame, and the multi-scale conditional context features are obtained through the following method: Construct a feature extraction network, and use the feature extraction network to map the reference frames into high-dimensional features, and further perform Gaussian blur through a set of Gaussian smoothing kernels with different scales to generate the scale space volume in the feature domain; Construct an optical flow network, and use the optical flow network to estimate the scale space flow between the interpolated frame and the reference frames to obtain the motion information between the interpolated frame and the reference frames; Perform a feature space warping operation on the scale space flow and the scale space volume to obtain multi-scale conditional context features.
[0038] Through the above method, more refined multi-scale context features can be obtained, thereby optimizing network training.
[0039] Represent the key frame encoder and the key frame decoder as: key frame codec , Represent the interpolated frame encoder and the interpolated frame decoder as: interpolated frame codec , The th GoP of the video frame sequence to be encoded, that is , where the last frame is regarded as a key frame, which is encoded and processed by the key frame codec, and the remaining video frames are regarded as interpolated frames, is the frame index, and together with two reference frames with an interval of are processed by the interpolated frame codec; Among them, the reference frame with frame index 0 is recorded as the key frame of the th GoP, that is: ; The key frame encoder directly maps the key frames within any GoP to the key frame compressed representation, that is: , wherein, is the estimated channel noise power; The OFDM transmitter processes the compressed representation of the key frame to obtain a transmission signal ; The OFDM receiver processes the received signal to obtain the received frequency-domain symbols and received pilot symbols after the multipath fading channel ; The key frame decoder uses the received frequency-domain symbols, received pilot symbols, and original pilot symbols to reconstruct the key frame, that is: ; The optical flow network estimates the scaled space flow of the interpolation frames and reference frames within any GoP and reference frames to model the motion information between the interpolation frame and the reference frame, that is: ; The feature extractor network maps the reference frame to high-dimensional features and further performs Gaussian blur through a set of Gaussian smoothing kernels with different scales to generate a scale-space volume in the feature domain , that is , wherein, is the feature dimension, represents the convolution of the high-dimensional feature with the Gaussian smoothing kernel of scale , is the number of layers of the scale-space volume; Performs a feature space warping operation using the scale-space flow and the scale-space volume to obtain the feature domain context, that is: , Another reference frame follows the same processing steps to finally obtain the feature domain context ; The interpolation frame encoder uses the feature domain context as a condition to map the interpolation frame to a compressed representation, that is: ; The OFDM transmitter processes the compressed representation of the interpolation frame to obtain a transmission signal The OFDM receiver processes the received signal to obtain the received frequency-domain symbols and received pilot symbols after passing through the multipath fading channel. ; Interpolation frame decoder Estimates the scale-space flow from the received frequency-domain symbols and pilot symbols And the preliminary decoded representation , that is: ; The conditional context decoder uses the feature-domain context generated by the receiver reference frame in the same steps as the transmitter to Generate the conditional reconstruction interpolation frame, that is: .
[0040] The above is the detailed step of the specific context acquisition method. It should be further noted that the reason for introducing the optical flow network in this application is: Estimate the interpolation frames within the GoP And the motion information between the two corresponding reference frames through the Scale-Space Flow network . The motion information is included in the scale-space flow field calculated by the network (including the horizontal displacement field , the vertical displacement field and the scale field in three dimensions):
[0041] Based on the optical flow, the Scale-Space Flow network adds a continuous and differentiable scale parameter to enhance the network's ability to model motion uncertainty. This design overcomes the limitations of traditional optical flow in dealing with occlusion and fast or complex motion scenarios. Compared with ordinary optical flow that only uses a two-dimensional displacement field (i.e., horizontal and vertical displacements , ) for bilinear warping to predict inter-frame motion, the Scale-Space Flow forms a three-dimensional displacement + scale joint field by introducing an additional scale field . This enables the network to adaptively blur the source image to an appropriate extent when the motion prediction is inaccurate or there is a large uncertainty, thereby generating more robust predictions.
[0042] Use the feature extractor The network maps the reference frame to a high-dimensional feature representation , where is the feature dimension. Subsequently, convolution operations are performed using a series of Gaussian kernels with different scales. Feature maps with different blurring scales are stacked to form a continuous space in the scale dimension, forming a Scale-Space Volume in the feature domain.
[0043] This volume contains multiple levels of feature maps with different degrees of blurring and can perform feature sampling through trilinear interpolation. Compared with the two-dimensional interpolation of ordinary optical flow, this three-dimensional interpolation method can better capture and process the changes and uncertainties in image content caused by motion.
[0044] Scale-space flow field (including the horizontal displacement field , the vertical displacement field and the scale field in three dimensions) and the scale-space Gaussian blur volume perform a joint Feature-Space Warping (FSW) operation to finally obtain the context in the feature domain, that is,
[0045] Another reference frame follows the same processing steps to finally obtain the context in the feature domain ; Specifically, when the scale field is close to 0, the operation degenerates into bilinear interpolation of ordinary optical flow; when , is 0 and is large, the operation degenerates into pure Gaussian blur; in the intermediate state, the warping operation is the combined effect of image space displacement and scale space blur, which can effectively balance the relationship between displacement prediction error and image detail retention.
[0046] Interpolation frame encoder uses the context in the feature domain as a condition to map the interpolation frame to a compressed representation, that is, ; It can be seen that the feature domain context is used as a condition instead of the traditional pixel domain residual calculation. The traditional pixel domain residual calculation only removes the inter-frame redundancy through simple subtraction operations, which limits the effective utilization of inter-frame correlation and results in suboptimal compression performance. In contrast, the feature domain context contains higher-dimensional and richer information, which can help the encoder and decoder better reconstruct high-frequency detailed content and improve the compression quality and efficiency of the interpolated frames. Therefore, the interpolated frame encoder uses the feature domain context as a condition to map the interpolated frames into compressed representations. The introduction of this conditional context enables the network to adaptively process regions with inaccurate motion or large uncertainties, improving the robustness of overall video compression.
[0047] Among them, referring to Figure 3 , 5 as described, the video codec transmission system further includes a denoising network ( Figure 3 the red square in the receiver in ). The method for constructing the video codec transmission system further includes the steps of: Constructing a denoising network, which is set at the receiver of the OFDM video transmission system and is used to receive the noisy pilot symbols and data symbols obtained after OFDM demodulation through the multipath fading channel;
[0048] Among them, the denoising network receives the OFDM demodulation signals after passing through the multipath fading channel, including the received frequency-domain symbols and received pilot symbols after channel fading and noise interference , as well as the original pilot symbols , and uses the channel state information implicit in the pilot symbols for adaptive feature weighting to efficiently filter out the noise and interference introduced by the channel from the received noisy data symbols, so as to finally obtain a clean compressed representation , that is: ; Referring to Figure 5 as shown, it is a schematic diagram of the lightweight denoising network structure. The output symbols after being processed by the denoising network form a clean compressed representation to reduce the learning complexity of the subsequent decoder and improve the reconstruction quality of video frames. The key frame and interpolated frame decoders no longer need to receive pilot inputs to implicitly complete channel estimation and equalization, and the decoding task is simplified, that is: , .
[0049] Among them, the step of jointly training the denoising network and the video codec using the original training dataset via the OFDM video transmission system until convergence includes: First, fix the interpolation frame codec, and train the key frame codec and the denoising network separately. Using the OFDM system model, optimize the peak signal-to-noise ratio (PSNR) and multi-scale structural similarity (MS-SSIM) of key frame reconstruction under simulated multipath fading channel conditions; After the training of the key frame codec converges, then introduce the relevant networks of the interpolation frame codec, including the scale space flow estimation network, the conditional context feature extraction network, and the interpolation frame decoding network, and conduct end-to-end training of the interpolation frame path; The training objectives include the denoising loss of the compressed representation and the video frame reconstruction loss. Use the weighted mean square error (MSE) to construct the loss function, and achieve the convergence of the final model through joint optimization. For the video frame sequence The training loss of, that is , where, is the weighted coefficient of the denoising loss.
[0050] Table 1 Communication system parameter settings
[0051] Reference Figure 6 As shown, it is a comparison chart of the training convergence performance between this application and the comparative method. Among them, "DeepWiVe + fading channel" is the training loss curve of the traditional DeepJSCC video transmission method under the multipath fading channel. It can be seen that its convergence speed is slow and the convergence error is large, indicating the lack of training stability of this method under complex channel conditions. "OFDM + context + fading channel" is the training loss curve of only using the OFDM modulation and conditional context coding mechanism proposed in this application. Compared with the traditional method, it has a faster convergence speed and a lower final convergence error. The curve after further adding the lightweight denoising network designed in this application is marked as "proposed method + fading channel". It can be seen that the convergence speed of the training curve is further accelerated and the error is significantly reduced, indicating that the complete system proposed in this application can achieve more stable and rapid convergence.
[0052] Figure 7Shows the performance comparison of video reconstruction quality (PSNR) of different methods under different signal-to-noise ratio (SNR) conditions, with an emphasis on image reconstruction quality. Among them, "DeepWiVe + Gaussian channel" is the performance curve of the traditional DeepJSCC method under an ideal additive white Gaussian noise (Gaussian channel) channel, providing a reference upper limit under ideal channel conditions, with an average PSNR reaching 29.68 dB. "DeepWiVe + fading channel" is the performance of the same method in a multipath fading channel, and the average PSNR significantly drops to 22.03 dB, indicating insufficient adaptability to complex channels. After introducing the OFDM modulation technology, "OFDM + fading channel" significantly improves the PSNR to 24.79 dB, indicating that OFDM effectively mitigates the impact of frequency-selective fading. "OFDM + context + fading channel" further integrates the conditional context coding mechanism on this basis, and the PSNR performance further improves to 26.59 dB, showing that the conditional context coding mechanism effectively compresses the semantic redundancy between video frames. Especially under high SNR conditions, the gain brought by the conditional context coding mechanism is more significant because this mechanism extracts finer-grained inter-frame redundancy features and presents richer semantic content at the same compression rate. Therefore, under better channel conditions, these semantic details significantly improve the reconstruction quality. "Proposed method + fading channel" is the complete system solution proposed in this application, which combines the advantages of OFDM, conditional context coding, and denoising network, achieving the highest average PSNR performance of 27.16 dB, and the overall performance is better than all comparison methods.
[0053] Figure 8 Further shows the performance comparison trend of different methods under the multi-scale structural similarity (MS-SSIM) metric, with an emphasis on pixel reconstruction quality, where the curve correspondence is the same as Figure 7 Remains consistent. The results show that the method "Proposed method + fading channel" proposed in this application is also significantly superior to other comparison methods in terms of the subjective perceptual quality of videos in a multipath fading channel, showing stronger anti-noise robustness and video detail fidelity capabilities.
[0054] In summary, the robust DeepJSCC video transmission scheme proposed in this application, by integrating OFDM technology, multi-scale conditional context coding, and lightweight denoising network, shows significant advantages in multiple metrics such as training convergence speed, video reconstruction quality (PSNR), and subjective visual quality (MS-SSIM), and is particularly suitable for efficient video transmission applications in complex multipath fading channel environments.
[0055] On the other hand, a video codec transmission system constructed by any of the above methods is provided, and the video codec transmission system is used to achieve accurate codec for video transmission in a wireless multipath fading environment.
[0056] In another aspect, an electronic device is provided, which includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of any of the above methods are implemented.
[0057] Specifically, please refer to Figure 9 , the electronic device described in the embodiments of the present application may specifically include a processor 210 and a memory 220. The memory 220 is coupled to the processor 210.
[0058] The processor 210 is used to control the operation of the electronic device. The processor 210 may also be referred to as a CPU (Central Processing Unit, central processing unit). The processor 210 may be an integrated circuit chip with signal processing capabilities. The processor 210 may also be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The general-purpose processor may be a microprocessor, or the processor 210 may also be any conventional processor, etc.
[0059] The memory 220 is used to store computer programs, which may be RAM, ROM, or other types of storage devices. Specifically, the memory may include one or more computer-readable storage media, and the computer-readable storage media may be non-transitory. The memory may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices and flash storage devices. In some embodiments, the non-transitory computer-readable storage medium in the memory is used to store at least one program code.
[0060] The processor 210 is used to execute the computer program stored in the memory 220 to implement the methods described in the method embodiments of the present application.
[0061] In some embodiments, the electronic device may further include: a peripheral device interface 230 and at least one peripheral device. The processor 210, the memory 220, and the peripheral device interface 230 may be connected through a bus or signal lines. Each peripheral device may be connected to the peripheral device interface 230 through a bus, signal lines, or a circuit board. Specifically, the peripheral device includes at least one of a radio frequency circuit 240, a display screen 250, an audio circuit 260, and a power supply 270.
[0062] The peripheral device interface 230 can be used to connect at least one I / O (Input / Output) related peripheral device to the processor 210 and the memory 220. In some embodiments, the processor 210, the memory 220, and the peripheral device interface 230 are integrated on the same chip or circuit board; in some other embodiments, any one or two of the processor 210, the memory 220, and the peripheral device interface 230 can be implemented on a separate chip or circuit board, and this embodiment does not limit this.
[0063] The radio frequency circuit 240 is used to receive and transmit RF (Radio Frequency) signals, also known as electromagnetic signals. The radio frequency circuit 240 communicates with a communication network and other communication devices through electromagnetic signals, and the radio frequency circuit 240 is the communication circuit of the electronic device. The radio frequency circuit 240 converts an electrical signal into an electromagnetic signal for transmission, or converts the received electromagnetic signal into an electrical signal. Optionally, the radio frequency circuit 240 includes: an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, a codec chipset, a subscriber identity module card, and so on. The radio frequency circuit 240 can communicate with other terminals through at least one wireless communication protocol. The wireless communication protocol includes but is not limited to: the World Wide Web, a metropolitan area network, an intranet, each generation of mobile communication networks (2G, 3G, 4G, and 5G), a wireless local area network, and / or a WiFi (Wireless Fidelity) network. In some embodiments, the radio frequency circuit 240 may further include a circuit related to NFC (Near Field Communication), and this application does not limit this.
[0064] The display screen 250 is used to display the UI (User Interface). The UI may include graphics, text, icons, videos, and any combination thereof. When the display screen 250 is a touch display screen, the display screen 250 also has the ability to collect touch signals on or above the surface of the display screen 250. The touch signals can be input to the processor 210 for processing as control signals. At this time, the display screen 250 can also be used to provide virtual buttons and / or a virtual keyboard, also known as soft buttons and / or a soft keyboard. In some embodiments, there can be one display screen 250, which is disposed on the front panel of the electronic device; in other embodiments, there can be at least two display screens 250, which are respectively disposed on different surfaces of the electronic device or are in a folded design; in other embodiments, the display screen 250 can be a flexible display screen, which is disposed on a curved surface or a folding surface of the electronic device. Even further, the display screen 250 can be set to an irregular non-rectangular shape, that is, a special-shaped screen. The display screen 250 can be prepared using materials such as LCD (Liquid Crystal Display) and OLED (Organic Light-Emitting Diode).
[0065] The audio circuit 260 may include a microphone and a speaker. The microphone is used to collect sound waves of the user and the environment, and convert the sound waves into electrical signals and input them to the processor 210 for processing, or input them to the radio frequency circuit 240 to achieve voice communication. For the purpose of stereo collection or noise reduction, there can be multiple microphones, which are respectively disposed at different parts of the electronic device. The microphone can also be an array microphone or an omnidirectional collection microphone. The speaker is used to convert the electrical signals from the processor 210 or the radio frequency circuit 240 into sound waves. The speaker can be a traditional thin film speaker or a piezoelectric ceramic speaker. When the speaker is a piezoelectric ceramic speaker, it can not only convert electrical signals into sound waves audible to humans, but also convert electrical signals into sound waves inaudible to humans for uses such as ranging. In some embodiments, the audio circuit 260 can also include a headphone jack.
[0066] The power supply 270 is used to supply power to each component in the electronic device. The power supply 270 can be alternating current, direct current, a disposable battery, or a rechargeable battery. When the power supply 270 includes a rechargeable battery, the rechargeable battery can be a wired rechargeable battery or a wireless rechargeable battery. A wired rechargeable battery is a battery charged through a wired line, and a wireless rechargeable battery is a battery charged through a wireless coil. The rechargeable battery can also be used to support fast charging technology.
[0067] For a detailed description of the functions and execution processes of each functional module or component in the embodiments of the intelligent control platform device of this application, reference can be made to the descriptions in the above method embodiments of this application, and details will not be repeated here.
[0068] In several embodiments provided in this application, it should be understood that the disclosed intelligent control platform devices and methods can be implemented in other ways. For example, the various embodiments of the intelligent control platform devices described above are merely illustrative. For example, the division of modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces, and the indirect coupling or communication connection of devices or units can be in electrical, mechanical or other forms.
[0069] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0070] In addition, in each embodiment of this application, the various functional units can be integrated in one processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit.
[0071] On the other hand, a readable storage medium is provided. A computer program is stored on the readable storage medium. When the computer program is executed by a processor, the method for constructing a video codec transmission system as described in any one of the above is implemented.
[0072] Please refer to Figure 10, when the above integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in the computer-readable storage medium 300. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions / computer programs to enable an intelligent control platform device (which can be a personal computer, server, or network device, etc.) or a processor to execute all or part of the steps of the methods in various embodiments of this application. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs, as well as electronic devices such as computers, mobile phones, laptop computers, tablet computers, cameras, etc. with the above storage media.
[0073] The description of the execution process of the program data in the computer-readable storage medium can refer to the description in the above method embodiments of this application, and will not be elaborated here.
[0074] The above are only the embodiments of this application, and do not limit the patent scope of this application. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of this application, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of this application.
[0075] Those skilled in the art can understand that in the above methods of the specific implementation manner, the writing order of each step does not mean a strict execution order that constitutes any limitation to the implementation process. The specific execution order of each step should be determined by its function and possible internal logic. To sum up, this application has the following beneficial effects: The method for constructing the video codec transmission system of this application effectively combats frequency-selective fading in the multipath fading channel by introducing OFDM modulation in the training system, and realizes high-quality transmission and reconstruction of video data under complex channel conditions.
[0076] Furthermore, it fully compresses the inter-frame redundancy and uses the feature-domain context as a condition instead of the traditional pixel-domain residual calculation, enabling the network to adaptively process regions with inaccurate motion or large uncertainties, and improving the robustness of the overall video compression.
[0077] Furthermore, a denoising network is introduced, which has lower computational complexity and faster convergence speed at the decoding end to meet the strict requirements of future wireless video communication, and improves the video transmission performance and robustness in complex channel environments.
Claims
1. A method for constructing a video encoding / decoding transmission system, characterized in that, The video codec transmission system includes a video codec. The method for constructing the video codec transmission system includes the steps: Construct an OFDM video transmission system; Construct a video codec, which includes a video encoder and a video decoder. The video encoder is arranged at the sending end of the OFDM video transmission system, and the video decoder is arranged at the receiving end of the OFDM video transmission system; Use the original training data set to train the video codec via the OFDM video transmission system until convergence.
2. The method for constructing a video codec transmission system according to claim 1, wherein The OFDM video transmission system includes an OFDM transmitter, a wireless multipath fading channel, and an OFDM receiver; Among them, the OFDM transmitter converts the compressed video output by the video encoder into frequency-domain symbols, applies the inverse discrete Fourier transform, and adds a cyclic prefix to obtain an OFDM transmission signal. The compressed video is obtained after the video encoder processes the original training data set; The OFDM transmission signal is transmitted through the wireless multipath fading channel and reaches the OFDM receiver; The OFDM receiver removes the cyclic prefix of the received signal and applies the discrete Fourier transform to finally obtain the received frequency-domain symbols. The received frequency-domain symbols are used to train the video decoder.
3. The method for constructing a video codec transmission system according to claim 2, wherein The frame sequence of the video is divided into key frames and interpolation frames; The constructed video encoder includes a key frame encoder and an interpolation frame encoder, and the video decoder includes a key frame decoder and an interpolation frame decoder; Among them, the key frame encoder is used to compress and encode the image of the key frame, and the key frame decoder is used to decode and reconstruct the image of the key frame; The interpolation frame encoder performs conditional compression encoding on the image of the interpolation frame using the multi-scale conditional context features of the reference frame, and the interpolation frame decoder is used to reconstruct the image of the interpolation frame.
4. The method for constructing a video codec transmission system according to claim 3, wherein, The reference frame includes the two adjacent frames before and after the interpolation frame. The multi-scale conditional context features are obtained through the following method: Construct a feature extraction network, which maps the reference frame into high-dimensional features, and further performs Gaussian blur through a group of Gaussian smoothing kernels with different scales to generate a scale space volume in the feature domain; Construct an optical flow network, which estimates the scale space flow between the interpolation frame and the reference frame to obtain the motion information between the interpolation frame and the reference frame; Perform a feature space warping operation on the scale space flow and the scale space volume to obtain the multi-scale conditional context features.
5. The method for constructing a video codec transmission system according to claim 4, characterized in that, The video codec transmission system further includes a denoising network. The method for constructing the video codec transmission system further includes the steps: Construct a denoising network, which is arranged at the receiving end of the OFDM video transmission system and is used to receive the noisy pilot symbols and data symbols obtained after OFDM demodulation through the multipath fading channel; Use the original training data set to jointly train the denoising network and the video codec via the OFDM video transmission system until convergence.
6. The method for constructing a video codec transmission system according to claim 5, characterized in that, The denoising network receives the OFDM demodulation signal after passing through the multipath fading channel, including the received frequency-domain symbols and received pilot symbols after channel fading and noise interference , as well as the original pilot symbols , and adaptively weights features using the channel state information implicit in the pilot symbols to efficiently filter out the noise and interference introduced by the channel from the received noisy data symbols, so as to finally obtain a clean compressed representation , that is: ; The output symbols processed by the denoising network form a clean compressed representation to reduce the learning complexity of the subsequent decoder and improve the reconstruction quality of video frames. The key-frame and interpolation-frame decoders no longer need to receive pilot inputs to implicitly perform channel estimation and equalization, simplifying the decoding task, that is: , 。 7. The method for constructing a video codec transmission system according to claim 6, wherein The step of jointly training the denoising network and the video codec via the OFDM video transmission system using the original training dataset until convergence includes: First, fix the interpolation-frame codec and separately train the key-frame codec and the denoising network. Using the OFDM system model, optimize the peak signal-to-noise ratio and multi-scale structural similarity of key-frame reconstruction under simulated multipath fading channel conditions; After the training of the key-frame codec converges, introduce the relevant networks of the interpolation-frame codec, including the scale-space flow estimation network, the conditional context feature extraction network, and the interpolation-frame decoding network, to perform end-to-end training of the interpolation-frame path; The training objectives include the denoising loss of the compressed representation and the video frame reconstruction loss. The loss function is constructed using the weighted mean squared error, and the final convergence of the model is achieved through joint optimization for the video frame sequence The training loss of , Among them, is the weighted coefficient of the denoising loss.
8. A video codec transmission system constructed by using the method according to any one of claims 1-7, characterized in that, The video codec transmission system is used to achieve accurate codec for video transmission in a wireless multipath fading environment.
9. An electronic device, characterized in that, It includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the steps of the method according to any one of claims 1-7.
10. A readable storage medium, characterized in that, A computer program is stored on the readable storage medium. When the computer program is executed by the processor, it implements the method for constructing the video codec transmission system according to any one of claims 1-7.
Citation Information
Patent Citations
Video SoftCast method based on residual distributed compressed sensing
CN105357536A
Generative multi-mode mutual benefit enhancement video semantic communication method
CN116939320A
OFDM channel estimation method and system based on deep learning
CN118784406A
Scale-spatial stream and hierarchical hyper-prior-based end-to-end video coding method research
CN119484877A
Channel estimation method based on MSM-DReEsNet in OFDM (Orthogonal Frequency Division Multiplexing) system
CN119865400A