Video communication method and apparatus, electronic device, and readable storage medium

An adaptive video communication method that selects and encodes keyframes at the transmitting end and decodes and interpolates them at the receiving end solves the problems of excessive latency and cliff effect in wireless video semantic communication, and achieves low-latency and high-efficiency video transmission.

CN119854564BActive Publication Date: 2026-08-25ZGC INSTITUTE OF UBIQUITOUS-X INNOVATION & APPLICATIONS
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411489462.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-24
Publication Date
2026-08-25
Estimated Expiration
2044-10-24

AI Technical Summary

Technical Problem

Existing wireless video semantic communication methods suffer from excessive processing latency when processing video data streams, especially in harsh channel environments where they cannot meet the low latency requirements of 6G. Furthermore, traditional source-channel separation coding schemes cannot effectively address the cliff effect.

Method used

By selecting keyframes at the sending end, jointly encoding them with the source channel, and then transmitting them, and performing joint decoding of the source channel and video frame interpolation at the receiving end, an adaptive video communication method is achieved, reducing processing latency and improving transmission efficiency.

Benefits of technology

It significantly reduces processing latency, improves transmission efficiency and communication reliability while ensuring video reconstruction quality, and can effectively cope with the cliff effect in harsh channel environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119854564B_ABST
    Figure CN119854564B_ABST
Patent Text Reader

Abstract

The application provides a video communication method and device, electronic equipment and readable storage medium. The method comprises: selecting key frames from a sequence of video frames to be transmitted according to the characteristics of the video and channel conditions, and obtaining a sequence of key frames; jointly encoding the sequence of key frames in a source channel; and transmitting the encoded sequence of key frames to a receiving end device. In the application, the transmitting end device selects key frames based on the characteristics of the video itself and the channel conditions, and transmits the encoded key frames to the receiving end device. The adaptive method of selecting key frames can ensure the reliability of communication. During video communication, only the encoded key frames are transmitted, which reduces the processing delay of the video frames, saves channel bandwidth, and improves transmission efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of communication technology, and in particular to a video communication method, apparatus, electronic device and readable storage medium. Background Technology

[0002] As human society becomes increasingly reliant on digital and networked lifestyles, the demand for networks is experiencing explosive growth. For specific scenarios served by future 6G technologies, such as Virtual Reality (VR) and telemedicine, the need for real-time video communication is particularly urgent. In these scenarios, wireless video communication not only needs to provide high-definition, low-latency video transmission, but also needs to ensure high reliability and efficiency throughout the transmission process.

[0003] However, traditional source-channel separation coding schemes cannot meet the low latency requirements of 6G, and they exhibit a cliff effect when the channel environment is poor. In recent years, the rapid development of AI technology and machine learning (ML) has brought revolutionary changes to the field of wireless communication technology, providing new ideas and possibilities for the research and development of 6G communication systems. The deep JSCC method, through an end-to-end deep learning model, tightly integrates source coding and channel coding, achieving joint optimization of source and channel characteristics.

[0004] Current wireless video semantic communication methods based on deep JSCC utilize deep neural networks to directly map video signals to channel symbols for compression, thus overcoming the cliff effect. However, most of these methods employ complex transmitter designs, focusing on improving video reconstruction quality while neglecting the problem of excessive processing latency. When processing video data streams, efficiency and real-time performance are crucial for user experience and the overall performance of the communication system. Therefore, effectively reducing processing latency while ensuring video reconstruction quality has become a major challenge that urgently needs to be addressed in the field of wireless video semantic communication. Summary of the Invention

[0005] The purpose of this application is to provide a video communication method, apparatus, electronic device, and readable storage medium that solves the problem of excessive latency in existing video communication processes.

[0006] To achieve the above objectives, embodiments of this application provide a video communication method applied to a transmitting device, comprising:

[0007] Based on the characteristics of the video and the channel conditions, key frames are selected from the video frame sequence to be transmitted to obtain a key frame sequence.

[0008] The keyframe sequence is jointly coded from the source and the channel.

[0009] The encoded keyframe sequence is sent to the receiving device.

[0010] Optionally, based on the characteristics of the video and channel conditions, key frames are selected from the video frame sequence to be transmitted to obtain a key frame sequence, including:

[0011] The transmission rate is determined based on the characteristics of the video.

[0012] The number and location of keyframes are determined based on the transmission rate and channel conditions.

[0013] Keyframes are selected from the video frame sequence based on the quantity and position to obtain a keyframe sequence.

[0014] Optionally, the characteristics of the video include: the video's frame rate and / or resolution;

[0015] The process of determining the transmission rate based on the characteristics of the video includes:

[0016] Calculate the transmission rate based on the video's frame rate, resolution, the inter-frame compression factor for keyframe selection, and the intra-frame compression factor for encoding.

[0017] Optionally, the channel conditions include: channel bandwidth and / or signal-to-noise ratio;

[0018] Determining the number and location of keyframes based on the transmission rate and channel conditions includes:

[0019] Based on the transmission rate and channel capacity, determine the minimum inter-frame compression factor for selecting key frames, wherein the transmission rate is less than or equal to the channel capacity, and the channel capacity is calculated based on the channel bandwidth and signal-to-noise ratio;

[0020] The number of keyframes is calculated based on the minimum inter-frame compression factor.

[0021] The position of the keyframe in the video frame sequence is determined based on the number of keyframes and the preset sampling method.

[0022] Optionally, the keyframe sequence is subjected to joint source-channel coding, including:

[0023] Semantic extraction is performed on the keyframe sequence to obtain the semantic feature map of the keyframe sequence;

[0024] The semantic feature map is jointly encoded by the source and channel to obtain the encoded sequence.

[0025] Optionally, before sending the encoded keyframe sequence to the receiving device, the method further includes:

[0026] The encoded sequence is quantized;

[0027] Sending the encoded keyframe sequence to the receiving device includes:

[0028] Send the quantized encoded sequence to the receiving device.

[0029] Optionally, before selecting keyframes from the video frame sequence to be transmitted based on its characteristics and channel conditions to obtain a keyframe sequence, the method further includes:

[0030] Video capture is performed to determine the video's frame rate and resolution, thereby obtaining a video frame sequence.

[0031] To achieve the above objectives, embodiments of this application provide a video communication method applied to a receiving device, comprising:

[0032] Receive keyframe sequence;

[0033] The keyframe sequence is subjected to joint source-channel decoding to obtain the reconstructed keyframe sequence;

[0034] Based on the number and position of the keyframes in the reconstructed keyframe sequence, video frame interpolation is performed on the reconstructed keyframe sequence to obtain a reconstructed video frame sequence;

[0035] The number and location of the keyframes are related to the characteristics of the video and the channel conditions.

[0036] Optionally, based on the number and position of keyframes in the reconstructed keyframe sequence, video frame interpolation is performed on the reconstructed keyframe sequence to obtain a reconstructed video frame sequence, including:

[0037] The position and number of video interpolated frames are determined based on the position and number of keyframes in the reconstructed keyframe sequence.

[0038] Based on the position and number of the video interpolation frames, the reconstructed keyframe sequence is interpolated to obtain a reconstructed video frame sequence.

[0039] Optionally, the characteristics of the video include the video's frame rate and / or resolution;

[0040] The channel conditions include: channel bandwidth and / or signal-to-noise ratio;

[0041] The number and location of the keyframes are determined based on the transmission rate and channel conditions; the transmission rate is calculated based on the frame rate and resolution of the video.

[0042] Optionally, the keyframe sequence is subjected to joint source-channel decoding to obtain a reconstructed keyframe sequence, including:

[0043] The keyframe sequence is subjected to joint source-channel decoding to obtain a reconstructed semantic feature map;

[0044] Semantic recovery is performed on the reconstructed semantic feature map to obtain the reconstructed keyframe sequence.

[0045] Optionally, the method further includes:

[0046] The reconstructed video frame sequence is played according to the video's frame rate.

[0047] To achieve the above objectives, embodiments of this application provide a video communication device applied to a transmitting end device, comprising:

[0048] The keyframe selection module is used to select keyframes from the video frame sequence to be transmitted based on the characteristics of the video and channel conditions, thereby obtaining a keyframe sequence.

[0049] The encoding module is used to perform source-channel joint coding on the keyframe sequence;

[0050] A wireless transmission module is used to send the encoded keyframe sequence to the receiving device.

[0051] To achieve the above objectives, embodiments of this application provide a video communication device applied to a receiving end device, comprising:

[0052] The receiving module is used to receive keyframe sequences;

[0053] The decoding module is used to perform joint source-channel decoding on the keyframe sequence to obtain the reconstructed keyframe sequence;

[0054] The video frame interpolation module is used to perform video frame interpolation on the reconstructed keyframe sequence according to the number and position of keyframes in the reconstructed keyframe sequence to obtain a reconstructed video frame sequence.

[0055] The number and location of the keyframes are related to the characteristics of the video and the channel conditions.

[0056] To achieve the above objectives, embodiments of this application provide an electronic device, including a transceiver, a processor, a memory, and a program or instructions stored in the memory and executable on the processor; when the processor executes the program or instructions, it implements the steps of the video communication method described above.

[0057] To achieve the above objectives, embodiments of this application provide a readable storage medium storing a program or instructions thereon, which, when executed by a processor, implement the steps of the video communication method described above.

[0058] To achieve the above objectives, embodiments of this application provide a computer program product that stores a program or instructions, including computer instructions, which, when executed by a processor, implement the steps of the video communication method described above.

[0059] The beneficial effects of the above technical solution in this application are as follows:

[0060] In embodiments of this application, the transmitting device selects key frames based on the characteristics of the video itself and channel conditions, encodes the key frames, and sends them to the receiving device. This adaptive key frame selection method ensures communication reliability. During video communication, only the encoded key frames are transmitted, reducing processing latency for video frames, saving channel bandwidth, and improving transmission efficiency. Attached Figure Description

[0061] Figure 1 This is one of the flowcharts illustrating the video communication method according to an embodiment of this application;

[0062] Figure 2 This is a second schematic flowchart of the video communication method according to an embodiment of this application;

[0063] Figure 3 This is a convolutional network used for semantic extraction in an embodiment of this application;

[0064] Figure 4 This is a convolutional network used for joint source-channel coding in an embodiment of this application;

[0065] Figure 5 This is a convolutional network used for joint source-channel decoding in an embodiment of this application;

[0066] Figure 6 This is a convolutional network used for semantic recovery in an embodiment of this application;

[0067] Figure 7 This is a schematic diagram of video frame interpolation according to an embodiment of this application;

[0068] Figure 8 This is one of the schematic diagrams illustrating the reconstruction quality and latency effects of the video communication method according to an embodiment of this application compared to a traditional video encoding and transmission scheme;

[0069] Figure 9 This is a second schematic diagram illustrating the difference in reconstruction quality and latency between the video communication method of this application and a traditional video encoding and transmission scheme.

[0070] Figure 10 This is the third flowchart illustrating the video communication method according to an embodiment of this application;

[0071] Figure 11 This is one of the structural schematic diagrams of a video communication device according to an embodiment of this application;

[0072] Figure 12 This is a second schematic diagram of the structure of the video communication device according to an embodiment of this application;

[0073] Figure 13 This is a schematic diagram of the structure of an electronic device according to an embodiment of this application. Detailed Implementation

[0074] To make the technical problems, technical solutions and advantages of this application clearer, a detailed description will be provided below in conjunction with the accompanying drawings and specific embodiments.

[0075] It should be understood that the phrase "one embodiment" or "an embodiment" throughout the specification means that a specific feature, structure, or characteristic related to the embodiment is included in at least one embodiment of this application. Therefore, "in one embodiment" or "in an embodiment" appearing throughout the specification does not necessarily refer to the same embodiment. Furthermore, these specific features, structures, or characteristics can be combined in any suitable manner in one or more embodiments.

[0076] In the various embodiments of this application, it should be understood that the sequence number of each process described below does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.

[0077] In addition, the terms "system" and "network" are often used interchangeably in this article.

[0078] In the embodiments provided in this application, it should be understood that "B corresponding to A" means that B is associated with A, and B can be determined based on A. However, it should also be understood that determining B based on A does not mean determining B solely based on A; B can also be determined based on A and / or other information.

[0079] like Figure 1 As shown, this application provides a video communication method applied to a transmitting device, including:

[0080] Step 11: Based on the characteristics of the video and the channel conditions, select key frames from the video frame sequence to be transmitted to obtain a key frame sequence;

[0081] Step 12: Perform source-channel joint coding on the keyframe sequence;

[0082] Step 13: Send the encoded keyframe sequence to the receiving device.

[0083] In this embodiment, the transmitting device selects a subset of key frames from the video frame sequence to be transmitted, processes them, and then sends them to the receiving device. The selected key frames are related to the characteristics of the video and channel conditions. On one hand, channel conditions affect communication reliability. For example, when channel conditions are good, the channel capacity is large, and the code length can be appropriately increased to improve the transmission rate, but the transmission rate must always be kept below the channel capacity. When channel conditions are poor, the channel capacity decreases, and the code length should be reduced to lower the transmission rate to ensure communication reliability. On the other hand, the characteristics of the video source itself also affect the transmission rate, thus impacting communication reliability.

[0084] In embodiments of this application, the transmitting device selects key frames based on the characteristics of the video itself and channel conditions, encodes the key frames, and sends them to the receiving device. This adaptive key frame selection method ensures communication reliability. During video communication, only the encoded key frames are transmitted, reducing processing latency for video frames, saving channel bandwidth, and improving transmission efficiency.

[0085] As an optional embodiment, based on the characteristics of the video and channel conditions, key frames are selected from the video frame sequence to be transmitted to obtain a key frame sequence, including:

[0086] The transmission rate is determined based on the characteristics of the video.

[0087] The number and location of keyframes are determined based on the transmission rate and channel conditions.

[0088] Keyframes are selected from the video frame sequence based on the quantity and position to obtain a keyframe sequence.

[0089] Optionally, the video characteristics include the video frame rate and / or resolution. Since the frame rate and resolution of the video source itself also affect the transmission rate, in this embodiment, the sending device determines the transmission rate based on the video characteristics, ensuring communication reliability. Optionally, the channel conditions include channel bandwidth and / or signal-to-noise ratio. Channel conditions affect channel capacity, and the transmission rate needs to be less than the channel capacity. Based on the transmission rate and channel capacity conditions, the number and position of keyframes to be selected are determined, thereby obtaining a keyframe sequence. This achieves adaptive transmission rate determination and keyframe selection.

[0090] When transmitting video, the transmitting device can adaptively select the required key frames based on the characteristics of the video and the channel conditions, and transmit the key frames to the receiving end, reducing the processing steps of video frames and saving transmission resources.

[0091] Optionally, the characteristics of the video include: the video's frame rate and / or resolution;

[0092] The method of determining the transmission rate based on the characteristics of the video includes: calculating the transmission rate based on the video's frame rate, resolution, the inter-frame compression factor of the selected keyframes, and the intra-frame compression factor of the encoding.

[0093] Specifically, the transmission rate R is calculated using the following formula: R = r × w × h ÷ ρ ÷ θ. Where r is the video frame rate, w × h is the video resolution, ρ is the inter-frame compression factor of the keyframe selector, and θ is the intra-frame compression factor of the source-channel joint encoder. The inter-frame compression factor can be used to indicate how many frames are needed to select a keyframe.

[0094] Optionally, the channel conditions include: channel bandwidth and / or signal-to-noise ratio;

[0095] Determining the number and location of keyframes based on the transmission rate and channel conditions includes:

[0096] Based on the transmission rate and channel capacity, determine the minimum inter-frame compression factor for selecting key frames, wherein the transmission rate is less than or equal to the channel capacity, and the channel capacity is calculated based on the channel bandwidth and signal-to-noise ratio;

[0097] The number of keyframes is calculated based on the minimum inter-frame compression factor.

[0098] The position of the keyframe in the video frame sequence is determined based on the number of keyframes and the preset sampling method.

[0099] In this embodiment, since the transmission rate R is less than or equal to the channel capacity C, the minimum inter-frame compression factor ρ can be solved according to the above formula for calculating the transmission rate R. min Given the total number of video frames and the minimum inter-frame compression factor length ρ min The number of keyframes selected, k, can be calculated.

[0100] Among them, channel capacity Where W is the channel bandwidth. For signal-to-noise ratio, its conversion formula is: When the transmission rate R > C, reliable transmission cannot be achieved. Therefore, it is necessary to ensure that the code rate R is lower than the channel capacity C.

[0101] Optionally, the device used to determine channel conditions can be a spectrum analysis device at the transmitting end, used to obtain channel bandwidth, signal-to-noise ratio, etc. For example, a spectrum analyzer can be used for measurement; it can receive the input signal and convert it into a spectrum diagram or spectrum density diagram for display.

[0102] As an optional embodiment, source-channel joint coding is performed on the keyframe sequence, including:

[0103] Semantic extraction is performed on the keyframe sequence to obtain the semantic feature map of the keyframe sequence;

[0104] The semantic feature map is jointly encoded by the source and channel to obtain the encoded sequence.

[0105] In this embodiment, the encoding device can be a semantic encoding device at the transmitting end, which is used to convert the key frame sequence to be transmitted into an electrical signal suitable for transmission. Based on the relevant parameters of the video determined by the sensor device and the channel bandwidth and signal-to-noise ratio measured by the spectrum analysis device, the transmission rate is adaptively adjusted. After determining the number of key frames to be selected, the key frame sequence to be transmitted is selected, the key frame sequence is encoded and quantized, the bandwidth of the transmitted data is reduced, and the transmission efficiency is improved.

[0106] Optionally, before sending the encoded keyframe sequence to the receiving device, the method further includes:

[0107] The encoded sequence is quantized;

[0108] Sending the encoded keyframe sequence to the receiving device includes:

[0109] Send the quantized encoded sequence to the receiving device.

[0110] In this embodiment, the transmitting end can transmit the electrical signal to be transmitted from the transmitting end to the receiving end through a wireless transmission device, such as through Wi-Fi, Bluetooth, Near Field Communication (NFC), 5G, etc.

[0111] As an optional embodiment, before selecting key frames from the video frame sequence to be transmitted based on the characteristics and channel conditions to obtain the key frame sequence, the method further includes: performing video acquisition, determining the video frame rate and resolution, and obtaining the video frame sequence.

[0112] In this embodiment, the device for video acquisition can be a sensor device used to acquire video from the environment and determine the relevant parameters of the video. One or more cameras can be used to acquire video.

[0113] The transmitting device performs video acquisition, determines the frame rate and resolution of the video, and obtains a video frame sequence. Keyframe selection is performed on the video frame sequence to obtain a keyframe sequence. Semantic extraction is performed on the keyframe sequence to obtain a semantic feature map; source-channel joint coding is performed on the semantic feature map to obtain a coded sequence; the coded sequence is quantized to obtain a quantized sequence, which is then sent to the receiving device. The quantized sequence is the quantized coded sequence.

[0114] The keyframe selection step of the transmitting device includes: determining an adaptive transmission rate based on the characteristics of the video frame sequence to be transmitted and the channel conditions; determining the number and position of the keyframes to be selected based on the adaptive transmission rate, and obtaining the keyframe sequence.

[0115] After receiving the quantized sequence, the receiving device decodes it to obtain a reconstructed semantic feature map; it then performs semantic recovery on the reconstructed semantic feature map to obtain a reconstructed keyframe sequence; finally, it performs video frame interpolation on the reconstructed keyframe sequence to obtain a complete reconstructed video frame sequence. Optionally, video display is performed after the video frame interpolation step, and the reconstructed video frame sequence is played according to the frame rate.

[0116] The step of the receiving device performing video frame interpolation on the reconstructed keyframe sequence includes: determining the number and position of video frames to be interpolated based on the number and position of the selected keyframes, thereby obtaining a complete reconstructed video frame sequence.

[0117] Specifically, the receiving end can use a semantic decoding device to recover the reconstructed video from the received electrical signal, decode the received electrical signal to recover the reconstructed keyframe sequence, and then recover the complete video using the reconstructed keyframe sequence. The receiving end can also play the reconstructed video through an output device, which can be displayed on a monitor.

[0118] In this embodiment, the transmitting end determines the channel conditions through a spectrum analysis device, collects video through a sensor device, converts the video into an electrical signal through a semantic encoding device, transmits it through a wireless transmission device, and at the receiving end, the video is restored and reconstructed through a semantic decoding device and played on a display. This realizes wireless video semantic communication, greatly reduces processing latency, improves transmission efficiency, and ensures communication reliability.

[0119] To make the objectives, technical solutions, and advantages of this application clearer, the video communication method of this application will be described below with reference to specific embodiments and accompanying drawings.

[0120] The video communication method in this application embodiment is as follows: Figure 2 As shown, for a video X to be transmitted, the transmitting end selects keyframes from the video frame sequence using a keyframe selector to obtain a keyframe sequence; it then performs semantic extraction on the keyframe sequence using a semantic extractor to obtain a semantic feature map; finally, it performs source-channel joint coding on the semantic feature map to obtain a coded sequence; and then it quantizes the coded sequence using a quantizer and transmits the quantized coded sequence to the receiving end via a wireless channel. It should be noted that before selecting keyframes, video characteristics such as bitrate and resolution are obtained.

[0121] The receiving device decodes the received sequence using a source-channel joint decoder to obtain a reconstructed semantic feature map; it then performs semantic recovery on the reconstructed semantic feature map using a semantic restorer to obtain a reconstructed keyframe sequence; and finally, it performs video frame interpolation on the reconstructed keyframe sequence using a frame interpolator to obtain a complete reconstructed video frame sequence.

[0122] Example 1

[0123] Instructions regarding the video to be transmitted.

[0124] In this embodiment, it is assumed that the transmitted video uses videos from the UVG dataset, with each video containing 600 frames. The UVG dataset is renowned for its high-quality video material and extensive scene diversity, providing an ideal testing environment. The frame rate of the test video is set to 25fps, a common video frame rate standard used in movies, television, and many other video applications. Furthermore, the resolution is set to 1920×1080, i.e., the high-definition 1080p standard, to ensure that the test video has sufficient detail and clarity.

[0125] For the transmitting device, the video communication method includes the following steps:

[0126] Step 1: Train the transmission model.

[0127] In this embodiment, the Vimeo-90k dataset was used as the training set during model training. It provides a wide variety of video footage suitable for tasks such as video frame reconstruction. This dataset contains 89,800 video clips, each showcasing various scenes and actions, and each clip consists of a sequence of 7 frames. During training, each frame from each video clip was treated as an independent video frame as input. To enhance the model's generalization ability and robustness, the input frames were randomly cropped to 256×256 pixels during training. This data augmentation technique helps the model learn more robust feature representations, as it requires processing image patches of different sizes and locations. The hyperparameter was set to 8192 during training. Additionally, the mini-batch size was set to 36. The choice of batch size is crucial for training stability and efficiency. Smaller batch sizes may lead to unstable training, while larger batch sizes may require more computational resources. Setting the minimum batch size to 36 ensures both stable and efficient training.

[0128] Step 2, Spectrum Analysis.

[0129] This step determines the channel conditions, including the channel bandwidth and signal-to-noise ratio.

[0130] Step 3: Determine the adaptive transmission rate.

[0131] The adaptive transmission rate is determined by both video characteristics and channel conditions. On one hand, when channel conditions are good and channel capacity is large, the code length can be appropriately increased to improve the transmission rate, but the transmission rate must always be kept below the channel capacity. Conversely, when channel conditions are poor and channel capacity decreases, the code length should be reduced to lower the transmission rate and ensure communication reliability. On the other hand, the frame rate and resolution of the video source itself also affect the transmission rate. In this example, the adaptive transmission rate is controlled by varying the inter-frame compression factor, which is controlled by the number of keyframes selected.

[0132] According to Shannon's second theorem, the channel capacity of a communication system... Where W is the channel bandwidth. For signal-to-noise ratio, its conversion formula is: When the transmission rate R > C, reliable transmission cannot be achieved; therefore, it is necessary to ensure that the bit rate R is lower than the channel capacity. In this embodiment, it is assumed that the communication bandwidth is fixed at 1MHz, the frame rate of the video source is r = 25fps, and the size (resolution) is h × w = 1920 × 1080. It is also assumed that the inter-frame compression factor of the keyframe selector is ρ, and the intra-frame compression factor of the source-channel joint encoder is θ. The transmission rate R can be expressed as: R = r × w × h ÷ ρ ÷ θ. The minimum inter-frame compression factor length ρ can be solved using R ≤ C. min Then, the minimum inter-frame compression factor length ρ is used. min Calculate the number of keyframes selected, k.

[0133] Step 4: Select the keyframe to be transmitted.

[0134] In this embodiment, the system input is a sequence of video frames, X = {x1, x2, ..., x...} T There are a total of T video frames, where the frame at time step t is modeled as a vector of pixel intensity. This embodiment selects a keyframe sequence of number k obtained by determining the keyframe selection through the adaptive transmission rate using an equally spaced sampling method, denoted as:

[0135] Step 5: Semantic extraction.

[0136] The keyframe sequence is semantically extracted using a convolutional network and transformed into a semantic feature map. The convolutional network, for example Figure 3 As shown.

[0137] Step 6: Joint coding of source and channel.

[0138] The semantic feature map is transformed into a continuous value symbol sequence through joint source-channel coding using a convolutional network. The convolutional network, for example Figure 4 As shown.

[0139] Step 7: Quantification.

[0140] The continuous value symbol sequence is transformed into a discrete value sequence with finite values ​​to be transmitted through soft quantization.

[0141] Step 8: Wireless transmission.

[0142] This embodiment primarily considers the widely used Additive White Gaussian Noise (AWGN) channel. Under an AWGN channel, the channel transfer function is: Where σ 2 It is the energy of noise; therefore, the sequence to be transmitted is transmitted through a wireless channel to obtain the received sequence.

[0143] For the receiving device, the video communication method includes the following steps:

[0144] Step 9: The receiving device performs joint decoding of the source and channel.

[0145] The receiving device performs joint source-channel decoding on the received sequence, transforming it into a reconstructed semantic feature map: The convolutional network used for decoding, for example Figure 5 As shown.

[0146] Step 10: The receiving device performs semantic recovery on the received sequence.

[0147] Semantic recovery is performed on the reconstructed semantic feature map, transforming it into a reconstructed keyframe sequence. The convolutional networks used for semantic recovery, such as Figure 6 As shown.

[0148] Step 11: The receiving device performs video frame interpolation on the keyframe sequence.

[0149] This embodiment uses an Intermediate Feature Refine Network (IFRNet) as the video frame interpolation network to interpolate virtual frames to approximate the original video frames. For example... Figure 7 As shown, assuming a keyframe is selected at 8-frame intervals, the reconstruction is first performed using two adjacent keyframes. and Frame interpolation Then by and as well as and Frame interpolation was performed separately. and This process continues until the entire reconstructed video frame sequence is obtained.

[0150] Step 12: After obtaining the reconstructed video frame sequence, the receiving device can play the video according to the video's frame rate.

[0151] This application embodiment can also evaluate the quality of the reconstructed video. For example, this embodiment uses Peak Signal-to-Noise Ratio (PSNR) and Multi-Scale Structural Similarity (MS-SSIM) as evaluation metrics for reconstruction quality. PSNR measures the objective difference between the reconstructed image and the original image, and can better reflect the objective quality of the image, measuring the accuracy of the frame reconstruction; MS-SSIM considers the human visual system's perception of structural information, and can better reflect the details and texture information of the image, measuring and evaluating the perceptual quality of the reconstructed frame.

[0152] This embodiment compares its reconstruction quality and latency with traditional video coding and transmission schemes and Model Division Video Semantic Communication (MDVSC) schemes. Traditional schemes include source coding H.264 and H.265, channel coding LDPC with half the coding efficiency, and modulation using Binary Phase Shift Keying (BPSK) and Quadrature Amplitude Modulation (QAM). For example... Figure 8 , Figure 9 As shown in the diagram, where WAFI-VSC represents the example, under PSNR metrics, compared to the 1 / 2LDPC+BPSK scheme, this embodiment outperforms the scheme combined with H.264 at high SNR, but is slightly lower than the scheme combined with H.265. At low SNR, traditional schemes suffer from the cliff effect, while this embodiment effectively addresses the cliff effect, achieving better recovery metrics. However, under MS-SSIM metrics, this embodiment almost completely surpasses the 1 / 2LDPC+BPSK scheme combining H.264 and H.265. This indicates that this embodiment performs better on metrics closer to human perception, exhibiting a recovery effect more suitable for the human visual system. In contrast, the 1 / 2LDPC+16QAM scheme, affected by the cliff effect, cannot transmit video normally across the entire SNR range.

[0153] like Figure 8 , Figure 9 As shown, the PSNR and MS-SSIM curves of this embodiment almost comprehensively surpass those of MDVSC. Furthermore, during the testing process, we incorporated a clock timing function to statistically analyze processing latency. The test results indicate that the processing latency per frame in this embodiment does not exceed 1ms, which is only 1 / 1800th of that of MDVSC.

[0154] This embodiment provides a video communication method for end-to-end adaptive frame selection of video. By adding an adaptive frame selection method and frame interpolation technology to the video semantic communication system, the reliability of the video semantic communication system can be guaranteed while improving its effectiveness and reducing processing latency, thus achieving better communication results.

[0155] Example 2: Taking the video communication method of this application as an example of applying it to an unmanned workshop, it can achieve the effect of video sharing between two unmanned workshops. For example, unmanned vehicle A is used as the transmitting device, and unmanned vehicle B is used as the receiving device. Compared with Example 1, the following part is added at the beginning of the entire process.

[0156] Video Acquisition: Unmanned vehicle A acquires real-time video, obtains the video frame sequence to be transmitted, and acquires parameters such as the frame rate and resolution of the video.

[0157] The other contents of this embodiment are similar to those of Example 1 above, and will not be repeated here.

[0158] In the embodiments of this application, video communication, compared to image communication, adds a process of removing inter-frame redundancy. This process mostly relies on traditional video communication architectures or other complex processing, resulting in high processing latency. AI technology has developed rapidly in recent years, and video frame interpolation technology has also developed accordingly. Video frame interpolation is an important low-level computer vision task, and traditional video frame interpolation mostly relies on optical flow estimation. This application fully utilizes the advantages of video frame interpolation technology and applies it to video semantic communication. At the sending end, only a small number of selected key frames need to be sent. These key frames are then transmitted after deep JSCC encoding. At the receiving end, other non-key frames are generated using the recovered key frames through a video frame interpolation network. This application improves transmission efficiency, saves communication bandwidth, greatly reduces complexity, and lowers processing latency at the sending end through key frame selection and frame generation methods.

[0159] The adaptive transmission rate method proposed in this embodiment allows the system to intelligently select or adjust the number of keyframes in real time based on the current channel conditions and the characteristics of the video content itself. This method avoids data loss or transmission interruption caused by the transmission rate exceeding the actual bandwidth of the channel during data transmission, thereby ensuring the stability and reliability of the communication process. Simultaneously, by adjusting the number of keyframes selected, communication bandwidth resources can be saved to the greatest extent, optimizing transmission efficiency. This design concept aligns with the core idea of ​​Shannon's second theorem, which states that reliable data transmission becomes impossible when the data transmission rate exceeds the channel's bandwidth limit. Therefore, the adaptive transmission rate method of this embodiment not only meets the requirements of communication reliability but also achieves efficient utilization of communication resources, improving the flexibility of the communication system.

[0160] Furthermore, 6G technology aims to achieve more stable and reliable communication, with stricter latency requirements, especially in scenarios with extremely high data transmission demands, such as remote medical surgery and autonomous driving. These scenarios require ensuring the accuracy, integrity, and immediacy of data transmission. Traditional methods, however, cannot provide sufficient redundancy and recovery mechanisms to handle errors and data loss during data transmission, nor can they provide sufficiently fast transmission speeds and response capabilities. They are also susceptible to the cliff effect under poor channel conditions. With the continuous improvement of Direct Neural Networks (DNNs), DNN technology can be used to process the encoding process in communication, jointly processing source coding and channel coding. The video to be transmitted is used as input to the DNN network. Through a series of convolutional processes, the size is reduced and the dimensionality is increased to obtain the final encoded sequence to be transmitted. At the receiving end, the original video is recovered through the reverse operation. By training, the network parameters can be learned to better adapt to the channel, resulting in better recovery performance. The method in this application can solve the cliff effect problem in traditional communication and improve the quality of video reconstruction.

[0161] like Figure 10 As shown in the illustration, this application also provides a video communication method applied to a receiving device, comprising:

[0162] Step 101: Receive the keyframe sequence;

[0163] Step 102: Perform joint source-channel decoding on the keyframe sequence to obtain the reconstructed keyframe sequence;

[0164] Step 103: Based on the number and position of the keyframes in the reconstructed keyframe sequence, perform video frame interpolation on the reconstructed keyframe sequence to obtain a reconstructed video frame sequence; wherein, the number and position of the keyframes are related to the characteristics of the video and the channel conditions.

[0165] In this embodiment, the transmitting device selects a portion of key frames from the video frame sequence to be transmitted, processes them, and then sends them to the receiving device. The selected key frames are related to the characteristics of the video and the channel conditions.

[0166] The transmitting device performs video acquisition, determines the frame rate and resolution of the video, and obtains a video frame sequence. Keyframe selection is performed on the video frame sequence to obtain a keyframe sequence. Semantic extraction is performed on the keyframe sequence to obtain a semantic feature map; source-channel joint coding is performed on the semantic feature map to obtain a coding sequence; the coding sequence is quantized to obtain a quantized sequence, which is then sent to the receiving device. This quantized sequence is the quantized keyframe coding sequence.

[0167] After receiving the quantized sequence, the receiving device decodes it to obtain the reconstructed keyframe sequence. Based on the number and position of the keyframes in the reconstructed keyframe sequence, video frame interpolation is performed on the reconstructed keyframe sequence to obtain the reconstructed video frame sequence.

[0168] In embodiments of this application, the transmitting device selects keyframes based on the characteristics of the video itself and channel conditions, encodes the keyframes, and sends them to the receiving device. The receiving device receives the keyframe sequence, decodes it to obtain a reconstructed keyframe sequence, and then performs video frame interpolation on the reconstructed keyframe sequence according to the number and position of the keyframes to obtain a reconstructed video frame sequence. This adaptive keyframe selection method ensures communication reliability. During video communication, only encoded keyframes are transmitted; the receiving end can obtain the reconstructed video frame sequence by decoding and interpolating the keyframes, reducing video frame processing latency, saving channel bandwidth, and improving transmission efficiency.

[0169] As an optional embodiment, based on the number and position of keyframes in the reconstructed keyframe sequence, video frame interpolation is performed on the reconstructed keyframe sequence to obtain a reconstructed video frame sequence, including:

[0170] Based on the position and number of keyframes in the reconstructed keyframe sequence, determine the position and number of video interpolation frames; perform video interpolation on the reconstructed keyframe sequence based on the position and number of video interpolation frames to obtain a reconstructed video frame sequence.

[0171] In this embodiment, the receiving device determines the number and position of video interpolated frames based on the number and position of keyframes to obtain a complete reconstructed video frame sequence. This operation is simple and easy to implement, reducing the complexity of video frame processing and saving processing resources.

[0172] Optionally, the characteristics of the video include the video's frame rate and / or resolution;

[0173] The channel conditions include: channel bandwidth and / or signal-to-noise ratio;

[0174] The number and location of the keyframes are determined based on the transmission rate and channel conditions; the transmission rate is calculated based on the frame rate and resolution of the video.

[0175] Specifically, the transmission rate R is calculated using the following formula: R = r × w × h ÷ ρ ÷ θ. Where r is the video frame rate, w × h is the video resolution, ρ is the inter-frame compression factor of the keyframe selector, and θ is the intra-frame compression factor of the source-channel joint encoder. The inter-frame compression factor can be used to indicate how many frames are needed to select a keyframe.

[0176] Since the transmission rate R is less than or equal to the channel capacity C, the minimum inter-frame compression factor ρ can be solved using the formula for calculating the transmission rate R. min Given the total number of video frames and the minimum inter-frame compression factor length ρ min The number of keyframes selected, k, can be calculated.

[0177] Among them, channel capacity Where W is the channel bandwidth. For signal-to-noise ratio, its conversion formula is: When the transmission rate R > C, reliable transmission cannot be achieved. Therefore, it is necessary to ensure that the code rate R is lower than the channel capacity C.

[0178] As an optional embodiment, the keyframe sequence is subjected to joint source-channel decoding to obtain a reconstructed keyframe sequence, including:

[0179] The keyframe sequence is subjected to joint source-channel decoding to obtain a reconstructed semantic feature map; the reconstructed semantic feature map is then subjected to semantic recovery to obtain a reconstructed keyframe sequence.

[0180] Optionally, the method further includes: playing the reconstructed video frame sequence according to the video's frame rate.

[0181] In this embodiment, the transmitting device performs video acquisition, determines the frame rate and resolution of the video, and obtains a video frame sequence. Keyframe selection is performed on the video frame sequence to obtain a keyframe sequence. Semantic extraction is performed on the keyframe sequence to obtain a semantic feature map; source-channel joint coding is performed on the semantic feature map to obtain a coding sequence; the coding sequence is quantized to obtain a quantized sequence, and the quantized sequence is sent to the receiving device. The quantized sequence is the quantized coding sequence.

[0182] After receiving the quantized sequence, the receiving device decodes it to obtain a reconstructed semantic feature map; it then performs semantic recovery on the reconstructed semantic feature map to obtain a reconstructed keyframe sequence; finally, it performs video frame interpolation on the reconstructed keyframe sequence to obtain a complete reconstructed video frame sequence. Optionally, video display is performed after the video frame interpolation step, and the reconstructed video frame sequence is played according to the frame rate.

[0183] The step of the receiving device performing video frame interpolation on the reconstructed keyframe sequence includes: determining the number and position of video frames to be interpolated based on the number and position of the selected keyframes, thereby obtaining a complete reconstructed video frame sequence.

[0184] Specifically, the receiving end can use a semantic decoding device to recover the reconstructed video from the received electrical signal, decode the received electrical signal to recover the reconstructed keyframe sequence, and then recover the complete video using the reconstructed keyframe sequence. The receiving end can also play the reconstructed video through an output device, which can be displayed on a monitor.

[0185] In this embodiment, the transmitting end determines the channel conditions through a spectrum analysis device, collects video through a sensor device, converts the video into an electrical signal through a semantic encoding device, transmits it through a wireless transmission device, and at the receiving end, the video is restored and reconstructed through a semantic decoding device and played on a display. This realizes wireless video semantic communication, greatly reduces processing latency, improves transmission efficiency, and ensures communication reliability.

[0186] like Figure 11 As shown, this application embodiment also provides a video communication device 1100, applied to a transmitting end device, including:

[0187] The key frame selection module 1110 is used to select key frames from the video frame sequence to be transmitted based on the characteristics of the video and channel conditions, and obtain a key frame sequence.

[0188] Encoding module 1120 is used to perform source-channel joint coding on the keyframe sequence;

[0189] The wireless transmission module 1130 is used to send the encoded key frame sequence to the receiving device.

[0190] Optionally, the keyframe selection module includes:

[0191] A transmission rate determination unit is used to determine the transmission rate based on the characteristics of the video.

[0192] The key frame information determination unit is used to determine the number and location of key frames based on the transmission rate and channel conditions.

[0193] A keyframe selection unit is used to select keyframes from the video frame sequence according to the number and position to obtain a keyframe sequence.

[0194] Optionally, the characteristics of the video include: the video's frame rate and / or resolution;

[0195] The transmission rate determination unit is specifically used for:

[0196] Calculate the transmission rate based on the video's frame rate, resolution, the inter-frame compression factor for keyframe selection, and the intra-frame compression factor for encoding.

[0197] Optionally, the channel conditions include: channel bandwidth and / or signal-to-noise ratio;

[0198] The keyframe information determination unit is specifically used for:

[0199] Based on the transmission rate and channel capacity, determine the minimum inter-frame compression factor for selecting key frames, wherein the transmission rate is less than or equal to the channel capacity, and the channel capacity is calculated based on the channel bandwidth and signal-to-noise ratio;

[0200] The number of keyframes is calculated based on the minimum inter-frame compression factor.

[0201] The position of the keyframe in the video frame sequence is determined based on the number of keyframes and the preset sampling method.

[0202] Optionally, the encoding module is specifically used for:

[0203] Semantic extraction is performed on the keyframe sequence to obtain the semantic feature map of the keyframe sequence;

[0204] The semantic feature map is jointly encoded by the source and channel to obtain the encoded sequence.

[0205] Optionally, the device further includes:

[0206] A quantization module is used to quantize the encoded sequence;

[0207] The wireless transmission module is specifically used to send the quantized encoded sequence to the receiving device.

[0208] Optionally, the device further includes:

[0209] The video acquisition module is used to acquire video, determine the frame rate and resolution of the video, and obtain the video frame sequence.

[0210] It should be noted that the apparatus provided in this application embodiment can implement all the method steps implemented in the method embodiment applied to the transmitting device, and can achieve the same technical effect. Here, the parts that are the same as those in the method embodiment and the beneficial effects will not be described in detail.

[0211] like Figure 12 As shown, this application embodiment also provides a video communication device 1200, applied to a receiving end device, including:

[0212] Receiver module 1210 is used to receive keyframe sequences;

[0213] Decoding module 1220 is used to perform source-channel joint decoding on the key frame sequence to obtain the reconstructed key frame sequence;

[0214] The video frame interpolation module 1230 is used to perform video frame interpolation on the reconstructed keyframe sequence according to the number and position of keyframes in the reconstructed keyframe sequence to obtain a reconstructed video frame sequence.

[0215] The number and location of the keyframes are related to the characteristics of the video and the channel conditions.

[0216] Optionally, the video frame interpolation module is specifically used for:

[0217] The position and number of video interpolated frames are determined based on the position and number of keyframes in the reconstructed keyframe sequence.

[0218] Based on the position and number of the video interpolation frames, the reconstructed keyframe sequence is interpolated to obtain a reconstructed video frame sequence.

[0219] Optionally, the characteristics of the video include the video's frame rate and / or resolution;

[0220] The channel conditions include: channel bandwidth and / or signal-to-noise ratio;

[0221] The number and location of the keyframes are determined based on the transmission rate and channel conditions; the transmission rate is calculated based on the frame rate and resolution of the video.

[0222] Optionally, the decoding module is specifically used for:

[0223] The keyframe sequence is subjected to joint source-channel decoding to obtain a reconstructed semantic feature map;

[0224] Semantic recovery is performed on the reconstructed semantic feature map to obtain the reconstructed keyframe sequence.

[0225] Optionally, the device further includes:

[0226] The display module is used to play the reconstructed video frame sequence according to the video's frame rate.

[0227] It should be noted that the apparatus provided in this application embodiment can implement all the method steps implemented in the method embodiment applied to the receiving end device, and can achieve the same technical effect. Here, the parts that are the same as those in the method embodiment and the beneficial effects will not be described in detail.

[0228] like Figure 13 As shown in the embodiment of this application, an electronic device may be a transmitting device or a receiving device, including a transceiver 1310, a processor 1300, a memory 1320, and a program or instructions stored in the memory and executable on the processor; when the processor executes the program or instructions, it implements the steps of the video communication method described above, and the transceiver 1310 is used to receive and send data under the control of the processor 1300.

[0229] Wherein, if the electronic device is a transmitting device, the processor 1300 is used for:

[0230] Based on the characteristics of the video and the channel conditions, key frames are selected from the video frame sequence to be transmitted to obtain a key frame sequence.

[0231] The keyframe sequence is jointly coded from the source and the channel.

[0232] The transceiver 1310 is used to: send the encoded keyframe sequence to the receiving device.

[0233] Optionally, the processor 1300 selects key frames from the video frame sequence to be transmitted based on the characteristics of the video and channel conditions to obtain a key frame sequence, including:

[0234] The transmission rate is determined based on the characteristics of the video.

[0235] The number and location of keyframes are determined based on the transmission rate and channel conditions.

[0236] Keyframes are selected from the video frame sequence based on the quantity and position to obtain a keyframe sequence.

[0237] Optionally, the characteristics of the video include: the video's frame rate and / or resolution;

[0238] The processor 1300 determines the transmission rate based on the characteristics of the video, including:

[0239] Calculate the transmission rate based on the video's frame rate, resolution, the inter-frame compression factor for keyframe selection, and the intra-frame compression factor for encoding.

[0240] Optionally, the channel conditions include: channel bandwidth and / or signal-to-noise ratio;

[0241] The processor 1300 determines the number and location of keyframes based on the transmission rate and channel conditions, including:

[0242] Based on the transmission rate and channel capacity, determine the minimum inter-frame compression factor for selecting key frames, wherein the transmission rate is less than or equal to the channel capacity, and the channel capacity is calculated based on the channel bandwidth and signal-to-noise ratio;

[0243] The number of keyframes is calculated based on the minimum inter-frame compression factor.

[0244] The position of the keyframe in the video frame sequence is determined based on the number of keyframes and the preset sampling method.

[0245] Optionally, the processor 1300 performs source-channel joint coding on the keyframe sequence, including:

[0246] Semantic extraction is performed on the keyframe sequence to obtain the semantic feature map of the keyframe sequence;

[0247] The semantic feature map is jointly encoded by the source and channel to obtain the encoded sequence.

[0248] Optionally, the processor 1300 is further configured to:

[0249] The encoded sequence is quantized;

[0250] The transceiver is specifically used for:

[0251] Send the quantized encoded sequence to the receiving device.

[0252] Optionally, the processor 1300 is further configured to: perform video acquisition, determine the video frame rate and resolution, and obtain a video frame sequence.

[0253] When the electronic device is a receiving device, the transceiver 1310 is used to: receive a keyframe sequence;

[0254] The processor 1300 is used to: perform source-channel joint decoding on the keyframe sequence to obtain a reconstructed keyframe sequence;

[0255] Based on the number and position of the keyframes in the reconstructed keyframe sequence, video frame interpolation is performed on the reconstructed keyframe sequence to obtain a reconstructed video frame sequence;

[0256] The number and location of the keyframes are related to the characteristics of the video and the channel conditions.

[0257] Optionally, the processor 1300 performs video frame interpolation on the reconstructed keyframe sequence based on the number and position of keyframes in the reconstructed keyframe sequence to obtain a reconstructed video frame sequence, including:

[0258] The position and number of video interpolated frames are determined based on the position and number of keyframes in the reconstructed keyframe sequence.

[0259] Based on the position and number of the video interpolation frames, the reconstructed keyframe sequence is interpolated to obtain a reconstructed video frame sequence.

[0260] Optionally, the characteristics of the video include the video's frame rate and / or resolution;

[0261] The channel conditions include: channel bandwidth and / or signal-to-noise ratio;

[0262] The number and location of the keyframes are determined based on the transmission rate and channel conditions; the transmission rate is calculated based on the frame rate and resolution of the video.

[0263] Optionally, the processor 1300 performs joint source-channel decoding on the keyframe sequence to obtain a reconstructed keyframe sequence, including:

[0264] The keyframe sequence is subjected to joint source-channel decoding to obtain a reconstructed semantic feature map;

[0265] Semantic recovery is performed on the reconstructed semantic feature map to obtain the reconstructed keyframe sequence.

[0266] Optionally, the processor 1300 is further configured to:

[0267] The reconstructed video frame sequence is played according to the video's frame rate.

[0268] Among them, Figure 13 In this context, the bus architecture may include any number of interconnected buses and bridges, specifically linking various circuits together, represented by one or more processors (processor 1300) and memory (memory 1320). The bus architecture may also link together various other circuits such as peripheral devices, voltage regulators, and power management circuits, which are well known in the art and therefore will not be described further herein. The bus interface provides an interface. The transceiver 1310 may be multiple elements, including transmitters and receivers, providing a unit for communicating with various other devices over a transmission medium. The processor 1300 is responsible for managing the bus architecture and general processing, and the memory 1320 may store data used by the processor 1300 during operation.

[0269] An embodiment of this application provides a readable storage medium storing a program or instructions. When the program or instructions are executed by a processor, they implement the steps in the video communication method described above and achieve the same technical effect. To avoid repetition, further details are omitted here.

[0270] The processor is the processor in the electronic device described in the above embodiments. The readable storage medium includes computer-readable storage media, such as computer read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk.

[0271] This application also provides a computer program product that stores a program or instructions, including computer instructions. When the computer instructions are executed by a processor, they implement the steps in the video communication method described above and achieve the same technical effect. To avoid repetition, they will not be described again here.

[0272] It should be further noted that the terminal devices described in this specification include, but are not limited to, smartphones, tablets, etc., and many of the described functional components are referred to as modules in order to more specifically emphasize the independence of their implementation.

[0273] In this embodiment, the module can be implemented in software so that it can be executed by various types of processors. For example, an identified executable code module may include one or more physical or logical blocks of computer instructions, which may be constructed as objects, procedures, or functions. Nevertheless, the executable code of the identified module does not need to be physically located together, but may include different instructions stored in different bits, which, when logically combined, constitute the module and achieve the module's intended purpose.

[0274] In practice, an executable code module can be a single instruction or many instructions, and can even be distributed across multiple different code segments, different programs, and across multiple memory devices. Similarly, operational data can be identified within the module and can be implemented in any suitable form and organized within any suitable type of data structure. This operational data can be collected as a single dataset or distributed across different locations (including different storage devices), and can exist, at least in part, solely as electronic signals within the system or network.

[0275] When a module can be implemented using software, considering the current level of hardware technology, modules that can be implemented in software can be implemented using hardware circuits by those skilled in the art to achieve the corresponding functions, without considering cost. These hardware circuits include conventional very-large-scale integrated circuits (VLSI) or gate arrays, as well as existing semiconductors such as logic chips and transistors, or other discrete components. Modules can also be implemented using programmable hardware devices, such as field-programmable gate arrays, programmable array logic, and programmable logic devices.

[0276] The exemplary embodiments described above are with reference to the accompanying drawings. Many different forms and embodiments are feasible without departing from the spirit and teachings of this application. Therefore, this application should not be construed as limiting the exemplary embodiments set forth herein. Rather, these exemplary embodiments are provided to make this application complete and convey the scope of this application to those skilled in the art. In these drawings, component dimensions and relative dimensions may be exaggerated for clarity. The terminology used herein is for the purpose of describing particular exemplary embodiments only and is not intended to be limiting. As used herein, unless clearly indicated otherwise, the singular forms “a,” “an,” and “the” are intended to include all such forms. It will be further understood that the terms “comprising” and / or “including”, when used in this specification, indicate the presence of the stated features, integers, steps, operations, components, and / or elements, but do not exclude the presence or addition of one or more other features, integers, steps, operations, components, and / or groups thereof. Unless otherwise indicated, when stated, a range of values ​​includes the upper and lower limits of the range and any subranges in between.

[0277] The above description is the preferred embodiment of this application. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principles described in this application, and these improvements and modifications should also be considered within the scope of protection of this application.

Claims

1. A video communication method, characterized in that, Applied to transmitting devices, including: Calculate the transmission rate based on the video's frame rate, resolution, the inter-frame compression factor of the selected keyframes, and the intra-frame compression factor of the encoding. Based on the transmission rate and channel capacity, determine the minimum inter-frame compression factor for selecting key frames, wherein the transmission rate is less than or equal to the channel capacity, and the channel capacity is calculated based on the channel bandwidth and signal-to-noise ratio; The number of keyframes is calculated based on the minimum inter-frame compression factor. The position of the keyframe in the video frame sequence is determined based on the number of keyframes and the preset sampling method; Based on the quantity and position, key frames are selected from the video frame sequence to be transmitted to obtain a key frame sequence; The keyframe sequence is jointly coded from the source and the channel. The encoded keyframe sequence is sent to the receiving device.

2. The method according to claim 1, characterized in that, The keyframe sequence is subjected to joint source-channel coding, including: Semantic extraction is performed on the keyframe sequence to obtain the semantic feature map of the keyframe sequence; The semantic feature map is jointly encoded by the source and channel to obtain the encoded sequence.

3. The method according to claim 2, characterized in that, Before sending the encoded keyframe sequence to the receiving device, the method further includes: The encoded sequence is quantized; Sending the encoded keyframe sequence to the receiving device includes: Send the quantized encoded sequence to the receiving device.

4. The method according to claim 1, characterized in that, Before selecting keyframes from the video frame sequence to be transmitted based on its characteristics and channel conditions to obtain a keyframe sequence, the method further includes: Video capture is performed to determine the video's frame rate and resolution, thereby obtaining a video frame sequence.

5. A video communication method, characterized in that, Applied to receiving devices, including: Receive keyframe sequence; The keyframe sequence is subjected to joint source-channel decoding to obtain the reconstructed keyframe sequence; Based on the number and position of the keyframes in the reconstructed keyframe sequence, video frame interpolation is performed on the reconstructed keyframe sequence to obtain a reconstructed video frame sequence; The keyframe sequence is obtained by selecting keyframes from the video frame sequence to be transmitted based on the number and position of the keyframes; the number of keyframes is determined by the minimum inter-frame compression factor for selecting keyframes based on the transmission rate and channel capacity, and is calculated based on the minimum inter-frame compression factor, wherein the transmission rate is less than or equal to the channel capacity, and the channel capacity is calculated based on the channel bandwidth and signal-to-noise ratio; the position of the keyframes is determined based on the number of keyframes and the preset sampling method; the transmission rate is calculated based on the video frame rate, resolution, the inter-frame compression factor of the selected keyframes, and the intra-frame compression factor of the encoding.

6. The method according to claim 5, characterized in that, Based on the number and position of keyframes in the reconstructed keyframe sequence, video frame interpolation is performed on the reconstructed keyframe sequence to obtain a reconstructed video frame sequence, including: The position and number of video interpolated frames are determined based on the position and number of keyframes in the reconstructed keyframe sequence. Based on the position and number of the video interpolation frames, the reconstructed keyframe sequence is interpolated to obtain a reconstructed video frame sequence.

7. The method according to claim 5, characterized in that, The keyframe sequence is subjected to joint source-channel decoding to obtain a reconstructed keyframe sequence, including: The keyframe sequence is subjected to joint source-channel decoding to obtain a reconstructed semantic feature map; Semantic recovery is performed on the reconstructed semantic feature map to obtain the reconstructed keyframe sequence.

8. The method according to claim 5, characterized in that, The method further includes: The reconstructed video frame sequence is played according to the video's frame rate.

9. A video communication device, characterized in that, include: The keyframe selection module is used to calculate the transmission rate based on the video's frame rate, resolution, the inter-frame compression factor of the selected keyframes, and the intra-frame compression factor of the encoding. Based on the transmission rate and channel capacity, a minimum inter-frame compression factor for selecting keyframes is determined, wherein the transmission rate is less than or equal to the channel capacity, which is calculated based on channel bandwidth and signal-to-noise ratio; the number of keyframes is calculated based on the minimum inter-frame compression factor; the position of the keyframes in the video frame sequence is determined based on the number of keyframes and a preset sampling method; and keyframes are selected from the video frame sequence to be transmitted based on the number and position to obtain a keyframe sequence. The encoding module is used to perform source-channel joint coding on the keyframe sequence; A wireless transmission module is used to send the encoded keyframe sequence to the receiving device.

10. A video communication device, characterized in that, include: The receiving module is used to receive keyframe sequences; The decoding module is used to perform joint source-channel decoding on the keyframe sequence to obtain the reconstructed keyframe sequence; The video frame interpolation module is used to perform video frame interpolation on the reconstructed keyframe sequence according to the number and position of keyframes in the reconstructed keyframe sequence to obtain a reconstructed video frame sequence. The keyframe sequence is obtained by selecting keyframes from the video frame sequence to be transmitted based on the number and position of the keyframes; the number of keyframes is determined by the minimum inter-frame compression factor for selecting keyframes based on the transmission rate and channel capacity, and is calculated based on the minimum inter-frame compression factor, wherein the transmission rate is less than or equal to the channel capacity, and the channel capacity is calculated based on the channel bandwidth and signal-to-noise ratio; the position of the keyframes is determined based on the number of keyframes and the preset sampling method; the transmission rate is calculated based on the video frame rate, resolution, the inter-frame compression factor of the selected keyframes, and the intra-frame compression factor of the encoding.

11. An electronic device, characterized in that, It includes a transceiver, a processor, a memory, and a program or instructions stored in the memory and executable on the processor; when the processor executes the program or instructions, it implements the steps of the video communication method as described in any one of claims 1 to 4, or implements the steps of the video communication method as described in any one of claims 5 to 8.

12. A readable storage medium having a program or instructions stored thereon, characterized in that, When the program or instructions are executed by the processor, they implement the steps of the video communication method as described in any one of claims 1 to 4, or the steps of the video communication method as described in any one of claims 5 to 8.

13. A computer program product having a program or instructions stored thereon, characterized in that, It includes computer instructions that, when executed by a processor, implement the steps of the video communication method as described in any one of claims 1 to 4, or implement the steps of the video communication method as described in any one of claims 5 to 8.