Video stream transmission method and device, and storage medium
Edge computing and server-side processing with noise reduction and super-resolution techniques address video streaming issues, improving resolution and reducing frame loss for enhanced user experience.
Patent Information
- Application Number
- CN202510657664.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-21
- Publication Date
- 2025-07-15
AI Technical Summary
Existing video streaming technologies face issues with video image resolution discrepancies and frame loss due to the mismatch between video capture device updates and network bandwidth, leading to a degraded user experience.
Implementing edge computing and server-side processing to enhance video resolution and reduce frame loss through noise reduction, compression, super-resolution enhancement, and frame interpolation, utilizing lightweight neural networks and adaptive transmission protocols.
Improves video image resolution and reduces frame loss, enhancing the overall video playback quality and user experience in low-bandwidth environments.
Smart Images

Figure CN120321436A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of video transmission technologies. Specifically, this application relates to a video stream transmission method, device, and storage medium. Background Art
[0002] With the development of video acquisition, transmission, storage, and display technologies, videos are constantly evolving towards higher resolutions. People's requirements for videos are also getting higher and higher, constantly pursuing high-resolution and high-definition video images. At the same time, the emergence of high-resolution display devices (such as 4K and 5K TVs and monitors) has made the popularization of high-resolution videos possible.
[0003] However, the update of video acquisition devices and the change of network bandwidth often cannot keep up with the change of people's requirements for videos, resulting in problems such as poor resolution and frame loss in video images in the video stream, greatly reducing the user's video viewing experience. Summary of the Invention
[0004] Embodiments of this application provide a video stream transmission method, device, and storage medium, which can solve the problems of poor resolution and frame loss that easily occur in video images in the existing video stream. To achieve this purpose, the embodiments of this application provide the following several solutions.
[0005] According to one aspect of the embodiments of this application, a video stream transmission method for an edge computing end or a server end is provided. The method includes:
[0006] Receiving a transmitted first video stream, the acquisition process of the first video stream includes: the front end preprocessing the original video stream, and the preprocessing includes at least one of noise reduction and compression;
[0007] Performing super-resolution enhancement and frame interpolation processing on the first video stream to obtain a target video stream;
[0008] Transmitting the target video stream to the video receiving object based on the current network bandwidth and / or the video receiving object to which it is connected, and the transmission includes any one of hierarchical transmission and overall transmission.
[0009] In a possible implementation manner, the preprocessing includes noise reduction processing. The front end preprocessing the original video stream includes:
[0010] Performing noise reduction processing on the original video stream to obtain a first video stream;
[0011] The noise reduction processing includes:
[0012] The noise reduction model is used to perform noise reduction processing on each frame of the original video stream to obtain a noise-reduced image. The noise reduction model is trained based on a lightweight convolutional neural network, and the number of parameters of the lightweight convolutional neural network is less than a first preset value, and the type and structure of the lightweight convolutional neural network correspond to the scene where the front end is located.
[0013] In a possible implementation, the preprocessing includes compression processing. The front end preprocesses the original video stream, including:
[0014] Identify key points in the noise-reduced image, obtain region category information of the noise-reduced image according to the key points, compress the noise-reduced image according to the region category information to obtain a first video stream, and embed reconstruction data in the first video stream. The region category information includes important regions and background regions, and the reconstruction data is used for super-resolution enhancement.
[0015] In a possible implementation, the identifying key points in the noise-reduced image includes:
[0016] Use a key point recognition model to identify key points in the noise-reduced image. The architecture of the key point recognition model is a hybrid architecture of MobileNetV3 + LSTM. The key point recognition model extracts spatial features through the MobileNetV3 architecture and uses LSTM to process and capture temporal continuity.
[0017] In a possible implementation, the receiving the first video stream includes:
[0018] The server side receives the first video stream transmitted by the edge computing side. The first video stream is obtained by performing super-resolution enhancement processing and frame interpolation processing on the second video stream of the edge computing side, and the second video stream is the original video stream preprocessed by the front end.
[0019] In a possible implementation, the edge computing side generates the first video stream, including:
[0020] Obtain the second video stream after super-resolution enhancement processing, and use an interpolation model to perform frame interpolation processing on the second video stream after super-resolution enhancement processing to obtain a frame interpolation video stream;
[0021] Extract compact data from the frame interpolation video stream, and determine the compact data as the first video stream. The compact data includes key point data and audio data.
[0022] In a possible implementation, performing frame interpolation processing on the first video stream includes:
[0023] Determine the first video stream after super-resolution enhancement as the video stream to be frame-inserted, and determine the positions of the lost frames according to the frame information of each frame image in the video stream to be frame-inserted;
[0024] Obtain the frame prediction information of the lost frames according to the adjacent frames corresponding to the positions, and generate the lost frames and insert the lost frames into the video stream to be frame-inserted based on the frame prediction information. The frame prediction information includes at least one of motion trend, light change, background consistency, key point coordinates, key point displacement trajectory, expression parameters, and body movement information;
[0025] Perform lip synchronization processing and expression and action adjustment on the lost frames in the video stream to be frame-inserted.
[0026] In a possible implementation manner, the transmitting the target video stream to the video receiving object based on the current network bandwidth and / or the video receiving object connected to itself includes:
[0027] If it is determined that the video receiving object connected to itself is a client, obtain the hierarchical division information of the target video stream, where the hierarchical division information includes a base layer and an enhancement layer;
[0028] Dynamically transmit the target video stream according to the network bandwidth and the hierarchical division information, and the transmission protocol of the target video stream includes the WebRTC / QUIC protocol.
[0029] According to one aspect of the embodiments of the present application, the embodiments of the present application provide an electronic device, including a memory, a processor, and a computer program stored on the memory. The processor executes the computer program to implement the steps of any of the above methods.
[0030] According to one aspect of the embodiments of the present application, the embodiments of the present application provide a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed, the steps of the above method are implemented.
[0031] The beneficial effects brought by the technical solutions provided by the embodiments of the present application are:
[0032] The video stream transmission method provided by the present application receives the transmitted first video stream, and the acquisition process of the first video stream includes preprocessing of the original video stream; performing super-resolution enhancement and frame interpolation processing on the first video stream to obtain a target video stream, and transmitting the target video stream to the video receiving object based on the current network bandwidth and / or the video receiving object connected to itself; the embodiments of the present application can perform super-resolution enhancement and frame interpolation processing on the video stream during the video stream transmission process, improve the resolution of the video images in the video stream received by the client and reduce video frame loss, effectively improve the video playback effect, and enhance the user experience. Description of the Drawings
[0033] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the accompanying drawings required for the description of the embodiments of the present application.
[0034] Figure 1 It is a flowchart of the video stream transmission method provided by the embodiments of the present application;
[0035] Figure 2 It is a structural diagram of the video transmission system provided by the embodiments of the present application;
[0036] Figure 3 It is a structural diagram of the electronic device provided by the embodiments of the present application. Detailed implementation manners
[0037] The following describes the embodiments of the present application in combination with the accompanying drawings in the present application. It should be understood that the implementation manners described below in combination with the drawings are exemplary descriptions for explaining the technical solutions of the embodiments of the present application, and do not constitute a limitation to the technical solutions of the embodiments of the present application.
[0038] Those skilled in the art of the present technology can understand that, unless specifically stated otherwise, the singular forms "a", "an", "the", and "said" used herein may also include the plural forms. It should be further understood that the terms "including" and "comprising" used in the embodiments of the present application mean that the corresponding features can be implemented as the presented features, information, data, steps, operations, elements, and / or components, but do not exclude being implemented as other features, information, data, steps, operations, elements, components, and / or their combinations supported by the art of the present technology. It should be understood that when we say an element is "connected" or "coupled" to another element, this element can be directly connected or coupled to the other element, or it can mean that this element and the other element establish a connection relationship through an intermediate element. In addition, the "connection" or "coupling" used here can include wireless connection or wireless coupling. The term "and / or" used here indicates at least one of the items defined by the term, for example, "A and / or B" indicates being implemented as "A", or being implemented as "A", or being implemented as "A and B".
[0039] To make the purpose, technical solutions, and advantages of the present application clearer, the following will further describe the embodiments of the present application in detail in combination with the drawings.
[0040] The following illustrates the technical solutions of the embodiments of the present application and the technical effects produced by the technical solutions of the present application through the description of several exemplary implementation manners. It should be noted that the following implementation manners can be referred to, learned from, or combined with each other. For the same terms, similar features, and similar implementation steps in different implementation manners, they will not be described repeatedly.
[0041] The video stream transmission method, device, and storage medium provided by this application aim to solve at least one technical problem existing in the prior art.
[0042] Optionally, the device for executing the video stream transmission method of this application can be an edge computing terminal or a server terminal. The edge computing terminal or the server terminal performs super-resolution enhancement and frame interpolation processing on the transmitted video stream during the video stream transmission process, improving the resolution of the video image and reducing the phenomenon of lost frames.
[0043] Optionally, as Figure 1 、 Figure 2 shown, the edge computing terminal and the server terminal can be set in a video transmission system. The edge computing terminal is respectively connected to the front end and the server terminal, and moreover, the server terminal is connected to the client terminal. The server terminal transmits the video stream after super-resolution enhancement and frame interpolation processing to the client terminal, improving the video playback quality of the client terminal.
[0044] Optionally, the video stream transmission method of this application includes:
[0045] S101: Receive the transmitted first video stream
[0046] Optionally, the acquisition process of the first video stream includes: The front end preprocesses the original video stream. When the execution object of the method of this application is the edge computing terminal, the preprocessed original video stream is the first video stream, and the front end transmits the first video stream to the edge computing terminal. When the execution object of the method of this application is the server terminal, the preprocessed original video stream is the second video stream. After the edge computing terminal performs super-resolution enhancement and frame interpolation on the second video stream, the first video stream is obtained, and the first video stream is transmitted to the server terminal.
[0047] Optionally, the front end can be an object for data acquisition and generation of the original video stream, such as a video acquisition tool like a camera, a video conferencing terminal, etc., or it can also be an object connected to multiple video acquisition tools and used for processing the original video stream transmitted by the video acquisition tools, such as a console connected to multiple cameras.
[0048] Optionally, the preprocessing includes at least one of noise reduction and compression. Among them, when the preprocessing is noise reduction processing, the front end preprocesses the original video stream, including: performing noise reduction processing on the original video stream to obtain the first video stream; the noise reduction processing includes: using a noise reduction model to perform noise reduction processing on each frame image in the original video stream to obtain a noise-reduced image, and the noise reduction model is trained based on a lightweight convolutional neural network, and the number of parameters of the lightweight convolutional neural network is less than a first preset value and the type and structure of the lightweight convolutional neural network correspond to the scene where the front end is located.
[0049] In one embodiment, the noise reduction model can be a dynamic noise reduction model based on deep learning. Among them, the lightweight convolutional neural network can be DnCNN or CBDNet, and this lightweight convolutional neural network is used for noise modeling and removal. The first preset value can be 2MB to meet the real-time processing requirements of the front end, and the real-time processing requirements can be that the latency when obtaining the noise-reduced image is less than 5ms.
[0050] Optionally, the feature extraction layer (convolutional layer) of the noise reduction model can be used to perform multi-scale feature extraction on the video image to identify the noise and details in the video image and obtain the type of noise. The noise reduction model outputs the noise-reduced image through residual learning for the extracted noise and noise type. The noise type can include transmission noise, compression noise, and sensor noise. When performing noise processing, different noise removal methods can be adopted for different types of noise.
[0051] Optionally, to retain the high-frequency details in the image (such as pupils and hair), a loss function of Perceptual Loss can also be added to the lightweight convolutional neural network to avoid excessive smoothing.
[0052] Optionally, to adapt to the low-bandwidth environment, a compression noise simulation layer can be added during model training. This compression noise simulation layer is used to simulate the noise generated by video image compression, and this noise is used for model training to improve the denoising effect of the model. And when training the noise reduction model, random blocking artifacts and quantization noise are injected into the training data, and the model is trained using the training data with injected random effects and quantization noise, so as to improve the robustness of the model to transmission noise.
[0053] In one embodiment, the convolutional neural network used by the noise reduction model can be DnCNN-Lite, which has 5 convolutional layers (which can be 3*3 kernels). The number of channels of the 5 convolutional layers decreases layer by layer (the change in the number of channels can be 64→32→16→8→3), and the last layer outputs the noise-reduced image. The activation function of the noise reduction model can be PReLU to reduce parameter redundancy. Compared with the case where BM3D takes 100ms / frame for noise reduction, the speed of nCNN-Lite can reach 5ms / frame, with a 20-fold speed increase, and it can effectively retain the details in the image.
[0054] Optionally, when the preprocessing is compression processing, the front end preprocesses the original video stream, including: identifying the key points in the noise-reduced image, obtaining the region category information of the noise-reduced image according to the key points, compressing the noise-reduced image according to the region category information to obtain the first video stream and embedding the reconstruction data in the first video stream. The region category information includes important regions and background regions, and the reconstruction data is used for super-resolution enhancement.
[0055] In one embodiment, the key points can be predefined key points in video images such as human faces and gestures. The important regions and background regions in the video image are identified through key point detection, and different compression algorithms are used for compression processing of the important regions and background regions. Among them, the QP (Quantization Parameter) parameter related to the compression rate of the important region can be 40, and the QP parameter related to the compression rate of the background region can be 22. Compared with the bit rate when using H.264 global compression, this method can increase the bit rate by 30% and improve the PSNR (Peak Signal-to-Noise Ratio) of the key region by 4 dB.
[0056] Optionally, the reconstructed data can be the metadata of the super-resolution model (such as texture feature vectors). By using this metadata, the efficiency of super-resolution enhancement after image decoding can be improved, thereby obtaining a high-definition picture.
[0057] In one embodiment, the front end is connected to the edge computing end. The edge computing end uses a lightweight super-resolution enhancement model, and the front end adds the metadata of the lightweight super-resolution enhancement model to the first video stream for the edge computing end to reconstruct a high-definition picture after decoding. By using different compression methods for different regions and embedding super-resolution information (reconstructed data) in the first video stream, "low-bitrate transmission and high-definition restoration" are achieved.
[0058] Optionally, identifying the key points in the denoised image includes: using a key point recognition model to identify the key points in the denoised image. The architecture of the key point recognition model is a hybrid architecture of MobileNetV3 + LSTM. The key point recognition model extracts spatial features through the MobileNetV3 architecture and uses LSTM to process and capture temporal continuity. The motion trajectory of the key points is predicted by LSTM to reduce jitter. Among them, the region where the number or density of the key points is greater than or equal to the preset threshold can be determined as the important region, and the region where the number or density of the key points is less than the preset threshold can be determined as the background region.
[0059] In one embodiment, to reduce the computational amount, the number of parameters of the key point recognition model can be compressed to less than 1 MB and support real-time detection at 30 fps on the front end. Moreover, to improve the robustness of the key point recognition model, low-resolution and high-noise images can be used to train the key point recognition model to improve the robustness of the model in a weak network environment. Compared with the MediaPipe model, the detection accuracy of the key point recognition model obtained in this way is improved by 15%. Among them, the LSTM temporal calculation can also be allocated to the CPU, and the MobileNetV3 inference can be allocated to the GPU to achieve GPU-CPU heterogeneous scheduling and reduce latency.
[0060] Optionally, in the key point recognition model, MobileNetV3 can be used as the backbone network. To reduce the number of channels, the width coefficient of this backbone network can be 0.5. The LSTM layer in the model captures the key point movement trajectories between consecutive frames (such as lip shape changes), and precise face key points and gesture coordinates are output through this model. The number of these key points can be 68.
[0061] Optionally, the front end can determine the sending object of the preprocessed original video stream according to its communication status with the edge computing. Among them, after determining that the edge computing end is busy (such as the feedback time of the edge computing end is greater than the preset time or the edge computing end feedbacks that it cannot process the video stream currently) or the edge computing end has no response, the front end will send the preprocessed video stream to the server end; after determining that the edge computing end feedbacks that it can process, the front end can send the preprocessed video stream to the edge computing end.
[0062] Optionally, to meet the edge computing requirements in low-bandwidth environments, the model is lightweighted through the following technologies:
[0063] For the denoising model, redundant channels can be pruned through importance evaluation (such as L1 norm), and the parameter quantity of DnCNN is compressed from 5MB to 2MB.
[0064] For the key point recognition model, depthwise separable convolutions can be used in MobileNetV3 in the model to replace standard convolutions, and the computational amount is reduced to 1 / 8. Moreover, knowledge distillation processing can also be performed on the key point recognition model. For example, a high-precision model in the cloud (such as HRNet) is used as the teacher model to guide the training of the key point recognition model on the edge side (MobileNetV3-LSTM). During the training process, the KL divergence loss function is used to make the student model imitate the output distribution of the teacher model. Thus, the accuracy of the key point recognition model can be improved by 12%, and the parameter quantity is only 1MB. Compared with HRNet with a parameter quantity of 10MB, the key point recognition model of this application can reduce jitter by 50% and meet the real-time requirements on the edge side.
[0065] Through the lightweight processing of the above denoising model and key point recognition model, the total parameter quantity of the edge-side model can be less than 5MB, and the memory occupancy is less than 50MB, so that it can run smoothly on a 1GB RAM device. The full-process inference speed on the edge side is less than <30ms (meeting the requirement of real-time processing at 30fps). For key point recognition, at a 480p input, the accuracy reaches 94.5% (the traditional model HRNet is 82%).
[0066] S102: Perform super-resolution enhancement and frame interpolation processing on the first video stream to obtain the target video stream.
[0067] Optionally, when the object executing the method of this application is an edge computing terminal, the edge computing terminal is equipped with a lightweight super-resolution enhancement model. The edge computing terminal is connected to the front end, receives the first video stream transmitted by the front end, and transmits the first video stream to the lightweight super-resolution enhancement model. The lightweight super-resolution enhancement model can extract reconstruction data from the first video stream and restore high-definition details based on the reconstruction data.
[0068] In one embodiment, the lightweight super-resolution enhancement model can be Real-ESRGAN. To improve the inference speed, the model weights can be converted from FP32 to INT8, reducing the memory footprint by 75% and increasing the inference speed by 3 times. Also, keep the weights of the key layers (such as residual blocks) in the model as FP16 to balance accuracy and efficiency. The PSNR of the video images in the second video stream obtained by super-resolution is 28.6dB.
[0069] Optionally, the super-resolution enhancement model adapted to the edge computing terminal can also be replaced according to the network environment. Among them, if it is determined that the current is a weak network environment (such as the network speed is less than the preset speed), the model running on the edge computing terminal can be a lightweight model (such as Fast-SRNet); if it is determined that the network condition is good or the network condition improves (such as the network is greater than the preset speed), it can be switched to a high-precision model installed in the cloud (such as Real-ESRGAN), and incremental data is transmitted through differential coding.
[0070] Optionally, when using ARM chips at the edge computing terminal and the front end, the NEON instruction set can also be used for acceleration, such as optimizing convolution calculations for ARM chips to increase the end-side inference speed by 40%.
[0071] Optionally, the edge computing terminal generates the first video stream, including: obtaining the second video stream after super-resolution enhancement processing, performing frame interpolation processing on the second video stream after super-resolution enhancement processing using an interpolation model to obtain a frame interpolation video stream; extracting compact data from the frame interpolation video stream, and determining the compact data as the first video stream, where the compact data includes key point data and audio data.
[0072] Optionally, both the edge computing terminal and the server terminal can perform super-resolution enhancement and frame interpolation operations, reducing the performance requirements for the terminal and improving the operation effect by performing these two operations on both terminals.
[0073] In one embodiment, the process of the edge computing terminal and the server terminal performing angular resolution enhancement and frame interpolation operations can be the same, such as using the same method for super-resolution enhancement, only the number of parameters of the models used is different.
[0074] Optionally, the edge computing side and the server side perform frame interpolation processing in the same way. Among them, the server side performs frame interpolation processing on the first video stream, including: determining the super-resolution enhanced first video stream as the video stream to be interpolated, and determining the positions of the lost frames according to the frame information of each frame image in the video stream to be interpolated; obtaining the frame prediction information of the lost frames according to the adjacent frames corresponding to the positions, and generating the lost frames and inserting the lost frames into the video stream to be interpolated based on the frame prediction information. The frame prediction information includes at least one of motion trend, light change, background consistency, key point coordinates, key point displacement trajectory, expression parameters, and body movement information; performing lip synchronization processing and expression and movement adjustment on the lost frames in the video stream to be interpolated.
[0075] Optionally, the frame information can be the timestamp and frame number of each frame image, and the positions of the lost frames are identified through the timestamp and frame number. For example, the position of the lost frame n+1 is identified by using the nth frame and the (n+2)th frame.
[0076] In one embodiment, context modeling can be performed using the adjacent frames of the lost frames. For example, the motion trend, light change, and background consistency are extracted from the valid frames before and after the lost frame n+1 (such as the nth frame and the (n+2)th frame) to provide a reference for prediction. For the information related to the key points, key point detection can be performed first, and then the displacement trajectory of the key points can be obtained. Among them, when a human face is included in the video image, MediaPipe Face Mesh or OpenPose can be used to track 68 key points (lips, eyes, eyebrows, etc.) of the human face in real time, output the key point coordinates and confidence of each frame, and determine the key point coordinates corresponding to the key points by determining the key point coordinates with a confidence greater than the confidence threshold. Based on the determined key point coordinates, the (Recurrent All-Pairs Field Transforms) or FlowNet model is used to calculate the pixel-level motion vector between adjacent frames. The key point displacement trajectory in the lost frame (such as the opening and closing amplitude of the lips, the rotation angle of the head) is predicted according to the vector. The predicted image of the lost frame is generated according to the motion trend, light change, background consistency, and the key point displacement trajectory in the lost frame.
[0077] Optionally, after inserting the lost frames, the speech features of the lost frames (such as phonemes, syllable duration) can be extracted by combining the lost frames, the adjacent frames of the lost frames, and the audio stream in the first video stream. The Wav2Lip or LipGAN model is used to generate the corresponding lip animation according to the speech features, and the lip animation is fused with the predicted lost frames to achieve lip synchronization.
[0078] Optionally, expression and action adjustments can be made according to expression parameters and body movement information. Among them, an LSTM or Transformer model can be used to analyze the expression changes (such as smile → frown) between the frames before and after the lost frame and predict the expression parameters of the lost frame (such as the AU value of the Facial Action Coding System FACS). Based on this expression parameter, the expression of the lost frame is adjusted. When making action adjustments, the full-body actions (such as gestures, limb movements) can be predicted based on human pose estimation (such as AlphaPose) and the frames before and after the lost frame. A GAN is used to generate natural-looking details of the action (such as the amplitude of arm swing) in the lost frame.
[0079] Optionally, to achieve lightweight and real-time optimization of the model in the edge computing terminal, channel pruning and quantization can be performed on the frame interpolation model (RIFE), and the number of parameters is compressed from 50MB to 5MB. TensorRT is used to accelerate inference, and the end-side latency is <10ms / frame. Also, the cooperation between the edge computing terminal and the server can be controlled.
[0080] Optionally, the frame interpolation operation can also be divided into multiple parts. The edge computing terminal processes tasks with high real-time requirements, such as key point tracking and lightweight frame interpolation in frame interpolation, and the server processes complex tasks such as high-precision expression synthesis (expression parameter adjustment), multi-modal fusion (lip synchronization, action adjustment), etc.
[0081] Through the above operations, the image quality of the lost frame can be effectively improved. The lost frame obtained by the frame interpolation operation of this application is 3DB higher than the traditional difference, and the lip synchronization error is less than 50ms (less than the human face perception threshold). Also, at a 20% packet loss rate, users rate the smoothness as 4.1 / 5 (2.8 / 5 for the traditional method). Through the above technologies, the present invention realizes high-fidelity action restoration, natural expression transition, and precise audio-visual synchronization in a low-bandwidth environment, significantly improving the realism and user experience of video images.
[0082] Optionally, to reduce the data transmission volume between the edge computing terminal and the server, the edge computing terminal can only transmit compact data to reduce the overall transmission volume.
[0083] In one embodiment, when the video image includes a key face image, a lightweight model (such as MobileNetV3-LSTM) can be used to extract the coordinates of 68 key points of the face, expression parameters (AU values), and action vectors (such as head pose, gestures), and transmit the key point coordinates and action vectors to the server side, so that only about 1KB needs to be transmitted per frame (hundreds of KB for traditional video frames). For the audio data in the second video stream, low-bitrate Opus encoding can be used to transmit the encoded audio stream to the server side. The server side generates a matching lip animation in real time through the Wav2Lip model, avoiding the transmission of redundant video data.
[0084] Optionally, the edge computing terminal can also obtain the different parts between adjacent frames and only transmit the different parts to the server terminal. Specifically, the edge computing terminal can calculate the inter-frame motion vectors based on optical flow estimation (RAFT model) and only transmit the motion vectors and residual data, reducing the data volume by 70% (compared with full-frame transmission). It can also dynamically detect the changing areas in the picture (such as faces and gestures) and only update the pixel data of these areas, while keeping the background area static, realizing the incremental update of key areas.
[0085] In one embodiment, video frames with different resolutions can be generated through super-resolution enhancement by the edge computing terminal and the server terminal. Specifically, the edge computing terminal can be used to generate a 360p video frame, and then the server terminal and the 360p video frame can be used to generate a 1080p video frame. Among them, when generating lost frames at the edge computing terminal, a lightweight Fast-SRNet (with a parameter quantity of 1MB, which can be obtained by knowledge distillation of Real-ESRGAN) can be used to generate low-resolution lost frames (such as 360p), rather than high-resolution (1080p) lost frames. After transmission, they are enhanced by the super-resolution model (such as Real-ESRGAN) of the server terminal. The data volume can be reduced by 75% through the lightweight generation of lost frames. Moreover, the pixel values of the lost frames and other frame images in the second video stream can be 8-bit quantized, combined with H.265 encoding compression, reducing the bit rate by 40%.
[0086] Optionally, when the video stream is a video stream related to a video conference, there is temporal correlation between video frames. Directly performing super-resolution on a single frame may cause the picture to flicker or jitter. A temporal loss function guided by optical flow (such as Charbonnier loss, SSIN loss, etc.) can be introduced into the model for super-resolution enhancement to constrain the consistency of the super-resolution results of adjacent frames. A 3D convolutional layer or an LSTM module can also be used in the model to capture the inter-frame motion information and optimize the stability of dynamic scenes.
[0087] Optionally, to eliminate compression noise, during the training of the super-resolution model, simulated compression noise (such as JPEG / HEVC compression noise) can be injected into the training data related to the model to enhance the model's adaptability to real scenarios; alternatively, a denoising module (such as CBDNet) can be set up to denoise the received first video and audio, and by combining the denoising module with the super-resolution model, end-to-end compression-super-resolution joint optimization can be achieved. This application combines super-resolution enhancement with video temporal processing and compression noise removal to form a complete optimization chain for low-bandwidth video conferencing. In a simulated low-bandwidth environment (bitrate 500 kbps), the PSNR is increased by 6 dB (compared to directly using Real-ESRGAN), and the inter-frame jitter is reduced by 70% (subjective score 4.5 / 5).
[0088] Optionally, during the super-resolution enhancement process, TensorRT or ONNX Runtime can be used for inference acceleration, thereby reducing the single-frame processing time from 50 ms to 10 ms.
[0089] Optionally, the model for super-resolution enhancement can also be a combination of a CNN (Convolutional Neural Network) and a GAN (Generative Adversarial Network). Specifically, a generator is constructed based on the CNN, which converts a low-resolution (LR) image into a high-resolution (HR) image. The Residual-in-Residual Dense Block (RRDB) in its structure can be the core module of ESRGAN / Real-ESRGAN. This residual dense block extracts deep features through multi-layer dense connections and residual learning. The upsampling layers in the generator use subpixel convolution or deconvolution to gradually increase the resolution. The discriminator determines whether the image output by the generator is realistic, forcing the generator to generate a more realistic high-definition image. The PatchGAN in the discriminator discriminates the image in blocks to improve the authenticity of local details, and the discriminator stabilizes the training process through spectral normalization to prevent mode collapse. The loss function of the model can include a content loss function, an adversarial loss function, and a texture loss function. The content loss is based on pixel-level MSE or perceptual loss (VGG feature matching) to ensure the similarity between the generated image and the real image. The adversarial loss is fed back by the discriminator to enhance the visual realism of the generated image. The texture loss is introduced in Real-ESRGAN to optimize the generation of high-frequency details. Among them, when dealing with synthetic data or an ideal environment (such as laboratory videos), the ESRGAN is selected as the core module of the residual dense block. When dealing with real-scene noise (such as compression artifacts in video conferencing), Real-ESRGAN is selected.
[0090] Optionally, the generator can also use ESRGAN and Real-ESRGAN jointly. Specifically, in the first stage, Real-ESRGAN is used to eliminate compression noise and blocking artifacts, and an intermediate clear image is output. In the second stage, the intermediate clear image is input into ESRGAN to further restore texture and details, obtaining the super-resolution enhanced video frame image. During the model training process, synthetic data (ESRGAN) and real noise data (Real-ESRGAN) can be used simultaneously to construct a mixed dataset and train a more compatible model.
[0091] Optionally, in a scenario with sufficient resources on the server side, the super-resolution enhancement model composed of Real-ESRGAN can be preferentially used because of its higher robustness to real noise and its suitability for video conferencing scenarios with low-bandwidth transmission. For scenarios with high real-time requirements (such as the edge computing side), the lightweight improved Fast-RealESRGAN (parameter-compressed version) is adopted, sacrificing some accuracy to improve speed. For scenarios with extreme picture quality requirements (such as medical image enhancement), ESRGAN and Real-ESRGAN can be jointly used to process noise and details in stages.
[0092] S103: Transmit the target video stream to the video receiving object based on the current network bandwidth and / or the video receiving object to which it is connected.
[0093] Optionally, the transmission includes either hierarchical transmission or overall transmission. Transmitting the target video stream to the video receiving object based on the current network bandwidth and / or the video receiving object to which it is connected includes: if it is determined that the video receiving object to which it is connected is a client, obtain the hierarchical division information in the target video stream, and the hierarchical division information includes a base layer and an enhancement layer; dynamically transmit the target video stream according to the network bandwidth and the hierarchical division information, and the transmission protocol of the target video stream includes the WebRTC / QUIC protocol.
[0094] Optionally, when the object transmitting the target video stream is the edge computing side, the edge computing side determines the video stream after super-resolution enhancement and frame interpolation processing as the target video stream and transmits it to the server side. When the object transmitting the target video stream is the server side, the server transmits the target video stream to the client. When transmitting the target video stream, the network bandwidth is monitored in real time, and the transmission mode is dynamically switched (such as only transmitting the base layer in a weak network and retransmitting the enhancement layer after the bandwidth is restored) to achieve adaptive bitrate control.
[0095] Optionally, the base layer includes key points in the target video stream, the low-resolution images of each frame in the target video stream (such as 240p images), and audio information. When it is determined that the network bandwidth is poor (such as the network speed is less than 1 Mbps), only the information of the base layer is transmitted, and these information are used to ensure the smoothness of video transmission (such as the bitrate is 200 kbps).
[0096] Optionally, the enhancement layer includes the super-resolution feature vectors of each frame of image, such as the texture information of each frame of image after super-resolution enhancement by the super-resolution model Real-ESRGAN.
[0097] Optionally, the server side and the client side can generate high-definition pictures based on the data of the base layer. Among them, the server side can use high-precision models (such as Real-ESRGAN, RIFE) to complement the data of the base layer to obtain high-definition pictures.
[0098] In one embodiment, the data of the base layer can be transmitted by differential encoding, which can achieve a video stream transmission speed of 50 KB / frame. Compared with the full-frame transmission speed of 500 KB / frame, the data volume is reduced by 90%. Through the priority transmission of key data (base layer), differential encoding, model lightweighting, and hierarchical adaptive strategies, the present application compresses the data transmission volume to 10%-20% of the traditional scheme while predicting and completing lost frames, achieving a high-fluency, high-definition, and low-latency video transmission experience in a low-bandwidth environment.
[0099] Optionally, after receiving the video stream transmitted by the server, the client can perform real-time decoding and rendering processing on the video stream to obtain a super-resolution enhanced video. Additionally, the client can send bitrate adjustment information to the server according to the current network bandwidth, enabling the server to adjust the bitrate of the video stream to ensure smooth playback even at low network speeds. During video playback, the client can also perform lip synchronization operations in real-time based on the video frame images and audio information to ensure audio-visual synchronization. Through the above measures, the method of the present application can be applied to remote meetings in rural and remote areas (high-definition picture quality can still be guaranteed at low network speeds), disaster emergency command (ensuring real-time visual command), remote medical video diagnosis (doctors can clearly observe the patient's expressions and movements), and video conferencing on low-end devices (allowing low-performance devices to be used smoothly), expanding the application scenarios.
[0100] The video stream transmission method provided by the present application receives the transmitted first video stream. The acquisition process of the first video stream includes preprocessing of the original video stream; performing super-resolution enhancement and frame interpolation processing on the first video stream to obtain a target video stream, and transmitting the target video stream to the video receiving object based on the current network bandwidth and / or the video receiving object connected thereto; the embodiments of the present application can perform super-resolution enhancement and frame interpolation processing on the video stream during the video stream transmission process, improving the resolution of the video images in the video stream received by the client and reducing video frame loss, effectively improving the video playback effect and enhancing the user experience.
[0101] Based on the same inventive concept, an embodiment of the present application provides an electronic device, as Figure 3 shown, Figure 3 the electronic device 2000 shown includes: a processor 2001 and a memory 2003. Among them, the processor 2001 and the memory 2003 are communicatively connected, such as connected through a bus 2002.
[0102] The processor 2001 may be a CPU (Central Processing Unit), a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field-Programmable Gate Array), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It can implement or execute various exemplary logical blocks, modules, and circuits described in connection with the disclosure of this application. The processor 2001 may also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, etc.
[0103] The bus 2002 may include a path for transmitting information between the above components. The bus 2002 may be a PCI (Peripheral Component Interconnect) bus, an EISA (Extended Industry Standard Architecture) bus, or the like. The bus 2002 may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, only a thick line is shown in the figure, but it does not mean that there is only one bus or one type of bus.
[0104] The memory 2003 may be a ROM (Read-Only Memory) or other type of static storage device that can store static information and instructions, a RAM (random access memory), or other type of dynamic storage device that can store information and instructions. It may also be an EEPROM (Electrically Erasable Programmable Read Only Memory), a CD-ROM (Compact Disc Read-Only Memory), or other optical disc storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto.
[0105] Optionally, the electronic device 2000 may further include a communication unit 2004. The communication unit 2004 can be used for receiving and sending signals. The communication unit 2004 may allow the electronic device 2000 to communicate with other devices wirelessly or wiredly to exchange data. It should be noted that in practical applications, the communication unit 2004 is not limited to one.
[0106] Optionally, the electronic device 2000 may further include an input unit 2005. The input unit 2005 can be used for receiving input digital, character, image, and / or sound information, or generating key signal inputs related to the user settings and function controls of the electronic device 2000. The input unit 2005 may include, but is not limited to, one or more of a touch screen, a physical keyboard, function keys (such as volume control buttons, switch buttons, etc.), a trackball, a mouse, a joystick, a shooting device, a pickup, etc.
[0107] Optionally, the electronic device 2000 may further include an output unit 2006. The output unit 2006 can be used for outputting or presenting the information processed by the processor 2001. The output unit 2006 may include, but is not limited to, one or more of a display device, a speaker, a vibration device, etc.
[0108] Although the electronic device 2000 with various devices is shown in the figure, it should be understood that it is not required to implement or have all the shown devices. More or fewer devices can be alternatively implemented or had.
[0109] Optionally, the memory 2003 is used for storing a computer program for executing the solution of this application, and is controlled by the processor 2001 to execute. The processor 2001 is used for executing the computer program stored in the memory 2003 to implement the steps of any method provided by the embodiments of this application.
[0110] Based on the same inventive concept, the embodiments of this application provide a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by an electronic device / processor, it implements the steps of any method provided by this application / implements the steps of various alternative embodiments of the method provided by this application.
[0111] Based on the same inventive concept, the embodiments of this application provide a computer program product, which includes a computer program. When the computer program is executed by an electronic device / processor, it implements the steps of any method provided by this application / implements the steps of various alternative embodiments of the method provided by this application.
[0112] Those skilled in the art can understand that the various operations, methods, steps, measures, and solutions in the processes discussed in this application can be alternated, changed, combined, or deleted. Further, other steps, measures, and solutions in the various operations, methods, and processes discussed in this application can also be alternated, changed, rearranged, decomposed, combined, or deleted. Further, the steps, measures, and solutions in the related art that are the same as those disclosed in the various operations, methods, and processes in this application can also be alternated, changed, rearranged, decomposed, combined, or deleted.
[0113] In the description of this application, the directions or positional relationships indicated by the words "center", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", etc. are the exemplary directions or positional relationships based on the drawings, which are for the convenience of describing or simplifying the embodiments of this application, rather than indicating or implying that the device or component referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, it should not be construed as a limitation to this application.
[0114] The terms "first", "second", "third", "fourth", "1", "2", etc. (if any) in the description, claims, and the above drawings of this application are used to distinguish similar objects, and do not necessarily have to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of this application described here can be implemented in an order other than the illustrated or textually described order.
[0115] It should be understood that although the flowchart of the embodiments of this application indicates each operation step by an arrow, the execution order of these steps is not limited to the order indicated by the arrow. Unless there is a clear description in this article, in some implementation scenarios of the embodiments of this application, the implementation steps in each flowchart can be executed in other orders according to requirements. In addition, some or all of the steps in each flowchart may include multiple sub-steps or multiple stages based on the actual implementation scenario. Some or all of these sub-steps or stages can be executed at the same time, and each sub-step or stage of these sub-steps or stages can also be executed at different times respectively. In the scenario where the execution times are different, the execution order of these sub-steps or stages can be flexibly configured according to requirements, and the embodiments of this application do not limit this.
[0116] The above are only optional implementation manners of some implementation scenarios of this application. It should be noted that for those of ordinary skill in the art, without departing from the technical concept of the solution of this application, adopting other similar implementation means based on the technical idea of this application also belongs to the protection scope of the embodiments of this application.
Claims
1. A video stream transmission method, characterized in that, For the edge computing side or the server side, the method includes: Receiving a transmitted first video stream, and the acquisition process of the first video stream includes: the front end preprocessing the original video stream, and the preprocessing includes at least one of noise reduction and compression; Performing super-resolution enhancement and frame interpolation processing on the first video stream to obtain a target video stream; Transmitting the target video stream to the video receiving object based on the current network bandwidth and / or the video receiving object connected to itself, and the transmission includes any one of hierarchical transmission and overall transmission.
2. The video stream transmission method according to claim 1, wherein, The preprocessing includes noise reduction processing. The front end preprocessing the original video stream includes: Performing noise reduction processing on the original video stream to obtain a first video stream; The noise reduction processing includes: Using a noise reduction model to perform noise reduction processing on each frame image in the original video stream to obtain a denoised image. The noise reduction model is trained based on a lightweight convolutional neural network, and the number of parameters of the lightweight convolutional neural network is less than a first preset value, and the type and structure of the lightweight convolutional neural network correspond to the scenario where the front end is located.
3. The video stream transmission method according to claim 2, wherein, The preprocessing includes compression processing. The front end preprocessing the original video stream includes: Identifying key points in the denoised image, obtaining region category information of the denoised image according to the key points, compressing the denoised image according to the region category information to obtain a first video stream and embedding reconstruction data in the first video stream. The region category information includes important regions and background regions, and the reconstruction data is used for super-resolution enhancement.
4. The video stream transmission method according to claim 3, wherein The identifying key points in the denoised image includes: Using a key point identification model to identify key points in the denoised image. The architecture of the key point identification model is a hybrid architecture of MobileNetV3+LSTM, and the key point identification model extracts spatial features through the MobileNetV3 architecture and uses LSTM to process and capture temporal continuity.
5. The video stream transmission method according to claim 1, characterized in that The receiving the first video stream includes: The server side receives the first video stream transmitted by the edge computing side. The first video stream is obtained by performing super-resolution enhancement processing and frame interpolation processing on the second video stream of the edge computing side, and the second video stream is the original video stream preprocessed by the front end.
6. The video stream transmission method according to claim 5, wherein The edge computing side generating the first video stream includes: Obtaining a second video stream after super-resolution enhancement processing, and using an interpolation model to perform frame interpolation processing on the second video stream after super-resolution enhancement processing to obtain a frame interpolation video stream; Extracting compact data in the frame interpolation video stream and determining the compact data as the first video stream. The compact data includes key point data and audio data.
7. The video stream transmission method according to claim 1, wherein Performing frame interpolation processing on the first video stream includes: Determining the first video stream after super-resolution enhancement as the video stream to be frame interpolated, and determining the positions of the lost frames according to the frame information of each frame image in the video stream to be frame interpolated; Obtain the frame prediction information of the lost frame according to the adjacent frames corresponding to the position, generate the lost frame based on the frame prediction information and insert the lost frame into the video stream to be interpolated, where the frame prediction information includes at least one of motion trend, light change, background consistency, key point coordinates, key point displacement trajectory, expression parameters, and body movement information; Perform lip synchronization processing and expression and action adjustment on the lost frames in the video stream to be interpolated.
8. The video stream transmission method according to claim 1, wherein Transmitting the target video stream to the video receiving object based on the current network bandwidth and / or the video receiving object connected to itself includes: If it is determined that the video receiving object connected to itself is a client, obtain the hierarchical division information of the target video stream, where the hierarchical division information includes a base layer and an enhancement layer; Dynamically transmit the target video stream according to the network bandwidth and the hierarchical division information, and the transmission protocol of the target video stream includes the WebRTC / QUIC protocol.
9. An electronic device, comprising a memory, a processor, and a computer program stored on the memory, characterized in that, The processor executes the computer program to implement the steps of the method as described in any one of claims 1-8.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed, it implements the steps of the method as described in any one of claims 1-8.
Citation Information
Cited By
Video packet loss synchronous compensation method and system based on deep learning
CN120640087A
Photovoltaic module hot spot detection method based on unmanned aerial vehicle
CN120932135A