MVPS video coding method and device based on deep learning, and medium

Through the MVPS video encoding method based on deep learning, the problems of inefficient traditional video encoding technology, insufficient hardware coordination, cross-frame quality fluctuations and poor network adaptability are solved, and efficient, real-time and stable video encoding effects are achieved.

CN120223891AActive Publication Date: 2025-06-27BEIJING TOPMOO TECH

Patent Information

Application Number
CN202510685618.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-27
Publication Date
2025-06-27
Estimated Expiration
2045-05-27

AI Technical Summary

Technical Problem

Traditional video encoding technology is inefficient, insufficient hardware coordination, cross-frame quality fluctuations and poor network adaptability, making it difficult to meet the requirements of real-time and efficientness.

Method used

The MVPS video encoding method based on deep learning is adopted, and the video stream is extracted in time-domain and air-domain features through the feature extraction unit, and a joint feature map is generated through the adaptive weight fusion layer, bidirectional feature interaction is performed between levels, and quantitative parameters are generated in combination with network bandwidth feedback to realize dynamic coding strategy adjustment.

Benefits of technology

It improves video encoding efficiency, reduces calculation delay and power consumption, reduces inter-frame quality fluctuations, enhances network bandwidth adaptability, and supports collaborative encoding of audio, depth maps and videos.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120223891A_ABST
    Figure CN120223891A_ABST
Patent Text Reader

Abstract

The invention discloses an MVPS video coding method and device based on deep learning, and a medium, and relates to the technical field of video coding. The method comprises the following steps: acquiring a video stream, performing time domain and space domain feature extraction on the video stream through a feature extraction unit, and generating a joint feature map through an adaptive weight fusion layer; carrying out bidirectional feature interaction between hierarchies on the combined feature map to obtain a multi-scale feature set; generating a frame-level quantization parameter initial value according to the multi-scale feature set in combination with current network bandwidth feedback; performing time sequence consistency correction on the quantization parameter initial value to generate a target quantization parameter; and performing block coding on the current frame according to the target quantization parameter to generate a coded data stream. According to the method, through dynamic parameter adjustment, hardware algorithm collaborative optimization, cross-frame consistency enhancement and network awareness, the problems of low efficiency, insufficient real-time performance, unstable quality and poor adaptability of a traditional scheme are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of video coding technologies, and in particular, to an MVPS video coding method, device, and medium based on deep learning. Background Art

[0002] With the diversification of video application scenarios, such as more and more ultra-high-definition live broadcasts, intelligent security, VR / AR, etc., traditional video coding technologies are facing significant bottlenecks.

[0003] Traditional video coding technologies rely on fixed quantization tables and predefined coding block division rules, and are unable to dynamically adjust the coding strategy according to the video content, resulting in waste of bitrate in low-complexity scenarios and degradation of quality in high-dynamic scenarios. Coding schemes based on deep learning mostly focus on software algorithm optimization and lack co-design with dedicated hardware, resulting in high computational latency and low energy efficiency ratio, making it difficult to meet real-time requirements; in addition, traditional bitrate allocation models independently process single-frame data and ignore the temporal correlation of videos, resulting in quality fluctuations between adjacent frames and affecting the subjective visual experience; the ability to respond in real time to network bandwidth fluctuations is poor, and there is easy to appear stuttering or image quality degradation in weak network environments.

[0004] Through the above analysis, the problems and defects existing in the prior art are as follows: The video coding technologies in the prior art are inefficient, lack hardware cooperation, have cross-frame quality fluctuations, and poor network adaptability. Summary of the Invention

[0005] Embodiments of this application provide an MVPS video coding method, device, and medium based on deep learning, which can solve the problems of low efficiency of video coding technologies, lack of hardware cooperation, cross-frame quality fluctuations, and poor network adaptability in the prior art.

[0006] In a first aspect, embodiments of this application provide an MVPS video coding method based on deep learning. The method includes: obtaining a video stream, extracting time-domain and spatial-domain features of the video stream through a feature extraction unit, and generating a joint feature map through an adaptive weight fusion layer; performing bidirectional feature interaction between levels on the joint feature map to obtain a multi-scale feature set; generating an initial value of frame-level quantization parameters according to the multi-scale feature set in combination with the current network bandwidth feedback; performing temporal consistency correction on the initial value of the quantization parameters to generate target quantization parameters; and performing block coding on the current frame according to the target quantization parameters to generate an encoded data stream.

[0007] In an implementation manner of the present application, a video stream is obtained, and the feature extraction unit extracts the time-domain and space-domain features of the video stream, and generates a joint feature map through the adaptive weight fusion layer, specifically including: deploying the feature extraction unit in parallel on the FPGA chip; parsing the video stream into MVPS video stream data in units of frames to obtain video frames; through the feature extraction unit, using a three-dimensional convolutional neural network to extract the time-domain motion features of the video frames, where the optical flow calculation result is used as the initial value of the convolutional weight; using a deformable convolutional kernel to extract the space-domain texture features of the video frames, and adjusting the sampling position according to the edge gradient of the video frames; performing channel splicing on the time-domain motion features and the space-domain texture features, calculating the fusion weights of the time-domain motion features and the space-domain texture features, and generating a joint feature map.

[0008] In an implementation manner of the present application, performing channel splicing on the time-domain motion features and the space-domain texture features, calculating the fusion weights of the time-domain motion features and the space-domain texture features, and generating a joint feature map, specifically including: obtaining spatial audio through a microphone array, obtaining a depth map through a ToF sensor, and parsing to obtain audio spectrum features and depth map edge features; aligning the audio spectrum features with the time-domain motion features to obtain a cross-modal similarity matrix; fusing the depth map edge features with the space-domain texture features to generate a geometrically aware space-domain feature field; inputting the cross-modal similarity matrix and the feature field into a differentiable neural architecture to generate a dynamically sparse-connected spatio-temporal attention network; performing multi-modal collaborative compression through the spatio-temporal attention network to generate a joint feature map.

[0009] In an implementation manner of the present application, performing hierarchical bidirectional feature interaction on the joint feature map to obtain a multi-scale feature set, specifically including: deploying a causal Transformer encoder and a non-causal Transformer decoder between adjacent levels of the feature pyramid; constraining the preset low-level features through the causal Transformer encoder; through the Transformer decoder, multiplying the preset high-level semantic features and the low-level features element by element to obtain an interaction feature map; applying channel attention weights to the feature map to suppress redundant feature channels and obtain a multi-scale feature set.

[0010] In an implementation manner of the present application, according to the multi-scale feature set, combined with the current network bandwidth feedback, generating an initial value of the frame-level quantization parameter, specifically including: establishing a network bandwidth monitoring thread to obtain network bandwidth data in real time, where the network bandwidth data includes the TCP congestion window size and the packet loss rate; encoding the network bandwidth data into a vector, splicing it with the multi-scale feature set, and inputting it into a gated attention network; generating an initial value of the quantization parameter according to the attention weight distribution to make the bit rate allocation meet the network transmission constraints.

[0011] In one implementation of the present application, after the initial value of the quantization parameter is corrected for temporal consistency to generate the target quantization parameter, the method further includes: initially dividing the video frame into coding blocks of the maximum size; extracting the motion vector field of the coding blocks of the previous frame to construct a motion consistency matrix; calculating the correlation coefficient between the coding blocks of the current frame and the motion consistency matrix to generate a temporal smoothing constraint factor.

[0012] In one implementation of the present application, after the temporal smoothing constraint factor is generated, the method further includes: performing a non-linear mapping on the initial value of the quantization parameter according to the temporal smoothing constraint factor; recursively calculating the rate-distortion cost of the coding blocks, and if the rate-distortion cost is higher than a preset threshold, continuing to divide the coding blocks according to the target quantization parameter.

[0013] In one implementation of the present application, after the current frame is divided into coding blocks according to the target quantization parameter to generate the encoded data stream, the method further includes: inputting the coding blocks into a rate-distortion optimization module to iteratively optimize the coding block division mode through a rate-distortion cost function; performing entropy coding on the optimized coding blocks to generate the final compressed bitstream; during the encoding process, monitoring the video content complexity in real time to update the weight parameters for bitrate allocation.

[0014] In a second aspect, an MVPS video encoding device based on deep learning provided by an embodiment of the present application includes at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to: obtain a video stream, perform temporal and spatial feature extraction on the video stream through a feature extraction unit, and generate a joint feature map through an adaptive weight fusion layer; perform bidirectional feature interaction between levels on the joint feature map to obtain a multi-scale feature set; generate an initial value of the frame-level quantization parameter according to the multi-scale feature set in combination with the current network bandwidth feedback; correct the initial value of the quantization parameter for temporal consistency to generate a target quantization parameter; perform block coding on the current frame according to the target quantization parameter to generate an encoded data stream.

[0015] In a third aspect, a non-volatile computer storage medium for MVPS video encoding based on deep learning provided by an embodiment of the present application stores computer-executable instructions, which are set to: obtain a video stream, perform temporal and spatial feature extraction on the video stream through a feature extraction unit, and generate a joint feature map through an adaptive weight fusion layer; perform bidirectional feature interaction between levels on the joint feature map to obtain a multi-scale feature set; generate an initial value of the frame-level quantization parameter according to the multi-scale feature set in combination with the current network bandwidth feedback; correct the initial value of the quantization parameter for temporal consistency to generate a target quantization parameter; perform block coding on the current frame according to the target quantization parameter to generate an encoded data stream.

[0016] A MVPS video coding method, device and medium based on deep learning provided by an embodiment of the present application, where an adaptive weight fusion layer interacts bidirectionally with features between levels to achieve dynamic parameter adjustment of video content perception; an FPGA parallel feature extraction unit and a deformable convolution kernel hardware acceleration design reduce video processing latency and power consumption; a motion consistency matrix and a temporal smoothing constraint factor reduce the range of inter-frame fluctuations and significantly reduce blocking effects and flickering phenomena; a bandwidth feedback mechanism and a rate-distortion optimization module achieve smooth quality transition during bandwidth fluctuations, and a cross-modal similarity matrix and a geometric perception spatial feature field support collaborative coding of audio, depth maps and videos. Through dynamic parameter adjustment, hardware-algorithm co-optimization, cross-frame consistency enhancement and network perception technology, the problems of low efficiency, insufficient real-time performance, unstable quality and poor adaptability of traditional solutions are solved. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments and descriptions thereof of the present application are used to explain the present application and do not constitute an improper limitation of the present application. In the drawings: Figure 1 is a flowchart of a MVPS video coding method based on deep learning provided by an embodiment of the present application; Figure 2 is a schematic internal structure diagram of a MVPS video coding device based on deep learning provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0018] To make the objectives, technical solutions and advantages of the present application clearer, the technical solutions of the present application will be clearly and completely described below in conjunction with the specific embodiments of the present application and the corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0019] An embodiment of the present application provides a MVPS video coding method, device and medium based on deep learning, which solves the problems of low efficiency, insufficient hardware collaboration, cross-frame quality fluctuations and poor network adaptability in the existing video coding technology.

[0020] The technical solutions proposed in the embodiments of the present application will be described in detail below with reference to the drawings.

[0021] Figure 1 is a flowchart of a MVPS video coding method based on deep learning provided by an embodiment of the present application. As Figure 1As shown in the figure, a MVPS video coding method based on deep learning provided by an embodiment of the present application specifically includes the following steps: Step 10: Obtain a video stream, extract time-domain and spatial-domain features of the video stream through a feature extraction unit, and generate a joint feature map through an adaptive weight fusion layer.

[0022] As an optional embodiment, obtaining a video stream, extracting time-domain and spatial-domain features of the video stream through a feature extraction unit, and generating a joint feature map may specifically include: Step 101: Parallelly deploy a feature extraction unit on an FPGA (Field-Programmable Gate Array) chip; Step 102: Parse the video stream into MVPS video stream data in units of frames to obtain video frames; Step 103: Through the feature extraction unit, use a three-dimensional convolutional neural network to extract the time-domain motion features of the video frames, where the optical flow calculation result is used as the initial value of the convolutional weight; Step 104: Use a deformable convolutional kernel to extract the spatial-domain texture features of the video frames and adjust the sampling position according to the edge gradient of the video frames; Step 105: Concatenate the time-domain motion features and the spatial-domain texture features in channels, calculate the fusion weight of the time-domain motion features and the spatial-domain texture features, and generate a joint feature map.

[0023] In this step, a time-domain feature extraction module and a spatial-domain feature extraction module are parallelly deployed on an FPGA (Field-Programmable Gate Array) chip. A dual-channel data pipeline is adopted in the hardware design: the first channel processes several consecutive frames of video data through a three-dimensional convolutional kernel to extract motion features; the second channel analyzes the texture details of a single-frame image through a deformable convolutional kernel; the feature maps output by the two channels are cached through a high-speed on-chip memory and are spliced in real time by the adaptive weight fusion layer.

[0024] Optionally, the time-domain feature extraction uses a pre-trained 3D convolutional network, and the initial weight is optimized through an optical flow field. Specifically, first calculate the optical flow field of adjacent frames, and convert the optical flow displacement into the initial weight of the three-dimensional convolutional kernel. For example, if the optical flow field shows that there is horizontal motion in a certain area, the horizontal weight of the corresponding convolutional kernel will be enhanced. In this way, the network can quickly capture the motion trend in the video.

[0025] Optionally, the spatial-domain feature extraction uses a deformable convolutional kernel, and the sampling position is dynamically adjusted by the image gradient. For the input video frames, first calculate the Sobel edge gradient map, and generate an offset control vector according to the gradient amplitude. For example, in areas with complex textures, such as leaves or hair, the gradient amplitude is high, and the sampling points of the convolutional kernel will shift towards the edge direction, thereby enhancing the ability to capture details. This process realizes the prediction of the offset through a two-layer fully connected network.

[0026] As an alternative embodiment, the time-domain motion features and the spatial-domain texture features are concatenated in channels, the fusion weights of the time-domain motion features and the spatial-domain texture features are calculated, and a joint feature map is generated, which may specifically include: Step 1051: Obtain spatial audio through a microphone array, obtain a depth map through a ToF (Time of Flight) sensor, and parse to obtain audio spectrum features and depth map edge features; Step 1052: Align the audio spectrum features with the time-domain motion features to obtain a cross-modal similarity matrix. In this step, based on the video frame rate, linear interpolation is performed on the audio features to ensure that each frame corresponds to a set of audio features. A similarity matrix is constructed through cosine similarity, the matrix is row-normalized, and then elements with a similarity greater than a preset value are retained to form a sparse association matrix.

[0027] Step 1053: Fuse the depth map edge features with the spatial-domain texture features to generate a geometrically aware spatial feature field.

[0028] In this step, depth edge weights are introduced in the calculation of the sampling offsets of deformable convolutions, and the fused features are mapped to the geometrically aware spatial feature field through a 3×3 convolutional layer, and the channel dimension is the same as that of the original spatial-domain features. In this way, for transparent objects, the geometric constraints provided by the depth edges can avoid misjudgment of RGB (Red Green Blue Texture) textures, making the feature field more conform to the real object boundaries.

[0029] Step 1054: Input the cross-modal similarity matrix and the feature field into a differentiable neural architecture to generate a dynamically sparse-connected spatio-temporal attention network.

[0030] In this step, the cross-modal matrix and the feature field are input into a neural architecture search module. Using the cross-modal matrix as the node connection weights and the feature field as the node features, the spatio-temporal graph structure is initialized. The top k strongly associated edges are selected through Gumbel-Softmax sampling to generate an attention network that only retains important connections. The searched network structure is compiled into a lookup table logic executable by an FPGA for dynamic reconstruction.

[0031] Step 1055: Perform multi-modal collaborative compression through the spatio-temporal attention network to generate a joint feature map.

[0032] In this step, an audio-driven attention mask is applied to the time-domain motion features to suppress the motion estimation error in the silent area, such as background random jitter. In the spatial feature field, lossless compression is performed on the texture features in the depth edge area, and lossy compression is performed on non-edge areas. The final joint feature map is output through weighted concatenation in the channel dimension.

[0033] That is to say, through the feature extraction of the temporal motion features and the spatial texture features, a joint feature map is generated by combining the audio and the depth map, providing a high-dimensional input for subsequent feature interaction.

[0034] Step 20: Perform two-way feature interaction between levels of the joint feature map to obtain a multi-scale feature set; As an optional embodiment, performing two-way feature interaction between levels of the joint feature map to obtain a multi-scale feature set may specifically include: Step 201: Deploy a causal Transformer encoder and a non-causal Transformer decoder between adjacent levels of the feature pyramid; Step 202: Constrain the preset low-level features through the causal Transformer encoder; Step 203: Multiply the preset high-level semantic features and the low-level features element by element through the Transformer decoder to obtain an interaction feature map; Step 204: Apply channel attention weights to the feature map to suppress redundant feature channels and obtain a multi-scale feature set.

[0035] In this step, a bidirectional Transformer is deployed between adjacent levels of the feature pyramid. On the encoder side, from bottom to top: and a causal attention mechanism is adopted, allowing only low-level features to be affected by historical frame features to avoid leakage of future information. For example, when encoding the t-th frame, the P3 layer only refers to the features of the t-1-th frame and previous frames; on the decoder side, from top to bottom: multiply the high-level semantic features such as object contours and the low-level detail features such as textures element by element, and fuse them through residual connections. For example, the weight of the face area in the high-level features will enhance the facial features at the corresponding low-level positions.

[0036] Furthermore, apply a channel attention mechanism to the interaction feature map, perform global average pooling on the feature map of each channel to obtain a channel description vector; calculate the importance weights of each channel through a two-layer fully connected network, and perform soft masking on the channels with weights lower than the threshold to reduce redundant calculations. For example, if a certain channel mainly responds to flat areas such as the sky, its weight will be reduced.

[0037] That is to say, the joint feature map contains multi-modal information. Through the interaction of the bidirectional Transformer between levels, the low-level details and the high-level semantics are fused, and a multi-scale feature set is output, providing a basis for generating dynamic quantization parameters.

[0038] Step 30: Generate an initial value of the frame-level quantization parameter according to the multi-scale feature set and in combination with the current network bandwidth feedback.

[0039] As an alternative embodiment, based on the multi-scale feature set and combined with the current network bandwidth feedback, an initial value of the frame-level quantization parameter is generated, which may specifically include: Step 301: Establish a network bandwidth monitoring thread to obtain network bandwidth data in real time. The network bandwidth data includes the TCP congestion window size and the packet loss rate; Step 302: Encode the network bandwidth data into a vector, splice it with the multi-scale feature set, and input it into the gated attention network; Step 303: Generate an initial value of the quantization parameter according to the attention weight distribution to make the bitrate allocation conform to the network transmission constraint.

[0040] In this step, a bandwidth monitoring thread is deployed at the network transport layer to obtain the TCP (Congestion Window) congestion window and the packet loss rate in real time. That is, when it is detected that the TCP congestion window shrinks and the packet loss rate rises, it is determined that the network bandwidth decreases, and the bitrate compression strategy is triggered. The bandwidth data is encoded into a 16-dimensional vector. The TCP congestion window is normalized to [0, 1], and the packet loss rate is mapped to a logarithmic scale. This vector is spliced with the multi-scale features and input into the gated recurrent unit network to predict the initial quantization parameter. When the network bandwidth decreases, the initial quantization parameter value is increased by 2 - 3 units to reduce the bitrate.

[0041] Furthermore, the initial quantization parameter can also be dynamically adjusted according to the complexity of different regions in the feature map. For high-complexity regions, such as regions with intense motion or rich texture, a lower initial quantization parameter is adopted to retain details; for flat or static regions, the initial quantization parameter is appropriately increased to reduce bit allocation; in specific implementation, the region complexity is calculated by combining the variance of the feature map and mapped into an initial quantization parameter offset.

[0042] That is to say, the above multi-scale feature set fuses semantic information at different levels, and then combines the network bandwidth feedback to generate an initial value of the quantization parameter through the gated attention network.

[0043] Step 40: Perform temporal consistency correction on the initial value of the quantization parameter to generate the target quantization parameter; As an alternative embodiment, after performing temporal consistency correction on the initial value of the quantization parameter to generate the target quantization parameter, the method may further include: initially dividing the video frame into the largest-sized coding blocks; extracting the motion vector field of the previous frame's coding blocks to construct a motion consistency matrix; calculating the correlation coefficient between the current frame's coding blocks and the motion consistency matrix to generate a temporal smoothing constraint factor.

[0044] In this step, QP (Quantization Parameter Smoothing) smoothing is performed using the correlation of motion vectors between consecutive frames. The motion vector of the previous frame's encoded block is extracted to construct a motion consistency matrix. For example, if the motion direction of the current block is the same as that of the previous frame, a higher consistency score is assigned. The correlation coefficient between each encoded block of the current frame and the historical motion trajectory is calculated to generate a temporal smoothing factor (0.0 - 1.0). That is, if the motion of a certain block mutates, assuming a suddenly emerging object, its smoothing factor is reduced to allow local fluctuations in the QP value.

[0045] Furthermore, QP correction is performed for the human eye sensitive regions. The regions of interest are extracted through a pre-trained visual saliency detection model, and the QP values of these regions are constrained by a lower limit, such as QP ≤ 32, to prevent blocking artifacts caused by excessive compression.

[0046] In summary, the initial value of the quantization parameter is adjusted based on content complexity and network status, etc., to ensure that the bitrate allocation not only meets the requirements of the video content but also adapts to real-time bandwidth changes.

[0047] Step 50: Perform block encoding on the current frame according to the target quantization parameter to generate an encoded data stream.

[0048] As an alternative embodiment, after generating the temporal smoothing constraint factor, the method may further include: performing a non-linear mapping on the initial value of the quantization parameter according to the temporal smoothing constraint factor; recursively calculating the rate-distortion cost of the encoded block. If the rate-distortion cost is higher than a preset threshold, continue to split the encoded block according to the target quantization parameter.

[0049] In this step, the initial quantization parameter may have an inter-frame jump due to network fluctuations or content mutations, and smoothing constraints need to be performed through the motion consistency matrix and visual saliency detection. The encoding block partitioning can adopt the rate-distortion optimization (RDO) strategy. Assuming that the rate-distortion cost J is initially calculated with a 64×64 block as a unit, if J exceeds the threshold, for example, due to complex texture resulting in increased distortion, the block is split into 4 32×32 sub-blocks for re-evaluation. Recursively execute until the minimum block size of 8×8 is reached to ensure fine encoding of complex regions.

[0050] As an alternative embodiment, after performing encoding block partitioning on the current frame according to the target quantization parameter to generate an encoded data stream, the method may further include: inputting the encoded block into a rate-distortion optimization module to iteratively optimize the encoding block partitioning mode through a rate-distortion cost function; performing entropy encoding on the optimized encoded block to generate a final compressed bitstream; during the encoding process, monitor the video content complexity in real time to update the weight parameters of the bitrate allocation.

[0051] In this step, further, the target quantization parameter takes into account content complexity, network status, and temporal smoothness, and recursively partitions the coding blocks through rate-distortion optimization to achieve fine coding of complex regions. Finally, the encoded data stream is compressed by entropy coding, and the bitrate distribution is matched with the network bandwidth in real time while ensuring the stability of subjective quality. Context Adaptive Binary Arithmetic Coding (CABAC) is adopted in the entropy coding stage, and the probability model is dynamically updated according to the statistical characteristics of the feature map. Differential coding is performed on the motion vectors, and the prediction context is constructed using historical MVs. The residual coefficients are grouped according to texture complexity, and different Huffman tables are used respectively.

[0052] The above is the method embodiment proposed in this application. Based on the same inventive concept, the embodiment of this application also provides a deep learning-based MVPS video coding device, and its structure is as Figure 2 shown.

[0053] Figure 2 It is a schematic internal structure diagram of a deep learning-based MVPS video coding device provided by the embodiment of this application. As Figure 2 shown, the device includes: At least one processor 201; And a memory 202 communicatively connected to the at least one processor; Wherein, the memory 202 stores instructions executable by the at least one processor. The instructions are executed by the at least one processor 201, so that the at least one processor 201 can: obtain a video stream, perform temporal and spatial feature extraction on the video stream through a feature extraction unit, and generate a joint feature map through an adaptive weight fusion layer; perform bidirectional feature interaction between levels on the joint feature map to obtain a multi-scale feature set; generate an initial value of the frame-level quantization parameter according to the multi-scale feature set in combination with the current network bandwidth feedback; perform temporal consistency correction on the initial value of the quantization parameter to generate a target quantization parameter; perform block coding on the current frame according to the target quantization parameter to generate an encoded data stream.

[0054] Some embodiments of this application provide a Figure 1 corresponding non-volatile computer storage medium for deep learning-based MVPS video coding, storing computer-executable instructions, and the computer-executable instructions are set to: obtain a video stream, perform temporal and spatial feature extraction on the video stream through a feature extraction unit, and generate a joint feature map through an adaptive weight fusion layer; perform bidirectional feature interaction between levels on the joint feature map to obtain a multi-scale feature set; generate an initial value of the frame-level quantization parameter according to the multi-scale feature set in combination with the current network bandwidth feedback; perform temporal consistency correction on the initial value of the quantization parameter to generate a target quantization parameter; perform block coding on the current frame according to the target quantization parameter to generate an encoded data stream.

[0055] Each embodiment in this application is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the embodiments of the Internet of Things devices and media, since they are basically similar to the method embodiments, the description is relatively simple, and for the relevant parts, reference can be made to the partial description of the method embodiments.

[0056] The systems and media provided by the embodiments of this application correspond one by one to the methods. Therefore, the systems and media also have beneficial technical effects similar to those of their corresponding methods. Since the beneficial technical effects of the methods have been described in detail above, the beneficial technical effects of the systems and media will not be elaborated here.

[0057] Those skilled in the art should understand that the embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, this application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.

[0058] This application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of this application. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0059] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that implement the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0060] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus, so that a series of operation steps are executed on the computer or other programmable apparatus to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable apparatus provide steps for implementing the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 steps of the functions specified in one block or multiple blocks.

[0061] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.

[0062] The memory may include non-permanent memory in the computer-readable medium, in the form of random access memory (RAM) and / or non-volatile memory such as read-only memory (ROM) or flash RAM. The memory is an example of a computer-readable medium.

[0063] Computer-readable media includes both permanent and non-permanent, removable and non-removable media implemented by any method or technology for storing information. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile discs (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices, or any other non-transmission media that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transitory media such as modulated data signals and carrier waves.

[0064] It should also be noted that the term "comprising", "including" or any other variation thereof is intended to cover non-exclusive inclusion, so that a process, method, article or apparatus comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or apparatus. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or apparatus comprising the element.

[0065] The above are only embodiments of the present application and are not intended to limit the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the scope of the claims of the present application.

Claims

1. An MVPS video coding method based on deep learning, characterized in that, The method includes: Obtain a video stream, extract time-domain and spatial-domain features of the video stream through a feature extraction unit, and generate a joint feature map through an adaptive weight fusion layer; Perform two-way feature interaction between levels on the joint feature map to obtain a multi-scale feature set; Generate an initial value of frame-level quantization parameters according to the multi-scale feature set in combination with the current network bandwidth feedback; Perform temporal consistency correction on the initial value of the quantization parameters to generate target quantization parameters; Perform block coding on the current frame according to the target quantization parameters to generate an encoded data stream.

2. The MVPS video coding method based on deep learning according to claim 1, characterized in that, Obtain a video stream, extract time-domain and spatial-domain features of the video stream through a feature extraction unit, and generate a joint feature map, specifically including: Parallelly deploy the feature extraction unit on an FPGA chip; Parse the video stream into MVPS video stream data in units of frames to obtain video frames; Through the feature extraction unit, use a three-dimensional convolutional neural network to extract the time-domain motion features of the video frames, where the optical flow calculation result is used as the initial value of the convolutional weight; Use a deformable convolutional kernel to extract the spatial-domain texture features of the video frames and adjust the sampling positions according to the edge gradients of the video frames; Perform channel splicing on the time-domain motion features and the spatial-domain texture features, calculate the fusion weights of the time-domain motion features and the spatial-domain texture features, and generate the joint feature map.

3. The MVPS video encoding method based on deep learning according to claim 2, wherein Perform channel splicing on the time-domain motion features and the spatial-domain texture features, calculate the fusion weights of the time-domain motion features and the spatial-domain texture features, and generate the joint feature map, specifically including: Obtain spatial audio through a microphone array, obtain a depth map through a ToF sensor, and parse to obtain audio spectrum features and depth map edge features; Align the audio spectrum features with the time-domain motion features to obtain a cross-modal similarity matrix; Fuse the depth map edge features with the spatial-domain texture features to generate a geometry-aware spatial feature field; Input the cross-modal similarity matrix and the feature field into a differentiable neural architecture to generate a dynamically sparse-connected spatio-temporal attention network; Perform multi-modal collaborative compression through the spatio-temporal attention network to generate a joint feature map.

4. A MVPS video encoding method based on deep learning according to claim 1, characterized in that, Perform two-way feature interaction between levels on the joint feature map to obtain a multi-scale feature set, specifically including: Deploy a causal Transformer encoder and a non-causal Transformer decoder between adjacent levels of the feature pyramid; Constrain the preset low-level features through the causal Transformer encoder; Through the Transformer decoder, multiply the preset high-level semantic features element-wise with the low-level features to obtain an interacted feature map; Apply channel attention weights to the feature map to suppress redundant feature channels and obtain a multi-scale feature set.

5. A MVPS video encoding method based on deep learning according to claim 1, characterized in that Generate an initial value of frame-level quantization parameters according to the multi-scale feature set in combination with the current network bandwidth feedback, specifically including: Establish a network bandwidth monitoring thread to obtain network bandwidth data in real time, where the network bandwidth data includes the TCP congestion window size and the packet loss rate; Encode the network bandwidth data into a vector, concatenate it with the multi-scale feature set, and input it into the gated attention network; Generate an initial value of the quantization parameter according to the attention weight distribution, so that the bitrate allocation conforms to the network transmission constraint.

6. A method for MVPS video encoding based on deep learning according to claim 1, characterized in that After performing temporal consistency correction on the initial value of the quantization parameter to generate the target quantization parameter, the method further includes: Initially divide the video frame into coding blocks of the maximum size; Extract the motion vector field of the coding block of the previous frame and construct a motion consistency matrix; Calculate the correlation coefficient between the coding block of the current frame and the motion consistency matrix to generate a temporal smoothing constraint factor.

7. The MVPS video encoding method based on deep learning according to claim 6, wherein, After generating the temporal smoothing constraint factor, the method further includes: Perform a non-linear mapping on the initial value of the quantization parameter according to the temporal smoothing constraint factor; Recursively calculate the rate-distortion cost of the coding block. If the rate-distortion cost is higher than the preset threshold, continue to divide the coding block according to the target quantization parameter.

8. A MVPS video coding method based on deep learning according to claim 5, characterized in that, After performing coding block division on the current frame according to the target quantization parameter to generate the encoded data stream, the method further includes: Input the coding block into the rate-distortion optimization module and iteratively optimize the coding block division through the rate-distortion cost function; Perform entropy coding on the optimized coding block to generate the final compressed bitstream; During the encoding process, monitor the video content complexity in real time to update the weight parameters of the bitrate allocation.

9. An MVPS video encoding device based on deep learning, characterized in that, The device includes: At least one processor; And a memory communicatively connected to the at least one processor; Wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can: Obtain a video stream, perform temporal and spatial feature extraction on the video stream through a feature extraction unit, and generate a joint feature map through an adaptive weight fusion layer; Perform bidirectional feature interaction between levels on the joint feature map to obtain a multi-scale feature set; Generate an initial value of the frame-level quantization parameter according to the multi-scale feature set in combination with the current network bandwidth feedback; Perform temporal consistency correction on the initial value of the quantization parameter to generate the target quantization parameter; Perform block coding on the current frame according to the target quantization parameter to generate the encoded data stream.

10. A non-volatile computer storage medium for MVPS video coding based on deep learning, storing computer-executable instructions, characterized in that, The computer-executable instructions are set to: Obtain a video stream, perform temporal and spatial feature extraction on the video stream through a feature extraction unit, and generate a joint feature map through an adaptive weight fusion layer; Perform bidirectional feature interaction between levels on the joint feature map to obtain a multi-scale feature set; Generate an initial value of the frame-level quantization parameter according to the multi-scale feature set in combination with the current network bandwidth feedback; Perform temporal consistency correction on the initial value of the quantization parameter to generate the target quantization parameter; Perform block coding on the current frame according to the target quantization parameter to generate the encoded data stream.

Citation Information

Patent Citations

  • Video coding convolution filtering method based on attention mechanism fusion unit division

    CN112261414A

  • Depth image video generation method and device based on data fusion

    CN116485863A

  • MVPS series video processing system and method

    CN117411978A

  • Video encoding and decoding method and device, vehicle, storage medium and program product

    CN118573860A

  • Block optimization method and device based on ultra-high-definition video transmission and medium

    CN120017832A

Cited By

  • Video processing method and device, readable storage medium and program product

    CN120877176A

  • Image-to-video generation method and device, computer equipment and storage medium

    CN121037649A

  • An image-to-video generation method, device, computer equipment and storage medium

    CN121037649B

  • Video coding and decoding adaptive optimization method based on generative adversarial network

    CN121585818A

  • Multilayer adaptive rate distortion optimization coding method and device, and storage medium

    CN122093565A