An intelligent video coding optimization method and optimization system
Through motion vector estimation and resolution pyramid processing combined with deep learning models, the motion area and background area in the video frame are identified, and differentiated quantitative parameter encoding is adopted, which solves the problem of inefficiency in traditional video encoding and achieves efficient and high-quality video encoding.
Patent Information
- Application Number
- CN202411073357.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-06
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2044-08-06
AI Technical Summary
When traditional video encoding methods deal with high dynamic range video and complex scene video, the encoding efficiency is inefficient and the image quality loss is serious, so they cannot fully utilize the spatial and temporal redundancy of the video content.
The motion and non-motion areas in the video frame are identified through motion vector estimation, and the resolution pyramid is constructed for multi-scale processing, and the pre-trained deep learning model is used to differentiate coding, and the foreground and background areas are encoded using different quantization parameters.
The compression ratio and encoding efficiency of video encoding are improved, ensuring high quality and low distortion of the video after encoding.
Smart Images

Figure CN119071489B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of video coding technology, and in particular, to an intelligent video coding optimization method and an optimization system. Background Art
[0002] With the rapid development of video technology, the explosive growth of video data volume has put forward higher requirements for video coding technology. Traditional video coding methods often face problems such as low coding efficiency and large loss of image quality when dealing with high-dynamic-range videos, complex-scene videos, and real-time video transmissions. To solve these problems, intelligent video coding technology has emerged. It realizes the efficient compression and high-quality reconstruction of video data by combining advanced image processing technology and deep learning algorithms.
[0003] In the prior art, video coding usually relies on the uniform processing of video frames. Regardless of how the video content changes, the same coding parameters are used for coding. However, video content often contains a large number of dynamically changing regions (such as moving objects) and relatively static background regions. Using the same coding parameters for these two types of regions not only fails to make full use of the spatial and temporal redundancy of video content but may also lead to low coding efficiency and a decline in image quality.
[0004] Therefore, it is necessary to provide an intelligent video coding optimization method and an optimization system to solve the above technical problems. Summary of the Invention
[0005] To solve the above technical problems, the present invention provides an intelligent video coding optimization method and an optimization system. By combining steps such as preliminary motion detection, video frame segmentation, multi-scale processing, deep learning prediction, and differential coding, it realizes the intelligent processing and efficient coding of video content, not only improving the compression ratio and coding efficiency of video coding but also ensuring the high quality and low distortion of the encoded video.
[0006] The present invention provides an intelligent video coding optimization method, and the optimization method includes the following steps:
[0007] S1: Based on adjacent video frames of the current video frame in the video to be encoded, perform preliminary motion detection on the current video frame through a motion vector estimation method to obtain the motion region and the non-motion region belonging to the current video frame;
[0008] S2: Segment the current video frame based on the selected coding standard to obtain a plurality of sub-blocks;
[0009] S3: Construct a resolution pyramid of the current video frame, and perform multi-scale processing on the plurality of sub-blocks belonging to the motion region based on the resolution pyramid to obtain the motion region after multi-scale processing;
[0010] S4: Use the motion region after multi-scale processing as input, and use a pre-trained deep learning model to predict the current video frame to obtain the foreground and background of the motion region and the first quantization parameter and the second quantization parameter corresponding to the foreground and background respectively, where the first quantization parameter is less than the second quantization parameter;
[0011] S5: Use the first quantization parameter and the second quantization parameter to encode the foreground and background regions in the motion region respectively, and use the second quantization parameter to encode the non-motion region.
[0012] Preferably, step S1 includes the following steps:
[0013] S101: Use a feature point detection algorithm to identify feature points in the current frame and adjacent video frames of the video to be encoded, determine the corresponding relationship of the identified feature points between frames, and calculate the motion vectors of each feature point between the current frame and the adjacent video frames;
[0014] S102: By analyzing the changes in the direction and magnitude of the motion vectors, identify and locate the motion region in the current video frame, and obtain the non-motion region through the exclusion method;
[0015] S103: Mark the obtained motion region and non-motion region with visual markers, where the visual markers include boundary markers and color markers.
[0016] Preferably, step S2 includes the following steps:
[0017] S201: Perform parameter initialization settings according to the selected coding standard;
[0018] S202: Use the block division mechanism corresponding to the selected coding standard to divide the current video frame into multiple sub-blocks.
[0019] Preferably, step S3 includes the following steps:
[0020] S301: Starting from the original resolution of the current video frame, construct a resolution pyramid with different resolution levels by gradually reducing the resolution;
[0021] S302: In each layer of the resolution pyramid, select the sub-blocks belonging to the motion region;
[0022] S303: Perform multi-scale processing on all the sub-blocks belonging to the motion region in each selected layer of the pyramid;
[0023] S304: Merge all the sub-blocks of the motion region in each layer of the pyramid after multi-scale processing to obtain the motion region after multi-scale processing.
[0024] Preferably, step S4 includes the following steps:
[0025] S401: Extract features from the multi-scale processed motion region of the input using a pre-trained deep learning model to obtain the foreground features and background features in the motion region, where the deep learning model is a neural network model trained based on a dataset of annotated video frames;
[0026] S402: Generate a probability map based on the foreground features and background features, and obtain the foreground and background based on the probability map, where the probability map is used to represent the probabilities of sub-blocks in the motion region belonging to the foreground and background;
[0027] S403: Perform differential allocation of quantization parameters for the obtained foreground and background to obtain a first quantization parameter and a second quantization parameter corresponding to the foreground and background respectively.
[0028] Preferably, step S5 includes the following steps:
[0029] S501: Configure corresponding coding parameters according to the first quantization parameter and the second quantization parameter;
[0030] S502: Use the configured first quantization parameter and corresponding coding parameters to encode the foreground of the motion region through an encoder, and use the configured second quantization parameter and corresponding coding parameters to encode the background of the non-motion region and the motion region through the encoder;
[0031] S503: Perform fusion processing on the encoded motion region and non-motion region to obtain video frame encoded data.
[0032] The present invention also provides an intelligent video coding optimization system for executing an intelligent video coding optimization method, and the optimization system includes:
[0033] A motion region acquisition module, configured to perform preliminary motion detection on the current video frame based on adjacent video frames of the current video frame to be encoded through a motion vector estimation method to obtain a motion region and a non-motion region belonging to the current video frame;
[0034] A segmentation module, configured to segment the current video frame based on a selected coding standard to obtain a plurality of sub-blocks;
[0035] A multi-scale processing module, configured to construct a resolution pyramid of the current video frame and perform multi-scale processing on the plurality of sub-blocks belonging to the motion region based on the resolution pyramid to obtain a multi-scale processed motion region;
[0036] A prediction module, configured to use the motion region after multi-scale processing as input, and utilize a pre-trained deep learning model to predict the current video frame, obtaining the foreground and background of the motion region and first and second quantization parameters respectively corresponding to the foreground and background, wherein the first quantization parameter is less than the second quantization parameter;
[0037] An encoding module, configured to respectively encode the foreground and background regions in the motion region by using the first quantization parameter and the second quantization parameter, and encode the non-motion region by using the second quantization parameter.
[0038] Compared with the related art, an intelligent video encoding optimization method and optimization system provided by the present invention have the following beneficial effects:
[0039] The present invention identifies the motion region and the non-motion region by using a motion vector estimation method; secondly, divides the video frame based on an encoding standard and constructs a resolution pyramid for multi-scale processing; then, uses a pre-trained deep learning model to predict the foreground and background and corresponding quantization parameters; finally, encodes the foreground and background of the motion region and the non-motion region respectively according to the differential quantization parameters, uses a low quantization parameter for the foreground to maintain details, and uses a high quantization parameter for the background and the non-motion region to improve the compression ratio, realizing intelligent processing and efficient encoding of video content, not only improving the compression ratio and encoding efficiency of video encoding, but also ensuring the high quality and low distortion of the encoded video. Description of the Drawings
[0040] Figure 1 It is a flowchart of an intelligent video encoding optimization method provided by the present invention;
[0041] Figure 2 It is a module structure diagram of an intelligent video encoding optimization system provided by the present invention. Detailed Embodiments
[0042] The present invention will be further described in detail below with reference to the drawings and embodiments. It can be understood that the specific embodiments described herein are only used to explain the present invention, rather than limiting the present invention. Additionally, it should be noted that for the sake of description, only parts related to the present invention are shown in the drawings rather than all structures. Furthermore, without conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other.
[0043] It should also be noted that, for the sake of convenience of description, only the parts related to the present invention rather than all the content are shown in the drawings. Before discussing the exemplary embodiments in more detail, it should be mentioned that some exemplary embodiments are described as processes or methods depicted as flowcharts. Although the flowcharts describe the operations (or steps) as being processed sequentially, many of the operations can be performed in parallel, concurrently, or simultaneously. In addition, the order of the operations can be rearranged. The processing can be terminated when its operations are completed, but there can also be additional steps not included in the drawings. The processing can correspond to a method, function, procedure, subroutine, subprogram, etc.
[0044] Embodiment 1
[0045] The present invention provides an intelligent video coding optimization method. Referring to Figure 1 as shown, the optimization method includes the following steps:
[0046] S1: Based on the adjacent video frames of the current video frame in the video to be encoded, perform preliminary motion detection on the current video frame through a motion vector estimation method to obtain the motion area and non - motion area belonging to the current video frame.
[0047] In this embodiment, the specific process of the preliminary motion detection is as follows: First, select the previous frame and the next frame of the current video frame as adjacent frames. Then, use a motion vector estimation method including but not limited to the optical flow method or the block matching method to analyze the pixel - level or block - level changes between the current frame and the adjacent frames, so as to identify the motion area in the current frame. The non - motion area is obtained by the exclusion method (that is, the part of the current video frame excluding the motion area).
[0048] S2: Segment the current video frame based on the selected coding standard to obtain a plurality of sub - blocks.
[0049] In this embodiment, based on the selected coding standard, the current video frame is evenly or unevenly segmented to obtain a plurality of sub - blocks. The size and shape of the segmentation are adjusted according to the requirements of the coding standard and the characteristics of the video content, aiming to divide the complex video frame into smaller units that are easier to process for subsequent coding optimization.
[0050] S3: Construct a resolution pyramid of the current video frame, and perform multi - scale processing on the plurality of sub - blocks belonging to the motion area based on the resolution pyramid to obtain the motion area after multi - scale processing.
[0051] In this embodiment, first, a resolution pyramid of the current video frame is constructed, that is, starting from the original resolution, the resolution of the image is gradually reduced to generate a series of images with different scales. Then, for the sub-blocks in the motion area obtained by preliminary motion detection, multi-scale processing is performed on the resolution pyramid, aiming to capture the feature information of the motion area at different scales.
[0052] S4: Use the motion area after multi-scale processing as input, and use a pre-trained deep learning model to predict the current video frame, obtaining the foreground and background of the motion area and the first quantization parameter and the second quantization parameter corresponding to the foreground and background respectively, where the first quantization parameter is less than the second quantization parameter.
[0053] In this embodiment, the motion area after multi-scale processing is used as input and fed into a pre-trained deep learning model (such as the neural network CNN). This model has been trained with a large amount of labeled data and has learned how to identify the foreground and background from the input data and obtain the first quantization parameter and the second quantization parameter corresponding to them respectively.
[0054] S5: Use the first quantization parameter and the second quantization parameter to encode the foreground and background areas in the motion area respectively, and use the second quantization parameter to encode the non-motion area.
[0055] In this embodiment, according to the first quantization parameter and the second quantization parameter obtained by deep learning prediction, the encoder is used to encode the non-motion area and the foreground and background areas in the motion area respectively. The foreground area is finely encoded using the smaller first quantization parameter to maintain the image quality and details; the non-motion area and the background area are roughly encoded using the larger second quantization parameter to improve the compression ratio. This step finally realizes the efficient and high-quality encoding of the video.
[0056] Specifically, step S1 includes the following steps:
[0057] S101: Use a feature point detection algorithm to identify feature points in the current frame and adjacent video frames of the video to be encoded, determine the correspondence of the identified feature points between frames, and calculate the motion vectors of each feature point between the current frame and adjacent video frames.
[0058] In this embodiment, the specific implementation process of step S101 is as follows:
[0059] Feature point detection: First, apply feature point detection algorithms, including but not limited to SIFT (Scale-Invariant Feature Transform), SURF (Speeded-Up Robust Features), or ORB (Oriented FAST and Rotated BRIEF), to the current frame and its adjacent frames (i.e., the previous frame and the next frame) of the video to be encoded. These algorithms can identify points with significant features (such as corner points, edge points, etc.) in the image as feature points.
[0060] Determine correspondence: Through feature matching algorithms, including but not limited to FLANN matcher, BF matcher, find the correspondence of feature points between the current frame and adjacent frames. These correspondences reflect the motion trajectories of feature points between frames.
[0061] Motion vector calculation: According to the determined correspondence of feature points, calculate the motion vector of each feature point between the current frame and adjacent frames. The motion vector contains information about the direction and distance of the feature point's movement between frames.
[0062] S102: By analyzing the changes in the direction and magnitude of the motion vectors, identify and locate the motion regions in the current video frame, and obtain the non-motion regions through the exclusion method.
[0063] In this embodiment, the specific implementation process of step S102 is as follows:
[0064] Analyze motion vectors: Conduct statistical analysis on the calculated motion vectors, including direction, magnitude, and change rate, etc. By comparing the motion vectors of feature points between adjacent frames, identify regions with significant motion characteristics.
[0065] Identify motion regions: According to the analysis results of the motion vectors, identify the motion regions in the current frame through a preset threshold. Exemplarily, the motion regions can be divided based on the consistency of the magnitude and direction of the motion vectors.
[0066] Locate motion regions: After identifying the motion regions, accurately locate the motion regions in the image space of the current frame through coordinate mapping. The non-motion regions are obtained through the exclusion method (i.e., the part of the current video frame excluding the motion regions).
[0067] S103: Mark the obtained motion regions and non-motion regions with visual markers, where the visual markers include boundary markers and color markers.
[0068] In this embodiment, the specific implementation process of step S103 is as follows:
[0069] Boundary marking: Use boundary lines or bounding boxes to mark the identified moving areas and non-moving areas. Among them, the boundary lines or bounding boxes should be drawn closely along the edges of the moving areas and non-moving areas so that the range of the moving areas can be clearly identified in subsequent processing.
[0070] Color marking: To more intuitively display the moving areas, specific colors or patterns can be filled in the moving areas and non-moving areas. The selection of colors or patterns should ensure a distinct contrast with the background for easy observation and identification.
[0071] Visual marking integration: Integrate the boundary marking and color marking into the current video frame to form a video frame image containing clear markings for the moving areas and non-moving areas.
[0072] Specifically, step S2 includes the following steps:
[0073] S201: Perform parameter initialization settings according to the selected coding standard.
[0074] In this embodiment, the specific implementation process of step S201 is as follows:
[0075] Select the coding standard: First, according to the application requirements and environmental conditions, select a suitable video coding standard, such as H.264 / AVC, H.265 / HEVC.
[0076] Initialize the coding parameters: According to the selected coding standard, initialize the relevant coding parameters. These parameters may include but are not limited to: coding rate, frame rate, resolution, quantization step size, prediction mode (such as intra-frame prediction, inter-frame prediction), number of reference frames, etc. The settings of these parameters will directly affect the quality and compression ratio of the encoded video.
[0077] Configure the block partitioning mechanism: According to the requirements of the coding standard, configure the corresponding block partitioning mechanism, which includes determining the size, shape, and partitioning method of the sub-blocks (such as uniform partitioning, content-based partitioning).
[0078] S202: Divide the current video frame into multiple sub-blocks using the block partitioning mechanism corresponding to the selected coding standard.
[0079] In this embodiment, the specific implementation process of step S202 is as follows:
[0080] Apply the block partitioning mechanism: Use the block partitioning mechanism corresponding to the coding standard configured in step S201 to divide the current video frame into multiple sub-blocks. These sub-blocks can be rectangular blocks of a fixed size (such as 16x16, 32x32, etc.) or blocks adaptively partitioned based on content (such as blocks partitioned according to texture complexity, motion characteristics, etc.).
[0081] Optimized Block Division: During the process of dividing sub - blocks, optimization and adjustment can be made according to the characteristics of the video content. Exemplarily, in areas with complex textures or intense movements, smaller sub - blocks can be divided to improve the encoding accuracy; while in areas with smooth textures or static states, larger sub - blocks can be divided to reduce the encoding complexity.
[0082] Record Block Information: Record information such as the position, size of each sub - block, and its relationship with other sub - blocks. This information will be used in the subsequent encoding process.
[0083] Specifically, step S3 includes the following steps:
[0084] S301: Starting from the original resolution of the current video frame, construct a resolution pyramid with different resolution levels by gradually reducing the resolution.
[0085] In this embodiment, the specific implementation process of step S301 is as follows:
[0086] Determine the Resolution Level: First, according to the original resolution of the video frame and the requirements of subsequent processing, determine the number of layers of the resolution pyramid and the resolution of each layer. Generally, the more layers the pyramid has, the richer the scale information that can be captured, but the computational complexity will also increase accordingly.
[0087] Gradually Reduce the Resolution: Starting from the original resolution of the current video frame, through image down - sampling techniques, including but not limited to bilinear interpolation, bicubic interpolation, and down - sampling after Gaussian filtering, gradually reduce the resolution of the image to generate a series of images with different resolution levels. These images are arranged in descending order of resolution to form the resolution pyramid.
[0088] S302: In each layer of the resolution pyramid, select the sub - blocks belonging to the motion region.
[0089] In this embodiment, the specific implementation process of step S302 is as follows:
[0090] Locate the Motion Region: Based on the results of the preliminary motion detection in step S1, locate the motion region in the current video frame.
[0091] Map to Each Layer of the Pyramid: Map the position information of the motion region to each layer of the resolution pyramid to determine which sub - blocks in each layer belong to the motion region.
[0092] Select Sub - blocks: In each layer of the pyramid, select the sub - blocks belonging to the motion region according to the mapping results. These sub - blocks will be used as the input for subsequent multi - scale processing.
[0093] S303: Perform multi - scale processing on all the sub - blocks belonging to the motion region in each selected layer of the pyramid.
[0094] In this embodiment, the specific implementation process of step S303 is as follows:
[0095] Apply the multi-scale analysis method: For all sub-blocks belonging to the motion region in each selected layer of the pyramid, apply the pyramid decomposition method for multi-scale processing, which can decompose the sub-blocks into components of different scales, thereby capturing the feature information of the sub-blocks at different scales.
[0096] Feature extraction: During the multi-scale processing, extract the feature information of the sub-blocks at different scales, such as edges, textures, motion vectors, etc. This feature information will be used for subsequent coding optimization and deep learning prediction.
[0097] S304: Merge all sub-blocks of the motion region in each layer of the pyramid after multi-scale processing to obtain the motion region after multi-scale processing.
[0098] In this embodiment, the specific implementation process of step S304 is as follows:
[0099] Merge the processing results: Merge the processing results of all sub-blocks of the motion region in each layer of the pyramid after multi-scale processing, which includes integrating and fusing the feature information at different scales to form a more comprehensive and in-depth understanding of the motion region.
[0100] Generate the motion region after multi-scale processing: Based on the merged processing results, generate an image or feature map of the motion region after multi-scale processing, and this image or feature map will be used as the input for subsequent deep learning prediction and coding optimization.
[0101] Specifically, step S4 includes the following steps:
[0102] S401: Use a pre-trained deep learning model to extract features from the input motion region after multi-scale processing to obtain the foreground features and background features in the motion region, where the deep learning model is a neural network model trained based on a dataset of annotated video frames.
[0103] In this embodiment, the specific implementation process of step S401 is as follows:
[0104] Input the motion region after multi-scale processing: Take the motion region after multi-scale processing in step S3 as the input and send it into the pre-trained deep learning model.
[0105] Pre-trained deep learning model: This model is a neural network model trained based on a dataset of annotated video frames (including foreground and background annotation information).
[0106] Feature extraction: Using a deep learning model to extract features from the input multi-scale motion regions. The model will automatically capture and learn the differences between the foreground and the background, and generate corresponding feature representations.
[0107] Output foreground and background features: After being processed by the deep learning model, the foreground and background features in the motion regions are output. These features will be used for subsequent foreground-background separation.
[0108] S402: Generate a probability map based on the foreground and background features, and obtain the foreground and the background based on the probability map, where the probability map is used to represent the probabilities of sub-blocks in the motion region belonging to the foreground and the background.
[0109] In this embodiment, the specific implementation process of step S402 is as follows:
[0110] Generate a probability map: Generate a probability map according to the foreground and background features extracted in step S401. This probability map is used to represent the probabilities of each sub-block in the motion region belonging to the foreground or the background. The closer the probability value is to 1, the higher the confidence that the pixel belongs to the corresponding category.
[0111] Threshold processing: Set a threshold (such as 0.5), perform threshold processing on the probability map, classify the pixels with probability values greater than the threshold as the foreground, and classify the pixels with probability values less than the threshold as the background.
[0112] Obtain the foreground and the background: After threshold processing, obtain the foreground and the background in the motion region. These regions will be used for subsequent differential allocation of quantization parameters.
[0113] S403: Perform differential allocation of quantization parameters for the obtained foreground and background to obtain a first quantization parameter and a second quantization parameter corresponding to the foreground and the background respectively.
[0114] In this embodiment, the specific implementation process of step S403 is as follows:
[0115] Analyze the foreground and background characteristics: Analyze the foreground and background obtained in step S402 to obtain the characteristics of the foreground and the background. These characteristics include motion complexity, texture complexity, and brightness contrast.
[0116] Determine the quantization parameters: Based on the characteristics of the foreground and the background, and the goals of video coding (such as compression ratio, video quality, etc.), determine the corresponding first quantization parameter and second quantization parameter respectively. For the foreground region with high motion complexity and rich texture details, a lower quantization parameter can be used to maintain high coding quality; while for the relatively static and simple-textured background region, a higher quantization parameter can be used to save coding bit rate.
[0117] Differential allocation: Apply the determined first quantization parameter and second quantization parameter to the encoding processes of the foreground and background regions respectively. In this way, while maintaining the overall video quality, the encoding efficiency can be further optimized.
[0118] Specifically, step S5 includes the following steps:
[0119] S501: Configure the corresponding encoding parameters according to the first quantization parameter and the second quantization parameter.
[0120] In this embodiment, the specific implementation process of step S501 is as follows:
[0121] Based on quantization parameter configuration: According to the first quantization parameter and the second quantization parameter determined in step S403, configure the corresponding encoding parameters for the foreground and background regions respectively. These encoding parameters may include but are not limited to prediction modes (intra-frame prediction, inter-frame prediction), motion vectors, number of reference frames, quantization step size of transform coefficients, etc.
[0122] Optimization and adjustment: During the process of configuring the encoding parameters, perform optimization and adjustment according to the characteristics of the video content and the encoding requirements. Exemplarily, for the foreground region with intense motion, more reference frames and a finer motion search algorithm can be configured; for the background region with simple texture, a larger quantization step size and fewer reference frames can be adopted to reduce the encoding complexity.
[0123] S502: Use the configured first quantization parameter and the corresponding encoding parameters to encode the foreground of the motion region through an encoder, and use the configured second quantization parameter and the corresponding encoding parameters to encode the non-motion region and the background of the motion region through the encoder.
[0124] In this embodiment, the specific implementation process of step S502 is as follows:
[0125] Foreground encoding: Use the configured first quantization parameter and the corresponding encoding parameters to encode the foreground region through an encoder. The encoder will perform compression processing on the pixel values, motion information, etc. of the foreground region to generate the encoded data of the foreground region.
[0126] Non-motion region and background encoding: Similarly, use the configured second quantization parameter and the corresponding encoding parameters to encode the non-motion region and the background of the motion region through the encoder. Since the characteristics and encoding parameters of the background region are different from those of the foreground region, the encoding process will also be different.
[0127] Parallel processing: To improve the encoding efficiency, the foreground and background regions can be encoded in parallel.
[0128] S503: Perform a fusion process on the encoded motion region and non-motion region to obtain video frame encoded data.
[0129] In this embodiment, the specific implementation process of step S503 is as follows:
[0130] Data integration: Integrate the encoded data of the encoded motion region and non-motion region, which includes merging the encoded bitstreams of both into a complete video frame encoded data.
[0131] Synchronization processing: During the integration process, it is necessary to ensure that the encoded data of the motion region and non-motion region are synchronized in both time and space.
[0132] Output encoded data: After the integration is completed, output the encoded data of the video frame, and these data can be used in subsequent application scenarios such as storage, transmission, or decoding and playback.
[0133] The working principle of an intelligent video coding optimization method provided by the present invention is as follows:
[0134] First, through the motion vector estimation method, this technical solution uses the pixel changes between the current video frame and its adjacent frames to initially detect the motion region and non-motion region in the current frame. The key to this step is to capture the dynamic changes between video frames and provide a basis for motion information for subsequent processing.
[0135] Next, according to the selected video coding standard, segment the current video frame into multiple smaller sub-blocks. This segmentation helps to process video content more finely, especially for complex scenes or high dynamic range videos, and can more effectively utilize coding resources.
[0136] Then, construct a resolution pyramid of the current video frame. The resolution pyramid is a multi-scale representation method that creates a series of images with different scales by gradually reducing the image resolution. In this step, special attention is paid to the sub-blocks in the motion region and they are processed at multiple scales. Multi-scale processing can reveal motion features at different scales and enhance the accuracy and robustness of motion detection.
[0137] Subsequently, use the motion region after multi-scale processing as input and feed it into a pre-trained deep learning model. This deep learning model has been trained with a large number of labeled video frames and has learned how to identify foreground (such as moving objects) and background in the motion region and output corresponding quantization parameters according to their characteristics. Specifically, the model will output two quantization parameters: the first quantization parameter is for the foreground region and its value is smaller, aiming to maintain image quality and details; the second quantization parameter is for the background region and its value is larger to increase the compression ratio within an acceptable range.
[0138] Finally, according to the quantization parameters output by the deep learning model, the encoder differentially encodes the foreground and background in the motion region and the non-motion region. The foreground region is finely encoded using smaller quantization parameters to preserve key information and details; while the background region and the non-motion region are roughly encoded using larger quantization parameters to reduce data redundancy and improve compression efficiency. This differential encoding strategy not only ensures the overall quality of the video but also significantly improves the encoding efficiency, making the video data more efficient during storage and transmission.
[0139] Embodiment 2
[0140] The present invention also provides an intelligent video encoding optimization system for performing an intelligent video encoding optimization method, as shown in Figure 2 The optimization system includes:
[0141] A motion region acquisition module 100, configured to perform preliminary motion detection on the current video frame based on adjacent video frames of the current video frame to be encoded by a motion vector estimation method, and obtain a motion region and a non-motion region belonging to the current video frame.
[0142] A segmentation module 200, configured to segment the current video frame based on a selected encoding standard to obtain a plurality of sub-blocks.
[0143] A multi-scale processing module 300, configured to construct a resolution pyramid of the current video frame and perform multi-scale processing on the plurality of sub-blocks belonging to the motion region based on the resolution pyramid to obtain a motion region after multi-scale processing.
[0144] A prediction module 400, configured to use the motion region after multi-scale processing as an input, and predict the current video frame using a pre-trained deep learning model to obtain the foreground and background of the motion region and first and second quantization parameters respectively corresponding to the foreground and background, wherein the first quantization parameter is smaller than the second quantization parameter.
[0145] An encoding module 500, configured to encode the foreground and background regions in the motion region using the first and second quantization parameters respectively, and encode the non-motion region using the second quantization parameter.
[0146] This application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, and the combination of flows and / or blocks in the flowchart and / or block diagram can also be implemented. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in one Figure 1 flow or multiple flows and / or blocks Figure 1 or means for implementing the functions specified in one block or multiple blocks.
[0147] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing relevant hardware through a program. This program can be stored in a computer-readable storage medium, and the storage medium includes read-only memory (ROM), random access memory (RAM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), one-time programmable read-only memory (OTPROM), electrically-erasable programmable read-only memory (EEPROM), compact disc read-only memory (CD-ROM), or other optical disc memories, magnetic disk memories, tape memories, or any other medium that can be used to carry or store data and is computer-readable.
[0148] It should also be noted that the term "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, commodity, or device comprising a series of elements not only includes those elements but also includes other elements not expressly listed, or still includes elements inherent to such process, method, commodity, or device. Without further limitation, an element defined by the statement "comprising one..." does not exclude the existence of additional identical elements in the process, method, commodity, or device comprising the element.
Claims
1. An intelligent video coding optimization method, characterized in that The optimization method includes the following steps: S1: Based on adjacent video frames of the current video frame in the video to be encoded, perform preliminary motion detection on the current video frame through a motion vector estimation method to obtain the motion area and non-motion area belonging to the current video frame; S2: Segment the current video frame based on the selected encoding standard to obtain multiple sub-blocks; S3: Construct a resolution pyramid of the current video frame, and perform multi-scale processing on the multiple sub-blocks belonging to the motion area based on the resolution pyramid to obtain the motion area after multi-scale processing; Step S3 includes the following steps: S301: Starting from the original resolution of the current video frame, construct a resolution pyramid with different resolution levels by gradually reducing the resolution; S302: In each layer of the resolution pyramid, select the sub-blocks belonging to the motion area; S303: Perform multi-scale processing on all sub-blocks belonging to the motion area in each selected layer of the pyramid, where the multi-scale processing includes extracting feature information of the sub-blocks at different scales, and the feature information includes at least one of edges, textures, and motion vectors; S304: Fuse the processing results of all sub-blocks of the motion area in each layer of the pyramid after multi-scale processing to obtain the motion area after multi-scale processing; S4: Use the motion area after multi-scale processing as input, and use a pre-trained deep learning model to predict the current video frame to obtain the foreground and background of the motion area and the first quantization parameter and the second quantization parameter corresponding to the foreground and background respectively, where the first quantization parameter is less than the second quantization parameter; S5: Use the first quantization parameter and the second quantization parameter to encode the foreground and background areas in the motion area respectively, and use the second quantization parameter to encode the non-motion area.
2. The intelligent video coding optimization method according to claim 1, characterized in that Step S1 includes the following steps: S101: Use a feature point detection algorithm to identify feature points in the current frame and adjacent video frames of the video to be encoded, determine the correspondence of the identified feature points between frames, and calculate the motion vectors of each feature point between the current frame and the adjacent video frames; S102: By analyzing the changes in the direction and magnitude of the motion vectors, identify and locate the motion area in the current video frame, and obtain the non-motion area through the exclusion method; S103: Mark the obtained motion area and non-motion area with visual markers, where the visual markers include boundary markers and color markers.
3. An intelligent video coding optimization method according to claim 2, characterized in that, Step S2 includes the following steps: S201: Perform parameter initialization settings according to the selected encoding standard; S202: Use the block division mechanism corresponding to the selected encoding standard to divide the current video frame into multiple sub-blocks.
4. An intelligent video coding optimization method according to claim 3, characterized in that Step S4 includes the following steps: S401: Use a pre-trained deep learning model to extract features from the input motion area after multi-scale processing to obtain the foreground features and background features in the motion area, where the deep learning model is a neural network model trained based on a dataset of annotated video frames; S402: Generate a probability map based on the foreground features and background features, and obtain the foreground and background based on the probability map, where the probability map is used to represent the probabilities of sub-blocks in the motion region belonging to the foreground and the background; S403: Perform differential allocation of quantization parameters for the obtained foreground and background to obtain a first quantization parameter and a second quantization parameter corresponding to the foreground and background respectively.
5. An intelligent video coding optimization method according to claim 4, characterized in that Step S5 includes the following steps: S501: Configure corresponding coding parameters according to the first quantization parameter and the second quantization parameter; S502: Use the configured first quantization parameter and corresponding coding parameters to encode the foreground of the motion region through an encoder, and use the configured second quantization parameter and corresponding coding parameters to encode the background of the non-motion region and the motion region through the encoder; S503: Perform a fusion process on the encoded motion region and non-motion region to obtain video frame encoded data.
6. An intelligent video coding optimization system for performing an intelligent video coding optimization method according to any one of claims 1 to 5, characterized in that, The optimization system includes: A motion region acquisition module, configured to perform preliminary motion detection on a current video frame of a video to be encoded based on adjacent video frames of the current video frame through a motion vector estimation method, and obtain a motion region and a non-motion region belonging to the current video frame; A segmentation module, configured to segment the current video frame based on a selected coding standard to obtain a plurality of sub-blocks; A multi-scale processing module, configured to construct a resolution pyramid of the current video frame, and perform multi-scale processing on the plurality of sub-blocks belonging to the motion region based on the resolution pyramid to obtain a multi-scale processed motion region; Specifically, the multi-scale processing module is configured to start from the original resolution of the current video frame and construct a resolution pyramid with different resolution levels by gradually reducing the resolution; In each layer of the resolution pyramid, select the sub-blocks belonging to the motion region; Perform multi-scale processing on all the sub-blocks belonging to the motion region in each selected layer of the pyramid, where the multi-scale processing includes extracting feature information of the sub-blocks at different scales, and the feature information includes at least one of edges, textures, and motion vectors; Fuse the processing results of all the sub-blocks of the motion region in each layer of the pyramid after multi-scale processing to obtain a multi-scale processed motion region; A prediction module, configured to use the multi-scale processed motion region as an input, and perform prediction on the current video frame by using a pre-trained deep learning model to obtain the foreground and background of the motion region and a first quantization parameter and a second quantization parameter corresponding to the foreground and background respectively, where the first quantization parameter is less than the second quantization parameter; An encoding module, configured to use the first quantization parameter and the second quantization parameter to encode the foreground and background regions in the motion region respectively, and use the second quantization parameter to encode the non-motion region.
Citation Information
Patent Citations
Method, device and system for regional division / coding of image
CN101882316A
Method and system for detecting picture offset of camera device
CN102609957A
Target tracking method based on multi-time-step pyramid codec
CN112288776A
Multistage ground wire defect identification method based on parallax auxiliary semantic segmentation
CN116385364A
Image processing method, electronic device, and computer storage medium
WO2024001345A1