Video encoding method, apparatus, device, storage medium and computer program product
By extracting features from multiple dimensions and analyzing content complexity, video coding parameters are dynamically adjusted, solving the problem of unreasonable allocation of coding resources in existing technologies and achieving reasonable allocation of video coding resources and improved image quality.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIJING QIHOOD TECHNOLOGY CO LTD
- Filing Date
- 2026-04-24
- Publication Date
- 2026-07-24
AI Technical Summary
Existing video coding methods use fixed or empirically based coding parameters, which cannot adapt to the differentiated coding needs of different video content. This leads to unreasonable allocation of coding resources and problems such as insufficient image quality in high-complexity segments and redundant bitrate in low-complexity segments.
By extracting multi-dimensional features from the video file to be processed, content complexity information is generated, and the encoding parameters are adjusted accordingly to dynamically match the complexity of the video content. This includes the temporal correlation and fusion mapping of motion features, texture features, and visual parameter change features, generating a content complexity quantification value, dividing the video into segments, and matching the encoding parameter combination.
It achieves a reasonable allocation of video encoding resources, improves encoding efficiency and image quality, avoids image quality fluctuations caused by sudden changes in encoding parameters, and adapts to the differentiated needs of different video content.
Smart Images

Figure CN122457772A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of video processing technology, and in particular to a video encoding method, apparatus, device, storage medium, and computer program product. Background Technology
[0002] Currently, relevant video encoding methods typically employ fixed or empirically based encoding parameters. However, different video content exhibits significant differences in spatiotemporal characteristics. Using fixed or empirically based encoding parameters cannot adapt to the differentiated encoding needs of different video content, resulting in unreasonable allocation of encoding resources (such as insufficient image quality and bitrate gaps in high-complexity segments; and bitrate redundancy and resource waste in low-complexity segments). Summary of the Invention
[0003] The main objective of this application is to provide a video encoding method, apparatus, device, storage medium, and computer program product, aiming to solve the technical problem that related video encoding methods use fixed or empirically based encoding parameters for video encoding, which cannot adapt to the differentiated encoding needs of different video content, resulting in unreasonable allocation of encoding resources.
[0004] To achieve the above objectives, this application provides a video encoding method, the video encoding method comprising: In response to a video encoding request, multi-dimensional feature extraction is performed on the video file to be processed to obtain multi-dimensional features; Based on the multi-dimensional features, generate the content complexity information corresponding to the video file to be processed; The encoding parameters are adjusted based on the content complexity, and the video file to be processed is encoded according to the adjusted encoding parameters.
[0005] Optionally, generating the content complexity information corresponding to the video file to be processed based on the multi-dimensional features includes: The multi-dimensional features are input into a preset content complexity prediction model; The multi-dimensional features are fused and analyzed using the preset content complexity prediction model to output the content complexity information corresponding to the video file to be processed.
[0006] Optionally, the step of fusing and analyzing the multi-dimensional features using the preset content complexity prediction model to output the content complexity information corresponding to the video file to be processed includes: The preset content complexity prediction model is used to perform temporal dimension feature association and fusion mapping on the multi-dimensional features to generate a content complexity quantification value corresponding to the video timeline. Based on the content complexity quantification value, the content complexity information corresponding to the video file to be processed is generated.
[0007] Optionally, the multi-dimensional features include motion features, texture features, and visual parameter change features. The step of performing temporal-dimensional feature association and fusion mapping on the multi-dimensional features using the preset content complexity prediction model to generate a content complexity quantification value corresponding to the video timeline includes: Based on the motion features, texture features, and visual parameter change features, the preset content complexity prediction model calculates the motion intensity, detail richness, and scene change degree of the video timeline. The motion intensity, detail richness, and scene change degree of the image are temporally correlated and weighted to generate a content complexity quantification value corresponding to the video timeline.
[0008] Optionally, before inputting the multi-dimensional features into the preset content complexity prediction model, the method further includes: Construct a training dataset containing video samples of multiple scene types, and label the content complexity ground truth labels of each video sample in the training dataset at the corresponding temporal position; Multi-dimensional feature samples are extracted from each video sample, and these multi-dimensional feature samples are used as input to the initial content complexity prediction model, while the corresponding ground truth labels of content complexity are used as supervised training targets. The initial content complexity prediction model is trained in a supervised manner using machine learning algorithms or deep neural network algorithms until the model loss converges, thereby obtaining the preset content complexity prediction model.
[0009] Optionally, adjusting the encoding parameters based on the content complexity and encoding the video file to be processed according to the adjusted encoding parameters includes: Based on the continuous distribution characteristics of the content complexity information on the video timeline, the video file to be processed is divided into multiple continuous video segments; Based on the content complexity level corresponding to each video segment, determine the combination of encoding parameters that matches the content complexity level; The encoding parameters are adjusted based on the combination of the encoding parameters, and each video segment is encoded according to the adjusted encoding parameters.
[0010] Optionally, determining the combination of encoding parameters matching the content complexity level based on the content complexity level corresponding to each video segment includes: Obtain the content importance score corresponding to each video segment on the video timeline of the video file to be processed; Based on the content importance score and content complexity level of each video segment, the encoding control weights of the video segments are generated. Based on the encoding control weights, the combination of encoding parameters corresponding to the video segments is dynamically adjusted.
[0011] Optionally, before adjusting the encoding parameters based on the combination of encoding parameters and encoding each video segment according to the adjusted encoding parameters, the method further includes: Identify the parameter differences between the encoding parameter combinations corresponding to adjacent video segments; Determine the frame connection boundary between adjacent video segments, and use the frame connection boundary as a reference to define a transition frame interval of a preset length; Based on the parameter difference information and the total number of frames in the transition frame interval, an intermediate coding parameter sequence that gradually changes within the transition frame interval is generated. The intermediate encoding parameter sequence is used to encode video frames within the transition frame interval.
[0012] Optionally, the multi-dimensional features include motion features, texture features, and visual parameter variation features. The step of extracting multi-dimensional features from the video file to be processed to obtain multi-dimensional features includes: Perform inter-frame difference analysis on consecutive video frames of the video file to be processed, and extract motion features that reflect the degree of change in content between frames; Image detail analysis is performed on single video frames of the video file to be processed to extract texture features that reflect the richness of the content of the picture. Pixel statistical analysis is performed on the video frames of the video file to be processed to extract visual parameter change features that reflect the fluctuation of visual attributes of the image.
[0013] Optionally, the step of performing inter-frame difference analysis on consecutive video frames of the video file to be processed, and extracting motion features reflecting the degree of change in content between frames, includes: Select adjacent video frame groups of the video file to be processed at a preset frame interval, and calculate the inter-frame motion vector and pixel difference matrix; Based on the inter-frame motion vector and pixel difference matrix, the inter-frame motion amplitude, motion area proportion and scene switching probability are statistically analyzed. Motion features corresponding to the video timeline are generated based on the inter-frame motion amplitude, the proportion of the motion region, and the scene switching probability.
[0014] Optionally, the step of performing image detail analysis on single video frames of the video file to be processed and extracting texture features reflecting the richness of the image content includes: The single-frame video file to be processed is uniformly divided into blocks to obtain multiple image sub-blocks of the same size. Spatial gradient statistics or frequency domain transform analysis are performed on each image sub-block to obtain the texture detail richness information of each sub-block; Based on the texture detail richness information of all image sub-blocks in the whole frame, the texture richness and spatial distribution characteristics of the whole frame are statistically analyzed to generate texture features corresponding to the video timeline.
[0015] Optionally, the step of performing pixel statistical analysis on the video frames of the video file to be processed and extracting visual parameter change features reflecting fluctuations in visual attributes of the image includes: Perform full-frame pixel value statistics on a single frame of the video file to be processed to obtain the core visual parameters of the single frame. The core visual parameters include at least one of the following: average brightness, contrast, and dynamic range. The core visual parameters of consecutive video frames are compared along the video timeline, and the fluctuation amplitude and frequency of parameter changes are statistically analyzed. Based on the fluctuation amplitude of the parameters and the frequency of mutations, visual parameter change features corresponding to the video timeline are generated.
[0016] Furthermore, to achieve the above objectives, this application also proposes a video encoding apparatus, which includes: The feature extraction module is used to extract multi-dimensional features from the video file to be processed in response to the video encoding request, and obtain multi-dimensional features. The information generation module is used to generate content complexity information corresponding to the video file to be processed based on the multi-dimensional features; The video encoding module is used to adjust the encoding parameters based on the content complexity and to encode the video file to be processed according to the adjusted encoding parameters.
[0017] Optionally, the information generation module is further configured to input the multi-dimensional features into a preset content complexity prediction model; perform fusion analysis on the multi-dimensional features through the preset content complexity prediction model, and output the content complexity information corresponding to the video file to be processed.
[0018] Optionally, the information generation module is further configured to perform temporal dimension feature association and fusion mapping on the multi-dimensional features through the preset content complexity prediction model to generate a content complexity quantification value corresponding to the video timeline; and generate content complexity information corresponding to the video file to be processed based on the content complexity quantification value.
[0019] Optionally, the multi-dimensional features include motion features, texture features, and visual parameter change features. The information generation module is further configured to calculate the image motion intensity, image detail richness, and scene change degree corresponding to the video timeline based on the motion features, texture features, and visual parameter change features using the preset content complexity prediction model; and to perform temporal correlation and weighted fusion of the image motion intensity, detail richness, and scene change degree to generate a content complexity quantification value corresponding to the video timeline.
[0020] Optionally, the video encoding device further includes: The model training module is used to construct a training dataset containing video samples of multiple scene types, and to label the content complexity ground truth labels of each video sample in the training dataset at the corresponding temporal positions; to extract multi-dimensional feature samples of each video sample, and to use the multi-dimensional feature samples as input to the initial content complexity prediction model, with the corresponding content complexity ground truth labels as supervised training targets; to use machine learning algorithms or deep neural network algorithms to perform supervised training on the initial content complexity prediction model until the model loss converges, thereby obtaining the preset content complexity prediction model.
[0021] In addition, to achieve the above objectives, this application also proposes a video encoding device, which includes a memory, a processor, and a video encoding program stored in the memory and executable on the processor, the video encoding program being configured to implement the video encoding method as described above.
[0022] In addition, to achieve the above objectives, this application also proposes a storage medium storing a video encoding program, which, when executed by a processor, implements the video encoding method as described above.
[0023] In addition, to achieve the above objectives, this application also provides a computer program product, which includes a video encoding program that, when executed by a processor, implements the video encoding method as described above.
[0024] One or more technical solutions proposed in this application have at least the following technical effects: This application discloses a method for responding to a video encoding request by extracting multi-dimensional features from a video file to be processed, obtaining multi-dimensional features, generating content complexity information corresponding to the video file to be processed based on the multi-dimensional features, adjusting encoding parameters based on the content complexity, and encoding the video file to be processed according to the adjusted encoding parameters. Because this application generates content complexity information corresponding to the video file to be processed based on the multi-dimensional features of the video file to be processed, and dynamically adjusts the encoding parameters of the video encoding based on the content complexity, it can adapt to the differentiated encoding requirements of different video content, thereby improving the rationality of encoding resource allocation. Attached Figure Description
[0025] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0026] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0027] Figure 1 This is a flowchart illustrating the first embodiment of the video encoding method of this application; Figure 2 This is a flowchart illustrating the second embodiment of the video encoding method of this application; Figure 3 This is a flowchart illustrating the third embodiment of the video encoding method of this application; Figure 4 This is a flowchart illustrating the fourth embodiment of the video encoding method of this application; Figure 5 This is a system structure diagram of an embodiment of the video encoding method of this application; Figure 6 This is a schematic diagram of the module structure of the video encoding device according to an embodiment of this application; Figure 7 This is a schematic diagram of the device structure of the hardware operating environment involved in the video encoding method in the embodiments of this application.
[0028] The realization of the purpose, functional features and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0029] It should be understood that the specific embodiments described herein are merely illustrative of the technical solutions of this application and are not intended to limit this application.
[0030] To better understand the technical solution of this application, a detailed description will be provided below in conjunction with the accompanying drawings and specific implementation methods.
[0031] This application provides a video encoding method, referring to... Figure 1 , Figure 1 This is a flowchart illustrating the first embodiment of the video encoding method of this application.
[0032] In the first embodiment, the video encoding method includes: Step S10: In response to the video encoding request, perform multi-dimensional feature extraction on the video file to be processed to obtain multi-dimensional features.
[0033] It should be noted that the executing entity in this embodiment can be a video encoding device with data processing, network communication, and program execution functions, such as a computer, or other electronic devices capable of performing the same or similar functions. This embodiment does not impose any limitations on this. The video encoding device may be equipped with a video encoder or a video processing software development kit (SDK), and the video encoding method may be integrated into the video encoder or video processing SDK.
[0034] It should be understood that a video encoding request can refer to an instruction signal that triggers the video encoding process, containing basic encoding configuration information such as the storage path of the video file to be processed, the encoding output specifications, and the target application scenario. The video file to be processed can refer to video data that needs to be compressed and encoded. Multi-dimensional feature extraction can refer to the process of extracting multiple types of feature information that can characterize the spatiotemporal attributes of the video content from the frame sequence of the video file to be processed using digital image processing and computer vision algorithms. Multi-dimensional features can refer to a set of multiple features extracted from the video that can comprehensively reflect the spatiotemporal variation characteristics of the video content. Multi-dimensional features include at least one of motion features, texture features, and visual parameter variation features.
[0035] In the specific implementation, the video encoding request serves as the trigger entry point. Before starting the encoding process, a multi-dimensional feature extraction operation is performed on the video file to be processed. This step is not an isolated extraction of a single feature, but rather a multi-dimensional feature acquisition that simultaneously covers the temporal and spatial dimensions of the video: For the temporal dimension, motion features are extracted using algorithms such as inter-frame differencing and motion estimation to capture the dynamic changes in the video sequence; for the spatial dimension, texture features are extracted using algorithms such as gray-level co-occurrence matrix and gradient calculation to capture the richness of detail in a single frame; at the same time, visual parameter change features within and between frames are extracted to capture key information affecting the visual experience, such as scene switching and fluctuations in screen brightness, forming a multi-dimensional feature set that can completely characterize the spatiotemporal characteristics of the video content, providing comprehensive and accurate input data for subsequent complexity assessment.
[0036] Step S20: Generate the content complexity information corresponding to the video file to be processed based on the multi-dimensional features.
[0037] It is understandable that content complexity information can refer to information generated after fusing and analyzing multi-dimensional features, which can quantitatively characterize the content complexity of the video file to be processed at each point in time on the timeline.
[0038] In the specific implementation, multi-dimensional features are fused and analyzed to generate the full-timeline content complexity information for the video file to be processed. This step is not an isolated complexity judgment of a single frame, but a continuous modeling of the entire time series of the video, generating a continuous complexity function with time as the independent variable. This enables a quantitative assessment of the content complexity at each time point within the full duration of the video, accurately distinguishing between high-complexity and low-complexity segments in the video, and identifying key nodes such as scene transitions and sudden changes in image, providing a precise decision-making basis for the adaptive adjustment of encoding parameters.
[0039] Step S30: Adjust the encoding parameters based on the content complexity, and encode the video file to be processed according to the adjusted encoding parameters.
[0040] It should be understood that encoding parameters can refer to control parameters that determine the encoding compression efficiency and output image quality during the video encoding process. Encoding parameters include at least one of the following: Group of Pictures (GOP) structure, Quantization Parameter (QP), bitrate, resolution, and frame rate. In this embodiment, the encoding parameters are explained using GOP structure, QP parameter, and bitrate as examples.
[0041] In its implementation, the encoding parameters are globally adaptively adjusted based on content complexity information. Then, based on the adjusted encoding parameters that highly match the video content, the entire encoding process is performed on the video file to be processed, outputting a compressed video stream. This step is not an independent adjustment of a single parameter, but a coordinated adjustment based on the complexity of the GOP structure, QP parameters, and bitrate allocation: for high-complexity segments, a combination strategy of increasing bitrate allocation, reducing QP quantization intensity, and shortening GOP length is executed simultaneously to ensure the image quality reproduction of complex scenes; for low-complexity segments, a combination strategy of decreasing bitrate allocation, increasing QP quantization intensity, and extending GOP length is executed simultaneously to maximize compression efficiency; at the same time, through continuous complexity modeling of the time axis, a smooth transition of encoding parameters is achieved, avoiding subjective image quality fluctuations caused by parameter abrupt changes, and realizing the rational allocation of encoding resources.
[0042] This embodiment generates content complexity information corresponding to the video file to be processed based on the multi-dimensional features of the video file to be processed, and dynamically adjusts the encoding parameters of the video encoding based on the content complexity, so as to adapt to the differentiated encoding requirements of different video content and improve the rationality of encoding resource allocation.
[0043] Reference Figure 2 , Figure 2 This is a flowchart illustrating the second embodiment of the video encoding method of this application, based on the above. Figure 1 The first embodiment shown illustrates a second embodiment of the video coding method of this application.
[0044] In the second embodiment, step S20 includes: Step S201: Input the multi-dimensional features into the preset content complexity prediction model.
[0045] It should be understood that the pre-defined content complexity prediction model can refer to an artificial intelligence model, including machine learning models or deep neural network models, that has been trained in advance using labeled video multi-dimensional feature datasets and corresponding coding complexity ground truth values to adapt to video coding scenarios. It has solidified the learning of the mapping relationship between "video features and coding complexity", can directly receive multi-dimensional feature inputs and complete end-to-end reasoning, and output content complexity information.
[0046] In the specific implementation, after obtaining multi-dimensional features, the multi-dimensional features are first imported into a pre-trained preset content complexity prediction model. This step completes the input adaptation of the original feature data to the AI inference model, ensuring that feature information of different dimensions such as motion, texture, and brightness enters the model inference stage without loss or bias. This provides a unified and standardized data foundation for subsequent intelligent analysis, eliminating the need to manually adjust feature input rules for different types of videos and achieving unified adaptation for videos across all scenarios.
[0047] Step S202: The multi-dimensional features are fused and analyzed using the preset content complexity prediction model to output the content complexity information corresponding to the video file to be processed.
[0048] It is understandable that fusion analysis can refer to the process of multimodal and nonlinear comprehensive operation and correlation analysis of the multidimensional features of the input by the preset content complexity prediction model. Specifically, it can be the normalization of motion features, texture features, and visual parameter change features, adaptive weight allocation, feature fusion and comprehensive evaluation to eliminate the one-sidedness of single-dimensional features and achieve a comprehensive and accurate determination of the complexity of video content.
[0049] In the specific implementation, after the preset content complexity prediction model receives multi-dimensional feature input, it performs end-to-end fusion analysis: First, it performs feature normalization processing on motion features, texture features, and visual parameter change features to eliminate the dimensional differences of features of different dimensions and ensure the consistency of analysis results; then, it performs non-linear correlation and adaptive weight allocation on the three types of features through the preset content complexity prediction model, and comprehensively evaluates the influence weight of different features on the complexity of video content by combining the temporal correlation of video content. For example, it automatically assigns higher weight to motion features for high dynamic motion videos and higher weight to texture features for high-texture game screens, avoiding the limitations of manually fixed weight rules; then, it generates continuous and quantified content complexity information of the video file to be processed on the time axis through the preset content complexity prediction model, and constructs a content complexity distribution model for the entire duration of the video, realizing the accurate positioning of high-complexity key areas and low-complexity redundant areas of the video.
[0050] Furthermore, to avoid insufficient accuracy in complexity determination due to isolated analysis of single-frame features and neglect of video temporal continuity, and to prevent image quality fluctuations caused by sudden changes in encoding parameters, and to provide accurate and continuous decision-making basis for subsequent adaptive and collaborative adjustment of encoding parameters, step S302 includes: performing temporal dimension feature association and fusion mapping on the multi-dimensional features through the preset content complexity prediction model to generate a content complexity quantification value corresponding to the video timeline; and generating content complexity information corresponding to the video file to be processed based on the content complexity quantification value.
[0051] It should be understood that the temporal dimension can refer to the time sequence dimension of video content, that is, the continuous timeline dimension formed by video frames arranged in playback order, reflecting the dynamic changes of video content over playback time. Feature association can refer to the correlation analysis of multi-dimensional features of consecutive frames in the temporal dimension performed by a pre-set content complexity prediction model, in order to uncover the continuous change trend, abrupt change nodes, motion trajectory, and other temporal correlation characteristics of features on the timeline, eliminating the one-sidedness of isolated single-frame analysis and restoring the temporal continuity of video content. Fusion mapping can refer to the nonlinear fusion operation performed on multi-dimensional features by the pre-set content complexity prediction model after completing temporal feature association. Through a pre-trained and fixed weight system, a mapping relationship between temporal correlation features and video content complexity is established, transforming the high-dimensional temporal feature set into a one-dimensional, quantifiable complexity value. The video timeline can refer to the continuous time dimension coordinates of the video file to be processed from the start point to the end point of playback, corresponding to the video frame sequence. Content complexity quantification values can refer to quantifiable values corresponding to each time point / frame position on the video timeline. The value directly reflects the motion intensity, detail richness, and scene change of the video content at the corresponding time point. The higher the value, the higher the content complexity.
[0052] In its implementation, the pre-defined content complexity prediction model, upon receiving multi-dimensional features arranged chronologically, does not process the feature data of each frame in isolation. Instead, it uses the temporal dimension of the video as a benchmark to perform cross-frame feature association and fusion mapping on the multi-dimensional features of consecutive frames. During the feature association stage, the model mines the continuous motion trajectories of motion features on the timeline, the continuous trends of inter-frame changes, identifies the subtle fluctuations of texture features over time, captures the abrupt scene transition nodes of brightness and contrast changes, and establishes temporal associations between multi-dimensional features of consecutive frames. This fully restores the continuous change characteristics of video content over time, completely avoiding the limitations of single-frame analysis. Based on this, the model performs non-linear fusion mapping on the temporally associated multi-dimensional features. Through a pre-trained and fixed weight system, it adaptively allocates weights and performs comprehensive calculations on features in different scenarios, transforming the high-dimensional temporal feature set into a one-dimensional content complexity quantification value C(t) corresponding to each time point / frame position on the video timeline. Here, C(t) represents the content complexity at time point t, achieving precise anchoring of video temporal content and complexity values.
[0053] After completing the temporal dimension calculations and generating quantization values, the quantization values of continuous content complexity covering the entire video timeline are structurally integrated and processed to generate complete content complexity information for the video file to be processed. This step is not a simple summary of quantization values, but rather, based on the temporally continuous quantization values, it further identifies the start and end times of high-complexity segments, the start and end times of low-complexity segments, abrupt change nodes in scene transitions, and trend curves of continuous complexity changes in the video, forming complete and structured video complexity distribution information for the entire duration, thus realizing the construction of complexity distribution information on the timeline.
[0054] Furthermore, to achieve accurate and complex quantification that aligns with human visual perception and adapts to temporal characteristics, the multi-dimensional features include motion features, texture features, and visual parameter change features. The step of using the preset content complexity prediction model to perform temporal feature association and fusion mapping on the multi-dimensional features to generate a content complexity quantification value corresponding to the video timeline includes: calculating the image motion intensity, image detail richness, and scene change degree corresponding to the video timeline based on the motion features, texture features, and visual parameter change features using the preset content complexity prediction model; and performing temporal association and weighted fusion on the image motion intensity, detail richness, and scene change degree to generate a content complexity quantification value corresponding to the video timeline.
[0055] Understandingly, motion features can refer to temporal characteristics that characterize the pixel variation amplitude, object trajectory, and camera movement state between consecutive video frames, reflecting the degree of dynamic change in video content over time. Texture features can refer to spatial characteristics that characterize the image detail density, texture distribution, and edge richness within a single video frame, while also including the temporal variation characteristics of texture between frames, reflecting the level of detail in the video content in the spatial dimension. Visual parameter variation features can refer to temporal characteristics that characterize the fluctuation amplitude of core visual parameters such as intra-frame brightness distribution and inter-frame brightness and contrast, reflecting sudden changes in brightness and scene transitions in the video. Image motion intensity can refer to quantifiable values corresponding to the video timeline, characterizing the dynamic motion amplitude and the severity of inter-frame changes at a corresponding point in time, and is an indicator that determines the human eye's perception of the smoothness of dynamic images. Image detail richness can refer to quantifiable values corresponding to the video timeline, characterizing the level of texture detail and image refinement at a corresponding point in time, and is an indicator that determines the human eye's perception of image clarity. Scene change level refers to a quantifiable value corresponding to the video timeline, used to characterize the magnitude of scene transitions and the degree of visual parameter mutations at a corresponding time point in the video. It serves as an indicator for identifying video scene boundaries and determining abrupt changes in complexity. Weighted fusion refers to the process by which a pre-trained weight system is used in a pre-defined content complexity prediction model to adaptively allocate and comprehensively fuse quantifiable values of dimensions such as motion intensity, detail richness, and scene change level, generating a comprehensive quantifiable value for content complexity. The weights can be dynamically adjusted according to the scene type of the video content.
[0056] In its implementation, the pre-defined content complexity prediction model receives motion features, texture features, and visual parameter change features arranged sequentially along the video timeline. It first decouples these features from their core attributes: based on motion features, and considering the inter-frame correlation within the temporal context, it calculates the intensity of motion corresponding to the video timeline, precisely quantifying the dynamic range of the image at each time point; based on texture features, and considering intra-frame texture distribution and inter-frame texture changes, it calculates the richness of detail corresponding to the video timeline, precisely quantifying the level of detail at each time point; based on visual parameter change features, and considering the fluctuation amplitude and abrupt changes in inter-frame visual parameters, it calculates the degree of scene change corresponding to the video timeline, precisely quantifying the scene transitions and abrupt changes at each time point. This stage achieves complete decoupling of complexity-influencing factors, with each core attribute precisely supported by a dedicated feature dimension, avoiding interference between different feature dimensions, and ensuring that the complexity calculation logic perfectly aligns with the subjective visual sensitivity of the human eye.
[0057] The pre-defined content complexity prediction model performs coherent processing of three sets of temporal data: motion intensity, detail richness, and scene change degree. First, temporal correlation analysis is performed to uncover the continuous change trends of the three core dimensions along the timeline, the correlation characteristics between consecutive frames, and abrupt change nodes in scene transitions. This fully restores the continuous change essence of video content along the time dimension, avoiding complexity jumps caused by isolated single-frame analysis. Based on this, the pre-trained content complexity prediction model adaptively allocates weights to the three core dimensions according to the scene type of the video content, assigning higher weights to motion intensity for high-dynamic scenes, higher weights to detail richness for high-texture scenes, and lower weights to scene change degree for static scenes. Finally, through weighted fusion calculation, a quantified content complexity value C(t) corresponding to each time point on the video timeline is generated, completing the timeline complexity quantification model.
[0058] Furthermore, to enable the model to autonomously learn the nonlinear mapping relationship between multi-dimensional video features and content complexity, before inputting the multi-dimensional features into the preset content complexity prediction model, the method further includes: constructing a training dataset containing video samples of multiple scene types, and labeling each video sample in the training dataset with ground truth labels for content complexity at corresponding temporal positions; extracting multi-dimensional feature samples from each video sample, and using the multi-dimensional feature samples as input to the initial content complexity prediction model, with the corresponding ground truth labels for content complexity as supervised training targets; and performing supervised training on the initial content complexity prediction model using machine learning algorithms or deep neural network algorithms until the model loss converges, thereby obtaining the preset content complexity prediction model.
[0059] It should be understood that multi-scene video samples refer to video data samples covering different content complexities and application scenarios, used to ensure that the trained model has the ability to generalize across all scenarios. The training dataset refers to a structured dataset used for supervised training of the artificial intelligence model, consisting of multi-scene video samples, corresponding extracted multi-dimensional feature samples, and ground truth labels for content complexity at corresponding temporal positions. A temporal position refers to the time axis coordinates corresponding to each frame in the video sample, corresponding to the time point of video playback, used to ensure accurate temporal anchoring of the ground truth labels for content complexity, multi-dimensional feature samples, and video content. A ground truth label for content complexity refers to a standardized quantitative value labeled for each temporal position in the video sample, which can truly reflect the complexity of the video content at that time point. It is the gold standard and optimization target for supervised training of the model; the label value is positively correlated with the motion intensity, detail richness, and scene change degree of the video content. Multi-dimensional feature samples refer to multi-dimensional feature data extracted from each temporal position of each video sample in the training dataset, matched with the ground truth labels for content complexity at the corresponding temporal positions. The initial content complexity prediction model can refer to an untrained artificial intelligence model architecture with randomly initialized weights and parameters. This can be a machine learning model or a deep neural network model, serving as the foundation for training. After supervised training and parameter optimization, a pre-defined content complexity prediction model is formed. The supervised training objective refers to the optimization target that needs to be fitted during model training. In this embodiment, the supervised training objective is the ground truth label of content complexity corresponding to each temporal position of the video sample. The logic of model training is to minimize the error between the output complexity prediction value and the ground truth label. The machine learning algorithm can refer to supervised machine learning algorithms, including but not limited to random forests, gradient boosting trees, and support vector machines, which can be used to build lightweight complexity prediction models suitable for low-latency real-time encoding scenarios on mobile devices. The deep neural network algorithm can refer to neural network algorithms with deep temporal feature learning capabilities, including but not limited to temporal neural networks, convolutional neural networks, and multimodal fusion networks, which can accurately capture the temporal correlation characteristics of video features and achieve higher accuracy in complexity prediction. Supervised training refers to an AI training paradigm that uses labeled training datasets as a foundation, allowing the model to autonomously learn the mapping relationship between input data and output labels. This involves continuously optimizing model parameters to minimize the error between predicted values and true labels. Model loss convergence refers to the process where, during model training, the loss function value, used to measure prediction error, decreases to a stable range with increasing training iterations, no longer showing a significant decrease. This indicates that the model has fully learned the mapping relationship between input features and output labels, the training has achieved its intended goal, and training can be stopped, with model parameters fixed.
[0060] In the specific implementation, the first step is to collect video samples covering all scenarios and types, including interviews, sports, and game footage. This also includes video material with different resolutions, frame rates, and shooting devices to ensure the diversity and coverage of the dataset, thus preventing overfitting and poor generalization in the model from the outset. Subsequently, for each temporal position of each video sample in the dataset, a ground truth label for content complexity is assigned. The labeling is based on the subjective perception of image quality by the human eye, combined with objective motion intensity, detail richness, and scene variation to achieve a comprehensive quantification. This ensures a precise correspondence between the label and the video's temporal position, providing an accurate and reliable standard for model training.
[0061] For each video sample in the training dataset, a multi-dimensional feature extraction operation, identical to that performed in the subsequent online encoding and inference stages, is executed. Motion features, texture features, and visual parameter variation features are extracted sequentially, generating multi-dimensional feature samples that match the ground truth labels at the corresponding temporal positions. This ensures complete consistency in feature extraction logic between the training and inference phases, preventing data distribution shifts between training and inference, and guaranteeing the stability and accuracy of the model. After feature extraction, the multi-dimensional feature samples are set as the input data for the initial content complexity prediction model, and the ground truth labels for content complexity at the corresponding temporal positions are set as the supervised training objectives of the model. This clarifies the model's learning direction and establishes a mapping relationship between multi-dimensional features and content complexity.
[0062] Based on the target application scenario, select an appropriate machine learning algorithm or deep neural network algorithm, build an initial content complexity prediction model, and randomly initialize the weight parameters. Then, use the prepared training dataset to perform iterative supervised training on the initial content complexity prediction model. During training, the model outputs a complexity prediction value based on the input multi-dimensional feature samples, calculates the loss error between the predicted value and the true label, and then continuously adjusts the model's weight parameters through the corresponding optimization algorithm to minimize the loss value. Continue iterative training until the model's loss function converges and the prediction accuracy on the validation set reaches the preset standard. At this point, stop training, solidify the model's weight parameters, and obtain the preset content complexity prediction model.
[0063] This embodiment uses a pre-trained artificial intelligence model to achieve multi-dimensional feature fusion analysis and quantitative output of content complexity, thereby enabling the fusion analysis of the correlation between encoding parameters and video content features, and accurately identifying the complexity of video content.
[0064] Reference Figure 3 , Figure 3 This is a flowchart illustrating the third embodiment of the video encoding method of this application. Based on the above embodiments, a third embodiment of the video encoding method of this application is proposed.
[0065] In the third embodiment, step S30 includes: Step S301: Based on the continuous distribution characteristics of the content complexity information on the video timeline, divide the video file to be processed into multiple continuous video segments.
[0066] It should be understood that continuous distribution characteristics can refer to the continuous variation of the content complexity quantification value C(t) on the video timeline, including the temporal distribution patterns such as the numerical range of complexity, stable range, fluctuation trend, abrupt change nodes, and the start and end positions of high and low complexity segments. Video segmentation can refer to dividing the video file to be processed into continuous video segments on the timeline based on the continuous distribution characteristics of content complexity, with each segment having the same level of content complexity.
[0067] In its implementation, based on the continuous distribution characteristics of content complexity information along the video timeline, video segmentation of non-fixed duration is performed. Continuous frames with stable complexity within the same numerical range along the timeline are divided into video segments. Scene transitions and abrupt changes in complexity are used as segment boundaries, dividing the complete video file into multiple continuous video units with highly consistent complexity characteristics within each segment. This step fundamentally avoids parameter adaptation conflicts caused by excessively large differences in complexity within the same encoding unit, laying a reasonable control unit foundation for accurate adaptation of subsequent encoding parameters.
[0068] Step S302: Based on the content complexity level corresponding to each video segment, determine the combination of encoding parameters that matches the content complexity level.
[0069] It's understandable that content complexity levels can refer to a standardized grading system pre-defined based on the numerical range of quantifiable content complexity values. This can be divided into three levels: high, medium, and low. Each level corresponds to a specific range of complexity values; higher levels represent greater motion intensity, richness of detail, and degree of scene variation in the video content. Encoding parameter combinations can refer to a pre-defined, collaboratively optimized set of multi-dimensional encoding parameters for different content complexity levels. This includes GOP structure, QP parameters, and bitrate allocation parameters, with different levels corresponding to differentiated parameter collaboration strategies.
[0070] In the specific implementation, the complexity feature value (such as average complexity and peak complexity) of each video segment is first calculated, and the complexity feature value is mapped to the preset content complexity level. Then, based on the pre-established "complexity level-encoding parameter combination" mapping relationship, a set of co-optimized encoding parameter combinations is matched for each level: high complexity level corresponds to the image quality priority combination of "short GOP length + low QP value + high bitrate allocation", low complexity level corresponds to the compression efficiency priority combination of "long GOP length + high QP value + low bitrate allocation", and medium complexity level corresponds to the balanced parameter combination.
[0071] Furthermore, in order to achieve a quantitative integration of technical compression requirements and content value requirements, step S302 includes: obtaining the content importance score corresponding to each video segment of the video file to be processed on the video timeline; generating encoding control weights for each video segment based on the content importance score and content complexity level of each video segment; and dynamically adjusting the encoding parameter combination corresponding to the video segment according to the encoding control weights.
[0072] It should be understood that adjusting encoding parameters based on content complexity, while addressing the "mismatch between compression difficulty and encoding resources" from a purely technical perspective of video compression, has certain limitations. Technically low-complexity segments may contain core, critical content for users (such as key statements in interviews or crucial UI prompts in games). Allocating low bitrates solely based on low complexity levels would result in degraded image quality of this critical content. Conversely, technically high-complexity segments may contain redundant, valueless content (such as meaningless camera shake in motion videos or ineffective movement sequences in games). Allocating high bitrates solely based on high complexity levels would lead to a significant waste of encoding resources. Therefore, this embodiment introduces a content importance scoring system based on multimodal content understanding, achieving coordinated control across both technical and content value dimensions.
[0073] Content importance score refers to a quantitative value corresponding to the video timeline, generated by a pre-trained, preset multimodal scoring model. It reflects the importance, emotional intensity, and information density of the video content at a given time point. It is an indicator that measures content priority from the dimensions of video semantics and user viewing value, generated by the fusion analysis of the video's visual, audio, and textual semantic multimodal features. The preset multimodal scoring model can be an artificial intelligence model pre-trained using massive amounts of labeled video samples. It can simultaneously extract and fuse the video's visual, audio, and textual semantic features to output a content importance score that matches the video timeline.
[0074] In the specific implementation, after segmenting the video based on content complexity, a pre-trained multimodal scoring model is used to perform full-duration analysis on the video file to be processed: extracting visual features (motion intensity, facial emotions, main subject), audio features (volume intensity, voice emotion, background music rhythm), and text semantic features (keywords after speech-to-text conversion, core semantics). Through multimodal fusion analysis, a content importance score corresponding to the video timeline is generated. Subsequently, the continuous content importance scores are matched with the segmented video segments, and the core importance indicators (such as average score, peak score, and core semantic weight) corresponding to each video segment are calculated. This completes the two-dimensional data alignment of "content complexity level - content importance score", providing standardized input data for subsequent weight calculation.
[0075] Based on the content complexity level, a basic weight coefficient is determined for each video segment. This basic weight is positively correlated with the complexity level, ensuring the basic image quality requirements for video compression. Then, the content importance score is used as a correction coefficient to dynamically adjust the basic weight: segments with high importance scores have their weight coefficient increased, while segments with low importance scores have their weight coefficient decreased. Finally, a weighted fusion calculation is used to generate the encoding control weight corresponding to each video segment. This step transforms two different dimensions of evaluation indicators into standardized control coefficients that the encoder can directly recognize and execute, completely establishing the connection between the technical attributes and value attributes of video content and encoding control.
[0076] First, based on the content complexity level of the video segments, the corresponding basic encoding parameter combinations are matched: high complexity levels correspond to the image quality-priority basic combination of "short GOP length + low QP value + high bitrate allocation", while low complexity levels correspond to the compression efficiency-priority basic combination of "long GOP length + high QP value + low bitrate allocation". Then, based on the encoding control weight of the segment, the basic encoding parameter combinations are dynamically fine-tuned: when the weight is higher than the baseline value, it is adjusted in the direction of image quality priority (further increase the bitrate, decrease the QP value, and shorten the GOP length); when the weight is lower than the baseline value, it is adjusted in the direction of compression efficiency priority (further decrease the bitrate, increase the QP value, and extend the GOP length). At the same time, the coordinated optimization of the three core parameters of GOP, QP, and bitrate is always ensured to avoid image quality fluctuations and compression efficiency losses caused by adjusting a single parameter.
[0077] Step S303: Adjust the encoding parameters based on the combination of encoding parameters, and encode each video segment according to the adjusted encoding parameters.
[0078] It is understandable that encoding can refer to the process of compressing video segments according to standard video encoding protocols such as H.264 and H.265, using matching combinations of encoding parameters, to remove spatiotemporal redundant information in the video and generate standardized compressed video files.
[0079] In the specific implementation, based on the encoding parameter combination corresponding to each video segment, the encoder configuration parameters are adjusted segment by segment to perform coherent encoding processing on each video segment. During the encoding process, a unified encoding parameter combination is used within the same segment to ensure the consistency of image quality within the segment; a smooth transition strategy is adopted for parameters between adjacent segments, combined with a time axis collaborative control mechanism to avoid image quality jumps caused by parameter abrupt changes; at the same time, the segmented encoding mode supports parallel processing, which greatly improves encoding efficiency and adapts to the low latency requirements of real-time encoding on mobile devices.
[0080] Furthermore, in order to eliminate the image quality fluctuation problem caused by segmented encoding while ensuring that the encoding parameters are accurately matched with the complexity of the video content, before step S303, the method further includes: identifying parameter difference information between the encoding parameter combinations corresponding to adjacent video segments; determining the frame connection boundary of adjacent video segments, and defining a transition frame interval of a preset length based on the frame connection boundary; generating an intermediate encoding parameter sequence that continuously and gradually changes within the transition frame interval according to the parameter difference information and the total number of frames in the transition frame interval; and encoding the video frames within the transition frame interval using the intermediate encoding parameter sequence.
[0081] It should be understood that parameter difference information can refer to the quantified information of the magnitude and direction of change of the differences in the core coding parameters in the coding parameter combinations corresponding to two adjacent video segments, including QP value differences, target bitrate differences, GOP length differences, etc. Frame transition boundary can refer to the frame sequence transition position of two adjacent video segments on the video timeline, that is, the boundary point between the last frame of the previous video segment and the first frame of the next video segment. Preset length can refer to the time length of the transition frame interval set in advance according to the video frame rate, the real-time requirements of the application scenario, and the characteristics of human vision. It can be measured in frames and is used to control the smoothness of parameter transition, balance the transition effect and computational overhead, and avoid transitions that are too short causing jumps or too long affecting the accuracy of parameter and content adaptation. Transition frame interval can refer to the time range containing multiple consecutive video frames extended from the frame transition boundary to the two adjacent video segments. It is the frame sequence interval for performing smooth transition of coding parameters and can adopt a symmetrical distribution design, with half of the frames at the end of the previous segment and half at the beginning of the next segment to ensure the continuity of parameter transition. The total number of frames can refer to the total number of video frames contained within the transition frame interval. The intermediate coding parameter sequence can refer to the set of coding parameters that is calculated based on parameter difference information and the total number of frames in the transition frame interval, and that changes continuously and gradually with the frame order within the transition frame interval. The coding parameters corresponding to each frame in the sequence smoothly transition from the coding parameter combination of the previous segment to the coding parameter combination of the next segment without abrupt jumps.
[0082] In the specific implementation, after matching the encoding parameter combinations for all video segments, the system iterates through all adjacent video segments on the video timeline group by group, identifies the encoding parameter combinations corresponding to each group of segments, extracts the magnitude and direction of the difference in QP value, target bitrate, and GOP length parameters, and generates quantified parameter difference information. At the same time, the parameter differences are compared with a preset threshold. If the difference does not exceed the threshold, it means that the parameter switching will not cause a perceptible change in image quality, and there is no need to start the transition process; the original parameter combination is directly used for encoding. If the difference exceeds the threshold, the subsequent complete smooth transition process is started to accurately locate all parameter switching scenarios that may cause image quality problems.
[0083] After locating the frame transition boundary between two adjacent video segments, and using this boundary as a reference, combined with the video frame rate and the real-time requirements of the application scenario, a preset length of frame range is defined for both the preceding and following segments, forming a transition frame interval. To balance smoothness and content adaptation accuracy, the transition frame interval can adopt a symmetrical distribution design, with half of the frames located at the end of the preceding segment and the other half at the beginning of the following segment. This ensures that the parameter transition completely covers the switching boundary while preventing the transition interval from exceeding the core content range of the segment, thus not affecting the accurate matching of the main content and encoding parameters of each segment.
[0084] Based on the quantified parameter difference information and the total number of frames in the transition frame interval, the single-frame gradient step size for each core coding parameter is calculated. Then, according to linear smoothing or non-linear gradient rules optimized for human vision, corresponding intermediate coding parameters are generated for each frame within the transition frame interval, forming a continuously gradient intermediate coding parameter sequence. The parameters of the starting frame of this sequence are completely consistent with the coding parameter combination of the previous video segment, and the parameters of the ending frame are completely consistent with the coding parameter combination of the next video segment. There are no abrupt jumps in parameters for each frame in between, eliminating the discontinuity points of parameter switching.
[0085] When performing full video encoding, a partition-adaptive encoding strategy is adopted: for the main non-transition areas of each video segment, the original encoding parameter combination that matches the content complexity is used for encoding to ensure the encoding efficiency and image quality of the core content; for all video frames in the transition frame interval, the generated continuous gradient intermediate encoding parameter sequence is used for frame-by-frame encoding, which not only achieves a smooth transition of encoding parameters between adjacent segments, but also does not affect the parameter adaptation effect of the main content of each segment, fundamentally avoiding image quality fluctuations caused by parameter mutations.
[0086] This embodiment determines segment boundaries based on the continuous distribution characteristics of content complexity information on the video timeline, accurately matching the natural temporal structure of the video content, thereby achieving complete alignment between the encoding control unit and the characteristics of the video content, and thus improving the video encoding effect.
[0087] Reference Figure 4 , Figure 4 This is a flowchart illustrating the fourth embodiment of the video encoding method of this application. Based on the above embodiments, the fourth embodiment of the video encoding method of this application is proposed.
[0088] In the fourth embodiment, the multi-dimensional features include motion features, texture features, and visual parameter change features. Step S10 includes: Step S101: In response to the video encoding request.
[0089] It should be understood that this embodiment constructs a multi-channel parallel feature extraction architecture. After receiving the video file to be processed, the time frame sequence of the original video is synchronously sent to multiple independent feature extraction channels to ensure that all extracted features are precisely aligned with the video timeline and integrated to form a standardized multi-dimensional feature set.
[0090] Step S102: Perform inter-frame difference analysis on consecutive video frames of the video file to be processed, and extract motion features that reflect the degree of change in content between frames.
[0091] It is understandable that consecutive video frames refer to a sequence of video frames in a video file to be processed, arranged in chronological order. Inter-frame difference analysis refers to a standardized image processing process that compares and analyzes the pixel data of adjacent frames in a consecutive video frame sequence to analyze the magnitude of pixel changes, motion trajectories, and displacement vectors between frames.
[0092] In the specific implementation, standardized image processing methods such as motion estimation algorithms and inter-frame pixel difference calculation are used to match the pixel data of adjacent frames block by block, calculate the core indicators such as motion vectors between adjacent frames, mean square error of pixel difference, and inter-frame content change rate, and finally extract the motion features corresponding to the video time axis.
[0093] Furthermore, to improve the accuracy of motion feature extraction, step S102 includes: selecting adjacent video frame groups of the video file to be processed at a preset frame interval, and calculating the inter-frame motion vector and pixel difference matrix; based on the inter-frame motion vector and pixel difference matrix, calculating the inter-frame motion amplitude, the proportion of motion region, and the scene switching probability; and generating motion features corresponding to the video timeline based on the inter-frame motion amplitude, the proportion of motion region, and the scene switching probability.
[0094] It should be understood that the preset frame interval refers to the interval between the number of frames selected for analysis, pre-set according to the video frame rate and the real-time and accuracy requirements of the application scenario. It can be set to 1 frame (i.e., selecting two temporally consecutive adjacent frames), and can be flexibly adjusted to 2-3 frames in high frame rate scenarios to balance the computational overhead of feature extraction with analysis accuracy, adapting to the performance requirements of real-time encoding on mobile devices. An adjacent video frame group refers to the smallest analysis unit selected from the temporal frame sequence of the video file to be processed, containing at least two temporally consecutive video frames, according to the preset frame interval. The inter-frame motion vector refers to the quantized values of the displacement direction and magnitude of corresponding pixel blocks in adjacent video frame groups, calculated using a block matching motion estimation algorithm, used to accurately characterize the motion trajectory and trend of objects and shots in the image. The pixel difference matrix refers to a two-dimensional numerical matrix generated after calculating the pixel-by-pixel difference of the luminance / chrominance components at corresponding positions in adjacent video frame groups, used to quantify the absolute degree of change of pixels between adjacent frames. Inter-frame motion amplitude can be a comprehensive quantitative value obtained by statistically analyzing the mean magnitude of the inter-frame motion vectors and the root mean square of the pixel difference matrix. It characterizes the intensity of the overall motion between adjacent frames, and the value is positively correlated with the motion intensity. Motion region proportion can be the proportion of pixel regions experiencing significant motion in the entire video frame, obtained by statistically analyzing the pixel difference matrix and applying a preset difference threshold. It characterizes the spatial coverage of inter-frame motion. Scene switching probability can be a probability value obtained by statistically analyzing the overall difference of the pixel difference matrix and the histogram distribution of pixel difference. The value ranges from 0 to 1, and it characterizes the likelihood of a full-screen scene switch occurring between adjacent frames. The closer the value is to 1, the higher the probability of a scene switch.
[0095] In the specific implementation, firstly, based on the frame rate of the video file to be processed and the real-time requirements of the target application scenario, an appropriate preset frame interval is set: for mobile real-time encoding scenarios, a default frame interval of 1 frame is used to ensure analysis accuracy; for high frame rate videos and low computing power devices, the interval can be adjusted to 2-3 frames to reduce computational overhead. Subsequently, following the playback order of the video timeline, the entire temporal frame sequence of the video is traversed, and multiple groups of consecutive adjacent video frames are selected at the preset frame interval to ensure that each frame group is precisely bound to a fixed position on the video timeline, providing a standardized minimum processing unit for subsequent analysis.
[0096] For each group of adjacent video frames, two core calculations are performed in parallel: First, the image is divided into blocks using a block-matching motion estimation algorithm, calculating the displacement direction and magnitude of each pixel block between adjacent frames to generate a complete set of inter-frame motion vectors, accurately capturing the motion trajectory and directionality of the image content; second, pixel-by-pixel differences are calculated for the luminance components of adjacent frames to generate a pixel difference matrix, quantifying the absolute change in the position of each pixel between adjacent frames. These two types of basic data complement each other, solving the problem that a single pixel difference cannot reflect the directionality of motion, and compensating for the deficiency that a single motion vector cannot quantify the absolute change of pixels, providing a comprehensive and accurate data foundation for subsequent multi-dimensional statistics.
[0097] Based on inter-frame motion vectors and pixel difference matrices, three core dimensions of quantitative indicators are simultaneously calculated: 1. Inter-frame motion amplitude: By calculating the weighted average of motion vector magnitudes and the root mean square of the pixel difference matrix, the intensity of overall motion between adjacent frames is comprehensively quantified, directly corresponding to the motion intensity of the video content; 2. Motion region proportion: By using a preset pixel difference threshold, pixel regions with significant motion are selected in the image, and their proportion of the entire image is calculated to accurately distinguish between local small-scale motion and global large-scale motion; 3. Scene switching probability: By analyzing the overall average of pixel differences across the entire image and the histogram distribution of pixel differences, it is determined whether a sudden change in the content of the entire image occurs between adjacent frames, quantifying the possibility of scene switching and accurately locating scene change nodes in the video. These three dimensions comprehensively cover all core scenarios of inter-frame content changes from the perspectives of motion intensity, spatial coverage, and scene change characteristics.
[0098] The statistical results of three dimensions—inter-frame motion amplitude, motion region proportion, and scene switching probability—corresponding to each adjacent video frame group are precisely bound to the position of the frame group on the video timeline. The results are then structurally integrated according to the temporal order of video playback to generate standardized motion features that cover the entire duration of the video file to be processed and correspond to the video timeline. These features are then output to the subsequent multi-feature fusion and content complexity prediction stages.
[0099] Step S103: Perform image detail analysis on single video frames of the video file to be processed, and extract texture features that reflect the richness of the content of the picture.
[0100] It is understandable that a single video frame can refer to an independent, complete single video image frame in the video file to be processed. Image detail analysis can refer to a standardized image processing process that performs spatial domain analysis on the image data of a single video frame to extract edge information, texture distribution, and detail density from the image.
[0101] In the specific implementation, spatial domain image processing algorithms such as edge detection operators, gray-level co-occurrence matrix, and wavelet transform are used to perform region-by-region analysis on each independent video frame, calculate core indicators such as texture distribution density, number of edges, proportion of high-frequency information, and detail richness of a single frame, and finally extract the texture features corresponding to the video timeline.
[0102] Furthermore, in order to improve the accuracy of texture feature extraction, step S103 includes: uniformly dividing the single-frame video of the video file to be processed into multiple image sub-blocks of the same size; performing spatial gradient statistics or frequency domain transformation analysis on each image sub-block to obtain the texture detail richness information of each sub-block; based on the texture detail richness information of all image sub-blocks in the whole frame, statistically analyzing the texture richness and spatial distribution characteristics of the whole frame, and generating texture features corresponding to the video time axis.
[0103] It should be understood that uniform block partitioning refers to a standardized image processing operation that divides a single video frame into multiple row- and column-aligned, uniformly sized, non-overlapping image sub-blocks according to a preset pixel size. An image sub-block can be the smallest texture analysis unit of uniform size obtained through uniform block partitioning, ensuring both the refinement of texture analysis and the spatial localization capability of texture information. Spatial gradient statistics refer to a standardized image processing method that calculates the magnitude and direction distribution of gray-level gradients of pixels within an image sub-block using gradient operators such as Sobel and Canny in the image spatial domain, quantifying image edge and detail information. Frequency domain transform analysis refers to a standardized image processing method that converts the spatial domain pixel data of an image sub-block into frequency domain coefficients using algorithms such as Discrete Cosine Transform (DCT) and wavelet transform, quantifying texture details through the distribution and energy proportion of high-frequency coefficients. Texture detail richness information refers to the quantized value obtained through spatial gradient statistics or frequency domain transformation analysis for each image sub-block. It is used to characterize the richness of image edges, textures, and details within that sub-block; a higher value indicates richer image details within the sub-block. Overall frame texture richness refers to the comprehensive quantized value calculated using weighted averaging, peak statistics, and other methods based on the texture detail richness information of all image sub-blocks within a single frame. It is used to characterize the overall image detail richness of a single video frame. Spatial distribution characteristics refer to the distribution characteristics of texture details in the image space, statistically obtained based on the texture detail richness information of all image sub-blocks within a single frame. This includes the proportion, location distribution, and concentration of high-texture areas, used to accurately locate high-detail key areas and low-detail redundant areas within the image.
[0104] In the specific implementation, for each single video frame of the video file to be processed, a pre-set, suitable pixel block size (such as 16×16 or 32×32 pixels) is determined based on the video resolution, the accuracy and performance requirements of the application scenario. The single frame is then divided into non-overlapping, uniform blocks, dividing the complete two-dimensional image into multiple row- and column-aligned image sub-blocks of identical size. This step decomposes the entire large frame into standardized, smallest analysis units, ensuring both fine-grained texture analysis and spatial localization of texture information. Furthermore, the standardized sub-block size supports parallel computation, significantly improving the efficiency of feature extraction and adapting to the low-computing-power, low-latency requirements of real-time encoding on mobile devices.
[0105] For each divided image sub-block, spatial gradient statistics or frequency domain transform analysis are flexibly selected based on the application scenario to perform quantitative analysis of local texture details: For low-latency scenarios such as real-time encoding on mobile devices, spatial gradient statistics, which has higher computational efficiency, are used. Gradient operators are used to calculate the gray-level gradient distribution of pixels within the sub-block, and key indicators such as the average gradient magnitude and the proportion of non-zero gradient pixels are statistically analyzed to generate texture detail richness information for the sub-block; For high-precision scenarios such as high-definition video-on-demand, frequency domain transform analysis, which has higher analytical accuracy, is used. The spatial domain pixels of the sub-block are converted into frequency domain coefficients through DCT transformation, and the energy proportion of high-frequency coefficients is statistically analyzed. The higher the high-frequency energy proportion, the richer the texture details within the sub-block, thus generating texture detail richness information for the sub-block. This step achieves accurate texture quantization of each local area within the image, capturing subtle edge, texture, and detail information.
[0106] Based on the texture detail richness information of all image sub-blocks within a single frame, two dimensions of statistical analysis are performed simultaneously: First, the overall frame texture richness is calculated, generating a comprehensive quantitative value of the overall detail richness of the single frame through weighted averaging, peak statistics, and other methods, directly corresponding to the overall complexity of the image; second, spatial distribution feature statistics are performed, filtering out high-texture sub-blocks through a preset richness threshold, and statistically analyzing their proportion, positional distribution, and concentration within the image to accurately locate high-detail key areas and low-detail redundant areas within the image. Finally, the overall frame texture richness and spatial distribution features corresponding to each frame are precisely bound to the frame's position on the video timeline, and structurally integrated according to the video playback sequence. This ultimately generates standardized texture features covering the entire duration of the video to be processed and corresponding to the video timeline, which are then output to the subsequent multi-feature fusion and content complexity prediction stages.
[0107] Step S104: Perform pixel statistical analysis on the video frames of the video file to be processed, and extract the visual parameter change features that reflect the fluctuation of visual attributes of the image.
[0108] Understandably, pixel statistical analysis refers to a standardized image processing process that performs statistical operations on video pixel data in single frames and between consecutive frames to calculate core visual parameters such as brightness distribution, contrast, and color gamut, as well as the fluctuation range of these parameters between frames.
[0109] In the specific implementation, full pixel statistics are performed on a single video frame to calculate core visual parameters such as average brightness, brightness distribution histogram, contrast, and color gamut range of the single frame; then, difference calculation and trend analysis are performed on the visual parameters between consecutive frames to calculate the fluctuation amplitude and abrupt change of the visual parameters between frames, and finally, the visual parameter change features corresponding to the video timeline are extracted.
[0110] Furthermore, to improve the accuracy of visual parameter change feature extraction, step S104 includes: performing full-frame pixel value statistics on a single frame of the video file to be processed to obtain the core visual parameters of the single frame, wherein the core visual parameters include at least one of average brightness, contrast, and dynamic range; comparing the core visual parameters of consecutive video frames along the video timeline to count the parameter fluctuation amplitude and abrupt change frequency; and generating visual parameter change features corresponding to the video timeline based on the parameter fluctuation amplitude and the abrupt change frequency.
[0111] It should be understood that full-frame pixel value statistics can refer to a standardized image processing operation that performs global statistical calculations on the luminance and chrominance components of all pixels within a single video frame to extract the overall visual attributes of the image. Core visual parameters can refer to quantitative indicators that directly reflect the overall visual attributes of a single frame and are highly correlated with human subjective perception. These include at least one of the following: mean luminance, contrast ratio, and dynamic range. They are data that characterize the visual characteristics of the image and help identify scene changes. The mean luminance can be the arithmetic mean of the luminance components of all pixels within a single video frame, used to quantify the overall brightness of the image. The contrast ratio can be the ratio of the maximum to the minimum luminance component of a pixel within a single video frame, used to quantify the contrast between light and dark levels and details in the image. The dynamic range can be the maximum range of grayscale levels that the luminance components of a pixel within a single video frame can cover, used to quantify the image's ability to render details from the darkest to the brightest. The parameter fluctuation amplitude can be the absolute value of the difference between core visual parameters between adjacent consecutive video frames along the video timeline, used to quantify the degree of continuous change in the image's visual attributes; a larger value indicates a more drastic change in visual attributes. The mutation frequency refers to the number of times the difference of the core visual parameters exceeds the preset mutation threshold within a preset sliding time window. It is used to quantify the frequency of abrupt changes in the visual attributes of the image and is an indicator for identifying video scene switching and scene transitions.
[0112] In its implementation, for each single frame of the video file to be processed, the image is first converted from the RGB color space to the YUV color space. The luminance component (Y component), which is most sensitive to the human eye, is extracted. Global statistical operations are then performed on the luminance components of all pixels in the entire frame to calculate core visual parameters such as the average luminance, contrast, and dynamic range of the single frame, thus achieving precise quantification of the visual attributes of the single frame. This step, starting from the spatial domain, obtains basic visual attributes of the image that highly match the subjective perception of the human eye. This provides standardized and unified basic data for subsequent temporal comparative analysis, avoiding comparison errors caused by inconsistent statistical standards between different frames from the outset.
[0113] Using the video timeline as a baseline, all consecutive video frames are traversed in playback order. The core visual parameters of adjacent frames are compared frame-by-frame, and the absolute value of the parameter difference between each group of adjacent frames is calculated to obtain the parameter fluctuation amplitude at each time point. Simultaneously, based on a preset abrupt change threshold according to the video scene and resolution, the number of times the parameter difference exceeds the threshold within a sliding time window is counted to obtain the abrupt change frequency for the corresponding time window. This step, starting from the time dimension, comprehensively captures the continuous change patterns and abrupt change nodes of the visual attributes of the image. It can accurately reflect the continuous parameter fluctuations caused by gradual changes in light and shadow, and accurately identify parameter abrupt changes caused by scene switching and image transitions.
[0114] The parameter fluctuation amplitude at each time point and the mutation frequency at each time window are precisely bound to the corresponding frame position on the video timeline. The data are then structurally integrated according to the video playback sequence to generate standardized visual parameter change features that cover the entire duration of the video to be processed and correspond to the video timeline. These features are then output to the subsequent multi-feature fusion and content complexity prediction stages.
[0115] This embodiment constructs a multi-channel parallel feature extraction architecture. After receiving the video file to be processed, the time-series frame sequence of the original video is synchronously sent to multiple independent feature extraction channels to ensure that all extracted features are precisely aligned with the video timeline and integrated to form a standardized multi-dimensional feature set.
[0116] For ease of understanding, please refer to Figure 5 This is for illustrative purposes only, and does not limit the scope of this application. As an example, Figure 5 This is a system structure diagram of an embodiment of the video encoding method of this application. This application introduces a content complexity prediction model into the video encoding process, analyzes the video through the content complexity prediction model, generates content complexity information on the timeline, and dynamically adjusts the encoding parameters based on the content complexity information to achieve content-driven encoding control. Figure 5 The modules can be integrated into a video encoder or video processing SDK. The modules include: 1. Video Input Module: Used to receive video files to be processed, either uploaded by the user through the user terminal or automatically selected by the system terminal.
[0117] 2. Feature Extraction Module: Used to extract multi-dimensional features from the video file to be processed. Multi-dimensional features include, but are not limited to: (1) motion features (reflecting changes between frames); (2) texture features (reflecting image details); and (3) brightness and contrast changes. Among them, multi-dimensional features are used to describe the changes of video content on the time axis.
[0118] 3. Content Complexity Prediction Model: Used to fuse and analyze multi-dimensional features to generate content complexity assessment results on the video timeline. Where C(t) represents the content complexity at time t, which reflects the intensity of motion, the richness of detail, and the degree of scene change.
[0119] In one implementation, the content complexity prediction model can be pre-trained using a machine learning model or a neural network; this embodiment does not impose any limitations on this.
[0120] 4. Image Processing Module: Used to dynamically adjust encoding parameters based on content complexity; the image processing module includes a GOP control module, a QP control module, and a bitrate allocation module. The GOP control module is used to adjust the GOP structure, the QP control module is used to adjust the quantization parameters (QP), and the bitrate allocation module is used to allocate the bitrate. The specific control strategies include: (1) for high-complexity segments: increasing bitrate allocation, decreasing QP value, and shortening GOP length to prioritize image quality; (2) for low-complexity segments: decreasing bitrate allocation, increasing QP value, and extending GOP length to prioritize compression efficiency. The control strategies can be generated through preset rules or based on preset large models, and this embodiment does not impose any restrictions on this.
[0121] 5. Encoder: Used to encode the video file to be processed according to the adjusted encoding parameters to obtain the processed video file.
[0122] 6. Video output module: Used to output the processed video file.
[0123] It should be noted that the above examples are only for understanding this application and do not constitute a limitation on the video coding method of this application. Any simple modifications based on this technical concept are within the protection scope of this application.
[0124] This application also provides a video encoding apparatus, please refer to... Figure 6 The video encoding device includes: Feature extraction module 10 is used to extract multi-dimensional features from the video file to be processed in response to the video encoding request, and obtain multi-dimensional features; Information generation module 20 is used to generate content complexity information corresponding to the video file to be processed based on the multi-dimensional features; The video encoding module 30 is used to adjust the encoding parameters based on the content complexity and to encode the video file to be processed according to the adjusted encoding parameters.
[0125] The video encoding apparatus provided in this application, employing the video encoding method described in the above embodiments, can solve the technical problem that related video encoding methods, which use fixed or empirically based encoding parameters, cannot adapt to the differentiated encoding needs of different video content, resulting in unreasonable allocation of encoding resources. Compared with the prior art, the beneficial effects of the video encoding apparatus provided in this application are the same as those of the video encoding method provided in the above embodiments, and other technical features in the video encoding apparatus are the same as those disclosed in the methods of the above embodiments, and will not be repeated here.
[0126] This application provides a video encoding device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, which are executed by the at least one processor to enable the at least one processor to perform the video encoding method in Embodiment 1 above.
[0127] The following is for reference. Figure 7 This document illustrates a structural diagram of a video encoding device suitable for implementing embodiments of this application. The video encoding device in these embodiments may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Description), PMPs (Portable Media Players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 7 The video encoding device shown is merely an example and should not impose any limitations on the functionality and scope of use of the embodiments of this application.
[0128] like Figure 7As shown, the video encoding device may include a processing unit 1001 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to a program stored in ROM (Read Only Memory) 1002 or a program loaded from storage device 1003 into RAM (Random Access Memory) 1004. RAM 1004 also stores various programs and data required for the operation of the video encoding device. The processing unit 1001, ROM 1002, and RAM 1004 are interconnected via bus 1005. Input / output (I / O) interface 1006 is also connected to the bus. Typically, the following systems can be connected to I / O interface 1006: input devices 1007 including, for example, touch screens, touchpads, keyboards, mice, image sensors, microphones, accelerometers, gyroscopes, etc.; output devices 1008 including, for example, LCDs (Liquid Crystal Displays), speakers, vibrators, etc.; storage devices 1003 including, for example, magnetic tapes, hard disks, etc.; and communication devices 1009. Communication device 1009 allows the video encoding device to communicate wirelessly or wiredly with other devices to exchange data. Although video encoding devices with various systems are shown in the figures, it should be understood that implementation or possession of all the systems shown is not required. More or fewer systems may be implemented alternatively.
[0129] Specifically, according to the embodiments disclosed in this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments disclosed in this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device, or installed from storage device 1003, or installed from ROM 1002. When the computer program is executed by processing device 1001, it performs the functions defined in the methods of the embodiments disclosed in this application.
[0130] The video encoding device provided in this application, employing the video encoding method described in the above embodiments, can solve the technical problem that related video encoding methods use fixed or empirically based encoding parameters, which cannot adapt to the differentiated encoding needs of different video content, resulting in unreasonable allocation of encoding resources. Compared with the prior art, the beneficial effects of the video encoding device provided in this application are the same as those of the video encoding method provided in the above embodiments, and other technical features of this video encoding device are the same as those disclosed in the previous embodiment method, and will not be repeated here.
[0131] It should be understood that the various parts disclosed in this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any suitable manner in one or more embodiments or examples.
[0132] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
[0133] This application provides a computer-readable storage medium having computer-readable program instructions (i.e., a computer program) stored thereon, the computer-readable program instructions being used to execute the video encoding method described in the above embodiments.
[0134] The computer-readable storage medium provided in this application may be, for example, a USB flash drive, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections having one or more wires, portable computer disks, hard disks, RAM (Random Access Memory), ROM (Read Only Memory), EPROM (Erasable Programmable Read Only Memory), or flash memory, optical fiber, CD-ROM (CD-Read Only Memory), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, system, or device. The program code contained on the computer-readable storage medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination thereof.
[0135] The aforementioned computer-readable storage medium may be included in the video encoding device; or it may exist independently and not be assembled into the video encoding device.
[0136] The aforementioned computer-readable storage medium carries one or more programs, which, when executed by a video encoding device, cause the video encoding device to perform the aforementioned video encoding method.
[0137] Computer program code for performing the operations of this application can be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, and C++, as well as conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including LAN (Local Area Network) or WAN (Wide Area Network)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0138] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0139] The modules described in the embodiments of this application can be implemented in software or hardware. The names of the modules do not necessarily limit the functionality of the unit itself.
[0140] The readable storage medium provided in this application is a computer-readable storage medium that stores computer-readable program instructions (i.e., a computer program) for executing the above-described video encoding method. This addresses the technical problem that related video encoding methods, which use fixed or empirically based encoding parameters, cannot adapt to the differentiated encoding needs of different video content, resulting in unreasonable allocation of encoding resources. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided in this application are the same as those of the video encoding method provided in the above embodiments, and will not be repeated here.
[0141] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the video encoding method described above.
[0142] The computer program product provided in this application can solve the technical problem that related video encoding methods use fixed or empirically based encoding parameters, which cannot adapt to the differentiated encoding needs of different video content, resulting in unreasonable allocation of encoding resources. Compared with the prior art, the beneficial effects of the computer program product provided in this application are the same as those of the video encoding methods provided in the above embodiments, and will not be repeated here.
[0143] The above description is only a part of the embodiments of this application and does not limit the scope of this application. All equivalent structural transformations made under the technical concept of this application and using the content of this application specification and drawings, or direct / indirect applications in other related technical fields, are included within the protection scope of this application.
[0144] It should be noted that the data collection, tag management, rule setting, and push decision-making processes involved in this application are designed to work with other technical features to solve technical problems. They do not involve or support any illegal activities. Any data processing that may violate laws and regulations (such as unauthorized collection of privacy data, generation of discriminatory tags, setting unfair rules, or pushing illegal information) is not within the scope of protection of this application's technical solution. Of course, the user data in this application will be encrypted, anonymized, or de-identified before storage to ensure user data security.
[0145] This application discloses A1, a video coding method, the video coding method comprising: In response to a video encoding request, multi-dimensional feature extraction is performed on the video file to be processed to obtain multi-dimensional features; Based on the multi-dimensional features, generate the content complexity information corresponding to the video file to be processed; The encoding parameters are adjusted based on the content complexity, and the video file to be processed is encoded according to the adjusted encoding parameters.
[0146] A2. The video encoding method as described in A1, wherein generating the content complexity information corresponding to the video file to be processed based on the multi-dimensional features includes: The multi-dimensional features are input into a preset content complexity prediction model; The multi-dimensional features are fused and analyzed using the preset content complexity prediction model to output the content complexity information corresponding to the video file to be processed.
[0147] A3. The video encoding method as described in A2, wherein the step of fusing and analyzing the multi-dimensional features through the preset content complexity prediction model to output the content complexity information corresponding to the video file to be processed includes: The preset content complexity prediction model is used to perform temporal dimension feature association and fusion mapping on the multi-dimensional features to generate a content complexity quantification value corresponding to the video timeline. Based on the content complexity quantification value, the content complexity information corresponding to the video file to be processed is generated.
[0148] A4. The video encoding method as described in A3, wherein the multi-dimensional features include motion features, texture features, and visual parameter change features, and the step of performing temporal-dimensional feature association and fusion mapping on the multi-dimensional features through the preset content complexity prediction model to generate a content complexity quantification value corresponding to the video timeline includes: Based on the motion features, texture features, and visual parameter change features, the preset content complexity prediction model calculates the motion intensity, detail richness, and scene change degree of the video timeline. The motion intensity, detail richness, and scene change degree of the image are temporally correlated and weighted to generate a content complexity quantification value corresponding to the video timeline.
[0149] A5. The video encoding method as described in A2, before inputting the multi-dimensional features into the preset content complexity prediction model, further includes: Construct a training dataset containing video samples of multiple scene types, and label the content complexity ground truth labels of each video sample in the training dataset at the corresponding temporal position; Multi-dimensional feature samples are extracted from each video sample, and these multi-dimensional feature samples are used as input to the initial content complexity prediction model, while the corresponding ground truth labels of content complexity are used as supervised training targets. The initial content complexity prediction model is trained in a supervised manner using machine learning algorithms or deep neural network algorithms until the model loss converges, thereby obtaining the preset content complexity prediction model.
[0150] A6. The video encoding method as described in any one of A1 to A5, wherein adjusting the encoding parameters based on the content complexity and encoding the video file to be processed according to the adjusted encoding parameters includes: Based on the continuous distribution characteristics of the content complexity information on the video timeline, the video file to be processed is divided into multiple continuous video segments; Based on the content complexity level corresponding to each video segment, determine the combination of encoding parameters that matches the content complexity level; The encoding parameters are adjusted based on the combination of the encoding parameters, and each video segment is encoded according to the adjusted encoding parameters.
[0151] A7. The video encoding method as described in A6, wherein determining the combination of encoding parameters matching the content complexity level based on the content complexity level corresponding to each video segment includes: Obtain the content importance score corresponding to each video segment on the video timeline of the video file to be processed; Based on the content importance score and content complexity level of each video segment, the encoding control weights of the video segments are generated. Based on the encoding control weights, the combination of encoding parameters corresponding to the video segments is dynamically adjusted.
[0152] A8. The video encoding method as described in A6, before adjusting the encoding parameters based on the encoding parameter combination and encoding each video segment according to the adjusted encoding parameters, further includes: Identify the parameter differences between the encoding parameter combinations corresponding to adjacent video segments; Determine the frame connection boundary between adjacent video segments, and use the frame connection boundary as a reference to define a transition frame interval of a preset length; Based on the parameter difference information and the total number of frames in the transition frame interval, an intermediate coding parameter sequence that gradually changes within the transition frame interval is generated. The intermediate encoding parameter sequence is used to encode video frames within the transition frame interval.
[0153] A9. The video encoding method as described in any one of A1 to A5, wherein the multi-dimensional features include motion features, texture features, and visual parameter variation features, and the step of extracting multi-dimensional features from the video file to be processed to obtain multi-dimensional features includes: Perform inter-frame difference analysis on consecutive video frames of the video file to be processed, and extract motion features that reflect the degree of change in content between frames; Image detail analysis is performed on single video frames of the video file to be processed to extract texture features that reflect the richness of the content of the picture. Pixel statistical analysis is performed on the video frames of the video file to be processed to extract visual parameter change features that reflect the fluctuation of visual attributes of the image.
[0154] A10. The video encoding method as described in A9, wherein performing inter-frame difference analysis on consecutive video frames of the video file to be processed and extracting motion features reflecting the degree of change in inter-frame content includes: Select adjacent video frame groups of the video file to be processed at a preset frame interval, and calculate the inter-frame motion vector and pixel difference matrix; Based on the inter-frame motion vector and pixel difference matrix, the inter-frame motion amplitude, motion area proportion and scene switching probability are statistically analyzed. Motion features corresponding to the video timeline are generated based on the inter-frame motion amplitude, the proportion of the motion region, and the scene switching probability.
[0155] A11. The video encoding method as described in A9, wherein the step of performing image detail analysis on single-frame video frames of the video file to be processed and extracting texture features reflecting the richness of the image content includes: The single-frame video file to be processed is uniformly divided into blocks to obtain multiple image sub-blocks of the same size. Spatial gradient statistics or frequency domain transform analysis are performed on each image sub-block to obtain the texture detail richness information of each sub-block; Based on the texture detail richness information of all image sub-blocks in the whole frame, the texture richness and spatial distribution characteristics of the whole frame are statistically analyzed to generate texture features corresponding to the video timeline.
[0156] A12. The video encoding method as described in A9, wherein performing pixel statistical analysis on the video frames of the video file to be processed to extract visual parameter change features reflecting fluctuations in visual attributes of the image includes: Perform full-frame pixel value statistics on a single frame of the video file to be processed to obtain the core visual parameters of the single frame. The core visual parameters include at least one of the following: average brightness, contrast, and dynamic range. The core visual parameters of consecutive video frames are compared along the video timeline, and the fluctuation amplitude and frequency of parameter changes are statistically analyzed. Based on the fluctuation amplitude of the parameters and the frequency of mutations, visual parameter change features corresponding to the video timeline are generated.
[0157] This application also discloses B13, a video encoding apparatus, the video encoding apparatus comprising: The feature extraction module is used to extract multi-dimensional features from the video file to be processed in response to the video encoding request, and obtain multi-dimensional features. The information generation module is used to generate content complexity information corresponding to the video file to be processed based on the multi-dimensional features; The video encoding module is used to adjust the encoding parameters based on the content complexity and to encode the video file to be processed according to the adjusted encoding parameters.
[0158] B14. In the video encoding apparatus described in B13, the information generation module is further configured to input the multi-dimensional features into a preset content complexity prediction model; perform fusion analysis on the multi-dimensional features through the preset content complexity prediction model, and output the content complexity information corresponding to the video file to be processed.
[0159] B15. In the video encoding apparatus described in B14, the information generation module is further configured to perform temporal dimension feature association and fusion mapping on the multi-dimensional features through the preset content complexity prediction model to generate a content complexity quantification value corresponding to the video timeline; and generate content complexity information corresponding to the video file to be processed based on the content complexity quantification value.
[0160] B16. The video encoding apparatus as described in B15, wherein the multi-dimensional features include motion features, texture features, and visual parameter change features, and the information generation module is further configured to calculate the image motion intensity, image detail richness, and scene change degree corresponding to the video timeline based on the motion features, texture features, and visual parameter change features through the preset content complexity prediction model; and to perform temporal correlation and weighted fusion of the image motion intensity, detail richness, and scene change degree to generate a content complexity quantification value corresponding to the video timeline.
[0161] B17. The video encoding apparatus as described in B14, wherein the video encoding apparatus further comprises: The model training module is used to construct a training dataset containing video samples of multiple scene types, and to label the content complexity ground truth labels of each video sample in the training dataset at the corresponding temporal positions; to extract multi-dimensional feature samples of each video sample, and to use the multi-dimensional feature samples as input to the initial content complexity prediction model, with the corresponding content complexity ground truth labels as supervised training targets; to use machine learning algorithms or deep neural network algorithms to perform supervised training on the initial content complexity prediction model until the model loss converges, thereby obtaining the preset content complexity prediction model.
[0162] This application also discloses C18, a video encoding device, the video encoding device comprising: a memory, a processor, and a video encoding program stored in the memory and executable on the processor, wherein the video encoding program, when executed by the processor, implements the video encoding method as described above.
[0163] This application also discloses D19, a storage medium storing a video encoding program, which, when executed by a processor, implements the video encoding method described above.
[0164] This application also discloses E20, a computer program product including a video encoding program that, when executed by a processor, implements the video encoding method described above.
Claims
1. A video encoding method, characterized in that, The video encoding method includes: In response to a video encoding request, multi-dimensional feature extraction is performed on the video file to be processed to obtain multi-dimensional features; Based on the multi-dimensional features, generate the content complexity information corresponding to the video file to be processed; The encoding parameters are adjusted based on the content complexity, and the video file to be processed is encoded according to the adjusted encoding parameters.
2. The video encoding method as described in claim 1, characterized in that, The step of generating content complexity information corresponding to the video file to be processed based on the multi-dimensional features includes: The multi-dimensional features are input into a preset content complexity prediction model; The multi-dimensional features are fused and analyzed using the preset content complexity prediction model to output the content complexity information corresponding to the video file to be processed.
3. The video encoding method as described in claim 2, characterized in that, The step of fusing and analyzing the multi-dimensional features using the preset content complexity prediction model to output the content complexity information corresponding to the video file to be processed includes: The preset content complexity prediction model is used to perform temporal dimension feature association and fusion mapping on the multi-dimensional features to generate a content complexity quantification value corresponding to the video timeline. Based on the content complexity quantification value, the content complexity information corresponding to the video file to be processed is generated.
4. The video encoding method as described in claim 3, characterized in that, The multi-dimensional features include motion features, texture features, and visual parameter change features. The step of performing temporal-dimensional feature association and fusion mapping on the multi-dimensional features using the preset content complexity prediction model to generate a content complexity quantification value corresponding to the video timeline includes: Based on the motion features, texture features, and visual parameter change features, the preset content complexity prediction model calculates the motion intensity, detail richness, and scene change degree of the video timeline. The motion intensity, detail richness, and scene change degree of the image are temporally correlated and weighted to generate a content complexity quantification value corresponding to the video timeline.
5. The video encoding method as described in claim 2, characterized in that, Before inputting the multi-dimensional features into the preset content complexity prediction model, the method further includes: Construct a training dataset containing video samples of multiple scene types, and label the content complexity ground truth labels of each video sample in the training dataset at the corresponding temporal position; Multi-dimensional feature samples are extracted from each video sample, and these multi-dimensional feature samples are used as input to the initial content complexity prediction model, while the corresponding ground truth labels of content complexity are used as supervised training targets. The initial content complexity prediction model is trained in a supervised manner using machine learning algorithms or deep neural network algorithms until the model loss converges, thereby obtaining the preset content complexity prediction model.
6. The video encoding method according to any one of claims 1 to 5, characterized in that, The step of adjusting the encoding parameters based on the content complexity and encoding the video file to be processed according to the adjusted encoding parameters includes: Based on the continuous distribution characteristics of the content complexity information on the video timeline, the video file to be processed is divided into multiple continuous video segments; Based on the content complexity level corresponding to each video segment, determine the combination of encoding parameters that matches the content complexity level; The encoding parameters are adjusted based on the combination of the encoding parameters, and each video segment is encoded according to the adjusted encoding parameters.
7. A video encoding device, characterized in that, The video encoding device includes: The feature extraction module is used to extract multi-dimensional features from the video file to be processed in response to the video encoding request, and obtain multi-dimensional features. The information generation module is used to generate content complexity information corresponding to the video file to be processed based on the multi-dimensional features; The video encoding module is used to adjust the encoding parameters based on the content complexity and to encode the video file to be processed according to the adjusted encoding parameters.
8. A video encoding device, characterized in that, The video encoding device includes: a memory, a processor, and a video encoding program stored in the memory and executable on the processor, wherein the video encoding program, when executed by the processor, implements the video encoding method as described in any one of claims 1 to 6.
9. A storage medium, characterized in that, The storage medium stores a video encoding program, which, when executed by a processor, implements the video encoding method as described in any one of claims 1 to 6.
10. A computer program product, characterized in that, The computer program product includes a video encoding program, which, when executed by a processor, implements the video encoding method as described in any one of claims 1 to 6.