Video compression processing method and device

By extracting the video feature and setting the target file size, dynamically generating and evaluating the compression parameters, the inefficiency problem caused by the fixation of compression parameters in the prior art is solved, and more efficient and accurate video compression is achieved.

CN120223908APending Publication Date: 2025-06-27CHINA UNICOM ONLINE INFORMATION TECHNOLOGY CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510445037.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-10
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The existing video compression algorithm cannot dynamically adjust the compression parameters according to the image characteristics, resulting in unsatisfactory compression effect, ignores the multi-dimensional characteristics of the image, cannot effectively control the target file size, and is inefficient in compression efficiency.

Method used

Video feature extraction is performed by extracting video features, including spatial complexity, time complexity, color complexity, and target file sizes, and determine the size of each target file according to specific scenarios and user needs. Through weighted calculations, compression parameters are generated dynamically, and compression effects are evaluated dynamically to determine the best compression parameters.

Benefits of technology

It realizes dynamic adjustment of compression parameters according to video characteristics and uses, improves the efficiency and accuracy of video compression, can effectively control the size of the target file, and improves the image quality and reliability of the video surveillance system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120223908A_ABST
    Figure CN120223908A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of image compression processing. The invention provides a video compression processing method and device. The method comprises the following steps: performing video feature extraction on a to-be-processed video, wherein video features comprise space complexity, time complexity and color complexity; configuring the size of a target file corresponding to the to-be-processed video, and determining the size of each target file according to a specific scene and a user demand; dynamically generating compression parameters through weighted calculation based on the extracted video features and the determined size of the target file; and performing dynamic evaluation according to the video compression effects under different compression parameters to determine the range of the optimal compression parameter, selecting the optimal compression parameter from the determined range of the optimal compression parameter, and performing video compression on the to-be-processed video according to the selected optimal compression parameter. According to the invention, the optimal compression parameter corresponding to the to-be-processed video can be determined, and more efficient and accurate video compression can be realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of video processing, and provides a video compression processing method and device. Background Art

[0002] Although traditional video compression algorithms can reduce the storage space and transmission bandwidth requirements to a certain extent, the compression ratio is limited. In commercial and enterprise video surveillance management, a large amount of high-definition video data needs to be stored for a long time and transmitted frequently. The existing technologies cannot well meet this demand, increasing the storage cost and network burden.

[0003] The loss of image quality is relatively large. Traditional algorithms often cause a significant decline in image quality during the compression process, which may affect the effectiveness and reliability of video surveillance. Especially in large-scale video surveillance systems, the loss of image quality may bring serious consequences.

[0004] Therefore, it is necessary to provide a new video compression processing method to solve the above problems. Summary of the Invention

[0005] The present invention provides a video compression processing method and device to solve the technical problems in the prior art that due to the use of fixed compression parameters, it is impossible to dynamically adjust according to the picture features, resulting in unsatisfactory compression effects; due to compression based on file size and resolution, multi-dimensional features such as picture format, color space, and bit depth are ignored, and it is impossible to dynamically adjust the compression parameters according to the target file size set by the user, resulting in the inability to effectively control the target file size, nor can the compression parameters be dynamically adjusted according to the use and content of the picture (such as noise and sharpness), and the compression efficiency is low. The technical problems to be solved by the present invention are achieved through the following technical solutions.

[0006] In the first aspect of the present invention, a video compression processing method is proposed. The video compression processing method includes: Performing video feature extraction on the video to be processed, where the video features include spatial complexity, temporal complexity, and color complexity; setting a target file size corresponding to the video to be processed, and determining each target file size according to specific scenarios and user requirements; based on the extracted video features and the determined target file size, dynamically generating compression parameters through weighted calculation; performing dynamic evaluation according to the video compression effects under different compression parameters to determine the range of the optimal compression parameters, selecting the optimal compression parameter from the determined range of the optimal compression parameters, and performing video compression on the video to be processed according to the selected optimal compression parameter.

[0007] In a second aspect of the present invention, a video compression processing device is proposed, which implements the video compression processing method described in the first aspect of the present invention. The video compression processing device includes: an extraction module for extracting video features from the video to be processed, where the video features include spatial complexity, temporal complexity, and color complexity; a configuration determination module for configuring a target file size corresponding to the video to be processed and determining each target file size according to the specific scenario and user requirements; a calculation module for dynamically generating compression parameters through weighted calculation based on the extracted video features and the determined target file size; a compression module for dynamically evaluating according to the video compression effects under different compression parameters to determine the range of the optimal compression parameters, selecting the optimal compression parameter from the determined range of the optimal compression parameters, and performing video compression on the video to be processed according to the selected optimal compression parameter.

[0008] In a third aspect of the present invention, an electronic device is provided, including: one or more processors; a storage device for storing one or more programs; when the one or more programs are executed by the one or more processors, the one or more processors implement the video compression processing method described in the first aspect of the present invention.

[0009] In a fourth aspect of the present invention, a computer-readable medium is provided, on which a computer program is stored, and when the computer program is executed by a processor, the video compression processing method described in the first aspect of the present invention is implemented.

[0010] The embodiments of the present invention include the following advantages: Compared with the prior art, the present invention extracts video features from the video to be processed, configures a target file size corresponding to the video to be processed, determines each target file size according to the specific scenario and user requirements, dynamically generates compression parameters through weighted calculation based on the extracted video features and the determined target file size, dynamically evaluates according to the video compression effects under different compression parameters to determine the range of the optimal compression parameters, selects the optimal compression parameter from the determined range of the optimal compression parameters, and performs video compression on the video to be processed according to the selected optimal compression parameter, so that the optimal compression parameter corresponding to the video to be processed can be determined, and more efficient and accurate video compression can be achieved. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] Figure 1 is a flowchart of the steps of an example of the video compression processing method of the present invention; Figure 2 is a block diagram of the structure of the video compression processing device of the present invention; Figure 3 is a schematic structural diagram of an embodiment of an electronic device according to the present invention; Figure 4Schematic structural diagram of an embodiment of a computer-readable medium according to the present invention. Detailed implementation manners

[0012] It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments may be combined with each other. The present invention will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.

[0013] In view of the above problems, the present invention proposes a video compression processing method. The method extracts video features from the video to be processed, and the video features include spatial complexity, temporal complexity, and color complexity; sets a target file size corresponding to the video to be processed, and determines each target file size according to specific scenarios and user requirements; based on the extracted video features and the determined target file size, dynamically generates compression parameters through weighted calculation, dynamically evaluates according to the video compression effects under different compression parameters to determine the range of the optimal compression parameters, selects the optimal compression parameter from the determined range of the optimal compression parameters, and performs video compression on the video to be processed according to the selected optimal compression parameter. By analyzing multi-dimensional features such as the file size, resolution, format, color space, bit depth, and noise of the video to be processed, and combining the target file size set by the user, the optimal compression parameters are dynamically calculated, so as to achieve efficient and accurate video compression.

[0014] Embodiment 1 The following refers to Figure 1 to describe the content of the present invention in detail.

[0015] Figure 1 is a step flowchart of an example of the video compression processing method of the present invention.

[0016] As Figure 1 shown, in step S101, video features are extracted from the video to be processed, and the video features include spatial complexity, temporal complexity, and color complexity.

[0017] Specifically, for example, FFmpeg is used to extract a quantization index of spatial complexity (specifically representing texture features or detail features), such as the Spatial Information Index (SI), and the Spatial Information Index represents texture features.

[0018] Specifically, the following method is adopted to extract the SI mean value: SI=$(ffmpeg -i "$INPUT" -vf "signalstats,metadata=print" -f null - 2>&1 | grep "SI_AVG" | awk '{print $2}').

[0019] SI_AVG is the average measure of the spatial complexity of video frames, which reflects the texture details and sharpness of the image by calculating the gradient magnitude of the pixels within the frame (Sobel edge detection). The higher the value, the richer the spatial information such as edges and textures within the frame, and the higher the complexity of the picture. Calculated through the signalstats filter, the SI_AVG field in the output result is this value.

[0020] For example, use ffmpeg to extract the quantization index of the temporal complexity (specifically representing dynamic changes): Temporal Information (TI for short).

[0021] Specifically, adopt the following method to extract the TI mean value: TI=$(ffmpeg -i "$INPUT" -vf "signalstats,metadata=print" -f null - 2>&1 | grep "TI_AVG" | awk '{print $2}').

[0022] TI_AVG is the average measure of the temporal complexity between video frames, which reflects the dynamic degree of the video (such as motion intensity, scene switching, etc.) by calculating the amplitude of pixel changes between consecutive frames. The higher the value, the greater the difference between frames, and the more intense the dynamic changes of the video. Calculated through the signalstats filter, the TI_AVG field in the output result is this value.

[0023] Use ffmpeg to extract the color complexity. Specifically, adopt the following expression to extract the color complexity: Colorfulness =$(ffmpeg -i "$INPUT" -vf "signalstats=metadata=1,histogram=display_mode=color2" -f null - 2>&1 | grep 'lavfi.colorfulness' | awk '{print $2}').

[0024] Colorfulness is an indicator to measure the vividness and diversity of video frame colors. The higher the value, the richer the colors and the more distinct the contrast in the picture. This indicator is based on psychology and color space models, quantifying the perceived intensity of color changes by the human eye. Generate color statistics through the color2 mode of the histogram filter and combine with the signalstats output metadata. Directly extract the value of the lavfi.colorfulness field.

[0025] The video to be processed includes, for example, entertainment videos released on short video platforms, surveillance videos on video surveillance platforms, etc. The video features include features such as spatial complexity, temporal complexity, and color complexity.

[0026] It should be noted that in the present invention, the application scenarios include entertainment videos released on short video platforms, surveillance videos on video surveillance platforms, etc. In other embodiments, tools such as OpenCV, VMA, and ImageMagick can also be used to extract video image features. The above are only illustrative examples and should not be construed as limitations to the present invention.

[0027] Next, in step S102, a target file size corresponding to the video to be processed is configured, and each target file size is determined according to specific scenarios and user requirements.

[0028] Specifically, the setting of video compression and the target file size needs to comprehensively consider factors such as application scenarios (such as short videos, short video playback duration), network loading speed, user experience, and image quality. The target file size is adjusted according to specific scenarios and user requirements.

[0029] In some specific embodiments, for example, the file size of a 5-minute short video is in the range of 50MB to 100MB. For videos with low definition (such as 480p) and fast loading speed on mobile devices, for example, the file size of the short video is in the range of 50MB to 100MB. Preferably, the size of the short video on the short video platform after compression is 100MB.

[0030] It should be noted that the above are only illustrative examples and should not be construed as limitations to the present invention.

[0031] Next, in step S103, based on the extracted video features and the determined target file size, compression parameters are dynamically generated through weighted calculation.

[0032] Specifically, the CRF (Constant Rate Factor) parameter is used to control the balance between compression quality and file size. Based on the CRF parameter, the compression parameters are determined. Further, by dynamically adjusting the compression parameters of each frame of the picture, the visual quality is kept consistent.

[0033] It should be noted that the CRF parameter is used to control the constant quality of the video. The smaller the value of the CRF parameter, the higher the picture quality and the larger the file size; the larger the value of the CRF parameter, the stronger the compression and the lower the picture quality.

[0034] Specifically, the usage example is as follows: ffmpeg -i "$INPUT" -c:v libx264 -crf "$CRF" -preset slow -c:a copy output_crf${CRF}.mp4.

[0035] When the video to be processed is a 5-minute video surveillance-related video, the extracted video features are spatial complexity (texture / detail), temporal complexity (dynamic changes), and color complexity. For the video to be processed, the configured target file size is in the range of 50MB to 100MB.

[0036] Next, based on the extracted video features, the configured target file size, and in combination with the compression quality, compression parameters are dynamically generated through weighted calculation. Specifically, the following expression is used to calculate the compression parameters of the video to be processed: CRF = BaseCRF 初始 −α⋅SI_norm−β⋅TI_norm−γ⋅Color_norm Among them, CRF represents the compression parameter obtained by weighted calculation based on the quantization index of the basic parameters and the extracted video features; BaseCRF 初始 represents the basic parameter, which is characterized by a constant rate factor (corresponding to the basic CRF value in Table 1 below), and can be specifically preset according to the scenario; SI_norm represents the quantization index of the extracted spatial complexity (specifically representing texture features or detail features), that is, the spatial information index; α represents the weight corresponding to the spatial information index; TI_norm represents the quantization index of the extracted temporal complexity (specifically representing dynamic changes), that is, the temporal information index; β represents the weight corresponding to the temporal information index; Color_norm represents the extracted color complexity; γ represents the weight corresponding to the color complexity.

[0037] Table 1

[0038] Table 1 is an example table showing the parameters, calculation methods, and weight ranges related to the calculation of compression parameters.

[0039] It can be seen from Table 1 the calculation methods for different application scenarios and examples of each weight range.

[0040] Next, in step S104, based on the video compression effects under different compression parameters, dynamic evaluation is performed to determine the range of the best compression parameters, the best compression parameter is selected from the determined range of the best compression parameters, and the video to be processed is compressed according to the selected best compression parameter. For the weights of the quantization metrics of the extracted video features, real-time weight adjustment is performed according to the weight adjustment principle.

[0041] Specifically, it includes: first determining the application scenario, and adjusting the weights of the quantization metrics corresponding to one or more video features according to the determined application scenario.

[0042] In the case where the application scenario is landscape video display and the video to be processed is a high-detail video (such as a 4K landscape video), the weight α of the spatial information index is increased, and the weight β of the time information index is decreased. Optionally, when the video to be processed is a high-detail video (such as a 4K landscape video), the weight α of the spatial information index is 0.6, and the weight β of the time information index is 0.3.

[0043] In the case where the application scenario is sports event display and the video to be processed is a high-dynamic video (such as a sports game video), the weight β of the time information index is increased, and the weight γ of the color complexity is decreased to obtain. Optionally, when the video to be processed is a high-dynamic video (such as a 1080p sports game video), the weight β of the time information index is 0.5, and the weight γ of the color complexity is 0.1.

[0044] In the case where the application scenario is animated video display or cartoon video display and the video to be processed is an animated video or a cartoon video, the weight γ of the color complexity is increased. Optionally, when the video to be processed is an animated video (such as 720p animation) or a cartoon video, the weight γ of the color complexity is 0.2.

[0045] When the video to be processed is a real-time monitoring video (such as a factory monitoring video, a home monitoring video), the weight β of the time information index is increased, and the weight γ of the color complexity is decreased. Optionally, when the video to be processed is a real-time monitoring video (such as a factory monitoring video, a home monitoring video), the weight β of the time information index is 0.6, and the weight γ of the color complexity is 0.1.

[0046] Table 2

[0047] Table 2 is an example table showing the respective relevant recommended parameters corresponding to different types of videos to be processed.

[0048] As can be seen from Table 2, the compression parameters calculated and determined by the method of the present invention, that is, the values of CRF corresponding to Table 2, are in the range of 18 to 26. Among them, BaseCRF 初始The value varies according to the video type in different application scenarios, specifically within the range of 22 to 26; the weight α of the spatial information index is within the range of 0.2 to 0.6, the weight β of the time information index is within the range of 0.2 to 0.6, and the weight γ of the color complexity is within the range of 0.0 to 0.5.

[0049] In an alternative embodiment, a surveillance video is selected as the test sample, i.e., the test video, for testing to verify the video compression effect.

[0050] For the same test video, different compression parameters are used for compression to verify the video compression quality. For specific data, refer to Table 3 and Table 4.

[0051] Table 3 Original video Original video size (MB) CRF value File size Compression ratio Test video 1 250 18 230 8% Test video 1 250 19 210 16% Test video 1 250 20 190 24% Test video 1 250 21 178 28.8% Test video 1 250 22 135 46% Test video 1 250 23 115 54% Test video 1 250 24 102 59.2% Test video 1 250 25 95 62% Test video 1 250 26 89 64.4% Test video 1 250 27 68 72.8% Test video 1 250 28 55 78% Table 3 shows a table of parameters related to the video compression quality of compressing the same test video with different compression parameters CRF.

[0052] It can be seen from Table 3 that when testing according to the video quality after compression with different compression parameters CRF, when the compression parameter (corresponding to the CRF value in Table 3) is increased, the file size formed by compression gradually decreases, and the compression parameter and the compression ratio have a non-linear relationship. Therefore, it is impossible to obtain the target compressed file by fixing the CRF parameter value.

[0053] Table 4 Original video Original video size (MB) CRF value File size Compression ratio Test video 1 250 22 99 60.4% Test video 2 250 23 102 59.2% Test video 3 250 23 100 60% Test video 4 250 24 95 62% Test video 5 250 25 98 60.8% Test video 6 250 22 97 61.2% Test video 7 250 24 105 58% Test video 8 250 23 103 58.8% Test video 9 250 22 104 58.4% Test video 10 250 25 96 61.6% Table 4 shows a table of parameters related to the video compression quality of compressing different test videos with different compression parameters CRF (calculated based on video image features).

[0054] It can be seen from Table 4 that for different video files, video feature extraction is performed. The video features include spatial complexity, temporal complexity, and color complexity. After calculating the compression parameters (corresponding to the CRF values in Table 4), different video files are compressed to obtain a relatively ideal compression effect.

[0055] Specifically, according to the video compression effects under different compression parameters (i.e., CRF values), dynamic adjustment evaluation is performed to determine the range of the optimal compression parameters.

[0056] By specifically analyzing the video compression effects under different compression parameters, i.e., CRF values, evaluation is performed to determine the range of the optimal compression parameters (also known as the optimal compression parameter range). The specific implementation steps include: Step S201: Perform multi-dimensional quality evaluation on the video compression effect after compressing the video to be processed to determine whether the video quality index meets the standard.

[0057] The following multiple subjective quality assessments are specifically carried out: Visual quality (detecting whether there are obvious blurs, mosaic effects or detail losses in the detection screen), color performance (evaluating color restoration and checking for color deviation or color distortion), audio quality (ensuring that the audio signal is clearly distinguishable without noise or distortion caused by compression), objective quality assessment (using professional media analysis tools, namely MediaInfo / FFmpeg / Adobe Premiere, etc. for quantitative detection), bitrate analysis (executing the ffmpeg -i output.mp4 command to obtain the actual bitrate and analyzing it against the recommended standard. For example, for a 480p low-definition video, the recommended range is 1.5 - 3Mbps), picture quality index (detecting PSNR, i.e., peak signal-to-noise ratio), and SSIM (structural similarity) index (the reference standard is that for a 480p video, PSNR > 30dB or SSIM > 0.95 is considered qualified).

[0058] Step S202: When it is determined that the video quality index meets the standard, based on the dynamic parameter optimization mechanism, through iterative testing until the requirements of the target file size are met, and at the same time, it is determined that the video quality index meets the standard.

[0059] Specifically, the dynamic parameter optimization mechanism includes: when the output file does not reach the target file size, start the parameter adjustment algorithm. According to different application scenarios, execute different weight adjustment strategies.

[0060] For the parameter adjustment algorithm, a progressive adjustment strategy is specifically adopted. Taking 0.1 as the step unit, at least two of the weights of the time information index, the space information index, and the color complexity weight are adjusted (increased or decreased) in turn, so that the output file meets the target file size to determine the range of each weight.

[0061] For example, when the application scenario is surveillance video, it includes the following scenario features: low SI (weak introverted sensing), much noise but few effective details; low TI (weak introverted thinking), long static time; extremely low color complexity, usually black and white or low saturation. When the application scenario is surveillance video, the following weight adjustment strategy is executed: First, increase the weight β of the time information index (hereinafter, also simply referred to as weight β). When the weight β gradually increases in steps of 0.1 and reaches the first specified value (for example, 0.8), then decrease the weight α of the spatial information index (hereinafter also referred to as weight α), and finally decrease the weight γ of the color complexity (hereinafter, also simply referred to as weight γ). Specifically, determine the specific value by which the weight α of the spatial information index is decreased according to the degree of noise suppression required. In the application scenario of surveillance video, sudden operating conditions need to be quickly responded to. Therefore, the amount (i.e., the second adjustment amount) by which the weight β of the time information index is increased needs to be larger than the other two weights. Suppressing noise is more important than retaining details. Therefore, the amount (i.e., the first adjustment amount) by which the weight α of the spatial information index is decreased is less than the amount (i.e., the second adjustment amount) by which the weight β of the time information index is increased, and the amount (i.e., the first adjustment amount) by which the weight α of the spatial information index is decreased is greater than the amount (i.e., the third adjustment amount) by which the weight γ of the color complexity is decreased. Optionally, the first adjustment amount is 0.1 to 0.4, the second adjustment amount is 0.4 to 0.8, and the third adjustment amount is 0 to 0.1.

[0062] For example, when the application scenario is a sports event (or a sports event display, specifically a sports game video), it includes the following scenario features: medium-high SI (weak introverted sensing), details of the edges of athletes / objects; high TI (weak introverted thinking), large inter-frame changes caused by fast movement; low color complexity, single-color venue (such as a green football field). When the application scenario is a sports event, the following weight adjustment strategy is executed: First, increase the weight β of the time information index. When the weight β is gradually adjusted in steps of 0.1 and reaches the second specified value (for example, 0.7), then decrease the weight γ. When the application scenario is a sports event, details are secondary to the smoothness of movement. Therefore, the weight α is not adjusted (for example, kept at 0.4). Optionally, the amount (i.e., the second adjustment amount) by which the weight β of the time information index needs to be increased is 0.3 to 0.7, and the amount (i.e., the third adjustment amount) by which the weight γ of the color complexity needs to be decreased is 0 to 0.3.

[0063] For example, when the application scenario is an animated video, it includes the following scenario features: low SI (weak introverted sensing), mainly flat color blocks; medium TI (weak introverted thinking), relatively regular character movements; high color complexity, and the edges of bright color blocks need to be sharp. When the application scenario is an animated video, the following weight adjustment strategy is executed: First, reduce the weight α. When the weight α is gradually adjusted downward in steps of 0.1 and reaches the third specified value (e.g., 0.1), then increase the weight γ. The impact on actions can be less concerned about and has little effect, so the weight β is not adjusted (e.g., remains at 0.2). Optionally, the amount by which the weight α for adjusting the spatial information index is reduced (i.e., the first adjustment amount) is 0.1 - 0.5, and the amount by which the weight γ for adjusting the color complexity is reduced (i.e., the third adjustment amount) is 0 - 0.4.

[0064] It should be noted that the above weights and specified values are only optional example illustrations, and in actual applications, they can be adaptively adjusted according to the specific codec characteristics and business requirements, and should not be regarded as a limitation to this solution.

[0065] Furthermore, according to the specific values or value ranges of the determined weights, calculate the specific values or value ranges of the compression parameters to determine the range of the optimal compression parameters.

[0066] Then, select the optimal compression parameter from the determined range of the optimal compression parameters, and perform video compression on the to-be-processed video according to the selected optimal compression parameter. For example, for a landscape video, the optimal compression parameter is 19. For a sports event video, the optimal compression parameter is 20.

[0067] Compared with the prior art, the present invention extracts video features from the to-be-processed video, sets the target file size corresponding to the to-be-processed video, determines the target file sizes according to the specific scenario and user requirements, dynamically generates compression parameters through weighted calculation based on the extracted video features and the determined target file sizes, dynamically evaluates according to the video compression effects under different compression parameters to determine the range of the optimal compression parameters, selects the optimal compression parameter from the determined range of the optimal compression parameters, and performs video compression on the to-be-processed video according to the selected optimal compression parameter, which can determine the optimal compression parameter corresponding to the to-be-processed video and can achieve more efficient and accurate video compression.

[0068] Embodiment 2 The following is an embodiment of the device of the present invention, which can be used to execute the method embodiment of the present invention. For details not disclosed in the device embodiment of the present invention, please refer to the method embodiment of the present invention.

[0069] Figure 2 is a schematic structural diagram of an example of a video compression processing device according to the present invention. The following will refer to Figure 2, a video compression processing device will be described. The video compression processing device is used to execute the video compression processing method described in the first aspect of the present invention.

[0070] As Figure 2 shown, the video compression processing device 300 includes an extraction module 310, a configuration determination module 320, a calculation module 330, and a compression module 340.

[0071] In a specific embodiment, the extraction module 310 is used to extract video features from the video to be processed, and the video features include spatial complexity, temporal complexity, and color complexity. The configuration determination module 320 is used to configure the target file size corresponding to the video to be processed, and determine each target file size according to the specific scenario and user requirements. The calculation module 330 dynamically generates compression parameters through weighted calculation based on the extracted video features and the determined target file size. The compression module 340 dynamically evaluates according to the video compression effects under different compression parameters to determine the range of the optimal compression parameters, selects the optimal compression parameter from the determined range of the optimal compression parameters, and performs video compression on the video to be processed according to the selected optimal compression parameter.

[0072] According to an optional embodiment, dynamically generating compression parameters through weighted calculation based on the extracted video features and the determined target file size includes: Using the following expression to calculate the compression parameters of the video to be processed: CRF = BaseCRF 初始 −α⋅SI_norm−β⋅TI_norm−γ⋅Color_norm; where, CRF represents the compression parameter obtained by weighted calculation according to the quantization index of the basic parameter and the extracted video features; BaseCRF 初始 represents the basic parameter, which is characterized by a constant rate factor and can be preset according to the scenario; SI_norm represents the quantization index of the extracted spatial complexity, that is, the spatial information index; α represents the weight corresponding to the spatial information index; TI_norm represents the quantization index of the extracted temporal complexity, that is, the temporal information index; β represents the weight corresponding to the temporal information index; Color_norm represents the extracted color complexity; γ represents the weight corresponding to the color complexity.

[0073] According to an optional embodiment, it further includes: performing real-time weight adjustment according to the weight adjustment principle, specifically including: first determining the application scenario, and adjusting the weights of the quantization indexes corresponding to one or more video features according to the determined application scenario.

[0074] In the case where the application scenario is landscape video display and the video to be processed is a landscape video, the weight α of the spatial information index is increased, and the weight β of the temporal information index is decreased.

[0075] In the case where the application scenario is sports event display and the video to be processed is a sports competition video, the weight β of the temporal information index is increased, and the weight γ of the color complexity is decreased.

[0076] In the case where the application scenario is animated video display or cartoon video display and the video to be processed is an animated video or a cartoon video, the weight γ of the color complexity is increased.

[0077] When the video to be processed is a real-time surveillance video, the weight β of the temporal information index is increased, and the weight γ of the color complexity is decreased.

[0078] According to an alternative embodiment, based on the video compression effects under different compression parameters, dynamic evaluation is performed to determine the range of the optimal compression parameters.

[0079] By specifically analyzing the video compression effects under different compression parameters, i.e., CRF values, evaluation is performed to determine the range of the optimal compression parameters, specifically including: Performing multi-dimensional quality evaluation on the video compression effect of the video to be processed to determine whether the video quality index meets the standard.

[0080] When it is determined that the video quality index meets the standard, based on the dynamic parameter optimization mechanism, iterative testing is performed until the requirement of the target file size is met, and at the same time, it is determined that the video quality index meets the standard.

[0081] According to an alternative embodiment, the dynamic parameter optimization mechanism includes: when the output file does not reach the target file size, starting the parameter adjustment algorithm.

[0082] For the parameter adjustment algorithm, a progressive adjustment strategy is specifically adopted, with a step size of 0.1, and at least two of the weights of the temporal information index, the spatial information index, and the color complexity weight are adjusted in sequence, so that the output file meets the target file size to determine the range of each weight.

[0083] In the case where the application scenario is a surveillance video, the following weight adjustment strategy is executed: preferentially increase the weight β of the temporal information index, when the weight β is gradually increased in steps of 0.1 and reaches the specified value, then decrease the weight α of the spatial information index, and finally decrease the weight γ of the color complexity; the amount of increase in the weight β of the temporal information index, i.e., the second adjustment amount, is 0.4 to 0.8.

[0084] When the application scenario is a sports event, the following weight adjustment strategy is executed: First, increase the weight β of the time information index. When the weight β is gradually adjusted in steps of 0.1 and reaches the second specified value, then decrease the weight γ. The amount of increase in the weight β of the time information index to be adjusted, that is, the second adjustment amount, is 0.3 to 0.7. When the application scenario is an animated video, the following weight adjustment strategy is executed: First, decrease the weight α. When the weight α is gradually adjusted downward in steps of 0.1 and reaches the third specified value, then increase the weight γ. The amount of decrease in the weight α of the spatial information index to be adjusted, that is, the first adjustment amount, is 0.1 to 0.5.

[0085] It should be noted that since Figure 2 the video compression processing method executed by the video compression processing device of Figure 1 is substantially the same as the video compression processing method in the example of

[0086] Compared with the prior art, the present invention extracts video features from the video to be processed, configures the target file size corresponding to the video to be processed, determines each target file size according to the specific scenario and user requirements, dynamically generates compression parameters through weighted calculation based on the extracted video features and the determined target file sizes, dynamically evaluates according to the video compression effects under different compression parameters to determine the range of the optimal compression parameters, selects the optimal compression parameter from the determined range of the optimal compression parameters, and performs video compression on the video to be processed according to the selected optimal compression parameter, can determine the optimal compression parameter corresponding to the video to be processed, and can achieve more efficient and accurate video compression.

[0087] Figure 3 is a schematic structural diagram of an embodiment of an electronic device according to the present invention.

[0088] As Figure 3 shown, the electronic device is presented in the form of a general-purpose computing device. The processor can be one or multiple and work cooperatively. The present invention does not exclude distributed processing, that is, the processors can be dispersed in different physical devices. The electronic device of the present invention is not limited to a single entity, but can also be the sum of multiple physical devices.

[0089] The memory stores computer-executable programs, usually machine-readable codes. The computer-readable programs can be executed by the processor so that the electronic device can execute the method of the present invention or at least part of the steps in the method.

[0090] The memory includes volatile memory, such as a random access storage unit (RAM) and / or a cache storage unit, and can also be non-volatile memory, such as a read-only storage unit (ROM).

[0091] Optionally, in this embodiment, the electronic device further includes an I / O interface for data exchange between the electronic device and external devices. The I / O interface may represent one or more of several bus structures, including a memory unit bus or a memory unit controller, a peripheral bus, a graphics acceleration port, a processing unit, or a local bus using any of the various bus structures.

[0092] It should be understood that Figure 3 the electronic device shown is merely an example of the present invention, and the electronic device of the present invention may further include elements or components not shown in the above examples. For example, some electronic devices further include a display unit such as a display screen, and some electronic devices further include human-computer interaction elements such as buttons and keyboards. As long as the electronic device can execute the computer-readable program in the memory to implement at least some steps of the method of the present invention, it can be considered as the electronic device covered by the present invention.

[0093] From the description of the above embodiments, those skilled in the art can easily understand that the example embodiments described herein can be implemented by software or by a combination of software and necessary hardware. Therefore, as Figure 4 shown, the technical solution according to the embodiment of the present invention can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which may be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, including several commands to enable a computing device (which may be a personal computer, a server, or a network device, etc.) to execute the above method according to the embodiment of the present invention.

[0094] The software product may employ any combination of one or more readable media. The readable media may be a readable signal medium or a readable storage medium. The readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive list) of the readable storage medium include: an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.

[0095] The computer-readable storage medium may include a data signal propagated in a baseband or as part of a carrier wave, in which the readable program code is carried. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the foregoing. The readable storage medium may also be any readable medium other than the readable storage medium, which may send, propagate, or transmit a program for use by or in conjunction with a command execution system, apparatus, or device. The program code contained on the readable storage medium may be transmitted using any appropriate medium, including but not limited to wireless, wired, optical fiber cable, RF, etc., or any suitable combination of the foregoing.

[0096] The program code for performing the operations of the present invention may be written in any combination of one or more programming languages, including object-oriented programming languages such as Java, C++, etc., and also including conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computing device, partially on the user's device, executed as a stand-alone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device may be connected to the user's computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or, alternatively, may be connected to an external computing device (e.g., by using an Internet service provider to connect through the Internet).

[0097] The above computer-readable medium carries one or more programs, which when executed by a device, cause the computer-readable medium to implement the data interaction method of the present disclosure.

[0098] Those skilled in the art can understand that the above-mentioned modules may be distributed in the device according to the description of the embodiments, or may be correspondingly changed and distributed in one or more devices that are different from the present embodiment only. The modules of the above embodiments may be combined into one module, or may be further split into multiple sub-modules.

[0099] Through the description of the above embodiments, those skilled in the art can easily understand that the example embodiments described herein can be implemented by software, or can be implemented by a combination of software and necessary hardware. Therefore, the technical solution according to the embodiments of the present invention can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which may be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, including several commands to cause a computing device (which may be a personal computer, a server, a mobile terminal, or a network device, etc.) to execute the method according to the embodiments of the present invention.

[0100] It should be noted that the above detailed description is exemplary and is intended to provide further illustration of the present application. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the technical field to which this application belongs.

[0101] In the above detailed description, reference has been made to the accompanying drawings, which form a part hereof. In the drawings, like symbols typically identify like components, unless the context indicates otherwise. The illustrated embodiments described in the detailed description, the drawings, and the claims are not meant to be limiting. Other embodiments may be used and other changes may be made without departing from the spirit or scope of the subject matter presented herein.

[0102] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention may have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A video compression processing method, characterized in that: The video compression processing method comprises: Extract video features from the video to be processed, including spatial complexity, temporal complexity, and color complexity; Configuring target file sizes corresponding to the video to be processed, and determining the target file sizes according to specific scenarios and user needs; Based on the extracted video features and the determined target file size, compression parameters are dynamically generated through weighted calculation; According to the video compression effect under different compression parameters, dynamic evaluation is performed to determine the range of the best compression parameters, the best compression parameters are selected from the determined range of the best compression parameters, and the video to be processed is compressed according to the selected best compression parameters.

2. The video compression processing method according to claim 1, characterized in that: The method of dynamically generating compression parameters through weighted calculation based on the extracted video features and the determined target file size includes: Use the following expression to calculate the compression parameters of the video to be processed: CRF=BaseCRF 初始 −α·SI_norm−β·TI_norm−γ·Color_norm; Among them, CRF represents the compression parameter obtained by weighted calculation based on the basic parameters and the quantitative indicators of the extracted video features; BaseCRF 初始 It represents the basic parameters, which are characterized by a constant rate factor and can be preset according to the scene; SI_norm represents the quantitative index of the extracted spatial complexity, namely the spatial information index; α represents the weight corresponding to the spatial information index; TI_norm represents the quantitative index of the extracted temporal complexity, namely the temporal information index; β represents the weight corresponding to the temporal information index; Color_norm represents the extracted color complexity; γ represents the weight corresponding to the color complexity.

3. The video compression processing method according to claim 2, characterized in that: Further including: According to the weight adjustment principle, real-time weight adjustment is performed, specifically including: first determining the application scenario, and adjusting the weight of the quantitative index corresponding to one or more video features according to the determined application scenario; wherein, When the application scenario is landscape video display and the video to be processed is a landscape video, the weight α of the spatial information index is increased, and the weight β of the temporal information index is reduced; When the application scenario is a sports event display and the video to be processed is a sports game video, the weight β of the time information index is increased and the weight γ of the color complexity is reduced; When the application scenario is animation video display or cartoon video display, and the video to be processed is an animation video or a cartoon video, the weight γ of the color complexity is increased; When the video to be processed is a real-time surveillance video, the weight β of the time information index is increased, and the weight γ of the color complexity is reduced.

4. The video compression processing method according to claim 1, characterized in that: The method of dynamically evaluating the video compression effects under different compression parameters to determine the range of the optimal compression parameters includes: By specifically analyzing the video compression effects under different compression parameters, namely CRF values, an evaluation is performed to determine the range of the optimal compression parameters, including: Perform multi-dimensional quality evaluation on the video compression effect after the video is compressed to determine whether the video quality index meets the standard; When determining whether the video quality indicators meet the standards, based on the dynamic parameter optimization mechanism, iterative testing is performed until the target file size requirements are met, and at the same time, it is determined that the video quality indicators meet the standards.

5. The video compression processing method according to claim 4, characterized in that: The dynamic parameter optimization mechanism includes: when the output file does not reach the target file size, starting the parameter adjustment algorithm; For the parameter adjustment algorithm, a progressive adjustment strategy is specifically adopted, with 0.1 as the step unit, and at least two of the weights of the time information index, the spatial information index, and the color complexity are adjusted in turn, so that the output file meets the target file size, in order to determine the range of each weight.

6. The video compression processing method according to claim 5, characterized in that: When the application scenario is surveillance video, the following weight adjustment strategy is implemented: Prioritize increasing the weight β of the time information index, and when the weight β gradually increases by 0.1 and increases to a specified value, reduce the weight α of the space information index, and finally reduce the weight γ of the color complexity; the amount by which the weight β of the time information index needs to be increased, i.e., the second adjustment amount, is 0.4 to 0.8; When the application scenario is a sports event, the following weight adjustment strategy is implemented: Prioritize increasing the weight β of the time information index, and when the weight β is gradually adjusted by 0.1 and adjusted to the second specified value, reduce the weight γ; the amount by which the weight β of the time information index needs to be adjusted, i.e., the second adjustment amount, is 0.3 to 0.7; When the application scenario is an animated video, the following weight adjustment strategy is implemented: The weight α is preferentially reduced, and when the weight α is gradually adjusted downward by 0.1 and adjusted to the third specified value, the weight γ is increased; the amount by which the weight α of the spatial information index needs to be adjusted to be reduced, that is, the first adjustment amount, is 0.1 to 0.

5.

7. A video compression processing device, characterized in that: The video compression processing method according to any one of claims 1 to 6 is executed, and the video compression processing device comprises: An extraction module is used to extract video features from the video to be processed, where the video features include spatial complexity, temporal complexity, and color complexity; A configuration determination module is used to configure a target file size corresponding to the video to be processed, and determine each target file size according to a specific scenario and user needs; A calculation module dynamically generates compression parameters through weighted calculation based on the extracted video features and the determined target file size; The compression module dynamically evaluates the video compression effect under different compression parameters to determine the range of the optimal compression parameters, selects the optimal compression parameters from the determined range of the optimal compression parameters, and compresses the video to be processed according to the selected optimal compression parameters.

8. The video compression processing device according to claim 7, characterized in that: include: Use the following expression to calculate the compression parameters of the video to be processed: CRF=BaseCRF 初始 −α·SI_norm−β·TI_norm−γ·Color_norm; Among them, CRF represents the compression parameter obtained by weighted calculation based on the basic parameters and the quantitative indicators of the extracted video features; BaseCRF 初始 It represents the basic parameters, which are characterized by a constant rate factor and can be preset according to the scene; SI_norm represents the quantitative index of the extracted spatial complexity, namely the spatial information index; α represents the weight corresponding to the spatial information index; TI_norm represents the quantitative index of the extracted temporal complexity, namely the temporal information index; β represents the weight corresponding to the temporal information index; Color_norm represents the extracted color complexity; γ represents the weight corresponding to the color complexity.

9. The video compression processing device according to claim 8, characterized in that: Further including: According to the weight adjustment principle, real-time weight adjustment is performed, specifically including: first determining the application scenario, and adjusting the weight of the quantitative index corresponding to one or more video features according to the determined application scenario; wherein, When the application scenario is landscape video display and the video to be processed is a landscape video, the weight α of the spatial information index is increased, and the weight β of the temporal information index is reduced; When the application scenario is a sports event display and the video to be processed is a sports game video, the weight β of the time information index is increased and the weight γ of the color complexity is reduced; When the application scenario is animation video display or cartoon video display, and the video to be processed is an animation video or a cartoon video, the weight γ of the color complexity is increased; When the video to be processed is a real-time surveillance video, the weight β of the time information index is increased, and the weight γ of the color complexity is reduced.

10. The video compression processing device according to claim 8, characterized in that: By specifically analyzing the video compression effects under different compression parameters, namely CRF values, an evaluation is performed to determine the range of the optimal compression parameters, including: Perform multi-dimensional quality evaluation on the video compression effect after the video is compressed to determine whether the video quality index meets the standard; When determining whether the video quality indicators meet the standards, based on the dynamic parameter optimization mechanism, iterative testing is performed until the target file size requirements are met, and at the same time, it is determined that the video quality indicators meet the standards.

Citation Information

Cited By

  • Video coding parameter determination method and device based on hardware performance portrait

    CN121442101A

  • Method and apparatus for determining video encoding parameters based on hardware performance profile

    CN121442101B