Video coding method for jointly considering human vision and machine task
By dividing the videos with areas of interest and areas of non-interest, a target detection accuracy and visual quality model is constructed, and the bit rate allocation is optimized, the problem of insufficient human vision and machine vision in the existing technology is solved, and the target detection accuracy and visual quality improvement of efficient video encoding is achieved.
Patent Information
- Application Number
- CN202510348015.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-24
- Publication Date
- 2025-07-11
AI Technical Summary
The existing video compression technology fails to take into account the needs of human vision and machine vision at the same time. The traditional method focuses on human vision optimization and ignores the special needs of machine vision, resulting in poor compression efficiency and visual effects in object detection tasks.
By CTU-level division of the target videos between the areas of interest and non-interested areas, a model of object detection accuracy and code rate is constructed, combined with the models of visual quality indicators and code rate, the trade-offs of detection accuracy and visual quality are designed, and the bit rate allocation of ROI and NROI regions is optimized to achieve joint optimization of video encoding.
Improve object detection accuracy and ensure visual quality under limited bit rate, comprehensively optimize the video compression process, and improve the overall performance of compressed video.
Smart Images

Figure CN120302044A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of video compression, and relates to a video encoding method that jointly considers human vision and machine tasks. Background Art
[0002] In recent years, with the rapid development of the fields of artificial intelligence, autonomous driving, and video analysis, the amount of video data has increased explosively, and the demand for the generation and transmission of video content has increased sharply, thereby posing higher requirements for bandwidth and storage. In order to improve the video transmission efficiency and reduce the storage consumption, video compression technology has developed rapidly. Existing video coding standards, such as H.265 / HEVC and H.266 / VVC, all take human vision (HVS) as the core optimization goal, aiming to minimize the amount of data while maintaining good visual quality. Based on the perceptual characteristics of the human eye, these compression algorithms preferentially retain information that is important for human visual perception. For example, during the encoding process, details sensitive to the human eye (such as color, edges, textures, etc.) are retained, while regions insensitive to the human eye are compressed at a higher rate. Although these compression algorithms are very effective in terms of visual effects, they do not consider the needs of machine vision. The perception methods of machine vision systems and the human eye are significantly different. In object detection tasks, machine vision relies more on the detailed information of images, especially information such as the edges, shapes, and positions of objects. Currently, relevant video compression methods have been studied for object detection. For example, region-aware compression that adjusts the compression ratio based on the importance of objects in video frames; compression methods that use deep learning to assist in adjusting the compression strategy; and adaptive encoding techniques that adjust parameters according to video content, etc.
[0003] Existing video compression technologies have not yet taken into account the needs of both human vision and machine vision. Traditional video coding methods focus on optimizing human vision and ignore the special needs of machine vision, while current compression methods optimized for object detection generally ignore the special needs of human vision. In the future, further research is still needed on how to achieve efficient compression while taking into account the high-quality visual effects of videos and the high accuracy of object detection. Summary of the Invention
[0004] In view of this, the purpose of the present invention is to provide a video encoding method that jointly considers human vision and machine tasks.
[0005] To achieve the above purpose, the present invention provides the following technical solutions:
[0006] A video encoding method that jointly considers human vision and machine tasks, including the following steps:
[0007] S1: Detect the target video and divide the CTU-level regions of the region of interest ROI and the non-region of interest NROI;
[0008] S2: Study the relationship between the bit rate of the NROI region and the detection accuracy;
[0009] S3: Build a model between the object detection accuracy and the bit rate in the ROI region;
[0010] S4: Build a model between the visual quality metric of the video sequence and the bit rate;
[0011] S5: Design a trade-off between the detection accuracy and the visual quality by building models between the object detection accuracy, the human visual quality and the bit rate.
[0012] Further, the CTU-level division of the region of interest ROI and the non-interested region in the target video described in step S1 specifically includes the following steps:
[0013] S11: Input the test sequence into the object detection algorithm YOLOv9, and output the labels of the detected objects, including the object detection boxes and their coordinates; determine the position, size and category of each object detection box through the coordinates;
[0014] S12: Determine the CTUs that overlap, intersect or are far from the detection box according to the position and size of the detection box; define the overlapping and intersecting CTUs as the region of interest ROI, and the CTUs far from the detection box as the non-interested region NROI.
[0015] Further, the study of the influence of the bit rate of the NROI region on the detection accuracy described in step S2 specifically includes the following steps:
[0016] Select the test sequences in HEVC for full I-frame encoding. For the CTUs in the ROI region, quantization is performed with QP ∈ {17, 22, 27, 32, 37, 42, 47, 51}; on the basis of the selected QP in the ROI region, the NROI region is quantized with QP = 42, 47, 51 respectively to study the influence of the NROI region on the detection accuracy under different QP quantizations.
[0017] Further, for the ROI region of interest, building a model between the object detection accuracy and the bit rate described in step S3 specifically includes the following steps:
[0018] S31: Calculate that under the quantization of QP ∈ {17, 22, 27, 32, 37, 42, 47, 51} selected in the ROI region, the NROI region is encoded with QP = 42, 47, 51 respectively, and calculate the bit rate of the ROI region under each QP quantization;
[0019] S32: Input the encoded test sequence into YOLOv9 for object detection, and record the detection accuracy MAP under different QP quantizations;
[0020] S32: Model the bitrate R of the ROI region of the test sequence and the detection accuracy MAP, expressed as:
[0021]
[0022] In the formula, K is the value of the detection accuracy MAP, R is the bitrate of the ROI region, and P1, P2, P3, P4 are model parameters.
[0023] Furthermore, the step of constructing the model between the visual quality index and the bitrate described in step S4 specifically includes the following steps:
[0024] S41: Select QP ∈ {17, 22, 27, 32, 37, 42, 47, 51} for full I-frame intra-coding of the entire test sequence:
[0025] S42: Calculate the bitrate R of the test sequence under different QP values, and calculate the mean square error MSE between the Y component of each pixel point of the source video and the encoded video, and obtain the model of the video quality evaluation index MSE and the test sequence bitrate R as:
[0026] D = α·(ln(R)) β
[0027] In the formula, D represents the value of the video quality evaluation index MSE, R is the test sequence bitrate, and α, β are model parameters.
[0028] Furthermore, the step of designing the trade-off between the detection accuracy and the human visual quality by constructing the model between the target detection accuracy, the human visual quality and the bitrate described in step S5 specifically includes:
[0029] S51: Design the loss function J, optimize the loss function J to obtain the optimal bitrate allocation a for the ROI and NROI regions. By modeling a and substituting it into the loss function J for calculation, obtain the bitrates allocated to the ROI and NROI regions; the loss function J is:
[0030]
[0031] In the formula, D ROI , D NROI are the MSE values of the ROI region and the NROI region respectively, and K is the detection accuracy MAP;
[0032] S52: By calculating the minimum loss function J, obtain the value of a, and allocate bitrates to the ROI and NROI regions according to the value of a:
[0033] R ROI = a·R
[0034] R NROI = (1 - a)·R
[0035] Wherein, a is the ratio assigned to the ROI region, and R ROI is the bitrate assigned to the ROI region, and R NROI is the bitrate assigned to the NROI region;
[0036] S53: Allocate bitrates for each CTU within each region according to the bitrates assigned to the ROI and NROI regions:
[0037]
[0038] Wherein R i is the target number of bits assigned to the i-th CTU, I is the index set of CTUs, and ω i is the weight of the i-th CTU. The weight ω i is defined as follows:
[0039]
[0040] Wherein, n i is the number of pixels of the i-th CTU.
[0041] The beneficial effects of the present invention are as follows:
[0042] The present invention proposes a video coding method that jointly considers human vision and machine tasks. By establishing a mathematical model among the visual metric MSE, detection accuracy MAP, and bitrate, video coding can be performed by allocating different bitrates to the ROI (region of interest) and NROI (non-region of interest) regions under limited bitrate conditions. While improving the detection accuracy, the visual quality is ensured. By innovatively combining object detection, visual quality, and coding efficiency, the video compression process is comprehensively optimized, and the overall performance of the compressed video is improved.
[0043] Other advantages, objectives, and features of the present invention will be described to some extent in the subsequent description, and to some extent, will be obvious to those skilled in the art based on the study of the following text, or can be taught from the practice of the present invention. The objectives and other advantages of the present invention can be realized and obtained through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] In order to make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be described in detail preferably with reference to the accompanying drawings, where:
[0045] Figure 1 is the flowchart of the video coding method that jointly considers human vision and machine tasks according to the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0046] The following specific examples illustrate the embodiments of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments. Various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the diagrams provided in the following embodiments only illustrate the basic concept of the present invention in a schematic manner. Without conflict, the following embodiments and the features in the embodiments can be combined with each other.
[0047] It should be noted that the diagrams provided in the following embodiments only illustrate the basic concept of the present invention in a schematic manner. Therefore, only the components related to the present invention are shown in the diagrams, rather than being drawn according to the number, shape, and size of the components in actual implementation. The type, quantity, and ratio of each component in its actual implementation can be arbitrarily changed, and the component layout type may also be more complex.
[0048] In the following description, a large number of details are explored to provide a more thorough explanation of the embodiments of the present invention. However, it is obvious to those skilled in the art that the embodiments of the present invention can be implemented without these specific details. In other embodiments, well-known structures and devices are shown in the form of block diagrams rather than in detail to avoid making the embodiments of the present invention difficult to understand.
[0049] Combined with Figure 1 , the specific embodiments of the present invention are described in detail. The present invention provides a video compression method for improving the accuracy of target detection and preserving the visual quality of videos, including:
[0050] Step 1: Perform CTU-level division of the region of interest (ROI) and non-interested regions in the target video, specifically including the following steps:
[0051] Step 1.1: First, input the test video sequence into the object detection algorithm YOLOv9. The output of YOLOv9 will provide the labels of each detected object, which contain the detection boxes of the objects and the coordinate information of the boxes. Through these coordinates, the position, size, and category to which each object belongs can be accurately determined. This process ensures that the key information in the video can be extracted for subsequent processing.
[0052] Step 1.2: Based on the position and size of the detection box in each frame of the image, it is possible to further identify the coding tree units (CTUs) that overlap, intersect, or are far from the detection box. Specifically, the overlapping or intersecting CTUs will be defined as the region of interest (ROI), while the CTUs far from the detection box will be defined as the non-region of interest (NROI). This CTU-level division helps to more precisely locate and optimize the key regions in the video coding process, thereby improving the coding efficiency and image quality.
[0053] Step 2: Study the impact of the bitrate in the NROI region on the detection accuracy. The specific steps include:
[0054] Step 2.1: Select the test video sequences in HEVC (High Efficiency Video Coding) and perform intra-frame full I-frame coding, that is, all frames use I-frames (intra-predicted frames). On this basis, for the CTUs (coding tree units) in the region of interest (ROI), multiple quantization parameters (QPs) are selected, specifically QP ∈ {17, 22, 27, 32, 37, 42, 47, 51}, and quantization processing is performed. These quantization parameters will be used to analyze the impact of ROI region detection on the detection accuracy under different quantization conditions.
[0055] Step 2.2: After the quantization parameters in the ROI region are selected, for the non-region of interest (NROI), different quantization values are used for coding, that is, quantization processing is performed at QP = 42, 47, 51. By studying the NROI region under different quantization conditions, analyze the impact of its bitrate on the object detection accuracy, so as to evaluate the optimization effect of different quantization strategies on video coding and detection performance.
[0056] Step 3: Build a model between the target detection accuracy and the bitrate in the ROI region, which specifically includes the following steps:
[0057] Step 3.1: Under the selected quantization parameter (QP) in the ROI region, first perform the coding process for the NROI region, where the NROI region is coded using QP = 42, 47, 51 respectively. Next, calculate the bitrate of the corresponding ROI region for each quantization group under the selected QP quantization. This process can obtain the coding efficiency of the ROI region under different quantization settings.
[0058] Step 3.2: Input the coded test video sequence into the YOLOv9 object detection algorithm for target detection, and record the detection accuracy (MAP) under different quantization parameters. This process will help to evaluate the impact of different quantization settings on the object detection results;
[0059] Step 3.3: Model the bitrate R of the ROI region and the detection accuracy MAP of the test sequence. The model can be expressed in the following form:
[0060]
[0061] Wherein, K is the value of the detection accuracy MAP, R is the bitrate of the ROI region, and P1, P2, P3, and P4 are model parameters. Through this modeling process, the relationship between the bitrate of the ROI region and the detection accuracy can be quantified, and the strategy can be further optimized to improve the object detection accuracy.
[0062] Step 4: For the entire video sequence, construct a model between the visual quality metric and the bitrate, which specifically includes the following steps:
[0063] Step 4.1: Encode the entire test video sequence, and select the quantization parameter (QP), specifically QP ∈ {17, 22, 27, 32, 37, 42, 47, 51}, and adopt the intra-frame coding mode with all I-frames. That is, during the encoding process, each frame is encoded using an I-frame (intra-predicted frame), which helps to simplify the encoding structure while ensuring the video quality.
[0064] Step 4.2: Under different QP values, first calculate the bitrate R of the test sequence. Then, for the Y component (luminance component) of each pixel, calculate the mean square error (MSE) between it and the source video sequence. This process evaluates the video quality by comparing the differences between the encoded video and the original video. Finally, the relationship between the video quality evaluation metric MSE and the bitrate R of the test sequence can be represented by the following model
[0065] D = α·(ln(R)) β
[0066] Wherein, D represents the video quality evaluation metric (mean square error), R is the bitrate of the test sequence, and α and β are model parameters. This model helps to quantify the relationship between video quality and encoding bitrate, providing a theoretical basis for optimizing the encoding strategy.
[0067] Step 5: By constructing models between the object detection accuracy, human visual quality, and the bitrate, design a trade-off between the detection accuracy and human visual quality, and the specific steps include:
[0068] Step 5.1: Design a loss function J, and through the optimization of this loss function, solve the optimal bitrate allocation ratio a between the ROI and NROI regions. By establishing a model related to the loss function and substituting a into the loss function for calculation, the specific bitrates allocated to the ROI and NROI regions can finally be obtained. The form of the loss function J is:
[0069]
[0070] Wherein, D ROI , D NROIThey are the mean squared error (MSE) of the ROI region and the NROI region respectively, and K is the detection accuracy MAP.
[0071] Step 5.2: By calculating the minimum loss function J, obtain the value of the optimal allocation ratio a. According to this optimal a, we allocate bitrates to the ROI and NROI regions respectively. The specific bitrate allocation formula is as follows:
[0072] R ROI = a·R
[0073] R NROI = (1 - a)·R
[0074] In the formula, a is the ratio allocated to the ROI region, R ROI is the bitrate allocated to the ROI region, and R NROI is the bitrate allocated to the NROI region;
[0075] Step 5.3: After determining the bitrate allocation for the ROI and NROI regions, further allocate bitrates to each CTU (Coding Tree Unit) within each region. Specifically, for each CTU, the target number of bits R i can be calculated by the following formula:
[0076]
[0077] In the formula, R i is the target number of bits allocated to the i-th ctu, I is the index set of CTUs, and ω i is the weight of the i-th CTU. The weight ωi is defined as follows:
[0078]
[0079] In the formula, n i is the number of pixels of the i-th CTU. Through this process, a reasonable number of bits can be allocated to each CTU, thereby further optimizing the coding quality and efficiency.
[0080] In the above embodiments, the reference to "this embodiment" in the specification means that the specific features, structures, or characteristics described in connection with the embodiments are included in at least some embodiments, but not necessarily all embodiments. Multiple occurrences of "this embodiment" do not necessarily all refer to the same embodiment.
[0081] In the above embodiments, although the present invention has been described in connection with specific embodiments of the present invention, many alternatives, modifications, and variations of these embodiments will be apparent to those of ordinary skill in the art based on the foregoing description. For example, other storage structures (e.g., dynamic RAM (DRAM)) may be used with the embodiments discussed. Embodiments of the present invention are intended to cover all such alternatives, modifications, and variations that fall within the broad scope of the appended claims.
[0082] This embodiment also provides a computer-readable storage medium having a computer program stored thereon, and when the program is executed by a processor, it implements any one of the methods in this embodiment.
[0083] This embodiment also provides an electronic terminal, including: a processor and a memory;
[0084] The memory is used to store a computer program, and the processor is used to execute the computer program stored in the memory so that the terminal executes any one of the methods in this embodiment.
[0085] For the computer-readable storage medium in this embodiment, those of ordinary skill in the art can understand that all or part of the steps of implementing the above method embodiments can be completed by hardware related to the computer program. The foregoing computer program can be stored in a computer-readable storage medium. When the program is executed, it executes the steps including the above method embodiments; and the foregoing storage medium includes: various media such as ROM, RAM, magnetic disk, or optical disk that can store program codes.
[0086] The electronic terminal provided in this embodiment includes a processor, a memory, a transceiver, and a communication interface. The memory and the communication interface are connected to the processor and the transceiver and complete communication therebetween. The memory is used to store a computer program, the communication interface is used for communication, and the processor and the transceiver are used to run the computer program so that the electronic terminal executes each step of the above method.
[0087] In this embodiment, the memory may include a random access memory (Random Access Memory, abbreviated as RAM), and may also include a non-volatile memory, such as at least one disk memory.
[0088] The above-mentioned processor may be a general-purpose processor, including a central processing unit (CPU for short), a network processor (NP for short), etc.; it may also be a digital signal processor (DSP for short), an application specific integrated circuit (ASIC for short), a field-programmable gate array (FPGA for short), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.
[0089] The present invention can be used in numerous general-purpose or special-purpose computing system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multi-processor systems, microprocessor-based systems, set-top boxes, programmable consumer electronic devices, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and so on.
[0090] The present invention can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The present invention can also be practiced in a distributed computing environment where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.
[0091] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the present technical solution, and they should all be covered within the scope of the claims of the present invention.
Claims
1. A video encoding method that jointly considers human vision and machine tasks, characterized in that: It includes the following steps: S1: Detect the target video and divide the CTU-level regions of the region of interest (ROI) and the non-region of interest (NROI). S2: Study the relationship between the bitrate of the NROI region and the detection accuracy. S3: Construct a model between the object detection accuracy and the bitrate of the ROI region. S4: Construct a model between the visual quality metric of the video sequence and the bitrate. S5: Design a trade-off between the detection accuracy and the visual quality by constructing models between the object detection accuracy, the human visual quality, and the bitrate.
2. The video encoding method that jointly considers human vision and machine tasks according to claim 1, wherein: The CTU-level division of the region of interest (ROI) and the non-region of interest in the target video detection described in step S1 specifically includes the following steps: S11: Input the test sequence into the object detection algorithm YOLOv9, and output the labels of the detected objects, including the object detection boxes and their coordinates; determine the position, size, and category of each object detection box based on the coordinates. S12: Determine the CTUs that overlap, intersect, or are far from the detection boxes according to the positions and sizes of the detection boxes; define the overlapping and intersecting CTUs as the region of interest (ROI), and the CTUs far from the detection boxes as the non-region of interest (NROI).
3. The video coding method that jointly considers human vision and machine tasks according to claim 1, wherein: The study of the influence of the bitrate of the NROI region on the detection accuracy described in step S2 specifically includes the following steps: Select test sequences in HEVC for all-I-frame encoding. For the CTUs in the ROI region, quantization is performed with QP ∈ {17, 22, 27, 32, 37, 42, 47, 51}; based on the selected QP in the ROI region, the NROI region is quantized with QP = 42, 47, 51 respectively to study the influence of the NROI region on the detection accuracy under different QP quantizations.
4. The video coding method that jointly considers human vision and machine tasks according to claim 1, wherein: The construction of a model between the object detection accuracy and the bitrate for the ROI region of interest described in step S3 specifically includes the following steps: S31: Calculate the bitrates of the ROI region under quantization with QP ∈ {17, 22, 27, 32, 37, 42, 47, 51} in the ROI region, and encode the NROI region with QP = 42, 47, 51 respectively, and calculate the bitrate of the ROI region under each QP quantization. S32: Input the encoded test sequence into YOLOv9 for object detection, and record the mean average precision (MAP) of the detection accuracy under different QP quantizations. S32: Model the bitrate R of the ROI region of the test sequence and the detection accuracy MAP, expressed as: In the formula, K is the value of the detection accuracy MAP, R is the bitrate of the ROI region, and P1, P2, P3, P4 are model parameters.
5. The video coding method considering both human vision and machine tasks according to claim 1, characterized in that: The construction of a model between the visual quality metric and the bitrate described in step S4 specifically includes the following steps: S41: Perform all-I-frame intra-encoding on the entire test sequence with QP ∈ {17, 22, 27, 32, 37, 42, 47, 51}. S42: Calculate the bitrate R of the test sequence under different QP values, and calculate the mean squared error (MSE) between the Y components of each pixel of the source video and the encoded video. The model of the video quality evaluation metric MSE and the bitrate R of the test sequence is: D = α·(ln(R)) β In the formula, D represents the value of the video quality evaluation metric MSE, R is the bitrate of the test sequence, and α, β are model parameters.
6. The video encoding method that jointly considers human vision and machine tasks according to claim 1, wherein: As described in step S5, by constructing a model between the target detection accuracy, human visual quality and bitrate, a trade-off between the detection accuracy and human visual quality is designed. The specific steps include: S51: Design a loss function J, optimize the loss function J to obtain the optimal bitrate allocation a for the ROI and NROI regions; by modeling a and substituting it into the loss function J for calculation, obtain the bitrates allocated to the ROI and NROI regions; S52: Obtain the value of a by calculating the minimum loss function J, and allocate bitrates to the ROI and NROI regions respectively according to the value of a; R ROI = a·R R NROI = (1 - a)·R Where a is the ratio assigned to the ROI region, and R ROI is the bit rate assigned to the ROI region, and R NROI is the bit rate assigned to the NROI region; S53: Allocate bitrates to each CTU in each region according to the bitrates allocated to the ROI and NROI regions: where R i is the target number of bits allocated to the \(i\)-th CTU, \(I\) is the set of indices of CTUs, and \(\omega\) i is the weight of the \(i\)-th CTU. The weight \(\omega\) i is defined as follows: where ni is the number of pixels of the i-th CTU.
7. The video encoding method that jointly considers human vision and machine tasks according to claim 1, characterized in that: In step S51, the loss function J is: where D ROI , D NROI are the MSE values of the ROI region and the NROI region respectively, and K is the detection accuracy MAP.